Intelligent contract management method and device for multi-modal large model, and medium
Through the multimodal large model, the multimodal feature extraction and differential positioning of contract files is solved, and the problem of insufficient information analysis in multi-version contract file management is realized, fine-grained version traceability and risk assessment are achieved, and contract management efficiency and risk prevention and control capabilities are improved.
Patent Information
- Application Number
- CN202510635443.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-05-16
Smart Images

Figure CN120523784A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of contract management technology, and in particular to a multimodal large-scale intelligent contract management method, device, and medium. Background Art
[0002] In the digital business environment, with the continuous expansion and diversification of corporate business, the types of contracts are becoming increasingly complex, covering multiple areas such as sales contracts, procurement contracts, service contracts, and lease contracts. At the same time, the content of the contract is no longer limited to traditional text descriptions, but also includes a large amount of non-text information such as tables, charts, and images. In the field of contract management, in order to improve the efficiency and accuracy of contract management, more and more companies have begun to introduce information technology. With the acceleration of corporate digital transformation, the need to compare multiple versions of contract documents (such as different revisions, industry standard versions, partner customized versions, etc.) in contract management scenarios is becoming increasingly complex. It is necessary to frequently handle the comparison of multiple revisions of the same contract (such as the adaptation of terms in different jurisdictions of cross-border procurement contracts), analysis of differences in contract terms among multiple suppliers, and version comparison needs between standard contracts and customized contracts.
[0003] Traditional manual comparison methods require legal personnel to verbatim check text and non-text information, such as tables of amounts and contract performance flowcharts. This not only takes weeks but is also prone to human errors. In particular, there is a significant lag in identifying subtle numerical differences between versions (such as changes in the decimal point of a tax rate) or semantically equivalent clauses (such as "Party A has the right to terminate" and "Party B's obligations are released"). While automated tools based on text matching can improve efficiency, they are unable to parse the nested table structures in PDF / Word contracts, resulting in missed comparisons of critical non-text information such as unit price matrices in purchase orders and KPI indicator charts in service contracts. Furthermore, relying on keyword string matching cannot identify the semantic consistency of clause wording across version iterations, such as the legal equivalence of "force majeure" and "accidental events."
[0004] Therefore, in the process of managing multi-version contract documents, due to the lack of parsing of multimodal information in the contract versions and insufficient understanding of the semantics of the terms, there are problems such as a high omission rate of key terms differences and ineffective non-text content comparison, which in turn leads to inefficient cross-version contract management. Summary of the Invention
[0005] One or more embodiments of this specification provide a multimodal large-model intelligent contract management method, device, and medium for solving the following technical problems: in the management process of multi-version contract documents, due to the lack of parsing of multimodal information in the contract version and insufficient understanding of the semantics of the terms, there are problems such as a high omission rate of key terms differences and failure of non-text content comparison, which in turn leads to low efficiency in cross-version contract management.
[0006] One or more embodiments of this specification adopt the following technical solutions:
[0007] One or more embodiments of the present specification provide a smart contract management method for a multimodal big model, the method comprising obtaining multiple pre-stored versions of contract documents to perform multimodal feature extraction on each of the contract documents and determine multimodal contract feature data, wherein the multimodal contract feature data comprises any one or more of a text feature vector, a table structured data set, and a chart quantitative data set; based on a pre-constructed multimodal big model, comparing the multimodal contract feature data corresponding to each of the contract documents and locating difference area information between the multiple contract documents; establishing contract revision evolution information between multiple historical versions based on the difference area information and the contract version information of each of the contract documents, so as to manage the multiple versions of the contract documents through the contract revision evolution information.
[0008] One or more embodiments of this specification provide a multimodal large-scale smart contract management device, including:
[0009] at least one processor; and,
[0010] a memory communicatively connected to the at least one processor; wherein,
[0011] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the above method.
[0012] One or more embodiments of this specification provide a non-volatile computer storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured to execute the above method.
[0013] At least one of the above-mentioned technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects: through the technical solutions of the embodiments of this specification, through the deep integration of multimodal large models, holographic parsing and intelligent correlation analysis of contract elements are realized. Traditional technology is limited by the processing capabilities of a single modality and can only perform shallow character comparison on the contract text. However, the embodiments of this specification can accurately capture the cross-modal correlation relationship implied in the contract terms by fusing text semantics, table structure, and multimodal feature extraction of chart data; at the level of version difference detection, it breaks through the limitations of traditional linear comparison and realizes three-dimensional tracking of contract revisions; traditional technology relies on manual marking or simple hash comparison, and cannot identify the actual changes in terms such as the break of formula dependency chain caused by the adjustment of table row and column structure, and the offset of chart data points. More complex scenarios; The embodiment of this specification can quantitatively evaluate the semantic deviation of text clauses, the topological change rate of table structure and the difference value of chart data distribution through the unified embedding space mapping of the multimodal large model, and accurately locate the substantive revision area; the traditional version control system (such as Git) only records file-level changes, while the embodiment of this specification uses the version alignment technology of the multimodal feature coordinate system to establish fine-grained revision tracks at the clause level, cell level, and data point level. The atomic version traceability capability enables legal personnel to conduct a penetrating analysis of the semantic drift path of a certain liability clause in multiple historical versions, or observe the parameter evolution law of a certain price calculation formula after multiple revisions; in addition, the embodiment of this specification builds an intelligent compliance protection system through cross-modal consistency verification and dynamic risk assessment models. Compared with the lagging risk control of traditional manual review relying on experience judgment, the embodiment of this specification can detect the implicit contradictions between text clauses, table formulas, and chart data in real time, achieving an exponential improvement in risk prevention and control efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are only some of the embodiments described in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without inventive work. In the drawings:
[0015] Figure 1 A flowchart of a multimodal large-scale smart contract management method provided in an embodiment of this specification;
[0016] Figure 2 A schematic diagram of the structure of a multimodal large-scale intelligent contract management device provided in an embodiment of this specification. DETAILED DESCRIPTION
[0017] To help those skilled in the art better understand the technical solutions in this specification, the following will provide a clear and complete description of the technical solutions in the embodiments of this specification, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this specification without creative work should fall within the scope of protection of this specification.
[0018] The embodiments of this specification provide a multimodal large-scale smart contract management method. It should be noted that the execution entity in the embodiments of this specification can be a server or any device with data processing capabilities. Figure 1 A flowchart of a multi-modal large-scale smart contract management method provided in an embodiment of this specification is shown as follows: Figure 1 As shown, it mainly includes the following steps:
[0019] Step S101: Acquire multiple versions of pre-stored contract documents to perform multimodal feature extraction on each contract document to determine multimodal contract feature data.
[0020] The multimodal contract feature data includes any one or more of a text feature vector, a table structured data set, and a chart quantitative data set;
[0021] In one embodiment of the present specification, multiple contract documents are collected, text information, table structure information, and chart content information are extracted from the contract documents, and the extracted information is cleaned and formatted. The steps of extracting text information, table structure information, and chart content information can be specifically implemented by: extracting image text information from the contract documents using optical character recognition (OCR) technology; identifying the table row and column structure and cell content in the contract documents using a table parsing algorithm; extracting chart areas in the contract documents using image segmentation technology, and parsing the text annotations and data relationships in the charts. The extracted information is cleaned and formatted, and redundant content, including headers, footers, and non-contract content, is removed based on preset noise rules. The text information, table structure information, and chart content information are uniformly converted into a structured data format, including text paragraph labels, table row and column coordinates, and chart area coordinates. The processed multiple versions of the contract documents are stored in a preset database. During the storage process, the version timestamp and revision user information should also be recorded. It should be noted that the database here is dedicated to storing contract documents, which can be multiple revisions of the same contract, contract terms from multiple suppliers, or versions of standard contracts and customized contracts.
[0022] A multimodal large model is constructed through the pre-processed contract documents stored in the database, and the model is trained through the contract semantic difference annotation data to enable it to have multimodal information fusion and deep semantic understanding capabilities. In the embodiment of this specification, a multimodal model based on the Transformer architecture can be selected as the basic framework, and the model includes a text encoder, a table encoder and a chart encoder. The pre-processed contract documents are input into the model, and the feature vectors of the text, table and chart modalities are cross-modally aligned and trained through a comparative learning algorithm, and the model is supervised and trained based on the semantic difference annotation data, and the annotation data includes the contract clause type label and the difference position label. Among them, the semantic difference annotation data can be obtained through manual annotation and automatic identification. The key differences in the contract clauses are manually annotated, including the amount, liability clauses and technical parameter differences; the logical relationship of the contract clauses is automatically annotated through the rule engine, including conditional statements, time nodes and the correlation between rights and obligations.
[0023] In one embodiment of the present specification, multiple versions of contract documents requiring comparison and management are obtained from a database used to store contract documents, and multimodal feature extraction is performed on each of the contract documents to determine multimodal contract feature data, wherein the multimodal contract feature data includes any one or more of a text feature vector, a table structured data set, and a chart quantitative data set. It should be noted that the multimodality herein includes text, tables, and charts. Tables are two-dimensional structured data based on rows and columns, with content primarily consisting of text or numerical values. Charts are unstructured or semi-structured data based on graphical elements (such as coordinate axes, graphics, and legends).
[0024] Multimodal feature extraction is performed on each contract document, and determination of multimodal contract feature data can be achieved through the following steps: multimodal feature decoupling is performed on each contract document, and the text clauses, table areas and chart areas corresponding to each contract document are determined; the responsible parties, amount values and clause reference relationships in the text clauses are extracted by a text parsing engine to generate text feature vectors; the row and column structure of the table area is identified by a table parsing engine, cell coordinates and merging relationships are extracted, and a table structured data set is generated, wherein the table structured data set includes cell hash values and formula dependency chains; the coordinate axes, data graphics and legend components in the chart area are segmented by a chart parsing engine, and the mapping relationship between data point coordinates and physical values is extracted to generate a chart quantitative data set.
[0025] In the process of parsing multimodal contract documents, it is first necessary to decouple the multimodal features of the contract documents. For the mixed text clauses, tables and charts in the contract documents, the document structure analysis engine automatically identifies the file format, such as PDF and Word, and performs regional segmentation based on page layout features (such as paragraph indentation and column structure) and visual element distribution (such as lines and graphic boundaries). The text clause area extracts continuous semantic blocks through the paragraph recognition algorithm in natural language processing technology, while filtering out non-main content such as headers and footers; the table area adopts a table detection model based on deep learning, identifies the table frame features through a convolutional neural network, and determines the row and column boundaries in combination with the cell alignment rules; the chart area uses an image segmentation algorithm to extract independent graphic areas unrelated to text and tables, and distinguish between types such as bar charts and line charts.
[0026] The text parsing engine uses a hybrid semantic analysis approach when extracting textual clauses. First, a named entity recognition model (such as a fine-tuned model based on BERT) is used to locate key elements such as the responsible parties (e.g., the entity names of Party A and Party B) and the amount values (including currency units and numerical combinations). Dependency parsing is then used to construct a clause reference relationship graph. For example, when the phrase "in accordance with Article 3.2" appears in a clause, a bidirectional link is established between the current clause and Article 3.2. The resulting text feature vector combines word vector representations, entity type labels, and reference relationship topology to form a dimensionally unified semantic representation.
[0027] The table parsing engine adopts a two-step processing flow. First, the intersection of table lines is detected through morphological operations in the Open CV library to build the initial row and column structure framework, and then the cell merge relationship recognition algorithm driven by the attention mechanism is used to process complex cells across rows and columns. For each cell content, the system will generate metadata containing row and column coordinates and merge status identifiers, and generate a unique hash value for the cell content through the SHA-256 algorithm. For cells containing formulas (such as "=SUM(B2:B5)"), the formula dependency chain is parsed and a cross-cell numerical calculation relationship map is established to ensure that subsequent comparisons can identify formula logic differences.
[0028] The chart parsing engine decomposes chart elements through computer vision technology. The Mask R-CNN model is used to segment the coordinate axis, data graphics (such as bars and lines), and legend components, and optical character recognition technology is used to extract the coordinate axis scale value and unit information. For bar chart data points, the relative position of each bar in the image coordinate system is calculated and mapped to an actual value, such as inferring the specific value through the proportional relationship between pixel height and vertical axis scale. For line charts, the turning point coordinates are extracted through the key point detection algorithm, and the horizontal and vertical axis scales are combined to convert them into a time-value series. The final generated quantitative data set contains the physical value of the data point, statistical dimensions (such as quarter, product category), and the relationship between visualization elements.
[0029] Through the above technical solution, through the structured analysis of the table, deep differences such as changes in cell merging status and broken formula dependencies can be accurately identified; for example, when a supplier's quotation cell is horizontally merged in the price list table of two contract versions, traditional tools will mistakenly judge it as content deletion, while the embodiment of this specification can accurately identify that this is a typesetting optimization rather than a substantive clause change through cell hash value comparison and merge relationship analysis; at the non-text information processing level, conventional methods only perform image similarity comparison on chart content and cannot discover substantive changes at the data level. For example, a certain contract version changes the vertical axis scale of the "quarterly sales growth rate" line chart from percentage to absolute value. Traditional comparison will determine it as the same chart, while the embodiment of this specification can detect fundamental changes in data presentation through coordinate axis range analysis and data point value mapping, avoiding misunderstanding of clauses due to visual deception; through responsible party identification, time expression normalization and The analysis of the clause reference relationship can determine the substantive consistency of such clauses that are semantically equivalent but have different expressions, thereby reducing the false alarm rate. At the same time, the topological analysis of the clause reference relationship can detect logical contradictions caused by the modification of serial clauses. When processing complex quotation sheets containing hundreds of related formulas, conventional tools may cause cascading misjudgments due to the modification of the value of a certain cell. However, the embodiment of this specification can accurately locate the root modification point and evaluate the scope of influence through the dynamic tracking of the formula dependency chain. The quantitative mapping mechanism of the chart data points eliminates the interference caused by non-substantive modifications such as image scaling and color adjustment, thereby improving the ability to resist interference. The embodiment of this specification ensures data integrity through hash value verification and formula dependency chain tracking. When processing complex quotation sheets containing hundreds of related formulas, conventional tools may cause cascading misjudgments due to the modification of the value of a certain cell. However, the embodiment of this specification can accurately locate the root modification point and evaluate the scope of influence through the dynamic tracking of the formula dependency chain.
[0030] Step S102 : Based on the pre-built multimodal large model, the multimodal contract feature data corresponding to each contract document is compared to locate the difference area information between the multiple contract documents.
[0031] Based on a pre-built multimodal macromodel, the multimodal contract feature data corresponding to each contract document is compared to locate the difference areas between the multiple contract documents. This can be achieved through the following steps: First, the text feature vector, table structured data set, and chart quantization data set are input into the multimodal macromodel. The feature representations of different modalities are aligned through a cross-modal attention mechanism to generate multimodal contract feature data in a unified embedding space. Based on this multimodal contract feature data, a similarity matrix is determined between the multimodal contract feature data of the multiple contract documents to locate the revised locations of text clauses, the changed areas of table cells, and the offset intervals of chart data points, thereby generating a set of coordinates for the difference areas.
[0032] Among them, the specific implementation process of aligning the feature representations of different modalities through the cross-modal attention mechanism to generate multimodal contract feature data in a unified embedding space is as follows: semantically match the clause number in the text feature vector with the reference identifier in the table cell to establish a text-table attention weight matrix; perform unit consistency check on the physical quantity units corresponding to the chart data points in the table area and the numerical descriptions in the text clauses to establish a text-chart attention weight matrix; based on the text-table attention weight matrix and the text-chart attention weight matrix, the multimodal attention weights are fused through a gating mechanism to generate multimodal contract feature data in a unified embedding space.
[0033] The specific implementation process of locating the revised position of text terms, the changed area of table cells and the offset interval of chart data points is as follows: based on the BERT model, the semantic similarity of multiple text terms in the contract document is calculated, and when the similarity is lower than the first preset threshold, it is marked as a revised area; according to the cell hash value in the table structured data set, the hash value difference of the cells with the same row and column indexes in multiple contract documents is compared, and the cell coordinates with inconsistent hash values are marked as the changed area of the table cell; the Euclidean distance of data points with the same physical quantity coordinates is calculated, and the offset interval whose distance exceeds the preset dynamic threshold is marked to determine the offset interval of the chart data point.
[0034] In one embodiment of the present specification, during the contract discrepancy detection phase, the pre-processed text feature vectors, table structured data, and chart quantitative data are first input into a multimodal large model. The model uses a cross-modal attention mechanism to semantically match the text clause numbers with the reference identifiers in the table cells. For example, when the text clause appears with "See attached quotation table A-3 ”, it will automatically be associated with the specific table marked as "A-3" in the table area. This association is achieved by calculating the cosine similarity between the text word vector and the table identifier vector, generating an attention weight matrix that reflects the strength of the association between the text and the table. At the same time, the model will perform unit consistency verification on the numerical descriptions appearing in the text terms (such as the total amount does not exceed 5 million yuan) and the physical quantity units marked on the chart coordinate axes (such as 10,000 yuan, percentage). When a unit mismatch is detected (such as the text uses 10,000 yuan and the chart uses yuan), a cross-modal anomaly mark will be generated.
[0035] During the feature fusion process, a gating mechanism is used to dynamically adjust the attention weights of different modalities. For example, when the association weight between text and tables reaches 0.8 (the threshold is configurable), the gating unit will enhance the comparison strength of this modality pair; if there is an order of magnitude difference between the numerical values of the chart data points and the text description, a downgrade is triggered. This dynamic adjustment enables the feature representation of the unified embedding space to retain the independence of each modality while reflecting the logical associations across modalities. After multiple rounds of iteration, the multimodal contract feature data generated by the model contains three-dimensional feature vectors of the text semantic fingerprint, the table structure topology, and the spatial distribution of the chart data.
[0036] A hierarchical detection strategy is employed in the discrepancy detection phase. For text clauses, the BERT model is used to calculate the semantic similarity of corresponding clauses across different contract versions. A revision flag is triggered when the similarity falls below 0.75 (the threshold is adjustable depending on the contract type). For example, a clause stating "liquidated damages are 20% of the total contract amount" versus "the breaching party must pay liquidated damages of 25% of the total price" in two versions would be considered materially different. Table comparison utilizes a hash value comparison mechanism, accurate down to the cell level. When a cell with the same row and column coordinates exhibits a different hash value (e.g., cell B5 changes from "SHA256:9f86d08..." to "SHA256:5d4140..."), the cell is marked as a changed region, and the impact of its formula dependency chain is traced. Chart data point discrepancy detection utilizes a dynamic threshold algorithm, automatically adjusting the criteria based on data distribution. For example, in a sales line chart, if the Euclidean distance between data points at the same time point exceeds three standard deviations of the overall data fluctuation range, a significant shift is flagged.
[0037] Traditionally, when processing scenarios where textual clauses reference table data, only the text and table contents can be compared in isolation. However, the embodiments of this specification use an attention weight matrix to accurately capture substantive changes, such as the "see Table 3 data" in the example and the deletion of Table 3 in the new version of the contract. This cross-modal association detection effectively improves the detection rate of important clause changes and avoids the omission of major clauses caused by fragmented analysis in traditional methods. In terms of dynamic threshold application, the embodiments of this specification break through the limitations of fixed thresholds. Traditional chart comparisons use preset pixel difference thresholds, which will generate a large number of false positives when the contract version only adjusts the chart color or zoom ratio. The embodiments of this specification use a dynamic threshold mechanism that links the Euclidean distance of data points with the overall data distribution, which can both identify substantive data changes and ignore non-substantive visual adjustments. The hierarchical processing of difference positioning significantly improves the interpretability of the results. In the traditional method, only the difference position is marked. However, the embodiment of this specification uses additional information such as formula dependency chain tracing (such as marking that the modification of a quotation table cell affects the total price calculation formula) and semantic similarity quantification value (such as showing that the clause similarity drops from 0.82 to 0.63), so that legal personnel can quickly determine the nature of the difference, especially when dealing with chain clause modifications (such as the modification of a liability clause triggering a chain change of three referenced clauses), and automatically generate an impact relationship map.
[0038] Step S103 : establishing contract revision evolution information between multiple historical versions based on the difference area information and the contract version information of each contract file, so as to manage the multiple versions of the contract file through the contract revision evolution information.
[0039] Based on the difference area information and the contract version information of each contract file, the contract revision evolution information between multiple historical versions is established, specifically including: obtaining the difference area information to determine the difference area coordinate set; setting unique revision identification information for each difference area in the difference area coordinate set according to the contract version information of each contract file to bind the version interval corresponding to the difference area; determining the revision operation type corresponding to the difference area through the version interval, so as to generate the contract revision evolution information between multiple historical versions based on the revision operation type and the unique revision identification information and according to the pre-acquired version timestamp.
[0040] In one embodiment of the present specification, the difference region coordinate set is first parsed, and the difference region coordinate set includes the text clause revision position, table cell change area, and chart data point offset interval. A unique revision identifier (such as D_Text Clause_Article 3.2_V2) is generated for each difference region. → V3), bound to the corresponding version number range (such as version 2 to version 3).
[0041] Compare the attribute changes in the difference areas between adjacent versions. If a clause exists in the old version but disappears in the new version, it is marked as a deletion operation; if a clause appears in the new version that does not exist in the old version, it is marked as a new addition operation; if the content is modified but the position remains unchanged, it is marked as a content change. For changes in table structure (such as the addition and subtraction of rows and columns, and cell merging), the type of structural change (such as inserting rows and merging cells) is identified by comparing the row and column topology diagrams of the previous and next versions. Chart data point offsets are determined as data corrections or trend adjustments based on the rate of change in data point distribution density (such as a data point in a certain quarter moving from a dense area to an abnormal area).
[0042] Extract the revision operation type (add / delete / modify) in the version repository and associate the operation type label with the difference area coordinate. For example, table cell B5 is labeled as " Value modification ” , whose operation type is " Data correction ” If the diff area of version 3 refers to the terms of version 2 (such as " According to Article 5.1 of Version 2 ” ), then a cross-version dependency edge (V3_D1 → V2_D4). Arrange all versions in timestamp order. All operation types are bound to the timestamp sequence, forming a revision event chain with time stamps. Compare the difference areas between versions Vn and Vn-1, and record the added / deleted clauses, table rows, and chart data points. For modification operations, record the content summary before and after the modification, such as clause 3.2 from " Penalty 10% ” Change to " Liquidated damages are one tenth of the total contract amount ” Generates a machine-readable timeline log file containing the timing, impact, and context of each revision operation.
[0043] With the initial version as the root node, the revision events of each version are linked in chronological order. For clauses that are continuously modified across versions, such as a liability clause that has undergone three revisions between v1.0 and v3.0, a vertical evolution path is established to show the semantic drift trajectory of the clause content with version iteration. For associative modifications, such as the adjustment of three referenced clauses caused by the modification of table formulas, horizontal association edges are established through the dependency relationship between revision identifiers. The final generated contract revision evolution map contains a triple information network of spatiotemporal dimension (version timeline), operation dimension (addition, deletion and modification type) and impact dimension (association change chain), which supports multi-condition retrieval by time, responsible person, clause type, etc.
[0044] Through the above technical solution, multi-dimensional evolution tracking of contract elements is achieved. Conventional methods can only record the addition, deletion and modification of text lines, while the version tracking capability of this specification embodiment is accurate to the cell and data point level, which can identify implicit changes such as table structure adjustment and chart data tampering; in addition, in the revision impact analysis solution, this specification embodiment realizes cross-modal association tracing. When a clause references table data that changes, a cross-modal impact chain of "clause-table-chart" is automatically generated, so that legal personnel can quickly understand the intention of the revision; in addition, temporal and spatial evolution visualization can be achieved. Through the interactive timeline, the evolution process of the responsible party of the liability clause in multiple revisions can be obtained, and the change trajectory of the compensation calculation formula of the relevant table can be displayed in conjunction. For contract versions involving multi-party negotiations (such as temporary versions such as v1.2a and v1.2b), the modification process of different stakeholders can be restored through the reviser tags and annotation information.
[0045] Before managing the multiple versions of the contract documents through the contract revision evolution information, the method also includes: performing a cross-modal consistency check on the multimodal contract feature data; when the consistency check does not meet the preset conditions, determining a cross-modal conflict evidence chain to update it to the contract revision evolution information.
[0046] In one embodiment of this specification, a multimodal feature association map is first constructed. Based on the numerical description in the text feature vector, the formula calculation results in the table structured data, and the visual numerical values in the chart quantitative data, a cross-modal numerical association is established through a semantic alignment engine. For example, when a text clause mentions "annual growth rate of 15% ”, the system automatically associates the slopes of the data points of the corresponding years in the chart, and verifies whether the output results of the growth rate calculation formula in the table match. This association is achieved through the cross-modal attention mechanism of the multimodal large model, calculating the spatial distance between the text word vector, the table formula vector and the chart data vector to form a three-dimensional verification matrix. When it is detected that the cross-modal numerical deviation exceeds the preset tolerance range (such as the difference between the text and the chart value >5%), the deep verification process is initiated: first, trace whether the responsible party in the text has a corresponding entity in the table contracting party list, secondly verify whether the time dimension of the chart data point matches the effective date of the terms, and finally review whether the input parameters of the table formula contain the latest revised values. For example, if the text description "down payment 30%" is found to be in conflict with the value of the first payment cell in the payment schedule of 25%, it will be marked as a major conflict, and the impact chain analysis will be used to locate the chain error caused by a formula modification in version v2.3. The version trajectory of the modal elements involved in the conflict is determined, such as the 30% down payment ratio established in the text clause in version 1.5, the calculation formula modified due to tax rate adjustments in the table in version 2.1, and the updated data visualization rules in the chart in version 2.3. By analyzing the revision timeline of each modal element, a three-dimensional evidence chain is generated, including temporal conflicts (such as chart data predating the effective date of the clause), logical conflicts (such as formula output inconsistent with the text description), and substantive conflicts (such as the contracting party not appearing in the table). This evidence is stored in a hypergraph structure, with nodes representing conflict points and hyperedges connecting cross-modal revision events to determine the cross-modal conflict evidence chain and update it in the evolutionary graph of the contract revision evolution information.
[0047] Through the above technical solution, three-dimensional contradiction detection of contract elements can be achieved, and hidden cross-modal conflicts can be captured; the generated chain of evidence not only includes the current conflict point, but also shows the evolution path of related elements in historical versions. This spatiotemporal correlation analysis enables legal personnel to quickly identify the root cause of the conflict, effectively improving tracing efficiency compared to traditional manual tracing methods.
[0048] The contract revision evolution information is used to manage the multiple versions of the contract documents, specifically including: generating a compliance risk assessment report between contract versions based on the revision operation type in the contract revision evolution information, wherein the compliance risk assessment report contains the clause location coordinates and associated evidence screenshots of high-risk revision items; in the contract management interface, the contract revision evolution information and the compliance risk assessment report are displayed to trace back the cross-modal impact path of a specific revision operation through user interaction.
[0049] In one embodiment of the present specification, based on the revision operation type in the contract revision evolution information, a compliance risk assessment report between contract versions is generated, a dynamic risk indicator model is constructed, a compliance rule base is preset according to the contract type (such as a procurement contract, a cooperation agreement), and the revision operation type is mapped to the risk dimension. The deletion of a clause may involve the risk of liability exemption, the modification of a formula may trigger a calculation logic risk, and the offset of chart data may imply the risk of data tampering. Each revision operation will be assigned a basic risk value, such as a clause deletion risk weight of 0.8 and a format adjustment weight of 0.2, and the derivative risk will be calculated through an impact factor matrix. For example, if a payment clause deletion operation is associated with three breach of contract liability clauses that reference the clause, the risk value will be exponentially amplified based on the length of the reference chain. At the same time, it is connected to an external regulatory database to match clause changes with the latest legal provisions in real time. When a high-risk revision is identified, such as a revision with a risk value greater than 0.9, the cross-modal evidence chain of the difference area is automatically intercepted, including the semantic comparison of the text clauses before and after the revision, the formula dependency tree change map caused by the table cell modification, and the three-dimensional heat map of the chart data point offset. Among them, the semantic comparison of the text clauses before and after the revision can be visualized through the BERT similarity change curve. The physical coordinates of the contract document are accurately located (such as area B on page 15 of the PDF), and a high-resolution annotated image is generated through the rendering engine, and the revision timeline mark (such as "v2.1 → v3.0" arrows indicate the direction of modification). The report is structured and stored as a composite document containing a risk level matrix, evidence package, and remediation recommendations, allowing drill-down viewing of the impact path of each risk item.
[0050] The evolutionary graph of contract revisions uses time as its primary axis, with spatial distribution showing the risk radiation range of each revision point. When a user clicks on a high-risk node (such as the liquidated damages clause in Article 8 of Version 3.0), the cross-modal impact path of the clause will pop up on the right side of the interface: drilling down to display the associated payment schedule formula changes, and tracing back to display the modification records of the arbitration clause in Version 2.5 that triggered this revision. The three-dimensional visualization view can be rotated and zoomed to view multi-version risk clustering, such as aggregating all standard clause adjustments over a period of 12 months into a low-risk cloud, while core clause modifications form a high-risk pulse peak. Touch gestures allow users to drag along the timeline to observe the risk propagation process, and pinch-to-zoom can focus on the upstream and downstream association network of a specific revision event.
[0051] Through the above technical solution, intelligent penetration of risk identification is achieved. Through the dynamic risk model, complex risks hidden in multimodal modifications can be discovered; it breaks through the one-dimensional limitations of traditional text reports. Through the visual integration of cross-modal evidence chains, reviewers can intuitively see the three-dimensional impact of risk points, and multimodal correlation display effectively improves risk understanding; through the time and space dimensions, the purpose of penetrating analysis of risk evolution laws can be achieved.
[0052] Through the technical solutions of the embodiments of this specification, through the deep integration of multimodal large models, holographic parsing and intelligent correlation analysis of contract elements are realized. Traditional technologies are limited by the processing capabilities of a single modality and can only perform shallow character comparison on contract texts. However, the embodiments of this specification can accurately capture the cross-modal correlation relationships implicit in contract terms by fusing text semantics, table structure, and multimodal feature extraction of chart data; at the level of version difference detection, it breaks through the limitations of traditional linear comparison and realizes three-dimensional tracking of contract revisions; traditional technologies rely on manual marking or simple hash comparison, and cannot identify complex scenarios such as the break of formula dependency chains caused by adjustments to table row and column structures, and substantive changes in terms implied by chart data point offsets; the embodiments of this specification pass The unified embedding space mapping of multimodal large models can quantitatively evaluate the semantic deviation of text clauses, the topological change rate of table structures, and the difference value of chart data distribution, and accurately locate the substantive revision area; traditional version control systems (such as Git) only record file-level changes, while the embodiment of this specification uses the version alignment technology of the multimodal feature coordinate system to establish fine-grained revision tracks at the clause level, cell level, and data point level. The atomic version traceability capability allows legal personnel to conduct a penetrating analysis of the semantic drift path of a certain liability clause in multiple historical versions, or observe the parameter evolution law of a certain price calculation formula after multiple revisions; in addition, the embodiment of this specification constructs an intelligent compliance protection system through cross-modal consistency verification and dynamic risk assessment models. Compared with the lagging risk control of traditional manual review relying on experience judgment, the embodiment of this specification can detect the implicit contradictions between text clauses, table formulas, and chart data in real time, achieving an exponential improvement in risk prevention and control efficiency.
[0053] The embodiment of this specification also provides a multi-modal large model intelligent contract management device, such as Figure 2 As shown, the device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the above method.
[0054] The embodiments of this specification also provide a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured to execute the above method.
[0055] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from the other embodiments. In particular, the device, apparatus, and non-volatile computer storage medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simplified. For relevant details, refer to the descriptions of the method embodiments.
[0056] The devices and media provided in the embodiments of this specification correspond one-to-one to the methods. Therefore, the devices and media also have similar beneficial technical effects to their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.
[0057] The foregoing description is merely one or more embodiments of this specification and is not intended to limit this specification. It will be apparent to those skilled in the art that various modifications and variations may be made to one or more embodiments of this specification. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of one or more embodiments of this specification are intended to be within the scope of the claims of this specification.
Claims
1. A multimodal large-scale smart contract management method, characterized in that: The method comprises: Acquiring multiple versions of pre-stored contract documents, performing multimodal feature extraction on each of the contract documents, and determining multimodal contract feature data, wherein the multimodal contract feature data includes any one or more of a text feature vector, a table structured data set, and a chart quantitative data set; Based on a pre-built multimodal large model, the multimodal contract feature data corresponding to each of the contract documents is compared to locate the difference area information between the multiple contract documents; Contract revision evolution information between multiple historical versions is established based on the difference area information and the contract version information of each contract document, so as to manage the multiple versions of the contract documents through the contract revision evolution information.
2. The multimodal large model smart contract management method according to claim 1, characterized in that: Performing multimodal feature extraction on each of the contract documents to determine multimodal contract feature data specifically includes: Performing multimodal feature decoupling on each of the contract documents to determine the text clauses, table areas, and chart areas corresponding to each of the contract documents; Extract the responsible party, amount value and clause reference relationship in the text clauses through a text parsing engine to generate a text feature vector; Identifying the row and column structure of the table area through a table parsing engine, extracting cell coordinates and merging relationships, and generating a table structured data set, wherein the table structured data set includes cell hash values and formula dependency chains; The coordinate axes, data graphics and legend components in the chart area are segmented by a chart parsing engine, a mapping relationship between data point coordinates and physical quantity values is extracted, and a chart quantitative data set is generated.
3. The multimodal large model smart contract management method according to claim 1, characterized in that: Based on the pre-built multimodal large model, the multimodal contract feature data corresponding to each of the contract documents is compared to locate the difference area information between the multiple contract documents, specifically including: Input the text feature vector, table structured data set, and chart quantization data set into a multimodal large model, align the feature representations of different modalities through a cross-modal attention mechanism, and generate multimodal contract feature data in a unified embedding space; Based on the multimodal contract feature data, a similarity matrix between the multimodal contract feature data of multiple contract documents is determined to locate the revised position of text terms, the changed area of table cells and the offset interval of chart data points, and generate a set of difference area coordinates.
4. The multimodal large model smart contract management method according to claim 3, characterized in that: The cross-modal attention mechanism is used to align the feature representations of different modalities and generate multimodal contract feature data in a unified embedding space. Specifically, Performing semantic matching on the clause numbers in the text feature vector and the reference identifiers in the table cells to establish a text table attention weight matrix; Perform unit consistency check on the physical quantity units corresponding to the chart data points in the table area and the numerical descriptions in the text clauses, and establish a text chart attention weight matrix; According to the text table attention weight matrix and the text chart attention weight matrix, multimodal attention weights are fused through a gating mechanism to generate multimodal contract feature data in a unified embedding space.
5. The multimodal large model smart contract management method according to claim 3, characterized in that: Position the revised position of text clauses, the changed area of table cells, and the offset range of chart data points, including: Calculating semantic similarity of text clauses in the plurality of contract documents based on the BERT model, and marking as a revision region when the similarity is lower than a first preset threshold; Comparing hash values of cells with the same row and column indexes in a plurality of contract documents based on the hash values of the cells in the table structured data set, marking the coordinates of cells with inconsistent hash values as the table cell change areas; The Euclidean distance of data points with the same physical quantity coordinates is calculated, and the offset intervals where the distance exceeds a preset dynamic threshold are marked to determine the offset intervals of the chart data points.
6. The multimodal large model smart contract management method according to claim 1, characterized in that: Establishing contract revision evolution information between multiple historical versions based on the difference area information and the contract version information of each contract document, specifically including: Acquire the difference area information to determine a difference area coordinate set; According to the contract version information of each of the contract files, unique revision identification information is set for each difference area in the difference area coordinate set to bind the version interval corresponding to the difference area; The revision operation type corresponding to the difference area is determined through the version interval, so as to generate contract revision evolution information between multiple historical versions based on the revision operation type and the unique revision identification information and according to the pre-acquired version timestamp.
7. The multimodal large model smart contract management method according to claim 1, characterized in that: Before managing the multiple versions of the contract documents using the contract revision evolution information, the method further includes: A cross-modal consistency check is performed on the multimodal contract feature data. When the consistency check does not meet a preset condition, a cross-modal conflict evidence chain is determined to update the contract revision evolution information.
8. The multimodal large model smart contract management method according to claim 1, characterized in that: The multiple versions of the contract documents are managed through the contract revision evolution information, specifically including: Based on the revision operation type in the contract revision evolution information, generating a compliance risk assessment report between contract versions, wherein the compliance risk assessment report includes clause location coordinates of high-risk revision items and screenshots of associated evidence; In the contract management interface, the contract revision evolution information and the compliance risk assessment report are displayed to trace back the cross-modal impact path of a specific revision operation through user interaction.
9. A multimodal large-scale intelligent contract management device, characterized in that: The device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.
10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions are configured to execute the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Text comparison method, contract review method and review system
CN112926299A
Contract historical version and execution progress management method and system
CN118278627A
Electronic contract auditing method and device, nonvolatile storage medium and electronic equipment
CN118918602A
Contract management method and equipment based on information retrieval and medium
CN118964443A
Contract auditing method, device and equipment and computer readable storage medium
CN119600634A
Cited By
Intelligent management and control method for port business contract rate
CN121281074A