Intelligent contract analysis method and system based on multi-modal feature fusion algorithm

By using a multimodal feature fusion algorithm, a collaborative model of visual, topological, and textual features is constructed, which solves the problem of insufficient accuracy in contract analysis in existing technologies and achieves more efficient contract content classification.

CN120822115AActive Publication Date: 2025-10-21LUDAN (SHANDONG) DATA TECH CO LTD

Patent Information

Application Number
CN202511341359.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2025-10-21
Estimated Expiration
2045-09-19

AI Technical Summary

Technical Problem

Existing technologies in contract analysis suffer from insufficient accuracy due to limitations in the ability of visual models to resolve ambiguities in legal terminology, high text localization bias in single-modal deep learning models, and limitations in topological feature extraction in multimodal large models.

Method used

A multimodal feature fusion algorithm is adopted. By constructing a multimodal collaborative feature extraction model, visual features, topological features, and textual features of contract image blocks are obtained respectively. Weighted fusion is performed using an attention mechanism, and topological features are captured by an edge enhancement graph convolutional network. This solves the problem of one-sidedness of single-modal feature information and enhances the expressive power of features.

Benefits of technology

It improved the accuracy of contract analysis, reduced layout errors, enhanced the understanding of complex layouts and structured information, and improved the precision of contract content classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120822115A_ABST
    Figure CN120822115A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent contract analysis method and system based on a multi-modal feature fusion algorithm, and relates to the technical field of intelligent contract analysis, and the method comprises the steps: obtaining a PDF contract document, and segmenting the PDF contract document into a plurality of image blocks; constructing a multi-modal collaborative feature extraction model to obtain a visual feature, a topological feature and a text feature of each image block; and performing weighted fusion on the three features of the same image block by using an attention mechanism, inputting the fused features into the pre-training classification model, and outputting a classification result, so that the accuracy of a contract analysis result can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of smart contract analysis, and in particular to a smart contract analysis method and system based on a multimodal feature fusion algorithm. Background Art

[0002] In today's highly intertwined world of business and legal affairs, contracts, as crucial legal documents that clarify the rights and obligations of all parties, ensure transaction security, and regulate transactions, are experiencing explosive growth in their number. Contracts play an indispensable role in numerous areas, from business collaborations between businesses to credit operations by financial institutions to administrative approvals by government agencies. With the advancement of economic globalization and the widespread application of digital technologies, the complexity and diversity of contracts are increasing, encompassing a variety of information formats, including text, images, and tables. Therefore, efficient and accurate contract analysis to extract key information, identify potential risks, and determine contract types has become a pressing need across various industries. Smart contract analysis technology has emerged, leveraging advanced artificial intelligence algorithms and computer technology to automate and intelligently process contracts, improving the efficiency and quality of contract management while reducing labor costs and error rates.

[0003] The current field of intelligent contract document analysis has developed a technology ecosystem dominated by rule engines, single-modal learning, and general large-scale models. These solutions exhibit significant performance differences in actual application scenarios. OCR-based rule-matching methods offer advantages in standardized form parsing scenarios, and their preset template mechanism ensures a stable 82%-87% accuracy rate for extracting key fields. However, when faced with unstructured legal clauses, the lack of semantic understanding results in an accuracy rate of less than 41% for complex rights and obligations clauses. While single-modal deep learning models have achieved breakthroughs in specific dimensions, the text model's neglect of layout features leads to a 29% deviation in clause positioning. Visual models are particularly vulnerable to the inability to resolve ambiguity in legal terminology.

[0004] The multimodal large model technology that has emerged in recent years theoretically has the potential for cross-modal fusion. For example, the Chinese invention patent with application publication number CN118734032A provides a contract information extraction method and device based on a multimodal large language model. The patent first obtains contract data in different formats to construct a model training data set; then constructs a latent space mapping model between contract images and text data; trains this mapping model with the training data set, and obtains a universal text and image data encoder based on the verification results; then constructs a contract extraction large language model based on these two encoders; finally, trains the large model with the training data set and verifies its information extraction accuracy.

[0005] However, this patent processes contract image data through patch segmentation, encoding and position encoding. Although it can capture local semantics and basic spatial relationships, it still has significant limitations in topological feature extraction (such as complex layout, structured information, and nonlinear spatial dependencies). If the contract contains a cross-page table, patch segmentation may cause the rows and columns of the same table to be independently encoded on different pages, destroying the topological consistency of the cross-page table. Summary of the Invention

[0006] In order to improve the accuracy of contract analysis results, this application provides an intelligent contract analysis method and system based on a multimodal feature fusion algorithm.

[0007] In the first aspect, this application provides a smart contract analysis method based on a multimodal feature fusion algorithm, which adopts the following technical solutions: The smart contract analysis method based on the multimodal feature fusion algorithm includes the following steps: Data collection and processing: Obtain PDF contract documents, segment them, and obtain multiple image blocks; Feature extraction: Build a multimodal collaborative feature extraction model and input each image block into the multimodal collaborative feature extraction model to obtain the visual features, topological features, and text features of each image block. Classification: For the visual features, topological features, and text features of the same image block, the attention mechanism is used to perform weighted fusion to obtain fused features. The fused features are input into the pre-trained classification model, and the classification results of the fused features are output.

[0008] This application constructs a multimodal collaborative feature extraction model to extract the visual, topological, and textual features of contract image blocks. It then assigns different weights to each feature based on its importance, highlighting the role of key features. The fused features are then fed into a pre-trained classification model. The fused features contain more comprehensive and effective information, helping the classification model to more accurately classify the contract content. PDF contract documents often contain complex hierarchical relationships, and traditional visual features are susceptible to typesetting formatting. By analyzing the topological relationships of PDF documents to restore the logical hierarchy and reduce typesetting deviations, the accuracy of contract analysis results can be improved.

[0009] Optionally, the multimodal collaborative feature extraction model includes a visual feature extraction sub-model for extracting visual features of image blocks, a topological structure feature extraction sub-model for extracting topological features between image blocks, and a text semantic understanding sub-model for extracting text features of image blocks.

[0010] The multimodal collaborative feature extraction model deconstructs the information of contracted image blocks from different dimensions through the division of labor and collaboration among three sub-models: the visual feature extraction sub-model, the topological structure feature extraction sub-model, and the text semantic understanding sub-model. This not only solves the one-sidedness problem of single-modal feature information, but also provides input data for the attention mechanism to perform weighted fusion processing of features from different modalities.

[0011] Optionally, the visual feature extraction submodel includes a CSPDarknet53 network integrated with a CBAM layer and a bidirectional feature pyramid network. The CSPDarknet53 network is used to perform five-level downsampling processing on the image block to obtain a multi-scale feature map. The bidirectional feature pyramid network is used to enhance the target area in the multi-scale feature map and output the enhanced multi-scale feature map. The enhanced multi-scale feature map is weighted by the CBAM layer to output the visual features.

[0012] This application uses the CSPDarknet53 network to perform five-level downsampling on image blocks, gradually expanding the receptive field and generating multi-scale feature maps from small-scale high resolution to large-scale low resolution. It can capture microscopic details while retaining macroscopic structural information, which helps to fully understand the image content and improve the detection and recognition capabilities of targets of different sizes.

[0013] Since the multi-scale feature maps generated by the CSPDarknet53 network have a hierarchical split problem, that is, the low-level feature maps are rich in details but lack global semantics, and the high-level feature maps have clear structures but lose details, this application adopts a bidirectional feature pyramid network to fuse features of different scales with the help of a bidirectional information flow mechanism, so that the information of the target area can be more fully transmitted and enhanced in the features of each scale, further highlighting the target area, and improving the positioning accuracy and classification accuracy of the target area. The CBAM layer automatically identifies and strengthens key visual features through the dual weighting of channel attention and spatial attention, reducing the risk of useful information being obscured by redundant information during feature fusion.

[0014] By adopting the above scheme, this application converts the unstructured visual information of the contract image block into a structured visual vector that fits the contract business needs, highlights key signals, and is multi-dimensionally complementary. It not only solves the problem that traditional visual extraction cannot achieve both details and structure and is interfered by redundant information, but also enables visual features to be accurately matched with topological features and text features, further improving the accuracy of contract analysis results.

[0015] Optionally, the steps of data collection and processing include: Parse the PDF contract document to obtain a page pixel map. Use the OpenCV contour detection algorithm to separate the image blocks in the page pixel map. Project the bounding boxes of all image blocks into the same coordinate system to obtain the coordinate values ​​of each vertex of each image block. Based on the coordinate values ​​of each vertex of each image block, a coordinate index of each image block is established.

[0016] This application parses the PDF contract document to generate a page pixel map, which can not only convert the logical content of the PDF into image data with a unified visual form, but also fully retain the visual details in the contract. If the PDF contract document is directly subjected to text extraction and structural analysis, non-text visual information such as color, annotations, and signatures will be lost, resulting in the loss of key dimensions in multimodal feature extraction. Subsequently, this application uses the OpenCV contour detection algorithm to separate the areas with continuous edges in the PDF contract document, thereby splitting each page into multiple independent image blocks, and projecting the bounding boxes of all image blocks into the same coordinate system to obtain the coordinate values ​​of each vertex of each image block, and establish a coordinate index based on the coordinate values ​​of each vertex of each image block, providing a mechanism for rapid search and positioning of image blocks. By establishing a coordinate index, this application can quickly find the corresponding image block based on the coordinate information of the image block, thereby improving the management efficiency and operational convenience of image blocks in the document.

[0017] Optionally, the text semantic understanding sub-model includes a CRNN sub-model and an MLP sub-model. The process of extracting text features by the text semantic understanding sub-model is as follows: The CRNN sub-model locates and identifies image blocks of text type, which are recorded as text blocks. The geometric parameters of the text blocks are obtained, including the center coordinates and width and height. The tilt angle of the text blocks is calculated based on the geometric parameters. Determine whether the tilt angle is zero. If so, do not process; if not, perform affine transformation on the text block based on the tilt angle to obtain a corrected text block, and update the corrected text block as the text block; The geometric parameters of the text block are normalized, and the MLP sub-model is used to map the normalized geometric parameters into a position vector. The coordinate index of the text block is obtained, and the coordinate index is encoded to obtain the original position encoding vector. The position vector is added to the original position encoding vector to obtain the text feature.

[0018] The present application uses a CRNN sub-model to locate and identify image blocks of text type (i.e., text blocks), which can effectively locate the area containing text in the image block and accurately identify the text content therein. Subsequently, the present application calculates the tilt angle based on the center coordinates and width and height of the text block to quantify the degree of tilt of the text block in the image. The present application determines whether the tilt angle is zero and performs an affine transformation on the text block with a non-zero tilt angle to obtain a corrected text block. The corrected text block is more in line with the normal reading direction, which is conducive to subsequent more accurate extraction of text features and reduces the recognition errors and feature extraction difficulties caused by text tilt. Subsequently, the present application normalizes the geometric parameters of the text block (i.e., center coordinates and width and height) to eliminate the differences between different text blocks due to factors such as size and position. Afterwards, the MLP sub-model is used to map the normalized geometric parameters into a 768-dimensional position vector, thereby converting the geometric information of the text block into a high-dimensional vector representation, adding position-related semantic information to the text features.

[0019] Subsequently, this application obtains the coordinate index of the text block and performs encoding processing to obtain the original position encoding vector, adds the 768-dimensional position vector to the original position encoding vector, and finally obtains the text feature. The above scheme can combine position information and geometric information, so that the text feature not only contains the semantic information of the text itself, but also incorporates the position and geometric features of the text in the image, enriching the connotation of the text feature and improving the expressiveness and semantic understanding ability of the text feature.

[0020] Optionally, the classification step further includes: According to the coordinate index of each image block, the image blocks belonging to the same page are integrated into an image block set, the dimension of the attention matrix in the attention mechanism is obtained, and a binary mask matrix with the same dimension as the attention matrix is ​​constructed. The element in the i-th row and j-th column of the binary mask matrix indicates whether the i-th image block and the j-th image block belong to the same page. If so, the value of the element is 1; otherwise, the value of the element is 0; Multiply the binary mask matrix and the attention matrix element-wise to obtain the multiplied matrix, and use the multiplied matrix as the new attention matrix.

[0021] This application first obtains the dimension of the attention matrix in the attention mechanism, and then constructs a binary mask matrix with the same dimension as the attention matrix. This application sets the value rules of the elements in the i-th row and j-th column of the matrix to clearly mark which image blocks are associated with each other on the same page. Subsequently, this application multiplies the constructed binary mask matrix and the attention matrix element by element. After multiplication, only the attention matrix elements corresponding to the image blocks on the same page maintain their original values, while the attention matrix elements corresponding to the image blocks on different pages become 0. The new attention matrix finally obtained focuses more on the relationship between the image blocks on the same page, highlighting the local structure and association information within the page.

[0022] This application constructs a binary mask matrix and adjusts the attention matrix so that the multimodal collaborative feature extraction model can focus more on local information within the same page when processing image blocks. In tasks such as document classification, elements within the same page usually have closer semantic and structural associations. The above method helps the multimodal collaborative feature extraction model to better capture these local features and improve the accuracy of understanding the page content.

[0023] When processing multi-page documents, image blocks across different pages may share some similar features, but their semantic connections are typically weak. The existing attention matrix can be disrupted by these cross-page similarities, leading to a bias in the multimodal collaborative feature extraction model's understanding of page content. However, by using a binary mask matrix, the new attention matrix eliminates the correlation between image blocks across pages, reducing this interference and enabling the multimodal collaborative feature extraction model to more accurately extract key information within the same page.

[0024] Optionally, the classification step further includes: Get the center coordinates of the page corresponding to the i-th image block set, calculate the Manhattan distance between the image block where the center coordinates of the page corresponding to the i-th image block set are located and the remaining image blocks in the i-th image block set, construct an attention bias term based on the Manhattan distance, and take the sum of the attention bias term and the attention score in the attention mechanism as the new attention score.

[0025] Based on the coordinate index of each image block, this application integrates the image blocks on the same page into an image block set, thereby grouping the image blocks belonging to the same visual level. Then, by obtaining the center coordinates of the i-th image block set and calculating the Manhattan distance between the image block where the center coordinates are located and the remaining image blocks in the set, by adopting the above scheme, this application can measure the relative position relationship of the image blocks on the page and quantify the degree of spatial correlation between different image blocks on the same page.

[0026] Subsequently, this application constructs an attention bias term based on the Manhattan distance. The introduction of the attention bias term enables the attention mechanism to take into account the spatial position relationship between image blocks, rather than just the similarity of the features themselves. This allows the model to pay more attention to image blocks that are spatially close to the central image block when processing image blocks, which is more in line with humans' cognitive habits of spatial layout when observing documents. Finally, this application adds the attention bias term to the attention score in the attention mechanism as a new attention score. The new attention score comprehensively considers feature similarity and spatial position relationship, so that the fused features can more comprehensively reflect the information of the image block, which helps to improve the accuracy of classification.

[0027] Optionally, the topology feature extraction sub-model includes an edge-enhanced graph convolutional network, and the process of extracting topology features by the topology feature extraction sub-model is as follows: Image blocks are used as nodes, and the features of the nodes include: the center coordinates, width, height and visual features of the image blocks. The Euclidean distance between image blocks is calculated based on the center coordinates, and edges are constructed between image blocks whose Euclidean distance is less than a preset distance threshold to obtain a dynamic graph structure. The edge-enhanced graph convolutional network is used to capture the topological features in the dynamic graph structure.

[0028] This application uses image blocks as nodes, and the node features include the center coordinates, width and height, and visual features of the image blocks. The center coordinates and width and height provide information about the position and size of the image blocks in the image, which helps the multimodal collaborative feature extraction model understand the spatial relationship between image blocks; the visual features include information such as the semantics and texture of the image blocks themselves. The above scheme enables the multimodal collaborative feature extraction model to describe image blocks from multiple dimensions.

[0029] The application then calculates the Euclidean distance based on the center coordinates of the image blocks and constructs edges between image blocks whose Euclidean distance is less than a preset distance threshold, thereby forming a dynamic graph structure. In an image, image blocks that are closer are often more likely to belong to the same object or have related semantic information. By constructing edges in this way, these related image blocks can be connected.

[0030] This application uses edge-enhanced graph convolutional networks to capture topological features in dynamic graph structures. Graph convolutional networks can update node features by aggregating information about nodes and their neighboring nodes, thereby learning local and global information in the graph structure. The edge enhancement mechanism can further highlight the importance of edges in the graph structure, allowing the multimodal collaborative feature extraction model to pay more attention to the connection relationship between nodes, thereby more accurately capturing topological features. In image classification tasks, combining visual features and topological features can enable the model to more accurately distinguish between images of different categories, especially in those image pairs with similar visual features but different topological structures. Topological features can complement visual features and jointly enhance the model's ability to understand images.

[0031] Optionally, the classification step further includes: The visual features, topological features, and textual features belonging to the same image block are aligned and spliced ​​to obtain a splicing vector. The splicing vector is input into a pre-built MLP model to obtain the gating weights of the visual features, topological features, and textual features respectively. The product of the visual feature and the corresponding gating weight is used as a new visual feature; the product of the topological feature and the corresponding gating weight is used as a new topological feature; and the product of the text feature and the corresponding gating weight is used as a new text feature.

[0032] This application first aligns the dimensions of the visual features, topological features, and text features belonging to the same image block, then concatenates the aligned features head to tail to obtain a concatenation vector, integrating features from different sources into a unified vector space. The concatenation vector is then input into a pre-built MLP model, which learns the importance of different features in the classification task and outputs the gating weights of the visual features, topological features, and text features, respectively. Subsequently, this application multiplies the visual features, topological features, and text features by the corresponding gating weights to obtain new visual features, new topological features, and new text features. Through this weighted operation, this application can dynamically adjust the contribution of each feature based on the importance of different features in the current classification task. This application uses a gating mechanism to dynamically adjust the weights of different features, thereby achieving dynamic feature fusion. Compared with traditional fixed weight fusion methods, the dynamic fusion method adopted by this application can better adapt to the needs of different image blocks and different classification tasks. Through weighted operations, this application can highlight features that are more valuable to the classification task and suppress the influence of irrelevant or redundant features, thereby improving the overall expressiveness of the features.

[0033] Secondly, this application provides a smart contract analysis system based on a multimodal feature fusion algorithm, which adopts the following technical solutions: The smart contract analysis system based on multimodal feature fusion algorithm includes: The data acquisition and processing module is used to obtain the PDF contract document, segment the PDF contract document, and obtain multiple image blocks; The feature extraction module is used to build a multimodal collaborative feature extraction model, input each image block into the multimodal collaborative feature extraction model, and obtain the visual features, topological features, and text features of each image block respectively; The classification module is used to perform weighted fusion of the visual features, topological features and text features of the same image block using the attention mechanism to obtain fused features, input the fused features into the pre-trained classification model, and output the classification results of the fused features.

[0034] In summary, this application includes at least one of the following beneficial technical effects: 1. This application constructs a multimodal collaborative feature extraction model to obtain the visual, topological, and textual features of contract image blocks. It then assigns different weights to each feature based on its importance, highlighting the role of key features. The fused features are then fed into a pre-trained classification model. The fused features contain more comprehensive and effective information, helping the classification model to more accurately classify the contract content. PDF contract documents often contain complex hierarchical relationships, and traditional visual features are susceptible to typesetting formatting. By analyzing the topological relationships of PDF documents, the logical hierarchy is restored, typesetting deviations are reduced, and the accuracy of contract analysis results is improved.

[0035] 2. The multimodal collaborative feature extraction model deconstructs the information of contracted image blocks from different dimensions through the division of labor and cooperation of three sub-models: the visual feature extraction sub-model, the topological structure feature extraction sub-model, and the text semantic understanding sub-model. It not only solves the one-sidedness problem of single modality feature information, but also provides input data for the attention mechanism to perform weighted fusion processing of features from different modalities.

[0036] 3. This application uses edge-enhanced graph convolutional networks to capture topological features in dynamic graph structures. Graph convolutional networks can update node features by aggregating information about nodes and their neighboring nodes, thereby learning local and global information in the graph structure. The edge enhancement mechanism can further highlight the importance of edges in the graph structure, allowing the multimodal collaborative feature extraction model to pay more attention to the connection relationship between nodes, thereby more accurately capturing topological features. In image classification tasks, combining visual features and topological features can enable the model to more accurately distinguish between images of different categories, especially in those image pairs with similar visual features but different topological structures. Topological features can complement visual features and jointly enhance the model's ability to understand images. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is a flow chart of Example 1 of the present application; Figure 2 This is a flowchart of Example 2 of the present application. DETAILED DESCRIPTION

[0038] The following combination Figure 1 and Figure 2 This application is described in further detail.

[0039] Example 1: This example discloses a smart contract analysis method based on a multimodal feature fusion algorithm. Figure 1The method includes: S11 data collection and processing, S12 feature extraction and S13 classification, obtaining a PDF contract document and dividing it into multiple image blocks; then building a multimodal collaborative feature extraction model to obtain visual features, topological features and text features of each image block; then using an attention mechanism to weightedly fuse the three features of the same image block, inputting the fused features into a pre-trained classification model and outputting a classification result. The execution process of each step in this embodiment is as follows: S11 data acquisition and processing, obtain PDF contract documents, and use non-uniform grid division technology based on specific elements in the PDF contract documents (such as titles, paragraphs, tables, etc.), initially divide them into units of 80×80 pixels, and perform secondary subdivision on areas containing high-frequency details (such as dense text or complex charts) to obtain multiple image blocks.

[0040] In other embodiments, the S11 data collection and processing further includes: Adaptively binarize the PDF contract document and use a minimum bounding rectangle-based skew correction algorithm to eliminate deformation caused by document rotation or wrinkles. A non-local means denoising algorithm is also used to reduce image noise, resulting in a preliminarily processed PDF contract document.

[0041] The PyMuPDF and other tools are used to parse the PDF contract document after preliminary processing to obtain the page pixel map. Based on the page pixel map, the OpenCV contour detection algorithm is used to separate the vector graphic area and the bitmap area, i.e., the image block, in the PDF document after preliminary processing.

[0042] In document analysis tasks, the consistency of the coordinate system is the basis for subsequent modeling. If pixel coordinates are used directly, different resolutions or devices will lead to inconsistent geometric descriptions of the same element. Therefore, this embodiment converts the pixel coordinates of the bounding box of the image block into physical coordinates by setting a scaling factor. The calculation model is as follows: ; in, is the scanning resolution, that is, dots per inch.

[0043] The pixel coordinates of the bounding box of the image patch are scaled by the scaling factor Convert to physical coordinates, the conversion process is as follows ; is the physical coordinate of the image block; is the center coordinate of the image block; w is the width of the image block; h is the height of the image block.

[0044] Project the bounding boxes of all image blocks to the same coordinate system, obtain the coordinate values ​​of each vertex of each image block, and establish the coordinate index of each image block based on the coordinate values ​​of each vertex of each image block, including: A coordinate index (including the page number, the horizontal and vertical coordinates of the upper left corner, and the horizontal and vertical coordinates of the lower right corner) is established for each image block, and the coordinate index is associated with the original page layout information to obtain the metadata of each image block.

[0045] Metadata is stored in JSON-LD format, and the types of image blocks (such as signature areas and clause tables) are annotated with semantic tags, supporting the rapid retrieval and association of subsequent multimodal features.

[0046] S12 feature extraction, constructing a multimodal collaborative feature extraction model, which includes a visual feature extraction sub-model for extracting visual features of image blocks, a topological structure feature extraction sub-model for extracting topological features between image blocks, and a text semantic understanding sub-model for extracting text features of image blocks.

[0047] First, the visual feature extraction sub-model gradually extracts visual features from low-level to high-level in the image block. The visual feature extraction sub-model includes: CSPDarknet53 network integrated with CBAM layer and Bidirectional Feature Pyramid Network (BiFPN). The process of extracting visual features using the visual feature extraction sub-model is as follows: The visual feature extraction sub-model uses the CSPDarknet53 network to perform five-level downsampling of image blocks. At each level, computational redundancy is reduced through cross-stage partial connections (CSP) to output processing results of different scales (gradually reducing the resolution of image blocks from 640×640 to 20×20).

[0048] The visual feature extraction sub-model uses a bidirectional feature pyramid network (BiFPN) to fuse features of different scales and enhances the perception of small objects (such as seals) and large-scale structures (such as tables) through a weighted addition mechanism.

[0049] The visual feature extraction sub-model embeds a spatial-channel dual attention module (CBAM) in the feature map, dynamically focuses on salient areas (such as handwritten annotations or bold text), and outputs visual features.

[0050] Secondly, a topological structure is constructed based on the spatial relationship of the image blocks, and a topological structure feature extraction sub-model is used to learn the topological features of the image blocks. The topological features can reflect the layout and connection mode of the image blocks.

[0051] The topology feature extraction sub-model includes an edge-enhanced graph convolutional network. The process of extracting topology features by the topology feature extraction sub-model is as follows: Image blocks are used as nodes, and the features of the nodes include: the center coordinates, width, height and visual features of the image blocks. The Euclidean distance between image blocks is calculated based on the center coordinates, and edges are constructed between image blocks whose Euclidean distance is less than a preset distance threshold to obtain a dynamic graph structure. Edge-enhanced graph convolution (EdgeConv) is used to learn the local geometric relationship between image blocks (such as horizontal juxtaposition or vertical hierarchy). Edge-enhanced graph convolution introduces edge feature calculation based on traditional GCN. This embodiment uses two layers of graph convolution and global mean pooling in EdgeConv to process image blocks and output 128-dimensional topological features.

[0052] In this embodiment, the preset distance threshold is set to 15% of the average value of all Euclidean distances.

[0053] In other embodiments, the value of the preset distance threshold may be set according to needs.

[0054] Finally, as a feasible method of this embodiment, if the image block contains text, the text content is first identified using OCR technology, and then the text content is input into the text semantic understanding sub-model to extract text features. The text features can capture the semantic information in the image block, and the semantic information can supplement the visual features and topological features.

[0055] As another achievable method of this embodiment, the text semantic understanding sub-model includes a CRNN sub-model and an MLP sub-model. The process of extracting text features by the text semantic understanding sub-model is as follows: The CRNN sub-model locates and identifies image blocks of text type, which are recorded as text blocks. The geometric parameters of the text blocks are obtained according to the coordinate index of the text blocks, and then the tilt angle is calculated based on the center coordinates and width and height of the text blocks.

[0056] Determine whether the tilt angle is zero. If so, do not process; if not, perform affine transformation on the text block based on the tilt angle to obtain a corrected text block, and update the corrected text block to the text block.

[0057] Normalize the geometric parameters of the text block, including center coordinates, width, height, etc.

[0058] The MLP sub-model maps the normalized geometric parameters to a position vector, obtaining the coordinate index of the text block. This coordinate index is then encoded using the pre-defined BERT model to obtain the original position encoding vector. The position vector is then added to the original position encoding vector to obtain the text feature. This text feature not only includes the global [CLS] feature but also retains the hidden state of each token for subsequent fine-grained alignment with the visual features.

[0059] After being processed by the multimodal collaborative feature extraction model, each image block will obtain three feature vectors: visual features, topological features, and text features.

[0060] For S13 classification, the visual features (512 dimensions), text features (768 dimensions), and topological features (128 dimensions) are mapped to a unified space (512 dimensions) through a deformable convolutional layer to obtain the mapped visual features, text features, and topological features.

[0061] In other embodiments, the deformable convolutional layer introduces a learnable spatial offset to adaptively align feature distributions of different modalities.

[0062] The visual features, topological features and text features that have been mapped and belong to the same image block are spliced ​​to obtain a splicing vector.

[0063] The concatenated vector is input into the pre-built MLP model. After nonlinear activation and Softmax normalization, the gating weights of visual features, topological features, and text features are obtained respectively. The calculation model of the gating weights is as follows: ; in, is the c-th gating weight; and is the weight matrix; is the temperature parameter, and its value range is The closer the temperature parameter value is to 0, the closer the Softmax is to the one-hot distribution. The closer the temperature parameter value is to 10, the closer the Softmax is to the uniform distribution. The function is used to ensure that the sum of the three-dimensional feature weights is 1; is the visual feature after mapping; is the topological feature after mapping; is the text feature after mapping processing, that is, the global [CLS] feature of BERT; || is the feature concatenation operator; ReLU is the activation function; is the concatenated vector, whose dimension is the sum of the dimensions of the three eigenvectors; As the activation function, the ReLU function is selected in this embodiment.

[0064] The product of the visual feature and the corresponding gating weight is used as a new visual feature; the product of the topological feature and the corresponding gating weight is used as a new topological feature; and the product of the text feature and the corresponding gating weight is used as a new text feature.

[0065] The new visual features are used as the query (Q), and the new text features and the new topological features are concatenated as key-value pairs (K / V). Cross-modal interaction is achieved through an 8-head attention mechanism. Each attention head focuses on a specific semantic relationship (such as "title-text" or "chart-caption"). Relative position bias is added during calculation to strengthen the association between spatially adjacent elements and output cross-modal attention features.

[0066] The new visual features, new topological features, and new text features are weightedly fused to obtain fused features (512 dimensions). The calculation model of the fused features is as follows: ; ; ; To fusion features, For new visual features, is the cross-modal attention feature, Indicates splicing the output of 8 heads. represents the output of the i-th head, Represents the vector after projecting the query and key-value pairs into the i-th subspace, is the relative position bias matrix of the i-th head, which represents the bias between the query position and the key-value pair position in the i-th head. The key information of the text feature refers to the maximum value of the text feature (i.e., BERT output) along the sequence dimension; Represents the dimensions of the key and query vectors.

[0067] The fused features are input into the pre-trained classification model and the classification results of the fused features are output.

[0068] Example 2: Reference Figure 2 The difference between this embodiment and embodiment 1 is that the S13 classification also includes: S21 divides the image blocks into sets, traverses the coordinate index of each image block, and integrates the image blocks belonging to the same page into an image block set.

[0069] S22 constructs a mask matrix, obtains the dimension of the attention matrix in the attention mechanism, and constructs a binary mask matrix with the same dimension as the attention matrix. The binary mask matrix can prohibit attention interactions between different page elements, so that the reasoning process conforms to the physical structure of the document.

[0070] The element in the i-th row and j-th column of the binary mask matrix indicates whether the i-th image block and the j-th image block belong to the same page. If so, the value of the element is 1; if not, the value of the element is 0.

[0071] S23 calculates a new weight, obtains the center coordinates of the i-th image block set, and calculates the Manhattan distance between the m-th image block where the center coordinates of the i-th image block set are located and the remaining image blocks in the i-th image block set. The Manhattan distance is calculated as follows: ; in, is the Manhattan distance between the mth image block where the center coordinates of the i-th image block set are located and the remaining n-th image blocks in the i-th image block set; is the coordinate value of the center coordinate of the mth image block; is the coordinate value of the center coordinate of the nth image block.

[0072] An attention bias term is constructed based on the Manhattan distance, and the sum of the attention bias term and the attention score in the attention mechanism is used as the new attention score. The calculation model of the new attention score is as follows: ; in, It is a learnable parameter with a value range of 0.3-0.7; is the maximum value of the Manhattan distance corresponding to the i-th image block set; is the new attention score; is the attention bias term.

[0073] Update the new attention score to the attention matrix to obtain the updated attention matrix.

[0074] S24 calculates a new attention matrix, multiplies the binary mask matrix and the updated attention matrix element by element, obtains a multiplied matrix, and uses the multiplied matrix as the new attention matrix.

[0075] In the S13 classification, when the attention mechanism is used to perform weighted fusion on the visual features, topological features, and text features of the same image block, the new attention matrix calculated in S24 is used to fuse the three feature vectors.

[0076] Preferably, the bounding box regression loss function and the contrast loss function are used to adjust the original loss function (cross entropy) of the classification model, and the classification model is retrained using the adjusted loss function. The adjusted loss function calculation model is as follows: ; in, is the adjusted loss function; is the original loss function The weight of is the bounding box regression loss function The weight of is the contrast loss function The weight of .

[0077] The calculation model of the original loss function is as follows: ; Where N is the number of training samples; E is the number of classification categories; is the one-hot encoding value corresponding to the true label of training sample i; Predict the probability that training sample i belongs to category e for the classification model.

[0078] The calculation model of the bounding box regression loss function is as follows: ; in, are the predicted bounding box coordinates; is the true boundary coordinate; f is the coordinate component; (x, y) is the center point coordinate; w is the width; h is the height.

[0079] The contrast loss function forces the multimodal features of the same sample to be aligned in the latent space. The calculation model is as follows: ; in, It is a preset boundary threshold with a value range of 0.1-2.0, which is used to enhance the consistency between modalities.

[0080] Example 3: This example discloses a smart contract analysis system based on a multimodal feature fusion algorithm, the system comprising: Data acquisition and processing module: This module first extracts the page layout information (such as text box position, table area, signature area) through a PDF parsing library (such as PyPDF2 and PDFMiner), and then divides the PDF contract document into image blocks based on the layout.

[0081] The feature extraction module is connected to the data acquisition and processing module and is used to construct a multimodal collaborative feature extraction model. Each image block is input into the multimodal collaborative feature extraction model to obtain the visual features, topological features and text features of each image block respectively.

[0082] The classification module communicates with the feature extraction module, inputs the visual features, topological features and text features belonging to the same image block into the attention weight network, outputs the weights of the three types of features, performs weighted fusion processing on the three types of features based on their respective weights to obtain fused features, inputs the fused features into the pre-trained classification model, and outputs the classification results of the fused features.

[0083] The classification model can adopt a structure such as CNN and RNN. As long as the classification model is trained with training data, the classification model can classify the fused features. The training data refers to the fused features with training labels corresponding to the historical image blocks after the above processing. The training labels include resident identity cards, household registration books, calculation sheets, etc.

[0084] The above are all preferred embodiments of the present application, and are not intended to limit the scope of protection of the present application. Therefore, any equivalent changes made based on the structure, shape, and principle of the present application should be included in the scope of protection of the present application.

Claims

1. The smart contract analysis method based on multimodal feature fusion algorithm is characterized by: include: Data collection and processing: Obtain PDF contract documents, segment them, and obtain multiple image blocks; Feature extraction: Build a multimodal collaborative feature extraction model and input each image block into the multimodal collaborative feature extraction model to obtain the visual features, topological features, and text features of each image block. Classification: For the visual features, topological features, and text features of the same image block, the attention mechanism is used to perform weighted fusion to obtain fused features. The fused features are input into the pre-trained classification model, and the classification results of the fused features are output.

2. The smart contract analysis method based on the multimodal feature fusion algorithm according to claim 1 is characterized in that: The multimodal collaborative feature extraction model includes a visual feature extraction sub-model for extracting visual features of image blocks, a topological structure feature extraction sub-model for extracting topological features between image blocks, and a text semantic understanding sub-model for extracting text features of image blocks.

3. The smart contract analysis method based on the multimodal feature fusion algorithm according to claim 2 is characterized in that: The visual feature extraction sub-model includes a CSPDarknet53 network integrated with a CBAM layer and a bidirectional feature pyramid network. The CSPDarknet53 network is used to perform five-level downsampling on image blocks to obtain multi-scale feature maps. The bidirectional feature pyramid network is used to enhance the target area in the multi-scale feature map and output the enhanced multi-scale feature map. The enhanced multi-scale feature map is weighted by the CBAM layer and then output as visual features.

4. The smart contract analysis method based on the multimodal feature fusion algorithm according to claim 2 is characterized in that: The steps of data collection and processing include: Parse the PDF contract document to obtain a page pixel map. Use the OpenCV contour detection algorithm to separate the image blocks in the page pixel map. Project the bounding boxes of all image blocks into the same coordinate system to obtain the coordinate values ​​of each vertex of each image block. Based on the coordinate values ​​of each vertex of each image block, a coordinate index of each image block is established.

5. The smart contract analysis method based on the multimodal feature fusion algorithm according to claim 4 is characterized in that: The text semantic understanding sub-model includes the CRNN sub-model and the MLP sub-model. The process of extracting text features by the text semantic understanding sub-model is as follows: The CRNN sub-model locates and identifies image blocks of text type, which are recorded as text blocks. The geometric parameters of the text blocks are obtained, including the center coordinates and width and height. The tilt angle of the text blocks is calculated based on the geometric parameters. Determine whether the tilt angle is zero. If so, do not process; if not, perform affine transformation on the text block based on the tilt angle to obtain a corrected text block, and update the corrected text block as the text block; The geometric parameters of the text block are normalized, and the MLP sub-model is used to map the normalized geometric parameters into a position vector. The coordinate index of the text block is obtained, and the coordinate index is encoded to obtain the original position encoding vector. The position vector is added to the original position encoding vector to obtain the text feature.

6. The smart contract analysis method based on the multimodal feature fusion algorithm according to claim 4 is characterized in that: The classification steps also include: According to the coordinate index of each image block, the image blocks belonging to the same page are integrated into an image block set, the dimension of the attention matrix in the attention mechanism is obtained, and a binary mask matrix with the same dimension as the attention matrix is ​​constructed. The element in the i-th row and j-th column of the binary mask matrix indicates whether the i-th image block and the j-th image block belong to the same page. If so, the value of the element is 1; otherwise, the value of the element is 0; Multiply the binary mask matrix and the attention matrix element-wise to obtain the multiplied matrix, and use the multiplied matrix as the new attention matrix.

7. The smart contract analysis method based on the multimodal feature fusion algorithm according to claim 6 is characterized in that: The classification steps also include: Get the center coordinates of the page corresponding to the i-th image block set, calculate the Manhattan distance between the image block where the center coordinates of the page corresponding to the i-th image block set are located and the remaining image blocks in the i-th image block set, construct an attention bias term based on the Manhattan distance, and take the sum of the attention bias term and the attention score in the attention mechanism as the new attention score.

8. The smart contract analysis method based on the multimodal feature fusion algorithm according to claim 4 is characterized in that: The topology feature extraction sub-model includes an edge-enhanced graph convolutional network. The process of extracting topological features by the topology feature extraction sub-model is as follows: Image blocks are used as nodes, and the features of the nodes include: the center coordinates, width, height and visual features of the image blocks. The Euclidean distance between image blocks is calculated based on the center coordinates, and edges are constructed between image blocks whose Euclidean distance is less than a preset distance threshold to obtain a dynamic graph structure. The edge-enhanced graph convolutional network is used to capture the topological features in the dynamic graph structure.

9. The smart contract analysis method based on the multimodal feature fusion algorithm according to claim 1 is characterized in that: The classification steps also include: The visual features, topological features, and textual features belonging to the same image block are aligned and spliced ​​to obtain a splicing vector. The splicing vector is input into a pre-built MLP model to obtain the gating weights of the visual features, topological features, and textual features respectively. The product of the visual feature and the corresponding gating weight is used as a new visual feature; the product of the topological feature and the corresponding gating weight is used as a new topological feature; and the product of the text feature and the corresponding gating weight is used as a new text feature.

10. An intelligent contract analysis system based on a multimodal feature fusion algorithm, the system being used to implement the method according to any one of claims 1 to 9, characterized in that: include: The data acquisition and processing module is used to obtain the PDF contract document, segment the PDF contract document, and obtain multiple image blocks; The feature extraction module is used to build a multimodal collaborative feature extraction model, input each image block into the multimodal collaborative feature extraction model, and obtain the visual features, topological features, and text features of each image block respectively; The classification module is used to perform weighted fusion of the visual features, topological features and text features of the same image block using the attention mechanism to obtain fused features, input the fused features into the pre-trained classification model, and output the classification results of the fused features.

Citation Information

Patent Citations

  • Contract information extraction method and equipment based on multi-modal large language model

    CN118734032A

  • Rumor detection method based on graph attention network

    CN117112786A

  • Image-text content matching generation system and method and storage medium

    CN120220162A

  • Credible traceability presentation file generation method and system based on document content

    CN120493894A

  • Blind sidewalk detection method and system based on machine vision

    CN120544024A

Cited By

  • Data processing method for specimen digital management platform

    CN121053659A

  • Intelligent cherry irrigation decision-making method based on multi-mode and space-time prediction

    CN121525882A

  • Long document intelligent retrieval method and system based on hierarchical analysis and multi-modal fusion

    CN121579677A

  • A long document intelligent retrieval method and system based on hierarchical analysis and multi-modal fusion

    CN121579677B

  • Method, device, processor and readable storage medium for realizing high-precision identification of small inclined character labels for electrical cabinet pressing plate

    CN121600517A