A contract information compliance verification method and device and related equipment

CN121330710BActive Publication Date: 2026-09-22AISINO CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511483535.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-16
Publication Date
2026-09-22
Estimated Expiration
2045-10-16

AI Technical Summary

Technical Problem

但现有技术大多停留在对单页合同或单个印章的逐页单点识别阶段,不能较好的将整份合同作为一个整体进行处理,缺乏对合同页面的智能分类与筛选机制,导致计算资源浪费在大量不含关键信息的页面上

Benefits of technology

[0008]本申请提供了一种合同信息的合规性校验方法、装置及相关设备,包括:接收合同文档,并将所述合同文档的每一页转化为标准化的彩色图像序列;利用预设的多模态大模型对所述彩色图像序列中的每一张图像进行内容分类,为每一张图像分配预设的分类标签,所述分类标签用于标识图像是否包含企业主体信息或印章信息; 根据所述分类标签,对包含企业主体信息或印章信息的图像,采用多条预设的信息提取链路中的一条或多条,提取甲方企业名称、乙方企业名称以及印章文本;其中,所述多条预设的信息提取链路配置有不同的信息处理策略,用于根据预设的触发条件自适应地选择最优的提取路径,以兼顾信息提取的效率与准确性;对比所述甲方、乙方企业名称与所述印章文本,判别合同签订主体与印章主体是否一致;对所述合同文档的页面进行跨页关联分析,以检测骑缝章的连续性与合规性;汇总主体一致性判别的结果和骑缝章合规校验的结果,生成最终的合同合规校验报告。通过这种方式,本方案构筑了从图像预处理、分类、信息提取到一致性比对及骑缝章校验的端到端自动化流程,显著提升了合同印章核验的全面性。其利用大模型进行智能分类和信息提取,并结合传统CV方法进行兜底,兼顾了处理效率与识别精度,大幅降低了人工核验的成本与风险。高效实现了对多页合同骑缝章的全自动合规检测,填补了现有技术的空白,有效防范因骑缝章问题引发的合同篡改风险,切实保障了多页合同的完整性与交易安全。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121330710B_ABST
    Figure CN121330710B_ABST
Patent Text Reader

Abstract

The application provides a contract information compliance verification method and device and related equipment, which converts the received contract document into a standardized color image sequence; uses a multi-modal large model to classify the content of each image in the color image sequence to assign a preset classification label, and then uses one or more of a plurality of preset information extraction links to extract the name of the first party, the name of the second party and the seal text according to the classification label of the image containing the enterprise subject information or the seal information; compare the names of the first and second parties with the seal text to determine whether the contract signing subject and the seal subject are consistent; perform cross-page association analysis on the pages to detect the continuity and compliance of the saddle seal; and generate a final contract compliance verification report by summarizing the subject consistency determination and saddle seal compliance verification results. The process is efficient and accurate, which not only guarantees the integrity and transaction safety of multi-page contracts, but also improves the comprehensiveness of contract seal verification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of computer vision and natural language processing technology, and in particular to a method, apparatus and related equipment for verifying the compliance of contract information. Background Technology

[0002] In business activities such as government and enterprise, finance, and supply chain management, contracts are legal documents that protect the rights and interests of all parties. High-frequency contract scenarios often involve contract documents with numerous pages (up to dozens of pages). In traditional contract review processes, legal or risk control personnel need to invest a significant amount of time manually verifying page by page whether the names of the parties (Party A and Party B) match the names engraved on the seal, and carefully checking whether the seal used to ensure contract integrity is continuous, clear, and unreplaced across all pages. This manual process is not only time-consuming and labor-intensive, but also highly susceptible to omissions and errors due to visual fatigue or negligence, creating potential legal risks and economic losses. To address this issue, the industry has proposed some automated solutions. However, most existing technologies remain at the stage of single-page, single-point recognition of single contracts or individual seals, failing to treat the entire contract as a whole and lacking intelligent classification and filtering mechanisms for contract pages. This results in wasted computational resources on numerous pages lacking key information. Furthermore, it cannot effectively compare and verify information distributed across different pages, particularly failing to solve the core challenge of cross-page continuity compliance detection of the seal. The verification process is incomplete, resulting in low efficiency and poor robustness of the verification results. Therefore, the industry urgently needs an intelligent solution that can perform a global analysis of the entire contract, efficiently and accurately identify seal information, compare subject consistency, and especially detect the compliance of the seal across the seam. Summary of the Invention

[0003] In view of this, embodiments of this application provide a method, apparatus, and related equipment for verifying the compliance of contract information, so as to at least partially solve the above problems.

[0004] In a first aspect, embodiments of this application provide a method for verifying the compliance of contract information, including: Receive the contract document and convert each page of the contract document into a standardized sequence of color images; The content of each image in the color image sequence is classified using a preset multimodal large model, and a preset classification label is assigned to each image. The classification label is used to identify whether the image contains corporate information or seal information. Based on the classification tags, for images containing corporate information or seal information, one or more of the preset information extraction links are used to extract the name of Party A, the name of Party B, and the seal text; wherein, the preset information extraction links are configured with different information processing strategies to adaptively select the optimal extraction path according to preset triggering conditions, so as to balance the efficiency and accuracy of information extraction. By comparing the company names of Party A and Party B with the seal text, it is determined whether the contract signing entity and the seal entity are consistent; cross-page association analysis is performed on the pages of the contract document to detect the continuity and compliance of the seal across the seam. The results of the consistency judgment of the main body and the compliance verification of the seal across the seam are combined to generate the final contract compliance verification report.

[0005] Secondly, based on the contract information compliance verification method described in the first aspect of this application, embodiments of this application also provide a contract information compliance verification device, comprising: A preprocessing module is used to receive the contract document and convert each page of the contract document into a standardized sequence of color images; The classification module is used to classify the content of each image in the color image sequence using a preset multimodal large model, and assign a preset classification label to each image. The classification label is used to identify whether the image contains corporate information or seal information. The extraction module is used to extract the name of the enterprise (Party A), the name of the enterprise (Party B), and the seal text from an image containing enterprise entity information or seal information based on the classification tags. The multiple preset information extraction links are configured with different information processing strategies to adaptively select the optimal extraction path according to preset triggering conditions, so as to balance the efficiency and accuracy of information extraction. The verification module is used to compare the company names of Party A and Party B with the seal text to determine whether the contract signing entity and the seal entity are consistent; and to perform cross-page association analysis on the pages of the contract document to detect the continuity and compliance of the seal across the seam. The generation module is used to summarize the results of the subject consistency judgment and the compliance verification of the seal across the seam, and generate the final contract compliance verification report.

[0006] Thirdly, embodiments of this application also provide a computer storage medium storing computer-executable instructions, which, when executed, perform any of the contract information compliance verification methods described in the first aspect of embodiments of this application.

[0007] Fourthly, embodiments of this application also provide an electronic device, including: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other through the communication bus; The memory is used to store at least one executable instruction that causes the processor to execute any of the contract information compliance verification methods described in the first aspect of the embodiments of this application.

[0008] This application provides a method, apparatus, and related equipment for verifying the compliance of contract information, comprising: receiving a contract document and converting each page of the contract document into a standardized sequence of color images; classifying the content of each image in the sequence of color images using a preset multimodal large model, and assigning a preset classification label to each image, the classification label being used to identify whether the image contains enterprise entity information or seal information; based on the classification label, for images containing enterprise entity information or seal information, extracting the name of Party A, the name of Party B, and seal text using one or more preset information extraction links; wherein, the multiple preset information extraction links are configured with different information processing strategies to adaptively select the optimal extraction path according to preset triggering conditions, so as to balance the efficiency and accuracy of information extraction; comparing the names of Party A and Party B with the seal text to determine whether the contract signing entity and the seal entity are consistent; performing cross-page association analysis on the pages of the contract document to detect the continuity and compliance of the seal across the pages; summarizing the results of the entity consistency determination and the results of the seal compliance verification to generate a final contract compliance verification report. In this way, this solution constructs an end-to-end automated process from image preprocessing, classification, information extraction to consistency comparison and seal verification, significantly improving the comprehensiveness of contract seal verification. It utilizes a large-scale model for intelligent classification and information extraction, combined with traditional computer vision (CV) methods as a fallback, balancing processing efficiency and recognition accuracy, and greatly reducing the cost and risk of manual verification. It efficiently achieves fully automated compliance detection of seals across multi-page contracts, filling a gap in existing technology, effectively preventing contract tampering risks caused by seal issues, and truly ensuring the integrity and transaction security of multi-page contracts. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings.

[0010] Figure 1A schematic diagram illustrating the workflow of a contract information compliance verification method provided in this application embodiment; Figure 2 A schematic diagram of the structure of a contract information compliance verification device provided in this application embodiment; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0011] To enable those skilled in the art to better understand the technical solutions in the embodiments of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art should fall within the protection scope of the embodiments of this application.

[0012] It should be understood that the steps described in the method embodiments of this application may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this application is not limited in this respect.

[0013] Example 1 This application provides a method for verifying the compliance of contract information, such as... Figure 1 As shown, Figure 1 This paper illustrates a flowchart of a compliance verification method for contract information provided in an embodiment of this application, including: Step S101: Receive the contract document and convert each page of the contract document into a standardized color image sequence. In this embodiment of the application, this step, as the first step in the method implementation, is responsible for standardizing input contract files of varying formats (such as scanned PDFs, JPG / PNG images taken with a mobile phone, and plain electronic PDFs). For example, the system can determine the input type by the file extension. For plain electronic PDFs, text and embedded images can be directly extracted; for scanned documents or photographed images, parallel image processing and preset image preprocessing processes can be initiated to efficiently convert contract documents of different sources and formats (such as PDF scans, photos, etc.) into a standardized color image sequence, providing a clean and consistent technical foundation for subsequent accurate analysis.

[0014] Specifically, in a preferred implementation of this application, the above-mentioned conversion of each page of the contract document into a standardized color image sequence includes: for the received original image of the contract document, creating a grayscale processing stream (G0) for morphological correction and a color retention stream (C0) that retains the original color information; performing skew or perspective correction on the grayscale processing stream (G0) to obtain a corrected grayscale image (G1), and simultaneously applying the geometric transformation method applied to G0 to the color retention stream (C0) to obtain a geometrically transformed color image, ensuring pixel-level correspondence between the two data streams; performing binarization processing on the corrected grayscale image (G1) and extracting the maximum contour to locate and crop the main contract area, and simultaneously applying this cropping operation to the grayscale image (G1) and the color image to remove irrelevant background; performing resolution normalization processing on the cropped grayscale image (G1) and the color image to unify them to a preset target resolution, and finally outputting a standardized color image. In this embodiment, a grayscale image is created for geometric correction, while a color image is used to retain color information. By ensuring that both operations are synchronized, crucial color information for seal color identification is avoided during the correction process. This is an optimized solution that balances geometric accuracy and information integrity, significantly improving the recognition accuracy of subsequent models (such as classification models and OCR models). Correction and cropping eliminate noise and irrelevant information; resolution normalization ensures that the data size and clarity of the input model meet optimal requirements.

[0015] This application embodiment innovatively designs an image processing flow of "dual image parallel processing + three-level preprocessing" through the above-described method, which is illustrated herein by way of example: Dual-image parallel stream creation: This step generates two data streams based on the original input image: one is a grayscale processing stream G0 converted to grayscale, used for subsequent geometric correction; the other is a color retention stream C0 that preserves the original RGB three-channel information, used for color-sensitive tasks such as stamp color analysis. The two data streams are lockstepped through a shared coordinate mapping table to ensure that any geometric transformation on G0 is accurately synchronized to C0, maintaining a pixel-level one-to-one correspondence. Specifically, the process of converting the original input image to the grayscale processing stream G0 can be performed using a preset weighted average conversion formula, as shown in the following equation: Gray=0.299×R+0.587×G+0.114×B, Gray (Grayscale) represents the calculated grayscale value. It is a single numerical value that indicates the brightness of a pixel, typically ranging from 0 (pure black) to 255 (pure white).

[0016] R (Red): Represents the red channel value of a pixel in the original color image.

[0017] G (Green): Represents the green channel value of a pixel in the original color image.

[0018] B (Blue): Represents the blue channel value of a pixel in the original color image.

[0019] In practical applications of the method described in this application, the human eye is most sensitive to green, followed by red, and least sensitive to blue. Therefore, this application embodiment limits the above formula to simulate the brightness perceived by the human eye by weighted summation of the values ​​of the three color channels R, G, and B, thereby obtaining a grayscale image that is more in line with human visual habits. These weights (0.299, 0.587, 0.114) are standard coefficient values ​​obtained from psychological and physiological experiments.

[0020] The process of locking the two data streams through a shared coordinate mapping table can be described as follows: by establishing a shared coordinate mapping table M, any geometric transformation of G0 is recorded as a 3×3 homogeneous matrix H and broadcast atomically to C0, ensuring that the two streams correspond one-to-one at the pixel level. This real-time process can effectively guarantee an error of ≤0.1 pixel, and the process of implementing lockstep does not require excessive data processing resources.

[0021] The first stage of preprocessing is skew / perspective correction, which includes the following steps: Edge and line detection: The Canny edge detection algorithm is used on G0 (for example, the grayscale threshold range is set to a low threshold of 50 and a high threshold of 150) to extract contours, focusing on capturing linear features such as contract borders and text line edges. Then, lines are detected by probabilistic Hough transform. For example, in this transformation process, the minimum line segment length L∈[40,60]pixel and the maximum gap G∈[8,12]pixel can be selected.

[0022] Branch judgment: If the area of ​​the quadrilateral formed by the four longest detected lines exceeds the preset proportion of the total image area (e.g., 60%), then the image is judged to have perspective distortion, and the process proceeds to the perspective correction branch in the subsequent image correction; otherwise, the process proceeds to the rotation correction branch.

[0023] Image Correction: In the perspective branch, the vertices Pᵢ (i=1…4) of the quadrilateral are found through corner detection (e.g., Shi-Tomasi), sorted clockwise, and a matrix Qᵢ that can convert it into a standard rectangle is calculated, maintaining the aspect ratio of the original image or, for example, an A4 ratio of 1:√2. The perspective matrix Hp is solved using RANSAC-least squares, with an interior point threshold of 2 pixels and ≥500 iterations. The warpPerspective function is called to complete the correction, and holes are filled with RGB[255,255,255]. When entering the rotation branch, the tilt angle θ of the longest straight line is calculated. If |θ|>0.5°, affine rotation correction is performed with the center of the input image as the origin, and the blank edges after rotation are also filled with white. Then, all transformation matrices are applied synchronously to the color-preserving stream C0 to ensure the geometric consistency between G0 and C0.

[0024] The second-level preprocessing step involves locating and cropping the main contract area: The corrected grayscale image is binarized using OTSU to obtain a black and white image B1 (black foreground, white background). Then, the contour with the largest area and a rectangularity (R = area / (boundary rectangle width × bounding rectangle height)) closest to 1.0 is identified (the closer R is to 1, the closer the contour's shape is to a perfect rectangle). This contour is then considered the main contract area. The grayscale and color images are simultaneously cropped using the bounding rectangle of this contour, removing irrelevant background elements such as margins.

[0025] The third-level preprocessing step is resolution normalization: the image is uniformly scaled to a preset target resolution, such as 300 DPI. If the input DPI is unknown, it is estimated based on the standard A4 paper size (2480×3508 pixels at 300 DPI), i.e., first detecting known size elements within the page (QR Code, standard font height); the default A4 vertical size is 2480×3508 pixels; then the scaling ratio s = target width (height) / current width (height) is calculated; when 0.33 ≤ s ≤ 3, a bicubic interpolation algorithm can be used to ensure the quality of the scaled image; otherwise, an image quality abnormality is indicated. The grayscale image G2 and the color image C2 are simultaneously resized, and the output C2 is a PNG 24-bit lossless compressed image with a resolution error ≤ 1%. This serves as the final output standardized color image.

[0026] Step S102: Each image in the color image sequence is classified using a pre-defined multimodal large model. A pre-defined classification label is assigned to each image, indicating whether the image contains enterprise information or seal information. This step utilizes a multimodal large model to perform accurate content-level classification of the pre-processed standardized color images, providing a decision-making basis for differentiated processing in subsequent steps. Optionally, in one implementation of this application embodiment, a preset multimodal large model is used to classify the content of each image in the color image sequence and assign a preset classification label to each image, including: constructing a prompt containing preset classification rules, wherein the preset classification rules define at least four preset categories: "containing only the main information of Party A and Party B", "containing only the seal", "containing both the main information of Party A and Party B and the seal", and "containing neither"; inputting the prompt and the image to be classified into the multimodal large model, obtaining and parsing the classification result returned by the large model, and assigning a corresponding classification label to the image. This application embodiment provides an illustrative description of this process: The following prompt, containing preset classification rules, is generated: "You are a contract image analysis expert. Based solely on the image content, determine which category the image belongs to below, and strictly follow the rules to output a single uppercase letter without explanation."

[0027] A = Contains only the main information of Party A and Party B (company name, etc.); S = Seal only; A+S = simultaneously includes the main information of Party A and Party B, as well as their seals; N = None of the above elements.

[0028] Output format example: A” The constructed prompt, along with the image, is fed into the multimodal large model. The first character returned by the model is the classification label. Data routing is performed based on the label: images with the label A or A+S can be directly processed in the subsequent information extraction step (step S103); images with the label S or A+S enter the subsequent stamp recognition and detection step; images with the label N are considered invalid pages and are discarded directly.

[0029] This embodiment of the application performs efficient page filtering and routing for the entire contract at this stage. The aim is to quickly identify which pages contain key information (company name or seal) and filter out invalid pages containing no key information, avoiding subsequent invalid calculations. In this step, each pre-processed image page, along with a carefully designed prompt, is input into a multimodal large model. The model outputs a preset classification label based on the image content. By avoiding subsequent invalid calculations, it narrows the processing scope of complex subsequent analysis tasks (such as information extraction and seal recognition) to a few pages containing only key information, significantly improving the processing efficiency and system throughput of the entire contract. Furthermore, this step clarifies how to interact with the large model. It defines the specific content of the prompt (four classification rules) and the subsequent data routing logic. This ensures that the implementation path of the classification step is unique and clear, guaranteeing the determinism and efficiency of the system process.

[0030] Step S103: Based on the classification labels, for images containing enterprise entity information or seal information, one or more of the preset information extraction links are used to extract the name of Party A, the name of Party B, and the seal text. These preset information extraction links are configured with different information processing strategies to adaptively select the optimal extraction path based on preset trigger conditions, balancing efficiency and accuracy. For example, one link can be based on a large language model with a Transformer architecture, while another can be based on a combination of a CNN (Convolutional Neural Network) detection network and an RNN (Recurrent Neural Network) recognition network (such as traditional OCR+NER). This enables the system to handle diverse scenarios. Different strategies have their advantages and disadvantages for different types of data; the coexistence of multiple strategies means greater system adaptability. The adaptive selection of the optimal extraction path based on preset trigger conditions defines the switching mechanism between the multiple preset information extraction links. The system does not randomly select links but follows a clear decision-making logic, achieving intelligent and optimized system operation.

[0031] This application embodiment limits the use of multiple preset information extraction links to target information with different attribute tags, so that the system no longer relies on a single technical path, but can intelligently select the most suitable "tool" to complete the task according to the specific situation of the data to be processed.

[0032] Specifically, in one implementation of this application embodiment, the step of extracting the name of Party A, the name of Party B, and the seal text from an image containing enterprise entity information or seal information based on the classification label includes: configuring a first preset information extraction link, which includes at least a large model end-to-end link, and a second preset information extraction link, which uses an optical character recognition (OCR) + entity recognition (NER) link; employing a dual-track collaborative strategy using the first and second preset information extraction links to extract the name of Party A, the name of Party B, and the seal text; wherein, the first preset information extraction link inputs the image or page text into the large model, directly extracts the full name of Party A, the full name of Party B, and the seal text using zero-shot learning, and outputs a confidence score; when the confidence score is lower than a preset threshold, the second preset information extraction link is triggered to extract information, and the second preset information extraction link adopts a step-by-step processing flow: first, all pixels in the image are converted into text using OCR technology, and then the NER model is used to accurately locate entities in these texts. While this method is cumbersome, it offers greater accuracy and recall when dealing with complex layouts, blurry text, or interference. This embodiment integrates these two approaches into a single system, creating a complementary and collaborative information extraction system. It employs a dual-track collaborative strategy of "large model priority, OCR+NER fallback" to ensure efficient and accurate information extraction. For example, an image (or PDF text) with classification labels A or A+S, along with a zero-shot prompt, is fed into the large model. Through the large model's end-to-end link (first preset link), the prompt instructs the model to extract the full name of the client / contractor company, the seal text, color, and type, returning the results and confidence level in JSON format. For the extraction results across all pages of the entire contract, a "frequency + confidence level" voting mechanism is used to uniquely determine the names of the client and contractor companies. If the highest confidence level is greater than or equal to 0.9, the result is adopted; otherwise, a fallback mechanism is triggered for information extraction, namely, the OCR+NER fallback mechanism (the second preset mechanism) is triggered for information extraction. During this process, the image is first fed into the PaddleOCR engine. Its built-in DBNet detection network and SVTR_LCNet recognition network work together to transcribe all text lines in the image into a character sequence with coordinates. Then, NER entity extraction is performed, which involves concatenating the text recognized by OCR and the text in the PDF in the order they are read. A model such as Chinese-RoBERTa-wwm-ext is used to convert it into a semantic vector, and then the GlobalPointer model is used to accurately mark the start and end positions of the company name.Finally, role matching is performed. The list of company names extracted by NER is matched with regular expressions (such as "Party A", "Party B", "Buyer", "Seller") to determine the names and locations of Party A and Party B companies. This dual-track strategy, which prioritizes the large model and uses OCR+NER as a fallback, first utilizes the zero-shot capability of the large model to directly extract information. If the confidence level returned by the model is lower than the threshold, it automatically switches to the backup path. That is, PaddleOCR is used to perform full-page text recognition first, and then an NER model (such as GlobalPointer) is used to extract company name entities, balancing efficiency and accuracy. The OCR+NER fallback path ensures zero missed detections and high accuracy even on complex, blurry, or specially formatted pages, further ensuring the robustness of information extraction.

[0033] Step S104: Compare the company names of Party A and Party B with the seal text to determine whether the contract signing entity and the seal entity are consistent; perform cross-page association analysis on the contract document pages to detect the continuity and compliance of the seal across the pages. In this embodiment of the application, through entity consistency determination and seal compliance verification, this stage effectively prevents the risk of contract invalidity due to inconsistent entities, while also effectively ensuring the legal validity and security of multi-page contracts.

[0034] Optionally, in one implementation of this application embodiment, comparing the company names of Party A and Party B with the seal text to determine whether the contract signing entity and the seal entity are consistent includes: summarizing all seal text extracted from the entire received contract and performing deduplication and normalization processing; calculating the Levenshtein distance between the overlapping and normalized seal text and the company names of Party A and Party B respectively; if the ratio of the shortest edit distance in the Levenshtein distance to the length of the company name string is less than or equal to a preset similarity threshold, it is determined to be a successful match; otherwise, it is determined to be a mismatch. In this embodiment of the application, this step normalizes the extracted company names of Party A / Party B and the extracted seal text (e.g., removing "Limited Company" and unifying capitalization), and then uses the Levenshtein edit distance algorithm to calculate their similarity. If the ratio of the edit distance to the string length is less than or equal to a preset threshold (e.g., 0.1), it is determined to be consistent. For example, in a scenario, suppose the extracted company name is "AB Intelligent Technology Co., Ltd." and the seal text is "Contract Seal of AB Intelligent Technology Co., Ltd." The system first extracts the subject "AB Intelligent Technology Co., Ltd." from the seal text. Then, it calculates the edit distance between "AB Intelligent Technology Co., Ltd." and "AB Intelligent Technology Co., Ltd." In this example, the distance is 1 (replacing "have" with "share"). The specific discrimination rule can be: divide the calculated shortest edit distance by the length of the original company name to obtain a ratio. If this ratio is less than or equal to a preset threshold (e.g., 0.1), it is considered that the two are successfully matched. If neither the name of Party A nor Party B can match either seal text, the system will trigger a "suspected anomaly" flag. Finally, the matching result (Boolean value), seal color (Boolean value), seal type (Boolean value) are merged with the original information to generate a structured output. The implementation of this step in the embodiment of this application completely automates the work that previously required manual word-by-word comparison, which was prone to errors. It provides an objective and quantitative discrimination standard, greatly improves the efficiency and accuracy of the review, and effectively prevents the risk of contract invalidity due to inconsistencies in the subject.

[0035] Optionally, in one implementation of this application embodiment, based on the classification label, for an image containing enterprise entity information or seal information, one or more of multiple preset information extraction links are used to extract the name of Party A, the name of Party B, and the seal text; wherein, the multiple preset information extraction links are configured with different information processing strategies to adaptively select the optimal extraction path according to preset triggering conditions, so as to balance the efficiency and accuracy of information extraction. Afterwards, the method further includes seal recognition, the recognition process specifically including: In the HSV color space, the image is filtered according to preset thresholds for red hue, saturation, and brightness to generate a rough stamp mask. Then, the stamp mask is refined using the object detection model (YOLO) to accurately locate and segment the stamp image. The segmented seal image is subjected to perspective-cylindrical unfolding transformation to correct the circularly arranged seal text into horizontal text lines, and then the seal text is extracted using optical character recognition (OCR) technology. The color and type of the seal image are determined. The color is determined by statistically analyzing the proportion of red pixels, and the type is determined by an independent image classification model to determine whether it is a "contract seal".

[0036] The recognition process described in this embodiment first uses the HSV color space to quickly filter out highly saturated red pixels by setting thresholds for hue, saturation, and value (e.g., hue H in the range of 0–8 or 175–180, saturation S ≥ 120), generating a rough mask. This mask is then fed into a lightweight YOLO-v8-seal object detection model, outputting a precise stamp bounding box. Further filtering using confidence level and the proportion of red pixels within the box yields clean, interference-free stamp candidate regions. The segmented stamp image is then geometrically transformed to unfold the circularly arranged text into a horizontal rectangular image. This process is called "perspective-cylindrical unfolding." The unfolded image is then fed into the PaddleOCR engine, which extracts the curved text from the stamp just like recognizing ordinary text. In the unfolded stamp image, the proportion of red pixels is statistically analyzed. If the proportion exceeds a threshold (e.g., 70%), it is considered red; otherwise, it is marked as an anomaly. The unfolded image is scaled to 224×224 pixels and fed into a pre-trained ResNet18 binary classification model, which determines whether it belongs to "contract seal" or another type. Finally, the seal text, type, color, bounding box, and confidence score are encapsulated into a JSON object and output. This step in the embodiment of this application supplements the specific technical details of extracting seal information from the image before subject consistency determination. It defines a three-stage process of "segmentation-extraction-determination": precise seal location using HSV+YOLO, recognition of curved text using "perspective-cylindrical unfolding" + OCR, and determination of color and type using pixel statistics and a ResNet classifier. This provides higher-quality and more structured seal data for subsequent comparisons.

[0037] Furthermore, in an optional implementation of this application embodiment, when performing seal recognition and detection, it is not necessary to perform threshold screening in the HSV color space. Instead, color distribution statistics can be performed on the candidate seal regions in the color retainer stream (CO), and the optimal color space (such as HSV, Lab, YCrCb) can be adaptively selected based on the statistical results for red pixel enhancement and extraction. Specifically, the optimal color space can be selected based on the statistical results, which can be a color space that maximizes the inter-class variance between red seal pixels and background pixels in that color space. In practical application scenarios of this application embodiment, a fixed HSV threshold may result in unstable segmentation effects for red seals under different lighting conditions, different printers, or different aging levels. For example, a dark maroon seal may not meet the preset saturation and brightness thresholds in the HSV space. This application embodiment introduces an adaptive color space selection mechanism, which can dynamically switch to the color space that best highlights the difference between the seal and the background based on the specific color characteristics of the current seal. This greatly enhances the seal segmentation algorithm's resistance to interference from external factors such as color changes and uneven lighting, thereby improving the accuracy and universality of seal positioning.

[0038] Optionally, in one implementation of this application embodiment, cross-page association analysis is performed on the pages of the contract document to detect the continuity and compliance of the binding seal. This includes: extracting feature points of adjacent pages in the contract document and matching them, solving the rigid transformation matrix, mapping all pages to a unified coordinate system to align the binding line position, and achieving cross-page image registration; in the registered page sequence, a preset instance segmentation model (Mask R-CNN) is used to detect a preset region near the binding line of each page to extract the mask outline of the binding seal; the geometric consistency, positional compliance, and textual consistency of the binding seal masks extracted from adjacent pages are jointly determined; wherein, geometric consistency is determined by calculating the intersection-over-union (IoU) ratio after mask alignment, positional compliance is determined by calculating the distance from the mask center to the binding line, and textual consistency is determined by splicing the aligned masks, recognizing their text, and then comparing it with the name of the contracting entity company.

[0039] The embodiments of this application employ a three-step method of "registration-segmentation-verification". First, adjacent pages are registered using algorithms such as ORB feature point matching (a fast and efficient corner detection and description algorithm) and RANSAC (Random Sample Consensus), aligning the binding line. Then, the Mask R-CNN instance segmentation model is used to accurately locate the seal fragments in the vicinity of the binding line. Finally, the fragments are aligned, and their geometric continuity (IoU), positional compliance (distance to the binding line), and textual consistency (identification and comparison after splicing) are jointly verified, effectively achieving fully automatic, end-to-end closed-loop verification of the seal. This step effectively ensures the legal validity and security of multi-page contracts. It efficiently and automatically detects whether the seals on multi-page contracts are continuous across pages, compliant in position, and consistent in text, ensuring the integrity of the contract and preventing page replacement or alteration.

[0040] Optionally, in one implementation of this application embodiment, the following three judgment criteria are followed in the process of jointly determining the geometric consistency, positional compliance, and textual consistency of the cross-stamp mask extracted from adjacent pages: When the Intersection over Union (IoU) of the masks on adjacent pages is greater than or equal to a first preset threshold, and the center point offset distance is less than or equal to a second preset threshold, the adjacent pages are determined to be geometrically continuous. This method aligns the detected overlapping seal masks of two adjacent pages (e.g., page n and page n+1) according to the page offset obtained in the registration step. Then, the IoU of these two masks is calculated. IoU is the ratio of the area of ​​the overlapping portion of two regions to their total area after merging. If the IoU is greater than or equal to a threshold (e.g., 0.75), and their center point offset distance is sufficiently small (e.g., ≤ 2mm), then the two seal portions are considered geometrically continuous for geometric consistency determination. When the distance from the center of the binding seal to the binding line is less than or equal to a third preset threshold, the position is deemed compliant. This step determines the position compliance by checking whether the center of the binding seal falls within a compliant area near the binding line (e.g., ≤5mm from the binding line). When the normalized edit distance between the identified seal text and the company name of Party A or Party B is less than or equal to the fourth preset threshold, the text is determined to be consistent. This step stitches the aligned adjacent masks into a complete seal image in virtual space. Then, the seal text extraction method described above in this application embodiment can be called to identify the complete seal text. The edit distance of this text is then compared with the company name of Party A or Party B output by the information extraction module (threshold ≤ 0.1) to accurately determine the text consistency. The paging seal is determined to be compliant if and only if all the above three criteria (the above three consistency determinations) are satisfied. This joint verification method provides more accurate and quantifiable evaluation criteria, so that the final conclusion of "paging seal compliance" is drawn based on a series of strict data indicators, which greatly enhances the reliability and persuasiveness of the result.

[0041] Step S105: Summarize the result of subject consistency determination and the result of paging seal compliance check, and generate a final contract compliance check report. After all pages of the entire contract are processed, all analysis results are output and summarized to generate a document-level contract compliance check report, such as in JSON format. The report includes Boolean fields such as the unique name of the first party / second party enterprise, whether the seal matches, whether the seal color is red, whether it is a special contract seal, whether the paging seal is compliant, as well as the number of processed pages and a timestamp, providing users with a clear, concise and structured final report that is convenient for users to make quick decisions or for system archiving. Users or upper-level management systems can easily conduct contract risk rating, automatic archiving or trigger manual review processes based on this report, realizing a closed loop of the entire verification process.

[0042] Optionally, in an implementation of the embodiment of the present application, before separately calculating the Levenshtein distance between the overlapped and normalized seal text and the enterprise names of the first party and the second party, the method further comprises: performing query and standardization processing on the seal text and the enterprise names of the first party and the second party through a preset enterprise industrial and commercial information knowledge base; wherein the enterprise industrial and commercial information knowledge base contains the standard full name, historical former names, common abbreviations and easily confused typo information of enterprises; if the extracted enterprise name or seal subject can fuzzy match an entity in the knowledge base, the standard full name of the entity is uniformly used for subsequent edit distance comparison. In the actual application scenario of the embodiment of the present application, misjudgment may occur for common abbreviations of enterprise names (e.g., "China Construction Bank" and "CCB"), or individual typos caused by seal wear and blurred recognition (e.g., "technology" is recognized as "techonlogy"). In this stage of the embodiment of the present application, by introducing a knowledge base, the simple string comparison is upgraded to semantic entity comparison based on background knowledge, which can greatly improve the accuracy and robustness of matching, significantly reduce false alarms caused by non-standard names or recognition errors, and enhance the practical value of the solution in real complex scenarios.

[0043] Optionally, in one implementation of this application embodiment, when accurately locating the seal fragment near the binding line using the Mask R-CNN instance segmentation model, the input of the Mask R-CNN model not only includes the binding neighborhood image of the current page, but also stitches together the corresponding neighborhood images of adjacent pages after registration and alignment, forming a cross-page joint image input. This allows the model to utilize the contextual visual information of adjacent pages as a reference when segmenting the seal mask of the current page. In the practical application scenario of this application embodiment, when processing the seal fragment of each page individually, if the seal is severely damaged or interfered with by page stains, the model may have difficulty segmenting accurately. In this stage of this application embodiment, by stitching together the corresponding areas of adjacent pages as input, the model can see the other half of the seal, using this strong spatial continuity prior knowledge to assist in the judgment. This significantly improves the segmentation accuracy and recall rate when the seal is unclear, incompletely covered, or contains interference, and enhances the stability and reliability of the entire seal compliance verification process.

[0044] This application provides a method for verifying the compliance of contract information. The method involves receiving a contract document and converting each page of the document into a standardized sequence of color images. A pre-defined multimodal model is used to classify the content of each image in the sequence, assigning a pre-defined classification label to each image. These labels identify whether the image contains enterprise entity information or seal information. Based on the classification labels, for images containing enterprise entity information or seal information, one or more pre-defined information extraction links are used to extract the name of the contracting party (Party A), the name of the contracting party (Party B), and the seal text. These pre-defined information extraction links are configured with different information processing strategies to adaptively select the optimal extraction path based on pre-defined triggering conditions, balancing efficiency and accuracy. The names of the contracting party (Party A) and the contracting party (Party B) are compared with the seal text to determine whether the contract signing entity and the seal entity are consistent. Cross-page association analysis is performed on the pages of the contract document to detect the continuity and compliance of the seal across the pages. The results of the entity consistency determination and the seal compliance verification are summarized to generate a final contract compliance verification report. In this way, this solution constructs an end-to-end automated process from image preprocessing, classification, information extraction to consistency comparison and seal verification, significantly improving the comprehensiveness of contract seal verification. It utilizes a large-scale model for intelligent classification and information extraction, combined with traditional computer vision (CV) methods as a fallback, balancing processing efficiency and recognition accuracy, and greatly reducing the cost and risk of manual verification. It efficiently achieves fully automated compliance detection of seals across multi-page contracts, filling a gap in existing technology, effectively preventing contract tampering risks caused by seal issues, and truly ensuring the integrity and transaction security of multi-page contracts.

[0045] Example 2 Based on the contract information compliance verification method provided in Embodiment 1 of this application, this embodiment also provides a corresponding contract information compliance verification device, such as... Figure 2 As shown, Figure 2 This application provides a schematic diagram of the structure of a contract information compliance verification device 20, which includes: The preprocessing module 201 is used to receive the contract document and convert each page of the contract document into a standardized sequence of color images; The classification module 202 is used to classify the content of each image in the color image sequence using a preset multimodal large model, and assign a preset classification label to each image. The classification label is used to identify whether the image contains corporate information or seal information. The extraction module 203 is used to extract the name of the enterprise, the name of the supplier, and the seal text from an image containing enterprise entity information or seal information by using one or more of multiple preset information extraction links based on the classification tags. The multiple preset information extraction links are configured with different information processing strategies to adaptively select the optimal extraction path according to preset triggering conditions, so as to balance the efficiency and accuracy of information extraction. The verification module 204 is used to compare the company names of Party A and Party B with the seal text to determine whether the contract signing entity and the seal entity are consistent; and to perform cross-page association analysis on the pages of the contract document to detect the continuity and compliance of the seal across the seam. The generation module 205 is used to summarize the results of the subject consistency judgment and the compliance verification of the seal across the seam, and generate the final contract compliance verification report.

[0046] Optionally, in one implementation of this application embodiment, the preprocessing module 201 is further configured to: create a grayscale processing stream (G0) for morphological correction and a color retention stream (C0) that retains the original color information for the received original image of the contract document; perform skew or perspective correction on the grayscale processing stream (G0) to obtain a corrected grayscale image (G1), and simultaneously apply the geometric transformation method applied to G0 to the color retention stream (C0) to obtain a geometrically transformed color image, ensuring pixel-level correspondence between the two data streams; perform binarization processing on the corrected grayscale image (G1) and extract the maximum contour to locate and crop the main contract area, and simultaneously apply this cropping operation to the grayscale image (G1) and the color image to remove irrelevant background; perform resolution normalization processing on the cropped grayscale image (G1) and the color image to unify them to a preset target resolution, and finally output a standardized color image.

[0047] Optionally, in one implementation of this application embodiment, the classification module 202 is further configured to construct a prompt containing preset classification rules, wherein the preset classification rules define at least four preset categories: "containing only the main information of Party A and Party B", "containing only the seal", "containing both the main information of Party A and Party B and the seal", and "containing neither". The prompt and the image to be classified are input into the multimodal large model, the classification result returned by the large model is obtained and parsed, and the corresponding classification label is assigned to the image.

[0048] Optionally, in one embodiment of this application, the extraction module 203 is further configured to configure the first preset information extraction link, which includes at least a large model end-to-end link, and the second preset information extraction link, which uses an optical character recognition (OCR) + entity recognition (NER) link. The first preset information extraction link and the second preset information extraction link employ a dual-track collaborative strategy to extract the name of Party A, the name of Party B, and the seal text. Specifically, the first preset information extraction link inputs the image or page text into the large model, uses zero-shot learning to directly extract the full name of Party A, the full name of Party B, and the seal text, and outputs a confidence score. When the confidence score is lower than a preset threshold, the second preset information extraction link is triggered to first extract all text in the image using optical character recognition (OCR) technology, and then uses a named entity recognition (NER) model to identify and extract the enterprise name entity from all the text.

[0049] Optionally, in one embodiment of this application, the device 20 further retains an identification module (not shown in the figures). This identification module is used to extract the name of the enterprise (Party A), the name of the enterprise (Party B), and the seal text from an image containing enterprise entity information or seal information based on the classification label, using one or more of multiple preset information extraction links. The multiple preset information extraction links are configured with different information processing strategies to adaptively select the optimal extraction path based on preset triggering conditions, balancing the efficiency and accuracy of information extraction. Then, the seal is identified. The specific identification process is as follows: in HSV color... In the color space, the image is filtered according to preset thresholds for red hue, saturation, and brightness to generate a rough seal mask. Then, the seal mask is refined using a target detection model (YOLO) to accurately locate and segment the seal image. The segmented seal image is then subjected to a perspective-cylindrical unfolding transformation to correct the circularly arranged seal text into horizontal text lines. The seal text is then extracted using optical character recognition (OCR) technology. The color and type of the seal image are determined, with the color determined by statistically analyzing the proportion of red pixels, and the type determined by an independent image classification model to identify whether it is a "contract seal".

[0050] Furthermore, in an optional implementation of this application embodiment, when performing seal recognition and detection, it is not necessary to perform threshold screening in the HSV color space. Instead, color distribution statistics can be performed on the candidate seal areas in the color retention stream (C0) first, and then the optimal color space (such as HSV, Lab, YCrCb) can be adaptively selected based on the statistical results to enhance and extract red pixels. The selection criteria are specifically based on maximizing the inter-class variance between red seal pixels and background pixels in that color space.

[0051] Optionally, in one embodiment of this application, the verification module 204 is further configured to: summarize all the seal text extracted from the received entire contract, and perform deduplication and normalization processing; calculate the Levenshtein distance between the overlapping and normalized seal text and the names of Party A and Party B respectively; if the ratio of the shortest edit distance in the Levenshtein distance to the length of the enterprise name string is less than or equal to a preset similarity threshold, it is determined to be a successful match; otherwise, it is determined to be a non-match.

[0052] Optionally, in one embodiment of this application, the verification module 204 is further configured to extract feature points of adjacent pages in the contract document and perform matching, solve the rigid transformation matrix, map all pages to the same coordinate system to align the binding line position and achieve cross-page image registration; in the registered page sequence, a preset instance segmentation model (Mask R-CNN) is used to detect a preset area near the binding line of each page to extract the mask contour of the binding seal; the geometric consistency, positional compliance and textual consistency of the binding seal mask extracted from adjacent pages are jointly determined; Geometric consistency is determined by calculating the intersection-union ratio (IoU) after mask alignment; positional compliance is determined by calculating the distance from the mask center to the binding line; and textual consistency is determined by splicing and aligning the mask, recognizing its text, and then comparing it with the name of the contracting entity.

[0053] Optionally, in one embodiment of this application, the verification module 204 follows the following three judgment criteria when jointly determining the geometric consistency, positional compliance, and textual consistency of the cross-stamp masks extracted from adjacent pages: When the intersection-over-union ratio (IoU) of the masks of adjacent pages is greater than or equal to a first preset threshold, and the center point offset distance is less than or equal to a second preset threshold, the adjacent pages are determined to be geometrically continuous. When the distance from the center of the binding seal to the binding line is less than or equal to the third preset threshold, the position is deemed compliant. When the text of the seal identified by splicing is less than or equal to the normalized edit distance of the name of Party A or Party B, the text is determined to be consistent. The seal across the seam is deemed compliant if and only if all three criteria are met.

[0054] Optionally, in one implementation of this application embodiment, the device 20 further includes a matching module (not shown in the figures). This matching module is used to query and standardize the seal text and the names of the A and B companies through a pre-set enterprise business information knowledge base before calculating the Levenshtein distance between the overlapping and normalized seal text and the names of the A and B companies respectively. The enterprise business information knowledge base contains the standard full name of the company, its historical former names, common abbreviations, and easily confused misspellings. If the extracted company name or seal body can be fuzzily matched to an entity in the knowledge base, the standard full name of the entity is used uniformly for subsequent edit distance comparison.

[0055] Optionally, in one implementation of this application embodiment, when the verification module 204 uses the Mask R-CNN instance segmentation model to accurately locate the seal fragment near the binding line, the input of the Mask R-CNN model not only includes the binding neighborhood image of the current page, but also stitches together the corresponding neighborhood images of its adjacent pages after registration and alignment, forming a cross-page joint image input. This allows the model to utilize the contextual visual information of adjacent pages as a reference when segmenting the seal mask of the current page. In the practical application scenario of this application embodiment, when processing the seal fragment of each page individually, if the seal is severely damaged or interfered with by page stains, the model may have difficulty accurately segmenting it.

[0056] Example 3 This application also provides a storage medium storing a computer program thereon, which, when executed by a processor, implements any of the contract information compliance verification methods described in the foregoing embodiment one of this application. Example 4 This application also provides an electronic device, such as... Figure 3 As shown, Figure 3 This application provides a schematic diagram of the structure of an electronic device 30, which includes: One or more processors 301, communication interface 302, memory 303 and communication bus 304, the processors 301, memory 303 and communication interface 302 communicate with each other through communication bus 304; Memory 303 is used to store one or more programs; When the one or more programs are executed by the one or more processors 301, the one or more processors 301 implement any of the contract information compliance verification methods described in Embodiment 1 of this application.

[0057] This application has now described specific embodiments of the subject matter. In some cases, the actions described in the claims can be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing can be advantageous.

[0058] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system layer onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.

[0059] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0060] The system layers, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0061] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.

[0062] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0063] Those skilled in the art will understand that embodiments of this application can be provided as methods, system-level, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0064] This application can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific transactions or implement specific abstract data types. This application can also be practiced in distributed computing environments where transactions are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0065] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system-level embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0066] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for verifying the compliance of contract information, characterized in that, include: The system receives a contract document and converts each page of the contract document into a standardized sequence of color images. Specifically, for the received original image of the contract document, a grayscale processing stream (G0) for morphological correction and a color retention stream (C0) that preserves the original color information are created. The grayscale processing stream (G0) undergoes skew or perspective correction to obtain a corrected grayscale image (G1). The geometric transformation applied to the grayscale processing stream (G0) is simultaneously applied to the color retention stream (C0) to obtain a geometrically transformed color image, ensuring pixel-level correspondence between the two data streams. The corrected grayscale image (G1) is binarized, and the maximum contour is extracted to locate and crop the main contract area. This cropping operation is simultaneously applied to both the grayscale image (G1) and the color image to remove irrelevant background. The cropped grayscale image (G1) and the color image are then normalized to a preset target resolution, ultimately outputting a standardized color image. The content of each image in the color image sequence is classified using a preset multimodal large model, and a preset classification label is assigned to each image. The classification label is used to identify whether the image contains corporate information or seal information. Based on the classification tags, for images containing corporate information or seal information, one or more of the preset information extraction links are used to extract the name of Party A, the name of Party B, and the seal text; wherein, the preset information extraction links are configured with different information processing strategies to adaptively select the optimal extraction path according to preset triggering conditions. The contract signing entity and the seal entity are compared with the names of the parties A and B to determine whether they are consistent. Cross-page association analysis is performed on the pages of the contract document to detect the continuity and compliance of the seal across the binding. Specifically, this step involves: extracting feature points from adjacent pages in the contract document and matching them; solving the rigid transformation matrix; mapping all pages to a unified coordinate system to align the binding line and achieve cross-page image registration; and then, in the registered page sequence, using a pre-defined instance segmentation model (Mask)... R-CNN detects a preset area near the binding line of each page to extract the mask outline of the seal. It then performs a joint judgment on the geometric consistency, positional compliance, and textual consistency of the extracted seal masks from adjacent pages. Specifically, if the intersection-over-union (IoU) ratio of the masks on adjacent pages is greater than or equal to a first preset threshold, and the center point offset distance is less than or equal to a second preset threshold, the adjacent pages are considered geometrically continuous. If the distance from the center of the seal to the binding line is less than or equal to a third preset threshold, the position is considered compliant. If the normalized edit distance between the identified seal text and the company name (either the client's or contractor's name) is less than or equal to a fourth preset threshold, the text is considered consistent. When all three criteria are met, the seal is deemed compliant. The results of the consistency judgment of the main body and the compliance verification of the seal across the seam are summarized to generate the final contract compliance verification report.

2. The method for verifying the compliance of contract information according to claim 1, characterized in that, The step of classifying the content of each image in the color image sequence using a preset multimodal large model and assigning a preset classification label to each image includes: Construct a prompt containing preset classification rules, wherein the preset classification rules define at least four preset categories: "containing only the main information of Party A and Party B", "containing only the seal", "containing both the main information of Party A and Party B and the seal", and "containing neither". The prompts and the image to be classified are input into the multimodal large model, the classification results returned by the large model are obtained and parsed, and the corresponding classification labels are assigned to the image.

3. The method for verifying the compliance of contract information according to claim 1, characterized in that, The step involves extracting the company name, company name, and seal text from images containing enterprise entity information or seal information based on the classification tags, using one or more of a set of preset information extraction links. This includes: The configuration includes at least a first preset information extraction link that includes a large model end-to-end link, and a second preset information extraction link that uses an optical character recognition (OCR) + entity recognition (NER) link. A dual-track collaborative strategy is adopted based on the first preset information extraction link and the second preset information extraction link to extract the name of Party A, the name of Party B, and the seal text; The first preset information extraction link inputs the image or page text into the large model, uses zero-shot learning to directly extract the full name of Party A, the full name of Party B, and the seal text, and outputs the confidence level; when the confidence level is lower than a preset threshold, the second preset information extraction link is triggered to first extract all the text in the image through optical character recognition (OCR) technology, and then use the named entity recognition (NER) model to identify and extract the enterprise name entity from all the text.

4. The method for verifying the compliance of contract information according to claim 1 or 3, characterized in that, Based on the classification tags, for images containing enterprise entity information or seal information, one or more of multiple preset information extraction links are used to extract the name of Party A, the name of Party B, and the seal text. These multiple preset information extraction links are configured with different information processing strategies to adaptively select the optimal extraction path based on preset triggering conditions, balancing efficiency and accuracy in information extraction. Subsequently, the method further includes seal recognition, specifically including: In the HSV color space, the image is filtered according to preset thresholds for red hue, saturation, and brightness to generate a rough stamp mask. Then, the stamp mask is refined using the object detection model (YOLO) to accurately locate and segment the stamp image. The segmented seal image is subjected to perspective-cylindrical unfolding transformation to correct the circularly arranged seal text into horizontal text lines, and then the seal text is extracted using optical character recognition (OCR) technology. The color and type of the seal image are determined. The color is determined by statistically analyzing the proportion of red pixels, while the type is determined by an independent image classification model to identify whether it is a contract seal.

5. The method for verifying the compliance of contract information according to claim 1, characterized in that, The comparison of the company names of Party A and Party B with the seal text to determine whether the contract signing entity and the seal entity are consistent includes: Summarize all the seal texts extracted from the entire received contract, and perform deduplication and normalization processing. Calculate the Levenshtein distance between the normalized seal text and the names of Party A and Party B, respectively. If the ratio of the shortest edit distance in the Levenstein distance to the length of the company name string is less than or equal to a preset similarity threshold, it is considered a successful match; otherwise, it is considered a non-match.

6. A device for verifying the compliance of contract information, characterized in that, include: A preprocessing module receives contract documents and converts each page of the contract documents into a standardized sequence of color images. Specifically, for the received original images of the contract documents, a grayscale processing stream (G0) for morphological correction and a color retention stream (C0) that preserves the original color information are created. The grayscale processing stream (G0) undergoes skew or perspective correction to obtain a corrected grayscale image (G1). The geometric transformation applied to the grayscale processing stream (G0) is simultaneously applied to the color retention stream (C0) to obtain a geometrically transformed color image, ensuring pixel-level correspondence between the two data streams. The corrected grayscale image (G1) is binarized, and the maximum contour is extracted to locate and crop the main contract area. This cropping operation is simultaneously applied to both the grayscale image (G1) and the color image to remove irrelevant background. The cropped grayscale image (G1) and the color image are then normalized to a preset target resolution, ultimately outputting a standardized color image. The classification module is used to classify the content of each image in the color image sequence using a preset multimodal large model, and assign a preset classification label to each image. The classification label is used to identify whether the image contains corporate information or seal information. The extraction module is used to extract the name of the enterprise (Party A), the name of the enterprise (Party B), and the seal text from an image containing enterprise entity information or seal information based on the classification tags. The multiple preset information extraction links are configured with different information processing strategies to adaptively select the optimal extraction path according to preset triggering conditions, so as to balance the efficiency and accuracy of information extraction. The verification module compares the company names of Party A and Party B with the seal text to determine whether the contract signing entity and the seal entity are consistent; it performs cross-page association analysis on the pages of the contract document to detect the continuity and compliance of the binding seal; specifically, it extracts feature points from adjacent pages in the contract document and matches them, solves the rigid transformation matrix, maps all pages to a unified coordinate system to align the binding line position, and achieves cross-page image registration; in the registered page sequence, it uses a preset instance segmentation model (Mask) R-CNN detects a preset area near the binding line of each page to extract the mask outline of the seal. It then performs a joint judgment on the geometric consistency, positional compliance, and textual consistency of the extracted seal masks from adjacent pages. Specifically, if the intersection-over-union (IoU) ratio of the masks on adjacent pages is greater than or equal to a first preset threshold, and the center point offset distance is less than or equal to a second preset threshold, the adjacent pages are considered geometrically continuous. If the distance from the center of the seal to the binding line is less than or equal to a third preset threshold, the position is considered compliant. If the normalized edit distance between the identified seal text and the company name (either the client's or contractor's name) is less than or equal to a fourth preset threshold, the text is considered consistent. When all three criteria are met, the seal is deemed compliant. The generation module is used to summarize the results of the subject consistency judgment and the compliance verification of the seal across the seam, and generate the final contract compliance verification report.

7. A computer storage medium, characterized in that, The computer storage medium stores computer-executable instructions, which, when executed, perform the contract information compliance verification method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Seal identification method fusing multiple features

    CN115661850A

  • Seal management method and platform, electronic equipment and storage medium

    CN117522094A