Document analysis method for qualification review and similarity comparison

By combining multimodal fusion technology and a rule engine, the problems of insufficient generalization and human-machine collaboration in the bidding process have been solved, realizing intelligent parsing and review of bid documents and improving the efficiency and accuracy of the bidding process.

CN121860733APending Publication Date: 2026-04-14THREE GORGES SMART WATER TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-10
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing bid clearing technologies lack generalizability and human-machine collaboration capabilities in the bidding and tendering field, and are unable to effectively identify and extract key information, resulting in low efficiency, poor accuracy, and insufficient information transparency.

Method used

It employs multimodal fusion technologies of OCV, CV, and NLP to perform document parsing and semantic association, constructs a multi-dimensional similarity assessment system, combines a rule engine to achieve dynamic risk judgment, supports user-defined review rules, and performs automated parsing and review through human-machine collaboration.

Benefits of technology

It has enabled automated, intelligent, and in-depth analysis and review of tender documents, improving the efficiency and accuracy of bid clearing, reducing the workload of manual review, and enhancing information transparency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121860733A_ABST
    Figure CN121860733A_ABST
Patent Text Reader

Abstract

The invention provides a document analysis method for qualification review and similarity comparison, and relates to the field of document intelligent analysis, the method adopts an OCR, CV and NLP multi-modal fusion technology to realize bidding document deep analysis, fuses text, layout and image features through an attention mechanism, and converts unstructured data into a structured semantic map. And fusing text, image and quotation multi-dimensional similarities, and dynamically generating a comprehensive risk score through a learnable function. And qualification review adopts a rule engine and NLP double-engine cooperation, a rigid rule guarantees a review base line, flexible semantic reasoning processes fuzzy conditions, and NLP weights are dynamically adjusted according to historical behaviors of an enterprise. Manual auditing feedback is used as continuous learning data, and the model and the rule base are optimized. According to the method, full-process automation of bidding clearing is realized, the review efficiency and accuracy are improved, dynamic configuration of rules is supported, the risk of range bidding is effectively prevented, the bidding clearing cost is reduced, and continuous evolution of the system is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent document analysis, specifically to a document analysis method for qualification review and similarity comparison. Background Technology

[0002] Bidding and tendering, as an important mechanism for resource allocation in a market economy, is widely used in government projects, state-owned enterprise procurement, and infrastructure construction. In the entire bidding and procurement process, bid evaluation, a crucial step before bid assessment, faces multiple challenges due to the traditional manual operation method, including efficiency bottlenecks, poor accuracy, insufficient information transparency, and high costs. To address these issues in the manual bid evaluation process, there is an urgent need to introduce advanced technologies to achieve an intelligent upgrade of the bid evaluation process.

[0003] To address the shortcomings of traditional manual bid review methods, the industry has explored various technological approaches for improvement, but significant limitations remain. Early solutions primarily focused on electronic storage and retrieval, such as PDF document management systems and OCR scanning and archiving systems. While these systems solved the document digitization problem, they lacked deep content understanding capabilities and could not automatically identify and extract key information. Some research attempted to apply NLP and deep learning technologies to improve bid review efficiency; however, these methods are mostly limited to single tasks and lack in-depth optimization tailored to the characteristics of the bidding and tendering field. This results in insufficient model generalization ability, poor performance when faced with bid documents containing dense technical terminology and complex formats, and a lack of human-machine collaboration mechanisms, failing to effectively balance automation and human judgment.

[0004] To address the shortcomings of existing technologies, this invention provides a document analysis method for qualification review and similarity comparison. It employs multimodal fusion technologies including OCV, Computer Vision (CV), and Natural Language Processing (NLP) to achieve unified parsing and semantic association of unstructured data such as text, tables, and images. By integrating text vector similarity, image hash similarity, and price quotation statistical similarity, a multi-dimensional similarity assessment system is constructed, and dynamic risk assessment is achieved through a rule engine. A visual rule configuration interface is provided, allowing users to dynamically define and adjust review rules. Rules are synchronized to the engine in real time, ultimately achieving a fully automated closed-loop process from document parsing, qualification review, similarity comparison to risk warning, replacing manual item-by-item verification. Summary of the Invention

[0005] The technical problem to be solved by this invention is to provide a document analysis method for qualification review and similarity comparison, which solves the problems of insufficient generalization and human-computer collaboration capabilities of existing bid clearing and auxiliary tools, and realizes automated, intelligent, and in-depth analysis and review of bid documents.

[0006] The technical solution adopted in this invention is to provide a document analysis method for qualification review and similarity comparison. The system receives multi-format bid documents uploaded by users, performs preprocessing operations such as virus scanning, format verification, and file segmentation, and uses technologies such as OCR, CV, and NLP to comprehensively parse the bid documents, extracting multimodal content such as text, tables, and images, and storing it in a structured manner. Based on the bid clearing requirements, the system distributes the parsed data to the qualification intelligent review module to verify the bidder's qualification compliance, the bid document similarity comparison module to detect possible plagiarism or collusion, and the credit risk assessment module to evaluate the bidder's performance capability and risk. Each of the three analysis modules generates professional results, which are then integrated by the comprehensive analysis report generation module to form a structured bid clearing report. The system presents the analysis results and evidence chain to human reviewers, who make the final judgment. The results and corrections from the human review serve as feedback data to optimize the rule engine and AI model, achieving continuous system evolution and accuracy improvement.

[0007] In a preferred embodiment, this invention provides a method for intelligent parsing of tender documents based on multimodal fusion. Through hierarchical processing and cross-modal fusion, it achieves a deep understanding of the tender document content. The system first intelligently identifies the file type and adopts the optimal initial processing strategy for different formats. For PDF / Word files, the logical structure is obtained through document structure analysis, while pure image files enter the image preprocessing process. CV technology is used to divide the document into different content types such as text areas, table areas, and image areas. For text areas, adaptive OCR technology is used, employing differentiated recognition strategies for text areas with different fonts, layouts, and qualities. For table areas, the physical structure of the table is identified, then the cell content is accurately extracted, and finally, semantic relationships between data are established through contextual understanding. For image areas, dedicated recognition models are used for specific image types such as qualification certificates, seals, and signatures. The system establishes a unified semantic space, associating relevant information scattered across different modalities, and organizes structured data into a semantic graph.

[0008] In a preferred embodiment, the present invention also provides a multi-dimensional similarity calculation method for bid documents for bid clearing. Addressing the risks of collusion and bid-rigging unique to bidding scenarios, a comprehensive multi-dimensional similarity evaluation system is constructed. The system simultaneously processes the three core content dimensions of the bid documents: text content, image content, and price data. It calculates the corresponding similarities and employs a weighted fusion strategy to combine the similarities of each dimension to calculate a comprehensive similarity score. The weights are dynamically adjusted according to the characteristics of the bidding project. A Bayesian network is constructed to comprehensively calculate the overall risk probability based on evidence from all dimensions. The system outputs a risk score, visualizes similar content, generates a similarity heatmap, and provides risk evidence chain tracing to accurately locate suspicious content, reducing the workload of manual review.

[0009] In a preferred embodiment, this invention provides a qualification intelligent review method based on rule engines and NLP. Through rule intelligence and human-machine collaboration, it achieves accurate and efficient qualification review. The system automatically parses the qualification requirements clauses in the bidding documents, identifies hard conditions and scoring standards, and receives structured bidding information output by the document parsing module, including information such as business licenses, qualification certificates, personnel certificates, and performance certificates. The system uses NLP technology to automatically extract review rules from the bidding documents. The system integrates three types of rules to form a comprehensive review capability. For clear conditions, it quickly judges and uses rule engines such as Drools to achieve efficient matching. For fuzzy expressions, it uses a pre-trained model for semantic understanding to convert natural language rules into computable conditions. It connects to an external credit database and evaluates the comprehensive credit status of enterprises through trained models. It uses evidence theory to integrate the evaluation results of multiple rules to handle rule conflicts, calculates the confidence level for each review conclusion, dynamically adjusts the decision threshold according to the importance of the bidding project, and adopts stricter standards for important projects. The system automatically processes high-confidence conclusions and only submits low-confidence or conflicting results for manual review, ultimately generating a review report containing a complete chain of evidence. Attached Figure Description

[0010] The present invention will be further described below with reference to the accompanying drawings and embodiments: Figure 1 This is an overall flowchart of the method of the present invention; Figure 2 This is a flowchart of a tender document intelligent parsing method based on multimodal fusion according to the present invention; Figure 3 This is a flowchart of a multi-dimensional calculation method for bid document similarity for bid clearing according to the present invention; Figure 4 This is a flowchart of an intelligent qualification review method based on rule engine and NLP according to the present invention. Detailed Implementation

[0011] To better understand the purpose, system architecture, and functional implementation of this embodiment, the embodiments and features described herein can be combined with each other without conflict. The exemplary embodiments disclosed herein will be described below with reference to the accompanying drawings, including specific technical details disclosed to aid understanding; however, these details should be considered exemplary rather than restrictive. Therefore, those skilled in the art should understand that various improvements and adjustments can be made to the embodiments described herein without departing from the scope and core ideas of the invention. Similarly, for clarity, detailed descriptions of well-known technologies, functions, and structures are omitted in the following description.

[0012] Example 1 Figure 1 This is the overall flowchart of the method of the present invention.

[0013] like Figure 1 As shown, users upload tender documents in various formats, including Word, PDF, and scanned images, through the system front end. The system uses a security scanning engine to detect malicious files and performs format verification, checking file extensions, encoding formats, and file integrity. The system uses OCR, CV, and NLP technologies to perform multimodal parsing of the tender documents. Text parsing uses OCR to recognize text in scanned documents and images, LayoutLM to detect text regions, and a layout understanding model to restore chapters, split paragraphs, and recognize titles. NER entity extraction includes information such as company names, certificate numbers, personnel information, and project performance. Table extraction involves detecting table borders, reconstructing row and column structures, recognizing merged cells, and performing semantic classification of table content, including automatic recognition of personnel and qualification tables. Image parsing extracts SIFT features from images to recognize seals and certificates. Text data is structured in a fielded JSON format, tables are represented using a two-dimensional table structure, and images are structured using feature vectors and meta-vectors.

[0014] In the qualification review module, the system calls the rule engine and AI model to conduct compliance checks on the bidders' qualifications. First, it verifies the business license and extracts information such as the unified social credit code, company name, establishment date, and validity period. This information is compared with government platforms or third-party credit databases to detect anomalies such as expiration, revocation, or name alteration. Then, it verifies the company's qualifications and extracts information such as qualification category, level, certificate number, and validity period. This information is automatically compared with the bidding requirements to determine whether the qualification level meets the requirements and whether the certificate has expired. Finally, it verifies the personnel qualifications and extracts information such as personnel name, ID number, professional category, registration status, and validity period. This information is matched with the project category. Finally, it integrates the business license verification, company qualification verification, and personnel qualification verification to generate a qualification review structure table for output.

[0015] The similarity comparison module identifies plagiarism, duplication, and bid-rigging. First, text similarity detection is performed using SBERT and cosine similarity calculations. SimHash is used to detect highly repetitive paragraphs to determine whether plagiarism or duplication occurs. SIFT features are extracted and compared with performance photos and scanned documents to detect whether the same image is used in multiple tender documents. Finally, the pricing structure is decomposed into numerical vectors, and highly consistent pricing and abnormally low prices are detected through Euclidean distance, proportional deviation, and structural consistency analysis.

[0016] The risk assessment module evaluates financial risk, performance risk, and bid-rigging risk. The system constructs a credit and risk score for bidders, extracts key indicators from historical financial reports and credit database interfaces to determine financial risk, assesses performance risk based on historical winning projects, acceptance status, and the existence of complaints and default records, and finally uses a multi-dimensional weighted model to determine bid-rigging risk by combining text, image, and quotation similarity, and generates a risk assessment report for output.

[0017] The system integrates the results of qualification review, similarity comparison, risk assessment, and anomaly list to form a complete bill clearing report. Humans only need to review the anomaly list highlighted by the system. At the same time, the system records the human modifications, rule items overturned by humans, and judgment logic added by humans. The above human feedback is used to optimize the OCR error recognition model, NER entity recognition model, similarity threshold, and rule base to achieve continuous learning and performance improvement of the system.

[0018] Figure 2 This is a flowchart of a tender document intelligent parsing method based on multimodal fusion according to the present invention.

[0019] like Figure 2 As shown, the input PDF file is paginated and its content objects are extracted, while the Word document is parsed using XML. The text, table, and image content in different format tender documents are split, and a dedicated processing flow is matched for different data types. Text preprocessing improves efficiency by removing useless information, table preprocessing improves data extraction efficiency by accurately identifying row and column relationships and merging cells, and image preprocessing further improves the accuracy of subsequent processes by solving the problems of blurriness and tilt in scanned documents.

[0020] Text parsing uses the LayoutLMv3 layout understanding model to perform region-level recognition of page content. The fusion representation of OCR and layout understanding is shown in the following formula (1).

[0021] (1) in, , The weights obtained during training, Used to represent the text encoding result of a region after OCR recognition. The vector of the layout understanding model is used to balance text features and layout features. The fusion representation of OCR and layout understanding can improve the parsing ability of complex typed tender documents.

[0022] The system performs image enhancement and feature extraction on non-text areas such as performance photos and certificate images, and extracts SIFT feature aggregation and image hash features respectively to obtain quantized features, as shown in the following formula (2).

[0023] (2) in, and For the weight vector, For local key point features, For SIFI feature aggregation, As an image hash feature, the feature quantization vector can be used for structured representation of image content and similarity comparison.

[0024] This invention fuses three types of vectors: text, layout, and image, to construct a unified cross-modal representation, as shown in equation (3) below.

[0025] (3) in, For text semantic vectors, Quantize the image feature vector. To understand the model vectors of the layout, an attention mechanism is used to fuse the model and automatically learn the importance and correlation of different modalities.

[0026] Based on the fused multimodal features, core business entities are extracted through Named Entity Recognition (NER). After data standardization, structured results are output. This invention adopts a two-way verification strategy, combining the NER model output with rule template matching to improve extraction accuracy. When the NER confidence is <0.85, a second extraction of rule templates is automatically triggered. The extracted information is formatted according to a unified standard and stored in a database or knowledge graph. The final output is structured data that can be used for subsequent qualification review, similarity comparison and risk assessment.

[0027] Figure 3 This is a flowchart of a multi-dimensional calculation method for bid document similarity for bid clearing purposes, as per the present invention. like Figure 3 As shown, the standard structured data output from the document parsing module is received, and text similarity, image similarity, and price similarity are calculated respectively. Text similarity is to compare the semantic similarity of the technical solutions, business terms, etc. in the tender documents. Semantic similarity is captured by Sentence-BERT, and cosine similarity is calculated to effectively identify hidden plagiarism such as synonym substitution and sentence restructuring. Image similarity is to convert the image into a hash code through algorithms such as dHash, calculate the Hamming distance and normalize it, and use it to detect the similarity or tampering of qualification certificates, seals, signatures, scheme diagrams, etc. Price similarity is to standardize the price data, calculate the Euclidean distance, and convert it into a similarity score through a negative exponential function. The closer the prices are, the higher the similarity, thereby capturing abnormally consistent pricing behavior. This invention integrates the similarity results of the three dimensions and generates a comprehensive similarity score by considering the context and the relationship between dimensions, as shown in the following formula (4).

[0028] (4) in, The adaptive multidimensional fusion function is a function with parameters of Learnable models For text similarity, For image similarity, For price similarity, The function provides contextual information, including data such as the type of bidding project and historical data on bid-rigging patterns. It can dynamically combine and cross-validate various dimensions based on context to generate a comprehensive similarity score for manual review, thereby achieving a comprehensive quantification of the risk of bid similarity.

[0029] Based on the comprehensive similarity score and the preset threshold, it is determined whether to trigger a cross-label warning, as shown in the following formula (5).

[0030] (5) Among them, threshold Based on historical data and expert experience, it can determine whether the overall similarity exceeds the threshold. If it exceeds the threshold, it will trigger a flag-sharing warning and generate a visual report containing various similarity scores, risk levels, and evidence chains for manual review.

[0031] Figure 4 This is a flowchart of an intelligent qualification review method based on rule engine and NLP according to the present invention.

[0032] like Figure 4 As shown, the system receives standardized structured data from the document parsing module. Based on the type of bidding project, the system loads the corresponding qualification inspection rule set. The rules are stored in the form of Drools rules. The rule base supports dynamic configuration, hot updates, and project-level template loading. The system reads the predefined qualification review rule set, which includes hard conditions, scoring items, and fuzzy requirements. The system compares the structured data rule by rule to check whether each qualification meets the requirements, as shown in equation (6) below.

[0033] (6) in, For structured fields, For the rules, To determine the compliance of the qualification items, the system uses exact matching, range matching, regular expression validation, and logical condition calculation.

[0034] For rules with fuzzy or natural language expressions that cannot be identified by precise matching, the NLP model is invoked for semantic understanding and reasoning, as shown in Equation (7).

[0035] (7) in, The more bids a company has submitted in its history, the higher its NLP weight. The number of times a company has submitted a bid is considered. The more bids a company submits, the lower its NLP weighting becomes. The NLP judgment threshold is relaxed for new bidders and tightened for existing bidders.

[0036] The system combines rule matching and NLP-assisted results to make intelligent judgments. If the requirements are met, the system marks the system as passable; otherwise, it marks the system as abnormal and synchronizes the cause of the abnormality, as shown in equation (8) below.

[0037] (8) in, The original data fragments supporting this conclusion, Based on natural language interpretation, the results of rule matching are combined with NLP-assisted analysis to make a final review decision.

[0038] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0039] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A document analysis method for qualification review and similarity comparison, characterized in that, Includes the following steps: S1. Receive multi-format tender documents uploaded by users and perform preprocessing operations; S2. Use OCR, CV and NLP technologies to perform multimodal parsing on the tender documents, extract multimodal content such as text, tables and images, and store them in a structured manner; S3. The parsed data is distributed to the qualification intelligent review module, the tender document similarity comparison module, and the credit risk assessment module for analysis; S4. Generate a structured bid clearing report.

2. The method according to claim 1, characterized in that, In step S2, the multimodal analysis specifically includes: Intelligently identify tender document types, perform document structure analysis on PDF / Word files to obtain logical structure, and perform image preprocessing on pure image files; The tender document was divided into text, table, and image regions using CV technology. Adaptive OCR technology is used to recognize text in text areas, the row and column structure is reconstructed and cell content is extracted in table areas, and a dedicated recognition model is used to extract features in image areas. Construct a unified semantic space to associate and organize multimodal information into a semantic graph.

3. The method according to claim 1, characterized in that, In step S3, the qualification intelligent review module uses a rule engine combined with NLP technology to automatically extract review rules from the bidding documents, realize qualification compliance checks, and support dynamic adjustment of decision thresholds.

4. The method according to claim 1, characterized in that, In step S3, the tender document similarity comparison module simultaneously processes the three core dimensions of text content, image content, and quotation data, calculates the corresponding similarity, and generates a comprehensive similarity score through a weighted fusion strategy.

5. The method according to claim 4, characterized in that, The similarity calculation also includes text similarity detection using SBERT and cosine similarity, image similarity detection using SIFT features to compare performance photos and scanned documents, and quotation similarity analysis through numerical vector analysis.

6. The method according to claim 1, characterized in that, In step S3, the credit risk assessment module assesses the overall credit status of enterprises based on external credit databases and training models, and constructs a risk scoring system for bidders.

7. The method according to claim 1, characterized in that, In step S3, the system provides a visual rule configuration interface, which supports users to dynamically define and adjust review rules and synchronize them to the rule engine in real time.

8. The method according to claim 1, characterized in that, In step S4, the comprehensive bid clearing report includes qualification review results, similarity comparison results, risk assessment results, and a list of anomalies, and allows human auditors to make the final judgment.

9. The method according to claim 1, characterized in that, In step S4, the system records manually modified points, manually overturned rule items, and manually added judgment logic, which are used as the basis for optimizing the OCR error recognition model, NER entity recognition model, similarity threshold, and rule base, so as to achieve continuous learning and performance improvement of the system.

10. The method according to claim 9, characterized in that, The continuous learning optimization mechanism triggers a model retraining task when the consistency rate between the system's judgment result and the manual review result is lower than a preset threshold.