Bidding document compliance auditing method based on multi-modal knowledge graph
By constructing a bid document compliance review method using a multimodal knowledge graph, this approach addresses the issues of low efficiency in manual review and the difficulty of rule engines in handling unstructured text in existing technologies. It enables intelligent and interpretable automatic compliance review of bid documents, thereby improving the accuracy and reliability of the review process.
Patent Information
- Application Number
- CN202511586458.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-01
- Publication Date
- 2026-02-10
AI Technical Summary
Existing tender document review relies on manual methods, which are inefficient and prone to omissions. Rule engines struggle to handle semantic ambiguity in unstructured text and lack unified correlation analysis capabilities for multi-source heterogeneous data.
A method for reviewing bid compliance based on a multimodal knowledge graph is constructed. Through natural language processing and image recognition technologies, a multimodal knowledge graph containing text entities, table structures and image templates is built. Modality separation and structuring are performed, semantic mapping and node matching are carried out, and automatic review reasoning and cross-modal consistency verification are executed.
It enables intelligent, traceable, and explainable automated compliance review of tender documents, improving the accuracy and reliability of the review and reducing the cost of manual review.
Smart Images

Figure CN121504484A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of tender document review technology, and more specifically, to a method for reviewing tender document compliance based on multimodal knowledge graphs. Background Technology
[0002] Compliance review of tender documents is a crucial step in ensuring the fairness, impartiality, and legality of the procurement process. Currently, tender document reviews primarily rely on manual methods. Reviewers must meticulously verify the content, format, eligibility requirements, and technical solutions of the tender documents based on relevant legal documents. However, tender documents typically contain a large amount of unstructured information, including text descriptions, tabular data, and image attachments, making the review process labor-intensive, time-consuming, and highly subjective. This process is not only time-consuming and labor-intensive but also easily affected by differences in individual knowledge and understanding, leading to the oversight of some compliance issues.
[0003] Meanwhile, while existing rule-engine-based automated review systems can perform format validation and keyword matching for some key fields, they struggle to accurately understand semantic ambiguity and contextual logic in natural language when processing unstructured text, failing to support deep semantic reasoning. Furthermore, tender document review involves multi-source heterogeneous data, including legal provisions, policy documents, historical precedents, industry standards, and past winning bid documents. These data sources are scattered, formatted inconsistently, and lack effective correlation and fusion mechanisms, making it difficult for existing systems to establish a unified knowledge framework to support automated analysis across documents and modalities.
[0004] Therefore, it is necessary to provide a bid compliance review method based on multimodal knowledge graphs to solve the above technical problems. In order to solve the above problems, a technical solution is provided. Summary of the Invention
[0005] To overcome the aforementioned shortcomings of existing technologies, this invention provides a bid compliance review method based on multimodal knowledge graphs. This method addresses the problems of existing manual reviews relying on experience, being inefficient and prone to missing legal clauses, the inability of unified rule engines to handle semantic ambiguity in unstructured text, and the lack of unified correlation analysis capabilities for heterogeneous source data.
[0006] To achieve the above objectives, the present invention provides the following technical solution: The method for reviewing tender document compliance based on multimodal knowledge graphs includes the following steps: Obtain standard reference documents from authoritative data sources, and use natural language processing and image recognition technologies to perform text entity extraction, visual template recognition and relationship fusion respectively, and construct a multimodal knowledge graph containing text entities, table structures and image templates; By performing modal separation and structuring on the input tender documents, the structured features of the tender documents are obtained; the structured features include text modal features, table modal features, and image modal features; Based on the text entities of the multimodal knowledge graph and the text modal features of the tender documents, semantic mapping and node matching are performed to obtain textual difference features and determine the clause nodes in the standard reference documents corresponding to the structured features. Automatic review and reasoning are performed based on the logical rules in the multimodal knowledge graph, and the compliance of the tender documents is initially judged based on the review results; For tender documents that meet the audit results, cross-modal consistency verification is performed. Based on the comprehensive verification results, the tender documents are labeled in a hierarchical manner, including "compliant", "partially compliant" or "non-compliant".
[0007] As a further aspect of the present invention, the structured features of the tender document are obtained by performing modal separation and structuring processing on the input tender document. The specific steps are as follows: Obtain the input tender document, parse and separate it into text modality, table modality and image modality; For the text modality, entity recognition, relation extraction, and semantic role labeling are performed to obtain the semantic information of the text expression as text modality features; for the table modality, structure recognition is performed to convert the table content into structured matrix features as table modality data; for the image modality, OCR text recognition and object detection are performed to extract the image expression information as image modality features.
[0008] As a further aspect of this invention, based on the text entities and relationships of a multimodal knowledge graph, semantic mapping and node matching are performed with the text modal features of the tender document to obtain textual difference features and determine the clause nodes in the standard reference document corresponding to the structured features. The specific steps are as follows: Obtain the set of text entities in a multimodal knowledge graph ,in, For the a-th text entity, For the m-th text entity, The total number of text entities and the set of text modal features of the tender document. ,in, For the k-th text modal feature, For the Kth text modality feature, This represents the total number of text modal features. Based on semantic mapping, the first text representation of each text entity and the second text representation of the text modality features are obtained respectively; Semantic similarity is obtained by matching the first and second text expressions respectively, and text difference features are calculated based on semantic similarity. Based on the set of relationships in the multimodal knowledge graph ,in, This describes the relationship between text entity a and text entity b. This describes the relationship between text entity A and text entity B. The text difference features are compared with a preset difference threshold. If the text difference features are less than the preset difference threshold, a relation check and matching is performed. If the text difference features are greater than or equal to the preset difference threshold, it indicates that the corresponding text entity and text modality features are not related.
[0009] As a further aspect of the present invention, the first text expression includes a first text semantic expression and a first text structural expression; the second text expression includes a second text semantic expression and a second text structural expression.
[0010] As a further aspect of the present invention, the first text expression includes a first text semantic expression and a first text structural expression; the first text semantic expression is the main semantic vector of the text entity, and the first text structural expression is the slot structured vector of the text entity.
[0011] The second text representation includes the second text semantic representation and the second text structural representation; the second text semantic representation is the main semantic vector of the text modality features, and the second text structural representation is the slot structured vector of the text modality features.
[0012] As a further aspect of the present invention, the specific steps for performing relationship verification and matching include: when the text difference feature is less than a preset difference threshold, searching within the context of the corresponding text modality feature, analyzing the corresponding text modality feature that has the same relational constraint as the corresponding text entity, and establishing a matching relationship based on the text entity and the text modality feature.
[0013] As a further aspect of the present invention, if, after analysis, no corresponding text modal feature with the same relational constraint as the corresponding text entity is detected, it indicates that there is no corresponding clause node in the standard reference file; if a corresponding text modal feature with the same relational constraint as the corresponding text entity is detected, it indicates that there is a corresponding clause node in the standard reference file.
[0014] As a further aspect of this invention, cross-modal consistency verification is performed on tender documents that meet the review results. The tender documents are then categorized into hierarchical levels based on the comprehensive verification results, including "compliant," "partially compliant," or "non-compliant." The specific steps are as follows: For tender documents that meet the review results, cross-modal consistency verification is performed, including text-table consistency verification, text-image consistency verification, and image-map template consistency verification. If all three modal conformity verifications are passed, the tender document will be marked as "passed"; if one or two of the three modal conformity verifications are passed, the tender document will be marked as "partially passed"; if none of the three modal conformity verifications are passed, the tender document will be marked as "unqualified".
[0015] As a further aspect of the present invention, the text-table consistency verification verifies whether the declared values in the tender document text are consistent with the summarized values in the financial details table. If they are consistent, it indicates that the text-table consistency verification is qualified.
[0016] As a further aspect of the present invention, the text-image consistency verification verifies whether the qualifications described in the tender document text are consistent with the qualification information identified in the attached certificate image. If they are consistent, it indicates that the text-image consistency verification is qualified. The technical effects and advantages of this invention's bid compliance review method based on a multimodal knowledge graph are as follows: This invention extracts standard reference documents from authoritative data sources and constructs a multimodal knowledge graph using natural language processing and image recognition technologies. This allows for the simultaneous understanding of the semantic logic, table structure, and format template features of regulatory provisions, providing a unified and reasonable knowledge system for subsequent bid review. Through modal separation and structuring, the text, tables, and images in the bid documents are transformed into computable structured features, achieving standardized expression of unstructured documents and significantly improving machine readability. Through semantic mapping and node matching, a one-to-one correspondence is established between the textual modal features of the bid documents and the standard clause nodes in the knowledge graph, and textual difference features are extracted. This enables accurate identification of semantic deviations, missing clauses, or logical conflicts in the bid documents, improving the accuracy and relevance of the review results. Based on the automatic reasoning mechanism of knowledge graph logical rules, a systematic analysis of the constraint relationships, citation relationships, and contextual dependencies between bid clauses can be performed, completing automated compliance judgments and reducing the subjectivity and omission risks of manual review. Cross-modal consistency verification enables logical correspondence and information consistency detection between text, table, and image modalities. Based on the verification results, the tender documents are annotated hierarchically to form an interpretable compliance assessment result.
[0017] This invention addresses the problems of heavy reliance on manual review, low review efficiency, and difficulty in handling unstructured information and semantic ambiguity in existing tender document review by introducing multimodal knowledge graphs, semantic mapping, and logical reasoning technologies. It achieves intelligent, traceable, and interpretable automatic compliance review of tender documents, effectively improving the accuracy and reliability of the review and reducing the cost of manual review. Attached Figure Description
[0018] Figure 1A flowchart of a tender compliance review method based on multimodal knowledge graph provided in an embodiment of the present invention. Detailed Implementation
[0019] The technical solutions of this invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described technical solutions are only a part of this invention, and not all of it. All other technical solutions obtained by those skilled in the art based on the technical solutions of this invention without inventive effort are within the scope of protection of this invention.
[0020] like Figure 1 The diagram shown is a flowchart of a tender document compliance review method based on a multimodal knowledge graph provided in an embodiment of the present invention. Figure 1 The execution entity of the method shown can be a software and / or hardware device. The execution entity of this application can include, but is not limited to, at least one of the following: user equipment, network equipment, etc. User equipment can include, but is not limited to, computers, smartphones, personal digital assistants (PDAs), and the aforementioned electronic devices. Network equipment can include, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of computers or network servers. Cloud computing is a type of distributed computing, consisting of a super virtual computer composed of a group of loosely coupled computers. This embodiment does not limit this. Steps 1 to 5 are detailed below: Step 1: Obtain standard reference documents from authoritative data sources, and use natural language processing and image recognition technologies to perform text entity extraction, visual template recognition, and relationship fusion to construct a multimodal knowledge graph containing text entities, table structures, and image templates. Step 2: By performing modal separation and structuring processing on the input tender document, the structured features of the tender document are obtained; the structured features include text modal features, table modal features and image modal features; Step 3: Based on the text entities of the multimodal knowledge graph and the text modal features of the tender document, perform semantic mapping and node matching to obtain text difference features and determine the clause nodes in the standard reference document corresponding to the structured features. Step 4: Perform automatic review and reasoning based on the logical rules in the multimodal knowledge graph, and make a preliminary judgment on the compliance of the tender documents based on the review results; Step 5: For tender documents that meet the audit results, perform cross-modal consistency verification. Based on the comprehensive verification results, label the tender documents in a hierarchical manner, including "compliant", "partially compliant" or "non-compliant".
[0021] Preferably, standard reference documents are obtained from authoritative data sources. Natural language processing and image recognition technologies are then used to perform text entity extraction, visual template recognition, and relationship fusion to construct a multimodal knowledge graph containing text entities, table structures, and image templates. The specific steps are as follows: Obtain standard reference documents from authoritative data sources, including legal documents, historical tender documents, winning bid announcements, industry standards, and expert experience databases, and standardize the data format of the standard reference documents; Natural language processing is used to extract text entities, attributes, and relations from standard reference documents to form structured triples. The table structure and image template in the standard reference file are extracted using image recognition technology and stored in the atlas as visual entities. Multimodal entities are obtained based on text entities and visual entities. Multimodal entities and relationships are then fused to construct a multimodal knowledge graph that includes text entities, table structures, and image templates.
[0022] In one embodiment of the present invention, taking the compliance review of tender documents for government procurement projects in construction engineering as an example, this invention illustrates how to construct a multimodal knowledge graph based on authoritative data sources to achieve automated semantic understanding and rule-based reasoning of tender document content.
[0023] First, standard reference documents are obtained from various authoritative data sources, including legal and regulatory documents, national and local engineering construction standards, sample winning bid announcements, historical bid document templates, and expert knowledge bases. To address the format differences between documents from different sources, text files, scanned copies, and PDF versions of the announcements undergo standardized processing, including format conversion, character recognition (OCR), encoding standardization, and layout correction, thereby ensuring consistency in subsequent feature extraction.
[0024] Secondly, based on natural language processing technology, the aforementioned standard reference documents are subjected to text semantic parsing. Key text entities, such as "bid bond," "warranty period," "legal representative's signature," "contract amount," and "bid opening date," are identified using named entity recognition and relation extraction models. Their corresponding attributes and semantic relationships are then extracted, for example, "bid bond—not less than—5% of the contract amount" and "warranty period—not less than—24 months." The extracted semantic information is organized into structured triples, constituting the text semantic layer of the knowledge graph.
[0025] Meanwhile, based on image recognition and table parsing technologies, visual features are extracted from tables and diagrams contained in standard reference documents. For structured image content commonly found in engineering bidding documents, such as "bid quotation table," "project schedule table," and "construction organization chart," the cell structure, title hierarchy, and field correspondence are identified, and they are parsed into table entities or visual template entities, such as "quotation table template" and "organization chart template," and added to the visual layer of the graph in the form of nodes.
[0026] Subsequently, the aforementioned text entities and visual entities are fused to form multimodal entity nodes, and cross-modal relationships are established through a relationship fusion algorithm. For example, the "Bidding Quotation Template" node establishes a "field mapping relationship" with the "Quotation Amount" text entity, and the "Construction Organization Chart" node establishes a "job correspondence relationship" with the "Project Manager" and "Technical Manager" entities. Ultimately, a multimodal knowledge graph is constructed, comprising a text semantic layer, a table structure layer, and a visual template layer.
[0027] In practical applications, when receiving construction project tender documents for review, this multimodal knowledge graph can be used to automatically match and semantically compare the tender content. For example, the statement "warranty period is 12 months" in the tender document can be matched with the standard clause node "warranty period is no less than 24 months" in the knowledge graph, and the difference calculation module can automatically identify any non-compliance risks. Simultaneously, the "bid price table" image in the tender document appendix can be visually matched with the standard price table template in the knowledge graph to detect field completeness and format compliance, thus achieving intelligent compliance review of the tender documents at both textual and visual levels.
[0028] As can be seen from the above embodiments, the embodiments of the present invention can achieve unified understanding and automatic reasoning of multimodal data in real projects, significantly improve the automation, accuracy and interpretability of tender review, reduce the workload of manual comparison, and improve the standardization and transparency of procurement review.
[0029] Preferably, the structured features of the tender document are obtained by performing modal separation and structuring processing on the input tender document; the structured features include text modal features, table modal features, and image modal features, and the specific steps are as follows: Obtain the input tender document, parse and separate it into text modality, table modality and image modality; For the text modality, entity recognition, relation extraction, and semantic role labeling are performed to obtain the semantic information of the text expression as text modality features; for the table modality, structure recognition is performed to convert the table content into structured matrix features as table modality data; for the image modality, OCR text recognition and object detection are performed to extract the image expression information as image modality features.
[0030] In one embodiment of the present invention, taking an electronic tender document for a construction engineering bidding project as an example, the process of modal separation and structuring of the input tender document file in practical applications is illustrated: First, the system receives the input tender documents, in formats including PDF, Word, or scanned copies. The file parsing module automatically performs layered parsing of the tender content. Based on layout characteristics, text distribution patterns, and file metadata, the tender content is divided into three modalities: text modality, table modality, and image modality. The text modality includes natural language paragraphs such as the main text, descriptions, and commitment letters; the table modality includes structured information tables such as price quotations, equipment lists, and personnel allocation tables; and the image modality includes visual content such as project site plans, construction organization diagrams, and scanned copies of qualification certificates.
[0031] For text modalities, a semantic parsing model based on natural language processing is employed for entity recognition, relation extraction, and semantic role labeling. For example, key text entities such as "bid bond," "legal representative's signature," "warranty period," and "bid deadline" are identified in the tender document body. Semantic relationships are then identified through relation extraction, such as "bid bond—amount—500,000 yuan" and "warranty period—no less than—12 months." Semantic role labeling further determines the action subject and logical relationships, such as "qualifications required from the construction unit" and "delivery time promised by the bidder." Finally, structured text modal features are formed, represented as entity-attribute-relation triples or embedded vectors, for subsequent semantic comparison and compliance analysis.
[0032] For table-based modalities, table structure recognition algorithms, such as table line detection and cell segmentation models, automatically parse the row and column structure and header levels of the table, converting the table content into structured matrix features. For example, from a "Project Bidding Price Table," the column fields "Serial Number," "Sub-item Name," "Quantity," "Unit Price," and "Total Price" are identified, and the corresponding data unit values are extracted. The parsed results are uniformly encoded into key-value pair format or relational table structure to preserve the internal hierarchy and logical relationships of the table, enabling programmatic data processing. This process effectively supports subsequent cross-form validation and amount consistency verification.
[0033] For image modalities, OCR technology is first used to extract text information from the scanned document. This includes identifying company names, certificate numbers, dates, and signatures. Simultaneously, object detection models, such as YOLO or Mask R-CNN, are used to identify key objects in the image, such as seals, signature areas, table areas, or drawing annotations. After feature extraction, the identified text content is fused with the location of the target objects to generate corresponding image modal features, such as "legal representative's signature exists and is in the correct position," "certificate number matches OCR text," and "complete annotations on the project floor plan."
[0034] Through the processing of the above three modalities, a unified structured feature set for tender documents is finally formed, which includes textual semantic features, table matrix features, and image recognition features. This structured feature not only retains the semantic logic and layout information of the original tender document, but also has computability and comparability, providing standardized input for subsequent semantic mapping, node matching, and automatic compliance reasoning.
[0035] In actual engineering projects, this step can automatically understand the semantic and structural features of various types of information in the tender documents. For example, when reviewing the "warranty period" clause, the statement description can be read directly from the text modality; if the information appears in tabular form, the corresponding fields can be extracted from the tabular modality matrix; if the warranty period only appears in scanned attachments, its content can also be obtained through OCR recognition and object detection. In this way, the three modalities complement each other, ensuring the completeness and accuracy of information extraction and providing a reliable data foundation for intelligent tender compliance analysis.
[0036] Preferably, based on the text entities and relationships of the multimodal knowledge graph, semantic mapping and node matching are performed with the text modal features of the tender document to obtain textual difference features and determine the clause nodes in the standard reference document corresponding to the structured features. The specific steps are as follows: Obtain the set of text entities in a multimodal knowledge graph ,in, For the a-th text entity, For the m-th text entity, The total number of text entities and the set of text modal features of the tender document. ,in, For the k-th text modal feature, For the Kth text modality feature, This represents the total number of text modal features. Based on semantic mapping, the first text representation of each text entity and the second text representation of the text modality features are obtained respectively; Semantic similarity is obtained by matching the first and second text expressions respectively, and text difference features are calculated based on semantic similarity. Based on the set of relationships in the multimodal knowledge graph ,in, This describes the relationship between text entity a and text entity b. This describes the relationship between text entity A and text entity B. The text difference features are compared with a preset difference threshold. If the text difference features are less than the preset difference threshold, a relation check and matching is performed. If the text difference features are greater than or equal to the preset difference threshold, it indicates that the corresponding text entity and text modality features are not related.
[0037] This invention uses the review of tender documents for a construction engineering procurement project as an example to illustrate how, in practical applications, text semantic mapping and node matching can be achieved based on multimodal knowledge graphs, and text difference features can be calculated to determine the correspondence between clauses: First, extract the set of text entities from the constructed multimodal knowledge graph. Each text entity This corresponds to a standard clause, specification requirement, or typical case node, such as "the bid security shall not be less than 5% of the contract amount," "the warranty period shall not be less than 24 months," or "the bid documents must be signed and sealed by the legal representative." Simultaneously, a set of text modal features is extracted from the structured bid documents. Each text modality feature The actual text description in the tender document should correspond to phrases such as "Our company's bid bond ratio is 3%", "The warranty period for this project is 12 months", and "Signed and stamped by the authorized representative".
[0038] Subsequently, a semantic mapping algorithm is used to calculate the first textual representation of each knowledge graph text entity and the second textual representation of each tender document text modal feature. For example, for the knowledge graph node "The bid security shall not be less than 5% of the contract amount", its first textual semantic representation is a legal clause semantic vector, and its first textual structure is {amount type: percentage, minimum value: 0.05}. Meanwhile, the second textual semantic representation of "Our company's bid security is 3% of the contract amount" in the tender document is a bidding clause semantic vector, and its second textual structure is {amount type: percentage, actual value: 0.03}.
[0039] Semantic matching and structural comparison are performed on the two expressions to calculate semantic similarity and difference features. For example, the clause has high semantic similarity but significant differences in structural parameters. The overall calculated text difference feature value is 0.02, which is compared with a preset difference threshold of 0.05. When 0.02 < 0.05, it indicates that the tender document statement is semantically consistent with the legal clause, with only slight numerical deviations, and will proceed to the relational verification stage.
[0040] During the relation verification phase, the relation set in the knowledge graph is invoked. This includes types such as "reference clauses," "constraints," and "exclusion relationships." For example, a "constraint relationship" exists between the nodes "bid security" and "warranty period," indicating that both should meet the specified amount and duration matching requirements. The system retrieves text modal features adjacent to or semantically related to the current text modal feature, such as "warranty period is 12 months." If a significant difference is detected between this text and the standard clause "warranty period is no less than 24 months," the overall tender content is deemed not to meet the clause association requirements. Conversely, if all relationship conditions are met, the semantic mapping is confirmed as a correct match.
[0041] If the semantic discrepancy is greater than or equal to a preset threshold, the statement in the tender document is considered to have a significant semantic or structural difference from the standard clause, and therefore no valid correspondence is determined. For example, if the statement in the tender document, "No bid security deposit is required for this project," completely conflicts with the standard clause "A bid security deposit should be paid," it will be directly marked as a mismatch node, and the relationship verification step will not be performed.
[0042] Through the above process, semantic matching and logical verification of text modal features and knowledge graph clause nodes can be automatically completed during actual tender document review. For example, when a tender document contains "the project manager holds a Level 2 Construction Engineer certificate," it can be mapped to the "the project manager must possess a Level 1 Construction Engineer qualification" node in the knowledge graph, and the difference features can be calculated to identify issues of inconsistent qualification levels. At the same time, implicit conflicts can be automatically identified by combining graph relationships, such as the constraint contradiction between "the bidder must possess a safety production license" and "the project manager must be an externally hired person."
[0043] Therefore, the embodiments of the present invention can realize automatic semantic alignment, difference detection and logical consistency judgment between tender documents and standard clauses in real projects. This can not only significantly reduce the amount of manual review, but also improve the ability to identify hidden non-compliant clauses, semantic deviations and logical conflicts, thereby improving the accuracy and intelligence level of tender compliance review.
[0044] Preferably, the specific steps for performing relationship verification and matching include: When the text difference features are less than the preset difference threshold, the system searches within the context of the corresponding text modality features and analyzes the corresponding text modality features that have the same relational constraints as the corresponding text entity. A matching relationship is then established based on the text entity and the text modality features.
[0045] If, after analysis, no corresponding text modal feature with the same relational constraint as the corresponding text entity is detected, it indicates that there is no corresponding clause node in the standard reference document; if a corresponding text modal feature with the same relational constraint as the corresponding text entity is detected, it indicates that there is a corresponding clause node in the standard reference document.
[0046] Specifically, a matching relationship is established based on text entities and text modal features. If the text modal features With text entities If the textual difference features are less than a preset difference threshold, then in the text modality features Context-based retrieval, analysis, and text entities The same relational constraints exist. Corresponding text modal features At this point, based on text entities and text modal features Establish matching relationship .
[0047] This invention uses the electronic tender document review of a road engineering project as an example to illustrate how, in practical applications, this invention performs relationship verification and matching based on text semantic matching to further verify the logical consistency between the tender document content and standard clauses: In the aforementioned semantic mapping and node matching steps, a preliminary matching relationship has been established between the set of modal features of the tender text and the set of text entities in the multimodal knowledge graph, and the matching results for each pair have been calculated. The textual difference feature values.
[0048] In this embodiment, it is assumed that the text entity nodes in the knowledge graph are... The standard clause "Project manager must possess a Level 1 Construction Engineer qualification certificate" indicates the following relationship: That is, with text entity nodes There is a parallel constraint relationship between "project managers must have a safety production qualification certificate" and the corresponding tender documents contain text modal features. The text states: "The project manager holds a Level 2 Construction Engineer certificate." The modal features and entity nodes of this text are calculated. The text difference feature is 0.03, which is less than the preset threshold of 0.05, indicating that the semantics are basically the same but there may be structural differences, that is, the certificate levels do not match. At this time, the relationship verification and matching process is triggered.
[0049] First, text modality features are used in the tender document. Focusing on the core content, the search is performed within its context, such as adjacent paragraphs, the same table cell, or the same chapter, to analyze whether there are any matching text entities. Having the same relational constraints Other text modal features. Through context analysis, another text modal feature was retrieved. "The project manager has passed the safety production assessment and holds a valid certificate." (Judgment) Text entities in knowledge graphs The semantic similarity between "project manager must have a safety production assessment certificate" and "project manager must have a safety production qualification certificate" is high, and they satisfy the relational constraints. Therefore, based on and Based on the graph relationships, establish the mapping relationship between tender features and entity nodes: This indicates that the corresponding clauses in the tender documents have met the parallel constraints in the standard reference documents, confirming that this part of the content is compliant.
[0050] Conversely, if no text entity is detected in the context search There is a class relationship constraint. If the text modal characteristics are such that there is no description of safety production assessment, it indicates that there is no corresponding clause node mapping relationship in the standard reference document. In this case, the clause is marked as "incomplete relationship" in the audit results, indicating a potential compliance risk.
[0051] In another scenario of this invention, if the tender document contains the following description: "The project manager holds a Level 2 Construction Engineer certificate and is not required to undergo safety production assessment," the calculation will be... and The textual differences are small, but the description and relational constraints were detected. The next node The “Safety Assessment Requirements” have a semantic conflict relationship, i.e., an exclusion relationship. Based on the “exclusion” relationship type in the graph, the clause is automatically determined to be contradictory to the standard clause. In this case, there is no corresponding clause node in the output standard reference file.
[0052] Through the above process, this invention enables the use of predefined relation sets in a knowledge graph to perform contextual logic checks and relational consistency verification on tender documents, even when semantic similarity is high but logical differences exist. This mechanism not only identifies consistency and discrepancies at the textual level but also further uncovers hidden logical conflicts or omissions in constraints between tender clauses. For example, in the review of large projects, the related clauses of "bid security" and "performance bond" can be checked simultaneously: if the tender document states that "the bid security is 3% of the contract amount" but no description of a performance bond appears, the relevant clauses will be checked based on the knowledge graph relationships. The statement "The bid bond and performance bond must appear in pairs" indicates that this part of the content is incomplete, thus generating a "relationship missing" warning message.
[0053] Therefore, by introducing a relationship verification and matching mechanism, the embodiments of the present invention can achieve two-layer compliance analysis from semantic matching to logical verification, which significantly improves the accuracy, comprehensiveness and intelligence of tender document review, especially in complex tender document scenarios involving multiple clause dependencies or constraints.
[0054] Preferably, semantic similarity is obtained by matching and analyzing the first text expression and the second text expression respectively, wherein the first text expression includes the first text semantic expression and the first text structural expression; the second text expression includes the second text semantic expression and the second text structural expression. The specific steps of semantic similarity are as follows: Based on the first and second text semantic expressions, the semantic similarity between text entities and text modal features is calculated. The formula for calculating semantic similarity is as follows: ; In the formula: For semantic similarity, For the first textual semantic expression, For the semantic expression of the second text; Based on the first and second text structure representations, the structural similarity between text entities and text modal features is calculated. The formula for calculating structural similarity is as follows: ; In the formula: For structural similarity, This is the first text structure expression. This is for the expression of the second text structure; The semantic similarity is obtained by comprehensively weighting semantic similarity and structural similarity.
[0055] It should be noted that the first text expression includes the first text semantic expression and the first text structural expression; the first text semantic expression is the main semantic vector of the text entity, and the first text structural expression is the slot structured vector of the text entity.
[0056] The second text representation includes the second text semantic representation and the second text structural representation; the second text semantic representation is the main semantic vector of the text modality features, and the second text structural representation is the slot structured vector of the text modality features.
[0057] In one embodiment of the present invention, taking the review of tender documents for a construction project as an example, this invention illustrates how, in practical applications, the semantic similarity between the tender document text and standard clauses is calculated based on semantic vectors and structured vectors, thereby achieving intelligent semantic matching of tender document statements.
[0058] In this embodiment, text entities are first extracted from the knowledge graph. "The bid security shall not be less than 5% of the contract amount," and the corresponding text modal features are extracted from the bid documents. "Our company's bid security is 3% of the contract amount." For these two texts, generate the first text representation and the second text representation respectively.
[0059] The first textual expression includes: the first textual semantic expression : Through a semantic encoding model (mapping standard clause text into high-dimensional semantic vectors to reflect the meaning of the clause in the semantic space; first text structure expression) : Extract parameter information from the terms by slot extraction, such as amount type = proportion, minimum value = 0.05.
[0060] The second textual expression includes: the second textual semantic expression : Converting tender statements into semantic vectors using the same semantic encoding model; second text structure expression Structured features are extracted through numerical analysis and unit normalization, such as amount type = proportion and actual value = 0.03.
[0061] The semantic similarity calculated based on the semantic expressions of the first and second texts is 0.94, indicating that the two are highly similar at the semantic level, that is, the content of the two sentences is consistent, both involving the "ratio of bid bond to contract amount". Subsequently, the structural similarity calculated based on the structural expressions of the first and second texts is 0.6, indicating that there is a certain deviation at the structural level.
[0062] Finally, based on the set weighting parameters, such as semantic weight 0.7 and structural weight 0.3, the final weighted semantic similarity is calculated to be 0.83. This result indicates that the tender document text and the standard clauses have a high degree of consistency in overall semantics, but there is a slight deviation in the parameter structure. Through the above calculation process, this embodiment of the invention can achieve semantic understanding and structural parameter comparison of tender documents in real-world applications. Semantic similarity is used to determine whether the meanings of the texts are consistent, while structural similarity is used to analyze whether the content of the clauses meets the standard requirements. The weighted fusion of the two yields a comprehensive quantitative assessment of the tender document content. This approach breaks through the limitations of traditional keyword or rule-based matching, accurately identifying differences in expression, numerical deviations, and potential semantic contradictions, thereby significantly improving the automation level and accuracy of tender document review.
[0063] Preferably, automatic review and reasoning are performed based on the logical rules in the multimodal knowledge graph, and the compliance of the tender documents is initially judged based on the review results. The specific steps are as follows: Automated reasoning is performed using predefined logical rules in the knowledge graph or learned from standard reference documents, such as: "IF project type is 'government procurement' AND bidder is 'natural person', THEN ineligible". Based on the results of automated reasoning, a preliminary judgment is made as to whether the tender documents are compliant.
[0064] In practical engineering applications, this invention first maps the text modal features, table modal features, and image modal features obtained from the tender document through modal separation and structuring to corresponding nodes in a multimodal knowledge graph, forming preliminary matching relationships and differential features. For example, after extracting structured information such as "Project Type = Government Procurement," "Bidder Type = Natural Person," "Bid Bond = 0," and "Business License Validity Period = Expired" from the tender document, rule matching and chain reasoning are performed on these nodes and relationships based on logical rules predefined in the knowledge graph or automatically learned from standard reference documents. Specifically, if the graph contains the rule "IF Project Type = 'Government Procurement' AND Bidder = 'Natural Person', THEN ineligible for bidding," and the bid feature mapping indicates that both conditions are met simultaneously, the inference engine will trigger this rule and generate a ineligibility conclusion. Simultaneously, if the graph contains the rule "IF Bid Bond < Required Bond Ratio OR Bid Bond = 0, THEN ineligible," and the bid bond is calculated to be 0 from the table modal features, the corresponding rule will also be triggered, and the system will generate a non-compliance judgment regarding the bid bond. For certificate verification, if the graph rule is "IF Certificate Validity Period ≤ Current Date, THEN Qualification Expired," and the image modality matches the template via OCR to determine that the certificate has expired, then the qualification expiration conclusion will be triggered.
[0065] Through the aforementioned automatic reasoning mechanism based on knowledge graph logic rules, a preliminary compliance judgment can be made immediately upon receiving the tender documents, and corresponding discrepancy reports and rectification suggestions can be generated. This significantly reduces the time required for manual preliminary review, improves the accuracy of detecting rule-based violations, and provides a clear and traceable chain of evidence and handling guidelines for subsequent manual review.
[0066] Preferably, for tender documents that meet the audit results, cross-modal consistency verification is performed. The tender documents are then categorized into hierarchical levels based on the overall verification results, including "compliant," "partially compliant," or "non-compliant." The specific steps are as follows: For tender documents that meet the review results, cross-modal consistency verification is performed, including text-table consistency verification, text-image consistency verification, and image-map template consistency verification. If all three modal conformity verifications are passed, the tender document will be marked as "passed"; if one or two of the three modal conformity verifications are passed, the tender document will be marked as "partially passed"; if none of the three modal conformity verifications are passed, the tender document will be marked as "unqualified".
[0067] Specifically, the text-table consistency verification verifies whether the declared values, such as the price, in the tender document text are consistent with the summarized values in the financial details table. If they are consistent, the text-table consistency verification is considered successful. The text-image consistency verification verifies whether the qualifications described in the tender document text, such as "possessing a certain model certification", are consistent with the qualification information identified in the attached certificate image, such as the certificate name, issuing authority, and validity period. If they are consistent, the text-image consistency verification is considered successful. Image-graph template consistency verification involves comparing the qualification certificate images in the tender documents with the standard templates stored in the knowledge graph to verify their format compliance and authenticity, such as the location of the seal or the absence of key fields. If the verification is successful, it indicates that the image-graph template consistency verification is qualified.
[0068] In this embodiment of the invention, during the bidding process for a large-scale engineering project, the bidding party needs to rigorously review the bids submitted by bidding companies to ensure the authenticity, completeness, and compliance of the bidding materials. A company submitted a complete bid document. The review process, combined with cross-modal consistency verification, can be performed as follows: First, a preliminary compliance check is conducted on the bid document to confirm that it contains necessary chapters, has complete text content, and includes all appendices and qualification certificates, thus determining the review result to be "compliant." Subsequently, cross-modal consistency verification is performed on the bid document.
[0069] Extract the company's price quotation information from the main body of the tender document, such as "total price is 5 million yuan," and compare it with the summation of all expenses in the financial statement. If the sum of all expenses in the financial statement is indeed 5 million yuan, the verification passes; otherwise, mark the text as inconsistent with the table information. This step allows for the timely detection of price mismatch issues caused by manual data entry or table errors.
[0070] The system identifies the company's qualification information declared in the tender document, such as "the company has obtained quality management system certification," and analyzes the attached qualification certificate image using OCR or image recognition technology to extract information such as the certificate name, issuing authority, and validity period. If the image recognition result is completely consistent with the text description, the verification passes; if there are discrepancies between the certificate information and the text description, such as a mismatch in certificate type or validity period, it is marked as inconsistent. This step ensures that the company's qualification declaration matches the actual certificate content.
[0071] The qualification certificate images in the tender documents are compared with standard templates stored in the knowledge graph to verify the certificate's format, seal location, and the completeness of key fields. If the certificate format is standard and no key fields are missing, the verification passes; if there are abnormal seal locations or missing key fields, it indicates a potential risk of forgery or format violations. This step further reviews the authenticity and standardization of the certificate.
[0072] Finally, hierarchical annotation was performed based on the consistency verification results of the three modalities: If the text-table consistency verification, text-image consistency verification, and image-map template consistency verification all pass, the tender document will be marked as "compliant," indicating that the tender materials meet the requirements in all aspects. If only two or one verification passes, the tender document will be marked as "partially compliant," alerting the reviewers to potential issues that require manual review. If all three verifications fail, the tender document will be marked as "non-compliant," and an automatic warning will be issued indicating that there may be significant inconsistencies or violations, requiring strict handling or direct rejection.
[0073] Through this process, the bidding party can not only improve the efficiency of the review and reduce errors in manual verification, but also effectively prevent and control problems such as false information, format violations and data inconsistencies through multimodal cross-validation, providing a scientific and quantifiable review method for the bidding management of real projects.
[0074] Through the above embodiments, this invention extracts standard reference documents from authoritative data sources and constructs a multimodal knowledge graph using natural language processing and image recognition technologies. This allows for the simultaneous understanding of the semantic logic, table structure, and layout template features of regulatory provisions, providing a unified and reasonable knowledge system for subsequent tender document review. By modal separation and structuring, text, tables, and image information in tender documents are transformed into computable structured features, achieving standardized expression of unstructured documents and significantly improving machine readability. Semantic mapping and node matching establish a one-to-one correspondence between the textual modal features of tender documents and standard clause nodes in the knowledge graph, and extract textual difference features. This enables accurate identification of semantic deviations, missing clauses, or logical conflicts in tender documents, improving the accuracy and relevance of the review results. The automatic reasoning mechanism based on knowledge graph logical rules can systematically analyze the constraint relationships, citation relationships, and contextual dependencies between tender clauses, completing automated compliance judgments and reducing the subjectivity and omission risks of manual review. Cross-modal consistency verification enables logical correspondence and information consistency detection between text, table, and image modalities. Based on the verification results, the tender documents are annotated hierarchically to form an interpretable compliance assessment result.
[0075] This invention addresses the problems of heavy reliance on manual review, low review efficiency, and difficulty in handling unstructured information and semantic ambiguity in existing tender document review by introducing multimodal knowledge graphs, semantic mapping, and logical reasoning technologies. It achieves intelligent, traceable, and interpretable automatic compliance review of tender documents, effectively improving the accuracy and reliability of the review and reducing the cost of manual review.
[0076] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
[0077] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for reviewing the compliance of tender documents based on multimodal knowledge graphs, characterized in that: Includes the following steps: Obtain standard reference documents from authoritative data sources, and use natural language processing and image recognition technologies to perform text entity extraction, visual template recognition and relationship fusion respectively, and construct a multimodal knowledge graph containing text entities, table structures and image templates; By performing modal separation and structuring on the input tender documents, the structured features of the tender documents are obtained; Structured features include text modality features, table modality features, and image modality features; Based on the text entities of the multimodal knowledge graph and the text modal features of the tender documents, semantic mapping and node matching are performed to obtain textual difference features and determine the clause nodes in the standard reference documents corresponding to the structured features. Automatic review and reasoning are performed based on the logical rules in the multimodal knowledge graph, and the compliance of the tender documents is initially judged based on the review results; For tender documents that meet the audit results, cross-modal consistency verification is performed. Based on the comprehensive verification results, the tender documents are labeled in a hierarchical manner, including "compliant", "partially compliant" or "non-compliant".
2. The method for reviewing tender document compliance based on multimodal knowledge graphs according to claim 1, characterized in that, By performing modal separation and structuring on the input tender document, the structured features of the tender document are obtained. The specific steps are as follows: Obtain the input tender document, parse and separate it into text modality, table modality and image modality; For the text modality, entity recognition, relation extraction, and semantic role labeling are performed to obtain the semantic information of the text expression as text modality features; for the table modality, structure recognition is performed to convert the table content into structured matrix features as table modality data; for the image modality, OCR text recognition and object detection are performed to extract the image expression information as image modality features.
3. The method for reviewing tender document compliance based on multimodal knowledge graphs according to claim 1, characterized in that, Based on the text entities and relationships of the multimodal knowledge graph, semantic mapping and node matching are performed with the text modal features of the tender document to obtain textual difference features and determine the clause nodes in the standard reference document corresponding to the structured features. The specific steps are as follows: Obtain the set of text entities in a multimodal knowledge graph ,in, For the a-th text entity, For the m-th text entity, The total number of text entities and the set of text modal features of the tender document. ,in, For the k-th text modal feature, For the Kth text modality feature, This represents the total number of text modal features. Based on semantic mapping, the first text representation of each text entity and the second text representation of the text modality features are obtained respectively; Semantic similarity is obtained by matching the first and second text expressions respectively, and text difference features are calculated based on semantic similarity. Based on the set of relationships in the multimodal knowledge graph ,in, This describes the relationship between text entity a and text entity b. This describes the relationship between text entity A and text entity B. The text difference features are compared with a preset difference threshold. If the text difference features are less than the preset difference threshold, a relation check and matching is performed. If the text difference features are greater than or equal to the preset difference threshold, it indicates that the corresponding text entity and text modality features are not related.
4. The method for reviewing tender document compliance based on multimodal knowledge graphs according to claim 3, characterized in that, The first textual expression includes the first textual semantic expression and the first textual structural expression; the second textual expression includes the second textual semantic expression and the second textual structural expression.
5. The method for reviewing tender document compliance based on multimodal knowledge graphs according to claim 4, characterized in that, The first text representation includes a first text semantic representation and a first text structural representation; the first text semantic representation is the main semantic vector of the text entity, and the first text structural representation is the slot structured vector of the text entity; The second text representation includes the second text semantic representation and the second text structural representation; the second text semantic representation is the main semantic vector of the text modality features, and the second text structural representation is the slot structured vector of the text modality features.
6. The method for reviewing tender document compliance based on multimodal knowledge graphs according to claim 3, characterized in that, The specific steps for performing relational verification and matching include: when the text difference features are less than the preset difference threshold, searching within the context of the corresponding text modality features, analyzing the corresponding text modality features that have the same relational constraints as the corresponding text entity, and establishing a matching relationship based on the text entity and the text modality features.
7. The method for reviewing tender document compliance based on multimodal knowledge graphs according to claim 6, characterized in that, If, after analysis, no corresponding text modal feature with the same relational constraint as the corresponding text entity is detected, it indicates that there is no corresponding clause node in the standard reference document; if a corresponding text modal feature with the same relational constraint as the corresponding text entity is detected, it indicates that there is a corresponding clause node in the standard reference document.
8. The method for reviewing tender document compliance based on multimodal knowledge graphs according to claim 1, characterized in that, For tender documents that meet the audit results, cross-modal consistency verification is performed. Based on the comprehensive verification results, the tender documents are categorized into hierarchical levels, including "compliant," "partially compliant," or "non-compliant." The specific steps are as follows: For tender documents that meet the review results, cross-modal consistency verification is performed, including text-table consistency verification, text-image consistency verification, and image-map template consistency verification. If all three modal conformity verifications are passed, the tender document will be marked as "passed"; if one or two of the three modal conformity verifications are passed, the tender document will be marked as "partially passed"; if none of the three modal conformity verifications are passed, the tender document will be marked as "unqualified".
9. The method for reviewing tender document compliance based on multimodal knowledge graphs according to claim 8, characterized in that, The text-table consistency verification verifies whether the declared values in the tender document text are consistent with the summarized values in the financial details table. If they are consistent, the text-table consistency verification is considered successful.
10. The method for reviewing tender compliance based on multimodal knowledge graphs according to claim 8, characterized in that, The text-image consistency verification verifies whether the qualifications described in the tender document text are consistent with the qualification information identified in the attached certificate image. If they are consistent, the text-image consistency verification is considered successful.
Citation Information
Cited By
Document auditing blind area detection method and device, electronic equipment and storage medium
CN121706799A
An intelligent document standard compliance automatic checking method and system
CN122242482A
A method and system for intelligent matching and recommendation of bidding information
CN122388269A