Intelligent Review Method for Eligibility and Fairness of Procurement Documents Based on Multi-Technology Integration
By employing a multi-technology integrated intelligent review method for the eligibility and fairness of procurement documents, the problem of low efficiency, poor accuracy, and insufficient comprehensiveness in existing technologies has been solved. This method enables rapid and accurate eligibility review of procurement documents, generates standardized reports, and supports the rapid advancement of the procurement process.
Patent Information
- Application Number
- CN202510785301.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-06-12
AI Technical Summary
The existing methods for qualification review of procurement documents are inefficient, inaccurate, and incomplete, and are difficult to adapt to market changes, resulting in slow procurement process and inaccurate review results.
A multi-technology integration approach, including text parsing, rule engines, knowledge graphs, natural language processing, and decision tree algorithms, is used to construct a compliance judgment model to comprehensively examine the explicit and implicit restrictions in procurement documents.
It enables comprehensive review of procurement documents, rapid identification of potential unfair terms, shortens the review cycle, improves the accuracy and efficiency of the review, generates standardized inspection reports, and supports scientific decision-making.
Smart Images

Figure CN120278685B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of internet procurement, and in particular to an intelligent method for reviewing the fairness of procurement document eligibility based on the integration of multiple technologies. Background Technology
[0002] Ensuring the fairness of eligibility clauses in procurement documents is crucial in current procurement activities. Fair eligibility clauses in procurement documents create an open and level playing field, allowing all types of suppliers to participate in bidding and promoting healthy market development. However, there are currently many issues in the review of procurement documents that urgently need to be addressed:
[0003] Traditional procurement document review methods rely primarily on manual processes: reviewers must read through the documents word by word, searching for any unfair eligibility clauses. This method is significantly inefficient. As procurement scales increase, the number and complexity of procurement documents grow, requiring reviewers to expend considerable time and effort to analyze and process them. For example, in the review of procurement documents for some large infrastructure projects, a manual review of all documents can take weeks, severely impacting the project's progress.
[0004] The accuracy of manual review is also difficult to guarantee: due to differences in the professional level, experience, and subjective judgment standards of reviewers, different reviewers may reach different conclusions regarding the same qualification clauses. Furthermore, prolonged review work can easily lead to reviewer fatigue and decreased concentration, causing them to overlook some deeply hidden unfair clauses. Some procurement documents may also contain restrictions on company size, years of operation, equity structure, ownership form, location, organizational form, etc., through vague wording or indirect connections, which may be difficult to detect during manual review. At the same time, the market environment is constantly changing, with new qualifications and licenses constantly emerging, and unfair clauses are being set up in increasingly sophisticated ways, making it difficult for manual review to adapt quickly to these changes.
[0005] Many existing automated review tools are limited in function and lack comprehensiveness. For example, some tools can only find obvious restrictive clauses through simple keyword matching, but are powerless to handle synonyms, vague expressions, and restrictions hidden behind qualification requirements. These tools do not make full use of advanced artificial intelligence and knowledge graph technologies, and cannot deeply understand the semantics and logical relationships of procurement documents, resulting in inaccurate and incomplete review results.
[0006] In the current field of qualification fairness review of procurement documents, existing technologies have many undeniable shortcomings that seriously affect the quality and efficiency of the review process. These shortcomings are mainly reflected in the following aspects:
[0007] I. Low review efficiency:
[0008] 1. Primarily Manual Review: Currently, most review work relies on manual labor. Reviewers need to manually go through every page of the procurement documents, searching word by word for any potentially unfair eligibility clauses. As the scale of procurement operations expands, the number of procurement documents increases, and the content becomes increasingly complex. Manually reviewing such a large volume of documents often requires a significant investment of time and energy, resulting in lengthy review cycles and severely slowing down the procurement process.
[0009] 2. Lack of automation capabilities: Existing technologies have limited applications of automated review tools. Even when some automation exists, it can only handle simple tasks and cannot achieve systematic and comprehensive automated review of procurement documents. For example, some simple text search tools can only locate specific keywords, but are powerless against complex semantic understanding and clause correlation analysis, leaving reviewers with limited need to manually complete most of the work, resulting in limited efficiency improvements.
[0010] II. The accuracy of the review is not ideal:
[0011] 1. Subjective Judgment Differences: During manual review, different reviewers may have significantly different judgment standards for the same eligibility clause due to differences in professional background, experience level, and personal understanding. For example, for some vaguely worded clauses, one reviewer may believe that they do not constitute a restriction, while another reviewer may hold the opposite view. This inconsistency in subjective judgment leads to a lack of objectivity and accuracy in the review results.
[0012] 2. Difficulty in Identifying Complex Clauses: Unfair clauses in procurement documents are becoming increasingly complex and covert, employing techniques such as ambiguous semantics, indirect associations, and synonym substitutions. During manual review, reviewers may overlook these deeply hidden unfair clauses due to fatigue, negligence, or misunderstanding of complex language structures. Furthermore, manual reviewers struggle to keep up with and respond to constantly evolving clause formats, further reducing the accuracy of the review process.
[0013] III. Insufficient comprehensiveness of the review:
[0014] 1. Single-Dimensional Review: Existing review techniques mostly focus only on the surface text of procurement documents, lacking in-depth exploration of potential restrictions hidden behind qualification requirements. For example, they only review whether the documents directly mention enterprise size requirements / years of operation / equity structure / ownership form / location / organizational form, while ignoring the restrictions on enterprise size requirements / years of operation / equity structure / ownership form / location / organizational form in the qualification certificate acquisition conditions. This single-dimensional review method cannot comprehensively discover unfair issues in procurement documents, and some suppliers may be unreasonably excluded from procurement activities due to implicit restrictions.
[0015] 2. Lack of multi-technology integration: Most review tools and methods only focus on the application of a single technology, such as simple text retrieval technology or simple rule matching technology. They fail to fully integrate multiple advanced technologies such as natural language processing and knowledge graphs, making it difficult to conduct comprehensive analysis of procurement documents from multiple perspectives, resulting in loopholes and blind spots in the review process.
[0016] IV. Poor adaptability and scalability:
[0017] 1. Difficulty in Adapting to Change: The market environment is constantly changing, and the form and content of procurement documents are also evolving, with new unfair clauses emerging one after another. However, existing technologies often lack effective self-learning and updating mechanisms, making it difficult to quickly adapt to these changes. For example, when new qualification requirements or clause wording appear, traditional review techniques may fail to identify and analyze them in a timely manner, leading to delays in the review process.
[0018] 2. Weak system scalability: Existing review systems or methods lack forward-looking design and have relatively closed architectures, making it difficult to easily integrate new technologies or modules to improve review capabilities. For example, when more advanced natural language processing models emerge, they cannot be easily integrated into existing systems, limiting the further development and optimization of review technologies.
[0019] In conclusion, existing procurement document review methods cannot meet the current needs for efficient, accurate, and comprehensive review. There is an urgent need for a new technology to solve these problems and ensure the fairness of the eligibility clauses in procurement documents. Summary of the Invention
[0020] The purpose of this invention is to provide an intelligent review method for the fairness of procurement document eligibility based on the integration of multiple technologies, which can ensure the fairness of the eligibility terms in procurement documents.
[0021] This invention is achieved through the following technical solution:
[0022] The intelligent review method for the eligibility and fairness of procurement documents based on multi-technology integration includes the following steps:
[0023] Step 1: Preprocess the existing procurement documents; use a text parsing algorithm to process the existing procurement documents to obtain the procurement text.
[0024] Step 2: First, define a keyword library for enterprise size restrictions; then, build a rule engine to search for words in the keyword library within the procurement text, match the numerical threshold immediately following the word, and determine whether the numerical threshold exceeds the specified limit to obtain the judgment result; finally, merge the keywords and their numerical thresholds as explicit restriction conditions, and save the explicit restriction conditions and their judgment results as explicit restriction features.
[0025] Step 3: First, use the graph database tool Neo4j to construct a knowledge graph centered on qualification certificates; then, use natural language processing methods to identify qualification entities in the procurement text, and query the knowledge graph through the qualification entities. When the knowledge graph finds a node with enterprise size restrictions for the qualification entity, it is determined that there is a hidden compliance risk; finally, the query path and judgment results of the qualification entity in the knowledge graph are saved as implicit restriction features.
[0026] Step 4: Combine the explicit and implicit restriction features of an existing procurement document to form its procurement data; combine the procurement data of each existing procurement document to form a compliance judgment dataset; use the compliance judgment dataset to train a model based on the decision tree algorithm to obtain the compliance judgment model.
[0027] Step 5: Analyze the procurement documents to be judged using the compliance judgment model, generate an inspection report of the procurement documents to be judged, and enter the decision-making process of the compliance judgment model for the procurement documents to be judged and the final judgment result in the inspection report.
[0028] Compared with previous technologies, the beneficial effects of the present invention are as follows:
[0029] 1. This invention takes into account both explicit and implicit restrictions, comprehensively reviewing the qualification clauses of procurement documents from the surface description of the text to the deep association with qualifications. It can not only discover obvious restrictions that directly reflect the provisions of the clauses, but also find the enterprise size restrictions hidden behind industry qualification certificates through knowledge graph reasoning, ensuring that the review is thorough and assisting in the fairness review, fully protecting the fairness of market access for market entities, and preventing the unreasonable exclusion of some suppliers due to loopholes in the clauses.
[0030] 2. This invention can utilize multiple natural language processing methods to coordinate the processing of procurement documents, enabling the processing of large amounts of text in a short time. The inspection report displays important information such as basic project information, an overview of inspection results, explicit and implicit restrictions, significantly reducing the amount of text review required. In reviewing a large procurement document of hundreds of pages, manual review may take weeks, while the technical means provided by this invention can complete the preliminary review in just a few hours, greatly shortening the review cycle, accelerating the procurement process, facilitating reviewers' access to key points, and enabling projects to enter the bidding and implementation stage more quickly.
[0031] 3. The standardized inspection report generated by the compliance judgment model in this invention clearly presents important information such as the basic information of the procurement document project, an overview of the inspection results, explicit restrictions, and implicit restrictions. Bidding entities and regulatory departments can quickly understand the fair competition compliance status of the documents without professional technical knowledge, and make targeted adjustments to the procurement documents based on the report content to make scientific decisions. The standardization and professionalism of the report also provide strong support for regulatory work and improve the overall management level of procurement activities.
[0032] 4. This invention also utilizes the BERT model to expand the understanding capability, enabling the identification of statements and paragraphs in procurement documents that set enterprise size restrictions through fuzzy descriptions or indirect associations, and saving them as compliance risk data to be displayed in the inspection report for easy review and analysis by auditors. Attached Figure Description
[0033] Figure 1 This is a flowchart of the present invention;
[0034] Figure 2 Pseudocode for constructing regular expressions;
[0035] Figure 3 Pseudocode for a function that determines a digital threshold based on defined logical rules;
[0036] Figure 4 Pseudocode for functions used to build the main program of the rules engine. Detailed Implementation
[0037] The present invention will now be described in detail with reference to the accompanying drawings, but the scope of protection of the present invention is not limited to the following description:
[0038] like Figure 1-4 The diagram shown is an embodiment of an intelligent review method for the eligibility and fairness of procurement documents based on the integration of multiple technologies provided by the present invention.
[0039] The intelligent review method for the eligibility and fairness of procurement documents based on multi-technology integration includes the following steps:
[0040] Step 1: Preprocess the existing procurement documents; use a text parsing algorithm to process the existing procurement documents to obtain the procurement text.
[0041] Step 2: First, define a keyword library for enterprise size restrictions; then, build a rule engine to search for words in the keyword library within the procurement text, match the numerical threshold immediately following the word, and determine whether the numerical threshold exceeds the specified limit to obtain the judgment result; finally, merge the keywords and their numerical thresholds as explicit restriction conditions, and save the explicit restriction conditions and their judgment results as explicit restriction features.
[0042] Step 3: First, use the graph database tool Neo4j to construct a knowledge graph centered on qualification certificates; then, use natural language processing methods to identify qualification entities in the procurement text, and query the knowledge graph through the qualification entities. When the knowledge graph finds a node with enterprise size restrictions for the qualification entity, it is determined that there is a hidden compliance risk; finally, the query path and judgment results of the qualification entity in the knowledge graph are saved as implicit restriction features.
[0043] Step 4: Combine the explicit and implicit restriction features of an existing procurement document to form its procurement data; combine the procurement data of each existing procurement document to form a compliance judgment dataset; use the compliance judgment dataset to train a model based on the decision tree algorithm to obtain the compliance judgment model.
[0044] Step 5: Analyze the procurement documents to be judged using the compliance judgment model, generate an inspection report of the procurement documents to be judged, and enter the decision-making process of the compliance judgment model for the procurement documents to be judged and the final judgment result in the inspection report.
[0045] It should be noted that this invention examines both explicit and implicit restrictions, comprehensively reviewing the qualification clauses of procurement documents from the surface description of the text to the deep-seated connections to qualifications. It not only uncovers obvious restrictions directly reflected in the clauses but also uses knowledge graph reasoning to identify hidden restrictions on company size behind industry qualification certificates, ensuring a thorough review, assisting in fairness review, and comprehensively protecting the fairness of market access for all market participants. This prevents the unreasonable exclusion of some suppliers due to loopholes in the clauses. Utilizing multiple natural language processing methods to coordinate the processing of procurement documents, it can process large amounts of text in a short time. The inspection report displays important information such as basic project information, an overview of the inspection results, explicit restrictions, and implicit restrictions, significantly reducing the amount of text review required. In reviewing a large procurement document of hundreds of pages, manual review may take weeks, while the technical means provided by this invention can complete the preliminary review in just a few hours, greatly shortening the review cycle, accelerating the procurement process, facilitating reviewers' access to key points, and enabling projects to enter the bidding and implementation stage more quickly. The standardized inspection report generated by the compliance assessment model clearly presents important information such as basic project information, inspection results overview, explicit and implicit restrictions in the procurement documents. Bidding entities and regulatory authorities can quickly understand the fair competition compliance status of the documents without professional technical knowledge, and make targeted adjustments to the procurement documents based on the report content to make scientific decisions. The standardization and professionalism of the report also provide strong support for regulatory work and improve the overall management level of procurement activities.
[0046] Furthermore, step 1 also includes the following steps:
[0047] Step 101: Perform text preprocessing on the existing procurement documents to remove redundant symbols, special symbols, and spaces.
[0048] Step 102: First, use the n-gram model to perform grammatical and semantic checks on the existing procurement documents; then, use a text segmentation algorithm to divide the long text into appropriate paragraphs; and finally, use a title recognition algorithm to identify chapter titles or clause numbers based on font format, layout features, and text content information to obtain the existing procurement text.
[0049] It's important to note that the n-gram model is a commonly used language model in large-vocabulary continuous speech recognition. It determines whether text conforms to linguistic conventions by statistically analyzing the frequency of n adjacent words and corrects abnormal word sequences. Take a simple procurement document text, "The supplier should provide high-quality products," as an example: Assuming we use a 3-gram model (n=3), this model will sequentially divide the text into a sequence of words of length 3 (also called n-gram). First, the text is segmented into words, resulting in "supplier," "should," "provide," "high-quality," "of," and "products." Then, a 3-gram sequence is generated, starting from the first word:
[0050] "Suppliers should provide";
[0051] "High quantity and quality should be provided";
[0052] "Provide high-quality and high-quantity products";
[0053] "High-quality products";
[0054] For each 3-gram, the model counts the frequency of the sequence in its training data. In a large amount of normal procurement document text training data, 3-grams such as "suppliers should provide" and "provided products" may have a high frequency of occurrence because they conform to common language expression habits. However, the word sequences containing "high quality" in "should provide high quantity and quality" and "provide high quantity and quality" are not in normal expression habits and have a very low frequency of occurrence in the training data. Based on this, the n-gram model judges "high quantity and quality" as an anomalous word sequence. When the model corrects itself, it proposes correction suggestions based on high-frequency word sequences in similar contexts in the training data. For example, in a context like "provide... products", the high-frequency word might be "high quality", so the model will suggest correcting "high quantity and quality" to "high quality" to make the text "suppliers should provide high-quality products" conform to normal language expression habits, thereby completing the processing of the procurement document text and providing a more reliable textual basis for subsequent accurate analysis of the procurement document content.
[0055] Text segmentation algorithms can divide long texts into appropriate paragraphs based on punctuation, paragraph markers, and semantic coherence. Title recognition algorithms can identify chapter titles or clause numbers based on font format, layout features, and text content. For example, for text formatted as "1. Qualification Requirements," it can be identified as a chapter title by matching numbers and specific title keywords. For extremely long texts, a sliding window-based chunking method is used, dividing the text into blocks of a certain length. The text within each window is treated as an independent analysis unit, improving the efficiency of subsequent NLP analysis.
[0056] Further, in step 2, the following steps are also included:
[0057] Step 201, first define a keyword library for enterprise scale restrictions. The keyword library includes registered capital, paid-in capital, number of employees, tax payment amount, business operation years, location, and text clues related to education level and qualifications;
[0058] Step 202, then construct a rule engine based on regular expressions and logical rules; use regular expressions to search for the vocabulary in the keyword library within the procurement text; use logical rules to match the numeric threshold immediately following the keyword and determine whether the numeric threshold exceeds the specified limit of the logical rules:
[0059] Step 203, obtain the judgment result: when the numeric threshold exceeds the limit set by the logical rules, set the judgment result as non-compliant; otherwise, set the judgment result as compliant;
[0060] Step 204, finally, combine the keyword and its numeric threshold and set it as an explicit restriction condition; save the explicit restriction condition and its judgment result as an explicit restriction feature.
[0061] Specific pseudocode and example explanation for constructing a rule engine based on regular expressions and logical rules:
[0062] 1. Define the keyword library:
[0063] Here, a keyword library keyword_list = ["registered capital", "annual turnover", "tax payment amount"] with three words is used as an example;
[0064] 2. Construct regular expressions:
[0065] Such as Figure 2 The following is the function pseudocode for constructing regular expressions. The meaning of the regular expression regex constructed here is: first match the keyword keyword, '\s*' means match 0 or more whitespace characters (including spaces, tabs, etc.), '[::]?' means match 0 or 1 colon (Chinese or English colon), '\s*' is used to match whitespace characters again, '(\d+)' is used to capture the number immediately following the whitespace character, '\s*' continues to match whitespace characters, '(万|元|%)?' means match 0 or 1 unit (ten thousand, yuan, or %);
[0066] 3. Define logical rules to judge the numeric threshold:
[0067] Such as Figure 3The image shows the pseudocode of a function that defines logical rules for judging numerical thresholds. The function `check_threshold` determines whether the set threshold conditions are met based on different keywords and their corresponding numbers and units. Taking "registered capital" as an example, if the unit is "ten thousand" and the number is less than 100, or the unit is "yuan" and the number is less than 1,000,000, then the condition is considered not met and `False` is returned; otherwise, `True` is returned.
[0068] 4. Build the main program of the rules engine
[0069] like Figure 4 The image shows the pseudocode of the main program for building the rules engine. The function iterates through the keyword library `keyword_list` and its corresponding regular expressions, and uses the `find All Matches` function (assuming this function is used to find all content that matches the regular expression in the text) to find matching items `matches` in the text of the procurement document. For each matching item `matches`, its keywords, numbers, and units are extracted, and the `check_threshold` function is called to determine whether the numerical threshold meets the requirements. If it does not meet the requirements, a prompt message is output.
[0070] Furthermore, step 3 also includes the following steps:
[0071] Step 301: Use the graph database tool Neo4j to construct a knowledge graph centered on qualification certificates. This graph uses common domestic qualification certificates as core nodes and sets enterprise size-restricted nodes and non-enterprise size-restricted nodes.
[0072] Step 302: Based on the acquisition conditions of each qualification certificate, the corresponding acquisition conditions are associated with the corresponding enterprise size restriction nodes and non-enterprise size restriction nodes through edge relationships;
[0073] Step 303: Then, the keyword matching algorithm is used to extract the sentences with qualification requirements from the procurement text, and then the named entity recognition technology is used to identify the qualification entities in the sentences.
[0074] Step 304: Use a path search algorithm to find the path from the qualification certificate node corresponding to the qualification entity to the enterprise size restriction node in the knowledge graph. If a path exists, it is determined that there is a hidden compliance risk, and the judgment result is set as non-compliant; otherwise, the judgment result is set as compliant.
[0075] Step 305: Finally, the query path and judgment result of the qualified entity in the knowledge graph are saved as implicit restriction features.
[0076] It should be noted that in steps 301 and 302, the Neo4j graph database links qualification certificates and enterprise size restriction nodes through data collection, organization, and import operations. For example, for the "Contract Abiding and Creditworthy Certificate" node, its attributes include information on the required years of enterprise operation. In the knowledge graph, the relationship between nodes can be represented as "has_requirement," etc., to clarify the association between qualifications and restrictions such as enterprise size, years of operation, equity structure, ownership form, location, and organizational form.
[0077] In step 303, a keyword matching algorithm is used to extract statements with qualification requirements from the procurement documents, such as statements mentioning "the enterprise needs to have a corporate credit rating certificate," "the enterprise needs to have a CMM / CMMI certification certificate," or "the enterprise needs to have a CCRC information security service qualification certificate." Then, Named Entity Recognition (NER) technology from Natural Language Processing (NLP) is used to identify the qualification entities in the qualification requirement keywords, namely, corporate credit rating certificate, CMM / CMMI certification certificate, or CCRC information security service qualification certificate. Specifically, a feature function can be defined using an NER algorithm based on a Conditional Random Field (CRF) model. ,in and Here, x is the label of adjacent positions, i is the input text, and i is the position index. The optimal label sequence for each position under the given text x is calculated to identify qualified entities. After identifying a qualified entity, the knowledge graph is queried to determine whether the qualified entity contains requirements such as enterprise size, years of operation, equity structure, ownership form, location, and organizational form.
[0078] In more advanced scenarios, the reasoning capabilities of knowledge graphs are leveraged to perform inferences based on graph structure and node relationships. For example, a path search algorithm can be used to find a path from the "qualification requirements" node to the "enterprise size restriction" node in the knowledge graph. If a path exists, it indicates a potential compliance risk. Specific path search algorithms can employ breadth-first search (BFS) or depth-first search (DFS). Taking breadth-first search as an example, starting from the "qualification requirements" node, the search expands layer by layer until the "enterprise size restriction" node is found or all reachable nodes have been traversed.
[0079] In steps 301 and 302, the construction and application of the qualification knowledge graph includes the following processes:
[0080] 3021) will use a triplet framework for resource description within a knowledge graph centered on qualification certificates. Vector embedding representation, where resource description framework triples The corresponding word sequence is set as Y= ;
[0081] here , , A sequence of words representing the subject (such as a specific qualification category), the predicate (such as a relational expression like "possessing...conditions"), and the object (such as expressions related to limiting conditions such as the lower limit of enterprise size and the threshold of years of operation);
[0082] 3022) Define function n ;in, Used to represent the sequence of entity or relation words corresponding to n The i-th word in the definition; word embedding mapping (in, express The j-th subword is calculated as a vector product.
[0083] These steps enable the Resource Description Framework triplet to be formed. The word sequence Y is embedded into a word vector tensor space and represented as a function:
[0084] Where J represents the word vector tensor, , , It is a word embedding map Mapped , , The amount;
[0085] 3023) Determine the input of the first-layer graph neural network (GNN) block at time t: Define the position offset. As a position-related parameter of a word vector in the graph neural network model at time t;
[0086] 3024) Define the eigenvector This is a feature vector in the feature vector sequence generated by the graphical neural network model at time t-1;
[0087] (here This represents a specific vector fusion operation, such as bitwise addition or concatenation, to obtain the input of the first layer graph neural network block at time t;
[0088] 3025) Define the number of layers in the graph neural network model; define each layer of the graph neural network model as consisting of multiple graph neural network blocks, and perform the following calculations in each graph neural network block:
[0089] ;
[0090] ;
[0091] ;
[0092] ;
[0093] Here, GraphAttention represents the graph attention function, used to capture the relationship information between nodes in the graph structure; GraphNorm represents the graph normalization function, which normalizes the graph data to accelerate convergence and improve stability; GraphFF represents the graph feedforward function, used to perform nonlinear transformations on features; through this series of operations, the deep information contained in the nodes and relationships in the knowledge graph is mined.
[0094] Furthermore, step 4 also includes the following steps:
[0095] Step 401: Explicit restrictions in an existing procurement document and implicit limiting features This constitutes the combination of the procurement documents and procurement data. ;
[0096] Among the features For the explicit or implicit restrictive features of this existing procurement document, Features The tag and ;
[0097] Step 402: Combine the procurement data from each existing procurement document. To form a compliance judgment dataset Where p is the number of existing procurement documents, For the labels of existing procurement documents, and ;
[0098] Step 403: Train a model based on a decision tree algorithm using the compliance judgment dataset E, with procurement data as the model. Features For internal nodes, tags Decisions are made based on the edge relationships of internal nodes; finally, the labels in the existing procurement documents are used. The compliance judgment model is obtained by ending at a leaf node.
[0099] It's important to note that a decision tree is an algorithm that makes decisions based on a tree structure. In this compliance judgment model, the decision tree consists of nodes and edges. Nodes are divided into internal nodes and leaf nodes. Internal nodes represent tests on an attribute, edges represent test outputs, and leaf nodes represent the decision result. For example, in a procurement document, explicit restrictive features might include "whether there are registered capital restrictions" or "whether there are operating period restrictions," while implicit restrictive features might include "potential equity structure restrictions identified through knowledge graph analysis" or "potential operating period restrictions." These explicit and implicit restrictive features together constitute the internal nodes of the decision tree. If the edge relationship is "compliant," the decision tree continues along the corresponding edge to the next node; if it's "non-compliant," the decision tree continues along another edge to the next node. Finally, the decision tree reaches a leaf node, resulting in a "fair procurement" or "unfair procurement" decision, thus becoming a decision branch of the decision tree.
[0100] It should be noted that the training process of the decision tree in step 403 mainly involves constructing the decision tree by continuously selecting the optimal features and split points; at each internal node, metrics such as information gain or Gini impurity are used to select the optimal feature for splitting; taking information gain as an example, information gain... Measured usage characteristics The increase in information obtained by segmenting the compliance judgment dataset E; The calculation formula is:
[0101] ;
[0102] in is the information entropy of the compliance judgment dataset E, used to measure the uncertainty of the compliance judgment dataset E, and its calculation formula is:
[0103] ;
[0104] in, It represents all possible values of feature A, i.e., A = "fair procurement" or "unfair procurement"; This is the number of labels (i.e., "fair procurement" and "unfair procurement") in the compliance judgment dataset E, i.e., C=2; It is a subset of the compliance judgment dataset E when the label is u, that is The number of procurement documents for "fair procurement" or "unfair procurement"; A subset of the compliance assessment dataset E;
[0105] By selecting information gain The largest feature is used as the segmentation feature of the current node. Then, the above operation is recursively performed on the segmented child nodes until the stopping condition is met (such as reaching the maximum depth D or the number of node samples is less than the minimum number of sample segments S), thus obtaining the compliance judgment model.
[0106] Furthermore, in step 403, after obtaining the compliance judgment model, a test dataset containing M procurement documents will be used. ; for each feature vector in the test dataset The data is input into the trained compliance judgment model, where... It includes both explicit and implicit limiting features of the procurement documents used for testing. This document is labeled as either "fair procurement" or "unfair procurement." The compliance assessment model starts from the root node and traverses the decision tree downwards based on the test conditions of each node until it reaches a leaf node to obtain the prediction result. ;
[0107] By comparing prediction results and real labels The accuracy and recall of the model are calculated to measure its performance.
[0108] Generally, accuracy is the most intuitive performance indicator, and its calculation formula is as follows:
[0109] ;in It is an indicator function, if equal If the value is 1, the prediction is correct; otherwise, it is 0, indicating an incorrect prediction.
[0110] It should be noted that after obtaining the judgment result through the compliance judgment model, an inspection report is generated according to a standardized template. The report includes basic project information (such as project name, project number, etc.), which is extracted from the procurement documents as the content of the corresponding sections of the report. The inspection result summary section directly fills in the "fair procurement" or "unfair procurement" conclusion output by the model; and lists all explicit and implicit restrictive features found in the explicit and implicit inspections, while outputting the "compliant" and "non-compliant" corresponding edge relationships. Finally, the generated inspection report is presented in PDF or Word format for easy access by the tendering party, centralized procurement agency, procurement supervision department, and social agency.
[0111] Furthermore, step 4 also includes the following steps:
[0112] Step 404 introduces the reinforcement learning (RLHF) method based on human feedback. Professional reviewers manually review the inspection reports generated by the compliance assessment model, evaluating and marking the summary of inspection results, the results of explicit restrictions, and the results of implicit restrictions in the inspection reports. For example, if the compliance assessment model determines that a procurement document is compliant, but the reviewer finds implicit restriction risks that were not identified by the model, the report will be marked as "misjudgment," and the reasons for the misjudgment will be noted in detail.
[0113] Step 405 involves inputting the evaluation labels as key reward signals into the compliance judgment model in batches of a specific size for retraining. The purpose of setting the batch size is to balance training efficiency and model stability, avoiding oscillations in the model learning process caused by inputting too much data at once, while also preventing the model from learning too slowly due to excessively small batches. For example, every 10 evaluation labels can be set as one batch input to the model.
[0114] Step 406: The compliance judgment model calculates the policy gradient based on the reward signal. The policy gradient represents the direction of improvement of the model's current policy (i.e., decision logic) under the guidance of the reward signal. The model parameters are updated through the Adam optimizer to optimize the decision logic, thereby improving the accuracy of the compliance judgment model in judging the compliance of the qualification clauses of procurement documents. This allows the compliance judgment model to better adapt to various complex procurement document scenarios. If the model's previous judgment leads the reviewer to give a positive evaluation mark (such as correctly judging the procurement document as compliant), the policy gradient will be adjusted in the direction of strengthening that judgment logic; conversely, if a negative evaluation mark is obtained (such as misjudgment), the policy gradient will prompt the model to correct its decision logic.
[0115] It should be noted that the use of reinforcement learning based on human feedback (RLHF) is a secondary technical means in this invention, not a necessary technical feature. This method allows professional reviewers (procurement agencies, procurement regulatory departments, and social agencies) to manually review the inspection reports generated by the compliance assessment model. Reviewers can use their professional knowledge and experience to judge the overview of inspection results and detailed risk points in the inspection report, identify and mark any inaccuracies or incompleteness in their judgments, and provide explanations.
[0116] These judgments, markings, and terminations then serve as reward signals for human feedback, feeding them back into the compliance assessment model for retraining. For example, if the model misjudges a clause as compliant while a human reviewer deems it non-compliant, the model receives a negative reward signal; conversely, if the model correctly judges the clause, it receives a positive reward signal.
[0117] Furthermore, it also includes the following steps:
[0118] Step 601: Use text parsing technology to locate and extract qualification clauses in existing procurement documents; then use dependency parsing technology to analyze the dependency relationships in the qualification clauses, identify and save the sentences and paragraphs related to qualification, and obtain the qualification dataset.
[0119] Step 602: Label the eligibility dataset, labeling sentences and paragraphs with enterprise size restrictions as risky and other sentences and paragraphs as not risky; then feed this dataset into the pre-trained BERT model for fine-tuning, so that the BERT model can understand sentences or paragraphs with enterprise size restrictions.
[0120] Step 603: Use the BERT model to perform semantic matching on the procurement documents to be judged, and extract all statements and paragraphs that are judged to have compliance risks from the procurement documents to be judged as compliance risk data.
[0121] Step 604: Finally, according to the order in which the compliance risk data appears in the procurement documents to be judged, the compliance risk data is entered into the inspection report in sequence to obtain an inspection report with compliance risk analysis capabilities.
[0122] It should be noted that in step 601, the text parsing techniques include commonly used regular expression matching, named entity recognition (NER), relation extraction, and dependency parsing techniques, etc.
[0123] In dependency parsing techniques, it is assumed that dependency relations are represented as follows: ,in and These are two words in a sentence. By analyzing the dependencies in the procurement documents one by one, key phrases such as "the supplier should have... qualifications" can be accurately extracted to obtain a qualification dataset.
[0124] In step 602, BERT stands for Bidirectional Encoder Representations from Transformers. BERT is a pre-trained language model based on the Transformer architecture. Its core innovation lies in its ability to pre-train deep bidirectional representations based on the left and right context of all layers. This allows the model to consider both preceding and following information when processing text, thus capturing the semantic features of the text more comprehensively and accurately. Unlike traditional unidirectional language models (such as those that can only process text from left to right or right to left), BERT's bidirectional processing capability greatly enhances its ability to understand complex semantics. The BERT model is composed of multiple stacked Transformer blocks. Each Transformer block mainly contains two parts: a multi-head attention mechanism and a feed-forward neural network. The multi-head attention mechanism, by computing multiple different attention heads in parallel, can capture semantic relationships in the text from different angles, further enriching the model's understanding of the text. Feedforward neural networks perform nonlinear transformations on the output of the self-attention mechanism to learn more complex feature representations;
[0125] Specifically, suppose a qualification clause in the procurement document states: "This procurement requires suppliers to be state-owned enterprises that have been registered and operating locally for more than 5 years to participate." This invention can accurately identify restrictions imposed through synonym substitution, ambiguous semantics, and implicit restrictions linked to certificates. This significantly reduces the probability of misjudgments and omissions, providing a more reliable assessment of the fair competition in procurement documents.
[0126] First, the text is preprocessed, such as by word segmentation and stop word removal, to transform it into a format suitable for the BERT model. The preprocessed text is then input into the finely tuned BERT model, where the text is encoded using a multi-layer self-attention mechanism. Taking the self-attention calculation of one layer as an example, assuming that the current processing is of the text content "registered and operated locally for more than 5 years", the model will generate a corresponding query matrix Q, key matrix K, and value matrix V for each word.
[0127] For example, for the word "local," the BERT model calculates attention using self-attention, and the formula for attention is:
[0128] ;
[0129] in K and V represent the query matrix, key matrix, and value matrix, respectively. The dimension of the key matrix is used to calculate the degree of association between each position in the text and other positions, thereby understanding the semantics of the text. The dimension is used to calculate the degree of association between the position "local" and other positions in the text (such as the positions corresponding to words like "registered," "operating," and "5 years"). In this way, the model can understand the semantic relationships between words in the text, and thus grasp the semantics of the entire sentence. Because the BERT model learns a large amount of general language knowledge during the pre-training stage and has been fine-tuned on datasets containing both restrictive and non-restrictive clauses such as company size, years of operation, equity structure, ownership form, location, and organizational form, it can determine whether the text contains restrictive clauses such as company size, years of operation, equity structure, ownership form, location, and organizational form.
[0130] Therefore, in the above example, the BERT model can identify that "registered and operating locally for more than 5 years" includes the restriction of operating years, and "state-owned enterprise" includes the restriction of ownership form. Based on its learned knowledge and fine-tuned parameters, the BERT model scores or labels the paragraph. If a high score is given, it indicates that the text is likely to contain explicit restrictions, or it can be directly labeled "contains restrictions on operating years and ownership form," thus helping reviewers quickly determine whether there are compliance risks in statements or paragraphs in procurement documents.
[0131] It should be noted that the BERT model analysis of procurement texts is a subsequent operation of the compliance judgment model. To prevent the BERT model from containing explicit or implicit restrictive features, while the procurement document to be judged is being screened for explicit and implicit restrictive features, statements that are saved as explicit or implicit restrictive features can be deleted from the procurement text. Finally, after the compliance judgment model obtains its judgment result, the procurement text is then input back into the BERT model analysis process to avoid the repeated display of the same content.
[0132] It should be noted that, taking a procurement document with the project name "Procurement Project of a Certain Information System" and the project number "2025-001" as an example, the procurement document is first processed into a procurement text.
[0133] When processing procurement documents using the method of this invention, if clauses such as "the supplier's registered capital must reach 5 million yuan or more" are found, they are saved as several explicit restriction features according to the order in which the clauses appear. ;
[0134] When the procurement document contains qualification entities such as "CMM / CMMI certification certificate", the knowledge graph is used to query whether the acquisition conditions of the certificate have a tendency to restrict the enterprise size node; the qualification entities are then saved as several implicit restriction features according to their order of appearance. ;
[0135] These features are then used to construct the feature vector of the procurement data in the procurement document. The input is fed into the compliance judgment model for decision-making. It should be noted that within the compliance judgment model, the internal nodes of each branch have already been pre-defined during training to determine the order of explicit and implicit constraints; feature vectors... The system will follow the compliance judgment model to make judgments and classifications from the root node to each internal node to the final leaf node, and will sequentially input explicit and implicit restriction features. Judgments and classifications will be made at each internal node. Through layer-by-layer screening, the system will draw a conclusion of "fair procurement" or "unfair procurement" within the leaf node.
[0136] Finally, in the generated inspection report, the basic information section should include the project name "Procurement Project of a Certain Information System" and the project number "2025-001". The inspection results should be summarized as "fair procurement" or "unfair procurement". Then, all explicit and implicit restrictions should be listed in detail, as well as the judgments (i.e., edge relationships) within the decision tree regarding whether the explicit and implicit restrictions are compliant or not.
[0137] Optionally, before generating the inspection report, the procurement text can be input into the BERT model for analysis. The BERT model, based on its extended semantic capabilities, can save risky statements and paragraphs within the procurement text as compliance risk data. When necessary, it can display the compliance risk data and its probability for reviewers to analyze, preventing unscrupulous parties from using unclear wording in procurement documents to benefit specific companies, thus further ensuring fairness. After analysis using the BERT model, the inspection report will also display the risk allocation data obtained from the BERT model analysis. At this point, the inspection report includes basic project information, an overview of the inspection results, explicit and implicit restrictive features, compliance judgments of explicit and implicit restrictive features, and compliance risk data.
[0138] Although the present invention has been illustrated and described through specific embodiments and alternative methods, it should be understood that various changes and modifications may be made without departing from the spirit and scope of the invention. Therefore, it should be understood that the present invention is not limited in any sense except by the appended claims and their equivalents.
Claims
1. A method for intelligent review of the eligibility and fairness of procurement documents based on multi-technology integration, characterized in that: Includes the following steps: Step 1: Preprocess the existing procurement documents; use a text parsing algorithm to process the existing procurement documents to obtain the procurement text. Step 2: First, define a keyword library related to enterprise size restrictions; Then, a rule engine is built, which searches for words in the keyword library within the procurement text, matches the numerical threshold immediately following the word, and determines whether the numerical threshold exceeds the specified limit to obtain the judgment result; finally, the keywords and their numerical thresholds are merged and set as explicit restriction conditions, and the explicit restriction conditions and their judgment results are saved as explicit restriction features. Step 3: First, use the graph database tool Neo4j to construct a knowledge graph centered on qualification certificates; then, use natural language processing methods to identify qualification entities in the procurement text, and query the knowledge graph through the qualification entities. When the knowledge graph finds a node with enterprise size restrictions for the qualification entity, it is determined that there is a hidden compliance risk; finally, the query path and judgment results of the qualification entity in the knowledge graph are saved as implicit restriction features. Step 3 also includes the following steps: Step 301: Use the graph database tool Neo4j to construct a knowledge graph centered on qualification certificates. This graph uses common domestic qualification certificates as core nodes and sets enterprise size-restricted nodes and non-enterprise size-restricted nodes. Step 302: Based on the acquisition conditions of each qualification certificate, the corresponding acquisition conditions are associated with the corresponding enterprise size restriction nodes and non-enterprise size restriction nodes through edge relationships. Step 303: Then, the keyword matching algorithm is used to extract the sentences with qualification requirements from the procurement text, and then the named entity recognition technology is used to identify the qualification entities in the sentences. Step 304: Use a path search algorithm to find the path from the qualification certificate node corresponding to the qualification entity to the enterprise size restriction node in the knowledge graph. If a path exists, it is determined that there is a hidden compliance risk, and the judgment result is set as non-compliant; otherwise, the judgment result is set as compliant. Step 305: Finally, the query path of the qualified entity in the knowledge graph and its judgment result are saved as implicit restriction features. Step 4: Combine the explicit and implicit restriction features of an existing procurement document to form its procurement data; combine the procurement data of each existing procurement document to form a compliance judgment dataset; use the compliance judgment dataset to train a model based on the decision tree algorithm to obtain the compliance judgment model. Step 5: Analyze the procurement documents to be judged using the compliance judgment model, generate an inspection report of the procurement documents to be judged, and enter the decision-making process of the compliance judgment model for the procurement documents to be judged and the final judgment result in the inspection report.
2. The intelligent review method for the fairness of procurement document eligibility based on multi-technology integration as described in claim 1, characterized in that, Step 1 also includes the following steps: Step 101: Perform text preprocessing on the existing procurement documents to remove redundant symbols, special symbols, and spaces. Step 102: First, use the n-gram model to perform grammatical and semantic checks on the existing procurement documents; then, use a text segmentation algorithm to divide the long text into appropriate paragraphs; and finally, use a title recognition algorithm to identify chapter titles or clause numbers based on font format, layout features, and text content information to obtain the existing procurement text.
3. The intelligent review method for the fairness of procurement document eligibility based on multi-technology integration as described in claim 2, characterized in that, Step 2 also includes the following steps: Step 201: First, define a keyword library for enterprise size restrictions. The keyword library includes registered capital, paid-in capital, number of employees, tax payment, years of operation, location, and textual clues related to education and qualifications. Step 202: Then, a rule engine is built based on regular expressions and logical rules; regular expressions are used to search for words in the keyword library within the procurement text; logical rules are used to match the numerical threshold immediately following the keyword and determine whether the numerical threshold exceeds the limit specified by the logical rules: Step 203, obtain the judgment result: when the numerical threshold exceeds the limit set by the logic rule, the judgment result is set as non-compliant; otherwise, the judgment result is set as compliant. Step 204: Finally, combine the keywords and their numerical thresholds and set them as explicit restrictions; save the explicit restrictions and their judgment results as explicit restriction features.
4. The intelligent review method for the fairness of procurement document eligibility based on multi-technology integration as described in claim 3, characterized in that, Step 4 also includes the following steps: Step 401: Explicit restrictions in an existing procurement document and implicit limiting features Combine them to form the procurement data of this existing procurement document. ; Among the features For explicit or implicit restrictive features of the existing procurement document, Features The tag and ; Step 402: Combine the procurement data from each existing procurement document. To form a compliance judgment dataset Where p is the number of existing procurement documents, For the labels of existing procurement documents, and ; Step 403: Train a model based on a decision tree algorithm using the compliance judgment dataset E, with procurement data as the model. Features For internal nodes, tags Decisions are made based on the edge relationships of internal nodes; finally, the labels in the existing procurement documents are used. The compliance judgment model is obtained by ending at a leaf node.
5. The intelligent review method for the fairness of procurement document eligibility based on multi-technology integration as described in claim 4, characterized in that, Step 4 also includes the following steps: Step 404 introduces the reinforcement learning (RLHF) method based on human feedback, whereby professional reviewers manually review the inspection report generated by the compliance judgment model and evaluate and mark the inspection results summary, explicit constraints, and implicit constraints in the inspection report. Step 405: The evaluation mark is used as a key reward signal and input into the compliance judgment model with a specific batch size for retraining; Step 406: The compliance judgment model calculates the strategy gradient based on the reward signal, updates the model parameters through the Adam optimizer, and optimizes the decision logic to improve the accuracy of the compliance judgment model in judging the compliance of the qualification clauses of the procurement documents, so that the compliance judgment model can better adapt to various complex procurement document scenarios.
6. The intelligent review method for the fairness of procurement document eligibility based on multi-technology integration as described in claim 4, characterized in that, It also includes the following steps: Step 601: Use text parsing technology to locate and extract qualification clauses in existing procurement documents; then use dependency parsing technology to analyze the dependency relationships in the qualification clauses, identify and save the sentences and paragraphs related to qualification, and obtain the qualification dataset. Step 602: Label the eligibility dataset, labeling sentences and paragraphs with enterprise size restrictions as risky and other sentences and paragraphs as not risky; then feed this dataset into the pre-trained BERT model for fine-tuning, so that the BERT model can understand sentences or paragraphs with enterprise size restrictions. Step 603: Use the BERT model to perform semantic matching on the procurement documents to be judged, and extract all statements and paragraphs that are judged to have compliance risks from the procurement documents to be judged as compliance risk data. Step 604: Finally, according to the order in which the compliance risk data appears in the procurement documents to be judged, the compliance risk data is entered into the inspection report in sequence to obtain an inspection report with compliance risk analysis capabilities.
7. The intelligent review method for the fairness of procurement document eligibility based on multi-technology integration as described in claim 6, characterized in that: The inspection report includes basic project information, an overview of the inspection results, explicit restrictions, implicit restrictions, and compliance risk data.
Citation Information
Patent Citations
Bidding document review method based on artificial intelligence technology, computer device, medium and program product
CN118536473A
Purchase file compliance checking system based on difference algorithm under AI large model
CN118551760A
Green energy power industry purchase file compliance inspection method and system
CN119067457A
Bid evaluation method and system based on artificial intelligence
CN119295007A