Purchase file qualification fairness intelligent review method based on multi-technology fusion

Through the intelligent review method of qualification fairness of procurement documents with multi-technology integration, text analysis, rule engine, knowledge graph and decision tree algorithm are used to comprehensively identify explicit and implicit restrictions in procurement documents, solving the problems of low efficiency and poor accuracy in the existing technology, and achieving fast and accurate fairness review.

CN120278685AActive Publication Date: 2025-07-08BOSI DIGITAL TECH CO LTD

Patent Information

Application Number
CN202510785301.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-08
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

The existing procurement documents qualification fairness review methods are inefficient, have poor accuracy, insufficient comprehensiveness, and poor adaptability and scalability, and cannot effectively identify complex and concealed unfair terms, resulting in inaccurate review results and delays in procurement processes.

Method used

Multi-technical fusion methods are adopted, including text analysis, rules engine, knowledge graph, natural language processing and decision tree algorithms, to build a compliance judgment model, comprehensively check the explicit and implicit restrictions in procurement documents, and generate standardized inspection reports.

Benefits of technology

It has achieved a comprehensive review of procurement documents, quickly identified potential unfair terms, shortened the review cycle, improved the accuracy and efficiency of review, ensured fair access to market entities, and provided scientific decision-making support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278685A_ABST
    Figure CN120278685A_ABST
Patent Text Reader

Abstract

The invention relates to a purchase file qualification fairness intelligent review method based on multi-technology fusion. The method comprises the following steps: processing an existing purchase file to obtain a purchase text; defining a keyword library, constructing a rule engine to search for vocabularies of the keyword library, matching digital thresholds closely following the vocabularies, performing judgment, and finally storing the digital thresholds as dominant restriction features; constructing a knowledge graph, identifying a qualification entity query knowledge graph in the purchase text, performing judgment, and storing the qualification entity query knowledge graph as a hidden restriction feature; combining dominant restriction features and implicit restriction features of the existing purchase file; combining the purchase data of each existing purchase file to form a compliance judgment data set, and training a model constructed based on a decision tree algorithm to obtain a compliance judgment model; and analyzing the to-be-judged purchase file by using the compliance judgment model. Examination of dominant limiting conditions and implicit limiting conditions is taken into consideration, and from text surface expression to deep association of qualification, the qualification terms of purchase files are comprehensively examined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Internet procurement business, and in particular to an intelligent review method for fairness of procurement document qualifications based on multi-technology integration. Background Art

[0002] In current procurement activities, it is crucial to ensure the fairness of the qualification clauses in procurement documents. Fair qualification clauses in procurement documents can create an open and fair market environment, giving all types of suppliers the opportunity to participate in bidding and promoting the healthy development of the market. However, there are many problems that need to be solved in the review of procurement documents:

[0003] The traditional way of reviewing procurement documents mainly relies on manual work: reviewers need to read the procurement documents word by word to find possible unfair qualification clauses. This method has significant inefficiencies. As the scale of procurement continues to expand, the number of procurement documents is growing, and the content is becoming more and more complex. Reviewers need to spend a lot of time and energy to sort out and analyze them. For example, in the review of procurement documents for some large-scale infrastructure construction projects, it takes several weeks to manually review all procurement documents, which seriously affects the progress of procurement projects.

[0004] The accuracy of manual review is also difficult to guarantee: due to differences in the professional level, experience and subjective judgment standards of reviewers, different reviewers may come to different conclusions for the same qualification clause. Moreover, long-term review work can easily lead to fatigue and lack of concentration of reviewers, thus missing some deeply hidden unfair clauses. Some procurement documents may also set restrictions on enterprise scale conditions / years of operation / equity structure / ownership form / location / organizational form through vague expressions or indirect associations, which may be difficult to detect during manual review. At the same time, the market environment is constantly changing, new qualifications and licenses are constantly emerging, and unfair terms are set in an endless stream. Manual review is difficult to quickly adapt to these changes.

[0005] Most of the existing automated review tools are single-function and lack comprehensiveness: for example, some tools can only find obvious restrictive clauses through simple keyword matching, and are powerless against synonymous substitution, ambiguous expressions, and restrictive conditions hidden behind qualification requirements. These tools do not make full use of advanced artificial intelligence technology and knowledge graphs, and cannot deeply understand the semantics and logical relationships of procurement documents, resulting in inaccurate and incomplete review results.

[0006] In the current field of fairness review of procurement document qualifications, the existing technology has many shortcomings that are difficult to ignore, which seriously affect the quality and efficiency of the review work, mainly reflected in the following aspects:

[0007] I. Inefficient review:

[0008] 1. Predominantly manual review: Currently, most review work relies on manual efforts. Reviewers need to manually flip through every page of the procurement documents and search for potential unfair qualification terms word by word. With the expansion of the procurement business scale, the number of procurement documents is increasing, and the content is becoming more and more complex. Manually reviewing such a large volume of documents often requires a large amount of time and energy, resulting in a long review cycle and seriously delaying the progress of the procurement process.

[0009] 2. Lack of automated processing capabilities: In the existing technology, automated review tools are less applied. Even if there are some automated means, they can only handle simple tasks and cannot achieve systematic and comprehensive automated review of procurement documents. For example, some simple text search tools can only locate specific keywords, but are powerless for complex semantic understanding, clause correlation analysis, etc. This makes reviewers still need to manually complete most of the work, and the efficiency improvement is limited.

[0010] II. Poor review accuracy:

[0011] 1. Differences in subjective judgment: During the manual review process, due to different professional backgrounds, experience levels, and personal understandings of different reviewers, there may be significant differences in the judgment criteria for the same qualification clause. For example, for some clauses with ambiguous expressions, one reviewer may think it does not constitute a restriction, while another reviewer may hold the opposite view. This inconsistency in subjective judgment leads to the lack of objectivity and accuracy in the review results.

[0012] 2. Difficulty in identifying complex clauses: The setting methods of unfair clauses in procurement documents are becoming more and more complex and concealed, such as using ambiguous semantics, indirect associations, synonymous substitutions, etc. During manual review, reviewers may miss these deeply hidden unfair clauses due to fatigue, negligence, or misunderstanding of complex language structures. At the same time, in the face of constantly changing new clause setting methods, it is difficult for manual review to master and respond in a timely manner, further reducing the review accuracy.

[0013] III. Insufficient review comprehensiveness:

[0014] 1. Single - dimension review: Most of the existing review techniques only focus on the surface text expression of procurement documents and lack in - depth exploration of potential restrictive conditions hidden behind the qualification requirements. For example, they only review whether the document directly mentions enterprise scale conditions / business operation years / equity structure / ownership form / location / organizational form, etc., while ignoring the restrictions on enterprise scale conditions / business operation years / equity structure / ownership form / location / organizational form, etc. in the qualification certificate acquisition conditions. This single - dimension review method cannot comprehensively discover the unfair problems in procurement documents, and may unreasonably exclude some suppliers from procurement activities due to hidden restrictions.

[0015] 2. Lack of multi - technology integration: Most review tools and methods only focus on the application of a certain technology, such as simple text retrieval technology or simple rule - matching technology, and fail to fully integrate multiple advanced technologies such as natural language processing and knowledge graphs. It is difficult to conduct comprehensive analysis of procurement documents from multiple perspectives, resulting in loopholes and blind spots in the review work.

[0016] IV. Poor adaptability and scalability:

[0017] 1. Difficulty in adapting to changes: The market environment is constantly changing, and the forms and contents of procurement documents are also evolving continuously. New means of setting unfair terms emerge in an endless stream. However, existing technologies often lack effective self - learning and updating mechanisms and are difficult to quickly adapt to these changes. For example, when new qualification requirements or clause expression methods appear, traditional review technologies may not be able to identify and analyze them in a timely manner, resulting in lagging review work.

[0018] 2. Weak system scalability: Existing review systems or methods lack foresight in design, and the architecture is relatively closed. It is difficult to conveniently integrate new technologies or modules to enhance the review ability. For example, when a more advanced natural language processing model appears, it cannot be easily integrated into the existing system, restricting the further development and optimization of review technologies.

[0019] In summary, the existing procurement document review methods cannot meet the current requirements of efficient, accurate, and comprehensive review. There is an urgent need for a new technology to solve these problems and ensure the fairness of the qualification terms in procurement documents. Summary of the Invention

[0020] The purpose of the present invention is to provide an intelligent review method for the fairness of procurement document qualifications based on multi - technology integration, which can ensure the fairness of the qualification terms in procurement documents.

[0021] The present invention is realized through the following technical solutions:

[0022] An intelligent review method for the fairness of procurement document qualifications based on multi - technology integration includes the following steps:

[0023] Step 1: Preprocess the text of existing procurement documents; use a text parsing algorithm to process the existing procurement documents to obtain procurement text;

[0024] Step 2: First, define a keyword library regarding enterprise scale restrictions; then, construct a rule engine, use the rule engine to search for the vocabulary in the keyword library within the procurement text, and then match the numerical threshold following the vocabulary and determine whether the numerical threshold exceeds the specified limit to obtain a judgment result; finally, combine the keyword and its numerical threshold and set it as an explicit restriction condition, and save the explicit restriction condition and its judgment result as an explicit restriction feature;

[0025] Step 3: First, use the graph database tool Neo4j to construct a knowledge graph with qualification certificates as the core; then, use natural language processing methods to identify qualification entities in the procurement text, query the knowledge graph through the qualification entities, and when an enterprise scale restriction node of the qualification entity is found in the knowledge graph, it is determined that there is a hidden compliance risk; finally, save the query path and judgment result of the qualification entity in the knowledge graph as an implicit restriction feature;

[0026] Step 4: Combine the explicit restriction features and implicit restriction features of an existing procurement document to form its procurement data; combine the procurement data of each existing procurement document to form a compliance judgment data set, and use the compliance judgment data set to train a model constructed based on the decision tree algorithm to obtain a compliance judgment model;

[0027] Step 5: Use the compliance judgment model to analyze the procurement document to be judged, generate an inspection report for the procurement document to be judged, and enter the decision-making judgment process and the final judgment result of the compliance judgment model for the procurement document to be judged in the inspection report.

[0028] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0029] 1. The present invention takes into account the inspection of both explicit restriction conditions and implicit restriction conditions, comprehensively examines the qualification clauses of procurement documents from the surface expression of the text to the deep association of qualifications. It can not only find obvious restrictions directly reflected in the clause provisions, but also use knowledge graph reasoning to find the enterprise scale restrictions hidden behind industry qualification certificates, ensuring that there are no dead ends in the inspection, assisting in fairness inspection, comprehensively guaranteeing the fairness of market entity access, and preventing the situation where some suppliers are unreasonably excluded due to clause loopholes.

[0030] 2. The present invention can utilize a variety of natural language processing methods to coordinate the processing of procurement documents, can process a large amount of text in a short time, and display important information such as basic project information, inspection result overview, explicit restriction features and implicit restriction features in the inspection report, greatly shortening the amount of text review by reviewers; in the review of a large procurement document of hundreds of pages, manual review may take several weeks, while the technical means provided by the present invention can complete the preliminary review in only a few hours, which can greatly shorten the review cycle, speed up the procurement process, facilitate reviewers to review the key points, and enable the project to enter the bidding implementation stage faster.

[0031] 3. The standardized inspection report generated by the compliance judgment model in the present invention clearly presents important information such as the basic information of the procurement document items, an overview of the inspection results, explicit restriction features, and implicit restriction features. Tenderers and regulatory authorities can quickly understand the fair competition compliance of the documents without professional technical knowledge, and adjust the procurement documents in a targeted manner based on the content of the report to make scientific decisions. The standardization and professionalism of the report also provide strong support for regulatory work and improve the overall management level of procurement activities.

[0032] 4. The present invention also uses the BERT model to expand the understanding ability, and can find the sentences and paragraphs in the procurement documents that set enterprise size restrictions through fuzzy expressions or indirect associations, and save them as compliance risk data and display them in the inspection report for the convenience of reviewers to review and analyze. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is a flow chart of the present invention;

[0034] Figure 2 Pseudocode for constructing regular expression functions;

[0035] Figure 3 Pseudocode for defining the function of logic rules to determine the digital threshold;

[0036] Figure 4 This is the pseudo code for constructing the main program of the rule engine. DETAILED DESCRIPTION

[0037] The present invention is described in detail below in conjunction with the accompanying drawings, but the protection content of the present invention is not limited to the following:

[0038] like Figures 1-4 Shown is a schematic diagram of an embodiment of a method for intelligent review of fairness of procurement document qualifications based on multi-technology fusion provided by the present invention.

[0039] The intelligent review method for fairness of procurement document qualifications based on multi-technology integration includes the following steps:

[0040] Step 1, perform text preprocessing on existing procurement documents; use text parsing algorithms to process existing procurement documents to obtain procurement text;

[0041] Step 2, first define a keyword library regarding enterprise scale restrictions; then construct a rule engine, use the rule engine to search for the vocabulary in the keyword library within the procurement text, and then match the numerical threshold following the vocabulary and determine whether the numerical threshold exceeds the specified limit to obtain a judgment result; finally, combine the keyword and its numerical threshold to set it as an explicit restriction condition, and save the explicit restriction condition and its judgment result as an explicit restriction feature;

[0042] Step 3, first use the graph database tool Neo4j to construct a knowledge graph with qualification certificates as the core; then identify qualification entities in the procurement text through natural language processing methods, query the knowledge graph through the qualification entities, and when an enterprise scale restriction node of the qualification entity is found in the knowledge graph, it is determined that there is a hidden compliance risk; finally, save the query path and judgment result of the qualification entity in the knowledge graph as an implicit restriction feature;

[0043] Step 4, combine the explicit restriction features and implicit restriction features of an existing procurement document to form its procurement data; combine the procurement data of each existing procurement document to form a compliance judgment data set, and use the compliance judgment data set to train a model constructed based on the decision tree algorithm to obtain a compliance judgment model;

[0044] Step 5, use the compliance judgment model to analyze the procurement document to be judged, generate an inspection report for the procurement document to be judged, and enter the decision-making judgment process and the final judgment result of the compliance judgment model for the procurement document to be judged in the inspection report.

[0045] It should be noted that the present invention takes into account the inspection of both explicit and implicit restrictions, from the surface expression of the text to the deep association of qualifications, and comprehensively reviews the qualification clauses of the procurement documents. It can not only find the obvious restrictions that directly reflect the provisions of the clauses, but also find out the enterprise scale restrictions hidden behind the industry qualification certificates through knowledge graph reasoning, ensure that there are no blind spots in the review, assist in fairness review, comprehensively guarantee the fairness of market access, and prevent some suppliers from being unreasonably excluded due to loopholes in the clauses. Using a variety of natural language processing methods to coordinate the processing of procurement documents, a large amount of text can be processed in a short time, and important information such as basic project information, inspection result overview, explicit restriction features and implicit restriction features can be displayed in the inspection report, which greatly shortens the amount of text review by the reviewer; in the review of a large procurement document with hundreds of pages, manual review may take several weeks, while the technical means provided by the present invention can complete the preliminary review in only a few hours, which can greatly shorten the review cycle, speed up the procurement process, and facilitate the reviewer to review the key points, so that the project can enter the bidding implementation stage faster. The standardized inspection report generated by the compliance judgment model clearly presents important information such as the basic information of the procurement document items, an overview of the inspection results, explicit restriction features, and implicit restriction features. Tenderers and regulatory authorities can quickly understand the fair competition compliance of the documents without professional technical knowledge, and adjust the procurement documents in a targeted manner based on the content of the report to make scientific decisions. The standardization and professionalism of the report also provide strong support for regulatory work and improve the overall management level of procurement activities.

[0046] Furthermore, step 1 also includes the following steps:

[0047] Step 101, pre-processing the existing procurement document to remove redundant symbols, special symbols and space information in the existing procurement document;

[0048] Step 102, first use the n-gram model to check the syntax and semantics of the existing procurement document; then use the text segmentation algorithm to divide the long text into appropriate paragraphs, and then use the title recognition algorithm to identify the chapter title or clause number based on the font format, typesetting characteristics and text content information to obtain the existing procurement text.

[0049] It should be noted that the n-gram model is a commonly used language model in large vocabulary continuous speech recognition. It can judge whether the text conforms to language habits by statistically analyzing the frequencies of adjacent n words, and correct abnormal word sequences. Take a simple procurement document text "The supplier should provide products of high quantity and quality" as an example: Suppose we use a 3-gram model (n = 3), which will divide the text into word sequences of length 3 (also known as n-grams) in order. First, the text is tokenized to obtain words such as "supplier", "should", "provide", "high quantity and quality", "of", and "products". Then 3-gram sequences are generated, starting from the first word:

[0050] "The supplier should provide";

[0051] "Should provide high quantity and quality";

[0052] "Provide high quantity and quality of";

[0053] "High quantity and quality of products";

[0054] For each 3-gram, the model will count the frequency of this sequence in its training data. In a large number of normal procurement document text training data, 3-grams such as "The supplier should provide" and "Provide products of" may have a relatively high frequency of occurrence because they conform to common language expression habits. However, "Should provide high quantity and quality" and "Provide high quantity and quality of" contain the word sequence "high quantity and quality" which does not conform to normal expression habits, and their frequencies of occurrence in the training data will be very low. Based on this, the n-gram model determines that "high quantity and quality" is an abnormal word sequence. When the model makes corrections, it will propose correction suggestions based on the high-frequency word sequences in similar contexts in the training data. For example, in a context like "Provide... products", the high-frequency word may be "high quality", so the model will suggest correcting "high quantity and quality" to "high quality", making the text "The supplier should provide high-quality products" conform to normal language expression habits, thus completing the processing of the procurement document text and providing a more reliable text basis for subsequent accurate analysis of the procurement document content;

[0055] The text segmentation algorithm can divide long texts into appropriate paragraphs based on punctuation marks, paragraph marks, and semantic coherence. The title recognition algorithm can identify chapter titles or clause numbers based on information such as font formats, layout features, and text content. For example, for text in the format of "1. Eligibility requirements", by matching numbers and specific title keywords, it is identified as a chapter title. For extremely long texts, a sliding window-based chunking method is adopted to chunk the text into windows of a certain length, and the text within each window is used as an independent analysis unit to improve the efficiency of subsequent NLP analysis.

[0056] Further, in step 2, the following steps are also included:

[0057] Step 201: First, define a keyword library for enterprise scale restrictions. The keyword library includes registered capital, paid-in capital, number of employees, tax payment amount, operating years, location, and text clues related to education background and qualifications.

[0058] Step 202: Then, build a rule engine based on regular expressions and logical rules; use regular expressions to search for the vocabulary in the keyword library within the procurement text; use logical rules to match the numerical threshold immediately following the keyword and determine whether the numerical threshold exceeds the specified limit of the logical rules:

[0059] Step 203: Obtain the judgment result: When the numerical threshold exceeds the limit set by the logical rules, set the judgment result as non-compliant; otherwise, set the judgment result as compliant.

[0060] Step 204: Finally, merge the keyword and its numerical threshold and set it as an explicit restriction condition; save the explicit restriction condition and its judgment result as an explicit restriction feature.

[0061] Specific pseudocode and example explanations for building a rule engine based on regular expressions and logical rules:

[0062] 1. Define the keyword library:

[0063] Here, take the keyword library keyword_list = ["registered capital", "annual turnover", "tax payment amount"] with three words as an example for explanation;

[0064] 2. Build regular expressions:

[0065] As Figure 2 shown is the function pseudocode for building regular expressions. The meaning of the regular expression regex built here is: First, match the keyword keyword, '\s*' means to match 0 or more whitespace characters (including spaces, tabs, etc.), '[::]?' means to match 0 or 1 colon (Chinese or English colon), '\s*' is used to match whitespace characters again, '(\d+)' is used to capture the number immediately following the whitespace character, '\s*' continues to match whitespace characters, '(wan|yuan|%)?' means to match 0 or 1 unit (wan, yuan, or %).

[0066] 3. Define logical rules to judge the numerical threshold:

[0067] As Figure 3The following is the pseudocode of a function that defines a logical rule to judge a numerical threshold. The function check_threshold judges whether the set threshold conditions are met according to different keywords and the corresponding numbers and units. Taking "registered capital" as an example, if the unit is "ten thousand" and the number is less than 100, or the unit is "yuan" and the number is less than 1,000,000, it is considered that the condition is not met and False is returned; otherwise, True is returned;

[0068] 4. Build the main program of the rule engine

[0069] As Figure 4 The following is the pseudocode of a function that builds the main program of the rule engine. In the function, the keyword library keyword_list and the corresponding regular expressions are traversed, and the find All Matches function (assuming this function is used to find all contents in the text that match the regular expression) is used to find the matching items matches in the text of the procurement document; for each matching item matches, its keyword, number and unit are extracted, and the check_threshold function is called to judge whether the numerical threshold meets the requirements; if not, a prompt message is output.

[0070] Furthermore, the following steps are also included in step 3:

[0071] Step 301, use the graph database tool Neo4j to build a knowledge graph with qualification certificates as the core. This graph takes common domestic qualification certificates as the core nodes, and sets enterprise scale limit nodes and non-enterprise scale limit nodes;

[0072] Step 302, based on the acquisition conditions of each qualification certificate, associate the corresponding acquisition conditions to the corresponding enterprise scale limit nodes and non-enterprise scale limit nodes through edge relationships;

[0073] Step 303, then extract the statements with qualification requirements in the procurement text through the keyword matching algorithm, and use the named entity recognition technology to identify the qualification entities in the statements;

[0074] Step 304, use the path search algorithm to find the path from the qualification certificate node corresponding to the qualification entity to the enterprise scale limit node in the knowledge graph. If a path exists, it is determined that there is a hidden compliance risk, and the determination result is set to non-compliant; in other cases, the determination result is set to compliant;

[0075] Step 305, finally, save the query path of the qualification entity in the knowledge graph and its determination result as a hidden restriction feature.

[0076] It should be noted that in steps 301 and 302, the Neo4j graph database associates the qualification certificates and enterprise scale limit nodes through data collection, collation, and import operations. For example, for the "Contract-abiding and Creditworthy Certificate" node, its attributes include information on the enterprise operation years requirement. In the knowledge graph, the relationship between nodes can be expressed as "has_requirement", etc., to clarify the association between qualifications and restrictions such as enterprise scale conditions, operation years, equity structure, ownership form, location, and organizational form;

[0077] In step 303, in the procurement document, statements with qualification requirements are extracted through a keyword matching algorithm. For example, the procurement document mentions expressions such as "The enterprise needs to have an enterprise entity credit rating certificate", "The enterprise needs to have a CMM / CMMI certification certificate", "The enterprise needs to have a CCRC information security service qualification certification certificate", etc.; then, using the named entity recognition (NER) technology of NLP, the qualification entities in the qualification requirement keywords are identified, that is, the enterprise entity credit rating certificate, CMM / CMMI certification certificate, or CCRC information security service qualification certification certificate; specifically, the NER algorithm based on the conditional random field (CRF) model can be used to define the feature function , where and are adjacent position labels, x is the input text, i is the position index, calculate the optimal label sequence for each position under the given text x, so as to identify the qualification entity; after identifying the qualification entity, query the knowledge graph to determine whether the qualification entity contains requirements for enterprise scale conditions, operation years, equity structure, ownership form, location, organizational form, etc.

[0078] In a more advanced scenario, utilize the reasoning ability of the knowledge graph to perform reasoning based on the graph structure and node relationships. For example, through the path search algorithm, search for the path from the "qualification requirement" node to the "enterprise scale limit condition" node in the knowledge graph. If there is a path, it indicates the existence of potential compliance risks. The specific path search algorithm can adopt the breadth-first search (BFS) or depth-first search (DFS) algorithm. Taking the breadth-first search as an example, start from the "qualification requirement" node and expand the search nodes layer by layer until the "enterprise scale limit condition" node is found or all reachable nodes are traversed;

[0079] In steps 301 and 302, the construction and application of the qualification knowledge graph include the following processes:

[0080] 3021) Perform vector embedding expression on the resource description framework triples in the knowledge graph with qualification certificates as the core, where the resource description framework triples are The corresponding vocabulary sequence is set as Y= ;

[0081] Here 、 、 represent the vocabulary sequences of the subject (such as a specific qualification category), the predicate (such as a relational expression like "meet the conditions of..."), and the object (such as the lower limit of enterprise scale, the threshold of operation years, etc., related to restrictive conditions), respectively;

[0082] 3022) Define the function , n ; where, is used to represent the i-th vocabulary in the entity or relationship vocabulary sequence corresponding to n in; define the word embedding mapping (where, represents the j-th sub-word of, and the calculation method is the product of vectors);

[0083] Through these steps, the vocabulary sequence Y of the Resource Description Framework triple can be embedded into a word vector tensor space, represented as a function:

[0084] , where J represents the word vector tensor, 、 、 are the components mapped by the word embedding mapping ; 、 、 ;

[0085] 3023) Determine the input of the first-layer Graph Neural Network (GNN) block at time t: Define the position offset as the position-related parameter of a word vector of the graph neural network model at time t;

[0086] 3024) Define the feature vector as a feature vector in the feature vector sequence generated by the graph neural network model at time t-1;

[0087] (here represents a specific vector fusion operation, such as bitwise addition or concatenation, etc.), and thus obtain the input of the first-layer graph neural network block at time t;

[0088] 3025) Set the number of layers of the graph neural network model; Set that each layer of the graph neural network model is composed of multiple graph neural network blocks, and perform the following calculations in each graph neural network block:

[0089] ;

[0090] ;

[0091] ;

[0092] ;

[0093] Among them, GraphAttention represents the graph attention function, which is used to capture the correlation information between nodes in the graph structure; GraphNorm represents the graph normalization function, which normalizes the graph data to accelerate convergence and improve stability; GraphFF represents the graph feed-forward function, which is used to perform non-linear transformation on features; through this series of operations, the deep information contained in the nodes and relationships of the qualification knowledge graph is mined.

[0094] Furthermore, step 4 further includes the following steps:

[0095] Step 401: Combine the explicit restriction features and implicit restriction features of an existing procurement document to form a combination of this procurement document, and the procurement data ;

[0096] where the feature is the explicit restriction feature or implicit restriction feature of this existing procurement document, is the label of the feature and ;

[0097] Step 402: Combine the procurement data of each existing procurement document to form a compliance judgment data set , where p is the number of existing procurement documents, is the label of the existing procurement document, and ;

[0098] Step 403: Use the compliance judgment data set E to train a model constructed based on the decision tree algorithm. In the model, use the feature of the procurement data as the internal node, and use the label of as the edge relationship of the internal node for decision-making; finally, use the label

[0099] It should be noted that the decision tree is an algorithm for making decisions based on a tree structure. In this compliance judgment model, the decision tree consists of nodes and edges. The nodes are divided into internal nodes and leaf nodes. The internal nodes represent tests on an attribute, the edges represent test outputs, and the leaf nodes represent decision results. For example, in a procurement document, its explicit restrictive features may include "whether there is a registered capital limit", "whether there is an operating life limit", etc., and its implicit restrictive features may include "potential equity structure restriction tendencies analyzed through a knowledge graph", "potential operating life limit tendencies", etc.; these explicit and implicit restrictive features together constitute the internal nodes of the decision tree. If the edge relationship is "compliant", then continue to judge the next node along the corresponding edge; if it is "non-compliant", then follow the other edge to the next node for continued judgment. Eventually, reach the leaf node and obtain the decision result of "fair procurement" or "unfair procurement"; this becomes a decision branch of the decision tree.

[0100] It should be noted that in the training process of the decision tree in step 403, the decision tree is mainly constructed by continuously selecting the optimal features and splitting points; at each internal node, metrics such as Information Gain or Gini Impurity are used to select the optimal feature for splitting; taking Information Gain as an example, Information Gain measures the increase in information obtained by splitting the compliance judgment dataset E using the feature ; The calculation formula is:

[0101] ;

[0102] where is the information entropy of the compliance judgment dataset E, which is used to measure the uncertainty of the compliance judgment dataset E. The calculation formula is:

[0103] ;

[0104] where, are all possible values of the feature A, that is, A = "fair procurement" or "unfair procurement"; is the number of labels (i.e., "fair procurement" and "unfair procurement") in the compliance judgment dataset E, that is, C = 2; is the sample subset of the compliance judgment dataset E when belonging to the label u, that is, is the number of procurement documents with "fair procurement" or the number of procurement documents with "unfair procurement"; is a subset of the compliance judgment dataset E;

[0105] By selecting the Information Gain The largest feature is used as the segmentation feature of the current node, and then the above operation is recursively performed on the segmented child nodes until the stopping condition is met (such as reaching the maximum depth D or the number of node samples is less than the minimum sample segmentation number S), and the compliance judgment model is obtained;

[0106] Furthermore, in step 403, after obtaining the compliance judgment model, a test data set containing M procurement documents is used. ; Each feature vector in the test data set Input into the trained compliance judgment model, where The explicit and implicit restrictive features of the procurement document used for testing are included. is the true label of whether this procurement document is "fair procurement" or "unfair procurement"; the compliance judgment model can start from the root node and gradually traverse the decision tree downward according to the test conditions of each node until it reaches the leaf node and obtains the prediction result ;

[0107] By comparing the prediction results and the true label , calculate the model’s accuracy, recall and other evaluation indicators to measure the model’s performance;

[0108] Generally speaking, accuracy is the most intuitive performance indicator, and its calculation formula is:

[0109] ;in is an indicator function if equal , then the value is 1, indicating that the prediction result is correct; otherwise, it is 0, indicating that the prediction result is wrong.

[0110] It should be noted that after the judgment result is obtained through the compliance judgment model, an inspection report is generated according to the standardized template. The report includes basic project information (such as project name, project number, etc.), and this information is extracted from the procurement documents as the content of the corresponding part of the report. The "fair procurement" or "unfair procurement" conclusion output by the model is directly filled in the overview of the inspection results; and all explicit restriction features and implicit restriction features found in the explicit and implicit inspections are listed one by one, and the "compliance" and "non-compliance" of their side relationships are correspondingly output. Finally, the generated inspection report is presented in PDF or Word format for easy reference by tenderers, centralized procurement agencies, procurement supervision departments, and social agencies.

[0111] Furthermore, step 4 also includes the following steps:

[0112] Step 404: Introduce the Reinforcement Learning from Human Feedback (RLHF) method. Have professional reviewers conduct manual review on the inspection reports generated by the compliance judgment model, and make judgment marks on the inspection result summary, the results of explicit restrictive conditions, and the results of implicit restrictive conditions in the inspection reports. For example, if the compliance judgment model determines that a certain procurement document is compliant, but the reviewer discovers an implicit restrictive risk in it that was not identified by the model, then mark this report as "misjudged" and specify the reason for the misjudgment in detail.

[0113] Step 405: Use the judgment marks as key reward signals and input them into the compliance judgment model in a specific batch size for retraining. The purpose of setting the batch size is to balance training efficiency and model stability, avoiding oscillations in the model learning process caused by inputting too much data at once, and also preventing the model learning speed from being too slow due to too small a batch size. For example, it can be set that every 10 judgment marks are input into the model as a batch.

[0114] Step 406: The compliance judgment model calculates the policy gradient based on the reward signal. The policy gradient represents the improvement direction of the model's current policy (i.e., decision logic) under the guidance of the reward signal. Update the model parameters through the Adam optimizer to optimize the decision logic, so as to improve the accuracy of the compliance judgment model in judging the compliance of the qualification clauses of procurement documents and enable the compliance judgment model to better adapt to various complex procurement document scenarios. If the previous judgment of the model makes the reviewer give a positive judgment mark (such as correctly judging that the procurement document is compliant), the policy gradient will adjust in the direction of strengthening this judgment logic; conversely, if a negative judgment mark (such as misjudgment) is obtained, the policy gradient will prompt the model to correct the decision logic.

[0115] It should be noted that the use of the Reinforcement Learning from Human Feedback (RLHF) method in the present invention is a secondary technical means and not an essential technical feature. This method allows professional reviewers (centralized procurement institutions, procurement supervision departments, social agency institutions) to conduct manual review on the inspection reports generated by the compliance judgment model. Reviewers can, based on their professional knowledge and experience, judge the inspection result summary and detailed risk points in the inspection reports, judge and mark the inaccurate or incomplete parts and give explanations;

[0116] Then these judgments, marks, and explanations will be used as reward signals of human feedback and fed back to the compliance judgment model for retraining. For example, if the model misjudges that a certain clause is compliant while the manual review deems it non-compliant, then give the model a negative reward signal; conversely, if the model makes a correct judgment, give a positive reward signal.

[0117] Furthermore, it further includes the following steps:

[0118] Step 601: Use text parsing technology to locate and extract the eligibility clauses in existing procurement documents; then use dependency syntax analysis technology to analyze the dependency relationships in the eligibility clauses, identify and save the sentences and paragraphs related to eligibility to obtain an eligibility dataset.

[0119] Step 602: Label the eligibility dataset, label the sentences and paragraphs with enterprise scale restrictions as having risks, and label other sentences and paragraphs as not having risks; then input this dataset into a pre-trained BERT model for fine-tuning so that the BERT model can understand the sentences or paragraphs with enterprise scale restrictions.

[0120] Step 603: Use the BERT model to perform semantic matching on the procurement document to be judged, and extract all the sentences and paragraphs determined to have compliance risks from the procurement document to be judged as compliance risk data.

[0121] Step 604: Finally, according to the order of appearance of the compliance risk data in the procurement document to be judged, enter the compliance risk data into the inspection report one by one to obtain an inspection report with the ability to analyze compliance risks.

[0122] It should be noted that in Step 601, the text parsing technology includes common regular expression matching, named entity recognition (NER), relation extraction, and dependency syntax analysis technology, etc.

[0123] In the dependency syntax analysis technology, assume that the dependency relationship is expressed as , where and are two words in the sentence. Analyze the dependency relationships in the procurement document one by one, and key sentences such as "Suppliers should possess... qualifications" can be accurately extracted to obtain an eligibility dataset.

[0124] In step 602, the full name of the BERT model is Bidirectional Encoder Representations from Transformers, that is, bidirectional encoder representations from Transformer. The BERT model is a pre-trained language model based on the Transformer architecture. Its core innovation lies in the ability to pre-train deep bidirectional representations based on the left and right contexts of all layers, which enables the model to consider the information of both the previous and subsequent texts when processing text, thereby capturing the semantic features of the text more comprehensively and accurately. Different from traditional unidirectional language models (such as those that can only process text from left to right or from right to left), the bidirectional processing ability of BERT greatly enhances its ability to understand complex semantics. The BERT model is stacked by multiple layers of Transformer blocks, and each Transformer block mainly includes two parts: the multi-head self-attention mechanism (Multi-Head Attention) and the feed-forward neural network (Feed-Forward Neural Network). The multi-head self-attention mechanism can capture the semantic relationships in the text from different angles by calculating multiple different attention heads in parallel, further enriching the model's understanding of the text. The feed-forward neural network then performs a non-linear transformation on the results output by the self-attention mechanism to learn more complex feature representations;

[0125] Specifically: Suppose there is a qualification clause text in the procurement document: "This procurement requires that the supplier must be an enterprise registered and operating locally for more than 5 years and is a state-owned enterprise to participate." In the face of restrictive clauses set through synonymous replacement and fuzzy semantics, as well as implicit restrictions associated with certificates, the present invention can accurately identify them. This can greatly reduce the probability of misjudgment and missed judgment, providing a more reliable judgment for the fairness and competitiveness of the procurement document.

[0126] First, preprocess the text, such as performing operations like word segmentation and removing stop words, convert the text into a format suitable for input to the BERT model, and input the preprocessed text into the fine-tuned BERT model. Inside the model, the text is encoded through multiple layers of self-attention mechanisms; taking the self-attention calculation of one layer as an example, assume that the current text being processed is "registered and operating locally for more than 5 years", the model will generate corresponding query matrix Q, key matrix K, and value matrix V for each word;

[0127] For example, for the word "locally", the BERT model calculates through self-attention Attention, and the formula for Attention is:

[0128] ;

[0129] where ,K, V are the query matrix, key matrix, and value matrix respectively, is the dimension of the key matrix. In this way, the degree of association between each position in the text and other positions is calculated, so as to understand the text semantics. It is used to calculate the degree of association between the position of "local" and other positions in the text (such as the positions corresponding to words like "registered", "operated", "5 years", etc.); in this way, the model can understand the semantic relationships between words in the text and then grasp the semantics of the entire sentence. Since the BERT model has learned a large amount of general language knowledge during the pre-training stage and has been fine-tuned on a dataset containing restrictive clauses and non-restrictive clauses such as enterprise scale conditions, operating years, equity structure, ownership form, location, and organizational form, it can determine whether this text contains restrictive clauses such as enterprise scale restrictions, operating years, equity structure, ownership form, location, and organizational form;

[0130] Therefore, in the above example, the BERT model can identify that "registered and operated locally for more than 5 years" contains the restrictive condition of operating years, and "being a state-owned enterprise" contains the restrictive condition of ownership form. The BERT model will score or label this paragraph according to the knowledge it has learned and the fine-tuned parameters. If a relatively high score is given, it means that this text is very likely to contain explicit restrictive conditions, or directly label it as "containing restrictive clauses on operating years and ownership form", so as to help the reviewer quickly judge whether there are compliance risks in the statements or paragraphs in the procurement document;

[0131] It should be noted that the BERT model analyzes the procurement text for subsequent operations of the compliance judgment model. To prevent the BERT model from containing explicit restrictive features or implicit restrictive features, while screening for explicit restrictive features and implicit restrictive features in the procurement document to be judged, it is possible to choose to delete the statements in the procurement document to be judged that save explicit restrictive features and implicit restrictive features from its procurement text; finally, after the compliance judgment model obtains the judgment result, inputting this procurement text into the process of BERT model analysis can avoid the situation of repeating the display of the same content.

[0132] It should be noted that taking a procurement document with a project name of "a certain information system procurement project" and a project number of "2025-001" as an example; first, process this procurement document into procurement text;

[0133] Using the method of the present invention to process the procurement text, when it is found that there are clauses such as "the supplier's registered capital needs to reach more than 5 million yuan" in the procurement text, save them as a number of explicit restrictive features in the order of appearance of the clauses ;

[0134] When "CMM / CMMI certification certificate" and other qualification entities are found in the procurement text, query through the knowledge graph whether the acquisition conditions of the certificate have a tendency to be restricted by the enterprise scale limit node; save them as several implicit restriction features in the order of appearance of the qualification entities ;

[0135] Then, these features are combined into the feature vector of the procurement data of this procurement document Input it into the compliance judgment model for decision-making; it should be noted that in the compliance judgment model, the internal nodes of each branch have been divided into the order of judgment of explicit restriction conditions and implicit restriction features during the training process; the feature vector will input the explicit restriction features and implicit restriction features in turn according to the judgment and division in the process from the root node → each internal node → the final leaf node of the compliance judgment model, and make judgments and divisions in each internal node; through layer-by-layer screening, draw the conclusion of "fair procurement" or "unfair procurement" in the leaf node;

[0136] Finally, in the generated inspection report, fill in the project name "A certain information system procurement project" and project number "2025-001" in the basic project information part, and the inspection result summary is "fair procurement" or "unfair procurement"; then list all the explicit restriction conditions, implicit restriction conditions, and the judgments (i.e., edge relationships) of whether the explicit restriction conditions and implicit restriction conditions are judged to be compliant in the decision tree in detail.

[0137] Optionally, before generating the inspection report, the procurement text can be input into the BERT model for analysis. The BERT model can save the risky statements and paragraphs in the procurement text as compliance risk data based on its ability to expand semantics; the compliance risk data and its risk probability can be displayed for reviewers to analyze when necessary; prevent those with intentions from using unclear expressions in the procurement document to benefit certain specific enterprises, and further ensure fairness. After analyzing with the BERT model, the inspection report will also display the risk data obtained by the BERT model analysis; at this time, the inspection report includes the basic project information, inspection result summary, explicit restriction features, implicit restriction features, compliance judgments of explicit restriction features and implicit restriction features, and compliance risk data.

[0138] Although the present invention uses specific embodiments and their alternative ways to illustrate and explain the present invention, it should be understood that various changes and modifications can be implemented as long as they do not depart from the spirit of the present invention. Therefore, it should be understood that the present invention is not limited in any sense except by the accompanying claims and their equivalent conditions.

Claims

1. An intelligent review method for the fairness of procurement document qualifications based on the integration of multiple technologies, characterized in that, Including the following steps: Step 1, perform text preprocessing on existing procurement documents; use a text parsing algorithm to process the existing procurement documents to obtain procurement text; Step 2, first define a keyword library regarding enterprise scale restrictions; Then construct a rule engine, use the rule engine to search for the vocabulary in the keyword library within the procurement text, then match the numerical threshold following the vocabulary and determine whether the numerical threshold exceeds the specified limit to obtain a judgment result; finally, combine the keyword and its numerical threshold and set it as an explicit restriction condition, and save the explicit restriction condition and its judgment result as an explicit restriction feature; Step 3, first use the graph database tool Neo4j to construct a knowledge graph with qualification certificates as the core; then identify qualification entities in the procurement text through natural language processing methods, query the knowledge graph through the qualification entities, and when an enterprise scale restriction node of the qualification entity is found in the knowledge graph, it is determined that there is a hidden compliance risk; finally, save the query path and judgment result of the qualification entity in the knowledge graph as an implicit restriction feature; Step 4, combine the explicit restriction features and implicit restriction features of an existing procurement document to form its procurement data; combine the procurement data of each existing procurement document to form a compliance judgment data set, and use the compliance judgment data set to train a model constructed based on the decision tree algorithm to obtain a compliance judgment model; Step 5, use the compliance judgment model to analyze the procurement document to be judged, generate an inspection report for the procurement document to be judged, and enter the decision-making judgment process and the final judgment result of the compliance judgment model for the procurement document to be judged in the inspection report.

2. The intelligent review method for the fairness of procurement document qualifications based on multi-technology integration according to claim 1, wherein In Step 1, the following steps are also included: Step 101, perform text preprocessing on existing procurement documents, removing redundant symbols, special symbols, and blank information from the existing procurement documents; Step 102, first use the n-gram model to check the grammar and semantics of the existing procurement documents; then use a text segmentation algorithm to cut the long text into appropriate paragraphs, and then use a title recognition algorithm to identify chapter titles or clause numbers based on font formats, layout features, and text content information to obtain the existing procurement text.

3. The intelligent review method for the fairness of procurement document qualifications based on multi-technology integration according to claim 2, characterized in that, In Step 2, the following steps are also included: Step 201, first define a keyword library regarding enterprise scale restrictions, and the keyword library includes registered capital, paid-in capital, number of employees, tax payment amount, operating years, location, and text clues related to education and qualifications; Step 202, then construct a rule engine based on regular expressions and logical rules; use regular expressions to search for the vocabulary in the keyword library within the procurement text; use logical rules to match the numerical threshold following the keyword and determine whether the numerical threshold exceeds the specified limit of the logical rules: Step 203, obtain the judgment result: when the numerical threshold exceeds the limit set by the logical rules, set the judgment result as non-compliant; otherwise, set the judgment result as compliant; Step 204, finally, combine the keyword and its numerical threshold and set it as an explicit restriction condition; save the explicit restriction condition and its judgment result as an explicit restriction feature.

4. The intelligent review method for the fairness of procurement document qualifications based on multi-technology integration according to claim 3, characterized in that In Step 3, the following steps are also included: Step 301: Use the graph database tool Neo4j to construct a knowledge graph with qualification certificates as the core. This graph takes common domestic qualification certificates as the core nodes, and sets enterprise scale limit nodes and non-enterprise scale limit nodes. Step 302: Based on the acquisition conditions of each qualification certificate, associate the corresponding acquisition conditions to the corresponding enterprise scale limit nodes and non-enterprise scale limit nodes through edge relationships respectively. Step 303: Then, extract the sentences with qualification requirements in the procurement text through the keyword matching algorithm, and use the named entity recognition technology to identify the qualification entities in the sentences. Step 304: Use the path search algorithm to find the path from the qualification certificate node corresponding to the qualification entity to the enterprise scale limit node in the knowledge graph. If there is a path, it is determined that there is a hidden compliance risk, and the determination result is set as non-compliant; in other cases, the determination result is set as compliant. Step 305: Finally, save the query path of the qualification entity in the knowledge graph and its determination result as the implicit restriction feature.

5. The intelligent review method for the fairness of procurement document qualifications based on multi-technology integration according to claim 4, characterized in that, The following steps are also included in Step 4: Step 401: Combine the explicit restrictive features and implicit restrictive features of an existing procurement document to form the procurement data of this procurement document; where the feature is an explicit restriction feature or an implicit restriction feature of this existing procurement document, is the label of the feature and ; Step 402, combine the procurement data of each existing procurement document , to form a compliance judgment data set , where p is the number of existing procurement documents, is the label of the existing procurement document, and ; Step 403: Use the compliance judgment dataset E to train a model constructed based on the decision tree algorithm. In the model, use the features of the procurement data as internal nodes and use the labels as the edge relationships of the internal nodes for decision-making; finally, end with the labels of the existing procurement documents as leaf nodes to obtain a compliance judgment model. ​ 6. The intelligent review method for the fairness of procurement document qualifications based on multi-technology integration according to claim 5, characterized in that The following steps are also included in Step 4: Step 404: Introduce the Reinforcement Learning from Human Feedback (RLHF) method. Let professional reviewers conduct manual review on the inspection reports generated by the compliance judgment model, and make judgment marks on the inspection result summary, the results of explicit restriction conditions, and the results of implicit restriction conditions in the inspection reports. Step 405: Take the judgment marks as the key reward signals, and input them into the compliance judgment model in a specific batch size for retraining. Step 406: The compliance judgment model calculates the policy gradient according to the reward signals, updates the model parameters through the Adam optimizer, and optimizes the decision-making logic to improve the accuracy of the compliance judgment model in judging the compliance of the qualification terms in the procurement documents, so that the compliance judgment model can better adapt to various complex procurement document scenarios.

7. The intelligent review method for the fairness of procurement document qualifications based on multi-technology integration according to claim 5, characterized in that, The following steps are also included: Step 601: Use text parsing technology to locate and extract the qualification terms in the existing procurement documents; then use dependency syntax analysis technology to analyze the dependency relationships in the qualification terms, identify and save the sentences and paragraphs related to qualifications, and obtain the qualification dataset. Step 602: Label the qualification dataset, label the sentences and paragraphs with enterprise scale restrictions as risky, and other sentences and paragraphs as non-risky; then input this dataset into the pre-trained BERT model for fine-tuning, so that the BERT model can understand the sentences or paragraphs with enterprise scale restrictions. Step 603: Use the BERT model to perform semantic matching on the procurement document to be judged, and extract all the sentences and paragraphs determined to have compliance risks from the procurement document to be judged as compliance risk data. Step 604: Finally, according to the order of appearance of the compliance risk data in the procurement document to be judged, input the compliance risk data into the inspection report in sequence to obtain an inspection report with the ability to analyze compliance risks.

8. The intelligent review method for the fairness of procurement document qualifications based on multi-technology integration according to claim 7, characterized in that: The inspection report includes basic project information, inspection result summary, explicit restriction features, implicit restriction features, and compliance risk data.

Citation Information

Patent Citations

  • Bidding document review method based on artificial intelligence technology, computer device, medium and program product

    CN118536473A

  • Purchase file compliance checking system based on difference algorithm under AI large model

    CN118551760A

  • Green energy power industry purchase file compliance inspection method and system

    CN119067457A

  • Bid evaluation method and system based on artificial intelligence

    CN119295007A

  • Data processing system and method in qualification authentication and management consultation field

    CN119849990A

Cited By

  • Detection task and index matching method and system based on atlas analysis

    CN121117637A