Intelligent Bid Clearing Analysis Method and System
By constructing a content analysis graph using a standard terminology database and conducting similarity evaluation, the problem of identifying bid rigging and collusion in bid clearing analysis was solved, achieving efficient and accurate analysis of bid documents.
Patent Information
- Application Number
- CN202511641012.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-11
- Publication Date
- 2026-03-06
- Estimated Expiration
- 2045-11-11
AI Technical Summary
Existing technologies are difficult to effectively identify and prevent bid rigging and collusion in bid clearing analysis, especially in terms of content analysis, which presents technical difficulties and makes it hard to meet the requirements of high efficiency and accuracy in bid clearing.
The bid documents are extracted and organized using a standard terminology database, a content analysis graph is constructed, and the relationship between related content in the bid documents is identified through similarity evaluation to generate bid evaluation results.
It improves the accuracy of bid clearing analysis, can identify anomalies caused by local modifications and structural adjustments, prevents bid rigging and collusion, and improves the precision of analysis results.
Smart Images

Figure CN121093011B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to intelligent bid evaluation analysis methods and systems for bid documents. Background Technology
[0002] Bid review analysis is a crucial step in the engineering bidding process. It refers to the systematic review, verification, and comparison of bid documents by professionals before bid evaluation to determine their completeness, compliance, accuracy, and reasonableness, providing foundational data and reference for subsequent evaluation. Its core objectives are to eliminate invalid bids, identify potential problems, and clarify discrepancies between bids.
[0003] One of the key objectives of bid clearing analysis is to prevent bid rigging and collusion. Current methods for addressing bid rigging and collusion include content analysis and data tagging analysis (such as submission URLs, web locations, fund flows, and company affiliations). Data tagging is relatively easy to analyze due to its uniqueness; however, content analysis faces significant technical challenges due to issues such as adversarial modifications, insufficient sample size, and ambiguous wording, making it difficult to meet the demands of efficient and accurate bid clearing of a large volume of bid documents. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this application provides an intelligent bid evaluation and analysis method, apparatus, system, and storage medium for bid documents.
[0005] Firstly, this application provides an intelligent bid evaluation analysis method for tender documents, including:
[0006] The standard terms in the standard terminology library are used to extract the associated content of the standard terms. The associated content of each standard term is generated based on a single bid document.
[0007] Organize the related content to obtain a relationship diagram, which includes node content and connection relationships;
[0008] The similarity of two related content relationship graphs is evaluated to obtain the similarity evaluation results. Each related content relationship graph participates in the similarity evaluation.
[0009] The results of the cleanup are given based on the obtained similarity evaluation results.
[0010] In one possible implementation of the first aspect, standard terms from a standard terminology library are used to extract the relevant content from the tender documents, including:
[0011] The standard terms in the standard terminology library are used to extract the paragraph content from the tender document.
[0012] Content analysis graphs are constructed using paragraph content, and these graphs are unidirectional.
[0013] Based on the obtained content analysis map, the paragraph content is screened and organized to obtain the related content of standard words.
[0014] In one possible implementation of the first aspect, constructing a content analysis graph using paragraph content includes:
[0015] Extract entities from paragraph content and determine the relationships between entities;
[0016] The relationships are merged and organized so that the constructed content analysis map does not have a network structure;
[0017] Entities include words and numerical values;
[0018] When a network structure exists in the content analysis graph, multiple connections corresponding to the network structure are evaluated and connection values are obtained. Then, a connection is selected to be retained based on the connection values.
[0019] In one possible implementation of the first aspect, evaluating multiple connections corresponding to the mesh structure and obtaining connection values includes:
[0020] Determine the common starting point of the extended chain and the type of connection relationship on the extended chain. The type includes strong connection relationship and weak connection relationship. The strong connection relationship is assigned a value of one, and the weak connection relationship is assigned a value of zero.
[0021] The connection relationship values are obtained by accumulating them.
[0022] In one possible implementation of the first aspect, similarity evaluation of two related content relationship graphs includes:
[0023] Divide the relationship diagrams of the two related contents to obtain merged regions and completely different regions. The merged regions include the same regions and similar regions.
[0024] Merge identical regions;
[0025] The similarity of similar regions is calculated, and based on the calculation results, similar regions are classified into the same region or completely different regions.
[0026] In one possible implementation of the first aspect, calculating the similarity of similar regions includes:
[0027] Based on the node content and connection relationship, the same or similar parts in the two merged regions are overlapped and the non-overlapping regions are determined;
[0028] Identify the node content associated with the non-overlapping region, where the node content has a direct connection with the non-overlapping region;
[0029] The similarity between two non-overlapping regions is determined based on the content of the nodes associated with them.
[0030] In this process, the content of nodes in non-overlapping regions participates in the similarity calculation in sequence, and in each similarity calculation, only the content of nodes in one non-overlapping region participates in the similarity calculation.
[0031] In one possible implementation of the first aspect, determining the similarity between two non-overlapping regions based on the content of nodes associated with them includes:
[0032] Associate the contents of two nodes that belong to two non-overlapping regions to obtain an associated node content group and an unassociated node content group.
[0033] The judgment result of two non-overlapping areas is determined based on the proportion of content groups of related nodes;
[0034] Among them, association includes content association and attribute association;
[0035] When content association and attribute association are both unavailable, structural association is used to associate the same node content in two nodes that are not associated with the same non-overlapping region.
[0036] Structural associations include:
[0037] Determine the directly related content groups for the two nodes respectively. There are multiple directly related content groups, and each directly related content group includes the content of multiple nodes.
[0038] Analysis lines are generated using node content and corresponding associated content groups. Any two adjacent node contents on the analysis line have a direct connection relationship.
[0039] Based on the similarity of the analysis lines, determine whether the contents of two nodes belonging to two non-overlapping regions are related;
[0040] The relative positions of the two analysis lines can be moved;
[0041] When calculating the similarity between two analysis lines, only it is determined whether the attributes at corresponding positions on the two analysis lines are the same;
[0042] If at least one pair of analysis lines has a similarity that meets the requirements, it indicates that the contents of two nodes belonging to two non-overlapping regions are related.
[0043] Secondly, this application provides an intelligent bid evaluation and analysis device for bid documents, comprising:
[0044] The content extraction unit is used to extract the bid documents using standard terms from the standard terminology library, and obtain the associated content of each standard term. The associated content of each standard term is generated based on a bid document.
[0045] The content organization unit is used to organize related content to obtain a relationship diagram of related content, which includes node content and connection relationships;
[0046] The content evaluation unit is used to evaluate the similarity between two related content relationship graphs and obtain the similarity evaluation result. Each related content relationship graph participates in the similarity evaluation.
[0047] Thirdly, this application provides an intelligent bid evaluation and analysis system for tender documents, the system comprising:
[0048] One or more memories for storing instructions; and
[0049] One or more processors are configured to retrieve and execute the instructions from the memory to perform the methods described in the first aspect and any possible implementation thereof.
[0050] Fourthly, this application provides a computer-readable storage medium, the computer-readable storage medium comprising:
[0051] The program, when run by a processor, is executed as described in the first aspect and any possible implementation thereof.
[0052] This application provides an intelligent bid evaluation analysis method and system for bid documents. It uses content structure analysis to determine similarity, which solves the problem of abnormal results caused by local modifications and structural adjustments in the content evaluation process, and helps to improve the accuracy of the analysis results. Attached Figure Description
[0053] Figure 1 This is a flowchart illustrating the steps of an intelligent bid clearing and analysis method for bid documents provided in this application.
[0054] Figure 2 This is a schematic diagram illustrating the process of obtaining a relational graph of related content, as provided in this application.
[0055] Figure 3 This is a schematic diagram illustrating the pairwise comparison of related content provided in this application.
[0056] Figure 4 This is a schematic diagram of a similarity evaluation result displayed in a table, as provided in this application.
[0057] Figure 5This is a schematic diagram of a method for removing a mesh structure provided in this application.
[0058] Figure 6 This is a schematic diagram of another method for removing the mesh structure provided in this application.
[0059] Figure 7 This is a schematic diagram illustrating how this application determines the node content associated with non-overlapping regions.
[0060] Figure 8 This is a schematic diagram of an analysis pair using analysis lines provided in this application. Detailed Implementation
[0061] The technical solutions in this application will be further described in detail below with reference to the accompanying drawings.
[0062] This application discloses an intelligent bid evaluation and analysis method for bid documents. Please refer to [link / reference]. Figure 1 In some examples, the intelligent bid clearing analysis method for bid documents disclosed in this application includes the following steps:
[0063] S101, use standard terms from the standard terminology library to extract the bid documents and obtain the associated content of the standard terms. The associated content of each standard term is generated based on a bid document.
[0064] S102, Organize the related content to obtain a related content relationship diagram, which includes node content and connection relationships;
[0065] S103, perform similarity evaluation on any two related content relationship graphs belonging to the same standard word, and obtain similarity evaluation results. Each related content relationship graph participates in the similarity evaluation.
[0066] S104, Based on the obtained similarity evaluation results, the clearing results are given.
[0067] The technical solution in this application involves similarity analysis of multiple bids for the same subject matter. The purpose is to determine precisely whether these bids involve bid rigging or collusion. Therefore, the bid analysis using this technical solution is applied to bids after preliminary routine analysis, such as excluding invalid bids based on compliance checks or excluding completely identical bids based on deduplication. For example:
[0068] Compliance check: In accordance with relevant laws, regulations and the requirements of the bidding documents, check whether the bid documents meet all the necessary conditions in order to identify invalid bids;
[0069] Automated scoring: Based on the set standards and weights, the system automatically scores various indicators of the tender documents, such as price, technical capabilities, and commercial terms;
[0070] Risk warning: Identify potential risk points, such as abnormally low prices, unreasonable technical parameters, and related relationships;
[0071] The main contents of the preliminary review of the technical solution are as follows:
[0072] Technical solution anomalies, including plagiarism anomalies, are addressed through comparative plagiarism detection methods. For details, please refer to the methods for checking thesis plagiarism.
[0073] If the relevance (irrelevant content) is abnormal, such as the technical proposal prepared by the bidder being completely unrelated to the category of the bidding project, for example, if the bidder submits a technical proposal on municipal roads in a building construction project or a technical proposal on decoration in a new building construction project, this part will be directly marked, such as by displaying it with a special color, to facilitate review by the reviewers.
[0074] Next, using the technical solution provided in this application, an in-depth bid-clearing analysis will be conducted:
[0075] Specifically, in step S101, the standard terms in the standard terminology library are first used to extract the bid documents. At this time, the associated content of the standard terms will be obtained. The extraction method is that when a paragraph in the bid document contains the standard term, then this paragraph is put into the associated content of the standard term.
[0076] Standard terminology is generally determined based on the tender documents, specifically describing content strongly related to the subject matter. Taking furniture procurement tenders as an example, the standard terminology includes terms such as process, materials, and parameters. Furthermore, the standard terms in the terminology also need to be determined according to the specific subject matter; that is, the standard terms in the terminology will differ for different subjects. For example, in furniture procurement tenders, standard terms for process include edge banding and coating processes; standard terms for materials include board type, solid wood type, environmental protection level, and hardware brand; and parameters include dimensions, load-bearing capacity, and formaldehyde emission.
[0077] For example, a suitable extraction method is to use standard words for location and then extract the entire paragraph containing the standard words.
[0078] The purpose of using a standard thesaurus is to provide an index that can categorize the content of tender documents.
[0079] In addition, the associated content for each standard term is generated based on a tender document. For example, if there are five tender documents, then there will be five associated contents for the standard terms.
[0080] In step S102, the related content is organized to obtain a related content relationship diagram. The related content relationship diagram is a structural display of the related content. Specifically, the related content relationship diagram includes two parts: node content and connection relationship, which reflects the structural framework of the related content.
[0081] In step S103, the similarity of any two related content graphs belonging to the same standard word is evaluated to obtain a similarity evaluation result. The similarity evaluation result is a value less than one, which is used to evaluate the similarity between two related contents generated based on the same standard word.
[0082] There is no restriction on the specific value of this number here. In the actual implementation, it can be set according to the specific requirements of different bidding scenarios.
[0083] This section requires that every related content in the relationship graph participate in the similarity evaluation. For example, if a standard word mentioned earlier has five related contents, then these five related contents need to be compared pairwise. Figure 3 As shown, five similarity evaluation results (numerical values) will be obtained at this time.
[0084] The above method is equivalent to dividing a bid document into multiple different parts and then evaluating these different parts separately. The advantage of this method is that it avoids the situation where two bid documents have similar (copied) content but pass the inspection because the final result value is lower, thus bypassing the bid clearing analysis.
[0085] In addition, strategies that could reduce overall similarity through "partial modification + core copying" can also be avoided. For example, only some non-core content of the tender document is modified, but the core technical solution is completely copied.
[0086] This kind of "partial similarity" recognition requires the model to accurately locate the similarities in the core areas. However, if the core areas themselves account for a small proportion of the document, such as in an IT project involving model architecture where the "algorithm model architecture" in the tender document only accounts for 5% of the document, the AI is easily interfered with by a large number of non-core modifications during processing, resulting in missed similarity judgments.
[0087] This approach leverages the role of a standard terminology database, which can directly extract the necessary content from the tender documents while ignoring other content, thus significantly reducing the impact of non-core information.
[0088] Finally, in step S104, the bid clearing results are given based on the obtained similarity evaluation results. These results are typically displayed in a table. For example, one bid document might have ten related entries based on standard keywords, while other bid documents might have four. A table is created accordingly. Figure 4 As shown, the similarity evaluation results are displayed in the corresponding position in the middle of the table.
[0089] For simplification, similarity is represented by 1 and dissimilarity by blank space. This table can be used to show staff the information, and then staff can further verify and process the specific content. It can also be used to give scores directly. In the example, the total score is 40 points. The higher the score (the sum of all the 1s), the higher the possibility of bid rigging or collusion in the bid.
[0090] In some examples, the specific methods for extracting standard terms from the standard terminology library into the tender documents and obtaining the associated content of the standard terms are as follows:
[0091] S201, use standard terms from the standard terminology library to extract paragraph content from the tender document;
[0092] S202, Use paragraph content to construct a content analysis graph, the content analysis graph is unidirectional;
[0093] S203. Based on the obtained content analysis map, the paragraph content is screened and organized to obtain the related content of standard words.
[0094] In steps S201 to S203, standard terms from the standard terminology database are first used to obtain paragraph content. The first paragraphs obtained are those containing the standard terms. However, there may also be paragraphs that do not contain the standard terms but are associated with them. For example, some content may require multiple paragraphs. The processing method for multiple paragraphs is to use a judgment rule:
[0095] The same theme exists;
[0096] There is a logical relationship (cause and effect, progression, comparison, supplementation, example, etc.);
[0097] There is a correspondence relationship (repeated keywords, synonym substitution, referential relationship, etc.);
[0098] There are structural relationships (general-specific, specific-general, parallel, progressive, etc.).
[0099] After obtaining the paragraph content, a content analysis graph is constructed using this content. This graph needs to be unidirectional, meaning it always extends in one direction, such as from a starting point to the right. The advantage of this technique is that it clearly defines a starting point and allows the content analysis graph to have a clear hierarchical relationship.
[0100] Finally, based on the obtained content analysis map, the paragraph content is filtered and organized to obtain the related content of standard words. The purpose of filtering and organizing is to remove duplicate and similar content. For example, the paragraph content consists of multiple sentences (divided by periods). The sentences can be extracted into the main body, and these main bodies can be connected using relationships.
[0101] In this application, the relationships between subjects include only parallel and hierarchical relationships. For two adjacent sentences, the judgment rules described above can be used to determine whether a relationship exists, thereby determining whether a relationship exists between the subjects in these two sentences.
[0102] After processing all the sentences, similar and related sentences are merged, taking into account the presence of duplicate content, different representations of the same content, or the same content after being split and then distributed.
[0103] Please see Figure 5 The specific method for constructing a content analysis map using paragraph content is as follows:
[0104] Extract entities from paragraph content and determine the relationships between entities;
[0105] The relationships are merged and organized so that the constructed content analysis map does not have a network structure;
[0106] Entities include words and numerical values;
[0107] When a network structure exists in the content analysis graph, multiple connections corresponding to the network structure are evaluated and connection values are obtained. Then, a connection is selected to be retained based on the connection values.
[0108] The requirement that the constructed content analysis graph does not have a network structure is to address the issue that a closed network structure would prevent subsequent similarity determination.
[0109] This is because the application requires explicit linear relationships (A and B). When a network structure appears, it will directly lead to non-linear relationships (A and B, A and C). In this case, how to determine the relationship requires consideration of more factors (position, context, related content, etc.). Too many influencing factors will directly lead to misjudgment, or even an inability to make a judgment.
[0110] A, B, and C refer to the entities extracted from the paragraph content as described in the preceding steps, such as... Figure 5 and Figure 6As shown, entities extracted from paragraph content are represented by rectangles. For ease of understanding, the three rectangles are represented by A, B, and C respectively, to explain linear and non-linear relationships.
[0111] When a network structure exists, multiple connections corresponding to the network structure are evaluated and connection values are obtained. Then, based on the connection values, one connection is selected to be retained, evaluated, and its value is obtained again. The specific process is as follows:
[0112] Determine the common starting point of the extended chain and the type of connection relationship on the extended chain. The type includes strong connection relationship and weak connection relationship. The strong connection relationship is assigned a value of one (1), and the weak connection relationship is assigned a value of zero (0). The connection relationship value is obtained by accumulation.
[0113] The above method assigns values based on the type of connection relationship on the extended chain (a strong connection is assigned a value of one, and a weak connection is assigned a value of zero), and then obtains the connection relationship value based on the cumulative result of the assignments.
[0114] For example, if A and B are related, and A and C are also related, then one needs to be retained. First, based on B and C, the content analysis graph is extended to... Figure 5 For reference, we need to extend to the left to determine whether the common starting point is the same or there are two.
[0115] When there is only one common starting point, it means that B and C have the same origin. At this time, the relationships between the multiple subjects involved in the extension process are statistically analyzed. Strong connections are assigned a value of one, and weak connections are assigned a value of zero. Strong connections refer to explicit hierarchical relationships, such as (superior and subordinate, whole and part, cause and effect, etc.). Weak connections refer to relationships that are in the same sentence but are not strong connections.
[0116] Finally, the connectivity values are accumulated. When B and C have the same origin, the one with the higher connectivity value is retained; when the values are the same, one is selected or randomly retained. When B and C have different origins (e.g., ... Figure 6 When (as shown by the dashed line), the quantity of A is increased to two, and these two A's are assigned to B and C respectively, as shown in the example. Figure 6 As shown. This processing method will still result in branching structures to some extent. If the branching structure leads to a network structure, the above method will continue to be used until the network structure disappears.
[0117] In some examples, the similarity evaluation of two related content relationship graphs is performed as follows:
[0118] S301, divide the relationship diagrams of the two related contents to obtain merged regions and completely different regions. The merged regions include the same regions and similar regions.
[0119] S302, Merge identical areas;
[0120] S303 calculates the similarity of similar regions and, based on the calculation results, assigns similar regions to the same region or completely different regions. That is, it re-divides the current similar regions in the merged regions obtained based on the relationship graph based on the similarity.
[0121] The content in steps S301 to S303 is to determine whether the corresponding content in the two related content relationship diagrams is similar. Specifically, the two related content relationship diagrams are first divided into merged regions and completely different regions. The merged regions are divided into two types: identical regions and similar regions.
[0122] The above method can also be described as dividing the related content relationship graph into parts that are exactly the same as another related content relationship graph (same area), parts that are not exactly the same as another (similar area), and parts that are completely different (completely different area).
[0123] In other words, a relationship diagram of related content can be divided into three parts.
[0124] In a specific processing procedure, we can first identify an A and multiple Bs that have a direct connection relationship with A, where A is the superior of B. Then, we make a judgment: if there are identical A and B in another relationship graph, then that part is determined to be the same area. If A is different, but B is completely or partially the same, then it is determined to be the similar area. The remaining parts are all classified into completely different areas.
[0125] Merge identical regions;
[0126] The similarity of similar regions is calculated, and based on the calculation results, similar regions are classified into the same region or completely different regions.
[0127] Its main purpose is to address the specific application scenarios where there are different expressions of the same word in many professional fields. This is mainly due to the limited corpus in professional fields, coupled with insufficient model training, which means that in the word processing of professional fields, the only way to distinguish between them is by similarity and difference. This allows us to avoid this by changing the expression.
[0128] This is why similar regions are generated.
[0129] For example, this approach can also be used in the process of processing related content to address the problem of a standard term being described in multiple ways and divided into multiple paragraphs.
[0130] The same region refers to the existence of at least two levels, such as one A corresponding to multiple Bs, and another region (with two levels, one A corresponding to multiple Bs) where the descriptions of A and B are the same.
[0131] Similar regions refer to regions with at least two levels, such as one A corresponding to multiple Bs. There is another region (with two levels, one A corresponding to multiple Bs) where the descriptions of B are the same or partially the same, while the descriptions of A are different.
[0132] To address this issue, this application employs a method based on structural similarity for determination. Specifically, the two related content relationship graphs are first divided into merged regions and completely different regions. The merged regions include identical regions and similar regions. Here, identical regions refer to two identical nodes (on the content analysis graph) or nodes that express the same meaning (on the content analysis graph).
[0133] Next, the similarity of similar regions is calculated, as follows:
[0134] Based on the node content and connection relationship, the same or similar parts in the two merged regions are overlapped and the non-overlapping regions are determined;
[0135] Identify the node content associated with the non-overlapping region, where the node content has a direct connection with the non-overlapping region;
[0136] The similarity between two non-overlapping regions is determined based on the content of the nodes associated with them.
[0137] In this process, the content of nodes in non-overlapping regions participates in the similarity calculation in sequence, and in each similarity calculation, only the content of nodes in one non-overlapping region participates in the similarity calculation.
[0138] In the above method, it is first necessary to determine the non-overlapping regions, and then determine the node content associated with the non-overlapping regions. Figure 6 Within the dashed-lined area (in the diagram), the node content has a direct connection with the non-overlapping area, such as... Figure 6 As shown, this clearly illustrates the superior-subordinate relationship.
[0139] Finally, the similarity between two non-overlapping regions is determined based on the content of the nodes associated with them. This application provides the following method:
[0140] Associate the contents of two nodes that belong to two non-overlapping regions to obtain an associated node content group and an unassociated node content group.
[0141] The judgment result of two non-overlapping areas is determined based on the proportion of content groups of related nodes;
[0142] In the above method, the associated node content group and the non-associated node content group each include at least one node content. For example, if the total number of node content is ten and the total number of node content in the associated node content group is seven, then the proportion value of the associated node content group is 0.7. Finally, the similarity between two non-overlapping regions is determined based on the node content associated with the non-overlapping regions. This step is a comparison process, that is, comparing 0.7 with a set value. If 0.7 is greater than or equal to the set value, the two non-overlapping regions are judged to be the same; otherwise, they are different.
[0143] In the above methods, association includes two types: content association and attribute association. That is, when the content is different, attributes can be used to make a judgment.
[0144] When content association and attribute association are both unavailable, structural association is used to associate the same node content between two nodes that are not associated with the same non-overlapping region. The specific method is as follows:
[0145] Determine the directly related content groups for the two nodes respectively. There are multiple directly related content groups, and each directly related content group includes the content of multiple nodes.
[0146] Analysis lines are generated using node content and corresponding associated content groups. Any two adjacent node contents on the analysis line have a direct connection relationship.
[0147] Based on the similarity of the analysis lines, determine whether the contents of two nodes belonging to two non-overlapping regions are related.
[0148] The relative positions of the two analysis lines can be moved;
[0149] When calculating the similarity between two analysis lines, only it is determined whether the attributes at corresponding positions on the two analysis lines are the same;
[0150] When there is at least one pair of analysis lines whose similarity meets the requirements, it is determined that the contents of two nodes belonging to two non-overlapping regions are related, that is, it is determined that the contents of two nodes in two non-overlapping regions are related.
[0151] In the above method, it is first necessary to determine the directly related content groups of the two node contents. There are multiple directly related content groups, and each directly related content group includes multiple node contents. The directly related content groups are the same according to the related content relationship diagram.
[0152] Next, use the node content and the corresponding associated content group to generate analysis lines, such as... Figure 7As shown, any two adjacent nodes on the analysis line have a direct connection relationship. Then, the analysis line is generated using the node content and the corresponding associated content group. Here, it is required that any two adjacent nodes on the analysis line have a direct connection relationship.
[0153] Finally, based on the similarity of the analysis lines, it is determined whether the content of two nodes belonging to two non-overlapping regions is related. The method for determining the similarity of the analysis lines is to determine whether the content of the corresponding nodes on the analysis lines is related by attributes. Here, it is necessary that the attributes of each node on the shorter first analysis line are the same as the attributes of the corresponding node on the other analysis line.
[0154] During the judgment process, the relative positions of the two analysis lines can be moved, such as... Figure 7 As shown.
[0155] In the above process, when at least one pair of analysis lines meets the similarity requirement, the contents of two nodes belonging to two non-overlapping regions are related.
[0156] After processing in the above way, we can determine whether the two related content relationship graphs are similar (based on the ratio of the same area to the sum of the same area and completely different areas, with the area represented by the number of included nodes), and then we can obtain the table mentioned above.
[0157] However, it should be noted that using the technical solution in this application to judge the standard content will result in a large number of identical results. However, the differences in the standard content for different tender documents are mostly concentrated in the performance parameters.
[0158] Therefore, the technical solution in this application mainly focuses on the analysis of non-standard content. The non-standard content in different tender documents should be completely different or partially the same. Furthermore, the non-standard content can be extracted by using a standard terminology library, or the standard content can be deleted by using a standard terminology library, with the remaining content being the non-standard content.
[0159] This application also provides an intelligent bid evaluation and analysis device for bid documents, including:
[0160] The content extraction unit is used to extract the bid documents using standard terms from the standard terminology library, and obtain the associated content of each standard term. The associated content of each standard term is generated based on a bid document.
[0161] The content organization unit is used to organize related content to obtain a relationship diagram of related content, which includes node content and connection relationships;
[0162] The content evaluation unit is used to evaluate the similarity between any two related content graphs belonging to the same standard term and obtain the similarity evaluation result. Each related content graph participates in the similarity evaluation.
[0163] The result output unit is used to provide the clearing result based on the obtained similarity evaluation result.
[0164] Furthermore, using standard terms from the standard terminology database, the tender documents are extracted, and the associated content of the standard terms is obtained, including:
[0165] The standard terms in the standard terminology library are used to extract the paragraph content from the tender document.
[0166] Content analysis graphs are constructed using paragraph content, and these graphs are unidirectional.
[0167] Based on the obtained content analysis map, the paragraph content is screened and organized to obtain the related content of standard words.
[0168] Furthermore, constructing content analysis maps using paragraph content includes:
[0169] Extract entities from paragraph content and determine the relationships between entities;
[0170] The relationships are merged and organized so that the constructed content analysis map does not have a network structure;
[0171] Entities include words and numerical values;
[0172] When a network structure exists in the content analysis graph, multiple connections corresponding to the network structure are evaluated and connection values are obtained. Then, a connection is selected to be retained based on the connection values.
[0173] Furthermore, the multiple connections corresponding to the network structure are evaluated, and the connection values are obtained, including:
[0174] Determine the common starting point of the extended chain and the type of connection relationship on the extended chain. The type includes strong connection relationship and weak connection relationship. The strong connection relationship is assigned a value of one, and the weak connection relationship is assigned a value of zero.
[0175] The connection relationship values are obtained by accumulating them.
[0176] Furthermore, the similarity evaluation of the two related content relationship graphs includes:
[0177] Divide the relationship diagrams of the two related contents to obtain merged regions and completely different regions. The merged regions include the same regions and similar regions.
[0178] Merge identical regions;
[0179] The similarity of similar regions is calculated, and based on the calculation results, similar regions are classified into the same region or completely different regions.
[0180] Furthermore, the calculation of the similarity of similar regions includes:
[0181] Based on the node content and connection relationship, the same or similar parts in the two merged regions are overlapped and the non-overlapping regions are determined;
[0182] Identify the node content associated with the non-overlapping region, where the node content has a direct connection with the non-overlapping region;
[0183] The similarity between two non-overlapping regions is determined based on the content of the nodes associated with them.
[0184] In this process, the content of nodes in non-overlapping regions participates in the similarity calculation in sequence, and in each similarity calculation, only the content of nodes in one non-overlapping region participates in the similarity calculation.
[0185] Furthermore, determining the similarity between two non-overlapping regions based on the content of nodes associated with them includes:
[0186] Associate the contents of two nodes that belong to two non-overlapping regions to obtain an associated node content group and an unassociated node content group.
[0187] The judgment result of two non-overlapping areas is determined based on the proportion of content groups of related nodes;
[0188] Among them, association includes content association and attribute association;
[0189] When content association and attribute association are both unavailable, structural association is used to associate the same node content in two nodes that are not associated with the same non-overlapping region.
[0190] Structural associations include:
[0191] Determine the directly related content groups for the two nodes respectively. There are multiple directly related content groups, and each directly related content group includes the content of multiple nodes.
[0192] Analysis lines are generated using node content and corresponding associated content groups. Any two adjacent node contents on the analysis line have a direct connection relationship.
[0193] Based on the similarity of the analysis lines, determine whether the contents of two nodes belonging to two non-overlapping regions are related;
[0194] The relative positions of the two analysis lines can be moved;
[0195] When calculating the similarity between two analysis lines, only it is determined whether the attributes at corresponding positions on the two analysis lines are the same;
[0196] When there is at least one pair of analysis lines whose similarity meets the requirements, the contents of two nodes belonging to two non-overlapping regions are related.
[0197] In one example, the unit in the above device may be one or more integrated circuits configured to implement the above methods, such as: one or more application-specific integrated circuits (ASICs), or one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs), or a combination of at least two of these integrated circuit forms.
[0198] For example, when the units in the device can be implemented through a processing element scheduler, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processor capable of calling programs. Alternatively, these units can be integrated together to form a system-on-a-chip (SOC).
[0199] In this application, various objects such as messages / information / devices / network elements / systems / apparatus / actions / operations / processes / concepts may be named. It is understood that these specific names do not constitute a limitation on the relevant objects. The names may be changed depending on the scenario, context, or usage habits. The understanding of the technical meaning of the technical terms in this application should be mainly determined from their functions and technical effects embodied / performed in the technical solution.
[0200] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0201] In light of the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0202] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0203] It should also be understood that in the various embodiments of this application, the terms "first," "second," etc., are merely to indicate that multiple objects are different. The aforementioned terms "first," "second," etc., should not impose any limitation on the embodiments of this application.
[0204] It should also be understood that, in the various embodiments of this application, unless otherwise specified or in case of logical conflict, the terms and / or descriptions between different embodiments are consistent and can be referenced by each other, and the technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationships.
[0205] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a computer-readable storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned computer-readable storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0206] This application also provides an intelligent bid evaluation and analysis system for bid documents, the system comprising:
[0207] One or more memories for storing instructions; and
[0208] One or more processors are configured to retrieve and execute the instructions from the memory, performing the methods described above.
[0209] This application also provides a computer-readable storage medium containing a computer program that, when executed, causes the terminal device and the network device to perform operations corresponding to the methods described above.
[0210] The embodiments described in this specific implementation are preferred embodiments of this application and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.
Claims
1. A method for intelligent bid evaluation and analysis of tender documents, characterized in that, The method comprises the following steps: extracting the bid documents using standard terms in a standard term library to obtain associated content of the standard terms, the associated content of each standard term being generated based on a bid document; organizing the associated content to obtain an associated content relationship diagram, the associated content relationship diagram comprising node content and connection relationships; evaluating the similarity of two associated content relationship diagrams to obtain a similarity evaluation result, each associated content relationship diagram participating in the similarity evaluation; the similarity evaluation of the two associated content relationship diagrams comprises: merging the same areas and calculating the similarity of the similar areas, and according to the calculation result, the similar areas are divided into the same areas or completely different areas, wherein: the similarity calculation of the similar areas comprises: determining the non-overlapping areas by overlapping the same or similar parts in the two merged areas and determining the non-overlapping areas according to the node content and connection relationships; determining the node content associated with the non-overlapping areas, which has a direct connection relationship with the non-overlapping areas; determining the similarity of the two non-overlapping areas according to the node content associated with the non-overlapping areas, comprising: associating two node contents belonging to the two non-overlapping areas respectively to obtain an associated node content group and a non-associated node content group, and determining the judgment result of the two non-overlapping areas according to the proportion of the associated node content group; wherein the node content in the non-overlapping area participates in the similarity calculation in sequence, and only the node content in one non-overlapping area participates in the similarity calculation each time; 2. The method of intelligent de-bid analysis of bid documents of claim 1, wherein, wherein the association includes content association and attribute association; when both content association and attribute association cannot be used, structure association is used to associate the same node content in the two node contents associated with the non-overlapping areas; giving a clear label result according to the obtained similarity evaluation result. extracting the bid documents using standard terms in a standard term library to obtain associated content of the standard terms, comprising: extracting the bid documents using standard terms in a standard term library to obtain associated content of the standard terms, comprising:
3. The method of intelligent de-bid analysis of bid documents of claim 2, wherein, extracting the bid documents using standard terms in a standard term library to obtain associated content of the standard terms, comprising: screening and organizing the paragraph content according to the obtained content analysis graph to obtain the associated content of the standard terms. constructing a content analysis graph using the paragraph content, comprising: extracting entities in the paragraph content and determining the association relationship between the entities; merging and organizing the association relationship to make the content analysis graph constructed not exist a network structure; 4. The method of claim 3, wherein, wherein the entity includes a term and a numerical value; when the content analysis graph has a network structure, evaluating a plurality of connection relationships corresponding to the network structure to obtain a connection relationship value, and then selecting one connection relationship according to the connection relationship value to be retained. evaluating a plurality of connection relationships corresponding to the network structure to obtain a connection relationship value, comprising: Determine the common starting point of the extended chain and the type of connection relationship on the extended chain, the type including strong connection relationship and weak connection relationship, the strong connection relationship being assigned a value of one and the weak connection relationship being assigned a value of zero; Obtain the connection relationship value by accumulation.
5. The method of claim 1, wherein, The structural association includes: Determine the direct association content group of the two node contents respectively, the number of the direct association content group being multiple, and the direct association content group including multiple node contents; Use the node content and the corresponding association content group to generate an analysis line, and any two adjacent node contents on the analysis line have a direct connection relationship; Determine whether the two node contents belonging to the two non-overlapping areas respectively have an association according to the similarity of the analysis line; The relative position of the two analysis lines can be moved; When calculating the similarity of the two analysis lines, only determine whether the attributes at the corresponding positions of the two analysis lines are the same; When the similarity of at least one pair of analysis lines meets the requirements, it is determined that the two node contents belonging to the two non-overlapping areas respectively have an association.
6. An intelligent bid document de-bid analysis apparatus, characterized by, It includes: A content extraction unit is configured to extract a bid document using standard terms in a standard term library to obtain associated content of the standard terms, and the associated content of each standard term is generated based on one bid document; A content arrangement unit is configured to arrange the associated content to obtain an associated content relationship diagram, and the associated content relationship diagram includes node content and connection relationship; A content evaluation unit is configured to evaluate the similarity of two associated content relationship diagrams to obtain a similarity evaluation result, and each associated content relationship diagram participates in the similarity evaluation; The similarity evaluation of the two associated content relationship diagrams includes: dividing the two associated content relationship diagrams to obtain merged areas and completely different areas, and the merged areas include identical areas and similar areas; The identical areas are merged, the similarity of the similar areas is calculated, and the similar areas are divided into the identical areas or the completely different areas according to the calculation result, wherein the similarity of the similar areas is calculated by: According to the node content and the connection relationship, the same or similar parts in the two merged areas are overlapped and the non-overlapping areas are determined, the node content associated with the non-overlapping areas is determined, and the node content has a direct connection relationship with the non-overlapping areas; The similarity of the two non-overlapping areas is determined according to the node content associated with the non-overlapping areas, including: associating two node contents belonging to the two non-overlapping areas respectively to obtain an associated node content group and a non-associated node content group; and determining a judgment result of the two non-overlapping areas according to the proportion of the associated node content group; Wherein, the node contents in the non-overlapping areas participate in the similarity calculation in sequence, and only the node contents in one non-overlapping area participate in the similarity calculation each time; Wherein, the association includes content association and attribute association; when the content association and the attribute association cannot be used, the structural association is used to associate the same node contents in the two node contents associated with the non-overlapping areas; A result output unit is configured to output a clear label result according to the obtained similarity evaluation result.
7. An intelligent bid sheet analysis system for bid documents, characterized by, The system includes: one or more memories for storing instructions; and one or more processors for invoking and running the instructions from the memories, implementing the method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The computer readable storage medium includes: a program that, when run by a processor, implements the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
An industrial big data integration system based on ontology fusion
CN109635119A
Credit product intelligent recommendation method and system based on multi-source data fusion
CN119722296A