Method and device for detecting bid and string bid for bid procurement
Through the multi-dimensional similarity detection method, the correlation between bidding documents can be quickly identified, which solves the problem of difficult identification of bid rigging and collusion in traditional methods and achieves efficient and accurate detection results.
Patent Information
- Application Number
- CN202510736722.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-12
AI Technical Summary
In existing technologies, bid rigging and collusion are difficult to identify quickly and accurately. Traditional methods are inefficient and easily affected by subjective factors. They lack in-depth analysis capabilities and cannot effectively identify similarities between bid documents.
By collecting bidding document information, constructing a data set of bidding address information, high-frequency word information and chapter content, multi-dimensional similarity detection is performed, including a first similarity detection of the bidding address information set, a second similarity detection of the high-frequency word information set and a third similarity detection of the chapter content data set, forming a progressive detection chain.
It improves the accuracy and efficiency of bid rigging and collusion detection, can quickly identify related bidders, reduce manpower and time costs, provide a quantitative evaluation basis, and avoid the limitations of subjective judgment.
Smart Images

Figure CN120634019A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bidding and procurement and the field of data processing, and in particular to a method and device for detecting bid rigging and collusion in bidding and procurement. Background Art
[0002] In bidding and procurement activities, bid rigging and collusion are common forms of unfair competition. Bid rigging and collusion not only undermine fair competition in the bidding market but can also compromise the quality of tendered projects, resulting in significant financial losses for tendering parties. Traditional methods for detecting bid rigging and collusion rely primarily on manual review, requiring reviewers to individually examine numerous bid documents, including their content and background information. However, this approach has numerous limitations. Firstly, manual review is inefficient and difficult to process within a short timeframe. Secondly, manual review is susceptible to subjective factors, making it difficult to accurately identify covert bid rigging and collusion. For example, bidders may conceal their connections through carefully crafted document content or avoid detection by using different IP and MAC addresses. Furthermore, traditional detection methods lack the ability to conduct in-depth analysis of bid document content and are unable to effectively identify similarities between bids in terms of technical solutions, commercial terms, and other aspects, making it difficult to promptly detect and effectively curb bid rigging and collusion. Summary of the Invention
[0003] The present invention mainly solves the problem of how to quickly and accurately detect bid rigging and collusion based on bidding document information. The present invention discloses a bid rigging and collusion detection method and device for bidding and procurement.
[0004] In a first aspect, an embodiment of the present invention discloses a method for detecting bid rigging in tender procurement, comprising:
[0005] S1, collecting a set of bidding documents for tender procurement; the set of bidding documents includes a plurality of bidding documents;
[0006] S2, extracting and processing the bidding document set to obtain a bidding address information set, a high-frequency word information set, and a chapter content data set;
[0007] S3, performing bid rigging detection processing on the bidding address information set, high-frequency word information set and chapter content data set to obtain a detection result information set.
[0008] The extracting process of the bidding document set to obtain a bidding address information set, a high-frequency word information set, and a chapter content data set includes:
[0009] S21, for each bidding document in the bidding document set, obtaining corresponding bidding address information; the bidding address information includes an IP address and a Mac address;
[0010] S22, constructing a bidding address information set using the bidding address information of all bidding documents; each bidding address information in the bidding address information set has a corresponding bidding document;
[0011] S23, extracting a corresponding high-frequency word information subset for each bid document in the bid document set; the high-frequency word information subset includes high-frequency word information; the high-frequency word information is a word vector whose number of occurrences exceeds a preset number threshold;
[0012] S24, constructing a high-frequency word information set using the high-frequency word information subsets of all bidding documents; each high-frequency word information in the high-frequency word information set has a corresponding bidding document;
[0013] S25, extracting corresponding chapter content data for each bidding document in the bidding document set;
[0014] S26 , constructing a chapter content data set using the chapter content data of all bidding documents; each chapter content data of the chapter content data set has a corresponding bidding document.
[0015] The bidding address information set, the high-frequency word information set, and the chapter content data set are subjected to bid rigging detection processing to obtain a detection result information set, including:
[0016] S31, performing a first similarity check on the bidding document set based on the bidding address information set to obtain a first target information set, and performing an update process on the bidding document set;
[0017] S32, performing a second similarity check on the updated bidding document set based on the high-frequency word information set, obtaining a second target information set and a feedback value set, and updating the bidding document set;
[0018] S33: Based on the chapter content data set and the feedback value set, a third similarity test is performed on the updated bidding document set to obtain a test result information set.
[0019] The step of performing a first similarity check on the bidding document set based on the bidding address information set to obtain a first target information set and performing an update process on the bidding document set includes:
[0020] S311, based on the bidding address information set, searching for bidding documents with the same bidding address information in the bidding document set;
[0021] S312, constructing a first target information set using all bidding documents having the same bidding address information;
[0022] S313: Delete all bidding documents included in the first target information set from the bidding document set to obtain an updated bidding document set, thereby completing the updating process of the bidding document set.
[0023] The method of performing a second similarity check on the updated bidding document set based on the high-frequency word information set to obtain a second target information set and a feedback value set, and updating the bidding document set includes:
[0024] S321, performing word similarity calculation on each combination of two high-frequency word information subsets in the high-frequency word information set, to obtain a word similarity value of the combination and first feedback values of two bid documents corresponding to the combination;
[0025] S322, deleting the combinations whose word similarity values are greater than a preset first similarity threshold from the updated bid document set, thereby completing the update process of the bid document set and obtaining an updated bid document set;
[0026] S323: Perform statistical average calculation on all calculated first feedback values of each bidding document to obtain the feedback value of the bidding document.
[0027] The step of performing word similarity calculation on each combination of two high-frequency word information subsets in the high-frequency word information set to obtain a word similarity value of the combination and first feedback values of two bidding documents corresponding to the combination includes:
[0028] S3211, for each high-frequency word information subset in a combination of every two high-frequency word information subsets in the high-frequency word information set, sort all high-frequency word information contained in the subset according to the number of occurrences of the high-frequency word information from largest to smallest, to obtain a high-frequency word information sequence corresponding to the subset; the high-frequency word information sequence includes high-frequency word information;
[0029] S3212, performing high-frequency word similarity calculation processing on two high-frequency word information sequences corresponding to the combination of the two high-frequency word information subsets to obtain a word similarity value of the combination;
[0030] S3213: Calculate and obtain first feedback values of the two bidding documents corresponding to the combination based on the word similarity value of the combination.
[0031] The expression for calculating the similarity of high-frequency words is:
[0032]
[0033] Among them, wg1 ij The jth element of the ith high-frequency word information of the first high-frequency word information sequence corresponding to the combination of the high-frequency word information subset, wg2 ij The jth element of the ith high-frequency word information of the second high-frequency word information sequence corresponding to the combination of the high-frequency word information subset, wg1 i0 is the mean value of the ith high-frequency word information of the first high-frequency word information sequence corresponding to the combination of high-frequency word information subsets, wg2 i0 is the mean of the i-th high-frequency word information of the second high-frequency word information sequence corresponding to the combination of high-frequency word information subsets, gs is the word similarity value of the combination, N and M are the number of elements of the high-frequency word information and the number of high-frequency word information contained in the high-frequency word information sequence, respectively.
[0034] The expression of the high-frequency word similarity calculation process is the element-level difference of the high-frequency word vector through the double summation formula An inverse trigonometric transformation is performed to map numerical differences to the interval [0, 1]. This avoids the dimension-sensitivity issue caused by direct subtraction and makes the features of high-frequency words of different magnitudes comparable. By introducing a mean term, the overall distribution of high-frequency word vectors is normalized, eliminating the influence of absolute differences in the number of occurrences of high-frequency words (e.g., a word appears 10 times in file A and 5 times in file B) on similarity, and focusing more on the matching of relative frequency patterns.
[0035] By using sine and exponential decay functions, we can capture nonlinear correlations between high-frequency word vectors (such as proportional relationships and fluctuation trends). Compared with traditional linear methods such as cosine similarity, we can more sensitively detect the "plagiarism in disguise" commonly seen in bid rigging (such as adjusting the order of words or replacing synonyms while retaining the core semantic structure).
[0036] In step S3211, high-frequency words are sorted by their frequency of occurrence. The formula prioritizes the similarity of core words with higher frequency of occurrence, which aligns with the behavioral characteristics of bid rigging, where bidders plagiarize core terms in bidding documents or collude to fabricate key content. For example, if the high-frequency word sequences in core technical parameters or commercial terms across multiple bid documents are highly similar, even if minor content differs, they will be judged as highly similar.
[0037] In step S3213, the first feedback value is calculated in combination with the word similarity value, and the high-frequency word similarity can be cross-validated with the multi-dimensional detection results such as the bidding address and chapter content (such as the cascade detection process of S31-S33), forming a progressive detection chain of "address association → text feature matching → content structure comparison", effectively reducing the misjudgment rate of a single dimension and improving the overall accuracy of bid rigging detection.
[0038] In a second aspect of an embodiment of the present invention, a device for detecting bid rigging and collusion in tendering and procurement is disclosed, the device comprising:
[0039] a memory storing executable program code;
[0040] a processor coupled to the memory;
[0041] The processor calls the executable program code stored in the memory to execute the method for detecting bid rigging and collusion in bidding and procurement.
[0042] According to a third aspect of an embodiment of the present invention, a computer-storable medium is disclosed, wherein the computer-storable medium stores computer instructions. When the computer instructions are called by a computer, the computer instructions are used to execute the method for detecting bid rigging and collusion in bidding and procurement.
[0043] According to a fourth aspect of an embodiment of the present invention, an information data processing terminal is disclosed, which is used to implement the method for detecting bid rigging and collusion in bidding and procurement.
[0044] The beneficial effects of the present invention are:
[0045] The present invention detects bid rigging and collusion on a set of bidding documents from three dimensions: bidding address information, high-frequency word information, and chapter content data. First, by detecting bidding documents with the same IP address and Mac address through the bidding address information set, it is possible to quickly identify possible related bidders and preliminarily screen out suspicious bidding documents. Secondly, by using the high-frequency word information set to perform a second similarity detection, it is possible to find the similarities in the bidding documents in terms of word usage habits, professional terms, etc. Even if the bidders try to cover up the association relationship by modifying the document content, the similarity of high-frequency words is difficult to completely eliminate. Finally, a third similarity detection is performed based on the chapter content data set to further analyze the content structure and semantic similarity of the bidding documents, which can more comprehensively identify bid rigging and collusion. This multi-dimensional detection method greatly improves the accuracy of bid rigging and collusion detection compared to the traditional single-dimensional detection method.
[0046] The present invention uses an automated method to extract and process bid document collections and detect bid colluded bids. During the extraction phase, bid address information, high-frequency word information, and chapter content data can be quickly acquired, and a corresponding information set constructed. During the detection phase, by gradually updating the bid document collection, a first similarity test, a second similarity test, and a third similarity test are performed in sequence, effectively reducing the amount of data to be processed and improving detection efficiency. This efficient processing method can adapt to the detection needs of a large number of bid documents in bidding and procurement activities, saving a considerable amount of manpower and time costs, and making bid colluded bid detection more efficient and timely. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 4 is an implementation flow chart of the method of the present invention. DETAILED DESCRIPTION
[0048] In order to better understand the content of the present invention, an embodiment is given here.
[0049] Figure 1 4 is an implementation flow chart of the method of the present invention.
[0050] In a first aspect, an embodiment of the present invention discloses a method for detecting bid rigging in tender procurement, comprising:
[0051] S1, collecting a set of bidding documents for tender procurement; the set of bidding documents includes a plurality of bidding documents;
[0052] S2, extracting and processing the bidding document set to obtain a bidding address information set, a high-frequency word information set, and a chapter content data set;
[0053] S3, performing bid rigging detection processing on the bidding address information set, the high-frequency word information set, and the chapter content data set to obtain a detection result information set;
[0054] The detection result information set is a set of bidding documents that may indicate bid rigging or collusion;
[0055] The extracting process of the bidding document set to obtain a bidding address information set, a high-frequency word information set, and a chapter content data set includes:
[0056] S21, for each bidding document in the bidding document set, obtaining corresponding bidding address information; the bidding address information includes an IP address and a Mac address;
[0057] S22, constructing a bidding address information set using the bidding address information of all bidding documents; each bidding address information in the bidding address information set has a corresponding bidding document;
[0058] S23, extracting a corresponding high-frequency word information subset for each bid document in the bid document set; the high-frequency word information subset includes high-frequency word information; the high-frequency word information is a word vector whose number of occurrences exceeds a preset number threshold; the number of occurrences is the number of times the high-frequency word information appears in a bid document;
[0059] S24, constructing a high-frequency word information set using the high-frequency word information subsets of all bidding documents; each high-frequency word information in the high-frequency word information set has a corresponding bidding document;
[0060] S25, extracting corresponding chapter content data for each bidding document in the bidding document set;
[0061] S26 , constructing a chapter content data set using the chapter content data of all bidding documents; each chapter content data of the chapter content data set has a corresponding bidding document.
[0062] The bidding address information set, the high-frequency word information set, and the chapter content data set are subjected to bid rigging detection processing to obtain a detection result information set, including:
[0063] S31, performing a first similarity check on the bidding document set based on the bidding address information set to obtain a first target information set, and performing an update process on the bidding document set;
[0064] S32, performing a second similarity check on the updated bidding document set based on the high-frequency word information set, obtaining a second target information set and a feedback value set, and updating the bidding document set;
[0065] S33: Based on the chapter content data set and the feedback value set, a third similarity test is performed on the updated bidding document set to obtain a test result information set.
[0066] The step of performing a first similarity check on the bidding document set based on the bidding address information set to obtain a first target information set and performing an update process on the bidding document set includes:
[0067] S311, based on the bidding address information set, searching for bidding documents with the same bidding address information in the bidding document set; the same bidding address information means the same IP address or the same MAC address;
[0068] S312, constructing a first target information set using all bidding documents having the same bidding address information;
[0069] S313: Delete all bidding documents included in the first target information set from the bidding document set to obtain an updated bidding document set, thereby completing the updating process of the bidding document set.
[0070] The method of performing a second similarity check on the updated bidding document set based on the high-frequency word information set to obtain a second target information set and a feedback value set, and updating the bidding document set includes:
[0071] S321, performing word similarity calculation on each combination of two high-frequency word information subsets in the high-frequency word information set, to obtain a word similarity value of the combination and first feedback values of two bid documents corresponding to the combination;
[0072] S322, deleting the combinations whose word similarity values are greater than a preset first similarity threshold from the updated bid document set, thereby completing the update process of the bid document set and obtaining an updated bid document set;
[0073] S323: Perform statistical average calculation on all calculated first feedback values of each bidding document to obtain the feedback value of the bidding document.
[0074] The step of performing word similarity calculation on each combination of two high-frequency word information subsets in the high-frequency word information set to obtain a word similarity value of the combination and first feedback values of two bidding documents corresponding to the combination includes:
[0075] S3211, for each high-frequency word information subset in a combination of every two high-frequency word information subsets in the high-frequency word information set, sort all high-frequency word information contained in the subset according to the number of occurrences of the high-frequency word information from largest to smallest, to obtain a high-frequency word information sequence corresponding to the subset; the high-frequency word information sequence includes high-frequency word information;
[0076] S3212, performing high-frequency word similarity calculation processing on two high-frequency word information sequences corresponding to the combination of the two high-frequency word information subsets to obtain a word similarity value of the combination;
[0077] S3213: Calculate and obtain first feedback values of two bidding documents corresponding to the combination based on the word similarity values of the combination;
[0078] The expression for calculating the similarity of high-frequency words is:
[0079]
[0080] Among them, wg1 ij The jth element of the ith high-frequency word information of the first high-frequency word information sequence corresponding to the combination of the high-frequency word information subset, wg2 ij The jth element of the ith high-frequency word information of the second high-frequency word information sequence corresponding to the combination of the high-frequency word information subset, wg1 i0 is the mean value of the ith high-frequency word information of the first high-frequency word information sequence corresponding to the combination of high-frequency word information subsets, wg2 i0is the mean of the i-th high-frequency word information of the second high-frequency word information sequence corresponding to the combination of high-frequency word information subsets, gs is the word similarity value of the combination, N and M are the number of elements of the high-frequency word information and the number of high-frequency word information contained in the high-frequency word information sequence, respectively.
[0081] The calculation expression of the first feedback value is:
[0082]
[0083] Wherein, α1 and α2 are respectively the first feedback value of the first bidding document and the first feedback value of the second bidding document corresponding to the combination.
[0084] The expression for the statistical average calculation is:
[0085]
[0086] Among them, β k is the kth first feedback value calculated for a bidding document, K is the total number of first feedback values, μ and β0 are the variance and mean of all first feedback values of a bidding document, θ is the feedback value of a bidding document, T k () represents the kth order polynomial of the first kind Chebyshev polynomial.
[0087] The numerator of the calculation expression of the first feedback value calculates the degree of dispersion (variance) of the high-frequency word vector relative to its mean, reflecting the distribution stability of high-frequency words in different dimensions. By taking the Mth root (geometric mean), the degree of dispersion of multiple dimensions is integrated into a single indicator to avoid numerical deviations caused by different numbers of dimensions. The denominator gs is the word similarity value. When the high-frequency word sequences of the two bidding documents are highly similar (gs approaches 0), the values of α1 and α2 will increase significantly, indicating "abnormal dispersion under high similarity", which may imply that the bidder has plagiarized in disguise by adjusting the distribution of high-frequency words (such as splitting core terms into synonym combinations).
[0088] The expression for the statistical average calculation uses Chebyshev polynomials to perform a nonlinear transformation on the first feedback value, mapping the original feedback value into a higher-dimensional feature space and capturing periodic or nonlinear patterns in the feedback value sequence. For example, if multiple first feedback values exhibit regular fluctuations, the Chebyshev polynomials can convert them into significant features, while linear averaging smooths out such abnormal signals. As the order k increases, the Chebyshev polynomials become more sensitive to changes in the input value. In bid-rigging detection, this means a higher ability to identify "slight but systematic similarities" (such as multiple bidders using similar synonym replacement strategies), avoiding the missed detection of hidden bid-rigging behavior. The introduction of the variance term in the denominator allows the formula to assign higher weight to bid documents with larger fluctuations in feedback values. For example, if a bidder's first feedback value varies significantly (with a large variance) across different comparison groups, it indicates that its behavior pattern is unstable and there may be suspicion of bid-rigging. In this case, the formula will amplify the impact of these unstable feedback values. The term directly focuses on the point with the largest deviation in the feedback value. Even if other feedback values are normal, the presence of a single significant outlier will significantly increase the final feedback value θ. This design has a strong ability to detect "single-point breakthrough" tagging (for example, two documents deliberately using the same high-frequency word pattern).
[0089] In the second similarity detection process, the present invention introduces a feedback value mechanism. By performing a statistical average calculation on all the calculated first feedback values of each bid document, the feedback value of the bid document is obtained. The feedback value can be used as an important indicator for quantitatively evaluating the similarity between bid documents, providing a more intuitive basis for judgment for the tendering party and the regulatory authorities. For example, when the feedback value exceeds a certain threshold, it can be considered that the bid document has a high suspicion of bid rigging and collusion, and further investigation and verification are required. This quantitative evaluation method helps to improve the scientificity and objectivity of bid rigging and collusion detection, and avoids the limitations of relying solely on human subjective judgment.
[0090] The third similarity detection is performed on the updated bidding document set based on the chapter content data set and the feedback value set to obtain a detection result information set, including:
[0091] S331, using the chapter content data of each combination of two bid documents in the chapter content data set, construct a first content matrix and a second content matrix; a row vector of the first content matrix is a chapter content data of the first bid document of the combination of the two bid documents; a row vector of the second content matrix is a chapter content data of the second bid document of the combination of the two bid documents;
[0092] S332, subtracting the first content matrix from the second content matrix to obtain a difference matrix;
[0093] S333, performing statistical processing on each row vector of the difference matrix to obtain a statistical value set of each row vector; the statistical value set includes: mean, variance, range, and median;
[0094] S334, using the feedback value set, performing feedback calculation on the statistical value set of all row vectors to obtain a text similarity value of the two bid document combinations;
[0095] S335 , deleting the bidding document combination whose text similarity value is greater than a preset text threshold value from the updated bidding document set, and using the obtained bidding document set as the detection result information set.
[0096] The method of performing feedback calculation on the statistical value set of all row vectors using the feedback value set to obtain the text similarity value of the two bid document combinations includes:
[0097] Performing cross-correlation calculation on the difference matrix to obtain a cross-correlation matrix; the element in the i-th row and j-th column of the cross-correlation matrix is the cross-correlation value between the i-th row vector and the j-th row vector of the difference matrix;
[0098] Performing feature transformation calculation processing on the difference matrix and the cross-correlation matrix to obtain a feature matrix;
[0099] Extracting all diagonal elements of the characteristic matrix, and constructing a diagonal vector using all the diagonal elements;
[0100] Performing a first calculation process on the feedback value set, the diagonal vectors, and the statistical value set of all row vectors to obtain a text similarity value of the two bid document combinations;
[0101] The expression of the first calculation process is:
[0102]
[0103] Wherein, P represents the row dimension value of the difference matrix, L i () represents the i-th Laguerre polynomial, u i , δ i 、 k i Represent the mean, variance, range and median respectively, τ i is the i-th element of the diagonal vector, θ1 and θ2 represent the feedback value of the first bid document and the feedback value of the second bid document of the two bid document combinations respectively, and xs is the text similarity value of the two bid document combinations.
[0104] The expression for the feature transformation calculation process is:
[0105] A=C 1 / 2 YC -1 / 2 ,
[0106] Among them, C is the difference matrix, Y is the cross-correlation matrix, and A is the feature matrix.
[0107] The extraction obtains the corresponding high-frequency word information, which can be obtained by using a bag-of-words model or a text vectorization algorithm;
[0108] For each bidding document in the bidding document set, the corresponding chapter content data is extracted, and a bag-of-words model or a TF-IDF model can be used. Specifically, in the chapter content data identification, the chapter name of each chapter can be identified; the chapter name of each chapter is retained in a preset database.
[0109] The present invention updates the bid document collection at each detection stage, deleting any bid documents suspected of bid rigging. This update mechanism ensures the accuracy of subsequent detection stages and prevents identified bid rigging from interfering with subsequent detection results. Furthermore, the update mechanism makes the detection process clearer and more efficient, helping to improve the performance and reliability of the entire bid rigging detection system.
[0110] In a second aspect of an embodiment of the present invention, a device for detecting bid rigging and collusion in tendering and procurement is disclosed, the device comprising:
[0111] a memory storing executable program code;
[0112] a processor coupled to the memory;
[0113] The processor calls the executable program code stored in the memory to execute the method for detecting bid rigging and collusion in bidding and procurement.
[0114] According to a third aspect of an embodiment of the present invention, a computer-storable medium is disclosed, wherein the computer-storable medium stores computer instructions. When the computer instructions are called by a computer, the computer instructions are used to execute the method for detecting bid rigging and collusion in bidding and procurement.
[0115] According to a fourth aspect of an embodiment of the present invention, an information data processing terminal is disclosed, which is used to implement the method for detecting bid rigging and collusion in bidding and procurement.
[0116] The foregoing is merely an embodiment of the present invention and is not intended to limit the present invention. It will be apparent to those skilled in the art that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.
Claims
1. A method for detecting bid rigging in tender procurement, characterized in that: include: S1, collect the bidding documents for tendering and procurement; The bidding document set includes several bidding documents; S2, extracting and processing the bidding document set to obtain a bidding address information set, a high-frequency word information set, and a chapter content data set; S3, performing bid rigging detection processing on the bidding address information set, high-frequency word information set and chapter content data set to obtain a detection result information set.
2. The method for detecting bid rigging in tender procurement according to claim 1, wherein: The extracting process of the bidding document set to obtain a bidding address information set, a high-frequency word information set, and a chapter content data set includes: S21, for each bidding document in the bidding document set, obtaining corresponding bidding address information; the bidding address information includes an IP address and a Mac address; S22, constructing a bidding address information set using the bidding address information of all bidding documents; each bidding address information in the bidding address information set has a corresponding bidding document; S23, extracting a corresponding high-frequency word information subset for each bid document in the bid document set; the high-frequency word information subset includes high-frequency word information; the high-frequency word information is a word vector whose number of occurrences exceeds a preset number threshold; S24, constructing a high-frequency word information set using the high-frequency word information subsets of all bidding documents; each high-frequency word information in the high-frequency word information set has a corresponding bidding document; S25, extracting corresponding chapter content data for each bidding document in the bidding document set; S26 , constructing a chapter content data set using the chapter content data of all bidding documents; each chapter content data of the chapter content data set has a corresponding bidding document.
3. The method for detecting bid rigging in tender procurement according to claim 2, wherein: The bidding address information set, the high-frequency word information set, and the chapter content data set are subjected to bid rigging detection processing to obtain a detection result information set, including: S31, performing a first similarity check on the bidding document set based on the bidding address information set to obtain a first target information set, and performing an update process on the bidding document set; S32, performing a second similarity check on the updated bidding document set based on the high-frequency word information set, obtaining a second target information set and a feedback value set, and updating the bidding document set; S33: Based on the chapter content data set and the feedback value set, a third similarity test is performed on the updated bidding document set to obtain a test result information set.
4. The method for detecting bid rigging in tender procurement according to claim 3, wherein: The step of performing a first similarity check on the bidding document set based on the bidding address information set to obtain a first target information set and performing an update process on the bidding document set includes: S311, based on the bidding address information set, searching for bidding documents with the same bidding address information in the bidding document set; S312, constructing a first target information set using all bidding documents having the same bidding address information; S313: Delete all bidding documents included in the first target information set from the bidding document set to obtain an updated bidding document set, thereby completing the updating process of the bidding document set.
5. The method for detecting bid rigging in tender procurement according to claim 3, wherein: The method of performing a second similarity check on the updated bidding document set based on the high-frequency word information set to obtain a second target information set and a feedback value set, and updating the bidding document set includes: S321, performing word similarity calculation on each combination of two high-frequency word information subsets in the high-frequency word information set, to obtain a word similarity value of the combination and first feedback values of two bid documents corresponding to the combination; S322, deleting the combinations whose word similarity values are greater than a preset first similarity threshold from the updated bid document set, thereby completing the update process of the bid document set and obtaining an updated bid document set; S323: Perform statistical average calculation on all calculated first feedback values of each bidding document to obtain the feedback value of the bidding document.
6. The method for detecting bid rigging in tender procurement according to claim 5, wherein: The step of performing word similarity calculation on each combination of two high-frequency word information subsets in the high-frequency word information set to obtain a word similarity value of the combination and first feedback values of two bidding documents corresponding to the combination includes: S3211, for each high-frequency word information subset in a combination of every two high-frequency word information subsets in the high-frequency word information set, sort all high-frequency word information contained in the subset according to the number of occurrences of the high-frequency word information from largest to smallest, to obtain a high-frequency word information sequence corresponding to the subset; the high-frequency word information sequence includes high-frequency word information; S3212, performing high-frequency word similarity calculation processing on two high-frequency word information sequences corresponding to the combination of the two high-frequency word information subsets to obtain a word similarity value of the combination; S3213: Calculate and obtain first feedback values of the two bidding documents corresponding to the combination based on the word similarity value of the combination.
7. The method for detecting bid rigging in tender procurement according to claim 6, wherein: The expression for calculating the similarity of high-frequency words is: Among them, wg1 ij The jth element of the ith high-frequency word information of the first high-frequency word information sequence corresponding to the combination of the high-frequency word information subset, wg2 ij The jth element of the ith high-frequency word information of the second high-frequency word information sequence corresponding to the combination of the high-frequency word information subset, wg1 i0 is the mean value of the ith high-frequency word information of the first high-frequency word information sequence corresponding to the combination of high-frequency word information subsets, wg2 i0 is the mean of the i-th high-frequency word information of the second high-frequency word information sequence corresponding to the combination of high-frequency word information subsets, gs is the word similarity value of the combination, N and M are the number of elements of the high-frequency word information and the number of high-frequency word information contained in the high-frequency word information sequence, respectively.
8. A device for detecting bid rigging and collusion in bidding and procurement, characterized in that: The device comprises: a memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the method for detecting bid rigging and collusion in bidding and procurement according to any one of claims 1 to 7.
9. A computer storable medium, characterized in that The computer storable medium stores computer instructions, and when the computer instructions are called by a computer, they are used to execute the method for detecting bid rigging and collusion in bidding and procurement according to any one of claims 1 to 7.
10. An information data processing terminal, characterized in that: The information data processing terminal is used to implement the method for detecting bid rigging and collusion in bidding and procurement as described in any one of claims 1 to 7.