Multi-modal large model-based surrounding mark detection method and system
Through the multimodal large model-based concatenation detection method, combined with local and global feature extraction and text data analysis, the problems of inefficient and insufficient accuracy in concatenation detection are solved, efficient and accurate automated detection is achieved, and the fairness and impartiality of procurement activities are ensured.
Patent Information
- Application Number
- CN202510346055.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-22
AI Technical Summary
The prior art has problems in inefficient review, poor accuracy and lack of multi-dimensional review capabilities in the collusion testing, especially in large-scale procurement projects, manual comparisons cannot effectively identify improper collusion between suppliers.
The detection model is constructed through data preprocessing, image feature extraction, multimodal model verification and other modules. Combined with local and global feature extraction, the LSH algorithm is used to quickly cluster similar pictures, and combined with text data for multi-dimensional analysis, and set multiple loss functions for training and optimization.
It significantly improves the accuracy and efficiency of the joint bid testing, can quickly process a large number of bid documents, ensure the objectivity and fairness of the review results, and adapt to the needs of expanding procurement scale and changing methods of joint bidding.
Smart Images

Figure CN120356041A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bid rigging detection, and particularly to a bid rigging detection method and system based on a multimodal large model. Background Art
[0002] In commercial procurement activities, it is crucial to ensure a fair, just, and open competition environment. However, bid rigging behavior seriously disrupts this competition environment, damages the interests of the purchaser, and hinders the healthy development of the market. Bid rigging behavior refers to the act of collusion among suppliers to obtain the winning bid through improper means. Among them, when submitting bid documents, some suppliers try to get through by providing similar or even identical picture information, such as sample pictures, qualification certificate pictures, etc., to achieve the purpose of bid rigging. Therefore, there is a need to compare pictures for bid rigging detection.
[0003] Regarding the detection of bid rigging, traditionally it mainly relies on manual comparison, but there are the following drawbacks: 1. Low review efficiency: The efficiency of manual review is extremely low. As the scale of procurement projects continues to expand, the number of suppliers increases, and the volume of submitted picture materials grows exponentially. Manually comparing each picture one by one requires a large amount of time and labor costs. For example, in a large-scale engineering project procurement, it may involve hundreds of suppliers, and each supplier submits dozens of pictures. Manually reviewing these pictures may take several weeks, seriously affecting the progress of the procurement project. The reason for this is that the manual review method is essentially serial processing, unable to process multiple pictures simultaneously, and the human eye recognition speed and information processing ability are limited. When faced with a large number of pictures, the efficiency bottleneck is extremely obvious. 2. Poor review accuracy: The consistency of manual picture review is easily affected by subjective factors and physiological fatigue. Different review personnel have different judgment criteria for the similarity of pictures, which makes the review results lack a unified and objective scale. Moreover, long-term repetitive review work will lead to a decline in the attention and judgment of review personnel. For pictures that are slightly processed and similar, it is extremely easy to misjudge them as different pictures, thus missing clues of bid rigging. For example, some suppliers submit pictures of samples after slight color adjustment, cropping, etc. Manual review may be difficult to detect the similarity among them. This is mainly because the human visual system lacks a precise quantitative judgment standard when processing complex image information, and long-term work will cause visual fatigue, affecting the judgment accuracy. 3. Lack of multi-dimensional review ability: Only the pictures themselves are reviewed, without fully combining other text information in the bidding documents. And bid rigging behaviors are complex, and it is difficult to comprehensively and accurately judge only by picture comparison. For example, even if no obvious similarity is found in the pictures themselves, suppliers may reach a tacit agreement on bid rigging through specific codes or related information in the text description, and the existing review technologies cannot effectively discover these potential clues. The reason is that the traditional review method is limited to single-modal data processing and has not established an association analysis mechanism between multi-modal data such as images and texts, resulting in incomplete review information and inability to deeply explore bid rigging behaviors.
[0004] Although there are some automated image comparison technologies in the prior art, most of these technologies are applied in fields such as image retrieval and security monitoring, and are not suitable for direct application in bid document review. These technologies usually focus on general retrieval and matching of image content, and lack targeted optimization and adjustment for specific scenarios involved in the images in bid documents, such as sample image consistency detection, authenticity and consistency detection of qualification certificates, etc. Moreover, these technologies often do not consider the business logic of bid rigging detection. For example, after detecting similar images, there is no perfect processing process for how to further combine other information of the procurement project to determine whether it is a bid rigging behavior. That is, for the qualification certificate images in bid documents, not only the image similarity needs to be judged, but also it needs to be judged in combination with business rules such as qualification standards and project requirements. However, the prior art lacks this targeted processing ability because the requirements and focuses of image analysis in different fields are different, and general technologies do not fully consider the particularity and complexity of bid rigging detection business. If directly applied, it will affect the accuracy of bid rigging detection.
[0005] Therefore, how to provide a bid rigging detection method and system based on a multimodal large model to improve the accuracy and efficiency of bid rigging detection has become an urgent technical problem to be solved. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a bid rigging detection method and system based on a multimodal large model to improve the accuracy and efficiency of bid rigging detection.
[0007] In a first aspect, the present invention provides a bid rigging detection method based on a multimodal large model, including the following steps:
[0008] Step S1, create a bid rigging detection model based on a data preprocessing module, an image feature extraction module, an image analysis and comparison module, a multimodal large model verification module, and a detection report output module, and set the loss function of the bid rigging detection model;
[0009] Step S2, obtain a large number of historical bid documents, preprocess each historical bid document and construct a data set;
[0010] Step S3, divide the data set into a training set, a validation set, and a test set, train the bid rigging detection model through the training set until the loss value of the loss function is less than a preset loss threshold, and then verify and test the bid rigging detection model through the validation set and the test set respectively;
[0011] Step S4, deploy the bid rigging detection model that passes the test, and perform bid rigging detection through the deployed bid rigging detection model to obtain a bid rigging detection result;
[0012] Step S5: Continuously optimize the bid rigging detection model based on the bid rigging detection results.
[0013] Further, in step S1, the data preprocessing module, the image feature extraction module, the image analysis and comparison module, the multi-modal large model verification module, and the detection report output module are connected in sequence; the output end of the data preprocessing module is connected to the input end of the multi-modal large model verification module;
[0014] The data preprocessing module is used to extract image data and text data from the bidding documents, perform preprocessing on the image data including at least format conversion and size normalization, perform preprocessing on the text data including at least outlier removal, missing value filling, noise character removal, stop word removal, and HTML tag removal, input the preprocessed image data into the image feature extraction module, and input the preprocessed text data into the multi-modal large model verification module;
[0015] The image feature extraction module consists of a local feature extraction unit, a global feature extraction unit, and a feature fusion unit; the local feature extraction unit is used to extract local features from the image data and is constructed based on the SIFT network; the global feature extraction unit is used to extract global features from the image data and is constructed based on the ResNet-50 network; the feature fusion unit is used to fuse the local features and global features to obtain fused features and input the fused features into the image analysis and comparison module;
[0016] The image analysis and comparison module consists of a rough screening unit and a fine screening unit; the rough screening unit is used to map similar fused features into the same hash bucket through the LSH algorithm; the fine screening unit is used to calculate the cosine similarity of each fused feature in each hash bucket in sequence and output the fused features with the cosine similarity higher than the preset similarity threshold to the multi-modal large model verification module;
[0017] The multi-modal large model verification module is used to extract text features from the text data, perform fusion analysis on the text features and the fused features to obtain an analysis result, and output the analysis result to the detection report output module;
[0018] The detection report output module is used to output the bid rigging detection result based on the analysis result;
[0019] The loss function is constructed by weighting the L1 loss function, feature matching loss function, triplet loss function, contrastive loss function, and cross-entropy loss function; the L1 loss function is used to measure the difference between the bid documents before and after being processed by the data preprocessing module; the feature matching loss function is used to calculate the difference between the fused features extracted by the image feature extraction module and the target features; the triplet loss function is used to train the image analysis and comparison module to learn the similarity of images; the contrastive loss function is used to measure the similarity between different modality data input by the multi-modal large model verification module; the cross-entropy loss function is used to measure the accuracy of the bid-rigging detection results output by the detection report output module.
[0020] Further, step S2 is specifically as follows:
[0021] Obtain a large number of historical bid documents. After preprocessing the image data and text data carried by each historical bid document through data cleaning, label the bid-rigging behaviors of the preprocessed image data and text data, and construct a dataset based on the labeled image data and text data.
[0022] Further, step S3 is specifically as follows:
[0023] Divide the dataset into a training set, a validation set, and a test set according to a ratio of 7:2:1. Train the bid-rigging detection model through the training set until the loss value of the loss function is less than a preset loss threshold. During the training process, continuously optimize the hyperparameters of the bid-rigging detection model, including at least the learning rate, dropout rate, batch size, and number of iterations;
[0024] Verify the trained bid-rigging detection model through the validation set, and determine whether the detection accuracy is greater than a preset accuracy threshold. If not, the verification fails, and the dataset is expanded and training continues; if so, the verification passes, and:
[0025] Test the bid-rigging detection model that has passed the verification through the test set, and determine whether the confidence level is greater than a preset confidence level threshold. If not, the test fails, and the dataset is expanded and training continues; if so, the test passes, and the training ends.
[0026] Further, step S4 is specifically as follows:
[0027] Deploy the bid-rigging detection model that has passed the test through a distributed computing architecture, obtain real-time bid documents, and input the real-time bid documents into the deployed bid-rigging detection model to perform bid-rigging detection through the deployed bid-rigging detection model to obtain bid-rigging detection results.
[0028] Second aspect, the present invention provides a bid rigging detection system based on a multimodal large model, including the following modules:
[0029] A bid rigging detection model creation module, configured to create a bid rigging detection model based on a data preprocessing module, a picture feature extraction module, a picture analysis and comparison module, a multimodal large model verification module, and a detection report output module, and set a loss function of the bid rigging detection model;
[0030] A dataset construction module, configured to obtain a large number of historical tender documents, preprocess each of the historical tender documents, and construct a dataset;
[0031] A bid rigging detection model training module, configured to divide the dataset into a training set, a validation set, and a test set, train the bid rigging detection model through the training set until the loss value of the loss function is less than a preset loss threshold, and then verify and test the bid rigging detection model through the validation set and the test set respectively;
[0032] A bid rigging detection module, configured to deploy the bid rigging detection model that has passed the test, and perform bid rigging detection through the deployed bid rigging detection model to obtain a bid rigging detection result;
[0033] A bid rigging detection model optimization module, configured to continuously optimize the bid rigging detection model through the bid rigging detection result.
[0034] Further, in the bid rigging detection model creation module, the data preprocessing module, the picture feature extraction module, the picture analysis and comparison module, the multimodal large model verification module, and the detection report output module are connected in sequence; the output end of the data preprocessing module is connected to the input end of the multimodal large model verification module;
[0035] The data preprocessing module is configured to extract picture data and text data from the tender documents, perform preprocessing on the picture data including at least format conversion and size normalization, perform preprocessing on the text data including at least outlier removal, missing value filling, noise character removal, stop word removal, and HTML tag removal, input the preprocessed picture data into the picture feature extraction module, and input the preprocessed text data into the multimodal large model verification module;
[0036] The picture feature extraction module consists of a local feature extraction unit, a global feature extraction unit, and a feature fusion unit; the local feature extraction unit is used to extract local features from picture data and is constructed based on the SIFT network; the global feature extraction unit is used to extract global features from picture data and is constructed based on the ResNet-50 network; the feature fusion unit is used to fuse the local features and global features to obtain fused features, and input the fused features into the picture analysis and comparison module;
[0037] The picture analysis and comparison module consists of a rough screening unit and a fine screening unit; the rough screening unit is used to map similar fused features into the same hash bucket through the LSH algorithm; the fine screening unit is used to calculate the cosine similarity of each fused feature in each hash bucket in turn, and output the fused features with the cosine similarity higher than the preset similarity threshold to the multi-modal large model verification module;
[0038] The multi-modal large model verification module is used to extract text features from text data, perform fusion analysis on the text features and fused features to obtain an analysis result, and output the analysis result to the detection report output module;
[0039] The detection report output module is used to output the bid rigging detection result according to the analysis result;
[0040] The loss function is weighted and constructed based on the L1 loss function, feature matching loss function, triplet loss function, contrast loss function, and cross-entropy loss function; the L1 loss function is used to measure the difference before and after the data preprocessing module processes the tender documents; the feature matching loss function is used to calculate the difference between the fused features extracted by the picture feature extraction module and the target features; the triplet loss function is used to train the picture analysis and comparison module to learn the similarity of pictures; the contrast loss function is used to measure the similarity between different modal data input by the multi-modal large model verification module; the cross-entropy loss function is used to measure the accuracy of the bid rigging detection result output by the detection report output module.
[0041] Further, the dataset construction module is specifically used for:
[0042] Obtain a large number of historical tender documents, after preprocessing the picture data and text data carried by each historical tender document through data cleaning, label the bid rigging behavior of the preprocessed picture data and text data, and construct a dataset based on the labeled picture data and text data.
[0043] Further, the bid rigging detection model training module is specifically used for:
[0044] Divide the dataset into a training set, a validation set, and a test set based on a ratio of 7:2:1. Train the bid rigging detection model using the training set until the loss value of the loss function is less than a preset loss threshold. During the training process, continuously optimize the hyperparameters of the bid rigging detection model, including at least the learning rate, dropout rate, batch size, and number of iterations.
[0045] Validate the trained bid rigging detection model using the validation set. Determine whether the detection accuracy is greater than a preset accuracy threshold. If not, the validation fails, and the dataset is expanded for continued training. If so, the validation passes, and:
[0046] Test the bid rigging detection model that has passed the validation using the test set. Determine whether the confidence level is greater than a preset confidence threshold. If not, the test fails, and the dataset is expanded for continued training. If so, the test passes, and the training ends.
[0047] Furthermore, the bid rigging detection module is specifically used for:
[0048] Deploy the bid rigging detection model that has passed the test through a distributed computing architecture, obtain real-time tender documents, and input the real-time tender documents into the deployed bid rigging detection model to perform bid rigging detection through the deployed bid rigging detection model and obtain the bid rigging detection result.
[0049] The advantages of the present invention are as follows:
[0050] 1. Create a bid rigging detection model through a data preprocessing module, a picture feature extraction module, a picture analysis and comparison module, a multi-modal large model verification module, and a detection report output module, and set the loss function of the bid rigging detection model; then obtain a large number of historical tender documents, preprocess each historical tender document to construct a dataset; divide the dataset into a training set, a validation set, and a test set, train the bid rigging detection model using the training set until the loss value of the loss function is less than a preset loss threshold, and then validate and test the bid rigging detection model using the validation set and the test set respectively; then deploy the bid rigging detection model that has passed the test, perform bid rigging detection through the deployed bid rigging detection model to obtain the bid rigging detection result, and continuously optimize the bid rigging detection model based on the bid rigging detection result; that is, automatically detect bid rigging through the bid rigging detection model, overcoming the problems of low review efficiency, poor review accuracy, and lack of multi-dimensional review ability caused by relying on manual comparison traditionally. Moreover, the bid rigging detection model is trained using a dataset constructed from historical tender documents, fully considering the particularity and complexity of the bid rigging detection business, and combined with the continuous optimization of the bid rigging detection model, ultimately greatly improving the accuracy and efficiency of bid rigging detection.
[0051] 2. By setting up a data preprocessing module to extract image data and text data from tender documents, and perform preprocessing on the image data including at least format conversion and size normalization, and perform preprocessing on the text data including at least outlier removal, missing value filling, noise character removal, stop word removal, and HTML tag removal, the data quality of the image data and text data is effectively improved, facilitating subsequent detection and analysis.
[0052] 3. By setting up an image feature extraction module consisting of a local feature extraction unit, a global feature extraction unit, and a feature fusion unit; the local feature extraction unit is used to extract local features from the image data and is constructed based on the SIFT network; the global feature extraction unit is used to extract global features from the image data and is constructed based on the ResNet-50 network; the feature fusion unit is used to fuse the local features and global features; that is, the feature extraction stage combines local features and global features, effectively improving the feature extraction performance, and thus greatly improving the accuracy of bid rigging detection.
[0053] 4. By setting up an image analysis and comparison module consisting of a rough screening unit and a fine screening unit; the rough screening unit is used to map similar fusion features into the same hash bucket through the LSH algorithm; the fine screening unit is used to calculate the cosine similarity of each fusion feature in each hash bucket in turn, and output the fusion features with cosine similarity higher than the preset similarity threshold to the multi-modal large model verification module; that is, the comparison of image similarity is carried out in two stages, using the LSH hashing algorithm to quickly cluster similar images, greatly reducing the computational complexity, adapting to the scenario of massive data, and accurately matching through cosine similarity to ensure that high-similarity features enter multi-modal verification, taking into account both speed and accuracy.
[0054] 5. By setting up a multi-modal large model verification module to extract text features from the text data, and perform fusion analysis on the text features and fusion features to obtain an analysis result, that is, combining the two modal data of image data and text data for bid rigging detection, fully considering their relevance, and thus greatly improving the accuracy of bid rigging detection.
[0055] 6. The loss function is constructed by weighting based on the L1 loss function, feature matching loss function, triplet loss function, contrastive loss function, and cross-entropy loss function. The L1 loss function is used to measure the difference between the bid documents before and after being processed by the data preprocessing module. The feature matching loss function is used to calculate the difference between the fused features extracted by the image feature extraction module and the target features. The triplet loss function is used to train the image analysis and comparison module to learn the similarity of images. The contrastive loss function is used to measure the similarity between different modality data input to the multi-modal large model verification module. The cross-entropy loss function is used to measure the accuracy of the collusive bidding detection results output by the detection report output module. That is, different loss functions are adopted for different modules to match the calculation tasks of different modules, effectively improving the training effect of the collusive bidding detection model, strengthening the relevance between feature extraction and target tasks, enhancing the model's ability to distinguish different bid documents, and thus greatly improving the accuracy of collusive bidding detection.
[0056] 7. The dataset is divided into a training set, a validation set, and a test set in a ratio of 7:2:1. The collusive bidding detection model is trained using the training set until the loss value of the loss function is less than a preset loss threshold. During the training process, the hyperparameters of the collusive bidding detection model are continuously optimized, including at least the learning rate, random dropout rate, batch size, and number of iterations. The trained collusive bidding detection model is validated using the validation set, and the validated collusive bidding detection model is tested using the test set. That is, during the training process of the collusive bidding detection model, continuous optimization, validation, and testing are carried out, thus greatly improving the accuracy of collusive bidding detection.
[0057] 8. The collusive bidding detection model that has passed the test is deployed through a distributed computing architecture, that is, different modules in the collusive bidding detection model are deployed on different computing nodes for parallel processing, effectively improving the collusive bidding detection efficiency and meeting the needs of large-scale bid document review.
[0058] 9. Automated extraction and comparison of image features (fused features) are performed through the collusive bidding detection model, which can quickly process a large number of images. For example, after uniformly converting the format and normalizing the size of the images in the data preprocessing stage, the ResNet-50 network is used to extract global features in parallel, and the SIFT network is used to extract local features synchronously, greatly shortening the feature extraction time for a single image. In the comparison stage, the LSH algorithm is used for rough screening and quick grouping, and the cosine similarity calculation within the hash bucket can quickly locate suspected similar images. Compared with manual review, the review efficiency can be increased by several times or even dozens of times, greatly accelerating the procurement process, enabling the purchaser to complete the review of bid documents in a shorter time, and promoting the progress of the project.
[0059] 10. By comprehensively applying a variety of advanced algorithms, combining global features and local features to extract comprehensive feature information, calculating cosine similarity to provide an objective similarity measure, and using a multi-modal large model to verify bid rigging situations from multiple dimensions of pictures and texts; taking the review of qualification certificate pictures as an example, it can not only find out whether the certificates submitted by suppliers are the same through picture feature comparison, but also combine text data to judge whether the supplier qualifications match the project requirements, avoiding misjudgment and missed judgment, greatly improving the accuracy of bid rigging behavior detection, and effectively maintaining the fairness of bidding.
[0060] 11. Through the detection by the bid rigging detection model, all bid documents are reviewed according to the same standard. Whenever and wherever the same bid document is input, or the picture content is the same but only the shooting background or angle is different, the obtained detection results are the same, ensuring the objectivity of the review process and results, and providing a reliable and fair review conclusion for the purchaser.
[0061] 12. By adopting a distributed computing architecture, it is convenient to add computing nodes to process more bid documents to meet the needs of the expanding procurement scale; at the same time, it has good scalability. As the bid rigging means change, it can easily add new picture type processing modules or optimize existing algorithms; for example, if new picture tampering means appear, the feature extraction algorithm can be optimized for this situation, or relevant training data can be added to the multi-modal large model to continuously maintain its effective detection ability for various bid rigging behaviors, and long-term ensure the healthy and orderly progress of procurement activities.
[0062] 13. Extract multi-dimensional features of images through SIFT (local details) and ResNet-50 (global structure), combine with text semantic features to construct a more comprehensive representation of bid documents; at the same time, integrate pictures (local + global features) and text data, break through the limitations of a single modality, and can identify hidden bid rigging features such as contradictions in text content and inconsistency between pictures and texts, thereby greatly improving the accuracy of bid rigging detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] The present invention will be further described below with reference to the accompanying drawings in conjunction with embodiments.
[0064] Figure 1 is a flowchart of a method for detecting bid rigging based on a multi-modal large model of the present invention.
[0065] Figure 2 is a schematic structural diagram of a system for detecting bid rigging based on a multi-modal large model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0066] The overall idea of the technical solution in the embodiments of this application is as follows: An automatic detection of bid rigging is carried out through a bid rigging detection model, overcoming the problems of low review efficiency, poor review accuracy, and lack of multi-dimensional review capabilities caused by relying on manual comparison traditionally. Moreover, the bid rigging detection model is trained through a dataset constructed from historical bidding documents, fully considering the particularity and complexity of the bid rigging detection business, and combined with the continuous optimization of the bid rigging detection model to improve the accuracy and efficiency of bid rigging detection.
[0067] Please refer to Figures 1 to 2 As shown, a preferred embodiment of a bid rigging detection method based on a multi-modal large model of the present invention includes the following steps:
[0068] Step S1: Create a bid rigging detection model based on a data preprocessing module, a picture feature extraction module, a picture analysis and comparison module, a multi-modal large model verification module, and a detection report output module, and set the loss function of the bid rigging detection model;
[0069] Step S2: Obtain a large number of historical bidding documents, and construct a dataset after preprocessing each of the historical bidding documents;
[0070] Step S3: Divide the dataset into a training set, a validation set, and a test set, train the bid rigging detection model through the training set until the loss value of the loss function is less than a preset loss threshold, and then verify and test the bid rigging detection model through the validation set and the test set respectively;
[0071] Step S4: Deploy the bid rigging detection model that passes the test, and perform bid rigging detection through the deployed bid rigging detection model to obtain a bid rigging detection result;
[0072] Step S5: Continuously optimize the bid rigging detection model through the bid rigging detection result.
[0073] By comprehensively applying a variety of advanced algorithms, combining global features and local features to extract comprehensive feature information, calculating cosine similarity to provide an objective similarity measure, and using a multi-modal large model to verify bid rigging situations from multiple dimensions of pictures and texts; taking the review of qualification certificate pictures as an example, not only can it be found whether the certificates submitted by suppliers are the same through picture feature comparison, but also it can be judged whether the supplier qualifications match the project requirements in combination with text data, avoiding misjudgment and missed judgment, greatly improving the accuracy of bid rigging behavior detection, and effectively maintaining the fairness of bidding.
[0074] Detection is carried out through the bid rigging detection model, enabling all bid documents to be reviewed according to the same criteria. Whenever and wherever the same bid document is input, or the picture content is the same but only the shooting background or angle is different, the obtained detection results are consistent, ensuring the objectivity of the review process and results, and providing a reliable and fair review conclusion for the purchaser.
[0075] In the step S1, the data preprocessing module, the picture feature extraction module, the picture analysis and comparison module, the multi-modal large model verification module, and the detection report output module are connected in sequence; the output end of the data preprocessing module is connected to the input end of the multi-modal large model verification module;
[0076] The data preprocessing module is used to extract picture data and text data from the bid document, perform preprocessing on the picture data including at least format conversion and size normalization, and perform preprocessing on the text data including at least outlier removal, missing value filling, noise character removal, stop word removal, and HTML tag removal. The preprocessed picture data is input into the picture feature extraction module, and the preprocessed text data is input into the multi-modal large model verification module;
[0077] By setting the data preprocessing module to extract picture data and text data from the bid document, perform preprocessing on the picture data including at least format conversion and size normalization, and perform preprocessing on the text data including at least outlier removal, missing value filling, noise character removal, stop word removal, and HTML tag removal, the data quality of the picture data and text data is effectively improved, facilitating subsequent detection and analysis.
[0078] The picture feature extraction module consists of a local feature extraction unit, a global feature extraction unit, and a feature fusion unit; the local feature extraction unit is used to extract local features from the picture data and is constructed based on the SIFT network; the global feature extraction unit is used to extract global features from the picture data and is constructed based on the ResNet-50 network; the feature fusion unit is used to fuse the local features and global features to obtain a fused feature, and input the fused feature into the picture analysis and comparison module;
[0079] The SIFT network first detects the extreme points in the scale space through the Difference of Gaussian (DoG) pyramid. For the image I(x, y), in the scale space L(x, y, σ), the DoG operator is defined as:
[0080] D(x, y, σ) = (G(x, y, kσ) - G(x, y, σ)) * I(x, y) = L(x, y, kσ) - L(x, y, σ);
[0081] Among them, G(x, y, σ) represents the Gaussian kernel function; k represents the scale factor; the key point positions are determined by finding the extreme points of D(x, y, σ).
[0082] Then, the main direction is calculated for each key point, which is determined based on the gradient direction distribution within the neighborhood of the key point to ensure the rotational invariance of the feature.
[0083] Finally, 128-dimensional local features (SIFT feature descriptors) are generated according to the gradient information within the neighborhood of the key points. These descriptors can accurately describe the unique features of the local regions in the picture, such as the texture details and edge features of the object.
[0084] The picture feature extraction module is composed of a local feature extraction unit, a global feature extraction unit, and a feature fusion unit; the local feature extraction unit is used to extract local features from the picture data and is constructed based on the SIFT network; the global feature extraction unit is used to extract global features from the picture data and is constructed based on the ResNet-50 network; the feature fusion unit is used to fuse the local features and global features; that is, the feature extraction stage combines local features and global features, effectively improving the feature extraction performance, and thus greatly improving the accuracy of bid rigging detection.
[0085] The ResNet-50 network processes the picture data through a series of convolutional layers, pooling layers, and fully connected layers, and its output is 1280-dimensional global features (feature vectors); in the convolution operation, assuming the input picture data is I(x, y) and the convolution kernel is K(m, n), the global feature F(i, j) is obtained after convolution operation, and the calculation formula is:
[0086]
[0087] Among them, M represents the row of the convolution kernel; N represents the column of the convolution kernel; (i, j) represents the position in the feature map (global feature). This convolution operation can automatically learn various feature patterns in the picture, and after multiple layers of convolution and pooling, an effective representation of the global features of the picture is finally obtained.
[0088] The picture analysis and comparison module is composed of a rough screening unit and a fine screening unit; the rough screening unit is used to map similar fusion features into the same hash bucket through the LSH algorithm; the fine screening unit is used to calculate the cosine similarity of each fusion feature in each hash bucket in turn, and output the fusion features with the cosine similarity higher than the preset similarity threshold to the multi-modal large model verification module;
[0089] The image analysis and comparison module is composed of a rough screening unit and a fine screening unit; the rough screening unit is used to map similar fusion features into the same hash bucket through the LSH algorithm; the fine screening unit is used to calculate the cosine similarity of each fusion feature in each hash bucket in turn, and output the fusion features with cosine similarity higher than the preset similarity threshold to the multi-modal large model verification module; that is, the comparison of image similarity is carried out in two stages. The LSH hashing algorithm is used to quickly cluster similar images, greatly reducing the computational complexity and adapting to the massive data scenario. Through the accurate matching of cosine similarity, it is ensured that high-similarity features enter the multi-modal verification, taking into account both speed and accuracy.
[0090] The LSH algorithm (Locality-Sensitive Hashing algorithm) maps the fusion features (feature vectors) of similar image data into the same hash bucket; for a given fusion feature v, the LSH algorithm maps it into the hash bucket through the hash function h(v).
[0091]
[0092] Among them, <a, v> represents the inner product of the random vector a and the fusion feature v; b represents the random offset; w represents the width of the hash bucket. Similar vectors have a high probability of being mapped into the same hash bucket, so that a large number of images in the bidding documents of each supplier can be quickly grouped according to similarity, greatly reducing the number of images that need to be carefully compared later.
[0093] The calculation formula of cosine similarity is:
[0094]
[0095] Among them, <v1, v2> represents the inner product of the vector v1 and the vector v2; ||v1|| represents the L2 norm of the vector v1; ||v2|| represents the L2 norm of the vector v2.
[0096] The multi-modal large model verification module is used to extract text features from text data, perform fusion analysis on the text features and fusion features, obtain an analysis result, and output the analysis result to the detection report output module;
[0097] By setting the multi-modal large model verification module to extract text features from text data and perform fusion analysis on the text features and fusion features to obtain an analysis result, that is, combining the two modal data of image data and text data for bid rigging detection, fully considering their relevance, and thus greatly improving the accuracy of bid rigging detection.
[0098] The multi-modal large model combines the fusion features and text features and conducts comprehensive analysis through the joint embedding space. It can be a model based on the Transformer architecture. Assuming the fusion feature is I and the text feature is T, the model determines whether there is bid rigging behavior through the learning function f(I, T). For example, this function can be learned through a multi-layer neural network. The fusion feature and text feature are used as inputs, processed through a series of fully connected layers and activation functions, and finally a probability value is output to represent the possibility of bid rigging behavior.
[0099] The detection report output module is used to output the bid rigging detection result according to the analysis result.
[0100] By using the bid rigging detection model to automatically extract and compare the image features (fusion features), a large number of images can be processed quickly. For example, after uniformly converting the format and normalizing the size of the images in the data preprocessing stage, the ResNet-50 network is used to extract global features in parallel, and the SIFT network is used to extract local features synchronously, which greatly shortens the feature extraction time of a single image. In the comparison stage, the LSH algorithm is used for rough screening and fast grouping, and the cosine similarity calculation in the hash bucket can quickly locate suspected similar images. Compared with manual review, the review efficiency can be increased by several times or even dozens of times, greatly accelerating the procurement process, enabling the purchaser to complete the review of the bidding documents in a shorter time and promoting the progress of the project.
[0101] Extract multi-dimensional features of the image through SIFT (local details) and ResNet-50 (global structure), combine with the text semantic features, and construct a more comprehensive representation of the bidding document. At the same time, integrate the image (local + global features) and text data, break through the limitations of a single modality, and can identify hidden bid rigging features such as contradictions in text content and inconsistencies between pictures and texts, thereby greatly improving the accuracy of bid rigging detection.
[0102] The loss function is weighted and constructed based on the L1 loss function, feature matching loss function, triplet loss function, contrast loss function, and cross-entropy loss function. The L1 loss function is used to measure the difference before and after the data preprocessing module processes the bidding document. The feature matching loss function is used to calculate the difference between the fusion feature extracted by the image feature extraction module and the target feature. The triplet loss function is used to train the image analysis and comparison module to learn the similarity of images. The contrast loss function is used to measure the similarity between different modality data input by the multi-modal large model verification module. The cross-entropy loss function is used to measure the accuracy of the bid rigging detection result output by the detection report output module.
[0103] The loss function is constructed by weighting based on the L1 loss function, feature matching loss function, triplet loss function, contrastive loss function, and cross-entropy loss function; the L1 loss function is used to measure the difference between the bid documents before and after being processed by the data preprocessing module; the feature matching loss function is used to calculate the difference between the fused features extracted by the image feature extraction module and the target features; the triplet loss function is used to train the image analysis and comparison module to learn the similarity of images; the contrastive loss function is used to measure the similarity between different modality data input by the multi-modal large model verification module; the cross-entropy loss function is used to measure the accuracy of the bid-rigging detection results output by the detection report output module; that is, different loss functions are adopted for different modules to match the calculation tasks of different modules, effectively improving the training effect of the bid-rigging detection model, strengthening the correlation between feature extraction and the target task, enhancing the model's ability to distinguish different bid documents, and thus greatly improving the accuracy of bid-rigging detection.
[0104] The specific steps of step S2 are as follows:
[0105] A large number of historical bid documents are obtained. After preprocessing the image data and text data carried by each historical bid document through data cleaning, the preprocessed image data and text data are labeled for bid-rigging behaviors, and a data set is constructed based on the labeled image data and text data.
[0106] The specific steps of step S3 are as follows:
[0107] The data set is divided into a training set, a validation set, and a test set according to a ratio of 7:2:1. The bid-rigging detection model is trained through the training set until the loss value of the loss function is less than a preset loss threshold. During the training process, the hyperparameters of the bid-rigging detection model, including at least the learning rate, random dropout rate, batch size, and number of iterations, are continuously optimized;
[0108] The trained bid-rigging detection model is verified through the validation set to determine whether the detection accuracy is greater than a preset accuracy threshold. If not, the verification fails, and the data set is expanded and training continues; if so, the verification passes, and:
[0109] The bid-rigging detection model that has passed the verification is tested through the test set to determine whether the confidence level is greater than a preset confidence threshold. If not, the test fails, and the data set is expanded and training continues; if so, the test passes, and the training ends.
[0110] Divide the dataset into a training set, a validation set, and a test set in a ratio of 7:2:1. Train the bid-rigging detection model using the training set until the loss value of the loss function is less than a preset loss threshold. During the training process, continuously optimize the hyperparameters of the bid-rigging detection model, including at least the learning rate, dropout rate, batch size, and number of iterations. Validate the trained bid-rigging detection model using the validation set, and test the bid-rigging detection model that has passed the validation using the test set. That is, during the training process of the bid-rigging detection model, continuous optimization, validation, and testing are carried out, thereby greatly improving the accuracy of bid-rigging detection.
[0111] The specific steps of step S4 are as follows:
[0112] Deploy the tested bid-rigging detection model through a distributed computing architecture, obtain real-time tender documents, and input the real-time tender documents into the deployed bid-rigging detection model to perform bid-rigging detection through the deployed bid-rigging detection model, and obtain the bid-rigging detection result.
[0113] Deploy the tested bid-rigging detection model through a distributed computing architecture, that is, deploy different modules in the bid-rigging detection model on different computing nodes for parallel processing, effectively improving the bid-rigging detection efficiency and meeting the needs of large-scale tender document review.
[0114] By adopting a distributed computing architecture, it is convenient to add computing nodes to process more tender documents, meeting the needs of expanding the procurement scale; at the same time, it has good scalability. As the bid-rigging means change, it can easily add new picture type processing modules or optimize existing algorithms. For example, if new picture tampering means appear, the feature extraction algorithm can be optimized for this situation, or relevant training data can be added to the multi-modal large model to continuously maintain its effective detection ability for various bid-rigging behaviors and long-term ensure the healthy and orderly progress of procurement activities.
[0115] A preferred embodiment of a bid-rigging detection system based on a multi-modal large model according to the present invention includes the following modules:
[0116] A bid-rigging detection model creation module for creating a bid-rigging detection model based on a data preprocessing module, a picture feature extraction module, a picture analysis and comparison module, a multi-modal large model verification module, and a detection report output module, and setting the loss function of the bid-rigging detection model;
[0117] A dataset construction module for obtaining a large number of historical tender documents and constructing a dataset after preprocessing each of the historical tender documents;
[0118] The bid-rigging detection model training module is used to divide the data set into a training set, a validation set, and a test set, train the bid-rigging detection model with the training set until the loss value of the loss function is less than a preset loss threshold, and then verify and test the bid-rigging detection model with the validation set and the test set respectively;
[0119] The bid-rigging detection module is used to deploy the bid-rigging detection model that has passed the test, and perform bid-rigging detection through the deployed bid-rigging detection model to obtain the bid-rigging detection result;
[0120] The bid-rigging detection model optimization module is used to continuously optimize the bid-rigging detection model based on the bid-rigging detection result.
[0121] By comprehensively applying a variety of advanced algorithms, combining global features and local features to extract comprehensive feature information, calculating cosine similarity to provide an objective similarity measure, and using a multi-modal large model to verify bid-rigging situations from multiple dimensions of pictures and texts; taking the review of qualification certificate pictures as an example, it can not only find out whether the certificates submitted by suppliers are the same through picture feature comparison, but also combine text data to judge whether the supplier qualifications match the project requirements, avoiding misjudgment and missed judgment, greatly improving the accuracy of bid-rigging behavior detection, and effectively maintaining the fairness of bidding.
[0122] Detecting through the bid-rigging detection model enables all bid documents to be reviewed according to the same standard. Whenever and wherever the same bid document is input, or the picture content is the same but only the shooting background or angle is different, the obtained detection results are the same, ensuring the objectivity of the review process and results, and providing a reliable and fair review conclusion for the purchaser.
[0123] In the bid-rigging detection model creation module, the data preprocessing module, the picture feature extraction module, the picture analysis and comparison module, the multi-modal large model verification module, and the detection report output module are connected in sequence; the output end of the data preprocessing module is connected to the input end of the multi-modal large model verification module;
[0124] The data preprocessing module is used to extract picture data and text data from the bid documents, perform preprocessing on the picture data including at least format conversion and size normalization, perform preprocessing on the text data including at least outlier removal, missing value filling, noise character removal, stop word removal, and HTML tag removal, input the preprocessed picture data into the picture feature extraction module, and input the preprocessed text data into the multi-modal large model verification module;
[0125] By setting up a data preprocessing module to extract image data and text data from tender documents, and performing preprocessing on the image data including at least format conversion and size normalization, and performing preprocessing on the text data including at least outlier removal, missing value filling, noise character removal, stop word removal, and HTML tag removal, the data quality of the image data and text data is effectively improved, facilitating subsequent detection and analysis.
[0126] The image feature extraction module consists of a local feature extraction unit, a global feature extraction unit, and a feature fusion unit; the local feature extraction unit is used to extract local features from the image data and is constructed based on the SIFT network; the global feature extraction unit is used to extract global features from the image data and is constructed based on the ResNet-50 network; the feature fusion unit is used to fuse the local features and global features to obtain fused features, and input the fused features into the image analysis and comparison module;
[0127] The SIFT network first detects extreme points in the scale space through the Difference of Gaussian pyramid (DoG). For the image I(x, y), in the scale space L(x, y, σ), the DoG operator is defined as:
[0128] D(x, y, σ) = (G(x, y, kσ) - G(x, y, σ)) * I(x, y) = L(x, y, kσ) - L(x, y, σ);
[0129] where G(x, y, σ) represents the Gaussian kernel function; k represents the scale factor; the key point positions are determined by finding the extreme points of D(x, y, σ).
[0130] Then, the main direction is calculated for each key point, which is determined based on the gradient direction distribution within the key point neighborhood to ensure the rotation invariance of the feature.
[0131] Finally, 128-dimensional local features (SIFT feature descriptors) are generated based on the gradient information within the key point neighborhood. These descriptors can accurately describe the unique features of the local regions in the image, such as the texture details and edge features of objects.
[0132] By setting that the image feature extraction module consists of a local feature extraction unit, a global feature extraction unit, and a feature fusion unit; the local feature extraction unit is used to extract local features from the image data and is constructed based on the SIFT network; the global feature extraction unit is used to extract global features from the image data and is constructed based on the ResNet-50 network; the feature fusion unit is used to fuse the local features and global features; that is, the feature extraction stage combines local features and global features, effectively improving the feature extraction performance, and thus greatly improving the accuracy of bid rigging detection.
[0133] The ResNet-50 network processes image data through a series of convolutional layers, pooling layers, and fully connected layers, and its output is a 1280-dimensional global feature (feature vector); in the convolution operation, assuming the input image data is I(x,y) and the convolution kernel is K(m,n), the global feature F(i,j) is obtained after convolution operation, and the calculation formula is:
[0134]
[0135] where M represents the row of the convolution kernel; N represents the column of the convolution kernel; (i,j) represents the position in the feature map (global feature). This convolution operation can automatically learn various feature patterns in the image. After multiple layers of convolution and pooling, an effective representation of the global feature of the image is finally obtained.
[0136] The described image analysis and comparison module consists of a rough screening unit and a fine screening unit; the rough screening unit is used to map similar fusion features into the same hash bucket through the LSH algorithm; the fine screening unit is used to calculate the cosine similarity of each fusion feature in each hash bucket in turn, and output the fusion features with the cosine similarity higher than the preset similarity threshold to the multi-modal large model verification module;
[0137] By setting that the image analysis and comparison module consists of a rough screening unit and a fine screening unit; the rough screening unit is used to map similar fusion features into the same hash bucket through the LSH algorithm; the fine screening unit is used to calculate the cosine similarity of each fusion feature in each hash bucket in turn, and output the fusion features with the cosine similarity higher than the preset similarity threshold to the multi-modal large model verification module; that is, the comparison of image similarity is carried out in two stages. The LSH hash algorithm is used to quickly cluster similar images, greatly reducing the computational complexity and adapting to the scenario of massive data. Through the accurate matching of cosine similarity, it is ensured that high-similarity features enter the multi-modal verification, taking into account both speed and accuracy.
[0138] The LSH algorithm (Locality-Sensitive Hashing algorithm) maps the fusion features (feature vectors) of similar image data into the same hash bucket; for a given fusion feature v, the LSH algorithm maps it into the hash bucket through the hash function h(v).
[0139]
[0140] where <a,v> represents the inner product of the random vector a and the fusion feature v; b represents the random offset; w represents the width of the hash bucket. Similar vectors have a high probability of being mapped into the same hash bucket, so that a large number of images in the tender documents of each supplier can be quickly grouped according to similarity, greatly reducing the number of images that need to be carefully compared subsequently.
[0141] The calculation formula of cosine similarity is:
[0142]
[0143] Among them, <v1, v2> represents the inner product of vector v1 and vector v2; ||v1|| represents the L2 norm of vector v1; ||v2|| represents the L2 norm of vector v2.
[0144] The multimodal large model verification module is used to extract text features from text data, perform fusion analysis on the text features and the fusion features to obtain an analysis result, and output the analysis result to the detection report output module;
[0145] By setting the multimodal large model verification module to extract text features from text data and perform fusion analysis on the text features and the fusion features to obtain an analysis result, that is, combining the two modal data of picture data and text data for bid rigging detection, fully considering their relevance, and thus greatly improving the accuracy of bid rigging detection.
[0146] The multimodal large model combines the fusion features and the text features and performs comprehensive analysis through the way of joint embedding space, which can be a model based on the Transformer architecture; assuming the fusion feature is I and the text feature is T, the model judges whether there is bid rigging behavior through the learning function f(I, T); for example, this function can be learned through a multi-layer neural network, taking the fusion feature and the text feature as inputs, and after being processed by a series of fully connected layers and activation functions, finally outputting a probability value representing the possibility of bid rigging behavior.
[0147] The detection report output module is used to output the bid rigging detection result according to the analysis result;
[0148] Through the bid rigging detection model for automatic extraction and comparison of picture features (fusion features), a large number of pictures can be processed quickly; for example, after uniformly converting the format and normalizing the size of the pictures in the data preprocessing stage, using the ResNet-50 network to extract global features in parallel and the SIFT network to extract local features synchronously, the feature extraction time of a single picture is greatly shortened; in the comparison stage, the LSH algorithm is used for rough screening and fast grouping, and the cosine similarity calculation in the hash bucket can quickly locate suspected similar pictures, which can improve the review efficiency by several times or even dozens of times compared with manual work, greatly accelerating the procurement process, enabling the purchaser to complete the review of tender documents in a shorter time and promoting the progress of the project.
[0149] Extract multi-dimensional features of images through SIFT (local details) and ResNet-50 (global structure), combine with text semantic features to construct a more comprehensive representation of tender documents; at the same time, integrate image (local + global features) and text data, break through the limitations of single modality, and can identify hidden collusive bidding features such as contradictions in text content and inconsistencies between pictures and texts, thereby greatly improving the accuracy of collusive bidding detection.
[0150] The loss function is weighted and constructed based on the L1 loss function, feature matching loss function, triplet loss function, contrast loss function, and cross-entropy loss function; the L1 loss function is used to measure the difference in the tender documents before and after being processed by the data preprocessing module; the feature matching loss function is used to calculate the difference between the fused features extracted by the image feature extraction module and the target features; the triplet loss function is used to train the image analysis and comparison module to learn the similarity of pictures; the contrast loss function is used to measure the similarity between different modality data input to the multi-modal large model verification module; the cross-entropy loss function is used to measure the accuracy of the collusive bidding detection results output by the detection report output module.
[0151] By setting the loss function to be weighted and constructed based on the L1 loss function, feature matching loss function, triplet loss function, contrast loss function, and cross-entropy loss function; the L1 loss function is used to measure the difference in the tender documents before and after being processed by the data preprocessing module; the feature matching loss function is used to calculate the difference between the fused features extracted by the image feature extraction module and the target features; the triplet loss function is used to train the image analysis and comparison module to learn the similarity of pictures; the contrast loss function is used to measure the similarity between different modality data input to the multi-modal large model verification module; the cross-entropy loss function is used to measure the accuracy of the collusive bidding detection results output by the detection report output module; that is, different loss functions are adopted for different modules to match the calculation tasks of different modules, effectively improving the training effect of the collusive bidding detection model, strengthening the relevance between feature extraction and target tasks, enhancing the model's ability to distinguish different tender documents, and thereby greatly improving the accuracy of collusive bidding detection.
[0152] The dataset construction module is specifically used for:
[0153] Obtain a large number of historical tender documents, perform data cleaning preprocessing on the picture data and text data carried by each historical tender document, label the collusive bidding behaviors of the preprocessed picture data and text data, and construct a dataset based on the labeled picture data and text data.
[0154] The collusive bidding detection model training module is specifically used for:
[0155] Divide the dataset into a training set, a validation set, and a test set based on a ratio of 7:2:1. Train the bid-rigging detection model using the training set until the loss value of the loss function is less than a preset loss threshold. During the training process, continuously optimize the hyperparameters of the bid-rigging detection model, including at least the learning rate, dropout rate, batch size, and number of iterations.
[0156] Validate the trained bid-rigging detection model using the validation set. Determine whether the detection accuracy is greater than a preset accuracy threshold. If not, the validation fails, and the dataset is expanded and training continues. If so, the validation passes, and:
[0157] Test the bid-rigging detection model that has passed the validation using the test set. Determine whether the confidence level is greater than a preset confidence threshold. If not, the test fails, and the dataset is expanded and training continues. If so, the test passes, and training ends.
[0158] Divide the dataset into a training set, a validation set, and a test set based on a ratio of 7:2:1. Train the bid-rigging detection model using the training set until the loss value of the loss function is less than a preset loss threshold. During the training process, continuously optimize the hyperparameters of the bid-rigging detection model, including at least the learning rate, dropout rate, batch size, and number of iterations. Validate the trained bid-rigging detection model using the validation set and test the bid-rigging detection model that has passed the validation using the test set. That is, during the training process of the bid-rigging detection model, continuous optimization, validation, and testing are carried out, thereby greatly improving the accuracy of bid-rigging detection.
[0159] The bid-rigging detection module is specifically used for:
[0160] Deploy the bid-rigging detection model that has passed the test through a distributed computing architecture, obtain real-time tender documents, and input the real-time tender documents into the deployed bid-rigging detection model to perform bid-rigging detection through the deployed bid-rigging detection model and obtain the bid-rigging detection result.
[0161] Deploy the bid-rigging detection model that has passed the test through a distributed computing architecture, that is, deploy different modules of the bid-rigging detection model on different computing nodes for parallel processing, effectively improving the bid-rigging detection efficiency and meeting the needs of large-scale tender document review.
[0162] By adopting a distributed computing architecture, it is convenient to add computing nodes to process more tender documents, meeting the needs of expanding the procurement scale; at the same time, it has good scalability. As the means of bid rigging and collusion change, new image type processing modules can be easily added or existing algorithms can be optimized. For example, if new image tampering means appear, the feature extraction algorithm can be optimized for this situation, or relevant training data can be added to the multi-modal large model to continuously maintain its effective detection ability for various bid rigging and collusion behaviors, and long-term ensure the healthy and orderly progress of procurement activities.
[0163] In summary, the advantages of the present invention are as follows:
[0164] 1. A bid rigging and collusion detection model is created through a data preprocessing module, an image feature extraction module, an image analysis and comparison module, a multi-modal large model verification module, and a detection report output module, and the loss function of the bid rigging and collusion detection model is set; then a large number of historical tender documents are obtained, and a data set is constructed after preprocessing each historical tender document; the data set is divided into a training set, a validation set, and a test set, and the bid rigging and collusion detection model is trained through the training set until the loss value of the loss function is less than a preset loss threshold, and then the bid rigging and collusion detection model is verified and tested through the validation set and the test set respectively; then the bid rigging and collusion detection model that passes the test is deployed, and bid rigging and collusion detection is carried out through the deployed bid rigging and collusion detection model to obtain bid rigging and collusion detection results, and the bid rigging and collusion detection model is continuously optimized through the bid rigging and collusion detection results; that is, the automatic detection of bid rigging and collusion is carried out through the bid rigging and collusion detection model, overcoming the problems of low review efficiency, poor review accuracy, and lack of multi-dimensional review ability caused by relying on manual comparison traditionally. Moreover, the bid rigging and collusion detection model is trained through the data set constructed by historical tender documents, fully considering the particularity and complexity of the bid rigging and collusion detection business, and combined with the continuous optimization of the bid rigging and collusion detection model, ultimately greatly improving the accuracy and efficiency of bid rigging and collusion detection.
[0165] 2. By setting a data preprocessing module to extract image data and text data from tender documents, preprocess the image data including at least format conversion and size normalization, and preprocess the text data including at least outlier removal, missing value filling, noise character removal, stop word removal, and HTML tag removal, effectively improving the data quality of image data and text data, facilitating subsequent detection and analysis.
[0166] 3. The image feature extraction module is composed of a local feature extraction unit, a global feature extraction unit, and a feature fusion unit; the local feature extraction unit is used to extract local features from the image data and is constructed based on the SIFT network; the global feature extraction unit is used to extract global features from the image data and is constructed based on the ResNet-50 network; the feature fusion unit is used to fuse the local features and the global features; that is, the feature extraction stage combines local features and global features, effectively improving the feature extraction performance, and thus greatly improving the accuracy of bid rigging detection.
[0167] 4. The image analysis and comparison module is composed of a rough screening unit and a fine screening unit; the rough screening unit is used to map similar fusion features into the same hash bucket through the LSH algorithm; the fine screening unit is used to calculate the cosine similarity of each fusion feature in each hash bucket in turn, and output the fusion features with the cosine similarity higher than the preset similarity threshold to the multi-modal large model verification module; that is, the comparison of image similarity is carried out in two stages. The LSH hashing algorithm is used to quickly cluster similar images, greatly reducing the computational complexity and adapting to the massive data scenario. Accurate matching through cosine similarity ensures that high-similarity features enter the multi-modal verification, taking into account both speed and accuracy.
[0168] 5. The multi-modal large model verification module is used to extract text features from the text data, and fuse and analyze the text features and the fusion features to obtain the analysis result, that is, the bid rigging detection is carried out by combining the two modal data of image data and text data, fully considering their relevance, and thus greatly improving the accuracy of bid rigging detection.
[0169] 6. The loss function is constructed by weighting the L1 loss function, the feature matching loss function, the triplet loss function, the contrast loss function, and the cross-entropy loss function; the L1 loss function is used to measure the difference between the bid documents before and after being processed by the data preprocessing module; the feature matching loss function is used to calculate the difference between the fusion features extracted by the image feature extraction module and the target features; the triplet loss function is used to train the image analysis and comparison module to learn the similarity of images; the contrast loss function is used to measure the similarity between different modal data input to the multi-modal large model verification module; the cross-entropy loss function is used to measure the accuracy of the bid rigging detection result output by the detection report output module; that is, different loss functions are adopted for different modules to match the computational tasks of different modules, effectively improving the training effect of the bid rigging detection model, strengthening the relevance between feature extraction and the target task, enhancing the model's ability to distinguish different bid documents, and thus greatly improving the accuracy of bid rigging detection.
[0170] 7. Divide the dataset into a training set, a validation set, and a test set in a ratio of 7:2:1. Train the bid-rigging detection model using the training set until the loss value of the loss function is less than a preset loss threshold. During the training process, continuously optimize the hyperparameters of the bid-rigging detection model, including at least the learning rate, dropout rate, batch size, and number of iterations. Validate the trained bid-rigging detection model using the validation set, and test the bid-rigging detection model that has passed the validation using the test set. That is, during the training process of the bid-rigging detection model, continuous optimization, validation, and testing are carried out, thereby greatly improving the accuracy of bid-rigging detection.
[0171] 8. Deploy the bid-rigging detection model that has passed the test through a distributed computing architecture, that is, deploy different modules of the bid-rigging detection model on different computing nodes for parallel processing, effectively improving the bid-rigging detection efficiency and meeting the needs of large-scale tender document review.
[0172] 9. Automatically extract and compare image features (fusion features) through the bid-rigging detection model, which can quickly process a large number of images. For example, after uniformly converting the format and normalizing the size of the images in the data preprocessing stage, use the ResNet-50 network to parallelly extract global features and the SIFT network to synchronously extract local features, greatly shortening the feature extraction time for a single image. In the comparison stage, the LSH algorithm quickly groups through rough screening, and the cosine similarity calculation within the hash bucket can quickly locate suspected similar images. Compared with manual review, the review efficiency can be increased by several times or even dozens of times, greatly accelerating the procurement process, enabling the purchaser to complete the review of tender documents in a shorter time, and promoting the progress of the project.
[0173] 10. By comprehensively applying a variety of advanced algorithms, combining global features and local features to extract comprehensive feature information, the cosine similarity calculation provides an objective similarity measure, and the multi-modal large model verifies the bid-rigging situation from multiple dimensions of images and texts. Taking the review of qualification certificate images as an example, it can not only find whether the certificates submitted by suppliers are the same through image feature comparison, but also combine text data to judge whether the supplier qualifications match the project requirements, avoiding misjudgment and missed judgment, greatly improving the accuracy of bid-rigging behavior detection, and effectively maintaining the fairness of bidding.
[0174] 11. Detecting through the bid-rigging detection model ensures that all tender documents are reviewed according to the same standard. Whenever and wherever, as long as the same tender documents are input, or the image content is the same but only the shooting background or angle is different, the obtained detection results are the same, ensuring the objectivity of the review process and results, and providing a reliable and fair review conclusion for the purchaser.
[0175] 12. By adopting a distributed computing architecture, it is convenient to add computing nodes to process more tender documents, meeting the needs of the expanding procurement scale. At the same time, it has good scalability. As the means of bid rigging and collusion change, it can easily add new picture type processing modules or optimize existing algorithms. For example, if new picture tampering means appear, the feature extraction algorithm can be optimized for this situation, or relevant training data can be added to the multi-modal large model to continuously maintain its effective detection ability for various bid rigging and collusion behaviors, and long-term ensure the healthy and orderly progress of procurement activities.
[0176] 13. Extract multi-dimensional features of images through SIFT (local details) and ResNet-50 (global structure), combine text semantic features, and construct a more comprehensive tender document representation. At the same time, integrate pictures (local + global features) and text data, break through the limitations of a single modality, and can identify hidden bid rigging features such as contradictions in text content and inconsistencies between pictures and texts, thereby greatly improving the accuracy of bid rigging and collusion detection.
[0177] Although the specific implementation manners of the present invention have been described above, those skilled in the art of this technology should understand that the specific embodiments we described are illustrative rather than used to limit the scope of the present invention. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present invention should be covered by the scope protected by the claims of the present invention.
Claims
1. A method for detecting bid rigging based on a multimodal large model, characterized in that: It includes the following steps: Step S1: Create a bid-rigging detection model based on a data preprocessing module, a picture feature extraction module, a picture analysis and comparison module, a multi-modal large model verification module, and a detection report output module, and set the loss function of the bid-rigging detection model; Step S2: Obtain a large number of historical bidding documents, and preprocess each of the historical bidding documents to construct a data set; Step S3: Divide the data set into a training set, a validation set, and a test set. Train the bid-rigging detection model with the training set until the loss value of the loss function is less than a preset loss threshold, and then verify and test the bid-rigging detection model with the validation set and the test set respectively; Step S4: Deploy the bid-rigging detection model that passes the test, and perform bid-rigging detection with the deployed bid-rigging detection model to obtain a bid-rigging detection result; Step S5: Continuously optimize the bid-rigging detection model based on the bid-rigging detection result; 2. The method for detecting bid rigging based on a multi-modal large model according to claim 1, wherein: In the step S1, the data preprocessing module, the picture feature extraction module, the picture analysis and comparison module, the multi-modal large model verification module, and the detection report output module are connected in sequence; the output end of the data preprocessing module is connected to the input end of the multi-modal large model verification module; The data preprocessing module is used to extract picture data and text data from the bidding documents, preprocess the picture data including at least format conversion and size normalization, preprocess the text data including at least outlier removal, missing value filling, noise character removal, stop word removal, and HTML tag removal, input the preprocessed picture data into the picture feature extraction module, and input the preprocessed text data into the multi-modal large model verification module; The picture feature extraction module consists of a local feature extraction unit, a global feature extraction unit, and a feature fusion unit; the local feature extraction unit is used to extract local features from the picture data and is constructed based on the SIFT network; the global feature extraction unit is used to extract global features from the picture data and is constructed based on the ResNet-50 network; the feature fusion unit is used to fuse the local features and the global features to obtain a fused feature, and input the fused feature into the picture analysis and comparison module; The picture analysis and comparison module consists of a rough screening unit and a fine screening unit; the rough screening unit is used to map similar fused features into the same hash bucket through the LSH algorithm; the fine screening unit is used to calculate the cosine similarity of each fused feature in each hash bucket in sequence, and output the fused features with the cosine similarity higher than a preset similarity threshold to the multi-modal large model verification module; The multi-modal large model verification module is used to extract text features from the text data, perform fusion analysis on the text features and the fused features to obtain an analysis result, and output the analysis result to the detection report output module; The detection report output module is used to output a bid-rigging detection result based on the analysis result; The loss function is constructed by weighting the L1 loss function, feature matching loss function, triplet loss function, contrastive loss function, and cross-entropy loss function; the L1 loss function is used to measure the difference between the bid documents before and after being processed by the data preprocessing module; the feature matching loss function is used to calculate the difference between the fused features extracted by the image feature extraction module and the target features; the triplet loss function is used to train the image analysis and comparison module to learn the similarity of images; the contrastive loss function is used to measure the similarity between different modality data input by the multi-modal large model verification module; the cross-entropy loss function is used to measure the accuracy of the bid-rigging detection results output by the detection report output module.
3. The method for detecting bid rigging based on a multi-modal large model according to claim 1, wherein: The specific steps of step S2 are as follows: Obtain a large number of historical bid documents. After preprocessing the image data and text data carried by each historical bid document through data cleaning, label the bid-rigging behavior for each preprocessed image data and text data, and construct a dataset based on the labeled image data and text data.
4. The anti-collusive bidding detection method based on a multi-modal large model according to claim 1, wherein: The specific steps of step S3 are as follows: Divide the dataset into a training set, a validation set, and a test set according to a ratio of 7:2:
1. Train the bid-rigging detection model through the training set until the loss value of the loss function is less than a preset loss threshold. During the training process, continuously optimize the hyperparameters of the bid-rigging detection model, including at least the learning rate, random dropout rate, batch size, and number of iterations. Validate the trained bid-rigging detection model through the validation set, and judge whether the detection accuracy is greater than a preset accuracy threshold. If not, the validation fails, and the dataset is expanded and training continues; if so, the validation passes, and: Test the bid-rigging detection model that has passed the validation through the test set, and judge whether the confidence level is greater than a preset confidence level threshold. If not, the test fails, and the dataset is expanded and training continues; if so, the test passes, and the training ends.
5. The method for detecting bid rigging based on a multimodal large model according to claim 1, wherein: The specific steps of step S4 are as follows: Deploy the bid-rigging detection model that has passed the test through a distributed computing architecture, obtain real-time bid documents, and input the real-time bid documents into the deployed bid-rigging detection model to perform bid-rigging detection through the deployed bid-rigging detection model to obtain bid-rigging detection results.
6. A bid rigging detection system based on a multi-modal large model, characterized in that: It includes the following modules: A bid-rigging detection model creation module, which is used to create a bid-rigging detection model based on the data preprocessing module, image feature extraction module, image analysis and comparison module, multi-modal large model verification module, and detection report output module, and set the loss function of the bid-rigging detection model. A dataset construction module, which is used to obtain a large number of historical bid documents and construct a dataset after preprocessing each historical bid document. A bid-rigging detection model training module, which is used to divide the dataset into a training set, a validation set, and a test set, train the bid-rigging detection model through the training set until the loss value of the loss function is less than a preset loss threshold, and then validate and test the bid-rigging detection model through the validation set and the test set respectively. The bid rigging detection module is used to deploy the bid rigging detection model that has passed the test, and perform bid rigging detection through the deployed bid rigging detection model to obtain the bid rigging detection result; The bid rigging detection model optimization module is used to continuously optimize the bid rigging detection model based on the bid rigging detection result.
7. The anti-collusive tendering detection system based on a multi-modal large model according to claim 6, wherein: In the bid rigging detection model creation module, the data preprocessing module, the image feature extraction module, the image analysis and comparison module, the multi-modal large model verification module, and the detection report output module are connected in sequence; the output end of the data preprocessing module is connected to the input end of the multi-modal large model verification module; The data preprocessing module is used to extract image data and text data from the tender documents, perform preprocessing on the image data including at least format conversion and size normalization, and perform preprocessing on the text data including at least outlier removal, missing value filling, noise character removal, stop word removal, and HTML tag removal, input the preprocessed image data into the image feature extraction module, and input the preprocessed text data into the multi-modal large model verification module; The image feature extraction module consists of a local feature extraction unit, a global feature extraction unit, and a feature fusion unit; the local feature extraction unit is used to extract local features from the image data and is constructed based on the SIFT network; the global feature extraction unit is used to extract global features from the image data and is constructed based on the ResNet-50 network; the feature fusion unit is used to fuse the local features and global features to obtain fused features, and input the fused features into the image analysis and comparison module; The image analysis and comparison module consists of a rough screening unit and a fine screening unit; the rough screening unit is used to map similar fused features into the same hash bucket through the LSH algorithm; the fine screening unit is used to calculate the cosine similarity of each fused feature in each hash bucket in sequence, and output the fused features with the cosine similarity higher than the preset similarity threshold to the multi-modal large model verification module; The multi-modal large model verification module is used to extract text features from the text data, perform fusion analysis on the text features and the fused features to obtain an analysis result, and output the analysis result to the detection report output module; The detection report output module is used to output the bid rigging detection result based on the analysis result; The loss function is weighted and constructed based on the L1 loss function, the feature matching loss function, the triplet loss function, the contrast loss function, and the cross-entropy loss function; the L1 loss function is used to measure the difference before and after the data preprocessing module processes the tender documents; the feature matching loss function is used to calculate the difference between the fused features extracted by the image feature extraction module and the target features; the triplet loss function is used to train the image analysis and comparison module to learn the similarity of images; the contrast loss function is used to measure the similarity between different modal data input by the multi-modal large model verification module; the cross-entropy loss function is used to measure the accuracy of the bid rigging detection result output by the detection report output module.
8. The anti-collusive tendering detection system based on a multi-modal large model according to claim 6, characterized in that: The dataset construction module is specifically used for: Obtain a large number of historical tender documents. After preprocessing the image data and text data carried by each of the historical tender documents through data cleaning, label the preprocessed image data and text data for bid rigging behavior, and construct a dataset based on the labeled image data and text data.
9. The anti-collusive bidding detection system based on a multi-modal large model according to claim 6, characterized in that: The bid rigging detection model training module is specifically used for: Divide the dataset into a training set, a validation set, and a test set according to a ratio of 7:2:
1. Train the bid rigging detection model through the training set until the loss value of the loss function is less than a preset loss threshold. During the training process, continuously optimize the hyperparameters of the bid rigging detection model, including at least the learning rate, dropout rate, batch size, and number of iterations. Verify the trained bid rigging detection model through the validation set, and judge whether the detection accuracy is greater than a preset accuracy threshold. If not, the verification fails, and the dataset is expanded and training continues; if so, the verification passes, and: Test the bid rigging detection model that has passed the verification through the test set, and judge whether the confidence level is greater than a preset confidence threshold. If not, the test fails, and the dataset is expanded and training continues; if so, the test passes, and the training ends.
10. A bid-rigging detection system based on a multi-modal large model according to claim 6, characterized in that: The bid rigging detection module is specifically used for: Deploy the bid rigging detection model that has passed the test through a distributed computing architecture, obtain real-time tender documents, and input the real-time tender documents into the deployed bid rigging detection model to perform bid rigging detection through the deployed bid rigging detection model and obtain the bid rigging detection result.
Citation Information
Cited By
Purchase file verification method and device based on multi-mode and rule optimization
CN120806825A
Burgling behavior identification method and electronic equipment
CN121456514A