Building engineering bid risk identification method and system based on RAG-LLM

Through the RAG-LLM-based construction project bid-rigging risk identification method, comprehensive bidding features are collected and extracted. Combined with the bid-rigging case knowledge base and LLM evaluation, the problem of identifying bid-rigging behavior in a complex bidding environment is solved, and accurate identification and real-time early warning are achieved.

CN120765362AActive Publication Date: 2025-10-10CCCC(XIAMEN)INFORMATION CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511221424.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-10-10
Estimated Expiration
2045-08-29

AI Technical Summary

Technical Problem

Existing technologies make it difficult to accurately identify and warn of bidding collusion in a complex multi-bid environment. Traditional methods are prone to misjudgment or omissions when faced with small-margin bidding and staggered bidding.

Method used

A RAG-LLM-based method for identifying bid-rigging risks in construction projects is adopted. By collecting multi-source bidding data from a multi-source data platform, comprehensive bidding features are extracted, and the RAG-LLM system is used to recall matching knowledge segments from the bid-rigging case knowledge base, and risk level assessment is performed in combination with LLM.

Benefits of technology

It achieves precise identification and early warning of bidding collusion, improves the accuracy and robustness of identification, can output early warning reports in real time, adapt to new bidding collusion techniques and keep maintenance costs low.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765362A_ABST
    Figure CN120765362A_ABST
Patent Text Reader

Abstract

The invention discloses a construction engineering bidding risk identification method and system based on RAG-LLM, and relates to the field of construction engineering bidding analysis, and the method comprises the steps: collecting bidding multi-source data of a to-be-analyzed target bidder from a multi-source data platform, the multi-source data comprises a bidding quotation sequence of the target bidder in each section of the target project and a bidding submission timestamp sequence formed by timestamps of submitting quotations each time, and the multi-source data further comprises a cross-project participation record of the target bidder; comprehensive bidding features are extracted from the bidding multi-source data, wherein the comprehensive bidding features include the differential quotation anomaly degree, the staggered time sequence anomaly degree and the cross-project activeness degree; and inputting the bidding comprehensive features into an RAG-LLM system to recall a knowledge paragraph set from a bidding case knowledge base through the RAG, and driving the LLM to output a corresponding construction engineering bidding risk level. Therefore, fusion of structured behavior modeling and semantic analogy large model reasoning is realized, and the accuracy and interpretability of surrounding mark recognition are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of construction project procurement analysis, and in particular to a method and system for identifying the risk of bid rigging in construction projects based on RAG-LLM (Retrieval-Augmented Generation Large Language Model). Background Art

[0002] Before a project is implemented, contractors are typically selected through an open bidding process. In the procurement of multi-bid projects organized by large real estate developers, participants often use strategies such as "segmented bidding" and "micro-spread bidding" to conceal bid collusion. Several affiliated companies under the same controlling entity submit bids for different sub-bids. By keeping bid differences within a very small range (e.g., 0.1%-0.5%) or staggering bids (e.g., submitting bids within a few minutes of each other), traditional detection methods based on price clustering or legal entity relationship maps are unable to detect anomalies. Summary of the Invention

[0003] The present application provides a method, system, storage medium, computer program product and electronic device for identifying the risk of bid rigging in construction projects based on RAG-LLM, which is used to at least solve the problem in the current related technologies that it is impossible to accurately identify and warn of bid rigging in a complex multi-bid bidding environment.

[0004] In a first aspect, an embodiment of the present application provides a method for identifying bid-rigging risks in construction projects based on RAG-LLM, the method comprising: collecting multi-source bid data of a target bidder to be analyzed from a multi-source data platform; the multi-source data comprises a bid quotation sequence of the target bidder for each bid section in a target project and a bid submission timestamp sequence composed of a timestamp of each bid submission, and the multi-source data further comprises a cross-project participation record of the target bidder; extracting bid comprehensive features from the bid multi-source data; the bid comprehensive features comprise: a micro-difference bid anomaly corresponding to the bid quotation sequence, a time-staggered sequence anomaly corresponding to the bid submission timestamp sequence, and a cross-project activity corresponding to the cross-project participation record; inputting the bid comprehensive features into the RAG-LLM system, so as to recall a set of bid-rigging knowledge paragraphs matching the bid comprehensive features from a bid-rigging case knowledge base through RAG, and using preset prompt words to drive LLM to output a corresponding construction project bid-rigging risk level.

[0005] In a second aspect, an embodiment of the present application provides a construction project bid-rigging risk identification system based on RAG-LLM, the system comprising: a multi-source data acquisition unit, for collecting multi-source bid data of the target bidder to be analyzed from a multi-source data platform; the multi-source data comprises a bid quotation sequence of the target bidder for each bid section in the target project and a bid submission timestamp sequence composed of a timestamp of each bid submission, and the multi-source data further comprises a cross-project participation record of the target bidder; a comprehensive feature extraction unit, for extracting bid comprehensive features from the bid multi-source data; the bid comprehensive features comprise: the micro-difference bid anomaly corresponding to the bid quotation sequence, the staggered sequence anomaly corresponding to the bid submission timestamp sequence, and the cross-project activity corresponding to the cross-project participation record; a RAG-LLM risk analysis unit, for inputting the bid comprehensive features into the RAG-LLM system, so as to recall a set of bid-rigging knowledge paragraphs matching the bid comprehensive features from the bid-rigging case knowledge base through RAG, and use preset prompt words to drive LLM to output the corresponding construction project bid-rigging risk level.

[0006] In a third aspect, an electronic device is provided, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps of the construction project bid-rigging risk identification method based on RAG-LLM of any embodiment of the present application.

[0007] In a fourth aspect, an embodiment of the present application provides a storage medium on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the construction project bid rigging risk identification method based on RAG-LLM of any embodiment of the present application are implemented.

[0008] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the construction project bid-rigging risk identification method based on RAG-LLM in any embodiment of the present application.

[0009] The RAG-LLM-based construction project bid-rigging risk identification method and system provided in this application can produce at least the following technical effects: (1) By uniformly collecting and multi-dimensionally integrating bid quotation sequences, bid submission timestamp sequences, and cross-project participation records, we can accurately characterize key features of collusion patterns such as "slightly differential bidding," "off-peak bidding," and "rotating participation by affiliated companies." Based on these rich and complementary feature dimensions, we can simultaneously capture abnormal signals in price, time, and participation structure during risk identification. Multi-source and multi-dimensional bidding feature modeling significantly improves the detection rate and fine-grained location capabilities of hidden bid-rigging behaviors.

[0010] (2) By introducing comprehensive bidding features into the RAG retrieval module, we can recall in real time the knowledge segments of bid-rigging cases that best match these features. This serves as context to drive the LLM risk assessment. This not only makes the model's judgments more accurate than real cases, but also provides a traceable knowledge basis for each risk output. Thus, by integrating semantic-level case matching with reasoning, we significantly improve the robustness and interpretability of the bid-rigging identification system when dealing with new or variant bid-rigging techniques.

[0011] Through this technical solution, the entire process from data collection, feature extraction, case retrieval to risk assessment is automated. Once new bidding data is generated, an early warning report can be output in real time, greatly shortening the response delay of bid evaluation. At the same time, through the scalable knowledge base maintenance and prompt word adjustment mechanism, the system can quickly adapt to new industry rules and bid-rigging cases without retraining the model, maintaining long-term effectiveness and low maintenance costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0013] Figure 1 A flowchart illustrating an example of a method for identifying the risk of bid rigging in construction projects based on RAG-LLM according to an embodiment of the present application is shown; Figure 2 An operational flowchart of an example of extracting cross-project activity according to an embodiment of the present application is shown; Figure 3 An operational flowchart of an example of performing tensor decomposition based on a non-negative block tensor decomposition model to calculate cross-project activity according to an embodiment of the present application is shown; Figure 4 An operational flowchart illustrating an example of predicting and outputting a construction project bid-rigging risk level by driving the RAG-LLM system according to an embodiment of the present application is shown; Figure 5 An operational flow chart showing an example of fine-tuning for a RAG-LLM system; Figure 6 A schematic diagram showing an experimental comparison of Accuracy and F1 Score. Figure 7 A structural block diagram of an example of a construction project bid-rigging risk identification system based on RAG-LLM according to an embodiment of the present application is shown; Figure 8 A system framework diagram of an example of a construction project bid-rigging risk identification system based on RAG-LLM according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0014] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0015] It should be noted that among the current relevant technologies, the following three methods are mainly relied upon to identify bid-rigging risks: price feature analysis method, legal person and relationship mining method, and static rule screening method.

[0016] Price feature analysis methods primarily rely on statistical thresholds or distance / density-based clustering algorithms to detect bid concentration. For example, bids with bid price differences less than a predetermined percentage (e.g., 1%) are considered suspicious of bid rigging; or all bids are clustered in numerical space, with extreme clusters identified as anomalies. However, when bidders control their bids within a very narrow range (e.g., 0.1%–0.3%), and this range closely overlaps with the actual bid price, simple threshold detection or cluster boundaries can misinterpret these as normal fluctuations, leading to the overlooking of highly concealed bid rigging cases. Furthermore, when the bidding scale or price differences between project sections are inherently small, either loosening or tightening the threshold setting can result in significant underreporting or false positives.

[0017] In the legal person and related relationship mining method, the corporate relationship map constructed based on data such as legal person equity, director relationships, and historical cooperation records can reveal long-term and relatively stable control chains. However, bidders often quickly reconstruct legal person relationships through short-term equity transfers, the establishment of temporary shell companies, or the use of proxy holding accounts, which makes the conventional map have blind spots in terms of timeliness. Especially when multiple subsidiaries appear alternately between different projects and the equity transfer is completed before and after the bid deadline, this "dynamic shell" relationship is difficult to capture in real time, resulting in a lag in the relationship mining module and an inability to incorporate it into risk assessment.

[0018] Static rule-based screening methods often include "bidding from the same address," "registering with the same email address," and "excessively small bid-to-bid spreads." These rules are easy to implement and can quickly filter out some low-level bid-rigging behaviors. However, once the rules are fixed, they are difficult to adapt to new bid-rigging tactics. Given the diverse nature of bidders, rule maintenance requires frequent manual updates. Furthermore, the addition of new rules requires expert research and historical data verification, resulting in high costs and slow response times.

[0019] It should be understood that the purpose of the above description of the current related art is only to facilitate the public to better understand the inventive spirit and motivation of this application, and is not to be construed as limiting this application. In addition, the technical solutions described in the above-mentioned current related art are not prior art and may also be undisclosed technical solutions, such as solutions under research or in the laboratory stage.

[0020] Figure 1 A flowchart of an example of a method for identifying the risk of bid rigging in construction projects based on RAG-LLM according to an embodiment of the present application is shown.

[0021] Regarding the execution entity of the method of the embodiment of the present application, it can be any controller or processor with computing or processing capabilities. By adopting a fusion mechanism of structured modeling and generative language reasoning, structured data is used to provide a quantitative basis for behavior analysis, and the large language model uses retrieval enhancement and generation capabilities to conduct case-based understanding and inductive judgment of behavior patterns, thereby improving the accuracy of recognition.

[0022] In some examples, it can be integrated into an electronic device or terminal through software, hardware, or a combination of software and hardware, and the type of terminal or electronic device can be diverse, such as a mobile phone, tablet computer, or desktop computer, etc.

[0023] like Figure 1 As shown, in step S110, multi-source bidding data of the target bidder to be analyzed is collected from the multi-source data platform.

[0024] In the construction project bidding supervision scene, the information involved in the bidding behavior is relatively scattered, and there are data barriers or inconsistent data granularity between different systems. Therefore, by breaking through the data interface of heterogeneous platforms, the behavior information of the target bidder in multiple dimensions can be uniformly gathered. Specifically, the multi-source data platform includes but is not limited to government public resource transaction platform, large real estate enterprise procurement system, industry supervision platform and third-party evaluation agency database, etc. The data interface is standardized designed to support automatic retrieval and format unification of bidding information. In addition, the target bidder can refer to any bidder and can be specified by user input, such as bidding in the round or entities with bidding intention for a specific project, etc.

[0025] Here, the multi-source data includes the bidding submission timestamp sequence composed of the bidding price sequence of the target bidder in each bid section of the target project and the timestamp of each submission price, and the multi-source data also includes the cross-project participation record of the target bidder.

[0026] Specifically, the bidding price sequence refers to the valid bidding price data submitted by the target bidder in multiple bid sections of a target project, for example, the data is indexed by bid section number and organized as an ordered sequence according to the submission order of the price, which helps to identify price coordination behavior (such as 0.1%-0.5% micro-difference control). The bidding submission timestamp sequence refers to the submission time corresponding to each price operation, which is used to analyze the concentration or misplacement of bidding operations in the time dimension, and provides a data basis for identifying misplacement bidding behavior and other coordination behavior patterns. The cross-project participation record refers to the participation of the target bidder in other projects in the same period or near time window, including bidding time, project attributes, cooperation subjects, bid section quantity and other information, which facilitates the description of the breadth and frequency of its bidding behavior, and helps to identify whether it participates frequently, whether it participates with other subjects, whether it concentrates bidding on a certain bidding organization, etc. The above multi-source data is cleaned, desensitized and stored in an internal database structure, so as to construct a structured and timely bidding behavior portrait by uniformly collecting and standardizing the multi-dimensional data such as bidding price, timestamp and participation record.

[0027] In step S120, the bidding comprehensive features are extracted from the bidding multi-source data. Thus, the original bidding multi-source data collected is converted into structured numerical features that can be used for pattern recognition.

[0028] Here, the bidding comprehensive features include: the micro-difference price abnormality corresponding to the bidding price sequence, the misplacement sequence abnormality corresponding to the bidding submission timestamp sequence, and the cross-project activity corresponding to the cross-project participation record.

[0029] Regarding the degree of abnormality in bids with small differences, if the bid differences across multiple bid sections for a target bidder remain within a very small range (e.g., 0.1%-0.5%), this indicates the possibility of coordination through controlling bid spacing. For example, a sliding window method can be used to calculate local fluctuations. This method calculates multiple indicators, including the bid standard deviation, the mean of adjacent bid differences, the range (maximum-minimum), and relative volatility (relative to the tender control price), to create a comprehensive anomaly score. A higher score indicates abnormal bid aggregation and possible manipulation.

[0030] Regarding the abnormality of misaligned timing sequences, if a controlling entity manipulates multiple companies to bid in multiple bidding sections, they often precisely control the timing of bid submission for each bidder to prevent the system from detecting simultaneous manipulation. However, this manipulation may result in deviations from the normal bidding time distribution (for example, a "final sprint" pattern), which in turn generates a corresponding quantitative abnormality metric for the misaligned timing sequence.

[0031] Cross-project activity is used to assess the intensity and degree of correlation among target bidders in similar projects over the same period. If a company participates in a large number of projects within a short period of time and frequently appears with several other companies in the same bidding section or for the same tendering unit, it may be suspected of engaging in a networked bid-rigging operation. Activity can be calculated in a variety of ways. For example, it can be derived through a comprehensive calculation based on bidding frequency, project distribution density, and the proportion of repeat partners. A time-series weighting mechanism can also be introduced to weight short-term, concentrated, and bursty bidding behavior.

[0032] Therefore, the "behavior-anomaly" correspondence is established through the comprehensive bidding characteristics, which realizes the mapping transformation from raw data to structured knowledge, enabling the system to characterize complex and hidden collaborative behavior characteristics in a data-driven manner.

[0033] In step S130, the comprehensive bidding characteristics are input into the RAG-LLM system to recall a set of bid-rigging knowledge paragraphs that match the comprehensive bidding characteristics from the bid-rigging case knowledge base through RAG, and use preset prompt words to drive LLM to output the corresponding construction project bid-rigging risk level.

[0034] Here, we use the RAG (Retrieval-Augmented Generation) mechanism to match case knowledge and combine it with quantitative features at a semantic level, enabling analogical reasoning and graded judgment of bid-rigging risk. For example, the target bidder's comprehensive feature vector is input into a vector search engine, which then retrieves the bid-rigging case passages most similar to the proposed behavior pattern from the bid-rigging case knowledge base.

[0035] It should be noted that the bid-rigging case knowledge base pre-stores multiple bid-rigging case knowledge sections, each indexed using knowledge processing metadata. For example, these case knowledge sections can be pre-processed to extract features such as bid price distribution, temporal coordination patterns, and enterprise network relationships associated with previously identified bid-rigging behaviors. Based on the specific bid-rigging behavior patterns in the case (e.g., bid aggregation coordination, bid-rigging network behavior), one or more features, including micro-difference bid anomaly, time-staggered sequence anomaly, and cross-project activity, are selected to construct indexing metadata for the corresponding case knowledge section, thereby achieving feature matching with the input comprehensive bid features.

[0036] After obtaining a collection of knowledge passages that match the bid's comprehensive features, these are combined with the original comprehensive features to construct prompt words, which are then fed into a large language model (such as Deepseek or the Qwen series). The prompt words are clearly structured, including project background, behavioral patterns, knowledge references, and output requirements. This guides the model through case-based reasoning, determining whether the behavior poses a risk of bid rigging and outputting a risk level (e.g., low, medium, or high). This also generates reasoning evidence, such as "The bids are extremely aggregated, with low time intervals, suggesting suspected human coordination."

[0037] More specifically, in the prompt word template, the input format is designed as follows: Role definition: For example, "You are a bid evaluation analyst and need to analyze whether the target bidder has the risk of bid rigging." Input feature summary: including quote volatility, bid time distribution, activity value, etc. Retrieval paragraph summary: Displays the 3-5 most relevant cases of bid-rigging knowledge paragraphs; Question instructions: For example, "Please judge whether there is a risk of bid rigging based on the above content, and give the risk level and reasons."

[0038] This enables intelligent reasoning between structured abnormal behavior and semantic knowledge, overcoming the limitations of traditional models that can only make "yes / no" judgments but cannot explain the "why." By combining case semantics, behavioral characteristics, and generative understanding, the system not only identifies bid-rigging risks but also provides readable, traceable, and referenceable evidence for regulators, significantly enhancing the interpretability and credibility of the identification results.

[0039] Regarding the extraction of the micro-difference quotation anomaly in step S120, in some examples of the embodiments of the present application, the bid quotation sequence is segmented into sliding windows, the normalized values ​​of the bid quotation data in each window are calculated, and the distribution pattern of the window normalized value sequence is analyzed through a lightweight neural network model to output the corresponding micro-difference quotation anomaly.

[0040] In the context of identifying bid-rigging risks in construction projects, bidders often engage in bid-rigging by adjusting bids by very small amounts within a short period of time and within a specific bid section to circumvent global threshold detection, making them difficult to distinguish within the full bid section or full time series statistics. Using sliding window segmentation combined with in-window normalization, we can amplify "micro-difference" features within the local context. By learning the distribution pattern of the normalized series using a lightweight neural network, the model can automatically capture the subtle differences between "normal fluctuations" and "abnormal aggregations."

[0041] More specifically, in the sliding window partition, let the target bidder's bid sequence in a certain bidding section be , formula (1) Where, For the The quotation, is the total number of quotes. Take the window length and step length ,Will Divided into Overlapping subsequences:

[0042] , formula (2) in , formula (3) Where, Indicates the window length, represents the step length, Indicates the number of windows, For the A subsequence of quotes for the window.

[0043] It should be noted that bidders often conceal their intent to collude by creating "small clusters" of bids within a series of consecutive bids. This pattern can be easily masked by normal bid fluctuations from a global perspective. Sliding window segmentation can break down long sequences into several local segments, focusing on uncovering subtle variations in bid distribution within each segment. Decomposing long sequences into contextually relevant local signals significantly increases sensitivity to "short-term concentrated bids" patterns. The window overlap strategy ensures that anomalies are captured even at the cut boundaries, reducing missed detections.

[0044] In the within-window normalization, for each window Compute the local mean and standard deviation: , formula (4) Then we get the normalized sequence , formula (5) Composition vector .

[0045] Where, 、 Respectively The mean and standard deviation of the window, To prevent division by zero for small constants, is the normalized window feature vector.

[0046] It should be understood that the absolute range of bids for different projects and bid sections varies significantly. Directly comparing raw values ​​can lead to misjudgments due to different scales. Standardizing within each window can smooth out scale differences and highlight localized fluctuations.

[0047] Furthermore, by designing a lightweight network, the model structure can include a one-dimensional convolutional neural network (1D-CNN). This extracts sequence microstructure features between and within windows, making it suitable for sliding window sequential data. For example, a two-layer Conv1D+ReLU+Pooling+Flatten+Dense model can be used, with a sigmoid output. The training dataset can be constructed by combining historical normal bidding samples with identified bid-rigging samples, labeled as "normal" or "slightly abnormal." Furthermore, a binary cross-entropy loss function can be used, with the optimization objective being a binary classification of window-by-window slight differences.

[0048] Calculate the abnormality for each window: , formula (6) Where, Indicates the The degree of abnormality of the window's micro-difference quotes, Indicates that the parameter is Specifically, , and the larger the value, the more abnormal the "micro-difference" of the window.

[0049] In this way, the sequence obtained after window normalization carries information about the local fluctuation shape. The lightweight neural network can learn the subtle differences in the shape and spectrum of normal and abnormal fluctuations, thereby assigning an abnormality score to each window.

[0050] Finally, the local anomalies of all windows are aggregated into the differential quote anomaly of the entire sequence. The maximum value can be taken to capture the strongest abnormal signal:

[0051] , formula (7) Where, Indicates the output of the abnormality of the micro-difference quotes of the entire sequence.

[0052] This approach, through sliding windows and normalization, highlights small bid aggregations within the same context, capturing subtle bid-rigging intentions that are often overlooked by traditional global statistics. Furthermore, the adjustable step size and window length automatically adapt to projects with varying bid densities and bidding periods, achieving multi-scale compatibility.

[0053] Regarding the extraction of cross-project activity in step S130, in some examples of the embodiments of the present application, the interval values ​​of adjacent timestamps in the bid submission timestamp sequence are calculated to obtain a corresponding interval value sequence; the interval value sequence is reconstructed based on the autoencoder model to obtain a corresponding reconstruction error, and the autoencoder model is pre-trained based on a sample set of a normal mode; based on the reconstruction error, the abnormality of the time-delayed sequence is determined.

[0054] It should be noted that bidders who collude often deliberately split their submission times within the bid deadline—for example, by evenly distributing submissions a few minutes apart or by skipping submissions during off-peak hours. Computing the time interval sequence between adjacent submissions can transform this "staggered" behavior into a quantifiable timing signal.

[0055] Specifically, the target bidder’s bid submission timestamp sequence is extracted from the procurement platform log. ,in For the The exact time of the commit.

[0056] Calculate adjacent time differences .

[0057] If multiple bids are submitted concurrently, they must first be grouped by "project ID + bidder ID" and calculated separately for each group. This converts discrete timestamp data into a sequence of equal-length intervals, forming a unified foundation for time series analysis.

[0058] It should be noted that the bid submission intervals are often very unevenly distributed. Specifically, normal bids may be concentrated in the last few minutes, resulting in a large number of small bids; while the staggered strategy spreads the submissions throughout the bidding period, resulting in long-tail intervals. Modeling is difficult and requires normalization and transformation to highlight abnormal patterns.

[0059] Specifically, by logarithmic transformation Compress the long-tail distribution and enhance the model's ability to identify the three types of intervals: short, medium and long. This effectively alleviates the long-tail problem through logarithmic transformation, making the autoencoder easier to converge. In addition, missing or abnormal extreme values ​​(such as or exceeds the upper limit) is truncated to the trainable range according to the upper and lower limits.

[0060] Next, time series reconstruction is performed using an autoencoder. It should be noted that a normal sequence of bid submission intervals exhibits a distinct "one submission, one final sprint" pattern. Pre-training the autoencoder with only normal samples allows the model to reconstruct this normal pattern with low error. However, scattered or skipping bid submissions at different times are difficult to reconstruct, resulting in high error. Furthermore, the autoencoder's encoder and decoder can employ a symmetrical multi-layer LSTM structure.

[0061] In the pre-training of the autoencoder, the training set can select a large number of time series submitted by real "non-colluded" bidders to ensure representativeness and time period diversity, and the mean square reconstruction error can be used to determine the reconstruction level baseline by monitoring the reconstruction error distribution on independent normal samples.

[0062] After deployment, the model only performs forward reasoning, and the preprocessed input Sequence, output reconstructed sequence . Calculate the reconstruction error at each time point The LSTM autoencoder is trained with normal samples to enable it to capture sequence dependencies and periodic trends.

[0063] Furthermore, the timing anomaly can reflect the deviation of the entire bidding sequence from the normal pattern. By reconstructing the statistical distribution of the error, the local peak or the overall dispersion can be quantified into a single indicator.

[0064] Specifically, the error sequence is obtained: , formula (8) Error aggregation is performed by adopting the quantile strategy: , formula (9) Through error aggregation, both local and global deviations are included in the measurement, taking into account both sudden mistime and long-term spread risk modes. Exceeding the threshold When the high risk alert is triggered, the threshold It can be calculated using a dynamic statistical method based on the median absolute deviation (MAD) to achieve dynamic adaptive adjustment.

[0065] During the practice of this invention, we discovered that the timing of bids in construction project bidding often follows a "final submission sprint" pattern, which refers to the behavior of most bidders submitting all their bid documents in a short period of time before the bid deadline. Specifically, for a long period after the bidding begins, the frequency of submissions is low and the intervals are large. However, once the deadline approaches (for example, the last 5–10 minutes), the majority of bids suddenly converge, and the intervals between submissions shorten dramatically, resulting in a "sprint-like" high-density submission.

[0066] The main reasons for this model are as follows: bidders often wait and see, prepare documents in the early stages, and submit them all at once in the final stage after confirming that they are correct and the price strategy is finalized, in order to reduce the risk of being imitated or copied by other bidders. Represents the time series characteristics, when When you are away from the deadline, Large and fluctuating; when As the deadline approaches, A sharp drop forms a period of "sharp decline". By modeling this pattern as a "normal" baseline, the autoencoder can learn the typical "low-frequency wait-and-see + high-frequency sprint" time series structure; any behavior that significantly deviates from this structure (such as uniform submission throughout the day or deliberately staggered distribution) can be highlighted in the reconstruction error, helping to identify mistimed risks and avoid risks.

[0067] Figure 2 An operational flowchart of an example of extracting cross-project activity according to an embodiment of the present application is shown.

[0068] like Figure 2 As shown, in step S210, a three-dimensional project participation tensor of the target bidder is constructed.

[0069] It's important to note that in the construction sector, a large number of bidders repeatedly participate across multiple projects and time periods. Some unusual bidders often participate in multiple projects or bidding sections within the same timeframe (such as the same quarter or month). This high-frequency, cross-project participation is rare in normal business, but it's a common tactic used by bid-rigging rings to evade regulation. Therefore, by constructing a three-dimensional tensor of "bidder-project-timeframe," we can quantify bidders' behavioral trajectories in the construction market from a multi-dimensional, interactive perspective, capturing unconventional, group-driven patterns of activity.

[0070] Specifically, the three dimensions of the three-dimensional project participation tensor are bidders, bidding projects, and bidding time slices. Specifically, in the first dimension ( ) assigns a unique number to all bidders, and in the second dimension ( ) assigns unique numbers to all engineering projects / sections during the observation period. In the third dimension ( ) Divide the analysis period into uniform time slices (such as months, weeks, or quarters), and each time slice is uniquely numbered.

[0071] Constructing a 3D tensor by tensor padding ,element Indicates the bidder Is it in the time slice? Participated in the project . Specifically, Indicates participation record. Indicates no participation record. The multidimensional tensor structure breaks the bottleneck of two-dimensional tables in expressing multi-project, multi-period, and group behaviors.

[0072] In step S220 , the three-dimensional participation tensor is subjected to tensor decomposition to obtain a bidder factor matrix to calculate cross-project activity.

[0073] It should be noted that in a three-dimensional tensor, normal bidders exhibit a dispersed distribution with weak correlation between projects and time slots, while abnormal bidders (such as bid-rigging rings) often exhibit an implicit pattern of highly concentrated projects and synchronized bursts of time slots. Therefore, various non-restrictive high-order tensor decomposition methods, such as CP decomposition, Tucker decomposition, dynamic tensor decomposition, or block tensor decomposition, can be used to automatically extract the "participation pattern factors" hidden in the multidimensional data. This high-dimensional behavior is compressed into a low-dimensional feature matrix. This decomposition enables automatic "clustering" of participation patterns, thereby revealing the patterns of group activity among bidders and distinguishing the behavioral differences between active bidders and average bidders.

[0074] Specifically, after tensor decomposition, each row of the bidder factor matrix reflects the bidder exist The "intensity of participation" in high-level activity patterns. By integrating the intensity scores of each pattern, we can quantify the degree of activity across projects and time periods throughout the observation period, serving as a sensitive indicator of bid-rigging behavior.

[0075] For example, for bidders , its cross-project activity can be defined as: , formula (10) Where, For bidders In mode The next score.

[0076] Through the embodiments of the present application, the bidder-project-time slice is constructed as a three-dimensional tensor, and the tensor is decomposed at a high order, which can capture complex group collaborative behaviors and help automatically discover hidden group activity patterns. Normal bidders often show sparse participation in scattered projects and sporadic time, while abnormal groups show high-dimensional interactions with highly concentrated projects and synchronized time outbreaks under specific patterns. The bidder factor matrix obtained by decomposition can quantify each person's score on multiple potential active patterns, and then be refined into a single cross-project activity indicator, providing objective and explainable multi-dimensional behavioral characteristics for subsequent risk identification.

[0077] In some examples of the embodiments of the present application, cross-project activity can be obtained through non-negative block tensor decomposition (NBTD).

[0078] Figure 3 An operational flowchart of an example of performing tensor decomposition based on a non-negative block tensor decomposition model to calculate cross-project activity according to an embodiment of the present application is shown.

[0079] like Figure 3 As shown, in step S310, the three-dimensional participation tensor is divided according to the number of bidders and project categories to generate multiple non-overlapping sub-tensor blocks.

[0080] To address the computational complexity and significant pattern variation in large-scale bidder-project-time three-dimensional participation data, we first partition the original three-dimensional tensor based on bidder attributes and project categories. By dividing the bidder set into several groups and the project set into several categories, we obtain multiple non-overlapping sub-tensor blocks, each containing only the participation data of the corresponding bidder and project at each time slice.

[0081] Specifically, the three-dimensional participation tensor According to the number of bidders Group, divided by project category Group, generated non-overlapping sub-tensor blocks .

[0082] , formula (11) Where, represents the original participating tensor, represents the total number of bidders, Indicates the total number of items, Indicates the total number of time slices, and are the number of bidder groups and the number of project categories, Indicates the Group Bidder and The subtensor corresponding to the class item, For the Number of bidders in the group, For the Number of class items.

[0083] In this way, each sub-tensor block can be independently decomposed and analyzed, and the overall operation complexity can be significantly reduced by using distributed or multi-thread parallel computing. In addition, the bidders can be grouped according to the enterprise size, historical participation frequency or geographical region; and the project category can be defined based on the engineering type, bidding scale or bid section function. In this way, the participation modes of different groups and categories are often different, and by blocking, the interference between cross-group and cross-category modes can be avoided, and the sensitivity of tensor decomposition to each mode can be improved.

[0084] In step S320, two-stage non-negative decomposition is performed on each sub-tensor block.

[0085] On each sub-tensor block, the non-negative CP decomposition is used to extract the “bidder-project-time” three-dimensional mode factor. The non-negative constraint ensures that the factor matrix extracted by decomposition has an interpretable intensity meaning, which can intuitively reflect the participation degree of each bidder in a specific project and time dimension. The factor values are all non-negative, representing the participation frequency or activity level; therefore, the original data can be approximated with a lower rank, and the dominant behavior mode can be efficiently refined.

[0086] Specifically, in the time dimension decomposition of the first stage, for each sub-tensor block directly perform CP decomposition.

[0087] , formula (12) In the formula, is the CP decomposition rank, indicating how many groups of factors are extracted; represents the weight of the th group of factors, represents the bidder factor vector, represents the project factor vector, represents the time factor vector and , represents the vector outer product operation.

[0088] In the spatial dimension decomposition of the second stage, the factor set after CP decomposition is regarded as a tensor component, and non-negative spatial decomposition is directly performed.

[0089] Specifically, in the spatial dimension decomposition of the second stage, the bidder factor matrix and the project factor matrix obtained by CP decomposition are regarded as a composite matrix, and then the non-negative bidder subspace basis is directly multiplied on the left and the non-negative project subspace basis is directly multiplied on the right to perform non-negative spatial decomposition: , formula (13) In the formula, , represents the bidder subspace basis matrix, represents the item subspace basis matrix, represents the core interaction matrix, Represents the spatial decomposition rank.

[0090] Here, we further apply non-negative Tucker decomposition to the bidder and project factor matrices, compressing the CP decomposition factors in a subspace, reducing redundant factors and extracting higher-order coupling features. This allows us to map similar bidders or projects into the same subspace, thereby aggregating group behavior patterns. This effectively aggregates the coupling relationships between bidders and projects into a small number of high-meaning features, which can then be used to assist in identifying cross-project collaboration groups through group association recognition.

[0091] In step S330 , a spatiotemporal activity matrix of each sub-block is constructed.

[0092] , formula (14) Where, Indicates the spatiotemporal activity of the sub-block, represents the matrix transpose operation, , Indicates the spatiotemporal activity of sub-blocks.

[0093] In step S340 , the time-space activity matrix is ​​solved by introducing a time decay factor to obtain the cross-project activity of the bidder in different time slices.

[0094] After combining the subspace decomposition results with the time factor, exponential time decay is introduced to weight the activity of each bidder in different time slices. Specifically, the target bidder In time slice Activity It is expressed by the following formula:

[0095] , formula (15) Where, is a matrix Middle Rank Column elements, is the time attenuation coefficient, is the current analysis time point, Represents the time decay factor.

[0096] Here, the attenuation coefficient This can be determined based on the length of the project bidding cycle and the regulatory window, and can also be fitted with historical data. This gives higher weight to the most recent data at the time of regulation, enabling timely detection of sudden concentrated bidding by bid-rigging parties. Using decaying weights, the weighting gradually deemphasizes future participation, minimizing the impact of outdated behavior. Focusing analysis on spatiotemporal activity close to the current review time point is more aligned with the dynamics of bid-rigging on-site.

[0097] In step S350, the cross-project activities of each time slice are aggregated to obtain the target bidder Cross-project activity .

[0098] , formula (16) Where, represents the time weighting function, is the slope parameter of the weighting function, is the center offset time of the weighted function, Indicates the target bidder In time slice In some embodiments, the slope Offset from center It can be optimized through maximum likelihood estimation to match the typical project submission rhythm.

[0099] Here, the temporal and spatial activity results of bidders in all time slices are fused into a single score using logistic regression time weighting to reflect their overall cross-project participation intensity and concentration. , which can reflect the bidder’s cross-project gameplay throughout the entire life cycle.

[0100] Through the embodiments of this application, a non-negative block tensor decomposition model is used to divide the original "bidder-project-time" three-dimensional participation data into more structurally homogeneous sub-blocks, and a multi-level, non-negatively constrained decomposition operation is performed on each sub-block. This not only helps to reveal the potential collaboration patterns between local bidder groups and specific project categories, but also facilitates the separation and modeling of dynamic changes in the time dimension. As a result, the ability to detect cross-project and cross-time period bid collusion can be significantly improved in the large-scale and complex engineering bidding data environment, and an abnormal activity signal with clear spatiotemporal attribution and structural explanation can be output.

[0101] It should also be noted that in the behavior of collusion in construction network bidding, collusion is often distributed in different projects and multiple time periods in a coordinated and controlled manner in a grouped and block-like manner. Comparing the NBTD decomposition method adopted in the embodiment of the present application with the traditional global CP or Tucker decomposition, NBTD can first divide the sub-tensor blocks according to project category and bidder group, and then independently impose non-negative constraints on each block for decomposition, which perfectly fits the "local collaboration + global heterogeneity" strategy. By combining block division with non-negative decomposition, on the one hand, the collaborative behavior pattern within each block is highlighted, and on the other hand, the mutual interference of patterns between different blocks is avoided, providing clearer and more meaningful high-order features for risk quantification and large model reasoning.

[0102] Figure 4An operational flowchart of an example of predicting and outputting the risk level of bid rigging in a construction project by driving the RAG-LLM system according to an embodiment of the present application is shown.

[0103] like Figure 4 As shown, in step S410, the top K case knowledge paragraphs with the highest matching degree are retrieved from the bid-rigging case knowledge base for the abnormality of slight difference bidding, abnormality of staggered sequence and cross-project activity in the comprehensive bidding characteristics.

[0104] In identifying bid-rigging risks, "Narrow Bid Anomaly," "Time-Misaligned Sequence Anomaly," and "Cross-Project Activity" reflect typical risk signals in the price, timing, and network dimensions, respectively. Retrieving the most relevant historical cases for each signal ensures the model draws on the most representative evidence when reasoning about that signal, rather than relying on a generalized global case set. Multi-dimensional feature recall also avoids the blind spots of certain risk patterns that can be detected by a single search method.

[0105] In some implementations, three feature searches may be initiated in parallel and the results may be returned asynchronously to shorten the overall recall latency.

[0106] In step S420, duplicates are removed from the retrieved case knowledge paragraphs and combined to obtain a set of matching bid-rigging knowledge paragraphs; the total number of paragraphs in the knowledge paragraph set is less than or equal to 3K.

[0107] It's important to note that a single paragraph of knowledge related to a bid-rigging case may simultaneously possess multiple risk characteristics, resulting in the possibility of identical or highly similar case paragraphs being retrieved across the three branches. Without deduplication, redundant cases will cause subsequent reasoning modules to repeatedly consume the same evidence, wasting computing resources and weakening the diversity of the reasoning chain. By merging and limiting the total number of paragraphs to 3K (K being the number of single-feature retrievals), we can ensure coverage of all signals while controlling the input size.

[0108] After completing the RAG search content recall, these contents can be used as references and input into LLM in combination with preset prompt words.

[0109] To ensure that RAG-LLM performs reasoning in a sequential, branching, and hierarchical manner under multi-dimensional risk factors, rather than a one-time "black box" output, the preset prompt words adopt a chain-of-thought prompting structure. The RAG-LLM system also includes a trunk module and multiple branch reasoning modules, that is, a "one trunk + three parallel branches" structure is constructed within the model, so that each risk factor can be reasoned independently and can also participate in the final integrated judgment.

[0110] For example, the design idea of ​​the chain prompt word instruction is as follows: Unified prefix instructions: "Based on the following information, please analyze step by step and give a conclusion."

[0111] Step 1 (price dimension): "[First branch] based on the abnormality of micro-difference quotes Based on your knowledge of relevant cases, determine whether the target bidder is at risk of price collusion. Please provide your conclusion and the paragraph number of the cited case.

[0112] Step 2 (Time Series Dimension): "[Second Branch] Combined with the abnormality of the out-of-time series Based on your knowledge of timing cases, determine whether there is any mistimed avoidance behavior. Please explain your reasons and cite the passage.

[0113] Step 3 (Cross-project dimension): "[Third branch] Combine cross-project activity And related cases, determine whether there is cross-project bidding collusion? And provide a paragraph citing the case. "

[0114] Step 4 (Trunk Integration): "[Trunk] Summarizes the conclusions of the first three steps, outputs the target bidder's comprehensive bid-rigging risk level (high / medium / low), and lists the key influencing factors and corresponding case paragraphs to form a complete knowledge chain."

[0115] Specifically, during inference, the prompt and the recalled knowledge paragraph are fed into the model together. The branch modules generate their own conclusions in parallel, which are then aggregated into the main branch. The first three steps, Steps 1–3, are each bound to a corresponding branch adapter. The branch input only contains the branch factor and the full set of case paragraphs. The final prompt, Step 4, is bound to the main branch LoRA adapter for comprehensive judgment. This achieves a reasoning architecture that is both separate and distinct: branches focus on individual factors, while the main branch unifies attribution. Chaining prompts reduces contextual interference in simultaneous reasoning, improving the accuracy and interpretability of the model for multi-factor discrimination.

[0116] In step S431, based on the first branch reasoning module, the abnormality degree of the micro-difference quotation and the set of bid-rigging knowledge paragraphs are inferred to determine the corresponding price-rigging risk identification result.

[0117] Here, the price-rigging risk identification results are used to indicate whether there is a price-rigging risk and the price anomaly reference case knowledge cited in the reasoning.

[0118] Here, based on the "abnormality of micro-difference quotes" It reflects suspicious signals of bidders' small concentrated quotations in different sections of the same project. Through branch reasoning combined with a collection of knowledge paragraphs on bid-rigging cases and driven by prompts, it generates corresponding price factor identification results.

[0119] For example, the model output format can be: Conclusion: Existence / non-existence of price string risk.

[0120] Reason: … Reference case: Case #12, paragraphs 3-5.

[0121] Here, by setting an Adapter branch dedicated to the price factor, the model can still maintain high accuracy when facing mixed features, and the output reference cases enhance the auditability of the judgment, allowing supervisors to directly locate the case paragraphs for review.

[0122] In step S432, the second branch reasoning module reasons the time error sequence abnormality and the surrounding bid knowledge paragraph set to determine the corresponding time error avoidance behavior identification result.

[0123] Here, the time error avoidance behavior identification result is used to indicate whether there is a time error avoidance behavior risk and the time sequence abnormality reference case knowledge cited in the reasoning. Time error avoidance focuses on the abnormal distribution of bidding submission rhythm. The reconstruction error is quantified, and the dedicated Adapter branch analyzes this factor to accurately locate various time error avoidance behavior patterns. In addition, the second branch Adapter can only update the LoRA for the time sequence sensitive layer. Through the parallel design of independent branches, cascading errors caused by misjudgment of the previous branch are avoided.

[0124] In step S433, the third branch reasoning module reasons the cross-project activity and the surrounding bid knowledge paragraph set to determine the corresponding cross-project string identification result.

[0125] Here, the cross-project string identification result is used to indicate whether there is a cross-project string risk and the project association abnormality reference case knowledge cited in the reasoning. Cross-project activity Reflects the abnormally high-frequency participation behavior of bidders in multiple projects at different times, and the dedicated branch can focus on this factor and combine case explanations to group operation patterns.

[0126] In step S440, the main module integrates the price string risk identification result, the time error avoidance behavior identification result, and the cross-project string identification result to determine the construction engineering surrounding bid risk level, and generates the surrounding bid reasoning reference knowledge chain in combination with the price abnormality reference case knowledge, the time sequence abnormality reference case knowledge, and the project association abnormality reference case knowledge.

[0127] Here, the main module assumes the "comprehensive view" responsibility, fuses the three branch independent output results with the corresponding reference cases, gives the final surrounding bid risk level, and generates a complete reasoning knowledge chain.

[0128] The exemplary input features of the main module are as follows: 1) Three-branch label vector ; 2) A list of three-branch citation case paragraphs; 3) Prompt: "Based on the conclusions from the first three steps, output the target bidder's overall bid-rigging risk level (high / medium / low), list the key influencing factors and corresponding case paragraphs, and form a complete knowledge chain."

[0129] Then, the backbone adapter integrates the vectors of each branch, and its exemplary output format is as follows:

Comprehensive risk level

Core Factors

Knowledge Chain

[0130] Through the embodiments of the present application, in the identification of bid-rigging risks in construction projects, abnormal signals of different dimensions (small price differences, mistimed submissions, and cross-project activity) each reflect a specific bid-rigging strategy, which is difficult for a single model to take into account. The large language model chain reasoning of "multi-branch parallel + backbone integration" is adopted, combined with RAG's recall of the most relevant cases for multi-dimensional features and LoRA's injection of lightweight adapters into each branch, so that each branch module focuses on learning and interpreting a risk factor. The backbone module then summarizes the conclusions and cited evidence of each branch to form a hierarchical and traceable comprehensive judgment process. In this way, the independence of the reasoning of each factor is maintained, and on the basis of multi-dimensional parallel analysis, the accuracy of single-factor and comprehensive risk judgment is significantly improved, and accurate case references are provided for each step of reasoning, ensuring the logical coherence and end-to-end explainability of the final decision.

[0131] It's important to note that bid-rigging in construction project bidding often manifests as a highly concealed, coordinated pattern involving multiple bidders, multiple projects, and multiple time periods. Traditional static rules, single-point detection, and shallow machine learning methods struggle to effectively identify these complex, interconnected, and knowledge-reliant risk chains.

[0132] The intelligent risk identification system based on RAG-LLM provided by the embodiment of the present application has multi-modal, multi-source heterogeneous data fusion and knowledge-enhanced reasoning capabilities, and can more accurately realize the automatic identification and interpretation of bid-rigging risks.

[0133] However, only by targeting real samples and business characteristics of the construction industry, designing scientific data structuring methods, chain-branch reasoning structures, and efficient fine-tuning mechanisms suitable for the scenarios, can the industry intelligence potential of RAG-LLM be unleashed.

[0134] Figure 5 An operational flowchart of an example of fine-tuning the RAG-LLM system is shown, which uses efficient LoRA parameter adaptation fine-tuning and introduces novel links such as multi-task joint loss and evidence correlation optimization.

[0135] like Figure 5 As shown, in step S510, a structured training sample is constructed based on historical bid-rigging samples.

[0136] Here, the structured training samples include multi-source bid sample features for bid-rigging bidders, historical case descriptions that match these multi-source bid sample features, a main label for the bid-rigging risk level of the bid-rigging bidders, and multiple branch node risk sub-labels. These branch node risk sub-labels are used to indicate the risk status of price collusion, timing avoidance behavior, and cross-project collusion.

[0137] An efficient fine-tuning process first requires the large model to "understand" the chain structure and knowledge traceability logic of risk patterns within business scenarios. To achieve this, the original bid-rigging samples must be structured, mapping the "digital features" of bidding behavior and the text content of historical cases with multi-layered labels, thus establishing a ternary channel of data, knowledge, and labels.

[0138] Specifically, the quotation sequence, bidding timestamp sequence and cross-project participation records of each bidder in the project and its various bidding sections are collected. After normalization, sliding window, time series difference, tensor decomposition and other processing, three categories of digital features (micro-difference quotation anomaly, time-stamped sequence anomaly, and cross-project activity) are extracted.

[0139] In terms of label alignment of historical case content, due to the differentiated performance of the RAG mechanism, it is recommended to adopt an expert-defined approach, where the user inputs information to specify the original bid-rigging sample and retrieves the knowledge paragraph ID or quoted text of the case paragraph most relevant to the sample characteristics in the bid-rigging case knowledge base, including previous bid-rigging penalty announcements, industry risk reports and typical bid-rigging link descriptions, which are highly semantically matched with the current bidding characteristics to achieve evidence chain binding.

[0140] Regarding the annotation of multi-layer labels, each structured sample is assigned a main label (comprehensive bid-rigging risk level) and multiple branch sub-labels (price collusion risk, mistime avoidance risk, cross-project collusion risk). The branch labels directly correspond to each branch node output in the subsequent RAG-LLM reasoning chain.

[0141] In step S520, a set of knowledge paragraphs matching the characteristics of the multi-source bidding samples is retrieved from the bid-rigging case knowledge base through RAG, and semantic encoding is performed to obtain the corresponding RAG knowledge characteristics of the bidding samples.

[0142] It should be noted that while the evidence chain for historical case content is specified above, this content is only used as a label. The content recalled using the corresponding RAG should be used in the inference chain. This is because RAG recall is also an optimization direction for the system. Directly using this historical case content as background knowledge for large language model training would severely limit the system's generalization capabilities and reduce its performance when background knowledge is not fully matched. The specific details of RAG operation are described in other sections above and will not be repeated here.

[0143] In step S530 , the bidding multi-source sample features and the bidding sample knowledge features are concatenated to generate a sample input vector.

[0144] In step S540, LoRA parameter adapters are respectively inserted into the backbone module and each branch reasoning module of the RAG-LLM system, so that the backbone network is used to output the prediction result of the main label of the bid-rigging risk level of the sample input vector, and each branch reasoning module is used to output the prediction result of the risk sub-label of each branch node corresponding to the sample input vector.

[0145] It should be noted that fine-tuning all parameters based on a large model is extremely computationally intensive, has high implementation costs, and is prone to forgetting existing knowledge. Here, we adopt the LoRA low-rank adaptation mechanism, inserting LoRAAdapters into the main and branch inference modules, for example, into the attention layer or FFN layer of a multi-layer Transformer. This enables efficient local parameter updates while ensuring the learning capabilities and business sensitivity of the main branch and branch nodes.

[0146] During training, the structured sample input vector (features + knowledge) is fed simultaneously into the backbone and all branch modules. Each branch LoRA parameter is sensitive only to its branch loss. The backbone LoRA adapter optimizes the primary task loss, and all branch information is ultimately fed back to the backbone module to form a global decision.

[0147] Therefore, the LoRA low-rank adaptation mechanism enables independent fine-tuning channels for each inference chain node and main decision branch. This eliminates the need for full parameter feedback from large models, significantly reducing the number of fine-tuning parameters and engineering implementation complexity, and supports rapid model updates and iterations with limited computing power. Furthermore, each branch module independently adapts to improve sensitivity to local risk factors, while the backbone module globally attributes risk, enhancing the consistency and robustness of the overall decision. Furthermore, the system can also construct new parallel branch modules based on new bid-rigging behaviors, demonstrating excellent scalability.

[0148] In step S550, a joint loss function is constructed using the main task loss, branch node loss, and evidence correlation loss between the RAG knowledge features of the bidding samples and the historical case description content, and the LoRA parameter adapter is jointly trained, so that the RAG-LLM system can simultaneously improve the accuracy of the main label of the bid-rigging risk level, the sub-labels of each branch node, and the case knowledge reference.

[0149] Identifying construction bid-rigging risks is essentially a chained multi-objective optimization problem, requiring both global discriminative capabilities and ensuring the accuracy of risk assessments and cited evidence at each branch node. By jointly optimizing the main task loss, branch node loss, and evidence relevance loss, we can ensure the completeness and resolution of chained reasoning while strengthening the model's ability to reference and interpret business knowledge.

[0150] Specifically, the joint loss function is expressed as follows: , formula (17) Where, represents the main task loss, which is the cross entropy loss of the main label of the bid-rigging risk level; represents the branch node loss, which is the average cross entropy loss of all branch node sublabels; is the evidence relevance loss, which is the semantic relevance loss between the RAG knowledge features of the bidding sample and the description content of the historical case; is the loss weighting coefficient.

[0151] More specifically, the main task loss Multi-classification cross entropy is used to measure the error between the final risk level output by the model and the true main label: , formula (17) Where, Indicates the model outputs the main risk prediction, such as high, medium and low risk levels; Indicates the true primary label.

[0152] Loss of each branch Use binary or multi-class cross entropy respectively: , formula (18) Where, is the number of branches (such as price, time sequence, project), for example, 3. Indicates the branch prediction, Indicates the The actual label of the branch.

[0153] Evidence relevance loss You can use BERTScore or sentence vector cosine similarity to encourage the model to actively cite relevant historical case content in chain reasoning output: , formula (18) Where, A text snippet representing the inference chain output by the model, Indicates relevant historical case paragraphs.

[0154] Through the embodiments of this application, based on the fine-tuning method of LoRA, combined with the main task loss, branch node loss, and evidence relevance loss, through end-to-end joint optimization, the RAG-LLM system can simultaneously improve the comprehensive risk level judgment, branch risk type distinction, and the ability to reference historical case knowledge. As a result, intelligent risk control with multi-factor judgment, chain reasoning, and evidence-based attribution is achieved, enabling the RAG-LLM system to achieve more refined and explainable risk identification output in actual engineering procurement business, improving the level of intelligent supervision of engineering bidding.

[0155] In order to verify the superiority of the "slight difference - time difference - cross-project" three-dimensional chain risk identification solution proposed in the embodiment of this application over traditional methods and other single-dimensional methods, this paper designed the following comparative experiment.

[0156] Regarding the selection of the data set, we can select the bidding records of 15 large-scale infrastructure projects from a provincial public resource trading platform from 2019 to 2022, totaling 100,000 bidding flows, including 500 cases of collusion / bid rigging confirmed by real supervision and several normal cases.

[0157] The selection of evaluation indicators is as follows: Detection accuracy (Accuracy): the overall risk judgment accuracy; F1 score: takes into account both recall of bid-rigging cases and false positives of normal cases; Explanation coverage: The overlap ratio between the model output related case paragraphs and the expert annotated paragraphs is used to measure auditability; Average inference latency: The time required from input to output, used to evaluate the real-time performance of the system.

[0158] Table 1 Experimental methods and comparison schemes

[0159] Baseline A: A common global threshold method that compares the bid price for all sections, the time interval for all periods, and the frequency of all projects to fixed thresholds.

[0160] Baseline B / C / D: respectively verify the effect gain of the single-dimensional method and the addition of RAG-LLM.

[0161] Solution E: The three-dimensional features are input into the corresponding branches in parallel, and then the backbone module makes a comprehensive judgment under the drive of chain prompt words and outputs a knowledge chain.

[0162] Table 2 Experimental results

[0163] The experimental results in Table 2 show that the global threshold (A) has low precision and recall, failing to capture hidden local or cross-item patterns. Single-dimensional approaches (B / C) offer slight improvements within their own dimensions but still lack the ability to integrate multiple factors. The price dimension plus RAG-LLM (D) significantly improves accuracy through knowledge augmentation, but lacks explanation coverage. This solution (E), through three-dimensional parallel branching and chained hinting, improves overall accuracy by approximately 7%, improves F1 score by approximately 8%, and achieves 89.3% explanation coverage, significantly outperforming all baselines. Despite an increase in average inference latency, it still meets practical risk warning requirements.

[0164] Figure 6 The figure shows an example of an experimental comparison effect of Accuracy and F1 Score.

[0165] like Figure 6 Figure 1 shows the experimental results of five methods (A-E) on the same test set, with the horizontal axis representing the method number and the vertical axis representing the detection accuracy and F1 score. It can be seen that the proposed method (E) achieves the highest values ​​in both indicators, significantly outperforming the other compared methods.

[0166] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of combined actions, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application. In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0167] Figure 7 A structural block diagram of an example of a construction project bid-rigging risk identification system based on RAG-LLM according to an embodiment of the present application is shown.

[0168] like Figure 7As shown, the RAG-LLM-based construction project bid-rigging risk identification system 700 includes a multi-source data acquisition unit 710, a comprehensive feature extraction unit 720 and a RAG-LLM risk analysis unit 730.

[0169] The multi-source data collection unit 710 is used to collect multi-source bidding data of the target bidder to be analyzed from the multi-source data platform; the multi-source data includes the bid quotation sequence of the target bidder in each bidding section in the target project and the bid submission timestamp sequence composed of the timestamp of each bid submission, and the multi-source data also includes the cross-project participation record of the target bidder.

[0170] The comprehensive feature extraction unit 720 is used to extract comprehensive bidding features from the multi-source bidding data; the comprehensive bidding features include: the abnormality of the micro-difference quotation corresponding to the bidding quotation sequence, the abnormality of the time-stamp sequence corresponding to the bidding submission timestamp sequence, and the cross-project activity corresponding to the cross-project participation record.

[0171] The RAG-LLM risk analysis unit 730 is used to input the comprehensive bidding characteristics into the RAG-LLM system, so as to recall a set of bid-rigging knowledge paragraphs that match the comprehensive bidding characteristics from the bid-rigging case knowledge base through RAG, and use preset prompt words to drive LLM to output the corresponding construction project bid-rigging risk level.

[0172] Figure 8 A system framework diagram of an example of a construction project bid-rigging risk identification system based on RAG-LLM according to an embodiment of the present application is shown.

[0173] like Figure 8 The system architecture shown in the figure has four layers from bottom to top: multi-source data collection layer, knowledge base and localized large model layer, risk business process layer, and function access and result display layer.

[0174] In the multi-source data collection layer, the system automatically aggregates bidding-related data in various formats from internal enterprise relational databases (via SQL interfaces), third-party public interfaces (such as RESTful APIs), and publicly available internet websites (via crawlers). This includes bid quotations, submission logs, project metadata, contract documents, regulatory announcements, and public opinion reports. Missing value interpolation, field standardization, and one-hot encoding are performed on structured tabular data. For unstructured documents, PDFs, images, and other data, OCR and NLP technologies are used for text extraction and semantic parsing. The processed raw data is temporarily stored in a data lake, awaiting centralized scheduling and cleansing by lower-level layers.

[0175] In the knowledge base and localized large model layer, this layer is divided into two parts: 1) Knowledge base construction: Multi-source cleaned data is classified and indexed according to themes such as "enterprise dimension", "contract dimension", "bid dimension", and "policy public opinion dimension". The text slices are converted into vectors using embedding models (such as industry-adapted BERT or specialized vectorization tools), and stored in the enterprise vector knowledge base. This knowledge base provides efficient context completion support for subsequent RAG retrieval.

[0176] 2) Localized RAG-LLM fine-tuning: Based on the previously designed chain prompt framework and LoRA adapter mechanism, fine-tune the large model to the enterprise business context. After inputting the bid features (micro-differential bidding, timing anomalies, cross-project activity) and retrieved case vectors, the model performs branch reasoning and generates risk conclusions and reference evidence.

[0177] In the risk business process layer, it is used to embed AI capabilities into key steps of the bidding process: Requirement analysis: According to the tenderer's demand text, the model automatically identifies potential risk points such as unreasonable clauses or suspicious scoring rules.

[0178] Preparation of tender documents: When preparing the announcement and bid evaluation method, the model gives risk prompts and recommends compliance wording.

[0179] Bid pre-audit and auxiliary evaluation: Intelligent verification of bid eligibility, bidding strategy, and bid scoring, the model combines historical cases to identify risks.

[0180] Contract risk assessment: Before contract negotiation and signing, the model analyzes clause vulnerabilities and credit risks, and outputs review comments.

[0181] Each step calls the localized RAG-LLM, pulls the knowledge base cases and tool functions (such as document generation, clause comparison) in real time, and writes the analysis results back to the unified risk management platform.

[0182] In the functional access and result display layer, it faces business users and provides two sets of unified interfaces for PC and mobile. At any node in the process, the intelligent assistant can be awakened, and the user can ask questions in natural language, and the large model returns:

[0183] Risk report: structured table / card form to display risk level, core factors, and reference case paragraphs; Decision suggestion: provide revision schemes or compliance measures for each risk point; Visual audit: clicking on the case reference can jump to the original clause or regulatory document, supporting full-link traceability.

[0184] Therefore, through the coordination of the above four-layer architecture, the data association, knowledge retrieval and model reasoning of each link of the system are closely connected, realizing the full-link, explainable and auditable intelligent bid-rigging risk identification and management from multi-source data ingestion to intelligent decision support.

[0185] In some embodiments, an embodiment of the present application provides a non-volatile computer-readable storage medium, which stores one or more programs including execution instructions, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to execute any step of the above-mentioned RAG-LLM-based construction project bid rigging risk identification method of the present application.

[0186] In some embodiments, the embodiments of the present application also provide a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer is caused to perform any step of the above-mentioned RAG-LLM-based construction project bid rigging risk identification method.

[0187] In some embodiments, an embodiment of the present application also provides an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the construction project bid rigging risk identification method based on RAG-LLM.

[0188] The above-mentioned product can execute the method provided in the embodiment of this application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the method provided in the embodiment of this application.

[0189] The electronic devices of the embodiments of the present application exist in various forms, including but not limited to: mobile communication devices, ultra-mobile personal computer devices, portable entertainment devices or other onboard electronic devices with data interaction functions.

[0190] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0191] Through the description of the above embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a general hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the relevant technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0192] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for identifying bid-rigging risks in construction projects based on RAG-LLM, characterized by: The method comprises: Collecting multi-source bid data of the target bidder to be analyzed from a multi-source data platform; the multi-source data includes a bid quotation sequence of the target bidder for each bid section in the target project and a bid submission timestamp sequence consisting of a timestamp of each bid submission, and the multi-source data also includes a cross-project participation record of the target bidder; Extracting comprehensive bidding features from the multi-source bidding data; the comprehensive bidding features include: the abnormality of the micro-difference bid corresponding to the bidding quotation sequence, the abnormality of the staggered sequence corresponding to the bid submission timestamp sequence, and the cross-project activity corresponding to the cross-project participation record; The comprehensive bidding characteristics are input into the RAG-LLM system to recall a set of bid-rigging knowledge paragraphs that match the comprehensive bidding characteristics from the bid-rigging case knowledge base through RAG, and use preset prompt words to drive LLM to output the corresponding construction project bid-rigging risk level.

2. The method according to claim 1, characterized in that The extraction of the abnormality of the micro-spread quotation includes: Performing sliding window segmentation on the bidding quotation sequence; Calculate the normalized value of the bid price data within each window; The distribution pattern of the window normalized value sequence is analyzed by a lightweight neural network model to output the corresponding degree of anomaly of the micro-difference quotes.

3. The method according to claim 1, characterized in that The extraction of abnormality of the time-staggered sequence includes: Calculating interval values ​​between adjacent timestamps in the bid submission timestamp sequence to obtain a corresponding interval value sequence; Reconstructing the interval value sequence based on an autoencoder model to obtain a corresponding reconstruction error; the autoencoder model is pre-trained based on a sample set of a normal mode; The abnormality degree of the time-staggered sequence is determined according to the reconstruction error.

4. The method according to claim 1, wherein Extraction of cross-project activity includes: Constructing a three-dimensional project participation tensor of the target bidder, wherein the three dimensions of the three-dimensional project participation tensor are the bidder, the bidding project, and the bidding time slice; The three-dimensional participation tensor is subjected to tensor decomposition to obtain a bidder factor matrix to calculate cross-project activity.

5. The method according to claim 4, wherein The tensor decomposition of the three-dimensional participation tensor to obtain a bidder factor matrix to calculate cross-project activity includes: The three-dimensional participation tensor is decomposed based on a non-negative block tensor decomposition model to obtain a bidder factor matrix to calculate cross-project activity, and specifically includes the following operations: The three-dimensional participation tensor According to the number of bidders Group, divided by project category Group, generated non-overlapping sub-tensor blocks : , Where, represents the original participating tensor, represents the total number of bidders, Indicates the total number of items, Indicates the total number of time slices, and are the number of bidder groups and the number of project categories, Indicates the Group Bidder and The subtensor corresponding to the class item, For the Number of bidders in the group, For the Number of class items; Perform a two-stage non-negative decomposition on each sub-tensor block: In the first stage of time dimension decomposition, for each sub-tensor block Directly perform CP decomposition: , Where, is the CP decomposition rank, indicating how many groups of factors are extracted; Indicates the The weight of the group factor, represents the bidder factor vector, represents the item factor vector, represents the time factor vector and , Represents a vector outer product operation; In the second stage of spatial dimension decomposition, the factor set after CP decomposition is regarded as a tensor component, and the following non-negative spatial decomposition is directly performed: , Where, , represents the bidder subspace basis matrix, represents the item subspace basis matrix, represents the core interaction matrix, represents the spatial decomposition rank; Construct the spatiotemporal activity matrix of each sub-block: , Where, Indicates the spatiotemporal activity of the sub-block, represents the matrix transpose operation, , Indicates the spatiotemporal activity of the sub-block; By introducing the time decay factor to solve the spatiotemporal activity matrix, we can get the target bidder In time slice Cross-project activity : , Where, is a matrix Middle Rank Column elements, is the time attenuation coefficient, is the current analysis time point, represents the time decay factor; Aggregate the cross-project activity of each time slice to obtain the target bidder Cross-project activity : , Where, represents the time weighting function, is the slope parameter of the weighting function, is the center offset time of the weighted function, Indicates the target bidder In time slice Cross-project activity.

6. The method according to claim 1, characterized in that A set of bid-rigging case paragraphs matching the comprehensive bidding characteristics are retrieved from the bid-rigging case knowledge base through RAG, including: Retrieving the top K case knowledge sections with the highest matching degree from the bid-rigging case knowledge base for the degree of abnormality of small difference bids, abnormality of staggered time series, and cross-project activity in the comprehensive bidding characteristics; The retrieved case knowledge paragraphs are deduplicated and combined to obtain a set of matching bid-rigging knowledge paragraphs; the total number of paragraphs in the set of knowledge paragraphs is less than or equal to 3K.

7. The method according to claim 1, characterized in that The preset prompt words adopt a chain reasoning instruction structure, and the RAG-LLM system includes a trunk module and multiple branch reasoning modules; The method of using the preset prompt words to drive the LLM to output the corresponding construction project bid-rigging risk level includes: Based on the first branch reasoning module, reasoning is performed on the abnormality degree of the micro-difference bid and the set of bid-rigging knowledge paragraphs to determine a corresponding price-rigging risk identification result; the price-rigging risk identification result is used to indicate whether there is a price-rigging risk and the price anomaly reference case knowledge cited in the reasoning; The second branch reasoning module is used to reason about the abnormality degree of the timing sequence and the set of bid-rigging knowledge segments to determine a corresponding timing avoidance behavior identification result; the timing avoidance behavior identification result is used to indicate whether there is a timing avoidance behavior risk and the timing anomaly reference case knowledge cited in the reasoning; Based on the third branch reasoning module, the cross-project activity and the set of bid-rigging knowledge paragraphs are reasoned to determine a corresponding cross-project bid-rigging identification result; the cross-project bid-rigging identification result is used to indicate whether there is a risk of cross-project bid-rigging and the project-related abnormal reference case knowledge cited in the reasoning; Based on the backbone module, the price collusion risk identification results, the mistime avoidance behavior identification results and the cross-project collusion identification results are integrated to determine the risk level of collusion in construction projects, and the price anomaly reference case knowledge, the timing anomaly reference case knowledge and the project association anomaly reference case knowledge are combined to generate a collusion reasoning reference knowledge chain.

8. The method according to claim 7, characterized in that Fine-tuning of the RAG-LLM system, including: Constructing structured training samples based on historical bid-rigging samples; the structured training samples include multi-source bid sample features for bid-rigging bidders, historical case descriptions matching the multi-source bid sample features, a main label for the bid-rigging bidder's bid-rigging risk level, and multiple branch node risk sub-labels; the multiple branch node risk sub-labels are used to indicate the risk status of price-rigging bid-rigging risk, timing-avoidance behavior risk, and cross-project bid-rigging risk; Recalling a set of knowledge paragraphs that match the characteristics of the multi-source bidding samples from the bid-rigging case knowledge base through RAG, and performing semantic encoding processing to obtain the corresponding RAG knowledge characteristics of the bidding samples; splicing the bidding multi-source sample features and the bidding sample knowledge features to generate a sample input vector; Insert LoRA parameter adapters into the backbone module and each branch reasoning module of the RAG-LLM system, respectively, so that the backbone network is used to output the prediction result of the bid-rigging risk level main label of the sample input vector, and each branch reasoning module is used to output the prediction result of the risk sub-label of each branch node corresponding to the sample input vector; A joint loss function is constructed using the main task loss, branch node loss, and the evidence correlation loss between the RAG knowledge features of the bidding samples and the description content of the historical case. The LoRA parameter adapter is jointly trained, enabling the RAG-LLM system to simultaneously improve the accuracy of the main label of the bid-rigging risk level, the sub-labels of each branch node, and the case knowledge reference. The joint loss function is expressed as follows: , Where, represents the main task loss, which is the cross entropy loss of the main label of the bid-rigging risk level; represents the branch node loss, which is the average cross entropy loss of all branch node sublabels; is the evidence relevance loss, which is the semantic relevance loss between the RAG knowledge features of the bidding sample and the description content of the historical case; is the loss weighting coefficient.

9. A construction project bid-rigging risk identification system based on RAG-LLM, characterized by: The system comprises: a multi-source data collection unit configured to collect multi-source bid data of a target bidder to be analyzed from a multi-source data platform; the multi-source data comprising a bid quotation sequence of the target bidder for each bid section in the target project and a bid submission timestamp sequence consisting of a timestamp of each bid submission; and the multi-source data further comprising a cross-project participation record of the target bidder; A comprehensive feature extraction unit is configured to extract comprehensive bidding features from the multi-source bidding data; the comprehensive bidding features include: the abnormality of the micro-difference bid corresponding to the bidding quotation sequence, the abnormality of the staggered sequence corresponding to the bid submission timestamp sequence, and the cross-project activity corresponding to the cross-project participation record; The RAG-LLM risk analysis unit is used to input the comprehensive bidding characteristics into the RAG-LLM system, so as to recall a set of bid-rigging knowledge paragraphs that match the comprehensive bidding characteristics from the bid-rigging case knowledge base through RAG, and use preset prompt words to drive LLM to output the corresponding construction project bid-rigging risk level.

Citation Information

Patent Citations

  • Natural language processing-based fraud bidding behavior recognition method and device

    CN113129118A

  • Abnormal huddling bidding and bidding behavior identification method, device and equipment and medium

    CN116484231A

  • Bidding and tendering abnormal behavior identification method and system based on pre-training model

    CN118861698A

  • Civil engineering field knowledge large model retrieval enhancement generation method

    CN119829778A

  • Bidding and tendering data intelligent analysis method and system based on AI technology and storage medium

    CN119850315A