A bid behavior outlier detection method and system fusing multi-ecological modeling
Patent Information
- Application Number
- CN202611013264.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]现有增强学习技术主要依赖智能体在环境中交互、试错和奖励反馈实现策略优化,实际运作中通常以环境状态、动作选择和奖励值为核心进行循环更新,若环境状态只反映单一业务变量或局部行为记录,智能体获得的反馈容易偏向局部收益变化,难以完整呈现多主体之间的隐性关联结构
[0014]与现有技术相比,本发明的优点和积极效果在于:
Smart Images

Figure CN122839205A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of reinforcement learning technology, and in particular to a method and system for outlier detection in bidding behavior that integrates multi-ecosystem modeling. Background Technology
[0002] Reinforcement learning is a branch of artificial intelligence that studies the techniques by which intelligent agents optimize strategies and improve decision-making capabilities through continuous interaction, autonomous trial and error, and reward feedback in their environment.
[0003] Current reinforcement learning techniques primarily rely on agents interacting with the environment, undergoing trial and error, and receiving reward feedback to optimize strategies. In practice, these techniques typically involve iterative updates based on the environment state, action selection, and reward value. If the environment state only reflects a single business variable or local behavioral records, the feedback received by the agent is prone to bias towards localized profit changes, failing to fully represent the implicit relationships between multiple agents. In bidding scenarios, anomalous behavior often manifests not as a single bidder's abnormal pricing, but as multiple bidders exhibiting coordinated proximity in business relationships, network addresses, bid amounts, and bidding time intervals. Adjusting decision-making strategies based solely on single interaction feedback can easily misidentify deliberate price dispersion, staggered document submissions, and network address switching as normal fluctuations. Existing technologies, if focusing only on task gains or immediate feedback in their reward feedback design, may overlook changes in network topology dispersion, allowing groups to evade detection by disconnecting explicit edges, creating random connections, and reducing node concentration. Therefore, improvements are needed. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing a method and system for detecting outliers in bidding behavior that integrates multi-ecosystem modeling.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: a method for detecting outliers in bidding behavior that integrates multi-ecological modeling, comprising the following steps: Based on business registration characteristics, network address characteristics, bid amount characteristics, and bid time interval characteristics, a feature vector of the bidding entity is generated, and a joint state matrix of bidding nodes is constructed. Based on the feature distance of the feature vector of the bidding entity in the joint state matrix of the bidding nodes and the preset association limit, a connection edge is established to generate a multi-ecological bidding environment state map; Based on the multi-ecological bidding environment state map, hidden layer representation vectors are extracted to generate multi-agent communication variables; The multi-agent communication variables are input into the policy layer of the multi-agent reinforcement learning network, and the output includes spoofing action indicators including bid deviation modification amount, time offset modification rate and graph edge disconnection probability. The spoofing action indicators are combined to generate a set of gang-coordinated spoofing bidding actions. The gang's coordinated and disguised bidding actions are applied to the multi-ecological bidding environment state map to generate an adversarial mutation map; The expected profit reward of the gang is obtained based on the gang's comprehensive winning probability and preset profit value in the adversarial mutation graph. The topological entropy masquerading reward is obtained based on the Shannon information of the node degree distribution in the adversarial mutation graph. A comprehensive topological entropy reward feedback signal is generated based on the gang's expected profit reward and the topological entropy masquerading reward. The comprehensive topological entropy reward feedback signal is input into the evaluation layer of the multi-agent reinforcement learning network, and the network weight parameters of the policy layer and the evaluation layer are updated according to the time difference error. The adversarial anomaly map is extracted as an outlier masquerading negative sample and used to construct an adversarial training set with the real and normal bidding map positive samples to update the model parameters of the discriminator classifier; the real bidding business map data is imported into the discriminator classifier to output the probability confidence level. When the probability confidence level is higher than the anomaly judgment limit, the bidding node identification code sequence is output to generate the outlier detection result of the bidding behavior.
[0006] Preferably, the steps for obtaining the joint state matrix of the bidding nodes are as follows: The system obtains the business registration characteristics, network address characteristics, bid amount characteristics, and bid time interval characteristics of each bidding entity. It extracts the registered capital value, establishment time value, business scope code, and legal representative identifier from the business registration characteristics; it extracts the network address range, login device identifier, and access region code from the network address characteristics; it extracts the bid amount value and bid deviation ratio from the bid amount characteristics; and it extracts the file download interval and bid submission interval from the bid time interval characteristics. The system aligns each item according to the bidding entity identifier code and concatenates them in a fixed field order to generate a bidding entity feature vector. A row index is created based on the sorting position of the bid entity identification code, and a column index is created based on the concatenation position of the business registration feature, network address feature, bid amount feature, and bid time interval feature. Each bid entity feature vector is written into the corresponding row position. Missing fields are filled with the median of the existing values in the same field, and non-numerical fields are replaced with a unified coded value to generate a joint state matrix of bidding nodes.
[0007] Preferably, the steps for obtaining the multi-ecological bidding environment status map are as follows: Based on the joint state matrix of the bidding nodes, the feature vectors of any two bidding entities are called row by row, and the differences in business registration features, network address features, bid amount features, and bidding time interval features are compared respectively. The feature distance between the corresponding two feature vectors of the bidding entities is obtained by summarizing. The two bidding entity nodes whose feature distance is lower than the preset association limit are connected, and the two bidding entity nodes whose feature distance is not lower than the preset association limit are kept disconnected, thus generating a multi-ecological bidding environment state map.
[0008] Preferably, the step of obtaining the multi-agent communication variables is as follows: Based on the multi-ecological bidding environment status map, the bidding entity node identifier, connection edge relationship and node feature vector are input into the graph neural network. The first-order neighbor bidding entity nodes of each bidding entity node are read one by one according to the connection edge relationship. The business registration feature component, network address feature component, bid amount feature component and bidding time interval feature component in the first-order neighbor bidding entity nodes are extracted. According to the node adjacency position corresponding to the connection edge relationship, the feature components of the first-order neighbor bidding entity nodes are merged into the feature update position of the current bidding entity node. The hidden layer representation vector of each bidding entity node is extracted. The hidden layer representation vector of each bidding entity node is read according to the node identifier code of the bidding entity. The hidden layer representation vector is sent to the fully connected linear mapping layer. The hidden layer components of business registration, network address, bid amount and bidding time interval in the hidden layer representation vector are mapped one dimension at a time. According to the connection edge relationship in the multi-ecological bidding environment state graph, the one-dimensional mapping result is transmitted between the bidding entity nodes with connection edges. The one-dimensional mapping result is written into the shared information position of the corresponding bidding entity node to generate multi-agent communication variables.
[0009] Preferably, the steps for obtaining the set of actions of the gang in colluding to spoof bidding are as follows: The multi-agent communication variables corresponding to each bidding entity node are input into the policy layer of the multi-agent reinforcement learning network. The corresponding components of the multi-agent communication variables are read according to the bidding action channel, time action channel, and edge relationship action channel in the policy layer. The output results of the bidding action channel are constrained by interval to obtain the bidding deviation modification amount. The output results of the time action channel are constrained by proportion to obtain the time offset modification rate. The output results of the edge relationship action channel are constrained by probability to obtain the graph edge disconnection probability. The bidding deviation modification amount, time offset modification rate, and graph edge disconnection probability are bound and combined according to the bidding entity node identification code to generate a set of gang-colluded fake bidding actions.
[0010] Preferably, the step of obtaining the adversarial mutation map is as follows: The algorithm reads the price deviation modification amount, time offset modification rate, and graph edge disconnection probability of each bidding entity node in the group's collaborative fake bidding action set. It writes the price deviation modification amount into the price amount feature parameter of the corresponding bidding entity node in the multi-ecological bidding environment state graph, and writes the time offset modification rate into the bidding time interval feature parameter of the corresponding bidding entity node in the multi-ecological bidding environment state graph. It performs retention marking or disconnection marking on the corresponding connection edge according to the graph edge disconnection probability. It updates the environment state based on the modified node feature parameters and the connection edge relationship after the disconnection marking, and generates an adversarial mutation graph.
[0011] Preferably, the step of obtaining the comprehensive topological entropy reward feedback signal is as follows: Based on the adversarial mutation graph, the identification code of the bidding entity node corresponding to the gang's collaborative fake bidding action set is read. The bidding amount feature parameter, bidding time interval feature parameter and connection edge disconnection mark corresponding to the bidding entity node identification code are extracted. The winning bid tendency value of each bidding entity node in the adversarial mutation graph is calculated. The winning bid tendency values of each bidding entity node in the same gang are normalized and summarized to obtain the gang's comprehensive winning bid probability. The gang's comprehensive winning bid probability is multiplied by the preset profit value to obtain the gang's expected income reward. Based on the adversarial mutation graph, the connection edge relationships after the disconnection mark are read, the number of connection edges retained by each bidding entity node is counted one by one, the number of connection edges retained by each bidding entity node is written into the node degree position of the corresponding bidding entity node, the node degree distribution is counted according to the frequency of occurrence of the same node degree, the Shannon information is calculated based on the proportion of occurrence of each node degree in the node degree distribution, and the Shannon information is used as the quantification result of the topological random discreteness of the graph network to obtain the topological entropy camouflage reward; Read the profit weight corresponding to the expected profit reward of the gang, read the topological weight corresponding to the topological entropy disguised reward, weight the expected profit reward of the gang according to the profit weight, weight the topological entropy disguised reward according to the topological weight, sum the weighted expected profit reward of the gang and the weighted topological entropy disguised reward to generate a comprehensive topological entropy reward feedback signal.
[0012] Preferably, the steps for obtaining the outlier detection results of the bidding behavior are as follows: The comprehensive topological entropy reward feedback signal is input into the evaluation layer of the multi-agent reinforcement learning network. The state value output value corresponding to the current environment state in the evaluation layer is read, and the state value output value corresponding to the next environment state formed after the strategy layer executes the gang-cooperative fake bidding action set is read. The difference between the comprehensive topological entropy reward feedback signal, the state value output value corresponding to the current environment state, and the state value output value corresponding to the next environment state is calculated to obtain the time difference error. Based on the positive and negative direction and error magnitude of the time difference error, the network weight parameters of the bidding action channel, time action channel, and edge relationship action channel in the strategy layer are adjusted respectively, and the network weight parameters of the state value output position in the evaluation layer are adjusted simultaneously to generate updated network weight parameters of the strategy layer and the evaluation layer. Based on the updated network weight parameters of the strategy layer and evaluation layer, the strategy layer is called to continue outputting the gang-collaborative fake bidding action set. The gang-collaborative fake bidding action set is applied to the multi-ecosystem bidding environment state graph. The bid amount feature parameters of the bidding entity nodes are updated according to the bid deviation modification amount, the bid time interval feature parameters of the bidding entity nodes are updated according to the time offset modification rate, and the connection edge relationship between the bidding entity nodes is updated according to the graph edge disconnection probability. The adversarial mutation graph generated after strategy evolution is extracted, and the adversarial mutation graph is marked as outlier fake negative samples. Then, the real and normal bidding graph positive samples are read, and the outlier fake negative samples and the real and normal bidding graph positive samples are mixed and arranged according to the sample labeling type to construct an adversarial training set. The bidding entity node features, connection edge relationships, and sample label types from the adversarial training set are input into the discriminator classifier constructed based on the graph neural network. The bidding entity node features are aggregated according to the connection edge relationships to form a classification representation of each bidding entity node. The classification representation is sent to the binary classification output position of the discriminator classifier to output the outlier camouflage category prediction value and the true normal category prediction value. The outlier camouflage category prediction value and the true normal category prediction value are compared with the sample label type respectively. The model parameters of the discriminator classifier are adjusted according to the comparison difference to form the discriminator classifier after training convergence. Obtain the real bidding business graph data to be detected, read the bidding entity node features, connection edge relationships, and bidding node identification codes from the real bidding business graph data, import the bidding entity node features and connection edge relationships into the trained and converged discriminator classifier, output the probability confidence score of each bidding entity node belonging to an abnormal group, compare the probability confidence scores item by item with the anomaly judgment limit, select the bidding entity nodes with probability confidence scores higher than the anomaly judgment limit, read the corresponding bidding node identification codes, arrange the bidding node identification codes in descending order of probability confidence scores, output the corresponding bidding node identification code sequence, and generate the bidding behavior outlier detection results.
[0013] The present invention also provides a system comprising: Feature modeling module: Generates feature vectors of bidding entities based on business registration features, network address features, bid amount features, and bidding time interval features, and constructs a joint state matrix of bidding nodes; establishes connection edges based on the feature distance of the feature vectors of bidding entities in the joint state matrix of bidding nodes and preset association limits, and generates a multi-ecological bidding environment state map; Collaborative camouflage action generation module: Extract hidden layer representation vectors based on the multi-ecosystem bidding environment state graph to generate multi-agent communication variables; input the multi-agent communication variables into the policy layer of the multi-agent reinforcement learning network, and output camouflage action indicators including bid deviation modification amount, time offset modification rate and graph edge disconnection probability; combine the camouflage action indicators to generate a set of gang collaborative camouflage bidding actions. The adversarial reward generation module applies the set of gang-coordinated fake bidding actions to the multi-ecosystem bidding environment state graph to generate an adversarial mutation graph; it obtains the gang's expected profit reward based on the gang's comprehensive winning probability and preset profit value in the adversarial mutation graph; it obtains the topological entropy fake reward based on the Shannon information of the node degree distribution in the adversarial mutation graph; and it generates a comprehensive topological entropy reward feedback signal based on the gang's expected profit reward and the topological entropy fake reward. Outlier Detection Output Module: The integrated topological entropy reward feedback signal is input into the evaluation layer of the multi-agent reinforcement learning network, and the network weight parameters of the policy layer and the evaluation layer are updated according to the temporal difference error; the adversarial anomaly graph is extracted as an outlier masquerading negative sample, and an adversarial training set is constructed with the real normal bidding graph positive sample to update the model parameters of the discriminator classifier; the real bidding business graph data is imported into the discriminator classifier to output the probability confidence level, and when the probability confidence level is higher than the anomaly judgment limit, the bidding node identification code sequence is output to generate the outlier detection result of the bidding behavior.
[0014] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In this invention, a feature vector of bidding entities is generated based on business registration characteristics, network address characteristics, bid amount characteristics, and bidding time interval characteristics. Furthermore, a joint state matrix of bidding nodes is constructed, so that bidding behavior is no longer limited to a single bid deviation or a single entity record judgment, but rather integrates entity identity association, network access association, bid behavior association, and time behavior association into a unified state expression, enhancing the computability of implicit relationships between bidding entities. Connection edges are established based on the feature distance between the feature vectors of bidding entities and preset association limits, generating a multi-ecosystem bidding environment state graph. This transforms scattered bidding records into a graph state with node relationships, edge relationships, and group structures, enabling the identification of multi-dimensional collaborative behaviors such as similar business registration, similar networks, similar bids, and synchronized time. The association transmission between bidding entities is expressed through hidden layer representation vectors and multi-agent communication variables, allowing the behavioral characteristics of a single bidding entity to be combined with the association information of neighboring bidding entities. This approach enhances the ability to characterize collaborative, dispersed, and cross-entity behaviors of bidding groups. By combining bid deviation modification amount, time offset modification rate, and graph edge disconnection probability, a set of collaborative and disguised bidding actions is formed, enabling the proactive construction and training of potential bid adjustments, time misalignments, and relationship severance behaviors by anomalous groups. Through feedback signals of adversarial anomaly graphs, expected group revenue rewards, topological entropy disguise rewards, and comprehensive topological entropy rewards, the detection training process focuses not only on winning bid revenue drivers but also on the random discrete changes in graph topology, improving the ability to identify highly concealed bid rigging, weakly correlated collusion, and deliberately disrupted relationship structures. An adversarial training set is constructed using outlier disguised negative samples and genuine, normal bidding graph positive samples, and real bidding business graph data is imported into the discriminator classifier to output probability confidence levels. This allows the final result to locate risk entities in the form of bidding node identification code sequences, improving the accuracy of anomalous group identification. Attached Figure Description
[0015] Figure 1 This is a feature distance distribution map corresponding to the preset association limit. Figure 2 This is a graph showing the reward feedback signal for the comprehensive topological entropy. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0017] Please see Figure 1-2 This invention provides a technical solution: a method for detecting outliers in bidding behavior that integrates multi-ecosystem modeling, comprising the following steps: Based on business registration characteristics, network address characteristics, bid amount characteristics, and bid time interval characteristics, a feature vector of the bidding entity is generated, and a joint state matrix of bidding nodes is constructed. Based on the feature distance of the feature vector of the bidding entity in the joint state matrix of bidding nodes and the preset association limit, a connection edge is established to generate a multi-ecological bidding environment state map. Based on the state map of the multi-ecological bidding environment, the hidden layer representation vector is extracted to generate multi-agent communication variables. The multi-agent communication variables are input into the policy layer of the multi-agent reinforcement learning network, and the output includes the spoofing action index containing the bid deviation modification amount, the time offset modification rate and the probability of graph edge disconnection. The spoofing action index is combined to generate a set of gang-coordinated spoofing bidding actions. The gang's collaborative fake bidding action set is applied to the state map of the multi-ecosystem bidding environment to generate an adversarial mutation map; the gang's expected profit reward is obtained based on the gang's comprehensive winning probability and preset profit value in the adversarial mutation map; the topological entropy fake reward is obtained based on the Shannon information of the node degree distribution in the adversarial mutation map; and a comprehensive topological entropy reward feedback signal is generated based on the gang's expected profit reward and the topological entropy fake reward. The comprehensive topological entropy reward feedback signal is input into the evaluation layer of the multi-agent reinforcement learning network, and the network weight parameters of the policy layer and the evaluation layer are updated according to the temporal difference error. Adversarial anomaly graphs are extracted as outlier camouflage negative samples, and adversarial training sets are constructed with real and normal bidding graph positive samples to update the model parameters of the discriminator classifier. Real bidding business graph data are imported into the discriminator classifier to output probability confidence. When the probability confidence is higher than the anomaly judgment limit, the bidding node identification code sequence is output to generate the outlier detection result of bidding behavior.
[0018] The steps to obtain the joint state matrix of the bidding nodes are as follows: The system obtains the business registration characteristics, network address characteristics, bid amount characteristics, and bid time interval characteristics of each bidding entity. It extracts the registered capital value, establishment time value, business scope code, and legal representative identifier from the business registration characteristics; it extracts the network address range, login device identifier, and access region code from the network address characteristics; it extracts the bid amount value and bid deviation ratio from the bid amount characteristics; and it extracts the file download interval and bid submission interval from the bid time interval characteristics. The system aligns each item according to the bidding entity identifier code and concatenates them in a fixed field order to generate a bidding entity feature vector. A row index is created based on the sorting position of the bid entity identifier code, and a column index is created based on the concatenation position of the business registration feature, network address feature, bid amount feature, and bid time interval feature. Each bid entity feature vector is written into the corresponding row position. Missing fields are filled with the median of the existing values in the same field, and non-numeric fields are replaced with a unified coded value to generate a joint state matrix of bidding nodes.
[0019] Specifically, multi-dimensional feature data of each bidding entity is obtained. Business registration features are collected from publicly available data sources such as the National Enterprise Credit Information Publicity System; network address features are extracted from the backend access logs of the bidding system; and bid amount and bid time interval features are directly read from the project database of the electronic bidding platform. For business registration features, the specific value of registered capital is extracted, and the Unix timestamp of the establishment date is used as the value. The business scope text is converted into a multi-dimensional encoded vector according to national industry classification standards, and the legal representative's name or unique identifier is hashed to obtain a fixed-length identifier. For network address features, the first 24 bits of the IPv4 address are extracted as the network address segment, the device fingerprint information of the login client is recorded as the device identifier, and the access IP is mapped to the administrative division code according to the IP address location database. For bid amount features, the bid amount is directly recorded... The system calculates the price amount and the percentage deviation between the quoted price and the tender control price or average price. For the characteristics of the bidding time interval, it calculates the second interval between the bidder downloading the tender documents and the first upload of the bid, and the second interval between the last modification of the bid and the bid submission deadline. Then, it creates a data record for each bidding entity, using its unique bid entity identifier as an index. The extracted features, such as registered capital, establishment date, business scope code, legal representative identifier, network address range, login device identifier, access region code, quoted price, price deviation percentage, document download interval, and bid submission interval, are arranged in a predefined fixed order, such as four categories: business registration, network, price, and time. These features are then horizontally concatenated into a long vector to form the original feature description of each bidding entity. Finally, the long vectors of all bidding entities are collected to generate the bidding entity feature vector.
[0020] A row index is created based on the sorting position of the bidder's identifier. Specifically, all bidder identifiers are sorted in ascending lexicographical order, and each identifier is assigned a consecutive integer starting from 0 as its row number in the matrix. Simultaneously, a column index is created for each feature dimension according to the fixed field order of the feature concatenation in the previous step. For example, if a feature vector is composed of 12 feature fields, the column index ranges from 0 to 11. This constructs a two-dimensional matrix structure with clearly defined row and column indexes. Then, each bidder's feature vector is treated as a whole and written into the matrix at the corresponding row position determined by its identifier's sorting position. During the filling process, data preprocessing is performed to address missing fields in the matrix, such as missing file downloads for some bidders. For intervals, the entire column corresponding to the feature is traversed, the median of all existing values in the column is calculated, and the median is used to fill in all missing positions. This avoids the interference of extreme values on the overall data distribution. For non-numerical fields in the matrix, such as legal representative identification or business scope code, target encoding is used for replacement. That is, the mean of the target variable (e.g., whether the bid is successful) under the category is used to replace the original category identifier. If the target variable is not available, label encoding is used to map each unique non-numerical string to an independent integer value. For example, "X Hai" and "X Zhou" are encoded as 1, 2, and 3 respectively. After completing all missing value filling and non-numerical encoding replacement, a purely numerical matrix is obtained, and the joint state matrix of bidding nodes is generated.
[0021] The steps for obtaining the multi-ecological bidding environment status map are as follows: Based on the joint state matrix of bidding nodes, the feature vectors of any two bidding entities are called row by row. The differences in business registration features, network address features, bid amount features, and bidding time interval features are compared respectively. The feature distance between the feature vectors of the corresponding two bidding entities is obtained. The two bidding entity nodes whose feature distance is lower than the preset association limit are connected, and the two bidding entity nodes whose feature distance is not lower than the preset association limit are kept disconnected, thus generating a multi-ecological bidding environment state map.
[0022] Specifically, based on the joint state matrix of bidding nodes, a pairwise comparison loop is initiated. Each row in the matrix represents a bidding entity, and it is paired with all subsequent rows for comparison. For any pair of bidding entity feature vectors, the feature distance is calculated dimensionally. For the business registration feature, the absolute values of the difference in normalized registered capital and the absolute values of the difference in establishment time are calculated, and the difference in business scope code and legal representative identifier is calculated using Hamming distance. For the network address feature, the difference in network address range, login device identifier, and access region code set is calculated using the reciprocal of the Jaccard similarity coefficient. For the numerical features in the bid amount feature and the bid time interval feature, the absolute values of the differences after normalization are also calculated. Then, the feature distances calculated in each dimension are weighted and summed to obtain the comprehensive feature distance between the two bidding entities. The weight coefficients are set based on prior knowledge. For example, the weight of the network address feature is set to 0.4, and the weight of the bid amount feature is set to 0. 3. The business registration feature is set to 0.2, and the bidding time interval feature is set to 0.1, because the similarity between network and bidding behavior is more indicative in bid rigging. Then, a preset association limit is set. The method for determining this limit is to collect historically confirmed bid rigging case data and normal bidding data, calculate the feature distance between pairs of bidders, and form a "bid rigging distance distribution" and a "normal distance distribution". Select the distance value that can maximize the distinction between these two distributions as the limit. For example, select the mean of the 95th quartile of the normal distance distribution and the 5th quartile of the bid rigging distance distribution. For example, if the calculated limit is 0.65, then an undirected connection edge is established between two bidding entity nodes with a comprehensive feature distance of less than 0.65, indicating that there is a potential collaborative association between them. For node pairs with a feature distance of not less than 0.65, no connection is established, and their independent disconnected state in the graph is maintained. After traversing all node pairs to complete the edge establishment, a multi-ecosystem bidding environment state graph is generated.
[0023] The steps for obtaining variables in multi-agent communication are as follows: Based on the multi-ecological bidding environment status map, the bidding entity node identifier, connection edge relationship and node feature vector are input into the graph neural network. The first-order neighbor bidding entity nodes of each bidding entity node are read one by one according to the connection edge relationship. The business registration feature component, network address feature component, bid amount feature component and bidding time interval feature component in the first-order neighbor bidding entity nodes are extracted. According to the node adjacency position corresponding to the connection edge relationship, the various feature components of the first-order neighbor bidding entity nodes are merged into the feature update position of the current bidding entity node. The hidden layer representation vector of each bidding entity node is extracted. The hidden representation vector of each bidding entity node is read according to the node identifier code of the bidding entity. The hidden representation vector is sent to the fully connected linear mapping layer. The hidden components of business registration, network address, bid amount and bidding time interval in the hidden representation vector are mapped one dimension at a time. According to the connection edge relationship in the multi-ecological bidding environment state graph, the one-dimensional mapping result is transmitted between the bidding entity nodes with connection edges. The one-dimensional mapping result is written into the shared information position of the corresponding bidding entity node to generate multi-agent communication variables.
[0024] Specifically, based on the multi-ecosystem bidding environment state map, the adjacency matrix (representing connection edge relationships), node feature matrix (i.e., the aforementioned joint state matrix of bidding nodes), and the bidding entity node identifier code of each node are provided as input data to a pre-built graph neural network (GNN). This network adopts a graph convolutional network (GCN) architecture, containing two graph convolutional layers and a ReLU activation function. The network processing flow is as follows: First, for each bidding entity node in the graph, the first layer of the GCN locates all directly connected first-order neighbor bidding entity nodes based on the connection edge relationships. Then, the network automatically reads all feature components in the feature vectors of these neighbor nodes, including business registration feature components, network address feature components, bid amount feature components, and bidding time interval feature components, and performs an aggregation operation. Specifically, the feature vectors of all first-order neighbor nodes are averaged element-wise to obtain a... The aggregated neighbor feature vectors are then concatenated with the original feature vector of the current node. This concatenation is passed through a linear transformation layer with learnable weights, and finally, the ReLU nonlinear activation function is applied to update the first layer of features. The updated feature vectors are then fed into the second layer of the GCN, and the process of neighbor feature aggregation, concatenation, linear transformation, and activation is repeated. Through information propagation in the two layers, the feature update position of each node incorporates information from its second-order neighborhood, so that the updated features not only contain its own attributes but also incorporate the attributes of its neighbors. After completing two rounds of updates for all nodes, the vector output by the second-layer graph convolutional layer is regarded as the final hidden layer representation of the node. This representation is a high-dimensional abstract representation of the node in a multi-ecological environment. The output vectors of all bidding entity nodes are extracted to form a set of hidden layer representation vectors, and the hidden layer representation vectors of each bidding entity node are extracted.
[0025] According to the identifier code of the bidding entity node, the hidden layer representation vector of each bidding entity node generated by the graph neural network in the previous step is read one by one. These vectors are then input into a fully connected linear mapping layer for processing. The role of this fully connected layer is to perform dimensional transformation and feature decoupling, mapping the high-dimensional hidden layer representation vector to a semantic space of a predetermined dimension that is more suitable for multi-agent communication. In specific implementation, the fully connected layer contains multiple parallel linear transformation channels, which correspond to the hidden layer components of the hidden layer representation vector derived from business registration, network address, bid amount, and bidding time interval, respectively. For example, if the hidden layer representation vector is 128-dimensional, the first 32 dimensions correspond to the hidden layer component of business registration, the next 32 dimensions correspond to the hidden layer component of network address, and so on. Then, four independent linear mapping matrices are set to perform dimension-by-dimensional mapping on these four 32-dimensional components, transforming them into communication of the specified dimension. The encoding and mapping process does not change the relative order of the components within the vector; it only performs linear combinations and transformations of the values. After mapping, the communication information of each node is obtained. Then, based on the established connection edges in the multi-ecological bidding environment state graph, an information transmission mechanism is initiated between all connected bidding entity node pairs. Specifically, if there is a connection edge between node A and node B, node A copies its dimensionally mapped result and passes it to node B. At the same time, node B also passes its mapping result to node A. Each node writes the mapping results received from all its connected neighboring nodes, along with its own mapping result, into a preset shared information location. This location can be a list of vectors or a concatenated large vector. The content of this shared information location constitutes all the local environmental information that the current node can perceive at the next decision moment, generating multi-agent communication variables.
[0026] The steps to obtain the set of actions used by the gang to fabricate bidding are as follows: The multi-agent communication variables corresponding to each bidding entity node are input into the policy layer of the multi-agent reinforcement learning network. The corresponding components of the multi-agent communication variables are read according to the bidding action channel, time action channel, and edge relationship action channel in the policy layer. The output results of the bidding action channel are constrained by interval to obtain the bidding deviation modification amount. The output results of the time action channel are constrained by proportion to obtain the time offset modification rate. The output results of the edge relationship action channel are constrained by probability to obtain the graph edge disconnection probability. The bidding deviation modification amount, time offset modification rate, and graph edge disconnection probability are bound and combined according to the bidding entity node identification code to generate a set of gang-colluded fake bidding actions.
[0027] Specifically, the multi-agent communication variables corresponding to each bidding entity node are used as state inputs and fed into the policy layer of a pre-built multi-agent reinforcement learning network. This policy layer is constructed using a deep neural network, and its internal structure is designed with a shared encoder and multiple independent action output channels. The shared encoder first performs deep feature extraction on the input multi-agent communication variables, and then feeds the extracted features into three different fully connected network branches: the bidding action channel, the time action channel, and the edge relationship action channel. Each channel is independently responsible for generating a disguised action. In the bidding action channel, the network outputs a... The original value within the range is multiplied by a preset upper limit for price modification (e.g., 5% of the tender control price) and a base offset is added to constrain the range, resulting in the final price deviation modification amount. This value represents the specific amount added or subtracted from the original price. In the time action channel, the network outputs a value within the range... The original values within the range are linearly mapped to, for example, by applying the Tanh activation function. Within the specified range, a proportional constraint is applied to obtain the time offset modification rate. This value represents the original bidding time interval multiplied by this ratio to achieve earlier or later timing. In the edge relationship action channel, the network outputs an original logical value for each edge connecting the current node, which is then converted to a Sigmoid activation function. The probability values between them are used to impose probability constraints, which is the graph edge disconnection probability. This indicates how likely the connection edge is to be disinformed when generating adversarial examples. Finally, the bid deviation modification amount, time offset modification rate and graph edge disconnection probability of all related edges generated for each bidding entity node are bound and combined according to their bidding entity node identification code to form a structured action instruction. The action instructions of all nodes are put together to generate a set of gang-coordinated disinformation bidding actions.
[0028] The steps for obtaining the adversarial mutation map are as follows: The algorithm reads the price deviation modification amount, time offset modification rate, and graph edge disconnection probability of each bidding entity node in the group's collaborative fake bidding action set. It writes the price deviation modification amount into the price amount feature parameter of the corresponding bidding entity node in the multi-ecosystem bidding environment state graph, and writes the time offset modification rate into the bidding time interval feature parameter of the corresponding bidding entity node in the multi-ecosystem bidding environment state graph. It performs retention marking or disconnection marking on the corresponding connection edge according to the graph edge disconnection probability. It updates the environment state based on the modified node feature parameters and the connection edge relationship after disconnection marking, and generates an adversarial mutation graph.
[0029] Specifically, the code reads the specific action instructions generated for each bidding entity node in the group's collaborative fake bidding action set, including the price deviation modification amount, time offset modification rate, and graph edge disconnection probability. These instructions are then applied to the original multi-ecosystem bidding environment state graph to simulate fake behavior. First, for price modification, it iterates through each affected bidding entity node, locating its price amount feature parameter in the graph node features. The original price amount is summed with the price deviation modification amount corresponding to that node to obtain a new price amount. This new amount is used to update the feature parameters of the corresponding node in the graph. Simultaneously, the derived feature of the price deviation ratio is recalculated and updated. Second, for time modification, it iterates through each node, locating its bidding time interval feature parameter. The original file download interval and price submission interval are multiplied by the time offset modification rate corresponding to that node to obtain new time interval values. These new values replace the original parameters. Then, it processes the fake graph structure by iterating through each connecting edge in the graph and reading the graph edge disconnection probability corresponding to that edge in the action set. For example, if the probability is 0.8, a new value is generated. A random number is generated between the nodes. If the random number is less than 0.8, the edge is considered to be disconnected and marked as "disconnected". Otherwise, it is marked as "retained". This operation simulates a group of trolls trying to sever their obvious data connections by adjusting their behavior. After modifying the feature parameters of all nodes and updating the markings of all connecting edges, the state of the entire graph changes, forming a new graph that includes camouflage behavior, generating an adversarial mutation graph.
[0030] The steps for obtaining the comprehensive topological entropy reward feedback signal are as follows: Based on the adversarial mutation graph, the identification codes of the bidding entity nodes corresponding to the gang's collaborative fake bidding action set are read. The characteristic parameters of the bid amount, the characteristic parameters of the bidding time interval, and the disconnection mark of the connection edge corresponding to the identification code of the bidding entity node are extracted. The winning bid tendency value of each bidding entity node in the adversarial mutation graph is calculated. The winning bid tendency values of each bidding entity node in the same gang are normalized and summarized to obtain the gang's comprehensive winning bid probability. The gang's comprehensive winning bid probability is multiplied by the preset profit value to obtain the gang's expected profit reward. Based on the adversarial variation graph, the connection edge relationships after the disconnection mark are read, and the number of retained connection edges for each bidding entity node is counted one by one. The number of retained connection edges for each bidding entity node is written into the node degree position of the corresponding bidding entity node. The node degree distribution is counted according to the frequency of occurrence of the same node degree. Shannon information is calculated based on the proportion of occurrence of each node degree in the node degree distribution. Shannon information is used as the quantitative result of the topological random discreteness of the graph network to obtain the topological entropy camouflage reward. Read the profit weight corresponding to the expected profit reward of the gang, read the topological weight corresponding to the topological entropy disguised reward, weight the expected profit reward of the gang according to the profit weight, weight the topological entropy disguised reward according to the topological weight, sum the weighted expected profit reward of the gang and the weighted topological entropy disguised reward, and generate a comprehensive topological entropy reward feedback signal.
[0031] Specifically, based on the adversarial mutation graph, the identification codes of the bidding entity nodes corresponding to the group's coordinated fake bidding actions are first read to identify the members of the "gang" participating in this fake bidding operation. Then, for each member node within this gang, the modified bid amount feature parameters in the adversarial mutation graph, as well as other features that may affect the probability of winning the bid, such as the bidding time interval, are extracted. Next, a bid-winning tendency value is calculated for each bidding entity node. This value is calculated based on a preset scoring model. For example, this model can be set such that the closer the bid is to a certain optimal discount rate (e.g., 95%) of the tender control price, the higher the score; bids that are too high or too low will result in a lower score. Additionally, if the bidding time is in the middle of the submission deadline rather than near the deadline or submitted too early, bonus points will be awarded. The scores from these factors are weighted and summed to obtain the bidding propensity value for each node. Then, the bidding propensity values of all member nodes within the same group are normalized, for example, by using the Softmax function, so that the sum of the bidding propensity values of all members within the group is 1. This normalized value is the relative probability of winning the bid for each member within the group. These probability values are summed to obtain the overall probability of winning the bid for the group as a whole. Finally, this overall probability of winning the bid for the group is multiplied by a preset profit value for the bidding project. This preset profit value is usually determined by factors such as the project budget and the industry average profit rate. For example, for a project with a budget of 10 million, the preset profit value is 500,000. If the overall probability of winning the bid for the group is 0.7, the expected income is 350,000, and the expected income reward for the group is obtained.
[0032] Based on the adversarial mutation graph, the focus is first on the changes in its topology. All edges marked as "retained" after camouflage are read, while edges marked as "disconnected" are ignored. Based on this modified connectivity, the degree of each bidding entity node in the graph is counted, i.e., the number of currently retained edges for each node is calculated. This number is written to a temporary attribute of the bidding entity node, namely the node's degree position. After completing the degree count for all nodes, the entire graph's nodes are analyzed. The degree distribution is statistically analyzed, specifically counting the number of nodes with a degree of 0, 1, 2, and so on, until all degree values are covered. The frequency or proportion of each degree value is calculated. For example, if there are 100 nodes and 30 of them have a degree of 2, the proportion of degree 2 is 0.3. Then, based on Shannon's entropy principle in information theory, the entropy value of the degree distribution of each node is calculated. This calculation involves substituting the proportion of each node's degree into the Shannon entropy formula. ,in, It is the total number of different degree values. Representing the A degree value, The degree value is The proportion of nodes in the entire graph, representing all degrees. The summation of terms and the negative result of the Shannon information is used as a quantitative measure of the random discreteness of the current graph network topology. The larger the value, the more uniform and irregular the distribution of node degrees, the better the camouflage effect, and the topological entropy camouflage reward is obtained.
[0033] From the configuration parameters of the multi-agent reinforcement learning environment, we retrieve the reward weight set for the expected profit reward of the group, and the topological weight set for the topological entropy camouflage reward. These two weight coefficients are pre-set hyperparameters used to balance the agent's strategic inclination between pursuing economic gains and behavioral concealment. Their settings depend on the specific needs of the regulatory scenario. For example, in a scenario with strict supervision and an emphasis on clue discovery, the topological weight can be set higher, such as increasing the reward weight. Set to 0.4, topological weight Setting it to 0.6 encourages agents to generate more deceptive camouflage patterns; conversely, in scenarios simulating an attacker maximizing their gains, the gain weight can be adjusted. Set to 0.8, topology weight Set it to 0.2, then perform weighted processing, multiplying the expected group reward value calculated in the previous step by its corresponding reward weight. The weighted reward portion is obtained, and the calculated topological entropy masquerading reward is multiplied by its corresponding topological weight. After obtaining the weighted disguised reward, it is important to note that before weighting, the two rewards need to be normalized, for example, by using moving average and standard deviation to scale them to a similar numerical range, to avoid one weight completely dominating the final result due to different units. Finally, the weighted expected group profit reward and the weighted topological entropy disguised reward are summed to obtain a single, comprehensive scalar reward value, generating a comprehensive topological entropy reward feedback signal.
[0034] The steps for obtaining outlier detection results in bidding behavior are as follows: The comprehensive topological entropy reward feedback signal is input into the evaluation layer of the multi-agent reinforcement learning network. The state value output value corresponding to the current environment state in the evaluation layer is read, and the state value output value corresponding to the next environment state formed after the strategy layer executes the gang-cooperative fake bidding action set is read. The difference between the comprehensive topological entropy reward feedback signal, the state value output value corresponding to the current environment state, and the state value output value corresponding to the next environment state is calculated to obtain the time difference error. Based on the positive and negative direction and error magnitude of the time difference error, the network weight parameters of the bidding action channel, time action channel, and edge relationship action channel in the strategy layer are adjusted respectively, and the network weight parameters of the state value output position in the evaluation layer are adjusted simultaneously to generate the updated network weight parameters of the strategy layer and the evaluation layer. Based on the updated network weight parameters of the strategy layer and evaluation layer, the strategy layer is called to continue outputting the gang-coordinated fake bidding action set. The gang-coordinated fake bidding action set is applied to the state graph of the multi-ecosystem bidding environment. The bid amount feature parameters of the bidding entity node are updated according to the bid deviation modification amount, the bidding time interval feature parameters of the bidding entity node are updated according to the time offset modification rate, and the connection edge relationship between the bidding entity nodes is updated according to the graph edge disconnection probability. The adversarial mutation graph generated after strategy evolution is extracted, and the adversarial mutation graph is marked as outlier fake negative samples. Then, the positive samples of the real and normal bidding graph are read. The outlier fake negative samples and the real and normal bidding graph positive samples are mixed and arranged according to the sample labeling type to construct the adversarial training set. The features of bidding entity nodes, connection relationships, and sample label types from the adversarial training set are input into the discriminator classifier constructed based on the graph neural network. The features of bidding entity nodes are aggregated according to the connection relationships to form a classification representation of each bidding entity node. The classification representation is sent to the binary classification output position of the discriminator classifier, which outputs the predicted values of outlier disguised category and true normal category. The predicted values of outlier disguised category and true normal category are compared with the sample label type respectively. The model parameters of the discriminator classifier are adjusted according to the comparison difference to form the discriminator classifier after training convergence. The system acquires real bidding business graph data to be detected, reads the features of bidding entity nodes, connection edges, and bidding node identifiers from the real bidding business graph data, imports the features of bidding entity nodes and connection edges into the trained and converged discriminator classifier, outputs the probability confidence score of each bidding entity node belonging to an abnormal group, compares the probability confidence scores item by item with the anomaly judgment limit, selects bidding entity nodes with probability confidence scores higher than the anomaly judgment limit, reads the corresponding bidding node identifiers, arranges the bidding node identifiers in descending order of probability confidence scores, outputs the corresponding bidding node identifier sequence, and generates outlier detection results for bidding behavior.
[0035] Specifically, the comprehensive topological entropy reward feedback signal is input into the evaluation layer (Critic) of the multi-agent reinforcement learning network. This evaluation layer operates in parallel with the policy layer (Actor), forming the Actor-Critic framework. At the current time step, the evaluation layer first reads its state value output value for the current environmental state (i.e., the graph state before adversarial mutation) from the evaluation layer network. Simultaneously, the policy layer executes a set of cooperative spoofing bidding actions to affect the environment and obtain the next environmental state (i.e., the adversarial mutation graph). This next environmental state is then input into the evaluation layer to obtain its corresponding state value output value. Then, based on the principle of temporal difference learning, the temporal difference error is calculated, specifically the difference between the current reward, the discounted value of the next state, and the current state value. This error serves as a loss signal. Based on the direction and magnitude of the error, a gradient descent algorithm (such as the Adam optimizer) is used to update the network weight parameters of the evaluation layer, enabling it to more accurately predict state values. Simultaneously, this temporal difference error is also used to update the policy layer. The error signal is backpropagated to the three action output channels of the policy layer: the bid action channel, the time action channel, and the edge relationship action channel. If the error is positive, the probability of selecting the action that leads to this result is increased; otherwise, it is decreased. By adjusting the network weight parameters of each channel, the policy network is guided to optimize towards actions that can generate higher overall rewards. The network weight parameters of the evaluation layer and the policy layer are simultaneously adjusted and optimized during this process, generating updated network weight parameters for the policy layer and the evaluation layer.
[0036] Based on the updated network weight parameters of the strategy layer and evaluation layer, a new round of adversarial example generation is initiated. The improved strategy layer network is invoked, and the current multi-ecosystem bidding environment state graph is input, causing it to output a new and superior set of collaborative, deceptive bidding actions. This action set is then applied again to the original multi-ecosystem bidding environment state graph. According to the specified bid deviation modification amount, the bid amount and related feature parameters of the corresponding bidding entity nodes in the graph are updated. The bidding time interval feature parameter is adjusted according to the time offset modification rate. Based on the graph edge disconnection probability, random sampling is used to determine whether to update the connection edge relationships between bidding entity nodes. After all modifications are executed, the data is extracted. This newly generated adversarial variant graph, which is evolved from the latest strategy and has a higher degree of camouflage or better expected returns, is explicitly labeled as an outlier camouflaged negative sample in the dataset with a label value of 0. At the same time, a batch of real, verified, and normal bidding project graph data without collusion is read from the historical database and prepared. These graphs are labeled as real and normal bidding graph positive samples with a label value of 1. Finally, these dynamically generated outlier camouflaged negative samples and static real and normal graph positive samples are randomly mixed and arranged in a certain ratio (e.g., 1:1) to construct a balanced dataset containing positive and negative samples for training the downstream discriminator, thus constructing the adversarial training set.
[0037] The data from the adversarial training set, including the node features, edge relationships, and corresponding sample label types (0 or 1) of each graph, are input in batches into a discriminative classifier built on a graph neural network. This classifier adopts a graph isomorphic network (GIN) structure. Due to its powerful expressive ability of graph structures, the classifier's processing flow is as follows: For each input graph sample, its GIN layer iteratively aggregates and updates the features of each bidding entity node according to the edge relationships. Specifically, each GIN layer merges the features of the central node with the sum of the features of all its neighboring nodes, and then performs a nonlinear transformation through a multilayer perceptron (MLP) to form the representation of the node at the current layer. After multiple layers are stacked, the final representation of each node will aggregate its structural and feature information within a multi-hop range in the graph, resulting in... After classifying each bidding entity node, a graph pooling layer (e.g., global mean pooling) aggregates the representations of all nodes into a single vector representing the entire graph, i.e., the graph-level classification representation. This graph-level classification representation is then fed into the final binary classification output position of the discriminator classifier. This position is a fully connected layer with a sigmoid activation function, which outputs a probability value between 0 and 1, representing the probability that the graph is predicted as an outlier / masked category, i.e., the outlier / masked category prediction value. Subtracting this value from 1 gives the true normal category prediction value. The prediction value is compared with the true label type of the sample, and the binary cross-entropy loss is calculated. This loss is then used to adjust the model parameters of each layer of GIN and the final classification layer through the backpropagation algorithm. This process is iterated repeatedly until the loss function converges on the validation set, forming the discriminator classifier after training convergence.
[0038] Obtain one or more sets of real bidding business graph data to be tested. This data is unprocessed and comes from the latest actual bidding projects. First, following the same method as during training, read and construct the graph structure from this data, including extracting node features for each bidding entity, establishing connections between nodes, and recording the unique identifier for each bidding node. Then, the processed bidding entity node features and connection relationships are treated as a complete graph data and imported into the previously trained and converged discriminator classifier for inference. The discriminator classifier performs a forward propagation calculation on the entire input graph and outputs an overall judgment. Simultaneously, before the graph pooling layer of the classifier, the final classification representation of each independent bidding entity node can be extracted. These node-level representations are then passed through a shared classification head, thus classifying the graph... Each bidding entity node outputs an independent probability confidence score, representing the likelihood of this single node participating in abnormal group behavior. An anomaly detection threshold is then set, referencing the precision-recall curve of the classifier on the validation set, selecting a balance point or determining it based on business needs, for example, setting it to 0.8. Then, the probability confidence score output by each bidding entity node is compared one by one with this threshold of 0.8. All bidding entity nodes with a probability confidence score higher than 0.8 are initially identified as suspected anomaly nodes. These nodes are selected, and their corresponding bidding node identifiers are read. Finally, to facilitate subsequent manual review and analysis, these selected bidding node identifiers are sorted in descending order of their probability confidence scores, outputting this ordered identifier list, generating the outlier detection results for bidding behavior.
[0039] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A method for outlier detection in bidding behavior that integrates multi-ecosystem modeling, characterized in that, Includes the following steps: Based on business registration characteristics, network address characteristics, bid amount characteristics, and bid time interval characteristics, a feature vector of the bidding entity is generated, and a joint state matrix of bidding nodes is constructed. Based on the feature distance of the feature vector of the bidding entity in the joint state matrix of the bidding nodes and the preset association limit, a connection edge is established to generate a multi-ecological bidding environment state map; Based on the multi-ecological bidding environment state map, hidden layer representation vectors are extracted to generate multi-agent communication variables; The multi-agent communication variables are input into the policy layer of the multi-agent reinforcement learning network, and the output includes spoofing action indicators including bid deviation modification amount, time offset modification rate and graph edge disconnection probability. The spoofing action indicators are combined to generate a set of gang-coordinated spoofing bidding actions. The gang's coordinated and disguised bidding actions are applied to the multi-ecological bidding environment state map to generate an adversarial mutation map; The expected profit reward of the gang is obtained based on the gang's comprehensive winning probability and preset profit value in the adversarial mutation graph. The topological entropy masquerading reward is obtained based on the Shannon information of the node degree distribution in the adversarial mutation graph. A comprehensive topological entropy reward feedback signal is generated based on the gang's expected profit reward and the topological entropy masquerading reward. The comprehensive topological entropy reward feedback signal is input into the evaluation layer of the multi-agent reinforcement learning network, and the network weight parameters of the policy layer and the evaluation layer are updated according to the time difference error. The adversarial anomaly map is extracted as an outlier masquerading negative sample and used to construct an adversarial training set with the real and normal bidding map positive samples to update the model parameters of the discriminator classifier; the real bidding business map data is imported into the discriminator classifier to output the probability confidence level. When the probability confidence level is higher than the anomaly judgment limit, the bidding node identification code sequence is output to generate the outlier detection result of the bidding behavior.
2. The outlier detection method for bidding behavior based on multi-ecosystem modeling as described in claim 1, characterized in that, The steps for obtaining the joint state matrix of the bidding nodes are as follows: The system obtains the business registration characteristics, network address characteristics, bid amount characteristics, and bid time interval characteristics of each bidding entity. It extracts the registered capital value, establishment time value, business scope code, and legal representative identifier from the business registration characteristics; it extracts the network address range, login device identifier, and access region code from the network address characteristics; it extracts the bid amount value and bid deviation ratio from the bid amount characteristics; and it extracts the file download interval and bid submission interval from the bid time interval characteristics. The system aligns each item according to the bidding entity identifier code and concatenates them in a fixed field order to generate a bidding entity feature vector. A row index is created based on the sorting position of the bid entity identification code, and a column index is created based on the concatenation position of the business registration feature, network address feature, bid amount feature, and bid time interval feature. Each bid entity feature vector is written into the corresponding row position. Missing fields are filled with the median of the existing values in the same field, and non-numerical fields are replaced with a unified coded value to generate a joint state matrix of bidding nodes.
3. The outlier detection method for bidding behavior based on multi-ecosystem modeling as described in claim 1, characterized in that, The steps for obtaining the multi-ecological bidding environment status map are as follows: Based on the joint state matrix of the bidding nodes, the feature vectors of any two bidding entities are called row by row, and the differences in business registration features, network address features, bid amount features, and bidding time interval features are compared respectively. The feature distance between the corresponding two feature vectors of the bidding entities is obtained by summarizing. The two bidding entity nodes whose feature distance is lower than the preset association limit are connected, and the two bidding entity nodes whose feature distance is not lower than the preset association limit are kept disconnected, thus generating a multi-ecological bidding environment state map.
4. The outlier detection method for bidding behavior based on multi-ecosystem modeling as described in claim 1, characterized in that, The steps for obtaining the multi-agent communication variables are as follows: Based on the multi-ecological bidding environment status map, the bidding entity node identifier, connection edge relationship and node feature vector are input into the graph neural network. The first-order neighbor bidding entity nodes of each bidding entity node are read one by one according to the connection edge relationship. The business registration feature component, network address feature component, bid amount feature component and bidding time interval feature component in the first-order neighbor bidding entity nodes are extracted. According to the node adjacency position corresponding to the connection edge relationship, the feature components of the first-order neighbor bidding entity nodes are merged into the feature update position of the current bidding entity node. The hidden layer representation vector of each bidding entity node is extracted. The hidden layer representation vector of each bidding entity node is read according to the node identifier code of the bidding entity. The hidden layer representation vector is sent to the fully connected linear mapping layer. The hidden layer components of business registration, network address, bid amount and bidding time interval in the hidden layer representation vector are mapped one dimension at a time. According to the connection edge relationship in the multi-ecological bidding environment state graph, the one-dimensional mapping result is transmitted between the bidding entity nodes with connection edges. The one-dimensional mapping result is written into the shared information position of the corresponding bidding entity node to generate multi-agent communication variables.
5. The outlier detection method for bidding behavior based on multi-ecosystem modeling as described in claim 1, characterized in that, The steps for obtaining the set of actions used by the gang to fabricate bidding are as follows: The multi-agent communication variables corresponding to each bidding entity node are input into the policy layer of the multi-agent reinforcement learning network. The corresponding components of the multi-agent communication variables are read according to the bidding action channel, time action channel, and edge relationship action channel in the policy layer. The output results of the bidding action channel are constrained by interval to obtain the bidding deviation modification amount. The output results of the time action channel are constrained by proportion to obtain the time offset modification rate. The output results of the edge relationship action channel are constrained by probability to obtain the graph edge disconnection probability. The bidding deviation modification amount, time offset modification rate, and graph edge disconnection probability are bound and combined according to the bidding entity node identification code to generate a set of gang-colluded fake bidding actions.
6. The outlier detection method for bidding behavior based on multi-ecosystem modeling as described in claim 1, characterized in that, The steps for obtaining the adversarial mutation map are as follows: The algorithm reads the price deviation modification amount, time offset modification rate, and graph edge disconnection probability of each bidding entity node in the group's collaborative fake bidding action set. It writes the price deviation modification amount into the price amount feature parameter of the corresponding bidding entity node in the multi-ecological bidding environment state graph, and writes the time offset modification rate into the bidding time interval feature parameter of the corresponding bidding entity node in the multi-ecological bidding environment state graph. It performs retention marking or disconnection marking on the corresponding connection edge according to the graph edge disconnection probability. It updates the environment state based on the modified node feature parameters and the connection edge relationship after the disconnection marking, and generates an adversarial mutation graph.
7. The outlier detection method for bidding behavior based on multi-ecosystem modeling as described in claim 1, characterized in that, The steps for obtaining the comprehensive topological entropy reward feedback signal are as follows: Based on the adversarial mutation graph, the identification code of the bidding entity node corresponding to the gang's collaborative fake bidding action set is read. The bidding amount feature parameter, bidding time interval feature parameter and connection edge disconnection mark corresponding to the bidding entity node identification code are extracted. The winning bid tendency value of each bidding entity node in the adversarial mutation graph is calculated. The winning bid tendency values of each bidding entity node in the same gang are normalized and summarized to obtain the gang's comprehensive winning bid probability. The gang's comprehensive winning bid probability is multiplied by the preset profit value to obtain the gang's expected income reward. Based on the adversarial mutation graph, the connection edge relationships after the disconnection mark are read, the number of connection edges retained by each bidding entity node is counted one by one, the number of connection edges retained by each bidding entity node is written into the node degree position of the corresponding bidding entity node, the node degree distribution is counted according to the frequency of occurrence of the same node degree, the Shannon information is calculated based on the proportion of occurrence of each node degree in the node degree distribution, and the Shannon information is used as the quantification result of the topological random discreteness of the graph network to obtain the topological entropy camouflage reward; Read the profit weight corresponding to the expected profit reward of the gang, read the topological weight corresponding to the topological entropy disguised reward, weight the expected profit reward of the gang according to the profit weight, weight the topological entropy disguised reward according to the topological weight, sum the weighted expected profit reward of the gang and the weighted topological entropy disguised reward to generate a comprehensive topological entropy reward feedback signal.
8. The outlier detection method for bidding behavior based on multi-ecosystem modeling as described in claim 1, characterized in that, The steps for obtaining the outlier detection results of the bidding behavior are as follows: The comprehensive topological entropy reward feedback signal is input into the evaluation layer of the multi-agent reinforcement learning network. The state value output value corresponding to the current environment state in the evaluation layer is read, and the state value output value corresponding to the next environment state formed after the strategy layer executes the gang-cooperative fake bidding action set is read. The difference between the comprehensive topological entropy reward feedback signal, the state value output value corresponding to the current environment state, and the state value output value corresponding to the next environment state is calculated to obtain the time difference error. Based on the positive and negative direction and error magnitude of the time difference error, the network weight parameters of the bidding action channel, time action channel, and edge relationship action channel in the strategy layer are adjusted respectively, and the network weight parameters of the state value output position in the evaluation layer are adjusted simultaneously to generate updated network weight parameters of the strategy layer and the evaluation layer. Based on the updated network weight parameters of the strategy layer and evaluation layer, the strategy layer is called to continue outputting the gang-collaborative fake bidding action set. The gang-collaborative fake bidding action set is applied to the multi-ecosystem bidding environment state graph. The bid amount feature parameters of the bidding entity nodes are updated according to the bid deviation modification amount, the bid time interval feature parameters of the bidding entity nodes are updated according to the time offset modification rate, and the connection edge relationship between the bidding entity nodes is updated according to the graph edge disconnection probability. The adversarial mutation graph generated after strategy evolution is extracted, and the adversarial mutation graph is marked as outlier fake negative samples. Then, the real and normal bidding graph positive samples are read, and the outlier fake negative samples and the real and normal bidding graph positive samples are mixed and arranged according to the sample labeling type to construct an adversarial training set. The bidding entity node features, connection edge relationships, and sample label types from the adversarial training set are input into the discriminator classifier constructed based on the graph neural network. The bidding entity node features are aggregated according to the connection edge relationships to form a classification representation of each bidding entity node. The classification representation is sent to the binary classification output position of the discriminator classifier to output the outlier camouflage category prediction value and the true normal category prediction value. The outlier camouflage category prediction value and the true normal category prediction value are compared with the sample label type respectively. The model parameters of the discriminator classifier are adjusted according to the comparison difference to form the discriminator classifier after training convergence. Obtain the real bidding business graph data to be detected, read the bidding entity node features, connection edge relationships, and bidding node identification codes from the real bidding business graph data, import the bidding entity node features and connection edge relationships into the trained and converged discriminator classifier, output the probability confidence score of each bidding entity node belonging to an abnormal group, compare the probability confidence scores item by item with the anomaly judgment limit, select the bidding entity nodes with probability confidence scores higher than the anomaly judgment limit, read the corresponding bidding node identification codes, arrange the bidding node identification codes in descending order of probability confidence scores, output the corresponding bidding node identification code sequence, and generate the bidding behavior outlier detection results.
9. The system for outlier detection in bidding behavior based on multi-ecosystem modeling as described in any one of claims 1-8, characterized in that, include: Feature modeling module: Generates feature vectors of bidding entities based on business registration features, network address features, bid amount features, and bidding time interval features, and constructs a joint state matrix of bidding nodes; establishes connection edges based on the feature distance of the feature vectors of bidding entities in the joint state matrix of bidding nodes and preset association limits, and generates a multi-ecological bidding environment state map; Collaborative camouflage action generation module: Extracts hidden layer representation vectors based on the multi-ecological bidding environment state map to generate multi-agent communication variables; The multi-agent communication variables are input into the policy layer of the multi-agent reinforcement learning network, and the output includes spoofing action indicators including bid deviation modification amount, time offset modification rate and graph edge disconnection probability. The spoofing action indicators are combined to generate a set of gang-coordinated spoofing bidding actions. The adversarial reward generation module applies the set of gang-coordinated fake bidding actions to the multi-ecosystem bidding environment state graph to generate an adversarial mutation graph; it obtains the gang's expected profit reward based on the gang's comprehensive winning probability and preset profit value in the adversarial mutation graph, obtains the topological entropy fake reward based on the Shannon information of the node degree distribution in the adversarial mutation graph, and generates a comprehensive topological entropy reward feedback signal based on the gang's expected profit reward and the topological entropy fake reward. Outlier detection output module: Inputs the comprehensive topological entropy reward feedback signal into the evaluation layer of the multi-agent reinforcement learning network, and updates the network weight parameters of the policy layer and the evaluation layer according to the time difference error; The adversarial anomaly map is extracted as an outlier masquerading negative sample and used to construct an adversarial training set with the real and normal bidding map positive samples to update the model parameters of the discriminator classifier; the real bidding business map data is imported into the discriminator classifier to output the probability confidence level. When the probability confidence level is higher than the anomaly judgment limit, the bidding node identification code sequence is output to generate the outlier detection result of the bidding behavior.