An intelligent identification method for collusive behavior of power generation enterprises based on Info-DAGMM model
Through the Info-DAGMM model, the conspiracy identification indicator system for power generation enterprises is constructed, combined with the deep joint network, and the problem of inefficient identification of conspiracy behavior in the power market is solved, real-time risk monitoring and intelligent early warning of the power market are realized, and fair competition and stable operation of the market are ensured.
Patent Information
- Application Number
- CN202211010411.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-23
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2042-08-23
AI Technical Summary
In the existing power market, the method of identifying conspiracy behavior of power generation companies is inefficient, the existing technology cannot adapt to market changes in real time, and the data feature database is incomplete, and the supervision learning model lacks labeled data, which leads to difficulty in identification.
Using the Info-DAGMM model, by constructing a conspiracy identification index system for power generation enterprises, combining with deep joint networks, maximizing the mutual information of the original variables and latent variables, using an automatic encoder and a Gaussian hybrid model to separate the conspiracy samples to achieve real-time monitoring.
Real-time risk monitoring of the power market is realized, and accurate and efficient intelligent early warning methods are provided to ensure fair competition and stable operation of the market.
Smart Images

Figure CN115409548B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of risk identification method design for power market entities, and specifically relates to an intelligent identification method for collusion behavior of power generation enterprises based on the Info-DAGMM model. Background Art
[0002] In the early stages of the electricity spot market, collusion among power generation companies, a common form of market misconduct, was the abuse of market power by these companies. Typically, in centralized bidding, these alliances manipulate market prices and maximize profits by changing their bid prices and bid volumes by the same or opposite proportions within the same timeframe. Compared to power markets abroad, China's power market is highly concentrated, making malicious collusion more common. Therefore, it is necessary to strengthen the top-level design of power market credit supervision and directly identify collusive behavior.
[0003] In practice, identifying collusion among power generation companies has traditionally relied primarily on indicator analysis and expert decision-making. However, with the continued liberalization of power market access and the expansion of spot trading, this approach is no longer sufficient. Therefore, to maintain fair competition in the power market, intelligent identification is needed for collusion early warning, enabling real-time computer monitoring of power generation company collusion and assisting expert decision-making. Existing pattern recognition-based methods quantitatively identify market violations by designing feature libraries and cloud models. However, these methods' feature libraries are incomplete, and the cloud models lack self-updating capabilities, making them unable to adapt to market changes in real time. Game theory-based methods construct a model of the interest relationships between market participants at the bidding decision-making level and determine violations based on whether the model reaches a Nash equilibrium at the clearing level. However, my country's power market information disclosure mechanism is imperfect, making the marginal cost data required by these methods difficult to obtain. Supervised learning algorithms are used to train classifiers, but in practice, labeled data is very scarce, making the use of supervised learning models impractical. Therefore, the use of unsupervised learning clustering models for collusion early warning is more relevant to actual trading scenarios and offers greater practicality.
[0004] In the research on deep clustering models, deep federated networks (DJNs) are capable of handling complex, high-dimensional data from the electricity market. They are generally composed of two networks: a representation network and an estimation network. The representation network learns low-dimensional representations of complex, high-dimensional samples, while the estimation network performs density estimation, identifying samples in low-density areas as outliers. This concept allows for the separation of a small number of outliers from a large number of normal samples, making this clustering model suitable for the data characteristics of collusion in the electricity market. However, dimensionality reduction of high-dimensional raw data using simple networks results in a significant loss of information. Therefore, to address the characteristics of collusion data in the electricity market, the representation network was modified to maximize the mutual information between the original and latent variables in the loss function. Preserving more information about the original variables allows the estimation network to better isolate collusive samples, resulting in the proposed new DJN network. Summary of the Invention
[0005] In response to the shortcomings of the traditional collusion behavior identification method mentioned in the above background technology in processing large-scale power transaction data inefficiency, the present invention proposes an intelligent identification method for collusion behavior of power generation enterprises based on the Info-DAGMM model, which overcomes the shortcomings of the existing technology and has good results.
[0006] The present invention adopts the following technical solutions:
[0007] An intelligent identification method for collusive behavior of power generation enterprises based on the Info-DAGMM model includes the following steps:
[0008] S1. Obtain the original declared electricity volume and electricity price of the power generation enterprise;
[0009] The original declared electricity quantity and original declared electricity price are the declared electricity quantity and declared electricity price of the hth period of the mth power generation enterprise in a certain period;
[0010] S2. Build an indicator system for identifying collusion among power generation companies. The indicators include: average market share of reported electricity, consistency of bids, consistency of reported electricity, ratio of difference area of bid curves, average safety of bids, and average relative comparison of bids;
[0011] S3. Calculate and obtain a set of collusion indicators and perform normalization processing;
[0012] S4. Use the processed collusion indicator set to train the Info-DAGMM model to obtain the density estimation value of each sample in the low-dimensional space, which is the collusion suspicion degree;
[0013] S5. Inversely map the collusion suspicion back to the measurement matrix of step S3. Through the horizontal and vertical coordinates of the samples, the power generation companies that have colluded in the market can be determined.
[0014] Furthermore, in step S2,
[0015] The average market share of the declared electricity volume is index 1, which represents the proportion of the declared electricity volume of each power generation enterprise in the market share. The calculation formula is:
[0016]
[0017] Where S m and S n are the electricity quantities reported by the mth and nth power generation companies in the market for this bidding; k is the number of power generation companies participating in this bidding;
[0018] The quotation consistency is indicator 2, which represents the difference in the declared prices of two power generation companies. The calculation formula is:
[0019]
[0020] Where p mh and p nh are the bids of the mth and nth power generation companies in the hth segment of this bidding; p h is the average of the bids of all power generation companies in the hth segment in this bidding;
[0021] The consistency of the reported electricity is indicator 3, which represents the difference in the reported electricity of two power generation companies. The calculation formula is:
[0022]
[0023] Where S mh and S nh The electricity quantity reported by the mth and nth power generation companies in the hth segment of this bidding respectively; is the average of the electricity volume declared by all power generation companies in the hth period in this bidding;
[0024] The bid curve difference area ratio is indicator 4, which reflects the degree of parallel bidding between two power generation companies. The calculation formula is:
[0025]
[0026] Where, f m and f n are the bidding curve functions of the mth and nth power generation companies in this bidding respectively;
[0027] The average safety of the bid is index 5, which measures the deviation between the bids of power generation companies and the historical average bid. The calculation formula is:
[0028]
[0029] Where, and are the weighted average bid prices of the mth and nth power generation companies in this bidding; E is the expected value of the market marginal price, which is calculated using the marginal prices of historical transactions;
[0030] The relative bid comparison mean is indicator 6, which represents the difference between the bids of two power generation companies and the average price of this centralized bidding. The calculation formula is:
[0031]
[0032] Furthermore, in step S3, assuming that there are m power generation companies bidding for a total of h periods in the centralized bidding, the calculation process of the collusion index dataset X is:
[0033] According to the calculation formula of the collusion index system in step S2, the qth index between each pair of power generation enterprises is calculated to form a data matrix:
[0034]
[0035] The column vectors are obtained by tiling the upper triangle elements in sequence, which is an indicator feature of the collusion indicator set:
[0036]
[0037] All 6 indicator features are calculated and normalized by column combination to obtain the collusion indicator set X for deep joint network training:
[0038] X=(x (1) ,x (2) ,...,x (6) )(9).
[0039] Furthermore, in step S4, the construction process of the Info-DAGMM model is as follows:
[0040] The network framework of the Info-DAGMM model mainly consists of two modules: the expression network module and the density estimation module. First, the Encoder in the autoencoder DA is used to train the collusion indicator set X, and the mutual information is introduced into the loss function of the expression network to maximize the original variable and the low-dimensional space latent variable Z. l The information features are then decoded by the Decoder in DA to obtain the reconstruction index set X′; then, the Euclidean distance and cosine similarity are used to measure the collusion index set X and the reconstruction X′ to obtain the reconstruction error Z r ; Finally, the latent variable Z l and reconstruction error Z r The data is passed to the Gaussian mixture model (GMM) to calculate its sample energy and separate the collusion samples. During training, the parameters of the autoencoder and the Gaussian mixture model are optimized simultaneously.
[0041] In the expression network, the autoencoder DA will learn the latent variable Z of the collusion indicator set X l The optimization goal is to maximize the information of X retained in Z l , that is, the sample reconstruction error is minimized:
[0042]
[0043] Where, L(x i ,x i ') Take the L2 norm; N is the batch size of each iteration;
[0044] To maximize the collusion indicator set X and latent variable Z l The information characteristics of the network are introduced into the loss function of the expression network. The mutual information is defined as:
[0045]
[0046] Where p(x) and p(z l ) are X and Z respectively l The probability density function of p(x,z l ) is its joint distribution; p(z l |x) is the latent variable Z l The true conditional distribution of ; KL divergence measures the degree of difference between two distributions;
[0047] When estimating mutual information, JS divergence is used for estimation, and the mutual information is estimated as:
[0048]
[0049] In the formula, x represents the original sample, is a similar sample that obeys other distributions; the T function is a discriminator modeled by a neural network, which gives positive scores to the combination of original variables and latent variables and negative scores to the combination of interference variables and latent variables; S represents the softplus function: S(z)=log(1+e z ); E is the mathematical expectation of the S function;
[0050] Considering both global and local mutual information:
[0051]
[0052] Where α1 and α2 are the coefficients of global mutual information and local mutual information, respectively. The combination of α1 and α2 can produce a higher mutual information estimate of the neural network. Among them, good reconstruction is highly dependent on α1, and good classification performance is highly dependent on α2. G is the global discriminator, T L is the local discriminator;
[0053] In summary, the loss function of mutual information is as follows:
[0054]
[0055] Where, q(z l ) is the latent variable Z l The prior distribution of
[0056] In the estimation network, GMM learns the probability density distribution F(x) of low-dimensional features, and the output of the network is the collusion suspicion between any two power generation companies;
[0057]
[0058] Where, is the output of the entire network, which is the density estimate of the potential expression Z; softmax is the kernel function; Z r Take the Euclidean distance and cosine similarity between the collusion index set X and the reconstruction index set X′; θ m To estimate the parameters of the network; h(z, θ m ) is the hidden layer output;
[0059] The loss function of the estimation network is defined as:
[0060]
[0061] Where E(z) is the sample energy.
[0062] Finally, according to Equations (10), (14), and (16), the optimization objective of the Info-DAGMM model is:
[0063]
[0064] Where λ1 and λ2 are the proportions of the loss functions of the GMM network and mutual information in the joint loss function, respectively.
[0065] Furthermore, the evaluation index of the conspiracy suspicion is defined as:
[0066] CSI=F(x i );i=1...η (18)
[0067] Where η is the sample size, x i is the indicator set of the i-th group of power generation enterprises, and F(x) is the approximate probability distribution function of the collusion indicator set X, which is estimated by the GMM in the network.
[0068] Furthermore, in step S5, the collusion suspicion CSI is inversely mapped back to the measurement matrix (9) of step S3, and the power generation enterprises that have engaged in collusion in the market can be determined through the horizontal and vertical coordinates of the samples.
[0069] The present invention has the following beneficial effects:
[0070] The intelligent identification method for collusive behavior of power generation enterprises based on the Info-DAGMM model proposed in this invention can provide an accurate and efficient intelligent early warning method for real-time monitoring of actual power market risks. It has a very positive significance for ensuring fair competition in the market, maintaining safe and stable market operation, and building a good and orderly power market. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 This is the network structure diagram of the Info-DAGMM proposed in the present invention;
[0072] Figure 2 This is a schematic diagram of the global mutual information discriminator proposed in the present invention;
[0073] Figure 3 This is the principle diagram of the local mutual information discriminator proposed by the present invention; DETAILED DESCRIPTION
[0074] The specific implementation of the present invention will be further described below with reference to the accompanying drawings and specific embodiments:
[0075] An intelligent identification method for collusive behavior of power generation enterprises based on the Info-DAGMM model includes the following steps:
[0076] S1. Calculations were performed using monthly centralized bidding data from Gansu Province to Shandong Province. This data includes the declared electricity volume and prices of 528 renewable energy (wind and solar) power plants of varying installed capacities, covering three bidding sessions in the first three months of 2018.
[0077] S2. Based on the characteristics and features of previous collusion behavior indicators and the actual situation of the power market, a collusion identification indicator system for power generation companies is constructed. The indicators include: average market share of reported electricity, consistency of bids, consistency of reported electricity, ratio of bid curve difference area, average bid safety, and average bid relative comparison:
[0078] Specifically, the average market share of declared electricity is index 1, which represents the proportion of the declared electricity of each power generation enterprise in the market share. The calculation formula is:
[0079]
[0080] Where S m and S nare the electricity quantities reported by the mth and nth power generation companies in the market for this bidding; k is the number of power generation companies participating in this bidding;
[0081] The consistency of the bid is indicator 2, which represents the difference in the bid prices of two power generation companies. The calculation formula is:
[0082]
[0083] Where, P mh and P nh The bids of the mth and nth power generation companies in the hth segment of this bidding are respectively; is the average of the bids of all power generation companies in the hth segment in this bidding;
[0084] The consistency of reported electricity is indicator 3, which represents the difference in reported electricity between two power generation companies. The calculation formula is:
[0085]
[0086] Where S mh and S nh S is the electricity quantity declared by the mth and nth power generation enterprises in the hth period of this bidding respectively; h is the average of the electricity volume declared by all power generation companies in the hth period in this bidding;
[0087] The bid curve difference area ratio is indicator 4, which reflects the degree of parallel bidding between two power generation companies. The calculation formula is:
[0088]
[0089] Where, f m and f n are the bidding curve functions of the mth and nth power generation companies in this bidding respectively;
[0090] The average bid safety index is 5, which measures the deviation between the bids of power generation companies and the historical average bid. The calculation formula is:
[0091]
[0092] Where, and are the weighted average bid prices of the mth and nth power generation companies in this bidding respectively; E is the expected value of the market marginal price, which is calculated using the marginal prices of historical transactions;
[0093] The relative bid comparison mean is indicator 6, which represents the difference between the bids of two power generation companies and the average price of this centralized bidding. The calculation formula is:
[0094]
[0095] S3. After the index calculation in step S2, the collusion index set X for network training is obtained. The calculation process of the collusion index data set X is as follows:
[0096] First, the qth index data of each pair of power generation enterprises is calculated according to the calculation formula of the collusion index system in step S2, where q=6:
[0097]
[0098] The column vectors are obtained by tiling the upper triangle elements in sequence, which is an indicator feature of the collusion indicator set:
[0099]
[0100] Calculate all 6 indicator features and normalize them by column combination to obtain the collusion indicator set X for deep joint network training:
[0101] X=(x (1) ,x (2) ,...,x (6) ) (9)
[0102] The dataset contains 212,959 samples and six collusion indicators, but there are only 274 collusion samples, indicating an imbalance between positive and negative samples. The labels are marked with y, where 0 represents a normal sample and 1 represents a collusion sample. Table 1 shows 10 of these samples:
[0103] Table 1 Collusion Identification Dataset
[0104]
[0105] The dataset is divided into training, validation, and test sets in a ratio of 7:1.5:1.5, containing 170,367, 21,296, and 21,296 samples respectively;
[0106] S4. Use the training set and validation set to train and validate the Info-DAGMM model, and then use the Info-DAGMM model to test the test set to obtain the density estimate of each sample in the low-dimensional space, which is the conspiracy suspicion degree;
[0107] Specifically, if Figure 1 As shown in Figure 2, the network framework of the Info-DAGMM model consists of an expression network and a density estimation module. The expression network uses the Encoder in the autoencoder DA to train the original variable X and obtain the latent variable Z in the low-dimensional space. l, and then decode it with the Decoder in DA to obtain the reconstruction index set X′; then, use the Euclidean distance and cosine similarity to measure the collusion index set X and the reconstruction index set X′ to obtain the reconstruction error Z r ; In order to maximize the collusion indicator set X and latent variable Z l The information features of the network are introduced into the loss function of the expression network. Figure 2 The global mutual information shown and Figure 3 The local mutual information shown; finally, the latent variables and reconstruction errors in the expression network are passed to the estimation network Gaussian mixture model GMM to calculate its sample energy and then separate the collusion samples.
[0108] The expression network module of the Info-DAGMM model adopts the autoencoder DA to learn the latent variable Z of the conspiracy indicator set X l The above is the screening and reorganization process of complex collusion indicators; the Info-DAGMM model will calculate the error between the collusion indicator set X and the reconstructed sample X′. The larger the feature, the more likely it is to be considered a collusion sample.
[0109] Therefore, the optimization goal of the autoencoder DA is to maximize the information of X and retain it in Z. l , that is, the sample reconstruction error is minimized:
[0110]
[0111] Where, L(x i ,x i ') takes the L2 norm; N is the batch size of each iteration.
[0112] The collusion index set X obtained in step S3 has the characteristic of imbalanced positive and negative samples. After simple dimensionality reduction by DA, the negative sample information is likely to be eliminated by the network as noise, and a good low-dimensional representation cannot be obtained. In order to maximize the collusion index set X and the latent variable Z l The mutual information is introduced into the loss function of the expression network. The mutual information is defined as:
[0113]
[0114] Where p(x) and p(z l ) are X and Z respectively l The probability density function, p(x,z l ) is its joint distribution; p(z l |x) is the latent variable Z l The KL divergence measures how different two distributions are.
[0115] In order to make the low-dimensional space more regular, it is advisable to let the prior distribution q(z l) obeys the standard normal distribution, then the latent variable z l The distance between the true distribution and the prior can be defined using the KL divergence:
[0116]
[0117] According to formula (11) and formula (12), the optimization objective is:
[0118] min{-(β+1)I(X,Z l )+βKL(p(z l |x)||q(z l ))} (13)
[0119] Where β is the weighting coefficient.
[0120] Since KL divergence is unbounded, JS divergence is used to estimate mutual information. Compared with other estimators, JS divergence has better performance when the dataset has fewer negative samples. Therefore, according to formula (13), the mutual information is estimated as:
[0121]
[0122] In the formula, x represents the original sample, is a similar sample that obeys other distributions; the T function is a discriminator modeled by a neural network, which gives positive scores to the combination of original variables and latent variables and negative scores to the combination of interference variables and latent variables; S represents the softplus function: S(z)=log(1+e z ); E is the mathematical expectation of the S function.
[0123] In practice, some samples may contain enough discriminative information. Therefore, we will consider both global mutual information and local mutual information.
[0124]
[0125] Where α1 and α2 are the coefficients of global mutual information and local mutual information, respectively. The combination of α1 and α2 can produce higher mutual information estimates for neural networks. Among them, good reconstruction is highly dependent on α1, and good classification performance is highly dependent on α2. G is the global discriminator, T L is a local discriminator.
[0126] like Figure 2 As shown, the global discriminator T G The working principle is: M and The original variable collusion indicator set X and similarity variable are Compress them into latent variables Z lThe shape of M and Z is mapped to the constant space through the neural network, that is, the global score is obtained. l The mutual information of is the true score; With Z l is an improper fraction.
[0127] like Figure 3 As shown, the local discriminator T L The working principle is: copy the latent variable Z l , expand the shape to the intermediate feature M and shape, and respectively with M and Integration. Mapping to constant space through neural network, we can get local score. M and Z l The mutual information of is the true score; With Z l is an improper fraction.
[0128] In summary, according to formula (13) and formula (15), the loss function of mutual information is as follows:
[0129]
[0130] The estimation network of the Info-DAGMM model uses a GMM to learn the probability density distribution F(x) of low-dimensional features. This is an implicit learning process. Because the collusion indicator set X is characterized by an imbalance of positive and negative samples, the probability density distribution learned by the GMM can be viewed as an approximate probability distribution of normal samples. Therefore, the network output of the Info-DAGMM is the probability of each sample following the approximate sample distribution F(x) of the entire dataset. Clearly, samples with a low probability are collusion samples. This also represents the degree of deviation between the bidding data of two power generation companies and the overall market level, and can be considered the degree of collusion suspicion.
[0131]
[0132] Where, is the output of the entire network, which is the density estimate of the potential expression Z; softmax is the kernel function; Z r Take the Euclidean distance and cosine similarity between the collusion index X and the reconstruction index set X′; θ m To estimate the parameters of the network; h(z, θ m ) is the hidden layer output. The optimization goal of GMM is to maximize the likelihood of all samples. Assume With K-dimensional features, the parameters of GMM can be obtained:
[0133]
[0134] Where, and are the weighted probability, expectation, and variance of the K-th dimension, respectively, and N is the batch size of each iteration. Further, the definition of sample energy can be deduced as follows:
[0135]
[0136] Then, the loss function of the estimation network can be defined as:
[0137]
[0138] Finally, according to Equations (10), (16), and (19), the optimization objective of Info-DAGMM is:
[0139]
[0140] Where λ1 and λ2 are the proportions of the loss functions of the GMM network and mutual information in the joint loss function, respectively.
[0141] In general, a threshold is required to classify collusive enterprises. However, in practice, determining the threshold is very complex and highly subjective. To this end, this paper proposes a Collusion Suspicion Indicator (CSI), which represents the probability of collusion among power generation enterprises and is defined as follows:
[0142] CSI=F(x i );i=1...η (21)
[0143] In the formula, η is the sample size, x i is the collusion index set of the i-th group of power generation enterprises. F(x) is the approximate probability distribution function of the collusion index set X, which is estimated by the GMM in the network.
[0144] The Info-DAGMM network structure and parameter settings for this example are shown in Tables 2 and 3:
[0145] Table 2 Network structure of expression network
[0146]
[0147] Based on the input and output of the expression network, this example uses relative Euclidean distance and relative cosine similarity as the reconstruction error of the sample. The specific calculation formula is:
[0148]
[0149] Then, the input of the estimation network can be expressed as, Z = [Z l ,euc,cos].
[0150] Table 3 Network structure of the estimated network
[0151]
[0152] S5. Inversely map the collusion suspicion back to the measurement matrix of step S3. Through the horizontal and vertical coordinates of the samples, the power generation companies that have colluded in the market can be determined.
[0153] In order to demonstrate the efficiency of the Info-DAGMM model in identifying collusion among power generation enterprises, the model was tested using a data set containing 1000 normal samples and 274 collusion samples.
[0154] As can be seen from the Info-DAGMM model sample suspicion stratification diagram obtained in this example, the collusion suspicion of the collusion samples is as high as over 0.8. This clearly shows the possibility of collusion between power generation companies in the electricity market, which helps provide early warning of collusive behavior. This result shows that the Info-DAGMM model performs very well, and can use its anomaly detection concept to separate a small number of collusion samples from the majority of normal samples. This is due to its ability to learn a highly discernible latent space through mutual information. However, the DAGMM model without mutual information does not perform as well on the same problem.
[0155] Although the DAGMM model clusters normal and collusive points separately in the latent space, the downstream density estimator interprets sparse points as collusive, resulting in a lack of separation between the two clusters and a relatively dispersed distribution of normal points, which can easily interfere with the estimator. In contrast, the Info-DAGMM model completely separates collusive and normal points in the latent space, creating very compact clusters. This is because the Info-DAGMM model locks in the maximum possible information about the collusion indicator set X in the latent variables, making the low-dimensional representation of the samples more discernible. This low-dimensional representation is more conducive to density estimation using the GMM, as it assumes that collusive points are sparse among all points and their distribution does not correspond to the distribution of the majority of normal points. The Info-DAGMM model, incorporating mutual information, can effectively isolate a small number of collusive points, making it more suitable for the data characteristics of power market collusion and providing accurate early warning of collusive behavior among power generation companies.
[0156] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. An intelligent identification method for collusive behavior of power generation enterprises based on the Info-DAGMM model, characterized by: The steps include: S1. Obtain the original declared electricity volume and electricity price of the power generation enterprise; The original declared electricity quantity and original declared electricity price are the declared electricity quantity and declared electricity price of the hth period of the mth power generation enterprise in a certain period; S2. Build an indicator system for identifying collusion among power generation companies. The indicators include: average market share of reported electricity, consistency of bids, consistency of reported electricity, ratio of difference area of bid curves, average safety of bids, and average relative comparison of bids; S3. Calculate and obtain a set of collusion indicators and perform normalization processing; S4. Use the processed collusion indicator set to train the Info-DAGMM model to obtain the density estimation value of each sample in the low-dimensional space, which is the collusion suspicion degree; S5. Inversely map the collusion suspicion back to the measurement matrix calculated in step S3. By using the horizontal and vertical coordinates of the samples, the power generation companies in the market that have engaged in collusion can be identified. In the step S2, The average market share of the declared electricity volume is index 1, which represents the proportion of the declared electricity volume of each power generation enterprise in the market share. The calculation formula is: Where S m and S n are the electricity quantities reported by the mth and nth power generation companies in the market for this bidding; k is the number of power generation companies participating in this bidding; The quotation consistency is indicator 2, which represents the difference in the declared prices of two power generation companies. The calculation formula is: Where p mh and p nh The bids of the mth and nth power generation companies in the hth segment of this bidding are respectively; is the average of the bids of all power generation companies in the hth segment in this bidding; The consistency of the reported electricity is indicator 3, which represents the difference in the reported electricity of two power generation companies. The calculation formula is: Where S mh and S nh The electricity quantity reported by the mth and nth power generation companies in the hth segment of this bidding respectively; is the average of the electricity volume declared by all power generation companies in the hth period in this bidding; The bid curve difference area ratio is indicator 4, which reflects the degree of parallel bidding between two power generation companies. The calculation formula is: Where, f m and f n are the bidding curve functions of the mth and nth power generation companies in this bidding respectively; The average safety of the bid is index 5, which measures the deviation between the bids of power generation companies and the historical average bid. The calculation formula is: Where, and are the weighted average bid prices of the mth and nth power generation companies in this bidding; E is the expected value of the market marginal price, which is calculated using the marginal prices of historical transactions; The relative bid comparison mean is indicator 6, which represents the difference between the bids of two power generation companies and the average price of this centralized bidding. The calculation formula is:
2. The intelligent identification method of power generation enterprise collusion behavior based on the Info-DAGMM model according to claim 1 is characterized in that: In step S3, assuming that there are m power generation companies bidding for a total of h periods in the centralized bidding, the calculation process of the collusion index dataset X is: According to the calculation formula of the collusion index system in step S2, the qth index between each pair of power generation enterprises is calculated to form a data matrix: The column vectors are obtained by tiling the upper triangle elements in sequence, which is an indicator feature of the collusion indicator set: All 6 indicator features are calculated and normalized by column combination to obtain the collusion indicator set X for deep joint network training: X=(x (1) ,x (2) ,K,x (6) ) (9)。 3. The intelligent identification method of power generation enterprise collusion behavior based on the Info-DAGMM model according to claim 1 is characterized in that: In step S4, the construction process of the Info-DAGMM model is as follows: The network framework of the Info-DAGMM model mainly consists of two modules: the expression network module and the density estimation module. First, the Encoder in the autoencoder DA is used to train the collusion indicator set X, and the mutual information is introduced into the loss function of the expression network to maximize the original variable and the low-dimensional space latent variable Z. l The information features are then decoded by the Decoder in DA to obtain the reconstruction index set X′; then, the Euclidean distance and cosine similarity are used to measure the collusion index set X and the reconstruction X′ to obtain the reconstruction error Z r ; Finally, the latent variable Z l and reconstruction error Z r The data is passed to the Gaussian mixture model (GMM) to calculate its sample energy and separate the collusion samples. During training, the parameters of the autoencoder and the Gaussian mixture model are optimized simultaneously. In the expression network, the autoencoder DA will learn the latent variable Z of the collusion indicator set X l The optimization goal is to maximize the information of X retained in Z l , that is, the sample reconstruction error is minimized: Where, L(x i ,x i ') Take the L2 norm; N is the batch size of each iteration; To maximize the collusion indicator set X and latent variable Z l The information characteristics of the network are introduced into the loss function of the expression network. The mutual information is defined as: Where p(x) and p(z l ) are X and Z respectively l The probability density function of p(x,z l ) is its joint distribution; p(z l |x) is the latent variable Z l The true conditional distribution of ; KL divergence measures the degree of difference between two distributions; When estimating mutual information, JS divergence is used for estimation, and the mutual information is estimated as: In the formula, x represents the original sample, is a similar sample that obeys other distributions; the T function is a discriminator modeled by a neural network, which gives positive scores to the combination of original variables and latent variables and negative scores to the combination of interference variables and latent variables; S represents the softplus function: S(z)=log(1+e z ); E is the mathematical expectation of the S function; Considering both global and local mutual information: Where α1 and α2 are the coefficients of global mutual information and local mutual information respectively. The combination of α1 and α2 can produce higher mutual information estimation of neural network. Among them, good reconstruction is highly dependent on α1, and good classification performance is highly dependent on α2. G is the global discriminator, T L is the local discriminator; In summary, the loss function of mutual information is as follows: Where, q(z l ) is the latent variable Z l The prior distribution of In the estimation network, GMM learns the probability density distribution F(x) of low-dimensional features, and the output of the network is the collusion suspicion between any two power generation companies; Where, is the output of the entire network, which is the density estimate of the potential expression Z; softmax is the kernel function; Z r Take the Euclidean distance and cosine similarity between the collusion index set X and the reconstruction index set X′; θ m To estimate the parameters of the network; h(z, θ m ) is the hidden layer output; The loss function of the estimation network is defined as: Where E(z) is the sample energy; Finally, according to Equations (10), (14), and (16), the optimization objective of the Info-DAGMM model is: Where λ1 and λ2 are the proportions of the loss functions of the GMM network and mutual information in the joint loss function, respectively.
4. The intelligent identification method of power generation enterprise collusion behavior based on the Info-DAGMM model according to claim 3 is characterized in that: The evaluation index of conspiracy suspicion is defined as: CSI=F(x i );i=1Kη (18) In the formula, η is the sample size, x i is the indicator set of the i-th group of power generation enterprises, and F(x) is the approximate probability distribution function of the collusion indicator set X, which is estimated by the GMM in the network.
5. The intelligent identification method of power generation enterprise collusion behavior based on the Info-DAGMM model according to claim 4 is characterized in that: In step S5, the collusion suspicion CSI is inversely mapped back to the measurement matrix (9) of step S3. The horizontal and vertical coordinates of the samples can be used to determine the power generation companies that have engaged in collusion in the market.