A power generation enterprise identification method based on cost-sensitive support vector machine
By using a cost-sensitive support vector machine approach combined with K-means and a custom nearest neighbor algorithm, the problem of insufficient data labels and low computational efficiency in identifying abuse of market power by power generation companies was solved. This approach achieves efficient and accurate identification of abuse of market power, thus ensuring the healthy development of the electricity market.
Patent Information
- Application Number
- CN202111556421.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-17
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2041-12-17
AI Technical Summary
Existing technologies struggle to efficiently identify power generation companies abusing their market power, especially when data labels are scarce and unbalanced. Traditional methods suffer from overfitting and low computational efficiency.
A cost-sensitive support vector machine-based approach is adopted, combined with the K-means algorithm to initially assign labels to unlabeled samples. Incorrect labels are corrected by pairwise label swapping. The problem is transformed into a variational inequality problem, which is solved by a customized nearest neighbor algorithm, enabling rapid identification of power generation companies that abuse market power.
It enables efficient and real-time identification of abuses of market power in the electricity market, improves identification accuracy and recall rate, meets the real-time monitoring needs of the electricity market, and safeguards the interests of market participants and a fair environment.
Smart Images

Figure QLYQS_8 
Figure QLYQS_9 
Figure QLYQS_10
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of power market subject risk identification method design, and particularly relates to a power generation enterprise identification method based on a cost-sensitive support vector machine. BACKGROUND
[0002] Under the promotion of the new round of power reform, the marketization degree of China's power market is continuously strengthened, and market transactions are continuously increased, which makes the power market develop rapidly while accompanied by corresponding market risks.
[0003] At present, the methods for identifying the abuse of market power of power generation enterprises in China mainly include expert decision-making and constructing a supervision index system method. However, with the expansion of the transaction scale of the power market and the increase of the transaction frequency, the expert decision-making method cannot meet the demand of real-time identification of the abuse of market power of power generation enterprises, and it is necessary to find an intelligent identification method that can be calculated by a computer. In addition, the acquisition of indicators in the supervision index system is based on the premise that the real situation of the power generation enterprise is specifically and comprehensively understood, which is difficult to achieve in the actual power market, and only the bidding data can be easily obtained. Therefore, the application proposes an intelligent identification method for the abuse of market power of power generation enterprises based on bidding data.
[0004] The identification of the abuse of market power of power generation enterprises is essentially a two-classification problem. As an intelligent algorithm, the support vector machine has the characteristics of high accuracy and strong generalization ability, and has strong applicability in two-classification problems. However, since only a small amount of data in the data of power generation enterprises is labeled data, the use of traditional support vector machines may cause overfitting problems. The semi-supervised support vector machine can use the data structure information of unlabeled samples to well solve the problem of poor generalization ability caused by too few labeled samples. Therefore, the transductive support vector machine (TSVM) in the semi-supervised support vector machine is adopted, and the cost-sensitive transductive support vector machine (CTSVM) is selected in view of the characteristics of unbalanced violation data. In addition, based on the problem that the TSVM itself needs to preset the number of positive and negative class sample points, the K-means method is used to assign initial labels to unlabeled data. In order to avoid the possible labeling errors of the semi-supervised support vector machine, the paired label exchange method is used to obtain the predicted label. In terms of computational complexity, the time complexity of TSVM is much higher than that of ordinary support vector machines, and the paired label exchange method has the characteristics of high accuracy but slow speed. In view of this problem, a more efficient variational inequality method based on a customized nearest neighbor point algorithm is used to speed up the solution method to obtain the discriminant function, and realize the real-time identification of the abuse of market power of power generation enterprises. SUMMARY
[0005] In view of the fact that only a small amount of data in the power generation enterprise data mentioned in the background art is labeled data and the traditional abuse of market power behavior recognition method has low processing efficiency for transaction data, the application provides a power generation enterprise recognition method based on a cost-sensitive support vector machine, which overcomes the shortcomings of the prior art and has good effects.
[0006] The power generation enterprise recognition method based on the cost-sensitive support vector machine comprises the following steps:
[0007] S1, obtaining declared power and declared power price data of a power generation enterprise;
[0008] S2, constructing an abuse of market power behavior recognition index system of the power generation enterprise, the indexes comprising: power generation enterprise bidding data, each segment declared power share, each segment transaction power share, each segment declared success ratio, whether it is a high price in this segment, and high bidding frequency;
[0009] S3, when the bidding segment number is high, using principal component analysis to reduce the dimension of the data;
[0010] S4, using the K-means algorithm to assign initial "pseudo-labels" to unlabeled samples, combining a semi-supervised support vector machine, and using a pair-wise label exchange method to exchange incorrectly labeled labels to obtain predicted labels;
[0011] S5, substituting the processed abuse of market power data set into a cost-sensitive direct support vector machine, converting the solving problem into a variational inequality problem, and using a customized proximal point algorithm to solve the problem, and determining the sample label y i by a discriminant function to obtain the power generation enterprise that abuses market power.
[0012] Preferably, in step S1, the original declaration data of m power generation enterprises in a period is obtained and and respectively represent the declared power price and the declared power of the i-th power generation enterprise in the j-th segment; for example, in a three-segment bidding rule,
[0013] Preferably, in step S2, a perfect abuse of market power behavior recognition index system is constructed according to the characteristics and features of the past abuse of market power behavior recognition indexes and in combination with the actual situation of the electricity market;
[0014] The power generation enterprise bidding data is index 1, and the calculation formula is:
[0015]
[0016] wherein: x 1i denotes the bidding data of the power generation enterprise, denotes the bidding price, the bidding capacity and the transaction capacity of the i-th power generation enterprise in the j-th segment, respectively, i=1, 2, …, m, j=1, 2, …, n, m is the number of power generation enterprises, n is the number of bidding segments;
[0017] The bidding capacity share of each segment is index 2, which indicates the market power of the enterprise in the bidding segment, and the calculation formula is:
[0018]
[0019] wherein: x 2ij denotes the bidding capacity share of the i-th enterprise in the j-th segment;
[0020] The transaction capacity share of each segment is index 3, which indicates the successful bidding capacity of the power generation enterprise, and reflects the possibility of using market power, and the calculation formula is:
[0021]
[0022] wherein: x 3ij denotes the transaction capacity share of the i-th enterprise in the j-th segment;
[0023] The bidding success ratio of each segment is index 4, which reflects the success or failure of the bidding strategy of the enterprise in the bidding segment, and reflects the possibility of abusing market power, and the calculation formula is:
[0024]
[0025] wherein: x 4ij denotes the bidding success ratio of the i-th enterprise in the j-th segment;
[0026] Whether it is the highest price in the segment is index 5, which indicates whether the enterprise is willing to take greater risks to obtain higher benefits, and reflects the intention of using market power, and the calculation formula is:
[0027]
[0028] wherein: x 5ij denotes whether the i-th enterprise is the highest price in the j-th segment;
[0029] The number of high bidding is index 6, which indicates the intention of the enterprise to abuse market power, and the calculation formula is:
[0030]
[0031] wherein: x 6i denotes the number of high bidding of the i-th power generation enterprise in the j-th segment.
[0032] Preferably, in step S3, since the power generation companies' bidding data has a high dimensionality when there are many bidding segments, and the support vector machine has poor applicability in the case of high-dimensional data, the principal component analysis method PCA is used to reduce the dimensionality of the high-dimensional data.
[0033] Preferably, in step S4, the constructed sample data includes the indicators in the indicator system based on the quoted price data: power generation enterprise quoted price data x 1i Electricity share declared for each segment x 2i Share of transaction volume in each segment x 3i The success rate of each application segment is x 4i Is this the highest price in this segment? 5i Number of high bids x 6i And power generation companies abusing market power labels i .
[0034] To address the limited availability of labeled data for power generation company quotations and the highly imbalanced nature of the data, a Cost-Sensitive Transitive Support Vector Machine (CTSVM) was adopted. The TSVM algorithm itself requires a pre-defined number of positive and negative class samples in the unlabeled samples. However, since the actual number of labeled samples is small, the number of positive and negative class samples in the labeled samples is insufficient to represent the number of positive and negative class samples in the unlabeled samples. Therefore, the K-means algorithm is used to initially assign "pseudo-labels" to the unlabeled samples, and then incorrectly labeled samples are swapped in pairs to obtain the predicted labels.
[0035] Given a labeled sample set D across all samples. l ={(x1,y1),(x2,y2),…,(x l ,y l )} and unlabeled sample set D u ={(x l+1 ,y l+1 ),(x l+2 ,y l+2 ),…,(x l+u ,y l+u )}. Where l is the number of labeled samples, u is the number of unlabeled samples, and l + u = h is the total number of samples. i Let ∈{1,-1} be the sample label, where 1 indicates no violation and -1 indicates violation. Let the optimal classification boundary be ω. T x+b=0, the goal of CTSVM is to process the unlabeled sample set D. u The middle sample provides the predicted label. The quadratic programming model of CTSVM is as follows:
[0036]
[0037] In the formula: ξ is the relaxation vector, ξi (i = 1, 2, …, l) corresponding to the slack variable of the labeled sample, ξ i (i = l + 1, l + 2, …, h) corresponding to the slack variable of the unlabeled sample. C l ,C u respectively corresponding to the penalty coefficient of the labeled sample and the penalty coefficient of the unlabeled sample, indicating the importance of the two types of samples. Among them, the penalty coefficient C l and C u are defined as follows:
[0038]
[0039] and satisfy where num1 and num2 respectively represent the number of positive and negative samples in the labeled sample.
[0040] When the sample points are linearly inseparable, a kernel function needs to be introduced to convert them into linearly separable data for processing. The general definition of the kernel function is as follows:
[0041] K(x i ,x j ) = <φ(x i ), φ(x j )> = φ T (x i ) · φ(x j )
[0042] where <.,.> represents the inner product operation.
[0043] The specific training steps of the K-CTSVM are as follows:
[0044] 1) Step 1: Set the values of C l , C u , where C u ≤ C l , use the K-means algorithm to cluster the unlabeled samples into two classes, count the number of samples in the two classes, and label the class with fewer samples as and the other class as to obtain the initial "pseudo label" of the unlabeled sample;
[0045] 2) Step 2: Use the "pseudo label" sample in step 1 and the labeled sample together to substitute into CTSVM for training, find a pair of opposite "pseudo label" determine whether the corresponding slack variable satisfies ξ i > 0, ξ j > 0, ξ i + ξ j > 2, if the condition is met, exchange the label After the exchange is completed, re-substitute it into the CTSVM for training, repeat the above process until there is no "pseudo label" that meets the above conditions;
[0046] 3) Step 3: gradually increase the parameter C u , C u = min{2C u , C l}, repeat the operation in step 2 until C u ≥ C l , the algorithm ends, and the last training result is output.
[0047] Preferably, in step S5, due to the high time complexity of TSVM itself, the calculation is slow when the sample data volume is large. And TSVM adopts the way of exchanging labels in pairs, which is accurate but slow. In order to speed up the solving speed, the solving problem is transformed into a variational inequality problem, and a customized proximal point algorithm is used to solve the variational inequality problem, the specific steps are as follows:
[0048] Step 1: According to the Lagrange dual principle, formula (7) is transformed into:
[0049]
[0050] Let the matrix Transform formula (8) into matrix form, the specific form is as follows:
[0051]
[0052] In the formula: C is an h x 1 matrix, C(i) = C l (i = 1, 2,..., l), C(i) = C u (i = l + 1, l + 2,..., h), and e represents a column vector with all components being 1.
[0053] Step 2: Let the null space matrix of data label vector y T α = 0, then α in the original problem can be linearly represented by the matrix Z as α = Zυ. Then the original problem can be transformed into a convex optimization problem with linear inequality constraints:
[0054]
[0055] Step 3: First, write the convex optimization problem in the form of Lagrange function:
[0056] L(υ,λ) = θ(υ) - λ T (Aυ - b)
[0057] Then find the saddle point of its Lagrange function (υ* ,λ * ) corresponds to solving the following variational inequality problem:
[0058] θ(υ)-θ(υ * )+(ω-ω * ) T F(ω * )≥0,ω∈Ω
[0059] where:
[0060] Step 4: Solve the variational inequality problem using the customized proximal point algorithm. For a given ω k =(υ k ,λ k ), the proximal point algorithm is customized and specifically expressed as follows:
[0061] ω k =(υ k ,λ k )
[0062]
[0063] Substitute θ(υ) in (10) into the above formula to obtain the K-CTSVM iterative algorithm based on CPPA:
[0064]
[0065] where: P=[Z T HZ+diag(r)] -1 When r·s≥||A T A||, the CPPA algorithm converges.
[0066] Step 5: Use the discriminant function f(x) = sgn[w * x+b * ] to determine the sample label When the output result is -1, it means that the power generation enterprise abuses market power.
[0067] The specific process of the K-CTSVM iterative solution algorithm based on CPPA is as follows:
[0068] 1) Set the initial value k = 0, ω 0 =(υ 0 ,λ 0 ), and select r, s that satisfy the conditions;
[0069] 2) According to ω k , iterate according to the iteration method of formula (11) to obtain and ω k+1 ;
[0070] 3) If Record u k And go to 4, otherwise k=k+1 go back to 2;
[0071] 4) Calculate alpha by alpha=Zu
[0072] 5) Calculate the parameters of the classification hyperplane according to alpha:
[0073] 6) Determine the sample label using the discriminant function f(x)=sgn[w * x+b * ]When the output result is -1, it indicates that the power generation enterprise abuses market power.
[0074] The power generation enterprise identification method based on the cost-sensitive support vector machine has the beneficial effects that the power generation enterprise identification method based on the cost-sensitive support vector machine can efficiently and timely identify the abuse of market power in the power market transaction, better maintain the self-interest of the power market subject, and build a fair market environment, and has important significance for implementing market supervision, establishing a good market order, and guaranteeing the healthy development of the power market. DETAILED DESCRIPTION
[0075] Embodiment 1
[0076] The power generation enterprise identification method based on the cost-sensitive support vector machine comprises the following steps:
[0077] S1, using 4 sets of three-stage bidding data in a provincial power market centralized bidding as original data and Wherein and denote the bidding price and the bidding capacity of the i-th power generation enterprise in the j-th segment, respectively;
[0078] S2, constructing an abuse of market power behavior identification index system of the power generation enterprise, the indexes comprising: power generation enterprise bidding data, segment bidding capacity share, segment transaction capacity share, segment bidding success ratio, whether it is a high price in the segment, and high bidding times;
[0079] Specifically, the power generation enterprise bidding data is index 1, and the calculation formula is:
[0080]
[0081] In the formula, x 1i denotes the bidding data of the power generation enterprise, respectively represent the declared price, declared electricity quantity and traded electricity quantity of the i-th power generation enterprise in the j-th segment, i = 1, 2, …, m, j = 1, 2, …, n, m is the number of power generation enterprises, and n is the number of bidding segments;
[0082] The declared electricity quantity share of each segment is index 2, and the calculation formula is:
[0083]
[0084] In the formula, x 2ij represents the declared electricity quantity share of the i-th enterprise in the j-th segment;
[0085] The traded electricity quantity share of each segment is index 3, and the calculation formula is:
[0086]
[0087] In the formula, x 3ij represents the traded electricity quantity share of the i-th enterprise in the j-th segment;
[0088] The declared success rate of each segment is index 4, and the calculation formula is:
[0089]
[0090] In the formula, x 4ij represents the declared success rate of the i-th enterprise in the j-th segment;
[0091] Whether it is the highest price in this segment is index 5, and the calculation formula is:
[0092]
[0093] In the formula, x 5ij represents whether the i-th enterprise is the highest price in the j-th segment;
[0094] The number of high bids is index 6, and the calculation formula is:
[0095]
[0096] In the formula, x 6i represents the number of high bids of the i-th power generation enterprise in the j-th segment.
[0097] S3, using principal component analysis to reduce the dimension of the data, using K-means algorithm to assign initial "pseudo-labels" to unlabeled samples, combining semi-supervised support vector machine, and using pair-wise exchange labels to exchange incorrectly labeled labels to obtain predicted labels, see Table 1 for specific steps.
[0098] Table 1 Specific steps of intelligent identification method
[0099]
[0100]
[0101] S4, the processed abuse market power data set is substituted into the cost-sensitive direct support vector machine, the solving problem is converted into a variational inequality problem, and a customized proximal point algorithm is used for solving, specific steps are shown in Table 2, and the sample label y is determined through a discrimination function i , and the power generation enterprise that abuses market power is obtained;
[0102] Table 2 K-CTSVM iteration solving algorithm flow based on CPPA
[0103]
[0104] S5, the evaluation indexes adopt the correct rate and the recall rate, and the correct rate and the recall rate are defined as follows:
[0105] The correct rate is:
[0106] The recall rate is:
[0107] In the formula: TP, TN, FP and FN are defined in the confusion matrix shown in Table 3.
[0108] Table 3 Confusion matrix
[0109]
[0110] In order to reflect the identification efficiency of the improved cost-sensitive direct support vector machine for the power generation enterprise that abuses market power, three data set examples are compared with other intelligent algorithms:
[0111] UCI data set test:
[0112] K-CTSVM and SVM algorithms are used to test the UCI data set (Iris, Wine, Glass, Haberman). In order to fully simulate the actual power market data, the number of test data sets is set to 50, the proportion of negative class samples accounts for 10% of the total number of samples, and the number of labeled samples accounts for 20% of the total number of samples. The UCI data set test results are shown in Table 4.
[0113] Table 4 UCI data set test results
[0114]
[0115] As can be seen from Table 4, the recall rate of K-CTSVM is much higher than that of SVM under the condition that the correct rate of K-CTSVM is similar to that of SVM. It is shown that the fitting effect of the K-CTSVM algorithm is better than that of the traditional SVM, and the identification efficiency is higher, so the algorithm proposed in the application can effectively identify the violation samples which only account for a small part of the samples.
[0116] Electricity market data test:
[0117] The electricity market simulation experiment is carried out on the data set using K-CTSVM, TSVM and SVM algorithms. The data set is constructed based on three-stage bidding and uniform clearing price mechanism. The number of power generation enterprises is set to 39, and 5 power generation enterprises are determined to abuse market power, and 34 power generation enterprises do not abuse market power. 20% of the samples are selected as labeled samples, and the rest are unlabeled samples.
[0118] The sample set is constructed using the abuse market power identification index system, and the sample data is processed by PCA dimension reduction. The dimension-reduced data set is substituted into the support vector machine (SVM), direct support vector machine (TSVM) and improved cost-sensitive direct support vector machine (K-CTSVM) for testing, and the CPPA algorithm and the sequential minimal optimization algorithm (Sequential Minimal Optimization, SMO) are used for calculation. The accuracy, recall rate and time of the three algorithms are compared, and the comparison results are shown in Table 5.
[0119] Table 5 Test results of electricity market simulation data
[0120]
[0121] From Table 5, it can be seen that under the condition of knowing only a small amount of data sample labels, the K-CTSVM proposed in this paper has the same accuracy as SVM and TSVM, but the recall rate of K-CTSVM is much higher than that of SVM and TSVM, which shows that K-CTSVM is more accurate and comprehensive than the former two in identification, and is more suitable for identifying the abuse of market power of power generation enterprises in actual electricity market. In terms of time consumption, the process of increasing the penalty coefficient of unlabeled samples each time in TSVM contains multiple SVM training, and the time complexity is much higher than that of SVM. However, since K-CTSVM algorithm adopts a more efficient variational inequality solution, it takes less time than TSVM solved by SMO algorithm. In addition, the solving speed of SVM using CPPA algorithm is much faster than that of SMO algorithm, which also shows the superiority of CPPA solving algorithm. Therefore, the method proposed in this paper can quickly and effectively identify the abuse of market power of power generation enterprises.
[0122] Actual electricity market data:
[0123] In order to further verify the feasibility of the method in the actual power market, the bidding data of 50 power generation enterprises in a province are selected for testing. Since the existing data in the actual power market only has the transaction price and transaction capacity data, and no bidding capacity data, we delete the successful bidding ratio index. And the data is one-segment bidding, so the high bidding frequency index coincides with whether it is a high bidding index in this segment, and the high bidding frequency index is changed to the sum of whether it is a high bidding in this segment and whether it is a high share in this segment. The calculation method of whether it is a high bidding in this segment is: 1 for yes, 0 for no. The calculation method of whether it is a high share in this segment is: sort the transaction capacity share from high to low, select the top 30% and take the value as 1, and the rest take the value as 0. Among them, the labeled samples account for 20% of the total samples, the violation samples account for 10% of the total, and the rest are non-violation samples. The normalized data are respectively substituted into K-CTSVM and SVM to test the unlabeled samples, and the test results are shown in Table 6.
[0124] Table 6 Actual power market data test
[0125]
[0126] As can be seen from Table 6, the K-CTSVM algorithm proposed in this paper has higher accuracy and recall rate compared with SVM and TSVM. This shows that K-CTSVM as a semi-supervised algorithm is more suitable for testing data sets with fewer labeled samples than general support vector machines, and is more suitable for identifying the violation behavior of power generation enterprises in the actual power market. In terms of time, the K-CTSVM algorithm solved by variational inequality is much faster than the TSVM solved by SMO algorithm, so the solving method adopted by the present application is more efficient and can better meet the real-time monitoring needs of power generation enterprises abusing market power.
Claims
1. A method for identifying power generation enterprises based on cost-sensitive support vector machines, characterized in that, Includes the following steps: S1. Obtain data on the declared electricity volume and declared electricity price from power generation companies; S2. Construct an indicator system to identify the abuse of market power by power generation companies. The indicators include: power generation company bidding data, the share of electricity declared in each segment, the share of electricity traded in each segment, the success rate of bidding in each segment, whether it is the highest price in this segment, and the number of times the high bid is made. S3. When the number of price segments is high, use principal component analysis to reduce the dimensionality of the data; S4. The K-means algorithm is used to assign initial "pseudo-labels" to unlabeled samples. Combined with a semi-supervised support vector machine, the incorrectly labeled labels are swapped in pairs to obtain the predicted labels. In step S4, the constructed sample data includes the indicators in the indicator system based on the quoted price data: power generation enterprise quoted price data. Electricity share declared for each segment Share of transaction volume in each segment Success rate of each application segment Is this the highest price in this segment? High-quote frequency And power generation companies' abuse of market power labels ; In all samples, suppose there is a labeled sample set. and unlabeled sample set Where l represents the number of labeled samples and u represents the number of unlabeled samples. This represents the total number of samples; Let be the sample labels, where 1 indicates no violation and -1 indicates a violation; let the optimal classification interface be... The goal of CTSVM is to provide unlabeled sample sets The middle sample provides the predicted label. , The quadratic programming model for CTSVM is as follows: (7) In the formula: Let be the relaxation vector. Slack variables corresponding to labeled samples, Relaxed variables corresponding to unlabeled samples; , These correspond to the penalty coefficients for labeled samples and unlabeled samples, respectively, representing the importance of the two classes of samples; where the penalty coefficient... and The definition is as follows: , And satisfy , where num1 and num2 represent the number of positive and negative samples in the labeled samples, respectively; When the sample points are linearly inseparable, a kernel function is needed to transform them into linearly separable data for processing; the definition of the kernel function is as follows: ,in Indicates inner product operation; S5. Substitute the processed dataset of abused market power into a cost-sensitive transductive support vector machine, transforming the problem into a variational inequality problem, and solve it using a custom nearest neighbor algorithm. The sample labels are then determined using a discriminant function. Power generation companies that abuse market power; In step S5, TSVM suffers from high time complexity, leading to slow computation when dealing with large amounts of sample data. Furthermore, TSVM employs a pairwise label swapping method, which, while accurate, is slow. To accelerate the solution, the problem is transformed into a variational inequality problem, and a custom nearest neighbor algorithm is used to solve it. The specific steps are as follows: Step 1: Based on the Lagrange duality principle, transform equation (7) into: (8) Let matrix Equation (8) is transformed into matrix form, as follows: (9) In the formula: for The matrix, , ,and This represents a column vector whose components are all 1; Step 2: Set the data label vector The null space matrix is ,because In the original problem Can be derived from matrix Linear representation is Then the original problem can be transformed into a convex optimization problem with linear inequality constraints: (10) Step 3: First, write the convex optimization problem in Lagrange function form: Then find its Lagrange function saddle point. This is equivalent to solving the following variational inequality problem: In the formula: ; Step 4: Solve the variational inequality problem using a custom nearest neighbor algorithm; for a given The custom nearest neighbor algorithm is specifically represented as follows: ; In equation (10) Substituting into the above equation, we obtain the K-CTSVM iterative algorithm based on CPPA: (11) In the formula: ;when At that time, the CPPA algorithm converged; Step 5: Use the discriminant function Determine sample labels When the output result is -1, it indicates that the power generation company has abused its market power.
2. The power generation enterprise identification method based on cost-sensitive support vector machine as described in claim 1, characterized in that, In step S1, the original declaration data of m power generation companies during a certain period are obtained. and , and They represent the first The first power generation company The declared electricity price and declared electricity volume for each segment.
3. The power generation enterprise identification method based on cost-sensitive support vector machine as described in claim 2, characterized in that, In step S2 The power generation company's quoted price data is index 1, and the calculation formula is: (1) In the formula: This represents the price quotes from power generation companies. They represent the first The first power generation company The declared electricity price, declared electricity volume, and transaction volume for each segment. , m represents the number of power generation companies, and n represents the number of price ranges; The electricity share for each segment is index 2, and the calculation formula is as follows: (2) In the formula: Indicates the first The first company The declared electricity share for each segment; The share of transaction volume in each segment is index 3, and the calculation formula is as follows: (3) In the formula: Indicates the first The first company The share of electricity traded in each segment; The success rate of each segment's application is index 4, and the calculation formula is as follows: (4) In the formula: Indicates the first The first company The success rate of segment application; Whether a price is considered high within this segment is indicator 5, and the calculation formula is as follows: (5) In the formula: Indicates the first The company in the Is the price quoted in this segment the highest price in this segment? The number of high bids is indicator 6, and the calculation formula is as follows: (6) In the formula: Indicates the first The first power generation company Number of high-priced bids.
4. The power generation enterprise identification method based on cost-sensitive support vector machine as described in claim 3, characterized in that, In step S3, since the power generation companies' bidding data has a high dimensionality when there are many bidding segments, and the support vector machine has poor applicability in the case of high-dimensional data, principal component analysis (PCA) is used to reduce the dimensionality of the high-dimensional data.