Electric power spot market boundary data generation method, equipment and medium
By separating the dynamic and static features of historical data of the electricity spot market, eliminating strongly correlated features and performing sampling processing, accurate boundary data is generated, which solves the problem that existing technologies cannot characterize complex market behaviors and improves data accuracy and market optimization capabilities.
Patent Information
- Application Number
- CN202510683509.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies have difficulty generating accurate electricity spot market boundary data and are unable to effectively characterize complex market behavior characteristics.
By obtaining the initial feature set from historical data, dividing it into dynamic and static feature sets, eliminating strongly correlated features, using the random forest model to determine the importance, classifying and sampling to generate the boundary data set, and combining feature types and historical data to characterize market characteristics.
The accuracy of electricity spot market boundary data has been improved, which can better characterize complex market behavior characteristics and support market optimization and operation.
Smart Images

Figure CN120634599A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of power market, and in particular to a method, device and medium for generating boundary data of power spot market. Background Art
[0002] With the gradual opening of the power market and the continuous improvement of its operating mechanisms, the electricity spot market, where eligible operators conduct day-ahead, intraday, and real-time electricity transactions, has played a vital role in promoting the optimal allocation of energy resources, improving market competition efficiency, and ensuring the safe and stable operation of the power grid. As a core component of power market operations, boundary data is a crucial foundation for market clearing calculations, security constraint analysis, and power trading optimization. Therefore, scientifically and efficiently generating boundary data that meets actual operating conditions has become a key technical issue in optimizing the operation of the electricity spot market.
[0003] The electricity spot market aims to balance market supply and demand, and determines market operating conditions and constraints across multiple scenarios by generating boundary data. However, existing methods for generating boundary data rely on single mathematical models or rule definitions, making it difficult to characterize complex market behavior. Summary of the Invention
[0004] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention proposes a method for generating electricity spot market boundary data, which can generate accurate electricity spot market boundary data to depict complex market behavior characteristics.
[0005] The present invention also proposes a device and a medium having the above-mentioned method for generating boundary data of the electricity spot market.
[0006] A method for generating electricity spot market boundary data according to a first aspect of an embodiment of the present invention includes:
[0007] Obtain the initial feature set from historical data;
[0008] Obtaining a first feature set and a second feature set according to the initial feature set;
[0009] Obtaining a third feature set and a fourth feature set according to the first feature set and the second feature set respectively;
[0010] According to the third feature set, a fourth feature set, a fifth feature set and a sixth feature set are obtained;
[0011] The fourth feature set, the fifth feature set, and the sixth feature set are sampled according to the second feature set to obtain a boundary data set.
[0012] A method for generating electricity spot market boundary data according to an embodiment of the present invention has at least the following beneficial effects: classifying the historical data of the electricity spot market according to its characteristics to obtain a first feature set and a second feature set, then processing the first feature set and the second feature set respectively, combining the feature type and historical data to characterize complex market characteristics, thereby generating electricity spot market boundary data and improving data accuracy.
[0013] According to some embodiments of the present invention, obtaining a third feature set and a fourth feature set based on the first feature set and the second feature set, respectively, includes:
[0014] Obtaining a third feature set and a fourth feature set according to the first feature set and the second feature set respectively;
[0015] Strongly correlated feature pairs in the third feature set and the fourth feature set are respectively eliminated, and the third feature set and the fourth feature set are updated.
[0016] According to some embodiments of the present invention, obtaining a third feature set and a fourth feature set based on the first feature set and the second feature set, respectively, includes:
[0017] Obtaining, based on the first feature set and the target variable, a first importance of each feature in the first feature set to the target variable;
[0018] Obtaining, based on the second feature set and the target variable, a second importance of each feature in the second feature set to the target variable;
[0019] Obtaining the third feature set according to the first importance;
[0020] The fourth feature set is obtained according to the second importance.
[0021] According to some embodiments of the present invention, obtaining the fourth feature set, the fifth feature set, and the sixth feature set according to the third feature set includes:
[0022] The third feature set is classified according to the value type of each feature in the third feature set to obtain the fourth feature set, the fifth feature set, and the sixth feature set.
[0023] According to some embodiments of the present invention, obtaining a boundary dataset according to the second feature set and sampling the fourth feature set, the fifth feature set, and the sixth feature set respectively includes:
[0024] According to the fourth feature set, the fifth feature set, and the sixth feature set, sampling and integrating them respectively obtain a first boundary data set;
[0025] Obtain a second boundary dataset based on the first boundary dataset and the second feature set;
[0026] A boundary dataset is obtained according to the second boundary dataset.
[0027] According to some embodiments of the present invention, obtaining a second boundary dataset according to the first boundary dataset and the second feature set includes:
[0028] Obtaining a first related data set according to each boundary data in the first boundary data set;
[0029] Obtaining a second related data set according to each boundary data in the second feature set;
[0030] The first related data set, the second related data set, the first boundary data set and the second feature set are combined to obtain a second boundary data set.
[0031] According to some embodiments of the present invention, obtaining the boundary dataset according to the second boundary dataset includes:
[0032] A boundary data set is obtained according to each boundary data in the second boundary data set and its threshold value, wherein the threshold value is a preset threshold value for each boundary data in the second boundary data set.
[0033] According to some embodiments of the present invention, for features in the sixth feature set, sample values are obtained by the following method:
[0034] A distribution function is obtained according to the number of grid points and the discretization step size of the feature map; wherein the feature map is a distribution map of features in the sixth feature set, and the discretization step size is a discrete step size of the feature map;
[0035] Obtaining a preliminary sampling value according to the inverse function of the distribution function and a random number that conforms to a uniform distribution;
[0036] A sampling value is obtained according to the preliminary sampling value, the distribution function, and the random number, and the sampling value is integrated into the first boundary data set.
[0037] An electronic device according to an embodiment of a second aspect of the present invention includes:
[0038] Memory, used to store programs;
[0039] A processor is used to execute the program stored in the memory. When the processor executes the program stored in the memory, the processor is used to execute the method as described in any one of the first aspects.
[0040] According to a third aspect of an embodiment of the present invention, a storage medium stores computer-executable instructions, where the computer-executable instructions are used to execute the method as described in any one of the first aspects.
[0041] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purposes and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The accompanying drawings are used to provide a further understanding of the technical solution of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the technical solution of the present invention and do not constitute a limitation to the technical solution of the present invention.
[0043] Figure 1 This is a flow chart of a method for generating boundary data of a power spot market provided by one embodiment of the present invention;
[0044] Figure 2 is a schematic diagram of step S310 in a method for generating electricity spot market boundary data provided by another embodiment of the present invention;
[0045] Figure 3 This is a heat map of step S320 in a method for generating electricity spot market boundary data provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0046] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0047] It should be understood that in the description of the embodiments of the present invention, "multiple" (or multiple) means more than two, "greater than," "less than," and "exceed" are understood to exclude the number itself, and "above," "below," and "within" are understood to include the number itself. The terms "first," "second," and so on are used solely to distinguish technical features and are not to be construed as indicating or implying relative importance, or implicitly indicating the number of the indicated technical features, or implicitly indicating the order of the indicated technical features.
[0048] like Figure 1 As shown, an embodiment of the present invention provides a method for generating boundary data of an electricity spot market, comprising:
[0049] Step S100: obtaining an initial feature set from historical data;
[0050] Step S200: obtaining a first feature set and a second feature set according to the initial feature set;
[0051] Step S300: Obtain a third feature set and a fourth feature set based on the first feature set and the second feature set respectively;
[0052] Step S400: Obtain a fourth feature set, a fifth feature set, and a sixth feature set based on the third feature set;
[0053] Step S500: Sample the fourth feature set, the fifth feature set, and the sixth feature set according to the second feature set to obtain a boundary data set.
[0054] It should be noted that the first feature set includes dynamic features, which refer to a set of dynamic features that change in the short term, such as system load fluctuations, unit energy quotation trends, and changes in reserve capacity; the second feature set includes static features, which refer to a set of static features that are fixed in the short term, such as grid topology, engine rated power, and interconnection line rated capacity.
[0055] According to the characteristics of the historical data of the electricity spot market, the data is classified to obtain a first feature set including dynamic features and a second feature set including static features. The first feature set and the second feature set are then processed separately, and the complex market characteristics are characterized by combining the feature type and historical data, so as to generate electricity spot market boundary data and improve data accuracy.
[0056] In one embodiment, in step S300, obtaining a third feature set and a fourth feature set based on the first feature set and the second feature set respectively includes:
[0057] Step S310: Obtain a third feature set and a fourth feature set based on the first feature set and the second feature set respectively;
[0058] Step S320: Remove strongly correlated feature pairs from the third feature set and the fourth feature set, and update the third feature set and the fourth feature set.
[0059] The strongly correlated feature pairs in the third feature set and the fourth feature set are removed to reduce the redundancy of data generation. The third feature set and the fourth feature set are dynamic core feature sets and static core feature sets respectively, which clarify the input basis of subsequent feature analysis, distribution characterization and sampling processes.
[0060] In one embodiment, in step S310, obtaining a third feature set and a fourth feature set based on the first feature set and the second feature set respectively includes:
[0061] According to the first feature set and the target variable, obtaining a first importance of each feature in the first feature set to the target variable;
[0062] According to the second feature set and the target variable, obtaining the second importance of each feature in the second feature set to the target variable;
[0063] According to the first importance, the third feature set is obtained;
[0064] According to the second importance, the fourth feature set is obtained.
[0065] In one embodiment, the target variable is the average clearing calculation time of the safety constrained unit commitment (SCUC).
[0066] In one embodiment, obtaining, based on the first feature set and the target variable, a first importance of each feature in the first feature set to the target variable includes:
[0067] Discretize each feature in the first feature set according to the calculation time of each feature in the first feature set to obtain a first discrete set;
[0068] The first discrete set and the target variable are input into the first model, and the result output by the first model is the first importance.
[0069] In one embodiment, the first model is a random forest (RF) model. 70% of the data set in the first feature set and the target variable is used as a training set, and 30% of the data set is used as a test set.
[0070] like Figure 2 As shown, in one embodiment, discretizing each feature in the first feature set includes: dividing the calculation time range of each feature in the first feature set into three categories: "short", "medium" and "long", with labels of 0, 1 and 2.
[0071] It should be noted that the first feature set and target variable within 3 months are taken. Different time periods during this period correspond to different first feature sets and target variables. The feature calculation time in the first feature set is discretized, and the optimal classification is voted out through the first model. The optimal classification is the type of feature that has a greater impact on the target variable (that is, the feature whose first importance exceeds a specific proportion).
[0072] In one embodiment, obtaining the third feature set according to the first importance includes:
[0073] According to the first importance, features whose importance to the target variable exceeds a specific ratio are screened out from the first feature set, and the screened feature set is the third feature set.
[0074] In one embodiment, the specific ratio is 95%.The third feature set is a core feature set that has a greater impact on the target variable.
[0075] In one embodiment, obtaining, based on the second feature set and the target variable, the second importance of each feature in the second feature set to the target variable includes:
[0076] Discretize each feature in the second feature set according to the calculation time of each feature in the second feature set to obtain a second discrete set;
[0077] The second discrete set and the target variable are input into the second model, and the result output by the second model is the second importance.
[0078] In one embodiment, the second model is a random forest (RF) model. 70% of the data set in the second feature set and the target variable is used as a training set, and 30% of the data set is used as a test set.
[0079] like Figure 2 As shown, in one embodiment, discretizing each feature in the second feature set includes: dividing the calculation time range of each feature in the second feature set into three categories: "short", "medium" and "long", with labels of 0, 1 and 2.
[0080] It should be noted that the second feature set and target variable within 3 months are taken. Different time periods during this period correspond to different second feature sets and target variables. The feature calculation time in the second feature set is discretized, and the optimal classification is voted out through the second model. The optimal classification is the two types of features that have a greater impact on the target variable (that is, the features whose second importance exceeds a specific proportion).
[0081] In one embodiment, obtaining the fourth feature set according to the second importance includes:
[0082] According to the second importance, features whose importance to the target variable exceeds a specific ratio are screened out from the second feature set, and the screened feature set is the fourth feature set.
[0083] In one embodiment, the specific ratio is 95%.The fourth feature set is a core feature set that has a greater impact on the target variable.
[0084] like Figure 2 As shown, in one embodiment, in step S310, obtaining a third feature set and a fourth feature set based on the first feature set and the second feature set respectively includes:
[0085] Input the candidate feature sets (first feature set and second feature set) of the boundary data in the past three months and the original sample dataset of the target variable. Use 70% of the dataset as the training set and 30% of the dataset as the test set.
[0086] Build a random forest model based on Python sklearn and set relevant parameters such as the number of decision trees and maximum depth. Loop through the sample data after each split, train the model, and evaluate the model based on the F1-Score value.
[0087] After training the model, the feature weights are output, and the average of multiple training results is calculated to reduce model fluctuations, and finally the feature importance ranking is determined.
[0088] Use the feature importance cumulative threshold method: select features whose cumulative contribution to the target variable exceeds a certain proportion (such as 95%) as core features, and integrate them into the third feature set and the fourth feature set respectively.
[0089] In one embodiment, in step S320, the strongly correlated feature pairs in the third feature set and the fourth feature set are respectively eliminated, and updating the third feature set and the fourth feature set includes:
[0090] The correlation between each feature in the third feature set and the fourth feature set is calculated using the following formulas:
[0091]
[0092] Where n is the number of data points, d i is the difference between the ranks of the two features of the i-th data point; the data point is the feature currently being calculated;
[0093] According to the calculated correlation and strong correlation threshold, strongly correlated feature pairs in the third feature set and the fourth feature set are eliminated.
[0094] Exemplarily, the third feature set includes feature A, feature B, and feature C. Calculating the correlation requires calculating the correlations between A and B, A and C, and B and C, respectively.
[0095] In one embodiment, the strong correlation threshold is 0.8.
[0096] In one embodiment, when there are strongly correlated pairs of features, based on the heat map, the features in the lower left portion of the heat map are retained.
[0097] like Figure 3As shown, in one embodiment, in step S320, the third feature set includes feature A, feature B, feature C, and feature D. Green indicates weak correlation, and orange indicates strong correlation. Looking only at the lower left half of the heat map and only at the features on the horizontal axis, BA is strongly correlated, so A is deleted from the feature set. DC is strongly correlated, so C is deleted from the feature set. Finally, only B and D are left.
[0098] In one embodiment, in step S400, obtaining the fourth feature set, the fifth feature set, and the sixth feature set according to the third feature set includes:
[0099] The third feature set is classified according to the value type of each feature in the third feature set to obtain a fourth feature set, a fifth feature set, and a sixth feature set.
[0100] In one embodiment, the fourth feature set includes discrete features, the fifth feature set includes continuous features, and the sixth feature set includes features that cannot be described by conventional distribution functions.
[0101] In one embodiment, the fitting distribution of the fourth feature set adopts a binomial distribution model or a multinomial distribution model.
[0102] In one embodiment, a statistical test (QQ plot) is applied to the fitted distribution of the fifth feature set.
[0103] In one embodiment, the fitted distribution for the sixth feature set is calculated using the following formula:
[0104]
[0105] in, is the density estimate at feature point x, n is the total number of feature points, h is the bandwidth (preset smoothing parameter), which controls the range of the kernel function, K is the kernel function used to assign weights (such as Gaussian kernel, triangular kernel, etc.), x i is the i-th feature point.
[0106] Adaptive kernel functions such as Gaussian kernel and triangular kernel are used to estimate the probability density, and the smoothing parameter (bandwidth) is automatically adjusted to ensure the smoothness and accuracy of the fitting results.
[0107] This embodiment constructs a smooth density curve by taking a local weighted average of data points, and is applicable to data with complex or non-standard distributions.
[0108] Feature distributions reflect how variables change over time, across regions, and under different scenarios, helping to identify regularities, anomalies, and underlying trends in the data. Furthermore, preserving the true distribution of features helps the generated data more closely resemble actual market conditions, thereby improving its ability to simulate complex market behavior.
[0109] In one embodiment, obtaining a boundary data set according to the second feature set and sampling the fourth feature set, the fifth feature set, and the sixth feature set respectively includes:
[0110] According to the fourth feature set, the fifth feature set and the sixth feature set, sampling and integrating them respectively obtain a first boundary data set;
[0111] Obtain a second boundary data set based on the first boundary data set and the second feature set;
[0112] A boundary dataset is obtained according to the second boundary dataset.
[0113] Different sampling methods are used for the fourth feature set, the fifth feature set and the sixth feature set of different feature types to obtain a first boundary data set, which is a preliminary dynamic boundary data set.
[0114] In one embodiment, sampling and integrating the fourth feature set, the fifth feature set, and the sixth feature set to obtain the first boundary data set includes:
[0115] For the discrete features included in the fourth feature set, a generating function with a disturbance term ε is obtained according to the probability parameter of the distribution, thereby obtaining the sampling value;
[0116] For the continuous features in the fifth feature set that are distributed within a fixed interval, construct the upper and lower limits a and b of the interval based on their distribution, redefine the upper and lower limits of the interval as a′=a+ε, b′=b-ε, and construct a generating function based on the upper and lower limits of the interval to obtain the sampling value;
[0117] For the continuous features in the fifth feature set whose distribution parameters are mean μ and standard deviation σ, generate the coefficient of variation CV:
[0118]
[0119] The mean μ is obtained according to the mean value of the feature, which is redefined as μ′=μ(1±CV), and a distribution function with a disturbance term ε is constructed to obtain the sampling value. The obtained sampling values have different volatility.
[0120] For the features in the sixth feature set that cannot be described by conventional distribution functions, the sampling values are obtained by the following steps:
[0121] According to the number of grid points and discretization step size of the feature map, the distribution function is obtained;
[0122] According to the inverse function of the distribution function and the random number that conforms to the uniform distribution, the preliminary sampling value is obtained;
[0123] The sampling value is obtained according to the preliminary sampling value, the distribution function, and the random number that conforms to the uniform distribution.
[0124] It is easy to understand that the feature map is the distribution map of the features in the sixth feature set, and the discretization step size is the discrete step size of the feature map;
[0125] Through precise sampling of characteristic values, we can capture significant data points such as extreme values, peaks, and low-frequency events, thereby preserving the representativeness and integrity of the data and achieving full coverage of the scenario.
[0126] In one embodiment, the distribution function is obtained according to the number of grid points and the discretization step size of the feature map, including:
[0127] The distribution function is as follows:
[0128]
[0129] Where m is the number of grid points and Δx is the discretization step size.
[0130] In one embodiment, obtaining the preliminary sampling value according to the inverse function of the distribution function and the random number conforming to the uniform distribution includes:
[0131] The preliminary sampling values are as follows:
[0132]
[0133] Among them, u i is a random number from uniform distribution U to Uniform(0,1), that is, a uniform random number. is the inverse function of the distribution function, indicating that the uniform random number u is obtained through cumulative distribution i The corresponding sample value.
[0134] In one embodiment, obtaining the sampling value according to the preliminary sampling value, the distribution function, and the random number conforming to the uniform distribution includes:
[0135] By numerically interpolating the function I, we can approximate the distribution function so that At the same time, the disturbance term ε is added to obtain the sampling value:
[0136]
[0137] In one embodiment, obtaining the second boundary dataset according to the first boundary dataset and the second feature set includes:
[0138] Obtaining a first related data set according to each boundary data in the first boundary data set;
[0139] Obtaining a second related data set according to each boundary data in the second feature set;
[0140] A second boundary data set is obtained according to the first related data set, the second related data set, the first boundary data set and the second feature set.
[0141] In one embodiment, obtaining the first related data set according to each boundary data in the first boundary data set includes:
[0142] Determine whether each boundary data in the first boundary data set has a strong correlation feature. If not, do not process it. If so, generate its strong correlation feature Y based on the boundary data. The specific formula is as follows:
[0143] Y=βX+ε
[0144]
[0145] Where β is the preset regression coefficient, ε is the disturbance term, Var(X) is the variance of feature X, and Cov(X,Y) is feature X and feature Y.
[0146] It should be noted that the disturbance term is dynamically adjusted by technical personnel based on subsequent feedback results and actual conditions.
[0147] In the feature analysis, strongly correlated features are removed to reduce data redundancy. In the final generation step, the corresponding boundary data is supplemented based on the strong correlation, which reduces data redundancy and speeds up the generation speed to adapt to the electricity spot market with rapidly changing data.
[0148] In one embodiment, obtaining the boundary dataset according to the second boundary dataset includes:
[0149] A boundary data set is obtained according to each boundary data in the second boundary data set and its threshold value, wherein the threshold value is a preset threshold value for each boundary data in the second boundary data set.
[0150] Based on electricity market rules (such as power balance and generator unit operating restrictions), we remove or adjust unreasonable boundary data, such as load factors exceeding physical limits or negative electricity prices. In addition, we set heuristic rules based on expert knowledge to correct the upper and lower limits of feature values.
[0151] In one embodiment, obtaining the boundary dataset according to the second boundary dataset further includes:
[0152] The second boundary dataset is fed into the market simulation test, and the second boundary dataset is further modified based on the output results to obtain the boundary dataset. The boundary data is fed into the market simulation test to verify its applicability for market optimization. Based on the simulation feedback, the rules and sampling methods are modified to further improve the quality and usability of the boundary data.
[0153] In one embodiment, obtaining the second boundary dataset according to the first related dataset, the second related dataset, the first boundary dataset, and the second feature set includes:
[0154] Each feature value in the first related data set, the second related data set, the first boundary data set, and the second feature set is used as the feature value in the second boundary data set.
[0155] It should be noted that the first relevant data set, the second relevant data set, the first boundary data set and the second feature set respectively include different feature types and encompass the feature types of electricity spot market boundary data.
[0156] An embodiment of the present invention further provides an electronic device, which includes but is not limited to:
[0157] Memory, used to store programs;
[0158] The processor is used to execute the program stored in the memory. When the processor executes the program stored in the memory, the processor is used to execute the above-mentioned method for generating boundary data of the electricity spot market.
[0159] The processor and the memory may be connected via a bus or other means.
[0160] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs, such as the method described in the embodiments of the present invention. The processor implements the above method by executing the non-transitory software programs and instructions stored in the memory.
[0161] The memory may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store and execute the above method. In addition, the memory may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory may optionally include a memory remotely located relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0162] The non-transitory software program and instructions required to implement the above-mentioned terminal selection method are stored in the memory, and when executed by one or more processors, the above-mentioned method is executed.
[0163] An embodiment of the present invention further provides a storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to execute the above method.
[0164] In one embodiment, the storage medium stores computer-executable instructions that are executed by one or more control processors.
[0165] The embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the embodiments.
[0166] Those skilled in the art will appreciate that all or some of the steps and systems in the method disclosed above can be implemented as software, firmware, hardware, and appropriate combinations thereof. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, and the computer-readable medium can include computer storage media (or non-transitory media) and communication media (or temporary media). As known to those skilled in the art, the term computer storage media is included in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data) and is volatile and non-volatile, removable, and non-removable. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory, or other memory technology, CD-ROM, digital versatile disks (DVD), or other optical disk storage, magnetic cassettes, magnetic tapes, disk storage, or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, as is well known to those skilled in the art, communication media typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0167] Embodiments of the present invention are described herein, including preferred embodiments known to the inventor for performing the present invention. After reading the above description, variations of these described embodiments will become apparent to those skilled in the art. The inventors expect that the skilled person will adopt such variations as appropriate, and the inventors intend to practice the embodiments of the present invention in a manner different from that specifically described herein. Therefore, as permitted by applicable law, the scope of the present invention includes all modifications and equivalents of the subject matter recited in the claims appended hereto. In addition, the scope of the present invention encompasses any combination of the above-mentioned elements in all possible variations thereof, unless otherwise indicated herein or otherwise clearly contradicted by the context.
Claims
1. A method for generating boundary data of electricity spot market, characterized in that: include: Obtain the initial feature set from historical data; Obtaining a first feature set and a second feature set according to the initial feature set; Obtaining a third feature set and a fourth feature set according to the first feature set and the second feature set respectively; According to the third feature set, a fourth feature set, a fifth feature set and a sixth feature set are obtained; The fourth feature set, the fifth feature set, and the sixth feature set are sampled according to the second feature set to obtain a boundary data set.
2. The method for generating electricity spot market boundary data according to claim 1, characterized in that: The obtaining of a third feature set and a fourth feature set according to the first feature set and the second feature set respectively includes: Obtaining a third feature set and a fourth feature set according to the first feature set and the second feature set respectively; Strongly correlated feature pairs in the third feature set and the fourth feature set are respectively eliminated, and the third feature set and the fourth feature set are updated.
3. The method for generating electricity spot market boundary data according to claim 2, characterized in that: The obtaining of a third feature set and a fourth feature set according to the first feature set and the second feature set respectively includes: Obtaining, based on the first feature set and the target variable, a first importance of each feature in the first feature set to the target variable; Obtaining, based on the second feature set and the target variable, a second importance of each feature in the second feature set to the target variable; Obtaining the third feature set according to the first importance; The fourth feature set is obtained according to the second importance.
4. The method for generating electricity spot market boundary data according to claim 1, characterized in that: The obtaining of the fourth feature set, the fifth feature set, and the sixth feature set according to the third feature set includes: The third feature set is classified according to the value type of each feature in the third feature set to obtain the fourth feature set, the fifth feature set, and the sixth feature set.
5. The method for generating electricity spot market boundary data according to claim 1, characterized in that: The obtaining of the boundary data set according to the second feature set and sampling the fourth feature set, the fifth feature set, and the sixth feature set respectively includes: According to the fourth feature set, the fifth feature set, and the sixth feature set, sampling and integrating them respectively obtain a first boundary data set; Obtain a second boundary dataset based on the first boundary dataset and the second feature set; A boundary dataset is obtained according to the second boundary dataset.
6. The method for generating electricity spot market boundary data according to claim 5, characterized in that: The obtaining of a second boundary dataset according to the first boundary dataset and the second feature set includes: Obtaining a first related data set according to each boundary data in the first boundary data set; Obtaining a second related data set according to each boundary data in the second feature set; The first related data set, the second related data set, the first boundary data set and the second feature set are combined to obtain a second boundary data set.
7. The method for generating electricity spot market boundary data according to claim 5, characterized in that: The obtaining of the boundary dataset according to the second boundary dataset comprises: A boundary data set is obtained according to each boundary data in the second boundary data set and its threshold value, wherein the threshold value is a preset threshold value for each boundary data in the second boundary data set.
8. The method for generating electricity spot market boundary data according to claim 5, characterized in that: For the features in the sixth feature set, the sample values are obtained by the following method: A distribution function is obtained according to the number of grid points and the discretization step size of the feature map; wherein the feature map is a distribution map of features in the sixth feature set, and the discretization step size is a discrete step size of the feature map; Obtaining a preliminary sampling value according to the inverse function of the distribution function and a random number that conforms to a uniform distribution; A sampling value is obtained according to the preliminary sampling value, the distribution function, and the random number, and the sampling value is integrated into the first boundary data set.
9. An electronic device, characterized in that: include: Memory, used to store programs; A processor, configured to execute the program stored in the memory. When the processor executes the program stored in the memory, the processor is configured to execute the method according to any one of claims 1 to 8.
10. A storage medium, characterized in that: Computer-executable instructions are stored, and the computer-executable instructions are used to execute the method according to any one of claims 1 to 8.