Method, medium and device for generating a residual demand curve based on a neural network
Patent Information
- Application Number
- CN202611008306.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-09-25
AI Technical Summary
第一,不确定性量化机制存在本质性割裂
[0022]第三方面,本申请还提供了一种终端设备,包括存储器和处理器;所述存储器存储有可被处理器执行的程序代码;所述程序代码用于执行第一方面任意一项所述的基于神经网络的剩余需求曲线生成方法。
Smart Images

Figure CN122819792A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power system automation technology, and in particular to a method, medium, and device for generating residual demand curves based on neural networks. Background Technology
[0002] With the continuous deepening of power system optimization dispatch, large power plants are undergoing a fundamental shift from planned dispatch mode to optimized operation mode. In the optimized operation environment, power plants, as recipients or limited influencers of nodal marginal costs, directly determine their operational efficiency levels through their day-ahead and real-time output planning decisions. However, system nodal marginal costs are significantly influenced by multiple uncertainties, including system load fluctuations, intermittent renewable energy output, transmission network congestion, and the operational strategies of other grid-connected entities, exhibiting significant stochasticity and non-stationarity. Accurately characterizing and quantifying these uncertainties has become a core technical bottleneck for power plants in formulating robust operation strategies.
[0003] The Residual Demand Curve (RDC), a key tool for analyzing the operational status and optimizing power output strategies of power plants, is defined as the functional mapping between the operating power of a power plant and the corresponding nodal marginal cost under a specific system operating environment. This curve not only reflects the potential nodal marginal cost levels a power plant may face at different operating power levels, but also contains structural information about the system's supply and demand dynamics. Therefore, accurate prediction of the RDC, especially the effective quantification of its uncertainty range, is of crucial engineering value for power plants in assessing the distribution of operational risks under different power output strategies and achieving a balance between risk and operational efficiency.
[0004] Existing technologies generally employ a two-stage paradigm of "point prediction + post-processing to generate scenarios" when dealing with residual demand curve forecasting and uncertainty quantification. Specifically, the first stage uses traditional neural networks (such as Long Short-Term Memory networks (LSTM) and Convolutional Neural Networks (CNN)) to perform deterministic point predictions of node marginal costs; the second stage, based on statistical distribution assumptions of the prediction error (such as normal distribution, Laplace distribution, etc.), generates scenarios with uncertain node marginal costs through random sampling or Monte Carlo simulation. However, this paradigm has the following technical drawbacks: First, there is a fundamental disconnect in the uncertainty quantification mechanism. Point prediction models only output a single expected value and cannot intrinsically characterize the heteroscedasticity and quantile structure of the prediction condition distribution. This forces subsequent scenario generation to rely on subjectively pre-defined error distribution patterns. When the actual error distribution exhibits heavy tails, skewness, or multimodal characteristics, scenario generation based on parametric assumptions will produce systematic biases and fail to accurately reflect extreme operational risk scenarios of the system.
[0005] Second, the two-stage process leads to error propagation and accumulation. The residual distribution estimation of the point prediction model is independent of the prediction process. The prediction error and the scene generation error are propagated and superimposed in the serial structure, which may amplify the degree of deviation in the uncertainty representation and reduce the probability and reliability of the scene set in covering the real system operating state.
[0006] Third, it lacks the ability to model long-sequence dependencies. Traditional recurrent neural networks suffer from gradient vanishing or gradient exploding problems when processing node marginal cost sequences with a time scale of 24 hours or more for day-ahead scheduling, making it difficult to effectively capture the periodic dependence and dynamic correlation between load and renewable energy output in long-sequence dimensions.
[0007] Fourth, there is a lack of explicit correlation between scenario structure and operating power. Existing technologies typically generate node marginal cost scenarios as unconditional marginal distribution samples, without establishing a conditional and structured mapping relationship with specific operating power thresholds of the power station. This leads to difficulties in adaptation and optimization deviations when the power station inputs the scenario into the output optimization model, making it difficult to accurately control the operational risk exposure.
[0008] In summary, existing technologies suffer from key technical problems in generating residual demand curves for uncertain scenarios, such as the disconnect between prediction and quantification, strong dependence on distribution assumptions, limited long-sequence modeling capabilities, and insufficient scenario structuring. There is an urgent need for a new method that can directly, synchronously, and without distribution assumptions generate residual demand curves for different scenarios corresponding to operating power thresholds, in order to support power plants in making refined output decisions and managing operational risks under optimized operating environments. Summary of the Invention
[0009] Based on this, the purpose of this application is to provide a method, medium, and device for generating residual demand curves based on neural networks, so as to solve at least one of the technical problems mentioned in the background art.
[0010] In a first aspect, this application provides a method for generating residual demand curves based on neural networks, including: Historical power grid operation data for each moment within a set time period are collected, including wind speed, temperature, regional load forecast, operating power, hourly data and weekday data, as well as the node marginal cost at the corresponding moment, in order to construct a training dataset. A marginal cost prediction model is constructed and trained based on the training dataset, taking time-series sample sequences as input and node marginal costs as output, including: The input layer is used to receive time-series sample sequences; The encoding embedding layer, connected to the input layer, is used to perform linear projection on the temporal sample sequence and add position encoding to inject temporal values, resulting in a temporal encoded sequence. The encoder stack, connected to the encoding embedding layer, is used to perform multi-layer attention extraction and sequence compression on the temporal encoded sequence to obtain the encoded representation; The sequence summarization layer, connected to the encoder stack, is used to perform average pooling on the encoded representation along the time dimension to obtain the global representation vector; The output layer, connected to the sequence summarization layer, is used to map the global representation vector to the predicted values of multiple preset quantiles, and output the predicted values of the node marginal cost at different quantile levels. Obtain the operating power range of the power plant and discretize it to obtain several operating power threshold points; collect the predicted grid operation data t times before the prediction date, and replace the operating power t-1 times before the prediction date with the operating power t-1 times before the prediction date. Then, replace the operating power at the t-th time of the prediction date with each operating power threshold point in turn to construct several time series sample sequences. Input each time series sample sequence into the trained marginal cost prediction model to obtain the predicted value of the node marginal cost at different quantile levels corresponding to each operating power threshold point at time t of the prediction day. Then, determine the scenario probability weight according to the quantile interval corresponding to each quantile level to generate the residual demand curve under different scenarios at time t.
[0011] Furthermore, the encoding embedding layer includes linear projection units and position encoding units connected in sequence.
[0012] The linear projection unit is used to map the feature vector of each time step in the time series sample sequence to the model dimension through a linear transformation, so as to obtain the projected feature sequence. The position coding unit, connected to the linear projection unit, is used to add position coding information based on sine and cosine functions to each time step in the projected feature sequence in order to preserve the time order relationship and obtain the time-series coded sequence.
[0013] Furthermore, the encoder stack includes several layers of encoder layers and distillation layers stacked sequentially; Each encoder layer comprises a probabilistic sparse self-attention unit and a feedforward network unit connected in sequence. The probabilistic sparse self-attention unit receives the input temporal coding sequence of the current encoder layer, calculates the query matrix, key matrix, and value matrix, filters the core query vector of the power grid based on the probabilistic sparse metric, performs scaled dot product attention calculation only on the filtered queries, and outputs a sparse attention feature sequence. The feedforward network unit, connected to the probabilistic sparse self-attention unit, performs feedforward calculation and layer normalization operation on the sparse attention feature sequence to obtain the output temporal coding sequence of the current encoder layer. The output temporal coding sequence of the last encoder layer is the encoded representation. The distillation layer is set after the encoder layer whose output timing coding sequence is longer than a set sequence length. It is used to compress the output timing coding sequence of the current encoder layer to obtain the distilled output timing coding sequence.
[0014] Furthermore, the probabilistic sparse self-attention unit comprises a query key-value projection element, a sparsity measurement element, a query filtering element, and a sparse attention calculation element connected in sequence: The query key-value projection element is used to perform linear projections on the input time-series encoded sequence representation to obtain the query matrix, key matrix, and value matrix; The sparsity metric element, connected to the query key projection element, is used to calculate the sparsity score between each query vector and all key vectors, resulting in a sparsity metric sequence. The query filtering element is connected to the sparsity metric element to retain the query vectors with the highest sparsity scores in the sparsity metric sequence, thus obtaining the filtered query submatrix. The sparse attention computation element, connected to the query key-value projection element and the query filtering element, is used to perform scaled dot product attention computation on the filtered query submatrix and all key matrices and value matrices to obtain the sparse attention feature sequence.
[0015] Furthermore, the feedforward network unit includes a first-layer linear transformation element, a modified linear activation element, a second-layer linear transformation element, and a layer normalization element connected in sequence: The first layer of linear transformation elements is used to perform linear transformation on the output of the probabilistic sparse self-attention unit, mapping the model dimension to the feedforward hidden dimension to obtain the first transformation result. The modified linear activation element is connected to the first-layer linear transformation element and is used to perform an element-wise nonlinear transformation on the first transformation result, setting the negative value to zero to obtain the activation result; The second-layer linear transformation element, connected to the modified linear activation element, is used to perform a linear transformation on the activation result, mapping the feedforward hidden dimension back to the model dimension to obtain the second transformation result; The first layer normalization element is connected to the second layer linear transformation element. It is used to add the second transformation result to the input of the feedforward network unit as a residual, and calculate the mean and standard deviation of the addition result step by step, and perform normalization scaling to obtain the output time-series coded sequence.
[0016] Furthermore, the distillation layer comprises one-dimensional convolutional units and max-pooling units connected in sequence.
[0017] One-dimensional convolutional units are used to extract local features and downsample the stride of the output temporal encoded sequence to obtain convolutional features. The max pooling unit, connected to the one-dimensional convolutional unit, is used to extract local maxima from the convolutional features to obtain the distilled output temporal coding sequence.
[0018] Furthermore, the output layer includes a first fully connected unit, a modified linear activation element, and a second fully connected unit connected in sequence.
[0019] The first fully connected unit is used to map the global representation vector from the model dimension to the feedforward hidden dimension to obtain the first connection result; A modified linear activation element is connected to the first fully connected unit to perform an element-wise nonlinear transformation on the first connection result to obtain the activated connection result; The second fully connected unit, connected to the modified linear activation element, is used to map the activation connection result from the feedforward hidden dimension to the quantile number dimension, so as to obtain the predicted value of the node marginal cost at different quantile levels.
[0020] Further steps in constructing the training dataset include: Obtain the average value and standard deviation of various historical power grid operation data at each moment within a set time period, and determine the fluctuation range of various historical power grid operation data based on the set fluctuation coefficient, average value and standard deviation; Determine whether the values of various historical power grid operation data are within the corresponding fluctuation range. If so, they are judged as normal values and retained. If not, they are judged as abnormal values. Linear interpolation replacement is performed on several normal values adjacent to the abnormal values to obtain several optimized historical power grid operation data. Based on the mathematical characteristics of various historical power grid operation data, they are classified into continuous data, periodic data, and categorical data; continuous data includes wind speed, temperature, regional load forecasts, and operating power; periodic data includes hourly data; and categorical data includes weekly data. Various continuous data are standardized to map to a uniform numerical scale, eliminating dimensional differences and obtaining continuous feature sequences; periodic data are encoded using sine and cosine to obtain periodic feature sequences; and categorical data are encoded using one-hot encoding to obtain categorical feature sequences. By concatenating continuous feature sequences, periodic feature sequences, and categorical feature sequences, a power grid operation feature sequence is obtained. The power grid operation feature sequence is then segmented according to a historical window of a set length to construct the time series sample sequence corresponding to each time point. The node marginal cost at the corresponding time point is then obtained to construct the training dataset.
[0021] Secondly, this application also provides a computer storage medium storing executable program code; the executable program code is used to execute the residual demand curve generation method based on neural networks as described in any one of the first aspects.
[0022] Thirdly, this application also provides a terminal device, including a memory and a processor; the memory stores program code executable by the processor; the program code is used to execute the residual demand curve generation method based on neural networks as described in any one of the first aspects.
[0023] This application provides a method, medium, and device for generating residual demand curves based on neural networks. It collects historical power grid operation data at various times within a set time period, including wind speed, temperature, regional load forecasts, operating power, hourly data, and weekday data, as well as the corresponding node marginal costs, to construct a training dataset. Based on the training dataset, it constructs and trains a marginal cost prediction model that takes time-series sample sequences as input and node marginal costs as output. The model includes: an input layer for receiving time-series sample sequences; an encoding embedding layer connected to the input layer for linearly projecting the time-series sample sequences and adding positional encoding to inject time-series values, resulting in a time-series encoded sequence; an encoder stack connected to the encoding embedding layer for performing multi-layer attention extraction and sequence compression on the time-series encoded sequence, resulting in an encoded representation; and a sequence summarization layer connected to the encoder stack for average pooling of the encoded representation along the time dimension, resulting in... A global representation vector is generated; the output layer, connected to the sequence summarization layer, maps the global representation vector to predicted values of multiple preset quantiles, and outputs predicted values of node marginal costs at different quantile levels; the operating power range of power plants is obtained and discretized to obtain several operating power threshold points; predicted grid operation data at t times before the prediction date is collected, and the operating power at t-1 times before the prediction date is replaced with the operating power at t times before the prediction date, and then the operating power at t times before the prediction date is replaced with each operating power threshold point in turn to construct several time-series sample sequences; each time-series sample sequence is input into the trained marginal cost prediction model to obtain the predicted values of node marginal costs at different quantile levels corresponding to each operating power threshold point at t times before the prediction date, and the scenario probability weights are determined according to the quantile intervals corresponding to each quantile level to generate the residual demand curves under different scenarios at t times. A new method has been developed to improve upon existing technologies that cannot directly, synchronously, and without the assumption of distribution to generate residual demand curves for different scenarios corresponding to operating power thresholds. This method supports power plants in making refined output decisions and managing operational risks under optimized operating environments. Attached Figure Description
[0024] Figure 1 This is a flowchart of an embodiment of the residual demand curve generation method based on neural networks of the present invention; Figure 2 This is a schematic diagram of the marginal cost prediction model of an embodiment of the residual demand curve generation method based on neural networks of the present invention. Figure 3This is a schematic diagram of the encoder layer structure of an embodiment of the residual demand curve generation method based on neural networks of the present invention. Figure 4 This is a schematic diagram of residual demand curves under different scenarios, representing an embodiment of the residual demand curve generation method based on neural networks of the present invention. Detailed Implementation
[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0026] It should be noted that if the embodiments of the present invention involve directional indications, such as up, down, left, right, front, back, etc., these directional indications are only used to explain the relative positional relationships and movement of the components in a specific posture. If the specific posture changes, the directional indications will also change accordingly. Furthermore, if the embodiments of the present invention involve descriptions such as "first," "second," "S1," "S2," "step one," "step two," etc., these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance, or implicitly indicating the number of technical features indicated or the execution order of the method. Those skilled in the art will understand that anything that does not violate the inventive concept should be included within the scope of protection of the present invention.
[0027] like Figure 1 As shown, this invention provides a method for generating residual demand curves based on neural networks, comprising: S1: Collect historical power grid operation data at each time point within a set time period, including wind speed, temperature, regional load forecast, operating power, hourly data and weekday data, as well as the node marginal cost at the corresponding time point, in order to construct a training dataset; Specifically, since the marginal cost of a node is significantly affected by the operating power of the power plant, and the marginal cost curve corresponding to different output levels of the operating power exhibits non-linear changes, it is possible to collect historical power grid operation data at various times within a set time period, including wind speed, temperature, regional load forecasts, operating power, hourly data, and weekday data, as well as the corresponding node marginal cost, to construct a training dataset. This allows the model to learn the historical power-cost mapping relationship and then conditionally predict the node marginal cost under different operating power scenarios, providing multi-scenario cost references for subsequent scheduling decisions.
[0028] In a preferred embodiment, the step of constructing the training dataset includes: S11: Obtain the average value and standard deviation of various historical power grid operation data at each moment within a set time period, and determine the fluctuation range of various historical power grid operation data based on the set fluctuation coefficient, average value and standard deviation; S12: Determine whether the values of various historical power grid operation data are within the corresponding fluctuation range. If so, they are judged as normal values and retained. If not, they are judged as abnormal values. Linear interpolation replacement is performed on several normal values adjacent to the abnormal values to obtain several optimized historical power grid operation data. S13: Classify various historical power grid operation data according to their mathematical characteristics to obtain continuous data, periodic data, and categorical data; continuous data includes wind speed, temperature, regional load forecasts, and operating power; periodic data includes hourly data; categorical data includes weekday data. S14: Standardize various continuous data to map them to a uniform numerical scale, eliminate dimensional differences, and obtain continuous feature sequences; perform sine and cosine coding on periodic data to obtain periodic feature sequences; perform one-hot coding on categorical data to obtain categorical feature sequences. S15: Concatenate continuous feature sequences, periodic feature sequences, and categorical feature sequences to obtain power grid operation feature sequences. Then, segment the power grid operation feature sequences according to a set length of historical window to construct time series sample sequences corresponding to each time point. Obtain the node marginal cost at the corresponding time point and construct a training dataset.
[0029] Specifically, after optionally collecting historical power grid operation data at various times within a set time period, the average and standard deviation of various historical power grid operation data can be obtained. Based on the set fluctuation coefficient, average, and standard deviation, the fluctuation range of various historical power grid operation data can be determined. An adaptive dynamic threshold boundary is constructed based on sample statistics. The detection sensitivity is adjusted by the fluctuation coefficient, allowing the threshold to automatically adjust with the distribution scale of different physical quantity data. This avoids the insufficient adaptability and misjudgment problems of fixed thresholds in multi-dimensional scenarios, improving the accuracy and flexibility of anomaly detection. Then, it is determined whether the values at each time point in various historical power grid operation data are within the corresponding fluctuation range. Outliers are replaced by linear interpolation based on several adjacent normal values to eliminate outliers. While mitigating point disturbances, the system reconstructs local variation trends of the data using adjacent normal values, maintaining the continuity and smoothness of the sequence. This eliminates anomalous jumps while preserving the overall data trend, preventing outliers from disrupting time-series correlations and misleading subsequent model training, thus improving data quality and model training reliability. Furthermore, various historical power grid operation data are classified according to mathematical characteristics, resulting in continuous, periodic, and categorical data. Prior classification based on the inherent mathematical properties of the data provides a structured basis for subsequent differentiated coding, avoiding information distortion, feature redundancy, or semantic loss caused by uniform transformation of all features, thereby improving the targeting and effectiveness of feature engineering. Finally, various continuous data are standardized to map them to a unified numerical value. The scaling mechanism eliminates dimensional differences, avoiding gradient imbalance caused by variations in numerical ranges of different features, thus accelerating model convergence and improving training stability. Simultaneously, sine and cosine encoding is applied to periodic data, mapping discrete-time indices to a continuous cyclic space. This preserves the proximity of adjacent time points within a period and the periodic topological structure, avoiding discontinuities at periodic boundaries caused by simple numerical encoding. This ensures the model accurately captures periodic patterns. One-hot encoding is then applied to categorical data to eliminate spurious order relationships and distance metric biases between categories, ensuring the semantic independence of categorical features. Through this differentiated encoding strategy, it ensures that each feature retains its original semantic characteristics after transformation, avoiding biases or information loss introduced by improper encoding, and improving the accuracy of feature representation and model stability. The input is standardized; finally, continuous feature sequences, periodic feature sequences, and categorical feature sequences are concatenated to obtain the power grid operation feature sequence. The power grid operation feature sequence is then segmented according to a historical window of a set length to construct the time series sample sequence corresponding to each time moment. Through multi-type feature fusion and fixed-length historical window segmentation, supervised learning training samples with standard input dimensions are constructed to clarify the temporal mapping relationship between input features and future states. This avoids model learning obstacles caused by non-standard construction of time series samples, inconsistent dimensions, or insufficient contextual information, improves the consistency of training data and the generalization ability of the model, and obtains the node marginal cost at the corresponding time moment. The training dataset is then constructed based on the time series sample sequence at each time moment and its corresponding node marginal cost.
[0030] Preferably, the standardization process can be carried out using the Z-score standardization method, so that the mean of each feature is 0 and the variance is 1, thereby improving the training efficiency of deep networks; the sine and cosine encoding can be carried out using a trigonometric function mapping form based on the period length T, so as to ensure that the periodic features have strict cyclic symmetry and translation invariance after encoding.
[0031] In another preferred embodiment, the step of collecting historical power grid operation data and constructing a training dataset further includes: a01) Input Feature Collection and Outlier Cleaning. Collect historical power grid operation data to construct a training sample set. Each sample corresponds to a specific time point on a historical day. A total of D days are collected, with T time points per day, for a total sample size of [number missing]. .
[0032] Input features include: wind speed Temperature Hourly data Weekly data Regional load forecast Historical operating power The output label represents the marginal cost of the node at the corresponding time point. .
[0033] Outlier detection and cleaning are performed using the 3σ criterion: For each continuous feature sequence, the mean of its historical data is calculated. and standard deviation It will exceed Data points within a certain range are considered outliers and are replaced using linear interpolation of data from the preceding and following time points.
[0034] a02) Feature encoding and standardization. For categorical data (weekday data) One-hot encoding is performed to convert it into a binary vector, resulting in a sequence of class features represented as an encoded vector. ,in This represents a 7-dimensional real vector, where each dimension corresponds to a day of the week (e.g., the first dimension represents Monday, the second dimension represents Tuesday, and so on).
[0035] For continuous data (wind speed) Temperature Regional load forecast Historical operating power Z-score standardization is performed to eliminate dimensions, resulting in a continuous feature sequence. Let the vector representation of the original continuous data be... ,in For sample index, symbol This represents the transpose of a vector. The standardized formula is shown in A1: (A1); In the formula: This is the mean vector of each continuous data point over historical power grid operation data; This is the standard deviation vector of each continuous data point over historical power grid operation data. It is a standardized continuous feature sequence.
[0036] For hourly data Since it is a periodic feature, it is not standardized. Instead, the sine and cosine coding method is used to convert it into two periodic feature sequences as shown in A2 and A3, in order to preserve its periodic value. (A2); (A3); In the formula: and Hourly characteristics The sine and cosine coding features.
[0037] The category feature sequence after one-hot encoding Standardized continuous feature sequences and periodic characteristic sequences , The final combined feature vector, i.e., the power grid operation feature sequence, is obtained by concatenating the features, as shown in A4: (A4); In the formula: This represents the total dimension of the final feature vector, which is a 13-dimensional real vector.
[0038] a03) Sample Sequence Construction. To utilize time-series values, power grid operation data from several time points prior to the current time are collected to construct an input time-series encoded sequence with a length of [missing information]. Historical windows are used to predict the marginal cost of nodes at different quantile levels at the current moment. For the th The sample at time t (corresponding to the global sample index n) is its input time series sample sequence. From the current moment and before The features at each time point are shown in A5: (A5); In the formula: Represent a OK A real matrix of columns, where For the number of prediction time steps; symbol This indicates that vectors are stacked row-wise into a matrix. The input is a time-coded sequence. The corresponding output label is the first one. The marginal cost of the node at time t. .
[0039] If some samples are not long enough, they are discarded. The final result is... A valid time-series sample sequence and its label, denoted as _ ... .
[0040] S2: Construct and train a marginal cost prediction model based on the training dataset, taking time-series sample sequences as input and node marginal costs as output. Figure 2 As shown, it includes: The input layer is used to receive time-series sample sequences; The encoding embedding layer, connected to the input layer, is used to perform linear projection on the temporal sample sequence and add position encoding to inject temporal values, resulting in a temporal encoded sequence. The encoder stack, connected to the encoding embedding layer, is used to perform multi-layer attention extraction and sequence compression on the temporal encoded sequence to obtain the encoded representation; The sequence summarization layer, connected to the encoder stack, is used to perform average pooling on the encoded representation along the time dimension to obtain the global representation vector; The output layer, connected to the sequence summarization layer, is used to map the global representation vector to predicted values of multiple preset quantiles, and outputs predicted values of node marginal costs at different quantile levels.
[0041] Specifically, the input layer can receive time-series sample sequences, including time-series features of power system operation data such as wind speed, temperature, regional load forecasts, operating power, hourly data, and weekday data. This time-series representation of the system's operating state provides a structured time-series input foundation for subsequent encoding processing, ensuring the model can capture the dynamic evolution of marginal cost over time. Then, the encoding embedding layer performs linear projection and position encoding on the time-series sample sequences, mapping the high-dimensional sparse feature vectors of each time step to a unified dimensional space of the model and injecting position encoding information based on sine and cosine functions. This preserves the temporal order while enabling the model to distinguish the relative positions of different time steps, addressing the shortcomings of traditional time-series models in modeling long-distance temporal dependencies and providing a dense vector representation with position awareness for the encoder stack. Finally, the encoder stack performs multi-layer attention extraction and sequence compression on the time-series encoded sequences, using a probabilistic sparse self-attention mechanism to filter key query vectors, performing scaling dot product attention calculations only on important time steps, significantly improving performance. While reducing computational complexity, this method accurately captures long-range dependencies and key abrupt changes in time-series data. Through layer-by-layer distillation, it progressively compresses sequence length and extracts high-level semantic information, enabling the encoded representation to reflect both local time-series patterns and global operational status, effectively filtering out noise interference irrelevant to marginal cost prediction. Furthermore, a sequence aggregation layer performs average pooling along the time dimension on the encoded representation, aggregating the variable-length time-series encoded representation into a fixed-dimensional global representation vector. This achieves balanced aggregation of information across all time periods, preventing key time-step information from being buried, while ensuring the dimensional stability of the output vector, providing a compact and information-rich global state description for the output layer. Finally, the output layer maps the global representation vector to multiple preset quantiles (e.g., 10%, 50%, 90%) of the predicted value, outputting the node marginal cost prediction results in quantile form. This not only provides point estimates but also quantifies the uncertainty range of the prediction, enabling dispatchers to comprehensively assess cost risks at different confidence levels. This provides a scientific and quantitative probabilistic basis for power plant unit combination optimization and real-time dispatch decisions.
[0042] In a preferred embodiment, the encoding embedding layer includes a linear projection unit and a position encoding unit connected in sequence.
[0043] The linear projection unit is used to map the feature vector of each time step in the time series sample sequence to the model dimension through a linear transformation, so as to obtain the projected feature sequence. The position coding unit, connected to the linear projection unit, is used to add position coding information based on sine and cosine functions to each time step in the projected feature sequence in order to preserve the time order relationship and obtain the time-series coded sequence.
[0044] Specifically, a linear projection unit can be used to map the original feature vectors (which may contain multi-dimensional information such as load, output, and node marginal cost) of each time step in the time-series sample sequence to a unified model dimension space. This achieves normalized alignment of features with different dimensions and scales, eliminates gradient imbalance caused by differences in feature dimensions, and enables the model to handle various input features in a balanced manner. Then, a positional encoding unit injects positional information based on sine and cosine functions into the projected feature sequence. This encoding method has uniqueness, boundedness, and smoothness, enabling the model to effectively distinguish the relative positional relationships of different time steps without increasing the learnable parameters. It also supports extrapolation and generalization of longer sequences not seen during training, thereby ensuring the model accurately models temporal causal relationships and providing a complete sequence representation that contains both semantic and temporal positional information for the subsequent encoder stack.
[0045] In a preferred embodiment, the encoding embedding layer is used to encode the input time-series sample sequence. The feature vector at each time step is transformed to the model dimension through linear projection. And add positional encoding to inject timing values, as shown in B1: (B1); In the formula: For linear layers, 3D feature mapping to dimension; This is a sine and cosine position encoding matrix; This is a time-series coded sequence. The positional encoding calculation formula is as follows: For position... and dimensions and The position codes are defined as B2 and B3: (B2); (B3); In the formula: Represents the position index in the sequence (in this invention, from 0 to...). ); Indicates the dimension index (the value ranges from 0 to...). ).
[0046] In a preferred embodiment, the encoder stack includes several layers of encoder layers and distillation layers stacked sequentially; Each encoder layer includes sequentially connected probabilistic sparse self-attention units and feedforward network units, such as... Figure 3As shown: The probabilistic sparse self-attention unit receives the input temporal coding sequence of the current encoder layer, calculates the query matrix, key matrix, and value matrix, filters the core query vector of the power grid based on the probabilistic sparse metric, performs scaling dot product attention calculation only on the filtered queries, and outputs a sparse attention feature sequence; the feedforward network unit, connected to the probabilistic sparse self-attention unit, performs feedforward calculation and layer normalization operation on the sparse attention feature sequence to obtain the output temporal coding sequence of the current encoder layer; the output temporal coding sequence of the last encoder layer is the encoded representation; The distillation layer is set after the encoder layer whose output timing coding sequence is longer than a set sequence length. It is used to compress the output timing coding sequence of the current encoder layer to obtain the distilled output timing coding sequence.
[0047] Specifically, the output temporal coding sequence of each encoder layer becomes the input temporal coding sequence of the next encoder layer, and the output temporal coding sequence of the last encoder layer becomes the encoded representation. A deep encoder network can be constructed by stacking probabilistic sparse self-attention units and feedforward network units hierarchically. This allows the model to extract multi-scale temporal feature representations of the input temporal coding sequence from local to global levels layer by layer. The probabilistic sparse self-attention units at each layer select the core query vector of the power grid based on a probabilistic sparsity metric, effectively suppressing the interference of redundant information on the attention distribution, allowing the deep network to maintain computational controllability while expanding the receptive field. Furthermore, the feedforward network units perform nonlinear transformations and layer normalization on the sparse attention feature sequence, increasing... The strong feature representation capability and training stability of the model promote efficient information transfer across layers. Then, the distillation layer actively compresses the length of the temporal encoded sequence when it exceeds a set threshold, mapping the high-dimensional, long temporal encoded sequence into a low-dimensional, compact feature sequence, which significantly reduces the computational complexity and memory consumption of subsequent layers, while preserving the semantic information of key temporal patterns. Finally, through the alternating configuration of the encoder layer and the distillation layer, a progressive processing flow of "sparse attention-feature transformation-sequence compression" is realized. While maintaining the ability to model long-distance dependencies, the marginal cost prediction model has the ability to efficiently encode and hierarchically abstract the time series data of ultra-large-scale power systems, taking into account both the requirements of prediction accuracy and inference efficiency.
[0048] In a preferred embodiment, the probabilistic sparse self-attention unit includes a query key-value projection element, a sparsity measurement element, a query filtering element, and a sparse attention calculation element connected in sequence: The query key-value projection element is used to perform linear projections on the input time-series encoded sequence representation to obtain the query matrix, key matrix, and value matrix; The sparsity metric element, connected to the query key projection element, is used to calculate the sparsity score between each query vector and all key vectors, resulting in a sparsity metric sequence. The query filtering element is connected to the sparsity metric element to retain the query vectors with the highest sparsity scores in the sparsity metric sequence, thus obtaining the filtered query submatrix. The sparse attention computation element, connected to the query key-value projection element and the query filtering element, is used to perform scaled dot product attention computation on the filtered query submatrix and all key matrices and value matrices to obtain the sparse attention feature sequence.
[0049] Specifically, the input time-series encoded sequence can be mapped to three different feature subspaces—query, key, and value—using a query key-value projection element. This allows the model to examine time-series data from different perspectives: the query space is used to locate the time steps of interest, the key space provides reference information for the focus, and the value space carries the actual transmitted feature content. These three spaces have clear divisions of labor and work collaboratively. Then, a sparsity metric element calculates the sparsity score between each query vector and all key vectors. This score quantifies the "importance" or "representativeness" of the query vector to the attention distribution, identifying the key queries that contribute the most to the overall attention pattern. Next, a query filtering element performs Top-K filtering based on the sparsity score, retaining only the highest-scoring query vectors for subsequent attention calculations, effectively reducing computational load. Finally, a sparse attention calculation element performs scaled dot-product attention on the filtered query submatrix and all key and value matrices. This reduces computational overhead while ensuring the model can still effectively model long-distance time dependencies, enabling the marginal cost prediction model to handle large-scale, long-series power system operation data.
[0050] In a preferred embodiment, the encoder stack consists of The encoder layers are stacked, where the ProbSparse self-attention is considered, and the th encoder layer is calculated. ( The output sequence of layer B4 is shown below: (B4); In the formula: For the first The output temporal encoded sequence of the layer encoder layer; To reduce computational complexity for probabilistically sparse self-attention computation; For layer normalization, the feature vector at each time step is standardized to stabilize network training, as shown in formula B5: (B5); in: , For vectors The mean and standard deviation, , For learnable scaling and offset parameters, This is element-wise multiplication.
[0051] Detailed calculation process of ProbSparse self-attention: For the input timing encoded sequence of the current encoder layer First, calculate the query matrix. Key matrix Value matrix ,in The weight matrix is learnable. ProbSparse self-attention reduces computation by filtering out important queries. First, the query matrix is defined. The sparsity measure is shown in B6: (B6); In the formula: This is the i-th query vector; Let j be the j-th key vector. Then, only the vector with the highest sparsity metric is selected. Query vectors ( ,in (A constant) is used in attention calculation. The set of indexes corresponding to each query vector is denoted as . .
[0052] Finally, the output of ProbSparse self-attention is shown in B7: (B7); In the formula: From The index selected in A matrix composed of rows in the matrix.
[0053] In a preferred embodiment, the feedforward network unit includes a first-layer linear transformation element, a modified linear activation element, a second-layer linear transformation element, and a layer normalization element connected in sequence. The first layer of linear transformation elements is used to perform linear transformation on the output of the probabilistic sparse self-attention unit, mapping the model dimension to the feedforward hidden dimension to obtain the first transformation result. The modified linear activation element is connected to the first-layer linear transformation element and is used to perform an element-wise nonlinear transformation on the first transformation result, setting the negative value to zero to obtain the activation result; The second-layer linear transformation element, connected to the modified linear activation element, is used to perform a linear transformation on the activation result, mapping the feedforward hidden dimension back to the model dimension to obtain the second transformation result; The first layer normalization element is connected to the second layer linear transformation element. It is used to add the second transformation result to the input of the feedforward network unit as a residual, and calculate the mean and standard deviation of the addition result step by step, and perform normalization scaling to obtain the output time-series coded sequence.
[0054] Specifically, a feedforward network unit can be used to perform nonlinear transformation and dimension mapping on the sparse attention feature sequence. The first-layer linear transformation expands the feature dimension to improve the model's expressive power. The Rectified Linear Unit (ReLU) introduces nonlinear activation to capture complex feature interaction patterns. The second-layer linear transformation maps the features back to the original dimension to ensure the dimensionality consistency of the residual connections. Finally, the layer normalization element normalizes the sum of the residuals to stabilize the output distribution of each layer, effectively alleviate the gradient vanishing and gradient exploding problems in deep networks, accelerate model convergence, and enable the encoder stack to maintain stable training dynamics and excellent feature extraction performance even with deep stacking.
[0055] In a preferred embodiment, the output sequence of the probabilistic sparse self-attention unit further needs to undergo feedforward computation and layer normalization operations through a feedforward network unit, as shown in B8 and B9: (B8); (B9); In the formula: For feedforward network units; To modify the activation function of the linear unit, a nonlinear transformation is introduced; These are learnable network weights or bias parameters.
[0056] In a preferred embodiment, the distillation layer includes a one-dimensional convolutional unit and a max pooling unit connected in sequence.
[0057] One-dimensional convolutional units are used to extract local features and downsample the stride of the output temporal encoded sequence to obtain convolutional features. The max pooling unit, connected to the one-dimensional convolutional unit, is used to extract local maxima from the convolutional features to obtain the distilled output temporal coding sequence.
[0058] Specifically, one-dimensional convolutional units can be used to extract local features and downsample the stride of the output temporal encoded sequence. Utilizing the local receptive field of the convolutional kernel, while preserving local temporal pattern information, the sequence length is halved through stride operations, achieving initial compression of the feature dimension. Then, max pooling units are used to extract local maxima from the convolutional features, further compressing the sequence length and retaining the most significant feature responses, effectively suppressing interference from secondary information. The distillation layer allows the encoder stack to progressively refine high-level semantics and compress sequence size like a pyramid during layer-by-layer propagation, gradually aggregating fine-grained time-step information from the lower layers into coarse-grained time-segment features at higher layers. This not only significantly reduces the computational burden of subsequent layers but also allows the model to focus more on key time periods and patterns that decisively impact marginal costs, avoiding information dilution and overwhelming during the propagation of lengthy sequences, and significantly improving the model's processing efficiency and prediction accuracy for long temporal data.
[0059] In a preferred embodiment, after a portion of the encoder layer (optionally an encoder layer whose output temporal encoded sequence is longer than a set sequence length), a distillation operation is used to compress the sequence length and highlight important values. This operation utilizes a one-dimensional convolutional layer (kernel size 3, stride 2) and a max-pooling layer to compress the sequence length, reducing the output sequence length. It may be less than the input length. To highlight important values. Specifically, for the first... Layer output After distillation, the updated product is obtained. ,in ( (This is a function for rounding down).
[0060] Output of each layer encoder As the input to the next layer encoder, after After layer encoder, the encoded representation can be obtained. ,in The length of the sequence after distillation.
[0061] In a preferred embodiment, the output layer includes a first fully connected unit, a modified linear activation element, and a second fully connected unit connected in sequence.
[0062] The first fully connected unit is used to map the global representation vector from the model dimension to the feedforward hidden dimension to obtain the first connection result; A modified linear activation element is connected to the first fully connected unit to perform an element-wise nonlinear transformation on the first connection result to obtain the activated connection result; The second fully connected unit, connected to the modified linear activation element, is used to map the activation connection result from the feedforward hidden dimension to the quantile number dimension, so as to obtain the predicted value of the node marginal cost at different quantile levels.
[0063] Specifically, the global representation vector output from the sequence summarization layer can be mapped to a higher-dimensional feedforward hidden space via a first fully connected unit, expanding the capacity of feature representation and enabling the model to fully mine the multi-dimensional marginal cost information contained in the global representation. Furthermore, by modifying the linear activation elements and introducing a nonlinear transformation, the model's expressive power is enhanced, allowing it to fit the complex nonlinear mapping relationship between marginal cost and system operating state. Finally, the activated features are mapped to the quantile dimension via a second fully connected unit, directly outputting predicted values for multiple preset quantiles. This quantile output mechanism allows the model to transcend the traditional single-point prediction paradigm, simultaneously providing point estimates of marginal cost (median quantiles) and interval estimates at different confidence levels (such as an 80% confidence interval composed of the 10th and 90th quantiles). This provides users with complete probabilistic prediction information, enabling them not only to know the expected value of marginal cost but also to clearly grasp the risk range of cost fluctuations and the possibility of extreme situations. This allows for more prudent and scientific judgments in risk management and scheduling decisions, significantly improving the reliability of power system operation.
[0064] In a preferred embodiment, the output layer first performs average pooling on the encoded sequence along the time dimension to obtain a global representation vector. As shown in B10: (B10); In the formula: This is an average pooling operation, i.e., along the time dimension... All Take the average of the vectors.
[0065] Then the global representation vector is mapped to multiple fully connected layers. The output has a preset quantile. For example, Corresponding quantiles The calculation expressions are shown in B11 and B12: (B11); (B12); In the formula: , For the weights and biases of the first fully connected layer (for hidden layer dimensions) , The weights and biases for the second fully connected layer; For the sample of quantile predicted values, of which Corresponding quantiles .
[0066] In a preferred embodiment, for each sample The quantile loss is defined as follows: ;
[0067] In the formula: Let be the loss function for the nth sample; This represents the actual marginal cost at each node. This is the predicted value for the kth quantile; This represents the kth quantile level.
[0068] Loss function for the entire training set This is the weighted average of the losses for all samples: ;
[0069] In the formula: This is the loss weight for the nth sample.
[0070] The Adam algorithm is preferred, using all training samples as input and all network optimizable parameters (i.e., weights and biases) as decision variables, minimizing the quantile loss function. An early stopping strategy is employed to prevent overfitting. After training, a marginal cost prediction model is obtained. (Including all optimized network parameters), the model trains multiple quantile outputs using the quantile loss function to generate residual demand curve scenarios at different confidence levels.
[0071] S3: Obtain the operating power range of the power station and discretize it to obtain several operating power threshold points; collect the predicted grid operation data at t times before the prediction date, and replace the operating power at t-1 times before the prediction date with the operating power at t times before the prediction date. Then, replace the operating power at t times before the prediction date with each operating power threshold point in turn to construct several time series sample sequences. Specifically, since environmental factors vary greatly at different times of the day, such as the significant differences in temperature and wind speed between early morning and noon, the day can be divided into several time periods. Power grid operation data can be collected at each time period to analyze the temporal characteristics of the power grid operation data. At the same time, due to the regularity of social activities, daily electricity consumption fluctuates regularly throughout the week, such as weekday electricity consumption being significantly lower than weekend consumption. Therefore, hourly data and weekday data also need to be obtained. Based on the power grid operation data at each time period, a time-series sample sequence can be constructed as the input for the marginal cost prediction model in subsequent steps.
[0072] More specifically, since the operating power in the power grid operation data at each moment is an unknown value, it is necessary to obtain the relationship between the operating power at each moment and the node marginal cost through subsequent prediction steps. This allows the power plant to determine the required operating power based on the relationship between operating power and node marginal cost. Therefore, it is possible to obtain the operating power range of the power plant and discretize it to obtain several operating power threshold points, which are used to characterize all the adjustable operating power of the power plant. Then, the predicted power grid operation data at t moments before the prediction date is collected through any existing database (such as obtaining wind speed, temperature, etc. for the next few days from the meteorological bureau). The actual operating power at t-1 moments before the prediction date is replaced with the operating power at t moments before the prediction date. Then, the operating power at t moments before the prediction date is replaced with each operating power threshold point in turn, and several time series sample sequences are constructed to provide a data basis for subsequent predictions, thereby obtaining the possible node marginal cost of the power plant at t moments before the prediction date under different operating power.
[0073] In a preferred embodiment, the step of constructing a time-series sample sequence includes: c01) Discretization of the residual demand curve. For a power plant, its installed capacity is known to be... The possible operating power range Divided into equal parts The segment (corresponding to the number of discretized intervals) is obtained The operating power threshold points are shown in C1: (C1); In the formula: For the first One operating power threshold; That is, the pre-defined number of segments.
[0074] c02) For each moment of the predicted day ( To prepare the forecast, corresponding feature sequences are needed. Features for the forecast day include: wind speed. Temperature ,Hour Week type Regional load forecast And the operating power at each time point. The operating power at the predicted time point is calculated using the operating power threshold for each interval. Set the operating power threshold for each interval. By sequentially replacing the running power at the prediction time, the feature sequence corresponding to the prediction time can be constructed as shown in C2, and the feature sequences corresponding to several times before the prediction time are shown in C3. (C2); (C3); In the formula: The characteristic sequence of the operating power range of the i-th discrete residual demand curve at time t (i.e., the prediction time) on the prediction day; , , , , These are the standardized feature sequences at time t on the prediction day; For the standardized first One operating power threshold; To predict the characteristic sequence at time tl of the day; To predict the actual operating power at time tl of the day before, the feature sequences corresponding to the predicted time and several times before it are spliced together to obtain the time series sample sequence corresponding to the predicted time.
[0075] S4: Input each time series sample sequence into the trained marginal cost prediction model to obtain the predicted value of the node marginal cost at different quantile levels corresponding to each operating power threshold point at the t-th time of the prediction day, and determine the scenario probability weight according to the quantile interval corresponding to each quantile level to generate the residual demand curve under different scenarios at the t-th time. Specifically, since the marginal cost of nodes is affected by multiple random factors such as the uncertainty of renewable energy output and load fluctuations, a single deterministic prediction is difficult to effectively characterize the tail risk and fluctuation range of the cost distribution. Therefore, it is possible to input each time series sample sequence into a trained marginal cost prediction model to obtain the predicted value of the node marginal cost at different quantile levels corresponding to each operating power threshold point at time t of the prediction day. The scenario probability weights (e.g., ...) are then determined based on the quantile intervals corresponding to each quantile level. Figure 4As shown, curves of different colors represent residual demand scenarios at different quantile levels, indicating the marginal costs of different nodes at different quantiles that may occur at the same operating power threshold point, i.e., the probability of the corresponding node marginal cost occurring. Each curve corresponds to a scenario, and the color of each curve is used to indicate the different probabilities / confidence levels of each scenario (the probability of a scenario occurring is determined by adjacent quantile intervals, and the color is only used for illustration). This generates residual demand curves under different scenarios at time t, enabling scheduling decisions to simultaneously consider the median expected cost and extreme cost scenarios at different confidence levels, quantitatively assess the risk-reward trade-off relationship corresponding to each output level, and provide probabilistic cost boundary information for subsequent robust scheduling and risk avoidance, thereby obtaining the optimal operating power of the power plant at each time of the forecast day.
[0076] In a preferred embodiment, the step of generating several residual demand curves based on the predicted values of node marginal costs under different operating power levels includes: c03) Quantile prediction. The input time-coded sequence... Input the trained marginal cost prediction model The corresponding output is shown in C3: (C3) In the formula: Indicates at time And the operating power is At that time, the first quantiles The corresponding node marginal cost.
[0077] c04) Construct the residual demand curve scenario. For each quantile level ( We can obtain a residual demand curve, which is formed by... The points are composed as shown in C4: (C4) This curve represents the confidence level. Below is a diagram illustrating the predicted node marginal costs for different operating power levels. Figure 4 Different quantile levels correspond to different confidence levels, therefore we obtain The remaining demand curves form a set of scenarios as shown in C5: (C5) In this embodiment, a method for generating residual demand curves based on neural networks is presented, the core inventive point of which is: I. A three-stage progressive framework of "feature extraction - temporal modeling - uncertainty quantification" is introduced, consisting of encoding embedding, probabilistic sparse attention, sequence distillation, and quantile output. Compared to traditional residual demand curve generation methods based on statistical regression or simple neural networks, this framework offers unique technical advantages such as high prediction accuracy, comprehensive scene coverage, and superior computational efficiency. Specifically, this includes: 1. A feature representation mechanism combining encoding embedding and positional encoding enables unified representation and sequential injection of multi-source heterogeneous data: Traditional residual demand curve generation methods often perform simple numerical concatenation or standardization on multi-source heterogeneous input data such as wind speed, load, and temperature, neglecting the differentiated semantic requirements of periodic features (such as hourly data) and categorical features (such as weekday data) in time series modeling, and lack explicit encoding of temporal order relationships, making it difficult for the model to distinguish the temporal position differences of inputs at different times. This application linearly projects the time series sample sequence through an encoding embedding layer, mapping the multidimensional feature vectors of each time step to a unified model dimensional space; at the same time, it injects absolute positional information into each time step through positional encoding units based on sine and cosine functions, preserving the temporal order relationship. This collaborative mechanism of "unified projection + positional ordering" enables continuous variables, periodic variables, and categorical variables to maintain their respective numerical characteristics and semantic information in a unified representation space, effectively avoiding the temporal modeling bias caused by differences in feature scale or missing positional information in traditional methods, and laying a high-quality feature input foundation for subsequent deep attention extraction.
[0078] 2. An encoder architecture that alternates between probabilistic sparse self-attention and feedforward networks achieves efficient capture and noise filtering of long-term time-series dependencies: Traditional time-series forecasting models based on fully connected self-attention experience a quadratic increase in computational complexity with sequence length when processing long-window historical data of power systems. Furthermore, performing full attention calculations on all query-key pairs easily introduces redundant noise, making it difficult for the model to focus on the key historical moments that truly drive price changes in node marginal cost prediction. This application uses a probabilistic sparse self-attention unit to calculate the sparsity score between each query vector and all key vectors based on a sparsity metric element. Only the core query vectors with the highest sparsity scores are retained for scaled dot product attention calculations. Simultaneously, a feedforward network unit performs nonlinear transformation and layer normalization on the sparse attention feature sequence to further refine the feature dimensions. This alternating stacking mechanism of "sparse screening-deep refinement" reduces the computational complexity of attention from quadratic to near linear, enabling the model to maintain high computational efficiency when dealing with long time-series windows spanning hours or even days. At the same time, by adaptively filtering out historical time step interference with low relevance to the current prediction task through probabilistic sparsity metrics, it can still accurately focus on key historical moments driving price mutations even in extreme scenarios such as drastic fluctuations in new energy output and peaks in node marginal costs, effectively reducing prediction errors and improving the model's generalization ability.
[0079] 3. A sequence distillation mechanism combining one-dimensional convolution and max pooling achieves hierarchical compression and key information preservation in multi-scale temporal representations: Traditional Transformer-type models maintain a constant sequence length during deep layer stacking, leading to increased size of deep feature maps, computational redundancy, and difficulty in effectively extracting local pattern features at different time scales. This application addresses this by setting distillation layers in the encoder stack. One-dimensional convolutional units extract local features and downsample the stride of the encoded representation, followed by max pooling units to extract local maxima, achieving layer-by-layer compression of the sequence length. This "convolutional local extraction - pooling key preservation" distillation mechanism allows the model to gradually focus on the most discriminative temporal patterns during deep encoding. It captures the fluctuation features of node marginal costs at different time scales through the local receptive field of convolution, and preserves the extreme value information (corresponding to key states such as node marginal cost peaks) within each local window through max pooling. This significantly reduces the computational burden of subsequent layers while achieving a multi-level progressive representation of "detail preservation - scale compression - key focus," effectively improving the model's modeling depth and computational efficiency for long temporal dependencies.
[0080] 4. A collaborative output architecture combining global average pooling and quantile-based full connectivity enables precise characterization of node marginal cost distribution and direct generation of residual demand curves across multiple scenarios: Traditional point forecasting methods only output a single forecast value, failing to provide risk boundary information for electricity market participants' bidding decisions and system scheduling operations; traditional uncertainty quantification methods based on Monte Carlo simulation or parameterized probability distributions require additional post-processing steps to generate residual demand curves for different scenarios, resulting in lengthy processes and error accumulation. This application uses a sequence aggregation layer to perform average pooling of the encoded representation along the time dimension to obtain a global representation vector aggregating global context information; then, through the cascaded mapping of the first fully connected unit, the modified linear activation element, and the second fully connected unit in the output layer, the global representation vector is directly mapped to predicted values of multiple preset quantiles. This end-to-end architecture of "global aggregation-quantile output" enables the model to output price prediction distributions covering different confidence levels in a single forward propagation, without the need for additional sampling or distribution fitting steps. By discretizing the operating power range of power plants into several threshold points, and constructing a time-series sample sequence using the grid operation data corresponding to each threshold point and inputting it into the model, the predicted values of node marginal costs at different quantiles corresponding to each operating power threshold point can be directly obtained, thereby generating residual demand curves under different scenarios. This integrated generation mechanism of "prediction-quantile-curve" significantly improves the efficiency and physical consistency of residual demand curve generation, providing high-quality scenario inputs with both accuracy and risk coverage for power system dispatching decisions.
[0081] II. A data preprocessing mechanism based on fluctuation range detection and multi-type feature encoding is adopted. Compared with simple missing value imputation or uniform standardization methods, it has unique technical advantages such as high data quality, sufficient feature representation, and strong model robustness. Specifically, this is reflected in: 1. A collaborative outlier handling mechanism combining fluctuation range detection and linear interpolation enables precise control of training data quality: Traditional data preprocessing methods often employ fixed missing value imputation strategies or global outlier removal rules, neglecting the differentiated fluctuation characteristics of power system operation data across different feature dimensions. This leads to overly conservative or insufficient outlier handling, affecting model training quality. This application obtains the average value and standard deviation of various historical power grid operation data at different times within a set time period and determines the fluctuation range of each feature based on a set fluctuation coefficient. For outliers exceeding the fluctuation range, linear interpolation is performed to replace them based on adjacent normal values. This differentiated processing mechanism of "statistical fluctuation definition - local linear repair" ensures that features with different fluctuation characteristics, such as wind speed, temperature, and load, can obtain outlier judgment criteria that match their physical laws, avoiding misjudgment of normal fluctuations or missed detection of outliers caused by a globally uniform threshold. By maintaining the continuity and local trend characteristics of time-series data through linear interpolation, the quality and reliability of the training dataset are effectively improved, providing a high-quality data foundation for the stable training of subsequent deep learning models.
[0082] 2. A sample construction mechanism based on multi-type feature differential coding and historical window segmentation to achieve full expression and structured input of heterogeneous time-series features: Traditional methods use uniform standardization processing for continuous, periodic, and categorical data, which leads to the destruction of the cyclic semantics of periodic features (such as hourly data) and the forced mapping of the discrete attributes of categorical features (such as weekday data) to continuous values, resulting in loss of feature information. This application classifies various historical power grid operation data according to their mathematical characteristics. Continuous data (wind speed, temperature, regional load forecast, operating power) is standardized to eliminate dimensional differences. Periodic data (hourly data) is sine and cosine encoded to retain cyclic periodic characteristics. Categorical data (weekday data) is one-hot encoded to maintain discrete category distinguishability. The feature sequences encoded by the three types are then concatenated into a unified power grid operation feature sequence and segmented according to a historical window of a set length to construct the time-series sample sequence corresponding to each moment. This structured sample construction mechanism of "classification encoding-feature concatenation-window segmentation" allows different types of features to maintain their respective mathematical properties and physical semantics in a unified input space. Sine and cosine encoding ensures the continuity of the 24-hour cycle (such as the proximity relationship between 11 PM and 12 AM), and one-hot encoding avoids spurious order relationships between weekday categories. By transforming long-term time-series data into structured supervised learning samples through historical window segmentation, the model can make full use of historical multi-step information to predict the marginal cost of the node at the current moment, effectively improving the sufficiency of feature expression and the structured degree of model input.
[0083] Preferably, to further broaden the application scenarios of this application, the residual demand curve generation method based on neural networks may also include: S5: Based on the remaining demand curves, the optimal output value of each power station is obtained, and the optimal operating power of each power station at each time of the forecast day is obtained.
[0084] Specifically, since the residual demand curves under different scenarios reflect the marginal cost change trends of each power station at different output levels, relying solely on experience-based scheduling is insufficient to achieve global cost optimization and efficient resource allocation. Therefore, it is possible to obtain the optimal operating power of each power station at each moment of the forecast day based on the residual demand curves, so as to adjust the operating power of the corresponding power station in real time. This allows the output decision of each station to be optimized based on the multi-scenario residual demand curves constructed by quantile cost prediction, thereby achieving a balance between optimal power station output and risk controllability while meeting system safety constraints, and improving the overall efficiency of power grid operation and the level of intelligent scheduling.
[0085] On the other hand, the present invention also provides a computer storage medium storing executable program code; the executable program code is used to execute any of the above-mentioned neural network-based residual demand curve generation methods.
[0086] On the other hand, the present invention also provides a terminal device, including a memory and a processor; the memory stores program code that can be executed by the processor; the program code is used to execute any of the above-mentioned neural network-based residual demand curve generation methods.
[0087] For example, the program code can be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the program code in the terminal device.
[0088] The terminal device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The terminal device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that the terminal device may also include input / output devices, network access devices, buses, etc.
[0089] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0090] The memory can be an internal storage unit of the terminal device, such as a hard drive or RAM. The memory can also be an external storage device of the terminal device, such as a plug-in hard drive, SmartMediaCard (SMC), Secure Digital (SD) card, or FlashCard. Furthermore, the memory can include both internal and external storage units of the terminal device. The memory is used to store the program code and other programs and data required by the terminal device. The memory can also be used to temporarily store data that has been output or will be output.
[0091] The aforementioned computer storage medium and terminal device are created based on the aforementioned neural network-based residual demand curve generation method. Their technical functions and beneficial effects will not be elaborated here. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0092] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.
Claims
1. A method for generating residual demand curves based on neural networks, characterized in that, include: Historical power grid operation data for each moment within a set time period are collected, including wind speed, temperature, regional load forecast, operating power, hourly data and weekday data, as well as the node marginal cost at the corresponding moment, in order to construct a training dataset. A marginal cost prediction model is constructed and trained based on the training dataset, taking time-series sample sequences as input and node marginal costs as output, including: The input layer is used to receive time-series sample sequences; The encoding embedding layer, connected to the input layer, is used to perform linear projection on the temporal sample sequence and add position encoding to inject temporal values, resulting in a temporal encoded sequence. The encoder stack, connected to the encoding embedding layer, is used to perform multi-layer attention extraction and sequence compression on the temporal encoded sequence to obtain the encoded representation; The sequence summarization layer, connected to the encoder stack, is used to perform average pooling on the encoded representation along the time dimension to obtain the global representation vector; The output layer, connected to the sequence summarization layer, is used to map the global representation vector to the predicted values of multiple preset quantiles, and output the predicted values of the node marginal cost at different quantile levels. Obtain the operating power range of the power plant and discretize it to obtain several operating power threshold points; collect the predicted grid operation data t times before the prediction date, and replace the operating power t-1 times before the prediction date with the operating power t-1 times before the prediction date. Then, replace the operating power at the t-th time of the prediction date with each operating power threshold point in turn to construct several time series sample sequences. Input each time series sample sequence into the trained marginal cost prediction model to obtain the predicted value of the node marginal cost at different quantile levels corresponding to each operating power threshold point at time t of the prediction day. Then, determine the scenario probability weight according to the quantile interval corresponding to each quantile level to generate the residual demand curve under different scenarios at time t.
2. The method according to claim 1, characterized in that, The encoding embedding layer includes linear projection units and position encoding units connected in sequence; The linear projection unit is used to map the feature vector of each time step in the time series sample sequence to the model dimension through a linear transformation, so as to obtain the projected feature sequence. The position coding unit, connected to the linear projection unit, is used to add position coding information based on sine and cosine functions to each time step in the projected feature sequence in order to preserve the time order relationship and obtain the time-series coded sequence.
3. The method according to claim 1, characterized in that, An encoder stack consists of several layers of encoder layers and distillation layers stacked sequentially. Each encoder layer includes a probabilistic sparse self-attention unit and a feedforward network unit connected in sequence; The probabilistic sparse self-attention unit is used to receive the input temporal coding sequence of the current encoder layer, calculate the query matrix, key matrix and value matrix, filter the core query vector of the power grid based on the probabilistic sparse metric, perform scaling dot product attention calculation only on the filtered queries, and output sparse attention feature sequence. The feedforward network unit, connected to the probabilistic sparse self-attention unit, is used to perform feedforward calculation and layer normalization operations on the sparse attention feature sequence to obtain the output temporal coding sequence of the current encoder layer; the output temporal coding sequence of the last encoder layer is the coded representation. The distillation layer is set after the encoder layer whose output timing coding sequence is longer than a set sequence length. It is used to compress the output timing coding sequence of the current encoder layer to obtain the distilled output timing coding sequence.
4. The method according to claim 3, characterized in that, The probabilistic sparse self-attention unit comprises a query key-value projection element, a sparsity measurement element, a query filtering element, and a sparse attention calculation element connected in sequence. The query key-value projection element is used to perform linear projections on the input time-series encoded sequence representation to obtain the query matrix, key matrix, and value matrix; The sparsity metric element, connected to the query key projection element, is used to calculate the sparsity score between each query vector and all key vectors, resulting in a sparsity metric sequence. The query filtering element is connected to the sparsity metric element to retain the query vectors with the highest sparsity scores in the sparsity metric sequence, thus obtaining the filtered query submatrix. The sparse attention computation element, connected to the query key-value projection element and the query filtering element, is used to perform scaled dot product attention computation on the filtered query submatrix and all key matrices and value matrices to obtain the sparse attention feature sequence.
5. The method according to claim 3, characterized in that, The feedforward network unit comprises a first-layer linear transformation element, a modified linear activation element, a second-layer linear transformation element, and a layer normalization element connected in sequence. The first layer of linear transformation elements is used to perform linear transformation on the output of the probabilistic sparse self-attention unit, mapping the model dimension to the feedforward hidden dimension to obtain the first transformation result. The modified linear activation element is connected to the first-layer linear transformation element and is used to perform an element-wise nonlinear transformation on the first transformation result, setting the negative value to zero to obtain the activation result; The second-layer linear transformation element, connected to the modified linear activation element, is used to perform a linear transformation on the activation result, mapping the feedforward hidden dimension back to the model dimension to obtain the second transformation result; The first layer normalization element is connected to the second layer linear transformation element. It is used to add the second transformation result to the input of the feedforward network unit as a residual, and calculate the mean and standard deviation of the addition result step by step, and perform normalization scaling to obtain the output time-series coded sequence.
6. The method according to claim 3, characterized in that, The distillation layer consists of sequentially connected one-dimensional convolutional units and max pooling units; One-dimensional convolutional units are used to extract local features and downsample the stride of the output temporal encoded sequence to obtain convolutional features. The max pooling unit, connected to the one-dimensional convolutional unit, is used to extract local maxima from the convolutional features to obtain the distilled output temporal coding sequence.
7. The method according to claim 1, characterized in that, The output layer includes a first fully connected unit, a modified linear activation element, and a second fully connected unit connected in sequence. The first fully connected unit is used to map the global representation vector from the model dimension to the feedforward hidden dimension to obtain the first connection result; A modified linear activation element is connected to the first fully connected unit to perform an element-wise nonlinear transformation on the first connection result to obtain the activated connection result; The second fully connected unit, connected to the modified linear activation element, is used to map the activation connection result from the feedforward hidden dimension to the quantile number dimension, so as to obtain the predicted value of the node marginal cost at different quantile levels.
8. The method according to any one of claims 1 to 7, characterized in that, The steps to construct the training dataset include: Obtain the average value and standard deviation of various historical power grid operation data at each moment within a set time period, and determine the fluctuation range of various historical power grid operation data based on the set fluctuation coefficient, average value and standard deviation; Determine whether the values of various historical power grid operation data are within the corresponding fluctuation range. If so, they are judged as normal values and retained. If not, they are judged as abnormal values. Linear interpolation replacement is performed on several normal values adjacent to the abnormal values to obtain several optimized historical power grid operation data. Based on the mathematical characteristics of various historical power grid operation data, they are classified into continuous data, periodic data, and categorical data; continuous data includes wind speed, temperature, regional load forecasts, and operating power; periodic data includes hourly data; and categorical data includes weekly data. Various continuous data are standardized to map to a uniform numerical scale, eliminating dimensional differences and obtaining continuous feature sequences; periodic data are encoded using sine and cosine to obtain periodic feature sequences; and categorical data are encoded using one-hot encoding to obtain categorical feature sequences. By concatenating continuous feature sequences, periodic feature sequences, and categorical feature sequences, a power grid operation feature sequence is obtained. The power grid operation feature sequence is then segmented according to a historical window of a set length to construct the time series sample sequence corresponding to each time point. The node marginal cost at the corresponding time point is then obtained to construct the training dataset.
9. A computer storage medium, characterized in that, It stores executable program code; the executable program code is used to execute the residual demand curve generation method based on neural networks according to any one of claims 1 to 8.
10. A terminal device, characterized in that, It includes a memory and a processor; the memory stores program code that can be executed by the processor; the program code is used to execute the residual demand curve generation method based on a neural network as described in any one of claims 1 to 8.