Supercritical unit wide-load modeling method and device, electronic equipment and storage medium

By employing an attention residual bi-branch gated recurrent unit model and a two-stage planned sampling training strategy, the problems of autoregressive prediction accuracy and stability in supercritical coal-fired power plants under wide load conditions were solved, achieving high-precision and robust dynamic modeling.

CN120805096APending Publication Date: 2025-10-17GUODIAN NANJING ELECTRIC POWER TEST RES CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510827209.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

The data-driven models of existing supercritical coal-fired power plants have low autoregressive prediction accuracy, poor long-term stability and limited generalization ability under a wide range of load conditions, making them difficult to adapt to complex operating conditions.

Method used

An attention residual bi-branch gated recurrent unit model is adopted, and time-series feature extraction is enhanced through data preprocessing. Combined with a two-stage planned sampling training strategy, an autoregressive prediction model for supercritical units is constructed to capture multi-scale dynamic characteristics.

Benefits of technology

It significantly improves the model's fitting accuracy and robustness over a wide load range, alleviates exposure bias, enhances the stability and generalization ability of long-term autoregressive predictions, and has strong adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120805096A_ABST
    Figure CN120805096A_ABST
Patent Text Reader

Abstract

The invention relates to a wide-load modeling method and device for a supercritical unit, electronic equipment and a storage medium, and the method comprises the steps: collecting the operation data of the supercritical unit, and carrying out the data preprocessing of the operation data, so as to generate data meeting a preset standardization condition; constructing a time sequence data sample of the supercritical unit based on the data meeting the preset standardization condition; constructing an attention residual error double-branch gating circulation unit model for the supercritical unit according to the time sequence data sample; and based on a preset two-stage plan sampling training strategy, training a pre-constructed attention residual error double-branch gating circulation unit model to generate a model autoregression prediction result of the supercritical unit. Therefore, the problems of low autoregressive prediction precision, poor long-term stability and limited generalization ability caused by factors such as exposure deviation and the like of an existing supercritical generator set data driving model under a wide load working condition are solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent control and modeling of thermal power plants, and particularly relates to a supercritical unit wide load modeling method and device, an electronic equipment and a storage medium. BACKGROUND

[0002] Supercritical coal-fired power plants still play an important role in modern energy systems due to their high thermal efficiency and wide load regulation capability, especially providing flexible backup for the power grid when new energy fluctuates. However, the dynamic characteristics of supercritical units exhibit multivariable coupling, strong nonlinearity and time-varying characteristics, which bring challenges to modeling and control.

[0003] In related technologies, modeling methods based on physical mechanisms, such as partial differential equation models based on thermodynamic principles, are complex to calculate and difficult to adapt to wide load nonlinear dynamics, and have limited applicability in real-time prediction and control scenarios. In recent years, data-driven neural network models have shown potential in industrial system dynamic modeling due to their strong nonlinear fitting and time series feature learning capabilities. However, these models are susceptible to exposure bias in long-term autoregressive prediction, resulting in insufficient generalization ability and stability, especially in complex working conditions, and the prediction accuracy decreases, which needs to be improved. SUMMARY

[0004] The present application provides a supercritical unit wide load modeling method, device, electronic equipment and storage medium to solve the problems of low autoregressive prediction accuracy, poor long-term stability and limited generalization ability of existing supercritical generator data-driven models under wide load conditions due to exposure bias and other factors.

[0005] The first aspect of the present application provides a supercritical unit wide load modeling method, which is applied to the model construction stage, comprising the following steps: collecting operation data of a supercritical unit, and performing data preprocessing on the operation data to generate data meeting preset standardization conditions; based on the data meeting the preset standardization conditions, constructing a time series data sample of the supercritical unit; and constructing an attention residual double-branch gated recurrent unit model according to the time series data sample for generating model autoregressive prediction results of the supercritical unit.

[0006] Optionally, in an embodiment of the present application, the operation data of the supercritical unit is collected, and the operation data is preprocessed to generate data meeting preset standardization conditions, comprising: collecting any control variable of the throttle opening, the coal feed rate and the water feed rate of the supercritical unit, and collecting any state variable of the main steam pressure, the unit load and the separator temperature of the supercritical unit; and performing data preprocessing on the any control variable and the any state variable to generate the data meeting the preset standardization conditions.

[0007] Optionally, in an embodiment of the present application, the calculation formula of the data preprocessing is:

[0008]

[0009] wherein μ and σ are the mean and standard deviation of the input and output respectively, x is the original data, x * is the data satisfying the preset standardization condition.

[0010] Optionally, in an embodiment of the present application, the constructing an attention residual double-branch gated recurrent unit model according to the time series data sample comprises: weighting an input sequence by using a self-attention mechanism to determine a time step feature satisfying a preset condition; capturing long and short term dependencies in the input sequence based on the time step feature satisfying the preset condition, and normalizing the output of two layers of gated recurrent units based on the long and short term dependencies to generate final output features of the two layers of gated recurrent units; inputting the final output features into a main regression branch and an auxiliary regression branch of the supercritical unit, and constructing the attention residual double-branch gated recurrent unit model according to the main regression branch and the auxiliary regression branch.

[0011] Optionally, in an embodiment of the present application, the formula for weighting the input sequence by using the self-attention mechanism is:

[0012]

[0013] wherein Q, K, and V are query, key, and value matrices generated by linear transformation of the input sequence, and dk is the dimension of the key vector.

[0014] The calculation formula of the two layers of gated recurrent units is:

[0015] z t = σ s (W z ·[h t-1 , x t ]+b z )

[0016] r t = σ s (W r ·[h t-1 , x t ]+b r )

[0017]

[0018] wherein x t is the current input, h t-1is the hidden state of the previous time, W, U, b are learnable parameters, σ is the Sigmoid function, and is the element-wise product.

[0019] The second aspect embodiment of the present application provides a supercritical unit wide load modeling method, which is applied to a model application stage and includes the following steps: obtaining control variables of a supercritical unit; inputting the control variables into a pre-constructed attention residual double-branch gated recurrent unit model; training the pre-constructed attention residual double-branch gated recurrent unit model based on a preset two-stage plan sampling training strategy to generate a model autoregressive prediction result of the supercritical unit, wherein the pre-constructed attention residual double-branch gated recurrent unit model is composed of a main regression branch and an auxiliary regression branch of the supercritical unit.

[0020] Optionally, in an embodiment of the present application, the preset two-stage plan sampling training strategy is used to train the pre-constructed attention residual double-branch gated recurrent unit model to generate the model autoregressive prediction result of the supercritical unit, including: in a first stage of training of the attention residual double-branch gated recurrent unit model, obtaining an input of each time step and determining an actual previous time observation value according to the input; in a second stage of training of the attention residual double-branch gated recurrent unit model, obtaining a first preset probability dynamically adjusted with a training iteration number; determining the actual previous time observation value according to the first preset probability and determining a prediction value of the attention residual double-branch gated recurrent unit model at a previous time according to a second preset probability; and generating the model autoregressive prediction result of the supercritical unit based on the actual previous time observation value and the prediction value of the attention residual double-branch gated recurrent unit model at the previous time.

[0021] Optionally, in an embodiment of the present application, a calculation formula of the first preset probability is:

[0022]

[0023] wherein t is a current time step, n is a size of an input window, and N is a total time step of training data.

[0024] The third aspect embodiment of the present application provides a supercritical unit wide load modeling device, which is applied to a model construction stage and includes: a collection module configured to collect operation data of a supercritical unit and perform data preprocessing on the operation data to generate data meeting a preset standardization condition; a construction module configured to construct time series data samples of the supercritical unit based on the data meeting the preset standardization condition; and a generation module configured to construct an attention residual double-branch gated recurrent unit model based on the time series data samples to generate a model autoregressive prediction result of the supercritical unit.

[0025] Optionally, in an embodiment of the present application, the collecting module comprises: a collecting unit, configured to collect any control variable of a throttle opening degree, a coal feeding rate and a water feeding rate of the supercritical unit, and to collect any state variable of a main steam pressure, a unit load and a separator temperature of the supercritical unit; and a preprocessing unit, configured to perform data preprocessing on the any control variable and the any state variable to generate the data satisfying the preset standardization condition.

[0026] Optionally, in an embodiment of the present application, a calculation formula of the data preprocessing is as follows:

[0027]

[0028] wherein μ and σ are the average value and the standard deviation of the input and the output respectively, x is the original data, x * is the data satisfying the preset standardization condition.

[0029] Optionally, in an embodiment of the present application, the generating module comprises: a weighting unit, configured to weight an input sequence by using a self-attention mechanism to determine a time step feature satisfying a preset condition; a capturing unit, configured to capture a long-short term dependency in the input sequence based on the time step feature satisfying the preset condition, and to perform normalization processing on the output of two-layer gated recurrent units based on the long-short term dependency to generate a final output feature of the two-layer gated recurrent units; and a constructing unit, configured to input the final output feature into a main regression branch and an auxiliary regression branch of the supercritical unit, and to construct the attention residual double-branch gated recurrent unit model according to the main regression branch and the auxiliary regression branch.

[0030] Optionally, in an embodiment of the present application, a formula of weighting the input sequence by using the self-attention mechanism is as follows:

[0031]

[0032] wherein Q, K and V are query, key and value matrices generated by linear transformation of the input sequence, and dk is the dimension of a key vector;

[0033] A calculation formula of the two-layer gated recurrent units is as follows:

[0034] z t = σ s (W z ·[h t-1 , x t ]+b z )

[0035] rt = σ s (W r · [h t-1 , x t ]+b r )

[0036]

[0037] where x t is the current input, h t-1 is the previous hidden state, W, U, b are learnable parameters, σ is the Sigmoid function, and is the element-wise product.

[0038] The fourth aspect embodiment of the present application provides a supercritical unit wide load modeling device, which is applied to a model application stage and includes: an acquisition module configured to acquire control variables of a supercritical unit; and a prediction module configured to input the control variables into a pre-constructed attention residual double-branch gated recurrent unit model, train the pre-constructed attention residual double-branch gated recurrent unit model based on a pre-set two-stage planning sampling training strategy, and generate a model autoregressive prediction result of the supercritical unit, wherein the pre-constructed attention residual double-branch gated recurrent unit model is composed of a main regression branch and an auxiliary regression branch of the supercritical unit.

[0039] Optionally, in an embodiment of the present application, the prediction module includes: a determination unit configured to acquire an input of each time step and determine an actual previous time observation value according to the input in a first stage of training of the attention residual double-branch gated recurrent unit model; a probability acquisition unit configured to acquire a first preset probability dynamically adjusted with a training iteration number in a second stage of training of the attention residual double-branch gated recurrent unit model; a prediction value determination unit configured to determine the actual previous time observation value according to the first preset probability and determine a prediction value of the attention residual double-branch gated recurrent unit model in a previous time according to a second preset probability; and a result generation unit configured to generate the model autoregressive prediction result of the supercritical unit based on the actual previous time observation value and the prediction value of the attention residual double-branch gated recurrent unit model in the previous time.

[0040] Optionally, in an embodiment of the present application, a calculation formula of the first preset probability is:

[0041]

[0042] where t is a current time step, n is a size of an input window, and N is a total time step of training data.

[0043] The embodiments of the present application enhance the extraction of key time series features through an attention mechanism, optimize information flow through residual connections, and capture multi-scale dynamic characteristics through a dual-branch structure. This significantly improves the model's accuracy in fitting the complex nonlinear dynamics of supercritical units over a wide load range and its robustness to operating condition changes. It effectively mitigates exposure bias and enhances the stability of long-term autoregressive predictions, resulting in superior overall performance and strong adaptability to wide loads. This solves the problems of low autoregressive prediction accuracy, poor long-term stability, and limited generalization capabilities of existing data-driven models for supercritical generator sets under wide load conditions, which are caused by factors such as exposure bias.

[0044] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0046] Figure 1 A flow chart of a supercritical unit wide load modeling method provided in an embodiment of the present application applied to a model building phase;

[0047] Figure 2 A flow chart of a supercritical unit wide load modeling method provided in an embodiment of the present application applied to a model application stage;

[0048] Figure 3 This is a structural diagram of an attention residual dual-branch GRU model according to one embodiment of the present application;

[0049] Figure 4 Schematic diagram of the internal structure of a GRU unit according to one embodiment of the present application;

[0050] Figure 5 Schematic diagram of a planned sampling method in two-stage training according to one embodiment of the present application;

[0051] Figure 6 This is a graph showing the results of long-term autoregressive prediction using different neural network models according to one embodiment of the present application on real data;

[0052] Figure 7a This is a distribution diagram of absolute errors of long-term autoregressive predictions performed on real data using different GRU models according to one embodiment of the present application;

[0053] Figure 7b This is a distribution diagram of absolute errors of long-term autoregressive predictions performed on real data using different GRU models according to one embodiment of the present application;

[0054] Figure 7c This is a distribution diagram of absolute errors of long-term autoregressive predictions performed on real data using different GRU models according to one embodiment of the present application;

[0055] Figure 8 This is a comparison chart of the results of long-term autoregressive prediction performed on real data by the GRU3 model according to one embodiment of the present application and the true value;

[0056] Figure 9 This is a relative error curve diagram of the long-term autoregressive prediction of the GRU3 model on real data according to one embodiment of the present application;

[0057] Figure 10 This is a structural diagram of a supercritical unit wide load modeling device provided in an embodiment of the present application applied to a model building stage;

[0058] Figure 11 This is a structural schematic diagram of a supercritical unit wide load modeling device provided in an embodiment of the present application applied in the model application stage. DETAILED DESCRIPTION

[0059] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0060] The following describes the supercritical unit wide load modeling method, device, electronic device and storage medium of the embodiment of the present application with reference to the accompanying drawings. In response to the problems mentioned in the above background technology that the existing supercritical generator set data-driven model has low autoregressive prediction accuracy, poor long-term stability and limited generalization ability under wide load conditions due to factors such as exposure bias, the present application provides a supercritical unit wide load modeling method, in which the key time series feature extraction is enhanced by the attention mechanism, the residual connection optimizes the information flow, and the dual-branch structure captures multi-scale dynamic characteristics, which significantly improves the model's fitting accuracy for the complex nonlinear dynamics of the supercritical unit in a wide load range and its robustness to operating condition changes, effectively alleviates exposure bias, enhances long-term autoregressive prediction stability, and has superior comprehensive performance and strong wide load adaptability. As a result, the problems of the existing supercritical generator set data-driven model under wide load conditions, such as low autoregressive prediction accuracy, poor long-term stability and limited generalization ability due to factors such as exposure bias, are solved.

[0061] Specifically, Figure 1 A schematic flow chart of a supercritical unit wide load modeling method provided in an embodiment of the present application applied to the model building stage.

[0062] As Figure 1 shown, the wide load modeling method of the supercritical unit includes the following steps:

[0063] In step S101, the operation data of the supercritical unit is collected, and the operation data is preprocessed to generate data meeting preset standardization conditions.

[0064] It can be understood that the data meeting the preset standardization conditions in the embodiments of the present application can be data after z-score standardization processing.

[0065] In actual execution process, the embodiments of the present application can collect the operation data of the supercritical unit, and preprocess the operation data to generate data meeting the standardization processing, thereby providing support for subsequent accurate capturing of wide load dynamic characteristics of the supercritical unit.

[0066] It should be noted that the preset standardization conditions can be set by those skilled in the art according to actual conditions, which is not limited here.

[0067] Optionally, in an embodiment of the present application, collecting the operation data of the supercritical unit and preprocessing the operation data to generate data meeting the preset standardization conditions includes: collecting any control variable of the throttle opening, the coal feeding rate and the water feeding rate of the supercritical unit, and collecting any state variable of the main steam pressure, the unit load and the separator temperature of the supercritical unit; preprocessing any control variable and any state variable to generate data meeting the preset standardization conditions.

[0068] It can be understood that in the prediction of the key parameters (such as the main steam pressure, the unit load and the separator temperature) of the supercritical unit, the embodiments of the present application show significant advantages in R 2 , RMSE, MAPE, MAE and other evaluation indexes, for example, in autoregressive prediction, the model (GRU3) in the present application reduces the RMSE of the main steam pressure, the unit load and the separator temperature by about 18.7%, 48.1% and 4.7% respectively compared with the baseline GRU1 model.

[0069] In actual implementation process, the embodiment of the application can select a certain 350 MW supercritical generator set as the research object, and verify the effectiveness of the proposed supercritical unit wide load modeling method based on attention residual double branch GRU and two-stage planning sampling training based on its real operation data. The actual operation data of the supercritical unit is collected, including the control variables as the model input: throttle opening (u1), coal feeding rate (u2), and feed water rate (u3); and the state variables as the model output: main steam pressure (y1), unit load (y2), and separator temperature (y3). The sliding window strategy is used to construct the time series data sample, and the window size is set. The z-score standardization processing is performed on all input and output data based on the training set data to generate the data satisfying the standardization processing.

[0070] In one embodiment of the application, the calculation formula of data preprocessing is:

[0071]

[0072] Wherein, μ and σ are the mean and standard deviation of the input and output respectively, x is the original data, x * Data satisfying the preset standardization condition.

[0073] In step S102, the time series data sample of the supercritical unit is constructed based on the data satisfying the preset standardization condition.

[0074] Specifically, the embodiment of the application can construct the time series data sample of the supercritical unit based on the data satisfying the preset standardization condition, enhance the key time series feature extraction through the attention mechanism, optimize the information flow through the residual connection, capture the multi-scale dynamic characteristics through the double branch structure, and significantly improve the fitting accuracy of the model to the complex nonlinear dynamics of the supercritical unit in the wide load range and the robustness to the working condition changes.

[0075] In step S103, an attention residual double branch gated recurrent unit model is constructed according to the time series data sample, for generating the model autoregressive prediction result of the supercritical unit.

[0076] It can be understood that the embodiment of the application can accurately capture the wide load dynamic characteristics of the supercritical unit, effectively overcome the exposure bias problem in autoregressive prediction, and improve the long-term prediction stability and generalization ability of the model. The new modeling method has important theoretical significance and engineering application value.

[0077] In actual implementation process, the embodiment of the application can construct an attention residual double branch gated recurrent unit model according to the time series data sample, for generating the model autoregressive prediction result of the supercritical unit.

[0078] The wide load adaptability in the embodiments of the present application is strong: can be effectively applied to the wide load operation range of 30%-100% of the supercritical unit, and can maintain good prediction performance and fast self-adjusting ability when there is data anomaly or mutation, thereby providing a high-precision and high-reliability dynamic model basis for intelligent operation optimization, model prediction control and digital twin system construction of the supercritical unit.

[0079] Optionally, in an embodiment of the present application, the attention residual double-branch gate recurrent unit model is constructed according to the time series data samples, which includes: weighting the input sequence by using a self-attention mechanism to determine the time step features meeting a preset condition; capturing long and short term dependencies in the input sequence based on the time step features meeting the preset condition, and normalizing the output of the two-layer gate recurrent unit based on the long and short term dependencies to generate the final output features of the two-layer gate recurrent unit; inputting the final output features into the main regression branch and the auxiliary regression branch of the supercritical unit, and constructing the attention residual double-branch gate recurrent unit model according to the main regression branch and the auxiliary regression branch.

[0080] In actual execution process, the embodiment of the present application can design a neural network model structure containing attention enhancement and residual fusion, double-layer GRU time series modeling and double-branch regression prediction. The attention enhancement and residual fusion module includes: a self-attention mechanism (Self-Attention): weighting the input sequence to determine the time step features meeting a preset condition, thereby enhancing the expression of important time step features.

[0081] Further, the double-layer GRU time series modeling module in the embodiment of the present application includes two layers of gate recurrent units (GRU): capturing long and short term dependencies in the input sequence, and normalizing the output of the two-layer gate recurrent unit based on the long and short term dependencies to generate the final output features of the two-layer gate recurrent unit; inputting the final output features into the main regression branch and the auxiliary regression branch of the supercritical unit, and constructing the attention residual double-branch gate recurrent unit model according to the main regression branch and the auxiliary regression branch.

[0082] The embodiment of the present application can improve the model prediction accuracy and robustness, enhance the key time series feature extraction through the attention mechanism, optimize the information flow through the residual connection, capture the multi-scale dynamic characteristics through the double-branch structure, and significantly improve the fitting accuracy of the model to the complex nonlinear dynamics of the supercritical unit in the wide load range and the robustness to the working condition changes.

[0083] In an embodiment of the present application, the formula for weighting the input sequence by using a self-attention mechanism is:

[0084]

[0085] where Q, K, V are the query, key, value matrices generated by linear transformation of input sequence, d k is the dimension of the key vector, used to scale the dot product to avoid numerical instability; the output is the weighted feature representation, highlighting the relevance of different time steps in the sequence.

[0086] The sigmoid layer performs a non-linear transformation on the attention weights:

[0087]

[0088] The softmax layer normalizes the attention scores:

[0089]

[0090] The residual connection layer (ResidualAdd) adds the output of the self-attention mechanism to the original input, preserving the original information and facilitating gradient propagation.

[0091] The two-layer gated recurrent unit consists of an update gate z t and a reset gate r t , whose core calculation formula is:

[0092] z t = σ s (W z · [h t-1 , x t ] + b z )

[0093] r t = σ s (W r · [h t-1 , x t ] + b r )

[0094]

[0095] where x t is the current input, h t-1 is the previous hidden state, W, U, b are learnable parameters, σ is the sigmoid function, and ⊙ is the element-wise product.

[0096] The batch normalization layer (BatchNorm) normalizes the output of each GRU layer, accelerating training and improving stability.

[0097]

[0098] The dropout layer randomly drops some neurons after the GRU layer to prevent overfitting.

[0099] r~Bernoulli(1-p),Output=r⊙x

[0100] Furthermore, the dual-branch regression prediction module includes:

[0101] The final output features of the GRU module are input into the main regression branch and the auxiliary regression branch respectively.

[0102] Each branch consists of multiple fully connected layers (the number of neurons in the middle layer is set to 128 and 64 respectively, and the ReLU activation function is used).

[0103] Fusion layer: The outputs of the two branches are fused through the addition layer.

[0104] Output layer: Finally, a linear regression layer is used to output the predicted values ​​of main steam pressure, unit load and separator temperature.

[0105] Further, Figure 2 A flow chart of a supercritical unit wide load modeling method provided in an embodiment of the present application applied in the model application stage.

[0106] like Figure 2 As shown, the supercritical unit wide load modeling method includes the following steps:

[0107] In step S201, the control variables of the supercritical unit are obtained.

[0108] Among them, the embodiment of the present application can obtain the control variables of the supercritical unit: throttle valve opening (u1), coal feeding rate (u2), and water feeding rate (u3), and the control variables are used as model inputs.

[0109] In step S202, the control variable is input into a pre-constructed attention residual two-branch gated recurrent unit model to train the pre-constructed attention residual two-branch gated recurrent unit model based on a preset two-stage planned sampling training strategy to generate a model autoregressive prediction result of the supercritical unit, wherein the pre-constructed attention residual two-branch gated recurrent unit model consists of a main regression branch and an auxiliary regression branch of the supercritical unit.

[0110] Among them, the embodiment of the present application can input the control variables into the pre-built attention residual two-branch gated recurrent unit model to train the pre-built attention residual two-branch gated recurrent unit model based on a preset two-stage planned sampling training strategy to generate the model autoregressive prediction results of the supercritical unit, providing a high-precision and high-reliability dynamic model foundation for the intelligent operation optimization, model predictive control and digital twin system construction of the supercritical unit.

[0111] Optionally, in an embodiment of the present application, based on a preset two-stage plan sampling training strategy, a pre-constructed attention residual double-branch gated recurrent unit model is trained to generate a model autoregressive prediction result of a supercritical unit, including: in a first stage of training of the attention residual double-branch gated recurrent unit model, obtaining an input of each time step, and determining an actual last time step observation value according to the input; in a second stage of training of the attention residual double-branch gated recurrent unit model, obtaining a first preset probability dynamically adjusted with a training iteration number; determining the actual last time step observation value according to the first preset probability, and determining a prediction value of the attention residual double-branch gated recurrent unit model at the last time step according to a second preset probability; and based on the actual last time step observation value and the prediction value of the attention residual double-branch gated recurrent unit model at the last time step, generating the model autoregressive prediction result of the supercritical unit.

[0112] It can be understood that the first preset probability is θ, and the second preset probability is 1-θ.

[0113] In the embodiment of the present application, a two-stage plan sampling training strategy can be designed to alleviate the exposure bias and improve the model autoregressive prediction capability. In the first stage, teacher forced training is adopted, and in this stage, the input of the GRU unit at each time step is always the actual last time step observation value during model training. In the second stage, plan sampling training is adopted, and in this stage, the input of the GRU unit selects the actual last time step observation value with a probability θ that is dynamically adjusted with the training iteration number, and selects the prediction value of the model itself at the last time step with a probability of 1-θ. The sampling probability θ gradually decays from an initial value close to 1 (biased towards teacher forced) to a value close to 0 (biased towards complete autoregression).

[0114] The embodiment of the present application can effectively alleviate the exposure bias and enhance the long-term autoregressive prediction stability: the two-stage plan sampling training strategy makes the model gradually adapt to the accumulation of its own prediction error during training, effectively alleviates the exposure bias problem caused by traditional teacher forced training, and significantly improves the stability and generalization ability of the model in the long-term autoregressive prediction task.

[0115] In an embodiment of the present application, the calculation formula of the first preset probability is:

[0116]

[0117] Wherein, t is the current time step, n is the size of the input window, and N is the total time step of the training data.

[0118] During training, a suitable optimizer (such as Adam) and loss function (such as mean square error MSE) are used, and learning rate decay, early stopping and other mechanisms are set.

[0119] The application can effectively capture the complex dynamic characteristics of the supercritical unit, significantly alleviate the exposure bias problem in autoregressive prediction, greatly improve the prediction accuracy, long-term stability and robustness of the model under wide load conditions, and provide a reliable model basis for realizing intelligent operation and optimization control of the supercritical unit.

[0120] Specifically, the supercritical unit wide load modeling method in the embodiments of the application can be combined with Figures 3 to 9 The working principle of the supercritical unit wide load modeling method in the embodiments of the application is described in detail with a specific embodiment.

[0121] As shown in Figure 3 The attention residual double-branch GRU model structure is the basis for realizing superior prediction performance. As shown in Figure 3 The input data sequence first passes through the attention enhancement and residual fusion module, wherein the self-attention mechanism dynamically allocates the weights of different time steps, and the residual connection ensures effective information transmission and gradient stability. Subsequently, the feature sequence enters the double-layer GRU time series modeling module, the core of which is the GRU unit shown in Figure 3 (The internal structure diagram of the GRU unit).

[0122] As shown in Figure 4 The GRU unit can effectively control the flow and fusion of historical information and current input through its update gate z t and reset gate r t , thereby capturing long-term and short-term dependencies in the sequence, which is crucial for modeling systems such as supercritical units that have complex dynamics and inertia. The output of the double-layer GRU is then passed through a double-branch regression prediction module, which extracts features and maps them through two parallel fully connected network branches, and finally fuses the outputs of the two branches to obtain predictions of the main steam pressure, unit load and separator temperature. This multi-path, multi-level feature learning and fusion helps to improve the accuracy and robustness of the prediction.

[0123] As shown in Figure 5 In the second stage of the planned sampling training, the input of the model at each time step has a certain probability p from the real previous time observation value, and has a probability of 1-p from the model's own previous time prediction output. This probability gradually decays from high (close to 1) to low (close to 0) as the number of training iterations increases. This mechanism allows the model to gradually "adapt" to the errors that may be caused by its own prediction during the training phase, thereby effectively alleviating the "exposure bias" problem and significantly enhancing the stability and generalization ability of the model under pure autoregressive prediction mode.

[0124] As shown in Figure 6As shown, the GRU3 model proposed by the present application is superior to TCN-MLP, BiLSTM, Transformer and other GRU variants (GRU1 without attention residual, double-branch regression module, GRU2 without double-branch regression module, GRU4 without two-stage training) in long-term autoregressive prediction. For example, compared with the baseline GRU1 model, the RMSE of GRU3 for main steam pressure, unit load and separator temperature is reduced by 18.7%, 48.1% and 4.7% respectively; compared with the Transformer model, the RMSE is reduced by as high as 87.8% (y1), 90.8% (y2) and 84.8% (y3). These significant performance improvements fully demonstrate the significant beneficial effects of the attention residual double-branch GRU model structure and the two-stage plan sampling training strategy proposed by the present application in solving the wide load dynamic modeling of supercritical units, especially in improving the long-term autoregressive prediction accuracy and stability.

[0125] As shown in Figure 7a , 7b and 7c, the absolute error distribution diagram of the GRU model with different training modules proposed by the present application for long-term autoregressive prediction on real data clearly shows the statistical distribution of the absolute error of the GRU3 model of the present application compared with other GRU variants in long-term autoregressive prediction. As can be seen from the figure, the absolute error distribution curve representing the GRU3 model of the present application is more concentratedly distributed near zero, with a higher peak and closer to zero, and the standard deviation of the distribution is also relatively smaller. Specifically, when predicting the main steam pressure y1, the GRU3 model has an error mean μ of 0.10901 MPa, a standard deviation σ of 0.54798 MPa, and a 95% confidence interval of [-0.99, 1.2] MPa, all of which are better than or equal to other comparative models. When predicting the unit load y2, the error mean μ of GRU3 is -0.17412 MW, and the standard deviation σ is 7.09909 MW, which is significantly smaller than GRU1 and GRU2. When predicting the separator temperature y3, the error standard deviation σ of GRU3 is also the smallest, which is 5.70527℃, and the error mean μ is -0.68757℃, which is within an acceptable range. These statistical results directly prove the superiority of the model in reducing systematic bias and random fluctuations, and the prediction results are more accurate and stable.

[0126] As shown in Figure 8As shown, the results of the long-term autoregressive prediction of the GRU3 model proposed in this application on real data are compared with the true value, showing the comparison between the prediction curve and the real operating data curve. As shown in the figure, whether it is the main steam pressure, unit load or separator temperature, the prediction curve of the GRU3 model is highly consistent with the true value curve in terms of overall trend, and accurately tracks the dynamic changes of the unit in different stages such as load increase, decrease and stable operation. Even at some points where there are sudden changes in data or abnormal fluctuations, although the prediction curve may have a short-term deviation, the model can quickly adjust and re-close to the true trajectory, showing good dynamic response and robustness. This is due to the model structure's ability to capture complex dynamics and the anti-interference and self-adaptation capabilities brought by planned sampling training.

[0127] like Figure 9 As shown, the relative error curve of the GRU3 model proposed in this application for long-term autoregressive prediction on real data further quantifies the prediction performance of the GRU3 model from the perspective of relative error. As shown in the figure, for the main steam pressure y1 and the unit load y2, the relative error can be controlled within the range of 0% to 10% in most time periods. For the separator temperature y3, the relative error is even more ideal, and remains stable below 5%. As can be seen from the figure, even if peaks in the relative error occur at certain moments due to data anomalies or insufficient instantaneous capture of specific dynamics by the model, these error peaks are short-lived, and then the relative error can quickly drop and return to a lower level. This fully demonstrates that the GRU3 model of the present invention not only has a high average prediction accuracy in long-term autoregressive prediction, but also has the self-stabilizing ability to converge quickly after a disturbance occurs, ensuring the long-term reliability of the prediction results.

[0128] According to the supercritical unit wide load modeling method proposed in the embodiment of the present application, the attention mechanism is used to enhance the extraction of key time series features, the residual connection optimizes the information flow, and the dual-branch structure captures multi-scale dynamic characteristics. This significantly improves the model's fitting accuracy for the complex nonlinear dynamics of supercritical units within a wide load range and its robustness to operating condition changes, effectively alleviates exposure bias, enhances long-term autoregressive prediction stability, and has superior overall performance and strong adaptability to wide loads. This solves the problem of low autoregressive prediction accuracy, poor long-term stability, and limited generalization ability of existing supercritical generator set data-driven models under wide load conditions due to factors such as exposure bias.

[0129] Next, a supercritical unit wide load modeling device proposed according to an embodiment of the present application will be described with reference to the accompanying drawings.

[0130] Figure 10 It is a structural diagram of the supercritical unit wide load modeling device applied to the model building stage in an embodiment of the present application.

[0131] likeFigure 10 As shown, the supercritical unit wide load modeling device 10 includes an acquisition module 100, a construction module 200, and a generation module 300.

[0132] Specifically, the acquisition module 100 is configured to acquire operation data of the supercritical unit, and perform data preprocessing on the operation data to generate data satisfying a preset standardization condition.

[0133] The construction module 200 is configured to construct time series data samples of the supercritical unit based on the data satisfying the preset standardization condition.

[0134] The generation module 300 is configured to construct an attention residual double-branch gate recurrent unit model according to the time series data samples, so as to generate a model autoregressive prediction result of the supercritical unit.

[0135] Optionally, in an embodiment of the present application, the acquisition module 100 includes an acquisition unit and a preprocessing unit.

[0136] The acquisition unit is configured to acquire any control variable of a throttle opening degree, a coal feeding rate, and a water feeding rate of the supercritical unit, and acquire any state variable of a main steam pressure, a unit load, and a separator temperature of the supercritical unit.

[0137] The preprocessing unit is configured to perform data preprocessing on the any control variable and the any state variable, so as to generate data satisfying a preset standardization condition.

[0138] Optionally, in an embodiment of the present application, the calculation formula of the data preprocessing is as follows:

[0139]

[0140] Wherein, μ and σ are the average value and the standard deviation of the input and the output respectively, x is the original data, x * is the data satisfying the preset standardization condition.

[0141] Optionally, in an embodiment of the present application, the generation module 300 includes a weighting unit, a capturing unit, and a construction unit.

[0142] The weighting unit is configured to weight the input sequence by using a self-attention mechanism, so as to determine a time step feature satisfying a preset condition.

[0143] The capturing unit is configured to capture a long-short term dependency relationship in the input sequence based on the time step feature satisfying the preset condition, and perform normalization processing on the output of the two-layer gate recurrent unit based on the long-short term dependency relationship, so as to generate a final output feature of the two-layer gate recurrent unit.

[0144] A construction unit is used to input the final output features into the main regression branch and the auxiliary regression branch of the supercritical unit, and to construct an attention residual dual-branch gated recurrent unit model based on the main regression branch and the auxiliary regression branch.

[0145] Optionally, in one embodiment of the present application, the formula for weighting the input sequence using the self-attention mechanism is:

[0146]

[0147] Among them, Q, K, and V are query, key, and value matrices generated by linear transformation of the input sequence, respectively. k is the dimension of the key vector;

[0148] The calculation formula of the two-layer gated recurrent unit is:

[0149] z t =σ s (W z ·[h t-1 , x t ]+b z )

[0150] r t =σ s (W r ·[h t-1 , x t ]+b r )

[0151]

[0152] Among them, x t is the current input, h t-1 is the hidden state at the previous moment, W, U, b are learnable parameters, σ is the Sigmoid function, and ⊙ is the element-by-element product.

[0153] Figure 11 It is a structural diagram of the supercritical unit wide load modeling device in an embodiment of the present application applied in the model application stage.

[0154] like Figure 11 As shown, the supercritical unit wide load modeling device 20 includes: an acquisition module 400 and a prediction module 500.

[0155] The acquisition module 400 is used to obtain the control variables of the supercritical unit.

[0156] The prediction module 500 is configured to input the control variable into the pre-constructed attention residual double-branch gated recurrent unit model, train the pre-constructed attention residual double-branch gated recurrent unit model based on a preset two-stage plan sampling training strategy, and generate a model autoregressive prediction result of the supercritical unit.

[0157] Optionally, in an embodiment of the present application, the prediction module 500 comprises a determination unit, a probability acquisition unit, a prediction value determination unit, and a result generation unit.

[0158] The determination unit is configured to, in a first stage of training of the attention residual double-branch gated recurrent unit model, acquire an input of each time step, and determine an actual previous time observation value according to the input.

[0159] The probability acquisition unit is configured to, in a second stage of training of the attention residual double-branch gated recurrent unit model, acquire a first preset probability dynamically adjusted with a training iteration number.

[0160] The prediction value determination unit is configured to determine the actual previous time observation value according to the first preset probability, and determine a prediction value of the attention residual double-branch gated recurrent unit model at a previous time according to a second preset probability.

[0161] The result generation unit is configured to generate the model autoregressive prediction result of the supercritical unit based on the actual previous time observation value and the prediction value of the attention residual double-branch gated recurrent unit model at the previous time.

[0162] Optionally, in an embodiment of the present application, a calculation formula of the first preset probability is as follows:

[0163]

[0164] Wherein, t is a current time step, n is a size of an input window, and N is a total time step of training data.

[0165] It should be noted that the foregoing explanation and description of the embodiment of the supercritical unit wide load modeling method also apply to the supercritical unit wide load modeling device of the embodiment, which will not be described here again.

[0166] The supercritical unit wide load modeling device provided by the embodiment of the present application significantly improves the fitting accuracy of the model on the complex nonlinear dynamics of the supercritical unit in a wide load range and the robustness to working condition changes, effectively alleviates the exposure bias, enhances the long-term autoregressive prediction stability, and has superior comprehensive performance and strong wide load adaptability. Thus, the problem of low autoregressive prediction accuracy, poor long-term stability and limited generalization ability of the existing data-driven model of the supercritical generator under wide load conditions due to exposure bias and other factors is solved.

[0167] In the description of the present specification, the description referring to the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or N embodiments or examples. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the present specification and the features of the different embodiments or examples without contradiction.

[0168] In addition, the terms "first", "second" are only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "N" is at least two, for example, two, three, etc., unless otherwise specifically limited.

[0169] Any process or method descriptions in flow charts or otherwise described herein can be understood as representing code modules, segments, or portions of code that include one or more executable instructions for implementing the specified logic functions or processes, and the preferred embodiments of the present application also include additional implementation examples, in which the functions can be performed in different orders, in substantially simultaneous fashion, or in reverse order, depending on the functionality involved, as will be understood by those skilled in the art.

Claims

1. A supercritical unit wide load modeling method, characterized in that: Applied to the model building phase, it includes the following steps: Collecting operating data of the supercritical unit and performing data preprocessing on the operating data to generate data that meets preset standardization conditions; Constructing a time series data sample of the supercritical unit based on the data that meets the preset standardization conditions; An attention residual dual-branch gated recurrent unit model is constructed based on the time series data samples to generate a model autoregressive prediction result of the supercritical unit.

2. The method according to claim 1, characterized in that The collecting of the operating data of the supercritical unit and performing data preprocessing on the operating data to generate data that meets preset standardization conditions include: Collecting any control variable of the supercritical unit including the throttle valve opening, coal feeding rate, and water feeding rate, and collecting any state variable of the supercritical unit including the main steam pressure, unit load, and separator temperature; Data preprocessing is performed on any of the control variables and any of the state variables to generate data that meets the preset standardization conditions.

3. The method according to claim 1, characterized in that The calculation formula for the data preprocessing is: Among them, μ and σ are the mean and standard deviation of input and output respectively, x is the original data, and x * The data satisfying the preset standardization conditions.

4. The method according to claim 1, wherein The constructing of the attention residual dual-branch gated recurrent unit model according to the time series data sample includes: The self-attention mechanism is used to weight the input sequence to determine the time-step features that meet the preset conditions; Based on the time step features that meet the preset conditions, capturing the long-term and short-term dependencies in the input sequence, and normalizing the outputs of the two-layer gated recurrent unit based on the long-term and short-term dependencies to generate final output features of the two-layer gated recurrent unit; The final output feature is input into the main regression branch and the auxiliary regression branch of the supercritical unit, and the attention residual dual-branch gated recurrent unit model is constructed according to the main regression branch and the auxiliary regression branch.

5. The method according to claim 4, characterized in that The formula for weighting the input sequence using the self-attention mechanism is: Among them, Q, K, and V are query, key, and value matrices generated by linear transformation of the input sequence, respectively. k is the dimension of the key vector; The calculation formula of the two-layer gated recurrent unit is: z t =σ S (W z ·[h t-1 ,x t ]+b z ) r t =σ s (W r ·[h t-1 ,x t ]+b r ) Among them, x t is the current input, h t-1 is the hidden state at the previous moment, W, U, b are learnable parameters, σ is the Sigmoid function, and ⊙ is the element-by-element product.

6. A supercritical unit wide load modeling method, characterized in that: Applied to the model application phase, it includes the following steps: Obtain control variables of supercritical units; The control variable is input into a pre-constructed attention residual two-branch gated recurrent unit model to train the pre-constructed attention residual two-branch gated recurrent unit model based on a preset two-stage planned sampling training strategy to generate a model autoregressive prediction result of the supercritical unit, wherein the pre-constructed attention residual two-branch gated recurrent unit model consists of a main regression branch and an auxiliary regression branch of the supercritical unit.

7. The method according to claim 6, characterized in that The method of training the pre-built attention residual dual-branch gated recurrent unit model based on a preset two-stage planned sampling training strategy to generate a model autoregressive prediction result of the supercritical unit includes: In the first stage of training the attention residual dual-branch gated recurrent unit model, obtaining an input for each time step and determining an actual observation value at the same moment based on the input; In the second stage of training the attention residual dual-branch gated recurrent unit model, obtaining a first preset probability that is dynamically adjusted according to the number of training iterations; Determining the actual observation value at the previous moment according to the first preset probability, and determining the predicted value of the attention residual dual-branch gated recurrent unit model at the previous moment according to the second preset probability; Based on the actual observation value at the previous moment and the predicted value of the attention residual dual-branch gated recurrent unit model at the previous moment, a model autoregressive prediction result of the supercritical unit is generated.

8. The method according to claim 7, characterized in that The calculation formula of the first preset probability is: Where t is the current time step, n is the size of the input window, and N is the total time step of the training data.

9. A supercritical unit wide load modeling device, characterized in that: Applied in the model building phase, including: An acquisition module is used to collect operating data of the supercritical unit and perform data preprocessing on the operating data to generate data that meets preset standardization conditions; A construction module, configured to construct a time series data sample of the supercritical unit based on the data meeting the preset standardization conditions; A generation module is used to construct an attention residual double-branch gated recurrent unit model based on the time series data sample to generate a model autoregressive prediction result of the supercritical unit.

10. A supercritical unit wide load modeling device, characterized in that: Applied in the model application phase, including: An acquisition module is used to obtain the control variables of the supercritical unit; A prediction module is used to input the control variable into a pre-constructed attention residual two-branch gated recurrent unit model to train the pre-constructed attention residual two-branch gated recurrent unit model based on a preset two-stage planned sampling training strategy to generate a model autoregressive prediction result of the supercritical unit, wherein the pre-constructed attention residual two-branch gated recurrent unit model consists of a main regression branch and an auxiliary regression branch of the supercritical unit.

Citation Information

Patent Citations

  • Intelligent prediction method for power generation load and heat supply of supercritical unit

    CN111027258A

  • Sequence-to-sequence power load prediction method based on multi-dimensional gating circulation unit

    CN116384572A

  • Systems, methods, devices, and platforms for industrial internet of things

    WO2024155584A1