Thermal power generating unit load prediction method and device, computer equipment, readable storage medium and program product
By detecting the similarity between the operating status characteristics of thermal power units and candidate operating conditions, the parameters of the load forecasting model can be quickly adjusted, solving the problem of model gap caused by sudden changes in operating conditions in existing technologies, and ensuring the accuracy and stability of power grid dispatch.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHN ENERGY NEW ENERGY TECHNOLOGY RESEARCH INSTITUTE CO LTD
- Filing Date
- 2025-12-08
- Publication Date
- 2026-05-05
AI Technical Summary
Existing technologies require long-term samples for load forecasting models to converge when thermal power units are under maintenance or coal-fired conditions change abruptly, resulting in a model downtime of several hours to several days, which affects the accuracy of power grid dispatching decisions.
By detecting the similarity between the current operating status characteristics of thermal power units and candidate operating conditions, model parameter tuning and matching are performed. A small amount of historical load sample data is used to adjust the parameters of the pre-trained model, and a target prediction model that is suitable for the current state is quickly obtained.
It enables the immediate output of accurate load forecast results after sudden changes in operating conditions, avoiding forecast gaps and ensuring the accuracy and stability of power grid dispatch.
Smart Images

Figure CN121984079A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of unit load forecasting, and in particular to a method, apparatus, computer equipment, readable storage medium, and program product for forecasting the load of thermal power units. Background Technology
[0002] Currently, the common approach for load forecasting of thermal power units is to use a technical system of "independent modeling of units" plus "periodic retraining". That is, a separate model is built for each thermal power unit and retrained and updated regularly.
[0003] However, in this approach, the model parameters need to rely on long-term samples to converge. This means that once the thermal power unit is under maintenance or the coal-fired conditions change abruptly, the entire training process must be restarted. The model downtime can last for hours or even days. During this period, the load forecasting function will fail, which can easily lead to deviations in power grid dispatching decisions. Summary of the Invention
[0004] Therefore, it is necessary to provide a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for predicting the load of thermal power units that can ensure the accuracy of power grid dispatching decisions, in response to the above-mentioned technical problems.
[0005] Firstly, this application provides a method for predicting the load of thermal power units, including:
[0006] In response to load forecasting commands for thermal power units, the current operating status characteristics of the thermal power units are obtained;
[0007] Detect the similarity between the characteristics of the operating status and the characteristics of the candidate operating conditions;
[0008] Based on similarity, model parameter tuning is matched to obtain the model parameter adjustment amount that matches the similarity.
[0009] Obtain historical load sample data of thermal power units under the current shift;
[0010] Based on historical load sample data and model parameter adjustment, the parameters of the pre-trained load prediction model are adjusted to obtain a target prediction model that is adapted to the current operating status of the thermal power unit.
[0011] By analyzing the real-time operating data of thermal power units using a target prediction model, the load prediction results of the thermal power units are obtained.
[0012] In one embodiment, detecting the similarity between operating state features and operating condition features of candidate operating conditions includes:
[0013] The difference between the characteristics of the operating status and the characteristics of the candidate operating conditions is detected.
[0014] The difference values are weighted and summed to obtain the weighted distance between the operating state characteristics and the operating condition characteristics of the candidate operating conditions;
[0015] Based on weighted distance, the similarity between the operating state characteristics and the operating condition characteristics of the candidate operating conditions is determined.
[0016] In one embodiment, model parameter tuning matching is performed based on similarity to obtain model parameter adjustment amounts that match the similarity, including:
[0017] If the similarity is not less than the preset similarity threshold, then query the model parameter adjustment amount that matches the candidate operating condition;
[0018] If the similarity is less than the preset similarity threshold, the corresponding model parameter adjustment amount will be generated based on the running status characteristics.
[0019] In one embodiment, the pre-trained load prediction model includes a lightweight adaptation layer; based on historical load sample data and model parameter adjustments, the parameters of the pre-trained load prediction model are adjusted, including:
[0020] Based on the model parameter adjustment amount, the model parameters of the lightweight adaptation layer are adjusted to obtain the parameter-tuned lightweight adaptation layer.
[0021] Historical load sample data is used as training samples to update the gradient of the hyperparameter-tuned lightweight adaptation layer.
[0022] In one embodiment, the target prediction model includes a time domain channel and a frequency domain channel; by analyzing the real-time operating data of the thermal power unit through the target prediction model, the load prediction results of the thermal power unit are obtained, including:
[0023] By analyzing the real-time operating data of thermal power units through the time domain channel, the first prediction result reflecting the long-term load change trend of thermal power units is obtained.
[0024] By analyzing real-time operating data through frequency domain channels, a second prediction result reflecting the short-term load disturbance of thermal power units is obtained;
[0025] By combining the first and second forecast results, the load forecast results for thermal power units are obtained.
[0026] In one embodiment, the thermal power unit load forecasting method further includes:
[0027] Obtain feature vectors of multiple operating states of thermal power units within historical operating cycles;
[0028] The weighted distance between the feature vectors of each operating state is detected, and the multiple operating state feature vectors are grouped based on the weighted distance to obtain candidate operating conditions;
[0029] Detect the business risk level of candidate operating conditions, and construct the candidate operating conditions as training tasks based on the business risk level;
[0030] Based on the training task, the initial prediction model is trained to obtain the load prediction model.
[0031] Secondly, this application also provides a load prediction device for thermal power units, comprising:
[0032] The feature acquisition module is used to acquire the current operating status features of the thermal power unit in response to the load forecasting command for the thermal power unit;
[0033] The similarity detection module is used to detect the similarity between the operating status features and the operating condition features of the candidate operating conditions.
[0034] The parameter tuning matching module is used to perform model parameter tuning matching based on similarity, and obtain the model parameter adjustment amount that matches the similarity.
[0035] The load data acquisition module is used to acquire historical load sample data of thermal power units under the current shift;
[0036] The model update module is used to adjust the parameters of the pre-trained load prediction model based on historical load sample data and model parameter adjustment amounts, so as to obtain a target prediction model that is adapted to the current operating state of the thermal power unit.
[0037] The load forecasting module is used to analyze the real-time operating data of thermal power units through the target forecasting model to obtain the load forecasting results of the thermal power units.
[0038] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0039] In response to load forecasting commands for thermal power units, the current operating status characteristics of the thermal power units are obtained;
[0040] Detect the similarity between the characteristics of the operating status and the characteristics of the candidate operating conditions;
[0041] Based on similarity, model parameter tuning is matched to obtain the model parameter adjustment amount that matches the similarity.
[0042] Obtain historical load sample data of thermal power units under the current shift;
[0043] Based on historical load sample data and model parameter adjustment, the parameters of the pre-trained load prediction model are adjusted to obtain a target prediction model that is adapted to the current operating status of the thermal power unit.
[0044] By analyzing the real-time operating data of thermal power units using a target prediction model, the load prediction results of the thermal power units are obtained.
[0045] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0046] In response to load forecasting commands for thermal power units, the current operating status characteristics of the thermal power units are obtained;
[0047] Detect the similarity between the characteristics of the operating status and the characteristics of the candidate operating conditions;
[0048] Based on similarity, model parameter tuning is matched to obtain the model parameter adjustment amount that matches the similarity.
[0049] Obtain historical load sample data of thermal power units under the current shift;
[0050] Based on historical load sample data and model parameter adjustment, the parameters of the pre-trained load prediction model are adjusted to obtain a target prediction model that is adapted to the current operating status of the thermal power unit.
[0051] By analyzing the real-time operating data of thermal power units using a target prediction model, the load prediction results of the thermal power units are obtained.
[0052] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:
[0053] In response to load forecasting commands for thermal power units, the current operating status characteristics of the thermal power units are obtained;
[0054] Detect the similarity between the characteristics of the operating status and the characteristics of the candidate operating conditions;
[0055] Based on similarity, model parameter tuning is matched to obtain the model parameter adjustment amount that matches the similarity.
[0056] Obtain historical load sample data of thermal power units under the current shift;
[0057] Based on historical load sample data and model parameter adjustment, the parameters of the pre-trained load prediction model are adjusted to obtain a target prediction model that is adapted to the current operating status of the thermal power unit.
[0058] By analyzing the real-time operating data of thermal power units using a target prediction model, the load prediction results of the thermal power units are obtained.
[0059] The aforementioned method, apparatus, computer equipment, computer-readable storage medium, and computer program product for load forecasting of thermal power units, in response to load forecasting commands for thermal power units, acquire the current operating status characteristics of the thermal power units; detect the similarity between the operating status characteristics and the operating condition characteristics of candidate operating conditions; perform model parameter tuning matching based on the similarity to obtain the model parameter adjustment amount that matches the similarity; acquire historical load sample data of the thermal power units under the current shift; adjust the parameters of the pre-trained load forecasting model based on the historical load sample data and the model parameter adjustment amount to obtain a target forecasting model that is adapted to the current operating status of the thermal power units; and analyze the real-time operating data of the thermal power units through the target forecasting model to obtain the load forecasting results of the thermal power units. Thus, this solution, by detecting the similarity between the current operating status characteristics of the thermal power units and the characteristics of candidate operating conditions, can immediately locate the most matching operating condition prototype after a sudden change in operating conditions. This process does not rely on long-term data accumulation and can be completed solely through feature comparison, greatly shortening the time required for model adjustment. Matching model parameter adjustments based on similarity results essentially transfers validated parameter optimization experience from similar operating conditions to the current operating conditions, eliminating the need to explore parameter convergence directions from scratch and significantly reducing parameter adaptation time. Furthermore, this solution uses a pre-trained load forecasting model as a foundation, combined with a small amount of load sample data from the current shift, to perform targeted fine-tuning of the model. This allows for rapid parameter convergence without requiring full retraining, eliminating the dependence of traditional model training on long-term samples and effectively removing prediction gaps after sudden changes in operating conditions. The target forecasting model can immediately output accurate load forecast results based on real-time operational data, thereby avoiding grid dispatch decision deviations caused by prediction gaps and ensuring the accuracy of thermal power dispatch and the stability of grid operation. Attached Figure Description
[0060] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0061] Figure 1 This is an application environment diagram of the thermal power unit load prediction method in one embodiment;
[0062] Figure 2 This is a flowchart illustrating a method for predicting the load of thermal power units in one embodiment;
[0063] Figure 3 This is a flowchart illustrating the process of similarity matching between operating conditions in one embodiment;
[0064] Figure 4 This is a flowchart illustrating the process of determining the adjustment amount of model parameters in one embodiment;
[0065] Figure 5 This is a schematic diagram of the model update process in one embodiment;
[0066] Figure 6 This is a flowchart illustrating the dual-channel prediction process in one embodiment;
[0067] Figure 7 This is a schematic diagram of the model training process in one embodiment;
[0068] Figure 8 This is a flowchart of a thermal power unit load prediction method in a specific embodiment;
[0069] Figure 9 This is a flowchart of the inner and outer loops during the training process in a specific embodiment;
[0070] Figure 10 This is a flowchart of the model update process in a specific embodiment;
[0071] Figure 11 This is an architecture diagram of a load forecasting model in a specific embodiment;
[0072] Figure 12 This is a structural block diagram of a thermal power unit load prediction device in one embodiment;
[0073] Figure 13 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0074] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0075] It should be noted that the terms "first," "second," etc., used in this application may be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more.
[0076] Currently, load forecasting for thermal power units generally adopts a technical system of "independent modeling of each unit" plus "periodic retraining." That is, a separate model is built for each thermal power unit, and it is periodically retrained and updated. The mainstream process is usually as follows: first, parameter fitting is completed on samples from six months to one year to obtain a baseline model for each unit. When the unit is under maintenance, the load level, or the coal mix changes, the baseline model is supplemented with new data from the last few days to several weeks, or the entire model is replaced. This approach emphasizes learning long-term trends, therefore requiring a large amount of computational resources in the parameter preparation and hyperparameter search stages, and the update cycle is often measured in days or weeks. During this period, the old model cannot accurately predict the load under new operating conditions, directly affecting the accuracy of power grid dispatch decisions.
[0077] To address the issue of prediction gaps, some methods attempt to feed back model prediction errors to the deployed network for online fine-tuning. However, these updates often only correct the low-dimensional output layer, failing to simultaneously retain sensitivity to high-frequency disturbances and robustness to long-term trends. Another optimization approach is multi-scale decomposition followed by reconstruction. While this method can consider both trends and details, it separates the time and frequency domains into different models or post-processing stages, resulting in weak coupling between scales, long inference paths, and difficulty in controlling latency in real-time scenarios. Yet another approach utilizes meta-learning or transfer learning. Meta-learning or transfer learning often uses daily windows as a few-sample benchmark, requiring repeated iterations to achieve adaptation, making it difficult to meet the deployment time requirements at the shift-level. In summary, mainstream methods either have excessive computational burdens or can only cover a single time scale, lacking the ability to quickly and comprehensively restore model load prediction accuracy with very little new data.
[0078] Therefore, this application provides a load forecasting method for thermal power units, which can quickly update the model based on the real-time operating status of thermal power units and a small amount of new sample data, thereby outputting accurate load forecasting results. This effectively avoids grid dispatching decision deviations caused by forecasting gaps and ensures the accuracy of thermal power dispatching and the stability of grid operation.
[0079] In some embodiments, the thermal power unit load forecasting method can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and IoT devices. Server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The data storage system can store the data that server 104 needs to process. The data storage system can be integrated on server 104 or located in the cloud or on other network servers.
[0080] Specifically, server 104 responds to a load forecasting command for thermal power units initiated by a user through terminal 102, or server 104 can trigger a load forecasting command itself. Upon receiving the command, it acquires the current operating status characteristics of the thermal power units and detects the similarity between these operating status characteristics and the operating condition characteristics of candidate operating conditions. Based on this similarity, it performs model parameter tuning and matching to obtain the model parameter adjustment amount that matches the similarity. Next, it acquires historical load sample data of the thermal power units under the current shift. Based on the historical load sample data and the model parameter adjustment amount, it adjusts the parameters of the pre-trained load forecasting model to obtain a target forecasting model adapted to the current operating status of the thermal power units. Finally, it analyzes the real-time operating data of the thermal power units through the target forecasting model to obtain the load forecasting results for the thermal power units.
[0081] In one exemplary embodiment, such as Figure 2 As shown, a method for predicting the load of thermal power units is provided, and this method is applied to... Figure 1 Taking server 104 as an example, the following steps are included:
[0082] Step S202: In response to the load forecasting command for the thermal power unit, obtain the current operating status characteristics of the thermal power unit.
[0083] The thermal power unit can be one or more units. The load forecasting command is used to instruct the thermal power unit to perform load forecasting. Its triggering timing can occur in any scenario, such as unit grid connection, batch replacement of coal, or completion of unit maintenance. Operating status characteristics refer to a set of features characterizing the current operating condition of the thermal power unit. In this embodiment, operating status characteristics may include at least one of the following: boiler type, coal volatile matter range, load ramp-up efficiency, maintenance status, and weather category. Boiler type refers to the design model of the boiler in the thermal power unit; coal volatile matter range is the volatile content range of the combustible components of the coal during the shift, such as 10%-20% or 30%-35%; load ramp-up efficiency refers to the speed at which the thermal power unit increases from low load to high load; maintenance status refers to whether the thermal power unit is in the recovery phase after maintenance; and weather category characterizes the weather conditions in the area where the thermal power unit is located, such as sunny, cloudy, or rainy.
[0084] For example, when a thermal power unit experiences any of the following events: grid connection, batch replacement of coal, or completion of unit maintenance, the user can trigger a load forecasting command for the thermal power unit. This command is transmitted to the server via a scheduling data bus, which is a unified data interface integrating SCADA (Supervisory Control and Data Acquisition), AGC (Automatic Generation Control), coal quality analysis, and weather station data. Through this scheduling data bus, the server can read boiler type and load ramp-up efficiency from the SCADA system, maintenance status from the AGC system, volatile matter range from the coal quality analysis system, and weather category from the weather station, according to a preset rolling time window (e.g., 30 minutes). This information can be arranged in a fixed order and mapped to a feature vector representing the operating status of the thermal power unit through an encoding function.
[0085] For example, for a thermal power unit u, the output signals of the SCADA system, AGC system, coal quality analysis system, and weather station are monitored through the scheduling data bus. These output signals can be sequentially written to the same sequence index I at the millisecond level. u In the middle. For sequence I u Perform a fixed 30-minute rolling capture to obtain data segments. Write start and end timestamps to the data segment, and record the segment's valid duration, which is the time difference between the start and end timestamps. Once generated, immediately extract attribute information such as boiler type, coal volatile matter range, load ramp-up efficiency, maintenance status, and weather category, and use a fixed encoding function. After converting these attribute information into codes, they are concatenated in a fixed order to obtain the running state feature vector. F u τ =[ f 1 , f 2 , f 3 , f 4 , f 5 ] ,in, , , , , , Indicates the boiler type. This represents the boiler type coding function. Indicates the volatile matter content range of coal. This represents the interval mapping function for volatile matter in coal. Indicates load ramp-up efficiency. This represents the load ramp-up efficiency quantization function. Indicates maintenance status. This represents the maintenance status mapping function. Indicates the weather category, This represents the meteorological category mapping function. After obtaining the operational state feature vector... It can be further merged with the thermal power unit's number u to form an operating status label (u, This label will be associated with the data fragment. It is also written to a pre-built cache.
[0086] Step S204: Detect the similarity between the operating status features and the operating condition features of the candidate operating conditions.
[0087] In this context, candidate operating conditions refer to representative and typical operating conditions mined from massive historical data during the model pre-training phase. Each candidate operating condition is uniquely identified by a condition feature vector, which is also a vector composed of the codes corresponding to various attribute information. Candidate operating conditions can be associated with a set of pre-trained model parameter adjustments and operating condition memory information, such as the typical load ramp-up duration and disturbance energy peak under that condition. Similarity refers to the degree of similarity between the current operating state characteristics of the thermal power unit and the operating condition characteristics of the candidate operating conditions; in other words, it is the degree of similarity between the operating state feature vector and the operating condition feature vector. In this embodiment, the similarity is calculated by combining attribute information differences and business weights.
[0088] For example, the current operating status feature vector of the thermal power unit is compared with the operating condition feature vector of each candidate operating condition to obtain the weighted similarity between the current operating status feature vector of the thermal power unit and the feature vector of each operating condition.
[0089] Step S206: Based on the similarity, perform model parameter tuning matching to obtain the model parameter adjustment amount that matches the similarity.
[0090] The model parameter adjustment amount is a set of parameter adjustments used to quickly adapt the pre-trained load prediction model to the current operating conditions of the generating unit. It is not a random initialization or large-scale retraining of all model parameters, but rather incremental parameters used to correct specific parts of the model, such as the lightweight adaptation layer. In this embodiment, the model parameter adjustment amount specifically consists of a set of scaling factors and bias terms. The scaling factors are used to scale the original weights in the lightweight adaptation layer element-wise, and the bias terms are used to offset the original biases in the lightweight adaptation layer element-wise. It can be understood that the model parameter adjustment amount is a direct product of the pre-training process. During the model pre-training phase, the model learns on thousands of typical operating conditions. After the model completes rapid fine-tuning of the inner loop on a specific operating condition, the calculated scaling factors and bias terms of the lightweight adaptation layer become the model parameter adjustment amount corresponding to that operating condition. Therefore, each candidate operating condition is associated with a set of model parameter adjustment amounts. The core purpose of model parameter adjustment is to achieve rapid model migration and avoid lengthy retraining. When the server identifies the current unit operating condition as a known operating condition, it can directly call the model parameter adjustment amount pre-stored for that operating condition, which is equivalent to instantly switching the model from a general mode to a special mode for the current operating condition. When facing an unknown operating condition, the server can quickly generate a set of initial model parameter adjustment amounts based on the characteristics of the current operating state.
[0091] In some embodiments, the load forecasting model in this embodiment is specifically a dual-channel time-frequency Transformer network. The model structure includes a time-domain channel, a frequency-domain channel, a fusion head, and a lightweight adaptation layer. The time-domain channel is used to capture long-term load change patterns, the frequency-domain channel is used to capture short-term load fluctuation characteristics, the fusion head is responsible for weighted superposition of the outputs of the two channels, and the lightweight adaptation layer is a neural network layer in the load forecasting model with fewer parameters, a relatively simple structure, and can be quickly adjusted.
[0092] For example, a similarity threshold can be preset. If the similarity between the current operating state characteristics and the operating condition characteristics of a candidate operating condition is not less than the similarity threshold, the current operating condition of the thermal power unit is considered to be a normal operating condition, and the server can directly query and obtain the model parameter adjustment amount bound to the candidate operating condition from the pre-stored mapping relationship. If the similarity is less than the similarity threshold, the current operating condition of the thermal power unit is considered to be an unconventional operating condition, and the server can generate a set of initial model parameter adjustment amounts based on the current operating state characteristics.
[0093] Step S208: Obtain historical load sample data of thermal power units under the current shift.
[0094] The current shift refers to the operating and management period of the thermal power unit, typically consisting of an 8-hour shift. However, this can be adjusted based on actual circumstances, and this embodiment does not impose any restrictions on this. Historical load sample data refers to the newly generated operating data of the thermal power unit within the current shift, and its data volume is much smaller than the sample data volume during model pre-training.
[0095] For example, according to a fixed time window such as 30 minutes, the operating data generated by the thermal power unit in the current shift is extracted as load sample data and used to participate in the subsequent adjustment of the load forecasting model.
[0096] Step S210: Based on historical load sample data and model parameter adjustment, adjust the parameters of the pre-trained load prediction model to obtain a target prediction model that is compatible with the current operating state of the thermal power unit.
[0097] The target prediction model is a model adapted to the current operating status of the unit after fine-tuning historical load sample data and model parameter adjustments.
[0098] For example, the model parameter adjustment is loaded into the lightweight adaptation layer of the pre-trained load prediction model, and the historical load sample data of the thermal power unit in the current shift is used as the training sample. One or several steps of gradient descent optimization are performed on the lightweight adaptation layer into which the model parameter adjustment is injected. In this way, the model can quickly adapt to the current operating characteristics of the unit with a very small amount of new data.
[0099] Step S212: Analyze the real-time operating data of the thermal power unit through the target prediction model to obtain the load prediction result of the thermal power unit.
[0100] Real-time operating data refers to the unit operating data used as input to the model. Unlike historical load sample data of thermal power units under the current shift, real-time operating data is the instantaneous operating data being generated at the current moment, while historical load sample data is past operating data. Load forecasting results can include two parts: long-term load change trends and short-term load disturbances.
[0101] For example, by inputting the real-time operating data of the thermal power unit into the adjusted target prediction model, the long-term load change trend and short-term load disturbance of the thermal power unit can be obtained.
[0102] In this embodiment, in response to a load forecasting command for a thermal power unit, the current operating status characteristics of the thermal power unit are acquired; the similarity between the operating status characteristics and the operating condition characteristics of candidate operating conditions is detected; based on the similarity, model parameter tuning matching is performed to obtain the model parameter adjustment amount that matches the similarity; historical load sample data of the thermal power unit under the current shift is acquired; based on the historical load sample data and the model parameter adjustment amount, the parameters of the pre-trained load forecasting model are adjusted to obtain a target forecasting model that is adapted to the current operating status of the thermal power unit; the real-time operating data of the thermal power unit is analyzed through the target forecasting model to obtain the load forecasting result of the thermal power unit. Thus, this embodiment, by detecting the similarity between the current operating status characteristics of the thermal power unit and the characteristics of candidate operating conditions, can immediately locate the most matching operating condition prototype after a sudden change in operating conditions. This process does not rely on long-term data accumulation and can be completed solely through feature comparison, greatly shortening the time required for model adjustment. Matching model parameter adjustments based on similarity results essentially transfers validated parameter optimization experience from similar operating conditions to the current operating condition, eliminating the need to explore parameter convergence direction from scratch and significantly reducing parameter adaptation time. Furthermore, this embodiment uses a pre-trained load forecasting model as a foundation, combined with a small amount of load sample data from the current shift, to perform targeted fine-tuning of the model. This allows for rapid parameter convergence without full retraining, eliminating the dependence of traditional model training on long-period samples and effectively removing prediction gaps after sudden changes in operating conditions. The target forecasting model can immediately output accurate load forecast results based on real-time operational data, thereby avoiding grid dispatch decision deviations caused by prediction gaps and ensuring the accuracy of thermal power dispatch and the stability of grid operation.
[0103] In one exemplary embodiment, such as Figure 3 As shown, the similarity between the operating state characteristics and the operating condition characteristics of the candidate operating conditions is detected, including:
[0104] Step S302: Detect the difference between the operating status characteristics and the operating condition characteristics of the candidate operating conditions.
[0105] Step S304: The difference values are weighted and summed to obtain the weighted distance between the operating status characteristics and the operating condition characteristics of the candidate operating conditions.
[0106] Step S306: Based on the weighted distance, determine the similarity between the operating status features and the operating condition features of the candidate operating conditions.
[0107] The difference value refers to the quantified difference between two operating state features and operating condition feature vectors in each attribute dimension. In this embodiment, different difference functions can be used to calculate the difference value according to different attribute types. Attribute types include discrete attributes and continuous attributes. Discrete attributes include boiler type and maintenance status, and the corresponding difference function is a binary classification judgment function. That is, for a certain attribute information, if its encoded value in the two vectors is exactly the same, the difference value can be assigned 0; otherwise, it is assigned 1. Continuous attributes include coal volatile matter range, load ramp-up efficiency, and weather category, and the corresponding difference function is a difference function. The difference value is the absolute difference or normalized absolute difference of the encoded values of the attribute information in the two vectors. The weighted distance is the result of summing the difference values according to the business weight, which is used to comprehensively reflect the overall difference between the two feature vectors. The business weight is a weight value set in advance according to the importance of the business. In this embodiment, the business weight is proportional to the degree of influence of the attribute information on the unit's operating safety. For example, the business weight of the maintenance identifier is higher than that of the coal volatile matter range, and the business weight of the coal volatile matter range is higher than that of other attributes.
[0108] For example, the current operating state feature vector of the thermal power unit is compared with the operating condition feature vector of a candidate operating condition, attribute by attribute. For each attribute, the corresponding difference function is called to calculate the difference value. At the same time, the predefined business weights for each attribute are read, and each difference value is multiplied by the corresponding business weight. Finally, all the product results are added together to obtain the weighted distance between the operating state feature vector and the operating condition feature vector. Then, this weighted distance is normalized relative to the sum of all business weights to calculate the similarity between the operating state feature vector and each operating condition feature vector.
[0109] In one example, the similarity calculation expression is as follows:
[0110] Sim F curr , C i =1- ∑ p=1 m γ p . δ p ( F curr [p], C i [p]) ∑ p=1 m γ p
[0111] in, This represents the feature vector indicating the current operating state of the thermal power unit. Let p represent the i-th candidate operating condition, m represent the attribute information, and m represent the number of attribute information. In this embodiment, m can be set to 5. This represents the difference function corresponding to attribute information p. This represents the business weight corresponding to attribute information p. For all similarities... Find the maximum value and compare the maximum similarity with the preset similarity threshold to match the model parameter adjustment amount.
[0112] In this embodiment, the difference between the current operating status characteristics of the thermal power unit and the operating condition characteristics of the candidate operating conditions is detected, and the difference values are weighted and summed to calculate the similarity between the two feature vectors based on the weighted distance. This provides a precise basis for the subsequent model to quickly reuse the pre-stored model parameter adjustment amount of similar operating conditions or to make targeted fine-tuning, effectively avoiding indiscriminate retraining, shortening the model adaptation time, and ensuring the matching accuracy in high-risk scenarios. This improves the response efficiency and prediction accuracy of the load prediction model after sudden changes in operating conditions.
[0113] In one exemplary embodiment, such as Figure 4 As shown, based on similarity, model parameter tuning is matched to obtain the model parameter adjustment amounts that match the similarity, including:
[0114] Step S402: If the similarity is not less than the preset similarity threshold, then query the model parameter adjustment amount that matches the candidate operating condition.
[0115] Step S404: If the similarity is less than the preset similarity threshold, then generate the corresponding model parameter adjustment amount based on the running status features.
[0116] The preset similarity threshold is a judgment value used to distinguish between normal and non-normal operating conditions. For example, 0.85. When the similarity is not less than the threshold, the thermal power unit is considered to be in normal operating condition. When the similarity is less than the threshold, the thermal power unit is considered to be in non-normal operating condition.
[0117] For example, the similarity between the current operating state characteristics of the thermal power unit and the operating condition characteristics of candidate operating conditions is compared with a preset similarity threshold. If the similarity is not less than the preset similarity threshold, the current operating condition of the thermal power unit is marked as a normal operating condition, and the operating condition-parameter mapping library is called. This mapping library stores the mapping relationship between each candidate operating condition and the corresponding model parameter adjustment amount. Through this mapping library, the model parameter adjustment amount matching the candidate operating condition can be found, and the model parameter adjustment amount can be directly written into the parameter injection slot of the lightweight adaptation layer of the model. If the similarity is less than the preset similarity threshold, the current operating condition of the thermal power unit is marked as an unconventional operating condition, and the current operating state characteristics of the thermal power unit are input into a pre-trained parameter generator. The parameter generator performs forward propagation calculation on the input operating state characteristics. For example, through a linear transformation layer, the operating state characteristics are mapped to the parameter dimensions required by the lightweight adaptation layer, and a set of initial scaling factors and bias terms are output as the model parameter adjustment amount for this time.
[0118] In some embodiments, after obtaining the model parameter adjustment amount, the model parameter adjustment amount can be combined with the original weights of the lightweight adaptation layer, such as: current weight of lightweight adaptation layer = original weight of lightweight adaptation layer * scaling factor + bias term, to quickly obtain the weight state adapted to the current operating condition. Next, the weight adjustment magnitude of each layer in the lightweight adaptation layer (such as the convolutional layer, fully connected layer, etc. of the adaptation layer) is quantified to obtain the weight adjustment percentage. Based on the weight adjustment percentage and the operating condition identifier (normal operating condition and non-normal operating condition), a freezing strategy is generated. Specifically, if the thermal power unit is in a normal operating condition and the maximum weight adjustment percentage is not higher than a preset threshold, it means that the pre-stored model parameter adjustment amount can adapt well to the current operating condition, and there is no need to modify the model structure too much. Therefore, the backbone layer, i.e., the time domain channel, the frequency domain channel, and the fusion head can be frozen, and only the injector weights are retained as trainable. If the thermal power unit is operating under unconventional conditions, or if the maximum weight adjustment percentage is higher than the preset threshold, it indicates that the model parameter adjustment amount and the lightweight adaptation layer have low compatibility, and more layers need to participate in the optimization. Therefore, the time domain channel and frequency domain channel can be frozen, the fusion head and output layer can be unfrozen, and the learning step size can be reduced to avoid parameter oscillation caused by excessive step size.
[0119] After the freezing strategy takes effect, a small batch of load data from the current shift is selected as training samples to fine-tune the lightweight adaptation layer. In this embodiment, the learning step size can be adjusted according to the following formula: η eff = η base [1- λ r max ]+k * 1 { flag = unconventional }
[0120] in, This indicates the learning step size for this update. Indicates the basic learning step size. This indicates the maximum weight adjustment percentage. represents the suppression coefficient, used to reduce the step size by a percentage as the weight is adjusted; k represents the compensation coefficient, which only takes effect under unconventional operating conditions. It is an indicator function; this term is equal to 1 only under abnormal operating conditions, and 0 otherwise.
[0121] After each parameter update, the system analyzes the costs incurred by the current model in three key business metrics: missed peak load predictions, load ramp-up errors, and forecast window duration. This allows for a quantitative assessment of the new model's business risk. If any metric exceeds its corresponding threshold, the model is automatically reverted to its pre-update state and falls back to a more conservative parameter-only injection mode. If all metrics remain within their respective thresholds, the current model is adopted as the final target load forecasting model for subsequent real-time load forecasting. Simultaneously, changes to the model itself, such as weight increments, are recorded and visualized in real-time.
[0122] In this embodiment, under normal operating conditions, the model parameter adjustment amount stored in similar operating conditions is reused. Essentially, this reuses the adaptation experience of historical operating conditions, avoiding the need to calculate the model parameter adjustment amount from scratch and significantly shortening the parameter tuning time. Under unconventional operating conditions, the initial adjustment amount is generated through linear mapping, ensuring that even without completely matching historical experience, the basic adjustment amount can be quickly generated based on the characteristics of the current operating state, providing a starting point for subsequent fine-tuning.
[0123] In one exemplary embodiment, such as Figure 5 As shown, based on historical load sample data and model parameter adjustments, the parameters of the pre-trained load prediction model are adjusted, including:
[0124] Step S502: Based on the model parameter adjustment amount, adjust the model parameters of the lightweight adapter layer to obtain the parameter-tuned lightweight adapter layer.
[0125] Step S504: Using historical load sample data as training samples, perform gradient updates on the hyperparameter-tuned lightweight adaptation layer.
[0126] For example, the model parameter adjustments, namely the scaling factor and bias term, are written into the injection slots corresponding to the lightweight adaptation layer, and the original weights and biases in the lightweight adaptation layer are updated. After the update is completed, the parameter-tuned load prediction model is obtained. Next, the collected historical load samples are input as input features, and the corresponding actual load values are input as labels, into the parameter-tuned load prediction model. The model performs forward propagation to obtain predicted values. Subsequently, the loss function, such as the mean absolute error, between the predicted and actual values is calculated. Based on the loss function, the parameters of the lightweight adaptation layer are updated one or more times according to the calculated learning step size.
[0127] In one exemplary embodiment, such as Figure 6 As shown, the target prediction model includes time-domain and frequency-domain channels; by analyzing the real-time operating data of thermal power units through the target prediction model, the load prediction results of the thermal power units are obtained, including:
[0128] Step S602: Analyze the real-time operating data of the thermal power unit through the time domain channel to obtain the first prediction result reflecting the long-term load change trend of the thermal power unit.
[0129] Step S604: Analyze real-time operating data through frequency domain channels to obtain a second prediction result reflecting the short-term load disturbance of thermal power units.
[0130] Step S606: Combine the first prediction result and the second prediction result to obtain the load prediction result of the thermal power unit.
[0131] The time-domain channel is the component in the load forecasting model responsible for handling time-series dependencies. It captures the dependencies between distant points in the load sequence. The long-term load change trend refers to the changes in unit load over a future period, such as the direction, periodicity, and smoothness of load changes. The frequency-domain channel is the component in the load forecasting model responsible for analyzing load fluctuation characteristics. Short-term load disturbances refer to rapid fluctuations, spikes, and glitches in load on time scales ranging from minutes to hours. These changes are usually caused by rapid events such as instantaneous grid commands, equipment start-ups and shutdowns, and fuel changes. The first forecast result is the long-cycle load change trend curve output by the time-domain channel, and the second forecast result is the short-cycle load disturbance amplitude output by the frequency-domain channel. The load forecast result is obtained by fusing the first and second forecast results point-by-point.
[0132] For example, when real-time operating data is input into the target prediction model, it enters both the time domain and frequency domain channels. The time domain channel performs calculations through multiple self-attention layers and feedforward neural network layers to capture global, long-term dependencies and output the corresponding long-term load change trend curve. The frequency domain channel first performs a short-time Fourier transform or a series of short-window convolutional filters on the real-time operating data to convert the time series into a time-spectrum graph. Then, feature extraction is performed on the time-spectrum graph to obtain frequency domain features. These features are analyzed to capture rapid, localized change patterns and output the corresponding short-term load disturbance amplitude. A fusion weight is then obtained, which represents the similarity during load condition matching. This weight reflects whether the model is more inclined to believe in trend prediction or disturbance prediction at the current prediction time. For example, under stable operating conditions, trend prediction has a higher weight, while under rapidly changing operating conditions, disturbance prediction has a higher weight. After aligning the long-term load change trend curve and the short-term load disturbance amplitude on the time axis, they are superimposed point-by-point within the fusion header according to the fusion weight to obtain the complete load prediction curve. Perform a second-order differential threshold scan on the load forecast curve, mark the time periods that exceed the scheduling tolerance range, and issue an alarm.
[0133] In some embodiments, after obtaining the complete load forecast curve, it can be written into the scheduling interface according to the unit number, and then the curve is integrated according to a preset time sliding window (e.g., 15 minutes) to calculate the proportion of the average load of each time period to the total daily load, forming a load proportion matrix. The start-stop constraint rule set containing the physical operating limitations of the units is called, and the load proportion matrix is checked row by row. Infeasible items are first filtered out, such as when the predicted load for a certain time period is lower than the unit's minimum output, or the load ramp-up rate exceeds the maximum ramp-up rate. Then, conflicts are eliminated. If the start-stop times of multiple units overlap, they are sorted according to priority, such as prioritizing units with higher efficiency or units that have just undergone maintenance, and the start-stop times are adjusted to eliminate overlap. Finally, a feasible power generation scheme containing information such as the unit's target load and start-stop time is generated.
[0134] In some embodiments, peak load periods are marked from the complete load forecast curve, and composite data for these periods are written into the fuel management interface to define the heat load demand corresponding to the power generation required during peak periods. At the receiving end of the fuel management interface, the inventory database is accessed to retrieve information on all coal batches available during peak periods. Information such as calorific value, sulfur content, and volatile matter for each batch of coal is extracted, and a fuel attribute matrix is generated. With the goal of meeting the heat load during peak periods, and constrained by lower inventory limits (e.g., a batch of coal cannot be used beyond its inventory) and upper emission limits (e.g., sulfur content after blending cannot exceed safety standards), the minimum cost blending ratio is calculated using algorithms such as linear programming. This approach satisfies heat demand while ensuring the lowest cost and emission compliance.
[0135] In some embodiments, the complete load forecast curve is transmitted to the AGC system, allowing the AGC system to adjust generator output based on the load forecast curve to maintain grid frequency stability. Specifically, the AGC system analyzes the frequency margin (the allowable range of grid frequency fluctuations) and frequency regulation depth requirement (the maximum increase or decrease in load required by the generators) for the next shift based on the load forecast curve, thereby automatically generating AGC instructions for the next shift. Additionally, the peak load in the current dispatch plan can be compared with the predicted peak load. If the difference exceeds a tolerance threshold, it indicates that the dispatch plan may have underestimated the peak load, leading to insufficient AGC frequency regulation and excessive grid frequency fluctuations. In this case, a frequency regulation warning can be issued immediately.
[0136] In one exemplary embodiment, such as Figure 7 As shown, the load forecasting method for thermal power units also includes:
[0137] Step S702: Obtain multiple operating status feature vectors of the thermal power unit within the historical operating cycle.
[0138] Step S704: Detect the weighted distance between the feature vectors of each operating state. Based on the weighted distance, group the multiple operating state feature vectors to obtain candidate operating conditions.
[0139] Step S706: Detect the business risk level of candidate operating conditions, and construct the candidate operating conditions as training tasks based on the business risk level.
[0140] Step S708: Train the initial prediction model according to the training task to obtain the load prediction model.
[0141] Here, "historical operating cycle" refers to the time range covered by model training, such as the past year or longer. The operating state feature vectors are the various operating state feature vectors generated by the thermal power unit within the historical operating cycle, with the same meaning as the operating state feature vectors in step S202, and will not be repeated here. Weighted distance is an indicator that quantifies the difference between two operating state feature vectors. The candidate operating condition is a set of similar operating conditions obtained through formaldehyde distance aggregation. Each candidate operating condition can contain representative feature vectors of that type of operating condition, as well as the operating condition evolution path. The business risk level is a quantitative judgment of the risk degree of the candidate operating condition. The training task is the encapsulation of the candidate operating conditions for model training. The initial prediction model refers to the load prediction model that has not yet been trained.
[0142] For example, historical operating data of thermal power units within their historical operating cycles is obtained, including boiler type, coal volatile matter range, load ramp-up efficiency, maintenance status, and weather category. This data is then encoded according to coding mapping rules to generate corresponding feature vectors. For these operating status feature vectors, the weighted distance between any two vectors is calculated using the weighted distance formula:
[0143] D F i , F j = ∑ p=1 5 γ p . δ p ( F i [p], F j [p])
[0144] in, and Let p represent the attribute information of two running state feature vectors to be compared. Represents the difference function. Indicates business weight.
[0145] like This indicates that the difference between the two running state feature vectors is small, and the two running state feature vectors can be grouped together. This represents the similarity threshold; if If the difference between two operating state feature vectors is significant, they are assigned to different groups. This process continues until all operating state feature vectors are assigned to their corresponding groups, resulting in candidate groups for the thermal power unit. Each candidate group represents a candidate operating condition. The business risk level of each candidate operating condition is then assessed. Only operating conditions with higher business risk levels are retained, while those with lower levels are merged. In this embodiment, the business risk level detection logic determines whether an operating condition is high-risk based on the business weights of maintenance, coal quality, and start-up / shutdown, as shown in the following expression:
[0146]
[0147] in, This indicates the business risk level of candidate operating condition k. This is for maintenance indication; use 1 if maintenance is required, otherwise use 0. The volatile matter content of coal is encoded using a central feature vector, where the central feature vector refers to the earliest timestamp of the operating state feature vector. Indicates candidate operating conditions Operating condition change trajectory Frequency of occurrence in recent times, This is a rare medium detection function; it takes the value 1 if the medium falls within the rare interval, and 0 otherwise. This is a high-frequency switching detection function; it takes the value 1 when the value is above the threshold, and 0 otherwise. , , All are weights, satisfying .
[0148] If a candidate operating condition includes a maintenance indicator, that is If so, it can be directly classified as a high-risk level. In other words, regardless of other characteristics, maintenance scenarios are the key areas for load forecasting to adapt to. If a candidate operating condition is a rare medium and not a maintenance scenario, then... and If so, it is classified as a secondary risk level. If a candidate operating condition is a high-frequency start-stop condition and neither of the first two conditions is met, then it is classified as a secondary risk level. and , Weights are assigned based on start / stop frequency; the more frequent the start / stop, the higher the weight. Ultimately, only [the following is retained]... The candidate operating conditions are selected, and the remaining operating conditions are merged. This ensures that high-risk operating conditions are included in the model training and prevents the task pool from expanding uncontrollably.
[0149] For the retained high-risk operating conditions, a corresponding training task is constructed, which can be represented as: ,in: The retained candidate operating conditions are the operating state feature vectors. This is the trajectory of the operating condition changes for the candidate operating condition. The unique heat encoding vector for the unit number. This represents the risk weight for the candidate operating condition; if there is no high-frequency switching, it is recorded as... All training tasks are written into the task pool. During model training, the task pool is checked for completeness one by one. Only training tasks containing the operating state feature vector, operating condition change trajectory vector, and unit number one-hot encoded vector are considered valid. Training tasks with missing fields are discarded. After the check is completed, the task pool is set to read-only mode. The parameters of the initial prediction model are initialized, and the operating state feature vector, operating condition change trajectory vector, and unit number one-hot encoded vector of each training task are extracted and concatenated in a fixed order to obtain the task embedding vector. ,in, The vector concatenation operator is used as input for model training. During training, the peak penalty unit price is... Additional co-firing costs With manual scheduling costs Mapped to cost coefficient vector c=[ c 1 , c 2 , c 3 ] Among them, the peak load penalty unit price refers to the penalty amount per unit of load deviation imposed by the power grid dispatch center when thermal power units fail to output according to peak load requirements, such as insufficient output due to missed peak forecasts or excessive output due to exceeding predicted peak loads. Additional co-firing costs refer to the extra costs incurred by power plants to meet model-predicted load demands by deviating from the optimal coal co-firing scheme. Manual dispatching costs refer to the labor costs corresponding to the additional working time invested.
[0150] The training process includes an inner loop and an outer loop, where the outer loop is for training tasks. Predicted output Compared with the baseline curve The absolute difference is taken at each time point, and the model prediction error is divided into peak sections, climbing sections, and maintenance gap sections. Then, it is weighted piecewise with the cost coefficient vector to obtain the cost-weighted loss:
[0151]
[0152] in, For element-wise multiplication, For the model in the training task The predicted curve below, This corresponds to the actual load curve. This is the loss tensor that has been amplified according to business costs. For each training task... Summing is performed to obtain the total loss of the current external loop. At this point, only the backbone network and the fusion head are updated once using a fixed small step gradient descent, while the lightweight adaptation layer remains frozen.
[0153] The inner loop is a lightweight adaptor layer that freezes the outer loop during the training task. Once the training task is thawed and the system is in a trainable state, a small number of samples corresponding to the training task are used to perform one or several gradient updates on the lightweight adaptation layer, enabling it to quickly adapt to the training task. After adaptation, the load ramp-up duration and peak value of the maximum disturbance capability of the model under this task are recorded, and the load ramp-up duration and peak value of the maximum disturbance capability are concatenated into a cross-channel memory vector. After all training tasks are completed, the backbone network weights, the initial weight templates of the lightweight adaptation layer, and the cross-channel memory vector are packaged into a deployable load prediction model package, which can be directly loaded for subsequent real-time prediction.
[0154] In one specific embodiment, Figure 8 A flowchart illustrating the load forecasting method for thermal power units is shown, including:
[0155] S1. Obtain historical operating data of thermal power units within the historical operating cycle.
[0156] S2. Extract the operating state feature vectors and detect the weighted distance between each operating state feature vector. Based on the weighted distance, group multiple operating state feature vectors to obtain candidate operating conditions.
[0157] S3. Detect the business risk level of candidate operating conditions, and construct the candidate operating conditions as training tasks based on the business risk level.
[0158] S4. Based on the training task, train the initial prediction model to obtain the load prediction model. The training process includes an outer loop and an inner loop.
[0159] S5. Encapsulate the backbone network weights, the initial weight templates of the lightweight adaptation layer, and the cross-channel memory vectors to obtain a deployable load prediction model package.
[0160] S6. In response to the real-time load forecasting command for thermal power units, load the load forecasting model package and retrieve the current operating status characteristics of the thermal power units.
[0161] S7. Detect the similarity between the operating status characteristics and the operating condition characteristics of the candidate operating conditions.
[0162] S8. If the similarity is not less than the preset similarity threshold, query the model parameter adjustment amount that matches the candidate operating condition.
[0163] S9. If the similarity is less than the preset similarity threshold, the corresponding model parameter adjustment amount is generated based on the running status characteristics.
[0164] S10. Obtain a small amount of load sample data generated by the thermal power unit in the current shift.
[0165] S11. Based on a small amount of load sample data and model parameter adjustment, the parameters of the lightweight adaptation layer of the load prediction model are adjusted to obtain the adjusted target prediction model.
[0166] S12. Analyze the real-time operating data of thermal power units through the target prediction model to generate a first prediction result reflecting the long-term load change trend of thermal power units and a second prediction result reflecting the short-term load disturbance of thermal power units.
[0167] S13. By integrating the first and second prediction results, the load prediction results of the thermal power units are obtained.
[0168] S14. Based on the load forecast results, generate feasible power generation schemes, calculate the minimum cost co-firing ratio, and generate AGC instructions for the next shift.
[0169] In another specific embodiment, Figure 9 The flowchart of the inner and outer loops during training is shown, including:
[0170] S41. Task Embedding Initialization: Train accurate task embedding vectors for the model. The task embedding vectors are obtained by concatenating the running state feature vector, operating condition change trajectory vector, and unit number one-hot encoded vector of each training task in a fixed order.
[0171] S42. Business Cost Vector Generation: During the training process, the peak penalty unit price, additional co-firing cost and manual scheduling cost are mapped into a cost coefficient vector.
[0172] S43, Outer Loop: Summarize the cost-weighted loss of all training tasks to obtain the overall loss, update only the parameters of the backbone network and the fusion head, and freeze the lightweight adaptation layer.
[0173] S44, Inner Loop: Unfreeze the lightweight adaptation layer frozen in the outer loop, and use a small number of samples from this training task to perform one or several gradient updates on the lightweight adaptation layer.
[0174] S45. Cross-channel memory vector generation: Record the load ramping duration and the peak value of the maximum disturbance capability of the model under this task, and concatenate the load ramping duration and the peak value of the maximum disturbance capability into a cross-channel memory vector.
[0175] S46. Encapsulate the model package: After all training tasks are completed, encapsulate the backbone network weights, the initial weight templates of the lightweight adaptation layer, and the cross-channel memory vectors into a deployable load prediction model package.
[0176] In another specific embodiment, Figure 10 The flowchart of the model update is shown, including:
[0177] A1. Freeze Strategy Generation and Learning Step Size Setting: Combine the maximum weight adjustment percentage and operating condition label to generate a freeze strategy, and at the same time calculate the learning step size for model updates.
[0178] A2. Model fine-tuning based on a small number of samples: Based on a small amount of load sample data under the current shift and the adjustment amount of model parameters, perform a gradient update on the lightweight adaptation layer of the load prediction model.
[0179] A3. Business Cost Assessment and Rollback Control: After the update, assess the cost of the current model in three key business metrics: missed peak load predictions, load ramp-up errors, and forecast window duration. If any metric exceeds the corresponding threshold, the model will be automatically restored to its state before the update. If all metrics do not exceed the corresponding threshold, the current model will be used as the final target load forecasting model.
[0180] A4. Real-time data input: Input real-time running data into the time domain and frequency domain channels of the target prediction model.
[0181] A5. Dual-channel output: The time-domain channel performs calculations through multiple self-attention layers and feedforward neural network layers to capture global, long-term dependencies and output the corresponding long-term load change trend curves. The frequency-domain channel first performs short-time Fourier transform or a series of short-window convolutional filters on the real-time operating data to convert the time series into a time-spectrum graph. Then, it extracts features from the time-spectrum graph to obtain frequency-domain features. The frequency-domain features are analyzed to capture the rapid, local change models and output the corresponding short-term load disturbance amplitudes.
[0182] A6. Dual-channel fusion: After aligning the long-cycle load change trend curve and the short-cycle load disturbance amplitude on the time axis, the curves are superimposed point by point in the fusion header according to the fusion weight to obtain a complete load prediction curve. The load prediction curve is then subjected to a second-order differential threshold scan to mark the time periods that exceed the scheduling tolerance range and issue an alarm.
[0183] In another specific embodiment, Figure 11The diagram illustrates the architecture of the load forecasting model, including an input layer, a dual-channel branch, cross-channel fusion, and an output layer. The input layer comprises historical operating data of thermal power units within their historical operating cycles, as well as auxiliary features. These auxiliary features may include medium, meteorological data, and time stamps, serving as supplementary factors to the load's impact. The dual-channel branch includes a time-domain branch and a frequency-domain branch, extracting features at different time scales in parallel. The left side represents the time-domain branch. First, data splitting occurs, separating the time-domain load sequence and auxiliary feature vectors from the input data to focus on long-cycle analysis. Next, time-domain embedding is performed. This involves using a linear layer to map the original time-domain load sequence and auxiliary feature vectors to a unified feature dimension, adding temporal location information to the sequence so the model can perceive the load's order. The time-domain Transformer is the encoder, used to extract global time-domain dependencies. It includes an 8-head self-attention layer and a feedforward network. The attention layer captures long-distance correlations in the load sequence, while the feedforward network performs non-linear transformations on the self-attention results to enhance feature representation. This process repeats with four layers of self-attention and feedforward networks to progressively refine time-domain features. Finally, there's temporal pooling, which involves pooling the sequence features output by the temporal Transformer to obtain a global temporal feature vector. The right side shows the frequency domain branch. First, there's the STFT (short-time Fourier transform), used to convert the temporal load sequence into a time-frequency domain representation. Next is the frequency domain convolution stack, which extracts local frequency domain features through convolutional layers. Finally, there's frequency domain embedding, where linear layers map the pooled frequency domain features to a feature dimension consistent with the temporal features, preparing for subsequent fusion. In cross-channel fusion, the first step is cross-channel attention, which allows temporal and frequency domain features to exchange information. For example, long-term temporal trends guide which frequency domain perturbations are critical, while frequency domain perturbations supplement details in the temporal trend. The fusion MLP (Multilayer Perceptron) further non-linearly fuses the output of the cross-channel attention. The FiLM (Feature-wise LinearModulation) adaptation layer scales and biases the fused features dimensionally, allowing the model to flexibly adapt to various scenarios. The output layer maps the output features of the FiLM adaptation layer to predicted values of future loads.
[0184] In this embodiment, a dual-channel time-frequency Transformer prediction model is employed. Through dual-channel feature extraction, cross-channel fusion, service cost-weighted loss, and a rapid adaptation layer, minute-level migration is achieved without the need for long-term offline retraining. In contrast, traditional load forecasting schemes use a single time-domain feature dimension and a uniform mean square error loss, outputting a single load curve, which requires long-term offline retraining. Therefore, this embodiment addresses the pain points of traditional schemes, such as slow retraining after changes in operating conditions and the existence of prediction gaps, leading to errors in power grid dispatching decisions.
[0185] This embodiment also provides an experimental process for verification. Using annual SCADA data, dispatch AGC commands, online coal quality test results, and on-site meteorological observation data from four thermal power plants as raw inputs, the data is merged according to millisecond timestamps to form a unified data bus. The system extracts segments from each unit using a 30-minute rolling window, generating approximately 8.6 million data segments containing start and end times and five attribute information. Subsequently, sequential clustering is performed within a three-day rolling window, extracting 23,000 high-risk operating condition prototypes. Among these, maintenance switching types account for 41%, rare high-volatile content types account for 17%, and high-frequency start-stop types account for 42%, which are then encapsulated into a training task pool.
[0186] During the model training phase, the outer loop iterates 15 times across all training tasks, calibrating the loss using three costs: peak penalty, co-firing increase, and manual scheduling hourly wage. The inner loop performs only one small-batch fine-tuning on each training task and generates cross-channel memory vectors. After training, the backbone is frozen, and the backbone network weights, the initial weight templates of the lightweight adaptation layer, and the cross-channel memory vectors are encapsulated into a deployable load prediction model package. After a long-term shutdown maintenance of a boiler at a thermal power plant, the system completes operating condition matching and model parameter adjustment injection within 10 minutes of the unit's grid connection, performing minute-level fine-tuning using only 512 of the latest load samples, and then resumes online inference. This solution can stably output a dual-scale curve of long-term load change trends plus short-term load disturbances within a shift, without requiring full retraining. The scheduling chain downtime is reduced from the traditional 4-6 hours to less than 25 minutes.
[0187] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.
[0188] Based on the same inventive concept, this application also provides a thermal power unit load forecasting device for implementing the above-mentioned thermal power unit load forecasting method. The solution provided by this device is similar to the solution described in the above-described method. Therefore, the specific limitations of one or more thermal power unit load forecasting device embodiments provided below can be found in the limitations of the thermal power unit load forecasting method above, and will not be repeated here.
[0189] In one exemplary embodiment, such as Figure 12 As shown, a load forecasting device for thermal power units is provided, comprising:
[0190] The feature acquisition module 1202 is used to acquire the current operating status features of the thermal power unit in response to the load forecasting command for the thermal power unit;
[0191] The similarity detection module 1204 is used to detect the similarity between the operating status features and the operating condition features of the candidate operating conditions.
[0192] The parameter tuning matching module 1206 is used to perform model parameter tuning matching based on similarity to obtain the model parameter adjustment amount that matches the similarity.
[0193] The load data acquisition module 1208 is used to acquire historical load sample data of thermal power units under the current shift.
[0194] The model update module 1210 is used to adjust the parameters of the pre-trained load prediction model based on historical load sample data and model parameter adjustment amount, so as to obtain a target prediction model that is adapted to the current operating state of the thermal power unit.
[0195] The load forecasting module 1212 is used to analyze the real-time operating data of thermal power units through the target forecasting model to obtain the load forecasting results of the thermal power units.
[0196] In one embodiment, the similarity detection module 1204 is further configured to:
[0197] The difference between the characteristics of the operating status and the characteristics of the candidate operating conditions is detected.
[0198] The difference values are weighted and summed to obtain the weighted distance between the operating state characteristics and the operating condition characteristics of the candidate operating conditions;
[0199] Based on weighted distance, the similarity between the operating state characteristics and the operating condition characteristics of the candidate operating conditions is determined.
[0200] In one embodiment, the parameter tuning matching module 1206 is further configured to:
[0201] If the similarity is not less than the preset similarity threshold, then query the model parameter adjustment amount that matches the candidate operating condition;
[0202] If the similarity is less than the preset similarity threshold, the corresponding model parameter adjustment amount will be generated based on the running status characteristics.
[0203] In one embodiment, the pre-trained load prediction model includes a lightweight adaptation layer; the model update module 1210 is further configured to:
[0204] Based on the model parameter adjustment amount, the model parameters of the lightweight adaptation layer are adjusted to obtain the parameter-tuned lightweight adaptation layer.
[0205] Historical load sample data is used as training samples to update the gradient of the hyperparameter-tuned lightweight adaptation layer.
[0206] In one embodiment, the target prediction model includes a time domain channel and a frequency domain channel; the load prediction module 1212 is further configured to:
[0207] By analyzing the real-time operating data of thermal power units through the time domain channel, the first prediction result reflecting the long-term load change trend of thermal power units is obtained.
[0208] By analyzing real-time operating data through frequency domain channels, a second prediction result reflecting the short-term load disturbance of thermal power units is obtained;
[0209] By combining the first and second forecast results, the load forecast results for thermal power units are obtained.
[0210] In one embodiment, the device is further configured to:
[0211] Obtain feature vectors of multiple operating states of thermal power units within historical operating cycles;
[0212] The weighted distance between the feature vectors of each operating state is detected, and the multiple operating state feature vectors are grouped based on the weighted distance to obtain candidate operating conditions;
[0213] Detect the business risk level of candidate operating conditions, and construct the candidate operating conditions as training tasks based on the business risk level;
[0214] Based on the training task, the initial prediction model is trained to obtain the load prediction model.
[0215] Each module in the aforementioned thermal power unit load prediction device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0216] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 13 As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs in the non-volatile storage media to run. The database stores load forecasting data for thermal power units. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a thermal power unit load forecasting method.
[0217] Those skilled in the art will understand that Figure 13 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0218] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0219] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0220] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0221] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0222] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0223] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0224] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for predicting the load of thermal power units, characterized in that, The method includes: In response to a load forecasting command for a thermal power unit, the current operating status characteristics of the thermal power unit are obtained; Detect the similarity between the operating state features and the operating condition features of the candidate operating conditions; Based on the similarity, model parameter tuning matching is performed to obtain the model parameter adjustment amount that matches the similarity. Obtain historical load sample data of the thermal power unit under the current shift; Based on the historical load sample data and the model parameter adjustment amount, the parameters of the pre-trained load prediction model are adjusted to obtain a target prediction model that is adapted to the current operating state of the thermal power unit. By analyzing the real-time operating data of the thermal power unit using the target prediction model, the load prediction results of the thermal power unit are obtained.
2. The method according to claim 1, characterized in that, The detection of the similarity between the operating state features and the operating condition features of the candidate operating conditions includes: The difference between the operating state characteristics and the operating condition characteristics of the candidate operating conditions is detected; The difference values are weighted and summed to obtain the weighted distance between the operating state characteristics and the operating condition characteristics of the candidate operating conditions; Based on the weighted distance, the similarity between the operating state features and the operating condition features of the candidate operating conditions is determined.
3. The method according to claim 1, characterized in that, The step of performing model parameter tuning matching based on the similarity to obtain the model parameter adjustment amount that matches the similarity includes: If the similarity is not less than a preset similarity threshold, then query the model parameter adjustment amount that matches the candidate operating condition; If the similarity is less than the preset similarity threshold, then the corresponding model parameter adjustment amount is generated based on the running state features.
4. The method according to claim 1, characterized in that, The pre-trained load prediction model includes a lightweight adaptation layer; the parameter adjustment of the pre-trained load prediction model based on the historical load sample data and the model parameter adjustment amount includes: Based on the model parameter adjustment amount, the model parameters of the lightweight adapter layer are adjusted to obtain the parameter-tuned lightweight adapter layer. Using the historical load sample data as training samples, the gradient of the tuned lightweight adaptation layer is updated.
5. The method according to claim 1, characterized in that, The target prediction model includes a time domain channel and a frequency domain channel; The step of analyzing the real-time operating data of the thermal power unit through the target prediction model to obtain the load prediction result of the thermal power unit includes: By analyzing the real-time operating data of the thermal power unit through the time domain channel, a first prediction result reflecting the long-term load change trend of the thermal power unit is obtained. By analyzing the real-time operating data through the frequency domain channel, a second prediction result reflecting the short-term load disturbance of the thermal power unit is obtained; By combining the first prediction result and the second prediction result, the load prediction result of the thermal power unit is obtained.
6. The method according to claim 1, characterized in that, The method further includes: Obtain feature vectors of multiple operating states of thermal power units within historical operating cycles; The weighted distance between each of the aforementioned operating state feature vectors is detected, and the multiple operating state feature vectors are grouped based on the weighted distance to obtain candidate operating conditions; The business risk level of the candidate operating conditions is detected, and the candidate operating conditions are constructed as training tasks based on the business risk level; Based on the training task, the initial prediction model is trained to obtain the load prediction model.
7. A load prediction device for thermal power units, characterized in that, The device includes: The feature acquisition module is used to acquire the current operating status features of the thermal power unit in response to the load forecasting command for the thermal power unit; The similarity detection module is used to detect the similarity between the operating state features and the operating condition features of the candidate operating conditions. The parameter tuning matching module is used to perform model parameter tuning matching based on the similarity to obtain the model parameter adjustment amount that matches the similarity. The load data acquisition module is used to acquire historical load sample data of the thermal power unit under the current shift; The model update module is used to adjust the parameters of the pre-trained load prediction model based on the historical load sample data and the model parameter adjustment amount, so as to obtain a target prediction model that is adapted to the current operating state of the thermal power unit. The load forecasting module is used to analyze the real-time operating data of the thermal power unit through the target forecasting model to obtain the load forecasting result of the thermal power unit.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.