Thermal power plant multi-target economic dispatch optimization method, equipment, medium and product
By constructing a dynamic weight allocation model and mixed integer planning, the problem of multi-objective optimization imbalance in economic scheduling of thermal power plants is solved, and dynamic collaborative optimization of economy, environmental protection and flexibility is achieved, and the real-time and reliability of scheduling strategies are improved.
Patent Information
- Application Number
- CN202510756359.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-09
AI Technical Summary
The existing economic dispatching technology of thermal power plants is difficult to effectively balance the multi-objective optimization conflict between economy, environmental protection and flexibility, resulting in the dispatching plan being unsuitable to complex working conditions.
A dynamic weight allocation model based on real-time scene features is constructed, combined with digital twin simulation and phased reinforcement learning, dynamically adjust the economic, environmental protection and flexibility target weights, and generate Pareto optimal solution sets through mixed integer planning to achieve multi-objective collaborative optimization.
It significantly improves the real-time performance of the thermal power plant scheduling strategy and scenario generalization capabilities, generates high-density Pareto solution sets, reduces the computational complexity, improves the reliability and efficiency of scheduling decisions, and meets the multi-dimensional scheduling needs of thermal power plants.
Smart Images

Figure CN120258488A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a multi-objective economic dispatch optimization method, device, medium and product for thermal power plants. Background Art
[0002] The economic dispatch of a thermal power plant is a decision-making process that optimizes the operation strategy of units to achieve a balance between economy (such as fuel cost, operation and maintenance cost, etc.) and multiple objectives (such as environmental protection emissions, grid flexibility, equipment health, etc.) on the premise of meeting the grid load demand, equipment safety constraints (such as upper and lower limits of output, ramp rate), and environmental protection standards.
[0003] In related technologies, the economic dispatch technology of thermal power plants usually adopts mixed integer linear programming (MILP) or dynamic programming methods, with the minimum fuel cost as the objective function, and solves the optimal unit combination and output allocation scheme under hard constraint conditions such as upper and lower limits of unit output, ramp rate, start-stop time, etc.
[0004] However, the scheduling schemes generated by related technologies rely relatively heavily on preset weights or priorities, cannot cover the complex trade-off relationships between multiple objectives, and it is difficult to effectively balance the conflicts between multi-objective optimizations in the economic dispatch scheme. Summary of the Invention
[0005] Aiming at the above technical problems and deficiencies, the purpose of the present invention is to provide a multi-objective economic dispatch optimization method, device, medium and product for thermal power plants, which can alleviate the problem of multi-objective optimization imbalance in the economic dispatch of thermal power plants, resulting in multi-objective conflicts.
[0006] To achieve the above purpose, in the first aspect, the present invention provides a multi-objective economic dispatch optimization method for thermal power plants, including: constructing an initial data set based on the grid operation parameters, generator set status parameters, and environmental monitoring data of the thermal power plant; determining the characteristic parameters of the current operation dispatch requirements according to the initial data set, and the characteristic parameters include load demand volatility, unit adjustable margin, and environmental protection limit compliance rate; based on the characteristic parameters, determining the real-time target weight coefficients of the economic index, environmental protection index, and dispatch flexibility index through a preset dynamic weight allocation model; generating an initial candidate dispatch scheme set that meets the safety constraints of the generator set according to the real-time target weight coefficients; performing main objective directed optimization on the initial candidate dispatch scheme set to obtain a Pareto optimal solution set that meets the preset convergence conditions; generating a dispatch recommendation scheme based on the Pareto optimal solution set.
[0007] The present invention effectively solves the inherent conflicts among the goals of economy, environmental protection, and flexibility by constructing a dynamic weight allocation model based on real-time scenario features. First, according to key feature parameters such as load volatility, unit regulation potential, and environmental compliance rate, the real-time weight coefficients of each goal are dynamically adjusted, breaking through the limitation of the poor adaptability of the traditional fixed weight mode to complex working conditions. Subsequently, through the main goal-directed optimization and Pareto front search, a set of scheduling schemes that take into account multi-dimensional balance is generated in the solution space that satisfies safety constraints, avoiding the scheduling imbalance caused by the excessive dominance of a single goal and ensuring that the solution set covers the optimal trade-off surface of different goal priorities. Finally, the global optimal solution is recommended by combining multi-criteria decision-making, realizing the collaborative optimization of the economic cost, emission control, and grid response ability of thermal power scheduling. This method forms a closed-loop from dynamic weight mapping, multi-objective solution set generation to intelligent decision-making recommendation, providing a scenario-based adaptive solution for multi-objective conflicts and improving the comprehensive scheduling ability of thermal power plants.
[0008] Optionally, in some embodiments, the construction and training process of the dynamic weight allocation model includes: constructing a multi-modal hybrid model structure, which includes an input layer for fusing grid, generator set, and environmental data, a constraint encoding layer embedding physical mechanism equations, and a hierarchical policy network based on the multi-head attention mechanism; generating a simulation data set containing various working conditions through a digital twin system, and the simulation data set covers scenarios such as fuel shortage, environmental protection equipment failure, and grid frequency disturbance; annotating the Pareto optimal weight vectors of each scenario in the simulation data set to form a training label set; using the training label set and physical mechanism equations as network constraint rules to pre-train the initial parameters of the hierarchical policy network to obtain a pre-trained model; respectively performing staged reinforcement learning training on the hierarchical policy network of the pre-trained model under single-unit steady state, multi-unit coupling, and random failure scenarios to obtain the dynamic weight allocation model.
[0009] Adopting the technical solution of the above embodiment, by constructing a hybrid model architecture that integrates multi-modal data and physical mechanisms, combined with diverse simulation data generated by simulation technologies (such as digital twins) and reinforcement learning training, a full-life cycle construction method of the dynamic weight allocation model is defined. The multi-modal input layer realizes the deep integration of grid, unit, and environmental data. The constraint encoding layer embeds physical rules such as boiler efficiency equations and emission characteristic curves into the network to ensure the physical feasibility of model decisions; the hierarchical policy network dynamically captures the game relationship between goals based on the multi-head attention mechanism. The staged reinforcement learning (single-unit steady state → multi-unit coupling → random failure) enables the model to gradually adapt to complex scenarios, and the simulation data set covers extreme working conditions such as fuel shortage and equipment failure, ensuring the generalization ability of weight allocation. This method breaks through the scenario adaptability limitation of the traditional static weight model and provides a high-precision and interpretable decision-making core for dynamic multi-objective optimization.
[0010] Optionally, in some embodiments, after performing phased reinforcement learning training on the hierarchical strategy network of the pre-trained model in single-unit steady-state, multi-unit coupling and random failure scenarios to obtain the dynamic weight allocation model, it also includes: introducing an adversarial generative network to construct a disturbance sample set; performing robustness reinforcement training on the dynamic weight allocation model according to the disturbance sample set; and dynamically fine-tuning the output layer parameters of the dynamic weight allocation model based on the deviation between the actual scheduling results and the simulation prediction through an online transfer learning framework.
[0011] By adopting the technical solution of the above-mentioned embodiment, adversarial generative networks and online transfer learning are introduced after the dynamic weight model is trained to systematically improve the robustness and continuous optimization capabilities of the model. The adversarial network constructs hidden disturbance samples such as sensor failures and control signal interference, forcing the model to identify abnormal patterns and enhance anti-interference in adversarial training; robustness reinforcement training suppresses weight mutations through reward function design to ensure decision-making stability under extreme working conditions. The online transfer learning module monitors the actual scheduling deviation in real time and dynamically fine-tunes the output layer parameters to enable the model to quickly adapt to changes in the operating environment (such as coal quality fluctuations and policy adjustments). This method solves the problem of domain offset between offline training models and real-world scenarios, and realizes the long-term and reliable application of models in complex industrial environments.
[0012] Optionally, in some embodiments, the initial set of candidate scheduling schemes is optimized in a main target-oriented manner to obtain a Pareto optimal solution set that meets preset convergence conditions, including: determining the main optimization target parameters and the secondary constraint target threshold based on the real-time target weight coefficient; screening a subset of schemes to be optimized that meet the secondary constraint target threshold from the initial set of candidate scheduling schemes according to the main optimization target parameters; optimizing and solving the subset of schemes to be optimized through a mixed integer programming model to obtain a Pareto solution set; verifying whether the Pareto solution set meets preset convergence conditions, the convergence conditions including the objective function variance threshold and the solution set distribution uniformity index; and determining the Pareto solution set that meets the preset convergence conditions as the Pareto optimal solution set.
[0013] The technical solution of the above-mentioned embodiment is adopted to define the complete process of the main target directional optimization, and the solution set quality and computational efficiency are taken into account through the hierarchical optimization of the primary and secondary targets and the Pareto solution set verification mechanism. Based on the real-time weight, economy is determined as the main goal, and environmental protection / flexibility is determined as the constraint threshold, and a subset of high-potential candidate solutions is screened out; the mixed integer programming model is combined with the ε-constraint method to generate Pareto solutions, and the convergence is quantified by the objective function variance and distribution uniformity. This method significantly reduces the computational complexity of multi-objective optimization while ensuring the diversity of the solution set. At the same time, by dynamically adjusting the constraint threshold and iterative verification (such as tightening emission restrictions and uniformity completion), it avoids falling into the local optimum, providing a high-coverage non-dominated solution basis for scheduling recommendations.
[0014] Optionally, in some embodiments, an optimization solution is obtained for the subset of solutions to be optimized through a mixed-integer programming model, and a Pareto solution set is obtained, including: constructing a mixed-integer programming model including a primary objective function, secondary constraint conditions, and equipment operation rules; based on the mixed-integer programming model, calling a solver to perform multi-threshold relaxation solving to obtain a candidate solution set; extracting the Pareto solution set from the candidate solution set through non-dominated sorting and uniform screening strategies.
[0015] Adopting the technical solution of the above embodiment, the specific technical path for solving the Pareto solution set by mixed-integer programming is refined, and the balance between the solution set quality and efficiency is achieved through multi-threshold relaxation and non-dominated sorting. An MILP model including unit start-stop status, power output allocation, and environmental protection equipment operation is constructed, and the secondary objective is transformed into dynamic constraints; when calling the solver, parallel multi-threshold relaxation is adopted (such as gradually tightening the emission limit from 120% to 100%), and combined with warm start and inertia constraint callback to accelerate the solution. The non-dominated sorting and crowding degree screening strategy extracts a uniformly distributed Pareto front from the candidate solutions to ensure wide coverage of the solution set in the objective space. This method solves the problems of slow convergence and sparse solution set of traditional multi-objective optimization algorithms through model-based solution and post-processing strategies.
[0016] Optionally, in some embodiments, an initial candidate scheduling solution set that satisfies the safety constraints of the generator set is generated according to the real-time objective weight coefficient, including: constructing a multi-objective weighted aggregation function based on the real-time objective weight coefficient; according to the multi-objective weighted aggregation function, calling a fast heuristic algorithm to generate a basic feasible solution; expanding the basic feasible solution through parameter perturbation method to obtain multiple different candidate solutions; performing hard filtering of safety constraints on the multiple different candidate solutions to obtain the initial candidate scheduling solution set.
[0017] Adopting the technical solution of the above embodiment, a method for generating the initial candidate scheduling solution set is defined, and the quality and diversity of the solution are balanced through the weighted aggregation function and parameter perturbation strategy. A multi-objective weighted function is constructed based on the real-time weight to drive a heuristic algorithm (such as an improved genetic algorithm) to quickly generate a basic feasible solution; the parameter perturbation method applies directional perturbations to the unit power output, weight coefficient, and start-stop status to expand different candidate solutions. The hard filtering of safety constraints batch-checks hard constraints such as unit safety, grid stability, and environmental protection compliance through a rule engine to ensure the feasibility of the solution set. This method generates a high-quality initial solution set with economy, environmental protection, and flexibility within seconds, providing a sufficient search space for subsequent optimization.
[0018] Optionally, in some embodiments, the basic feasible solution is extended by the parameter perturbation method to obtain multiple differential candidate solutions, including: determining the perturbation parameter and perturbation direction based on the real-time objective weight coefficient and the basic feasible solution; generating multiple groups of parameter combinations according to the perturbation parameter and perturbation direction; adjusting the weight coefficients and output baselines of multiple groups of parameter combinations to obtain multiple groups of directional parameter combinations; and independently solving with each group of directional parameter combinations as the input to obtain differential candidate solutions.
[0019] Adopting the technical solution of the above embodiment, the implementation details of parameter perturbation are clarified, and the diversity of the solution set is improved through systematic perturbation parameter generation and directional adjustment. Based on the weight sensitivity analysis and unit priority, the perturbation direction is determined (such as economic weight ±5%, critical unit output ±3%), and multiple groups of parameter combinations are generated by Monte Carlo sampling; the physical feasibility adjustment of the parameter combinations is realized through weight normalization, output baseline compensation, and start-stop state synchronization, and a lightweight solver is called to quickly generate candidate solutions. This method efficiently explores the objective trade-off space within the neighborhood of the basic solution through a controllable perturbation strategy, avoiding computational waste caused by blind search, and providing a standardized process for generating differential high-quality solutions.
[0020] In a second aspect, an embodiment of the present invention provides an electronic device, including: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, and the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the electronic device to execute the method described in the first aspect or the second aspect, and any possible implementation manner in the first aspect or the second aspect.
[0021] In a third aspect, the present invention provides a computer-readable storage medium, including instructions, when the above instructions run on the above electronic device, causing the above electronic device to execute the method described in the first aspect or the second aspect, and any possible implementation manner in the first aspect or the second aspect.
[0022] In a fourth aspect, the present invention provides a computer program product including instructions, when the above computer program product runs on the above electronic device, causing the above electronic device to execute the method described in the first aspect or the second aspect, and any possible implementation manner in the first aspect or the second aspect.
[0023] It can be understood that the electronic device provided in the second aspect, the storage medium provided in the third aspect, and the computer program product provided in the fourth aspect are all used to execute the method provided by the present invention. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method, which will not be elaborated here.
[0024] One or more technical solutions provided by the present invention have at least the following technical effects or advantages: 1. Dynamic multi-objective weight adaption breaks through the rigid constraints of traditional optimization models: By constructing a dynamic weight allocation model that integrates physical mechanisms and deep learning, the pain point of fixed weights in traditional multi-objective scheduling being difficult to adapt to complex working conditions is solved. The model automatically adjusts the weights of economy, environmental protection, and flexibility based on real-time features (such as the compliance rate of environmental protection limits), and combines digital twin simulation and staged reinforcement learning to achieve precise matching of weight decisions and operating scenarios. For example, when fuel is in short supply, the economic weight is automatically reduced to ensure the safety of the unit, and when the demand for power grid frequency modulation surges, the flexibility weight is increased to give priority to the response. This adaptability upgrades multi-objective optimization from static trade-off to dynamic game, significantly improving the real-time performance and scenario generalization ability of scheduling strategies.
[0025] 2. An efficient Pareto solution set generation mechanism ensures the co-optimal of multiple objectives: Based on mixed integer programming and multi-threshold relaxation solution techniques, a method for main objective-directed optimization and solution set convergence verification is proposed, effectively balancing the quality of the solution set and computational efficiency. By dynamically tightening the sub-objective thresholds using the ε-constraint method, combined with non-dominated sorting and uniform screening strategies, a high-density Pareto solution set covering the three-dimensional trade-off surface of economy - environmental protection - flexibility is generated. Compared with traditional multi-objective algorithms, the hypervolume index (HV) of the solution set of this method is increased by more than 20%, and the convergence time is shortened by 50%, providing a rich and reliable non-dominated solution library for scheduling decisions and avoiding decision-making imbalances caused by sparse solution sets or local optima.
[0026] 3. Enhancement of full-link robustness enables industrial-level reliable applications: Through adversarial training, online transfer learning, and hard filtering of security constraints, a closed-loop reliability guarantee system from model training to solution generation is constructed. An adversarial generative network is introduced to construct complex perturbation samples such as sensor faults and control disturbances to enhance the anti-interference ability of the model; the online transfer framework is fine-tuned in real time based on actual scheduling deviations to solve the domain shift problem between digital twin simulation and physical systems; parameter perturbation expansion and security constraint verification ensure that the candidate solutions fully meet the requirements of unit safety and grid stability. This system enables the scheduling system to still output compliant solutions in extreme scenarios such as fuel fluctuations and equipment abnormalities, and the mean time between failures (MTBF) is increased by 3 times, meeting the stringent requirements of continuous and stable operation of thermal power plants. Description of the Drawings
[0027] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments in accordance with the present invention, and are used together with the specification to explain the principles of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts. In the drawings: Figure 1It is a schematic flow chart of a multi-objective economic dispatch optimization method for thermal power plants according to an embodiment of the present invention; Figure 2 It is a schematic flow chart of the construction and training of a dynamic weight allocation model according to an embodiment of the present invention; Figure 3 It is a schematic architecture diagram of an electronic device according to an embodiment of the present invention. Detailed implementation manners
[0028] The terms used in the following embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention, the singular forms "a", "an", "above-mentioned", "the", and "this" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present invention refers to any or all possible combinations including one or more of the listed items.
[0029] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present invention, unless otherwise stated, the meaning of "a plurality" is two or more.
[0030] It should also be noted that, unless otherwise clearly specified and limited, in the embodiments of the present invention, terms such as "set" and "connect" should be understood in a broad sense. For example, "connect" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the communication inside two components; it can be a wired communication connection or a wireless communication connection. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations. The embodiments of the present invention will be specifically described below.
[0031] The embodiments of the present invention provide a multi-objective economic dispatch optimization method for thermal power plants, as Figure 1 shown, including the following steps: Step 201, based on the grid operation parameters, generator set status parameters, and environmental monitoring data of the thermal power plant, construct an initial data set.
[0032] Specifically, first, through various sensors, monitoring systems, and data interfaces deployed in the thermal power plant and the grid, three major categories of core data are obtained in real time.
[0033] Grid operation parameters include real-time grid load power, peak-valley period division, electricity price fluctuation curve, real-time output of new energy (wind power, photovoltaic power), and grid dispatching instructions (such as ramp rate requirements, reserve capacity demand), etc. These data are synchronized to the dispatching center in real time through the power dispatching data network (such as the SCADA system).
[0034] Generator set status parameters cover the active / reactive power output of each unit, main steam temperature / pressure, turbine vibration amplitude, boiler efficiency, unit start-stop status, cumulative operation time, temperature field and stress monitoring data of key components (such as superheater, turbine blades), and fuel inventory (coal calorific value, sulfur content, blending ratio), etc. These are mainly collected through the unit DCS (Distributed Control System) and equipment Internet of Things sensors.
[0035] Environmental monitoring data include real-time monitoring values of SOx, NOx, and CO2 concentrations at the chimney outlet, the operation status and operation efficiency of environmental protection equipment (such as denitration devices, dust removal equipment), wastewater / slag discharge indicators, etc. These are sourced from on-line pollutant monitors and environmental protection ledger systems.
[0036] After data collection, preprocessing is required to ensure data quality: First, outliers (such as jump data caused by sensor failures) are removed through data cleaning, and high-frequency data is smoothed using a moving average or Kalman filter algorithm; Second, multi-source heterogeneous data is time-aligned and format-unified. For example, second-level and minute-level data from different systems are unified into a time series with a 5-minute resolution; Finally, dimensional differences are eliminated through data normalization (such as Z-score standardization) to construct a structured initial data set containing timestamps, unit numbers, parameter categories, and values. This data set not only covers the real-time status at the current moment but also includes historical data from the previous 1 hour to the previous 1 day, which is used to capture short-term trends (such as load fluctuation inertia, equipment thermal inertia) and provide a comprehensive data basis for subsequent feature parameter extraction.
[0037] Step 202: Determine the characteristic parameters of the current operation and dispatching requirements based on the initial data set. The characteristic parameters include load demand volatility, unit adjustable margin, and environmental protection limit compliance rate.
[0038] Among them, the load demand volatility refers to the percentage change in the real-time grid load relative to the reference load (such as the predicted value or historical average), which is used to quantify the urgency of the power supply-demand imbalance. It is usually calculated by dividing the load change amount within a unit time by the maximum regulation rate of the unit.
[0039] The unit adjustable margin characterizes the range of flexible power generation adjustment capabilities of the generator set at the current output. The calculation formula is (the difference between the current output and the maximum / minimum technical output) divided by the rated capacity of the unit, which reflects the potential of the unit to participate in frequency modulation and peak shaving.
[0040] The percentage ratio of the pollutant emission concentration (such as SO2, NOx) monitored in real time for meeting environmental protection limits to the environmental protection regulation limits. If the compliance rate is lower than 100%, it indicates that the emissions exceed the standard, and the environmental protection weight increase mechanism needs to be triggered immediately.
[0041] Specifically, the calculation of the load demand volatility needs to combine the real-time grid load and historical data. First, obtain the load forecast value and actual value within the current scheduling period (such as 15 minutes), calculate the ratio of the absolute deviation to the rated load, or use a sliding window to calculate the standard deviation of the load change in the past 30 minutes to reflect the severity of the load fluctuation. If the grid is in the peak load period and the new energy output drops suddenly, the load demand volatility will increase significantly, indicating that the grid has an urgent need for the rapid response of thermal power units, and the flexibility weight needs to be increased in the scheduling.
[0042] The evaluation of the adjustable margin of the unit needs to consider the constraints at both the single-unit and plant-wide levels: for a single unit, the adjustable margin = min (rated output - current output, current output - minimum technical output) / rated output, which reflects the up and down adjustment space of the unit at the current load; at the same time, consider the ramp rate constraint, that is, the maximum adjustable capacity in the next 15 minutes = current output + ramp rate × time interval. The plant-wide adjustable margin is the weighted sum of the adjustable margins of each unit (the weight is the ratio of the unit rated capacity), and if the plant-wide adjustable margin is lower than 10%, it indicates that the overall adjustment ability of the unit is approaching the limit, and the risk of equipment overload caused by excessive peak shaving needs to be avoided first in the scheduling.
[0043] The calculation of the environmental protection limit compliance rate is based on real-time emission data and environmental protection standards: for NOx, the compliance rate = real-time emission concentration / environmental protection limit (such as the ultra-low emission limit of 50 mg / m³). If the compliance rate exceeds 90% (i.e., the concentration is close to the limit), it indicates that the environmental protection constraints are tightening, and the weight of the environmental protection index needs to be increased. The emissions can be reduced by adjusting the combustion parameters or operating the in-depth purification equipment. This parameter can also be dynamically corrected in combination with the operating status of environmental protection equipment (such as the activity of denitration catalysts and the ammonia water reserve). For example, when the denitration system fails, the response ability of the standby environmental protection facilities needs to be included in the calculation of the compliance rate.
[0044] Step 203: Based on the characteristic parameters, determine the real-time target weight coefficients of the economic index, environmental protection index, and scheduling flexibility index through a preset dynamic weight allocation model.
[0045] Among them, the dynamic weight allocation model is a rule-driven algorithm that adaptively adjusts the target weights based on multi-dimensional scenario features. This model takes the real-time operating status of thermal power plants and external environmental parameters as inputs, and dynamically calculates the weight coefficients of economic, environmental, and flexibility indicators through a predefined target conflict resolution rule library and fuzzy logic inference mechanism. Its core lies in constructing a mapping relationship table between target weights and characteristic parameters. For example, when the compliance rate of environmental protection limits is lower than the preset threshold, the rule for increasing the environmental protection weight is triggered; when the demand for power grid frequency modulation surges, the flexibility weight is automatically increased.
[0046] The dynamic weight allocation model adopts a hierarchical decision-making architecture. The first layer screens key constraints through threshold judgment (such as whether the environmental protection limits are on the verge of exceeding the standard), the second layer calculates the weight ratios of non-critical targets based on linear interpolation or look-up table methods, and finally ensures that the sum of all weight coefficients is 1 through normalization processing. This model avoids complex mathematical optimization and relies on expert experience and historical data statistical laws, and can complete weight decision-making within 10 ms.
[0047] The specific implementation process of this step is as follows: First, three characteristic parameters are extracted from the real-time data: the load demand volatility (the ratio of the current load change rate to the maximum ramp rate of the unit), the unit adjustable margin (the ratio of the difference between the current output and the maximum adjustable output to the maximum output), and the compliance rate of environmental protection limits (the ratio of the actual emission concentration to the limit). Subsequently, the characteristic parameters are input into the dynamic weight allocation model: 1) If the compliance rate of environmental protection limits > 95%, the baseline value of the environmental protection weight is set to 0.2, otherwise the weight increases by 0.05 for each 1% decrease; 2) When the load demand volatility > 5%, the baseline value of the flexibility weight linearly increases from 0.1 to 0.3, with an increase amplitude of 2 times the excess of the volatility; 3) When the unit adjustable margin < 20%, the economic weight is reduced by 0.1 to reserve adjustment margin. Finally, the three baseline weights are normalized. For example, the initial baseline values are economic 0.6, environmental protection 0.2, and flexibility 0.2, and after adjustment, they become economic 0.55, environmental protection 0.25, and flexibility 0.2, and are output after the sum verification is 1. This process is implemented through pre-compiled rule scripts, and the weight is updated every 15 seconds to ensure that the scheduling strategy is adapted to the changes in the operating scenario in real time.
[0048] In this embodiment, the construction and training process of the dynamic weight allocation model is divided into four stages: First, an initial rule library is constructed based on historical scheduling data and expert experience. For example, the optimal weight distribution under different load demands and environmental protection limit scenarios in the past year is collected, and regular parameters such as the average value of the economic weight during the peak shaving period being 0.55 and the environmental protection weight increasing to 0.4 during the emission over-standard warning are statistically obtained.
[0049] Secondly, a fuzzy logic system is used to design the weight adjustment rules. The characteristic parameters are divided into three fuzzy sets: "low", "medium", and "high", and a fuzzy inference rule table is defined (for example, when the compliance rate of environmental protection limits is "low" and the load volatility is "high", trigger the adjustment action of "environmental protection weight +0.15, flexibility weight +0.05, economic weight -0.2"), and the membership function is used to quantify the parameter and rule matching degree.
[0050] Then, the rule parameters are optimized using offline historical data. The grid search method is adopted to traverse the candidate values of each threshold in the rule base (such as the environmental protection compliance rate trigger threshold is tested from 90% to 98% in steps of 2%), and the comprehensive economic and environmental protection cost of the actual scheduling result is used as the evaluation index to select the optimal parameter combination.
[0051] Finally, an online adaptive calibration module is deployed to compare the deviation between the weight output by the model and the scheduling effect in real time. When the target achievement degree in 5 consecutive scheduling plans is lower than the expected value, the rule base iteration is automatically triggered (such as relaxing the trigger sensitivity of the environmental protection weight increase), and at the same time, the latest data is reloaded monthly to batch update the rule base to ensure that the model continuously adapts to the changes in the power plant operation environment.
[0052] Step 204, generate an initial candidate scheduling plan set that meets the safety constraints of the generating units according to the real-time target weight coefficient.
[0053] First, a weighted single-objective function is constructed based on the real-time weight coefficient. For example, the total cost = economic weight × fuel cost + environmental protection weight × emission penalty + flexibility weight × frequency modulation cost. Subsequently, the greedy algorithm or relaxed linear programming is used to quickly generate a basic feasible solution: starting from the current unit state, preferentially select the units with low coal consumption rate and excellent emission performance to increase the output, while meeting the power balance constraint. For example, if the output upper limit of a certain unit is 300MW and the current output is 200MW, it is incremented in steps of 10MW until the ramp rate limit is reached (such as 5MW / min). During this process, the hard safety constraints (such as the minimum start-stop time of the unit and the no-entry vibration area) are verified in real time through logical judgment, and the illegal solutions are directly eliminated. To increase the diversity of the solution set, the weight coefficient can be perturbed within a small range (such as ±5%) to generate multiple groups of differentiated objective functions, and the solver is called to calculate respectively. For example, the original weights are economic 0.6, environmental protection 0.3, and flexibility 0.1. After perturbation, three groups of weights are generated (0.63 / 0.28 / 0.09, 0.57 / 0.31 / 0.12, 0.60 / 0.30 / 0.10), and each group of weights corresponds to a candidate solution. Finally, all feasible solutions are collected through parallel computing or iterative search to form an initial candidate set containing 10 - 20 groups of solutions for subsequent optimization. This process can usually be completed within 5 seconds, meeting the real-time scheduling timeliness requirements.
[0054] Step 205: Conduct primary objective-oriented optimization on the initial candidate scheduling plan set to obtain a Pareto optimal solution set that meets the preset convergence conditions.
[0055] The implementation of this step is based on a hierarchical optimization strategy. First, select the objective with the highest weight from the weight coefficients output by the dynamic weight allocation model as the primary optimization objective (for example, the economic weight is 0.6), and the remaining objectives as constraints.
[0056] In specific implementation, construct a mixed-integer linear programming (MILP) model with the primary objective as the core: 1) The primary objective function is the economic index (such as minimizing fuel cost), including parameters such as coal consumption rate and start-stop cost; 2) The secondary objectives are transformed into constraints. For example, the environmental protection index requires that the total emissions do not exceed 110% of the real-time limit, and the flexibility index requires that the total unit ramp rate is not lower than 90% of the grid frequency regulation demand; 3) The hard safety constraints include the upper and lower limits of unit output, the minimum start-stop time, and the vibration zone entry prohibition rule. Use the branch and bound method for accurate solution, and dynamically adjust the constraint boundaries during the calculation process to explore the solution space. For example, when the economically optimal solution causes environmental protection to exceed the standard, gradually tighten the emission constraint (such as from 110% to 105%), and re-solve until all constraints are met. To generate the Pareto front, iteratively adjust the relaxation range of the secondary objective constraints through the ε-constraint method: After the first solution, record the optimal value of environmental protection, relax this value by 5% as the new constraint, and solve the economic objective again. Repeat this process until all non-dominated solutions are covered. The convergence condition is set that the improvement amplitude of the objective function is less than 0.5% in three consecutive iterations or reaches the maximum number of iterations (such as 50 times), and finally output a Pareto optimal solution set containing 10 - 20 groups of non-dominated solutions.
[0057] Step 206: Generate a scheduling recommendation plan based on the Pareto optimal solution set.
[0058] Specifically, the comprehensive optimal plan can be screened from the Pareto solution set through multi-criteria decision analysis as the scheduling recommendation plan. The scheduling recommendation plan is an optimal unit output allocation and start-stop combination strategy that takes into account economy, environmental protection, flexibility, and has no hidden risks.
[0059] In this step, first construct an evaluation index system, including four dimensions: economy (unit power supply cost), environmental protection (emission intensity), flexibility (frequency modulation response time), and equipment health (turbine stress coefficient). Each dimension is assigned a standardized scoring weight (such as 40% for economy, 30% for environmental protection, 20% for flexibility, and 10% for health).
[0060] Quantify the indicators for each plan in the Pareto solution set: 1) Economic score = (historical lowest cost - current cost) / (historical highest cost - historical lowest cost) × 100; 2) Environmental protection score = 1 - (actual emissions / limit value) × 100 (if exceeding the standard, the score is 0); 3) Flexibility score = (theoretical maximum frequency modulation rate - actual rate) / theoretical value × 100; 4) Equipment health score = 1 - (stress accumulation / safety threshold) × 100.
[0061] Subsequently, TOPSIS (Technique for Order of Preference by Similarity to Ideal Solution) is used to calculate the closeness of each scheme to the ideal solution: determine the positive ideal solution (the highest score for each dimension) and the negative ideal solution (the lowest score for each dimension) of each index, calculate the Euclidean distances of each scheme to the positive and negative ideal solutions, and finally sort them from high to low according to the closeness.
[0062] For example, for a certain scheme with an economic score of 90 points, an environmental protection score of 85 points, a flexibility score of 70 points, and a health score of 80 points, its closeness is 0.82, ranking first. To enhance the reliability of decision-making, a digital twin verification link is added: input the top 3 schemes into the power plant digital twin system, simulate the operating conditions in the next 2 hours, and detect whether there are hidden constraint violations (such as delayed NOx concentration exceeding the standard). If the simulation results show that a certain scheme's emissions exceed the standard after 30 minutes, the weight of this scheme will be automatically reduced and the sub-optimal solution will be selected. Finally, the scheme with the highest comprehensive score and passing the simulation verification is output as the dispatching recommendation scheme, and a list of alternative schemes is provided for reference in manual decision-making.
[0063] Through the collaborative design of the dynamic weight allocation model and the multi-stage optimization strategy in this embodiment, the comprehensive efficiency and real-time decision-making ability of the economic dispatching of thermal power plants are significantly improved, and the conflicts between multi-objective optimizations in the economic dispatching scheme are effectively balanced.
[0064] This embodiment constructs multi-dimensional characteristic parameters based on the real-time collected power grid, unit and environmental data, accurately depicts the core contradictions of the current dispatching scenario. For example, the power grid stability pressure is quantified through the load demand volatility, the peak shaving ability is evaluated through the unit adjustable margin, and the emission constraint intensity is reflected through the environmental protection limit compliance rate, providing high-information-density inputs for dynamic weight allocation and solving the problem of poor scenario adaptability caused by traditional methods relying on manual experience or static rules. Secondly, the dynamic weight allocation model adopts a hybrid architecture that combines reinforcement learning and physical mechanisms, and can automatically adjust the weight coefficients of economic, environmental protection and flexibility objectives according to real-time scenario characteristics. For example, when the power grid frequency modulation commands are frequently issued, the flexibility weight is increased to more than 70%, and during the strengthening stage of environmental protection monitoring, the environmental protection weight is dynamically increased to the dominant position, breaking through the limitations of the traditional fixed weight model in scenarios with target conflicts, where the solution set is single and cannot cover multi-dimensional trade-off relationships.
[0065] Furthermore, the computational complexity is significantly reduced through a two-stage optimization strategy (initial candidate solution generation and main objective-oriented optimization). The initial candidate solution set is generated using heuristic rules and parameter perturbation methods, which can provide 20 - 30 feasible solutions that meet safety constraints within 3 seconds, achieving a speedup of more than 5 times compared to traditional multi-objective optimization algorithms. In the main objective-oriented optimization stage, the mixed-integer programming is solved by focusing on the objective with the highest weight, and the sub-optimal objectives are transformed into constraint conditions, which not only ensures the diversity of the Pareto front but also controls the overall optimization time within 10 seconds, meeting the time sensitivity requirements of real-time scheduling.
[0066] In addition, a device life loss prediction model and a digital twin verification mechanism are introduced in the generation process of the Pareto optimal solution set, which can avoid the implicit costs caused by traditional methods ignoring long-term device health. For example, taking the steam turbine stress accumulation as a constraint term can reduce the life loss caused by frequent start-stop by more than 40%.
[0067] Under the premise of ensuring computational real-time performance, this embodiment can reduce the comprehensive scheduling cost by 15% - 22%, reduce the number of environmental protection violations by 90%, improve the frequency modulation response speed of the unit by 35%, and at the same time increase the diversity index of the solution set (such as Hypervolume) by more than 50%, providing multi-dimensional decision-making support with economy, environmental protection, and reliability for dispatchers, and effectively solving the three core problems and contradictions of rigid optimization objectives, low computational efficiency, and uncontrollable long-term device loss in traditional technologies.
[0068] In some embodiments, a method for constructing and training a dynamic weight allocation model is also provided, as Figure 2 shown. The specific process is as follows: S301, construct a multi-modal hybrid model structure.
[0069] Among them, the multi-modal hybrid model structure includes an input layer for fusing power grid, generator set data, and environmental data, a constraint encoding layer embedded with physical mechanism equations, and a hierarchical policy network based on the multi-head attention mechanism. The physical mechanism equations are mathematical expressions established based on physical laws such as thermodynamics and fluid mechanics, which are used to quantitatively describe the internal laws of energy conversion, equipment efficiency, and pollutant generation during the operation of thermal power units. For example, the quadratic function relationship between boiler combustion efficiency and load, and the differential equation of steam turbine heat consumption rate with respect to steam parameters.
[0070] Specifically, the input layer is designed as a multi-channel data fusion module. The power grid data (frequency deviation, frequency modulation demand) is processed by sliding window normalization, the generator set data (coal quality calorific value, output curve) is normalized by Z-score, and the environmental data (emission concentration, meteorological parameters) is aligned by time series.
[0071] The constraint encoding layer integrates physical mechanism equations such as the thermodynamic equations of thermal power units. For example, the boiler efficiency-load relationship formula (η = aP² + bP + c) is transformed into an inequality constraint matrix and embedded into the forward propagation process of the network through automatic differentiation technology.
[0072] The hierarchical policy network adopts a multi-head attention mechanism. Among them, the economic attention head calculates the correlation weight between fuel cost and load demand, the environmental protection attention head focuses on the time series dependence of emission concentration and the status of environmental protection equipment, and the flexibility attention head analyzes the dynamic matching between frequency regulation instructions and the unit ramp rate. The outputs of each attention head are fused into the final weight vector after residual connection and layer normalization.
[0073] S302, a simulation dataset containing various working conditions is generated through simulation technology. The simulation dataset covers scenarios such as fuel shortage, environmental protection equipment failure, and power grid frequency disturbance.
[0074] Specifically, first, a high-fidelity simulation model is constructed based on the actual power plant design parameters (such as boiler capacity, steam turbine model, environmental protection equipment specifications). Among them, the core equipment of the unit adopts a mechanism model (for example, the boiler combustion process is modeled by the partial differential equation of mass-energy conservation, and the denitration reactor simulates the dynamic relationship between ammonia injection and NOx conversion rate based on the chemical reaction kinetics equation), and the power grid interaction module integrates the IEEE standard node model to simulate the frequency-power coupling characteristics.
[0075] After that, diverse working conditions are generated by the parameter perturbation method: 1) The fuel shortage scenario is realized by adjusting the coal quality parameters. For example, the received base low calorific value is randomly reduced from the design value of 20 MJ / kg to 15 - 18 MJ / kg, and at the same time, the ash content is increased (from 15% to 25%), and the fuel delivery delay caused by coal feeder blockage is simulated (random disturbance of 0 - 30 seconds); 2) The environmental protection equipment failure scenarios include abnormal ammonia slip rate in the selective catalytic reduction (SCR) system (suddenly increasing from the normal 3 ppm to 15 ppm), and the dust removal efficiency decreases due to the short circuit of the dust collector plate (from 99.9% to 85%). It is realized by modifying the equipment state variable parameters and injecting a step signal or a ramp failure curve; 3) The power grid frequency disturbance scenario uses a random signal generator to simulate the power grid frequency fluctuation, superimposing a normal distribution noise (standard deviation 0.1 Hz) and a periodic disturbance (amplitude ±0.5 Hz, period 10 - 60 seconds randomly) on the reference frequency of 50 Hz, and at the same time, associating the logic of issuing frequency regulation instructions (triggering the AGC instruction when the frequency deviation exceeds 0.2 Hz).
[0076] All scenarios are covered in combination, such as the compound abnormal operating condition of "coal calorific value decline + SCR ammonia pump failure + grid frequency drop of 0.3Hz" triggered simultaneously. During the data generation process, the simulation model is continuously run with a sampling interval of 1 second to record 300+ dimensional parameters such as grid frequency, unit output, emission concentration, equipment status, etc. Each scenario lasts for no less than 6 hours to capture transient and steady-state characteristics, and finally generates a simulation data set containing more than 100,000 groups of samples. Each sample is annotated with the corresponding operating condition type (such as fuel shortage level, equipment fault code), real-time target weight true value (pre-calculated through offline multi-objective optimization) and key constraint violation flag for subsequent model training and verification.
[0077] S303, labeling the Pareto optimal weight vector of each scene in the simulation data set to form a training label set.
[0078] Among them, the training label set refers to the set of optimal target weight vectors that are pre-calculated by the multi-objective optimization algorithm and match each operating scenario in the simulation data set. It contains the Pareto optimal solution of the economy, environmental protection and flexibility weights under different working conditions, and is used to supervise the model to learn the mapping relationship between scenario features and dynamic weights.
[0079] Specifically, for each simulation scenario, the improved NSGA-III algorithm can be called for multi-objective optimization, and the objective function is the three-dimensional Pareto front solution of economy (fuel cost), environmental protection (total emission), and flexibility (frequency response error). The constraints set include unit output limit, start-stop time, and NOx instantaneous concentration hard upper limit.
[0080] After the optimization is completed, the weight vectors of all non-dominated solutions on the Pareto front are extracted, and their distribution consistency is evaluated by KL divergence to eliminate abnormal solutions that deviate from the main cluster.
[0081] The final labeled data contains scenario feature vectors (load demand, coal quality parameters, equipment status) and the corresponding Pareto optimal weight set, forming a large number of training label sets.
[0082] S304, using the training label set and the physical mechanism equation as network constraint rules, pre-training the initial parameters of the hierarchical strategy network to obtain a pre-trained model.
[0083] Specifically, first, a composite loss function is constructed. For example, \(L_{total}=\alpha\cdot L_{label}+\beta\cdot L_{phy}\), where \(L_{label}\) is the cosine similarity loss between the predicted weight and the Pareto label, and its calculation formula is \(1 - \frac{y_{pred}\cdot y_{true}}{\|y_{pred}\|\cdot\|y_{true}\|}\), which is used to drive the network to approximate the optimal weight distribution; \(L_{phy}\) is the physical mechanism constraint loss, including three types of penalty terms: 1) The constraint based on the boiler efficiency-load curve, which calculates the mean square error between the theoretical coal consumption rate corresponding to the economic weight output by the network and the actual value; 2) The environmental protection weight constraint, which substitutes the predicted environmental protection weight into the SCR denitration efficiency equation (\(\eta = k\cdot C_{nox}\cdot T_{reactor}\)) to verify whether the emission concentration meets the strict requirement of \(\eta\geq85\%\), otherwise a gradient penalty is imposed; 3) The flexibility weight constraint, which checks whether the unit frequency modulation rate meets \(\frac{\Delta P}{\Delta t}\geq(weight\times theoretical maximum)\), and if violated, the square of the slack variable is calculated as a loss term.
[0084] During training, the projected gradient descent method is adopted, and after each parameter update, the weight coefficients are forced to be projected into the feasible region (such as \(\sum weights = 1\), each weight \(\geq0.1\)). In the data preprocessing stage, the input features are normalized by a sliding window (window length 60 seconds), and the physical equation parameters (such as the heat consumption rate of the steam turbine, the NOx generation coefficient) are encoded as 32-dimensional vectors and input into the constraint encoding layer. The pre-training is carried out for 200 epochs, the batch size is 128, the initial learning rate is 0.01, and it decays to 1 / 10 every 50 epochs. At the same time, the ratio of \(\alpha\) to \(\beta\) is dynamically adjusted (the initial \(\alpha:\beta = 1:1\), and it is adjusted according to the validation set loss after each epoch. If \(L_{label}\) decreases slowly, then \(\alpha\) is increased by 20%).
[0085] The pre-trained model obtained needs to meet the following when tested by the validation set: the cosine similarity \(\geq0.92\), the physical constraint violation rate \(<3\%\), and the weight normalization error \(<0.001\).
[0086] S305. Respectively, in the single-unit steady state, multi-unit coupling, and random fault scenarios, the hierarchical strategy network of the pre-trained model is trained by staged reinforcement learning to obtain a dynamic weight allocation model.
[0087] This step can be implemented through the following stages: Single-unit steady-state training phase: Under a fixed grid load (such as a constant 500 MW) and a fault-free state of the equipment, the environmental parameters are initialized (coal calorific value 20 MJ / kg, SCR denitrification efficiency 90%), and the Deep Deterministic Policy Gradient (DDPG) algorithm is used for training. The action space is the continuous value adjustment of the economy, environmental protection, and flexibility weights (Δw ∈ [-0.1, +0.1]), and the state space contains 20-dimensional features such as the current output of the unit, coal consumption rate, and NOx concentration. The reward function is designed as R = 0.7×(benchmark fuel cost - actual cost) / benchmark cost + 0.2×(1 - emission exceedance flag) + 0.1×frequency modulation response rate. The exploration noise is set as an OU process (θ = 0.15, σ = 0.2), and training is carried out for 500,000 steps until the average reward is stable above 0.85.
[0088] Multi-unit coupled training phase: Extended to the coordinated scheduling of 6 units, the state space increases to 120 dimensions (20-dimensional parameters for each unit), and the action space is the global weight vector + the output allocation ratio of each unit. The Multi-Agent Proximal Policy Optimization (MAPPO) framework is adopted. The central Critic network evaluates the global reward, and each unit's Actor network independently outputs local actions. The reward function is upgraded to R = 0.5×total economic benefit + 0.3×environmental compliance rate + 0.2×Σ (reciprocal of the frequency modulation synchronization error of unit i), and a coupling penalty term is added: If the output deviation between any two units exceeds 20% and lasts for 5 minutes, 0.1 reward value is deducted. The training adopts a curriculum learning strategy, gradually increasing the number of units (from 2 to 6), and the exploration rate linearly decays from 0.3 to 0.1, with a total of 800,000 steps of training.
[0089] Random fault training phase: Introduce dynamic environmental disturbances. There is a 30% probability of triggering one of the following events in each round of training: 1) sudden change in coal quality (step change in calorific value within ±15%); 2) failure of the SCR ammonia pump (denitrification efficiency drops to 50% within 10 seconds); 3) power grid frequency disturbance (superimposed ±0.5 Hz sinusoidal fluctuation). The Adversarial Reinforcement Learning framework is adopted, and an additional adversarial network is deployed to generate deceptive environmental monitoring signals (such as false NOx compliance data). The policy network needs to identify anomalies and adjust weights within 0.5 seconds. The reward function strengthens the long-term penalty term: R = R_base - 0.5×Σ (predicted equipment life loss rate in the next 1 hour). The training adopts priority experience replay, preferentially sampling data of fault scenarios, setting the exploration rate reset to 0.2, and training for 1.2 million steps until the average reward in the fault scenario ≥ 0.75. After training in each phase, the model performance is tested on an independent validation set (such as 10,000 unseen disturbance scenarios), requiring the weight decision delay < 50 ms and the Pareto solution coverage rate to reach more than 95%. Finally, through three-phase progressive training, the model is made to have multi-scenario robustness to obtain a dynamic weight allocation model.
[0090] S306. Introduce an adversarial generative network to construct a perturbation sample set.
[0091] Specifically, construct a conditional generative adversarial network (CGAN). The generator takes the real environmental state (such as unit output, emission concentration) and a random noise vector as inputs, and outputs forged samples with perturbation characteristics, including: 1) Sensor fault type perturbation (such as a ±20% drift in coal calorific value data); 2) Control signal interference (such as a false AGC frequency modulation command superimposed with Gaussian noise of ±5% amplitude); 3) Concealed environmental protection data tampering (injecting a slow upward trend when the NOx concentration meets the standard to make it seem to exceed the standard).
[0092] The discriminator adopts a hybrid structure of convolutional-long short-term memory network (CNN-LSTM), and judges whether the input data is a generated sample through time series pattern recognition. During the adversarial training process, the generator's objective function minimizes the weight volatility (Δw < 0.05) of the policy network on the forged samples, while maximizing the misjudgment probability of the discriminator; the discriminator maximizes the classification accuracy of real and fake samples. After each round of training, retain the perturbation samples that cause the weight of the policy network to deviate by more than 10%, forming a perturbation data set containing a large number of highly aggressive samples to construct a perturbation sample set.
[0093] S307. Perform robust reinforcement training on the dynamic weight allocation model according to the perturbation sample set.
[0094] Specifically, in the reinforcement learning environment, the dynamic weight allocation model (Agent) interacts with the simulation environment injecting perturbation samples. Each time of interaction, the environmental state is sampled from the real data set with a probability of 60% and from the adversarial generated perturbation data set with a probability of 40%. For example: 1) The sensor perturbation sample changes the actual unit output of 300 MW to a random value between 280 and 320 MW; 2) The control signal perturbation sample superimposes a sine interference with an amplitude of ±5% and a frequency of 0.1 - 1 Hz on the real AGC command; 3) The environmental protection data perturbation sample makes the NOx concentration fluctuate at the edge of meeting the standard (95% - 105% of the limit).
[0095] The Agent receives the perturbed state observation value, outputs the economic, environmental protection, and flexibility weight vectors, and obtains feedback based on the improved robust reward function: R = 0.6×R_base (basic reward, including economic benefits and environmental protection penalties) + 0.3×R_stability (stability reward, calculating the cosine similarity between the current weight and the average value of the previous 10 steps to suppress mutations) - 0.1×R_attack (attack detection penalty, imposing a 0.1-fold penalty when the discriminator network determines that the current state is an adversarial sample).
[0096] The training adopts the TD3 reinforcement learning algorithm. The Actor network is updated at each step, and the Critic network reduces the overestimation bias through double Q-learning. The update coefficient τ of the target network is 0.005. To improve the anti-interference ability, a state recovery mechanism is designed: when the variance of the Q value predicted by the Critic exceeds the threshold for 5 consecutive steps, the historical state is rolled back (the states of the last 10 steps are replaced with the undisturbed version), forcing the Agent to re-make decisions.
[0097] During the training process, a robustness stress test is conducted every 50,000 steps: the model is evaluated in an independent test set (including 2,000 groups of extreme disturbance scenarios, such as a 30% sudden drop in coal calorific value and a 40% drop in SCR efficiency at the same time). It is required that the volatility of the objective function ((maximum value - minimum value) / mean) < 10%, and the confidence level of the weight decision (calculated by the softmax probability entropy) > 0.7. The final model needs to pass 300 adversarial attack tests, and under perturbations, maintain the weight adjustment amplitude Δw < 0.15 and the Pareto solution coverage rate ≥ 85%.
[0098] S308, through an online transfer learning framework, dynamically fine-tunes the output layer parameters of the dynamic weight allocation model based on the deviation between the actual scheduling results and the simulation predictions.
[0099] Specifically, during the deployment phase, actual scheduling data can be collected in real time (such as the unit output, fuel consumption, and emission concentration per minute), and compared with the model prediction values to calculate multi-dimensional deviation indicators: economic deviation ΔC = (actual cost - predicted cost) / predicted cost, environmental protection deviation ΔE = MAX(0, actual emissions - predicted emissions) / emission limit, flexibility deviation ΔF = 1 - (actual frequency modulation response rate / predicted value).
[0100] When ΔC > 10% or ΔE > 5% in 5 consecutive groups of data, the parameter update process is triggered: 1) Freeze the parameters of the feature extraction layer and the attention mechanism layer of the policy network, and only open the fully connected weight matrix of the output layer (dimension 128×3) for fine-tuning; 2) Construct an online training set, extract the 500 samples with the largest deviations within the last 24 hours from the circular buffer. Each sample contains the environmental state (a 120-dimensional vector after normalization), the original weight output of the model, and the actual scheduling results; 3) Design a transfer loss function, such as: L = 0.5×MSE (predicted weight, corrected weight) + 0.3×|ΔC + ΔE| + 0.2×L2 regularization term; Among them, MSE (predicted weight, corrected weight) refers to the mean square error between the target weight coefficient output by the calculation model and the target weight after adjustment by the expert rules, which is used to constrain the weight adjustment direction in the model fine-tuning stage to make the prediction result approach the ideal value corrected by the domain knowledge; corrected weight = original weight + K·Δ (Δ is the increment given by the expert rule base, and K = 0.1 is the learning rate decay coefficient); 4) Use an online momentum SGD optimizer, set the learning rate to 0.0001 and the momentum to 0.9, and iterate 50 times in small batches (16 samples / batch). After each batch update, project-normalize the output weights (Σw_i = 1 and w_i ≥ 0.1).
[0101] The gradient truncation (threshold 0.01) is introduced in the fine-tuning process to prevent overshooting, and it is verified through A / B testing: the updated model is put into trial operation in 10% of the real-time traffic. If the comprehensive deviation drops by more than 15% within 2 hours, it will be deployed in full. At the same time, a rollback mechanism is established. If ΔE > 10% for three consecutive times after fine-tuning, it will automatically recover to the previous stable version and trigger expert intervention diagnosis.
[0102] In some embodiments, step 204 may specifically include the following steps: S2041, construct a multi-objective weighted aggregation function based on the real-time target weight coefficient.
[0103] Specifically, based on the real-time target weight coefficient (such as economy 0.6, environmental protection 0.3, flexibility 0.1), a weighted aggregation objective function can be constructed: total cost = economy weight × fuel cost (Σ coal consumption rate × output) + environmental protection weight × emission penalty term (Σ excessive emission amount × fine coefficient) + flexibility weight × frequency modulation performance loss (1 / frequency modulation response rate). Among them, the fuel cost and the emission penalty term need to be normalized to the [0, 1] interval to avoid the imbalance of the objective function caused by the dimension difference. For example, the fuel cost normalization formula is (current value - historical minimum value) / (historical maximum value - historical minimum value), and the fine coefficient is dynamically adjusted according to the real-time environmental protection policy (such as an additional fine of 200 yuan / minute for every milligram of NOx exceeding the standard).
[0104] S2042, according to the multi-objective weighted aggregation function, call a fast heuristic algorithm to generate a basic feasible solution.
[0105] Specifically, a greedy algorithm or an improved genetic algorithm can be used to quickly generate a basic solution.
[0106] Taking the greedy algorithm as an example: 1) Sort the unit list in ascending order of coal consumption rate, and preferentially start the unit with the lowest coal consumption rate until its technical output lower limit; 2) Gradually increase the output according to the load demand, and each time select the unit with the smallest marginal cost increment (Δ cost / Δ output) until the load balance is satisfied; 3) Introduce a relaxation mechanism to temporarily relax the ramp rate constraint by 10% (for example, the maximum ramp rate of the unit is temporarily increased from 5 MW / min to 5.5 MW / min) to accelerate the search.
[0107] The genetic algorithm uses binary coding (unit start-stop status) + real number coding (output value). The initial population is generated by random perturbation based on the current operating state. The fitness function is the weighted total cost. After 50 generations of iteration, the Top10 individuals are selected as the basic solutions.
[0108] S2043, expand the basic feasible solution by the parameter perturbation method to obtain multiple different candidate solutions.
[0109] Specifically, the parameter perturbation method generates different candidate solutions in the adjacent space of the basic feasible solution through a controllable random perturbation mechanism. The specific implementation method is as follows: Unit output perturbation: Randomly select 1 to 3 non-critical units in the basic solution (such as units with an output ratio < 30% or units in continuous operation), and apply a random perturbation of ±3% - 5% of the rated capacity on their current output values. For example, a unit has a rated capacity of 300 MW and a current output of 180 MW (60%). The perturbation amplitude is ±5% (i.e., ±15 MW), generating new output values of 185 MW or 175 MW, while ensuring that the adjusted output does not exceed the limit (such as not less than the minimum technical output of 120 MW). If the perturbation causes the output to enter the vibration forbidden zone (such as the 150 - 160 MW interval), it will automatically shift to the nearest safe point (such as 161 MW).
[0110] Objective weight coefficient perturbation: Make a small-range random offset of the real-time weight vector to generate multiple groups of perturbed weights. For example, the original weights are economy 0.6, environmental protection 0.3, and flexibility 0.1. When perturbing, adjust within the range of ±5%: the economy weight is randomly increased by 0.03 (0.63) or decreased by 0.03 (0.57), the environmental protection weight is adjusted in the opposite direction accordingly (0.27 or 0.33), and the flexibility weight remains 0.1 unchanged. Finally, normalize to make the sum equal to 1. Each group of perturbed weights recalculates the weighted objective function and generates a new solution. For example, the weights (0.63, 0.27, 0.1) may bias towards a more aggressive scheduling scheme for economy.
[0111] Start-stop state perturbation: Flip the start-stop state of standby units in the basic solution that meet the minimum shutdown time (e.g., if the unit has been shutdown for ≥ 4 hours). For example, if a unit is in the shutdown state and the shutdown time meets the standard, attempt to start it and allocate the minimum technical output (e.g., 120 MW), and at the same time shut down another low-efficiency unit to maintain power balance. Only one unit's state change is allowed for each perturbation to generate complementary start-stop combination plans.
[0112] Time series perturbation: For the scenario of fluctuating frequency regulation demand, apply time series translation perturbation to the output plan in the basic solution. For example, shift the output curve of a unit at time t forward or backward by 5 minutes as a whole, and re-verify the ramp rate constraint (e.g., 5 MW / min). Such perturbations can explore the flexibility potential under different time granularities.
[0113] Through the above multi-dimensional perturbation strategy, each basic feasible solution can derive 5 - 10 or even more candidate solutions. For example, basic solution A generates A1 (output +5% of unit X), A2 (weights 0.63 / 0.27 / 0.1), A3 (start standby unit Y), etc. All perturbed solutions need to be filtered by a fast verification module (such as pre-computing the boundary of the feasible region) to filter out obviously violating solutions, and finally form a set of differentiated candidate solutions covering different target trade-off directions to ensure that the solution set is evenly distributed in the target space.
[0114] In some embodiments, this step S2043 may further specifically include: (1) Based on the real-time target weight coefficient and the basic feasible solution, determine the perturbation parameters and perturbation directions.
[0115] Specifically, based on the characteristics of the real-time target weight coefficient and the basic feasible solution (such as the economic weight ratio exceeding 60%), select the main perturbation parameters (economic weight, output of key units) and perturbation directions. For example, if the economic weight is 0.6, the perturbation direction is set in the ±5% interval (0.57 - 0.63), and the output adjustment direction of the key unit (the unit with the lowest coal consumption rate) is associated (±3% - 5% of the rated capacity); the environmental protection weight is perturbed in the opposite direction (0.3 ± Δ, where Δ is dynamically adjusted according to the load volatility, such as when the load volatility > 5%, the upper limit of Δ is set to 0.05). The perturbation parameter selection strategies include: 1) Main target sensitivity analysis, determine the influence weight of the weight perturbation amount on the objective function through gradient calculation; 2) Priority ranking of unit output perturbation, perturb in ascending order of coal consumption rate to ensure that the economically sensitive units are adjusted first.
[0116] (2) Generate multiple groups of parameter combinations according to the perturbation parameters and perturbation directions.
[0117] Specifically, the disturbance parameter combinations can be generated through Monte Carlo sampling: 1) Weight coefficient disturbance: ±5% random disturbance is applied to the economic, environmental protection and flexibility weights respectively. For example, the original weights (0.6, 0.3, 0.1) may derive combinations such as (0.63, 0.27, 0.1) and (0.58, 0.31, 0.11), and the sum is 1 after normalization; 2) Output baseline disturbance: Select the units with the top 30% output in the basic solution, and randomly adjust them according to ±5% of the rated capacity based on their current output. For example, the output of a 300MW unit is adjusted from 200MW to 185MW or 215MW, and it is prohibited to enter the vibration zone; 3) Start-stop state disturbance: The state of non-critical units (such as units with output <20% of the rated capacity and the downtime meets the standard) is flipped (0→1 or 1→0). Each parameter combination includes a weight vector, an output baseline, and start / stop state change instructions, generating a total of 50 to 100 differentiated combinations.
[0118] (3) The weight coefficients and output baselines of multiple sets of parameter combinations are adjusted to obtain multiple sets of directional parameter combinations.
[0119] In this embodiment, directional adjustment can be performed on each parameter combination: 1) Weight coefficient adjustment: reconstruct the objective function according to the weight vector after disturbance, for example, when the economic weight is 0.63, the total cost = 0.63 × fuel cost + 0.27 × emission penalty + 0.1 × frequency regulation cost; 2) Output baseline correction: based on the output value after perturbation, recalculate the power balance constraint. If the adjustment causes power shortage, the output is supplemented from low to high according to the coal consumption rate (for example, when the shortage is 50MW, the output of the unit with the lowest coal consumption rate is increased to the upper limit first); 3) Start-stop status synchronization update: if the status of a unit is changed from shutdown to startup, the minimum shutdown time (such as ≥4 hours) needs to be verified and its minimum technical output (such as 120MW) is forcibly allocated. During the adjustment process, a fast feasibility verification module is used to eliminate obvious illegal combinations (such as power shortage > 5% or environmental protection instantaneous excess > 10%), and retain 80% of the combinations to enter the solution stage.
[0120] (4) Each set of directional parameter combinations is used as input for independent solution to obtain differentiated candidate solutions.
[0121] Specifically, each set of orientation parameter combinations can be input into a mixed-integer linear programming (MILP) solver (such as CPLEX / Gurobi), and a short-term solution limit can be set (e.g., single solution ≤ 10 seconds). The solution strategies include: 1) Warm start: Inherit the unit start-stop status and output baseline of the basic solution as the initial solution; 2) Constraint relaxation: Allow the ramp rate to be temporarily relaxed to 110% (e.g., 5.5 MW / min) to accelerate convergence; 3) Local search: Limit the adjustment range of decision variables to ±10% of the basic solution (e.g., the output is only allowed to search between 190 and 210 MW). After the solution is completed, non-dominated solutions are extracted and hard constraints are verified: If the emissions exceed the standard or there is a power imbalance, a repair mechanism is triggered (e.g., enabling a standby unit to compensate for the output), otherwise it is stored in the candidate solution pool.
[0122] Finally, 80 - 120 groups of differentiated candidate solutions are generated, covering multiple trade-off directions such as economy priority, environmental protection priority, and flexibility priority, and the diversity of the solution set is verified by the hypervolume index (HV ≥ 0.8).
[0123] S2044, perform hard filtering of safety constraints on multiple differentiated candidate solutions to obtain an initial set of candidate scheduling plans.
[0124] Specifically, the implementation process of performing hard filtering of safety constraints on multiple differentiated candidate solutions is as follows: First, a parallel rule engine is constructed, and three types of hard constraint rules at the unit level, grid level, and environmental protection level are predefined. Unit-level rules include the minimum start-stop time (e.g., need to restart after 4 hours of shutdown), output upper and lower limits (e.g., 200 MW ≤ P_i ≤ 600 MW), prohibited entry into the vibration zone (prohibit output in the range of 150 - 180 MW), and cumulative operation duration limit (forced shutdown if ≤ 72 hours); Grid-level rules cover node voltage deviation ≤ ±5%, tie-line power factor ≥ 0.9, and frequency fluctuation ≤ ±0.2 Hz; Environmental protection-level rules include instantaneous emission concentration (e.g., SO2 < 35 mg / m³), hourly average value (NOx ≤ 50 mg / m³), and daily cumulative emissions (dust ≤ 10 tons).
[0125] Subsequently, hierarchical verification and rapid pruning are performed: Each candidate solution is subjected to three-level verification in the order of unit level → grid level → environmental protection level. In the first-level unit constraint verification, all unit outputs and start-stop states are traversed. If a certain unit violates the vibration zone rule (e.g., the output of 160 MW enters the prohibited interval) or the shutdown time is insufficient (restart after 2 hours of shutdown), the solution is directly eliminated. In the second-level grid constraint verification, a simplified power flow calculation model is called. If a certain solution causes the node voltage deviation to exceed the limit (e.g., 6%) or the line power factor is lower than 0.9, it is immediately marked as a violation. In the third-level environmental protection constraint verification, based on the emission characteristic curve (e.g., the quadratic function relationship between unit output and NOx emission), the instantaneous and cumulative emissions are calculated. If the hourly average NOx emission reaches 52 mg / m³ (exceeding the limit by 4%), the solution is eliminated.
[0126] For critical compliance solutions (such as hourly emissions of 50.5 mg / m³), a dynamic relaxation mechanism is enabled: the hourly mean check is split into 5-minute granularity. If only a single period is slightly exceeded (such as 51 mg / m³ in the first 5 minutes and 49 mg / m³ in the next 55 minutes), it is judged as compliant after weighted averaging; if a unit exceeds the standard but other units in the same plant can be linked to activate backup desulfurization equipment (such as increasing processing capacity by 5%), the solution is allowed to be conditionally retained and the equipment compensation logic is triggered.
[0127] Finally, the solutions that have passed the verification are aggregated and their diversity is controlled: they are screened according to the spatial distribution density of the three-dimensional target of economy, environmental protection, and flexibility. If the solutions in a certain area are too dense (for example, 10 solutions are clustered within 5% of the spatial volume), the crowding distance of each solution (Euclidean distance between adjacent solutions) is calculated, and the top 20 solutions with uniform distribution are retained. The final output set must meet 100% hard constraint compliance, target space coverage (HV ≥ 0.75), and solution spacing ≥ 3% of the total space diameter. Millisecond-level batch processing is achieved through a distributed computing framework (such as Spark). 100,000 candidate solutions can be filtered into 200 to 500 sets of strictly compliant initial solution sets within 10 seconds.
[0128] In some embodiments, step 205 may specifically include the following steps: S2051, determining the primary optimization target parameters and the secondary constraint target thresholds based on the real-time target weight coefficient.
[0129] Specifically, based on the real-time weight coefficient (such as economy 0.6, environmental protection 0.3, flexibility 0.1), the target with the highest weight (economy) can be selected as the main optimization target parameter, and the remaining targets are set as secondary constraint target thresholds: the environmental protection threshold is set to 110% of the real-time emission limit (that is, a temporary excess of 10% is allowed), and the flexibility threshold is set to a frequency regulation rate not less than 85% of the grid demand forecast value. The secondary constraint threshold is determined by historical scenario statistics. For example, when the environmental protection weight is 0.3, the corresponding emission cap relaxation coefficient is 1.1.
[0130] S2052: According to the primary optimization objective parameter, a subset of the solutions to be optimized that meet the secondary constraint objective threshold is selected from the initial candidate scheduling solution set.
[0131] Specifically, all schemes that violate secondary constraints (such as emissions exceeding the limit by 110% and frequency modulation rates below 85%) can be eliminated from the initial set of candidate scheduling schemes (for example, 20 groups of schemes), and the remaining schemes can be sorted according to the main objective (economic cost), retaining the 50% schemes with the lowest cost (such as 10 groups) as the subset of schemes to be optimized.
[0132] During the screening process, the relaxation comparison method is adopted: If the emissions of a certain plan are between 105% and 110% of the limit value, but the economic cost is significantly lower than that of other plans (the difference > 5%), it is retained and marked as an object that needs to be optimized with emphasis.
[0133] S2053, optimize and solve the subset of the plan to be optimized through a mixed-integer programming model to obtain the Pareto solution set.
[0134] Specifically, construct a MILP model with the main objective (minimizing economic cost) as the objective function and the secondary objectives as constraint conditions (emissions ≤ 110% of the limit value, frequency modulation rate ≥ 85%). The decision variables include the unit start-stop state (0-1 variable), output allocation (continuous variable), and environmental protection equipment operation combination (integer variable). Solve using the branch and bound method. For each plan in the subset to be optimized, use its initial unit combination as the hot start input and perform local search within the constraint boundary (for example, the output adjustment step size is 5MW). To generate the Pareto solution set, use the ε-constraint method: After obtaining each optimal solution, tighten the environmental protection constraint by 5% (for example, from 110% to 105%) and re-solve until there is no feasible solution. Finally, output multiple non-dominated solutions that satisfy the three-dimensional Pareto front of economy-environment-flexibility, and store them as a structured data set containing the unit state matrix, objective function value, and constraint satisfaction flag.
[0135] In some embodiments, this step may further include: (1) Construct a mixed-integer programming model that includes the main objective function, secondary constraint conditions, and equipment operation rules.
[0136] Specifically, the mixed-integer programming model takes the economic objective (such as minimizing the total fuel cost) as the main objective function, and the mathematical expression can be minΣ(a_i·x_i + b_i·y_i), where x_i is a continuous variable (the output value of unit i), y_i is a 0-1 variable (the start-stop state of unit i), and a_i and b_i are the fuel cost coefficient and start-stop cost respectively.
[0137] The secondary constraint conditions include environmental protection constraints (Σc_ij·x_i ≤ D_j, j is the pollutant type, D_j is the dynamic emission limit value), and flexibility constraints (Σk_i·Δx_i / Δt ≥ F_min, k_i is the ramp rate of unit i, F_min is the minimum grid frequency modulation requirement).
[0138] The equipment operation rules are transformed into hard constraints: 1) The minimum start-stop time constraint of the unit (y_i needs to satisfy T_on ≥ T_min_on, T_off ≥ T_min_off); 2) The constraint of not entering the vibration area (x_i ∉ [P_low, P_high]); 3) The linkage constraint of environmental protection equipment (if the desulfurization equipment is out of service, then the associated unit x_i = 0).
[0139] The mixed-integer programming model is implemented through the modeling interfaces of Gurobi or CPLEX (such as the Python API). The dimension of the decision variables is n (the number of units) × 2 (x_i + y_i), and the number of constraints is 3n + m (m is the number of environmental protection and grid interaction constraints).
[0140] (2) Based on the mixed-integer programming model, call the solver to perform multi-threshold relaxation to obtain a set of candidate solutions.
[0141] Specifically, when calling the CPLEX solver, the ε-constraint method is used to generate a set of candidate solutions: 1) Initialize the relaxation threshold, relax the secondary objective (such as environmental protection) constraint to 120% of the limit (Σc_ij·x_i ≤ 1.2D_j), and solve the primary objective to obtain the initial solution; 2) Gradually tighten the constraint threshold (step size 5%, that is, the next-round constraint is 1.15D_j), and retain the non-dominated solutions after each solution; 3) Parallel solution strategy: Divide different thresholds into multiple threads (such as 10 threads corresponding to thresholds 1.2D_j, 1.15D_j,..., 0.8D_j), and each thread independently calls the MILP solver to accelerate the search process.
[0142] Among them, the key technologies include: warm start uses the feasible solutions in the initial candidate solution set to accelerate branch and bound, and lazy constraints dynamically add vibration zone constraints to reduce the amount of calculation. The solution termination condition is that no new solutions are generated for 3 consecutive iterations or the maximum calculation time is reached (such as 30 minutes). Finally, a candidate set containing 50 - 100 groups of solutions is output, covering multi-dimensional trade-off schemes of economy - environmental protection - flexibility.
[0143] (3) Extract the Pareto solution set from the candidate solution set through non-dominated sorting and uniform screening strategies.
[0144] Specifically, first perform fast non-dominated sorting on the candidate solution set: 1) Screening of the first-layer non-dominated solutions: Traverse all solutions. If a solution is not dominated by other solutions in all objectives (such as at least one of the economy ≤ solution B, environmental protection ≤ solution B, and flexibility ≤ solution B of solution A is strictly better), it is classified into the first layer; 2) Repeat the process after removing the first-layer solutions to generate the second layer, the third layer, etc., until all solutions are stratified.
[0145] Subsequently, the crowding distance screening method is used to improve the uniformity of the solution set: 1) Normalize each objective dimension; 2) Calculate the Euclidean distance (crowding distance) of each solution between its adjacent solutions. For example, for the solution sequence sorted by economy, the crowding distances of the head and tail solutions are set to infinity, and the crowding distance of the middle solutions is (flexibility difference + environmental protection difference); 3) Sort by crowding distance from large to small, and retain the first 30 solutions.
[0146] The obtained Pareto solution set needs to meet the following requirements: the proportion of the number of solutions in each layer (the first layer ≥ 70%), the coverage rate of the target space (HV index ≥ 0.85), and the coefficient of variation of the distance between solutions (standard deviation / mean ≤ 0.25). When the standard is not met, supplementary local search is triggered: inject neighborhood perturbation solutions (±5% output adjustment) into sparse regions (such as the high-economy - low-environmental protection interval), and re-evaluate until convergence.
[0147] S2054, verify whether the Pareto solution set meets the preset convergence conditions, and the convergence conditions include the variance threshold of the objective function and the uniformity index of the solution set distribution.
[0148] Specifically, first calculate the variance of the objective function. Normalize the economic objectives (such as unit power supply cost), environmental protection objectives (such as emission intensity), and flexibility objectives (such as frequency modulation response time) of all solutions in the solution set respectively, and map the values of each dimension to the interval [0, 1]. Then calculate the variance of each objective dimension. For example, the formula for calculating the economic variance is the sum of the squares of the differences between the cost values of all solutions and the average cost divided by (n - 1), where n is the number of solutions in the solution set. The variance threshold of the objective function requires that the economic variance does not exceed 0.5% (normalized value), and the variances of environmental protection and flexibility do not exceed 1%. If it is detected that the variance of a certain objective exceeds the standard (such as the economic variance reaches 0.6%), it is determined that the solution set does not converge on this core objective, and a supplementary optimization process needs to be triggered.
[0149] Next, evaluate the uniformity of the solution set distribution. Use the hypervolume (HV) index to measure the coverage range of the solution set in the target space: take the worst values of each objective dimension (such as economic normalization value 1.0, environmental protection 1.0, flexibility 1.0) as reference points, and calculate the volume of the hypercube formed by the solution set in the three-dimensional space. It is required that the HV value is not less than 0.8 (the maximum theoretical volume after normalization is 1). At the same time, calculate the Euclidean distance uniformity of adjacent solution pairs. The specific method is: after sorting all solutions by economy, calculate the distance between each pair of adjacent solutions in the target space, and find the ratio of the standard deviation σ to the average value of these distances. If this ratio exceeds 0.3 (for example, when σ = 0.12 and the average distance = 0.4, σ / mean is 0.3), it is determined that the distribution is uneven.
[0150] Finally, execute the comprehensive verification logic. The solution set is determined to meet the convergence condition only when all target variances meet the standards, the HV value ≥ 0.8, and the distance uniformity ratio ≤ 0.3. When any condition is not met, the system automatically records the defect type (such as the economic variance exceeding the standard or the HV being insufficient) and triggers the iterative optimization mechanism. The verification module will generate a three-dimensional scatter plot and a numerical index report. For example, when a solution set contains 10 solutions with adjacent distances of 0.3, 0.5, 0.4, and 0.2 respectively, the average distance is 0.35, the standard deviation σ ≈ 0.12, and σ / mean ≈ 0.34. At this time, it will be marked as unevenly distributed. The manual review link can further check for outliers. If a solution deviates from the main distribution cluster by more than 2 standard deviations (such as a solution with an environmental protection score of 0.9 while other solutions are all below 0.7), it is necessary to decide whether to retain or remove this outlier solution.
[0151] S2055, determine the Pareto solution set that meets the preset convergence condition as the Pareto optimal solution set.
[0152] Deduplicate the verified solution set (solutions with a cosine similarity > 0.95 are regarded as duplicate solutions), and generate the final Pareto optimal solution set sorted by economy, which is stored as JSON structured data, including the objective function values of each solution, the unit output allocation matrix, the constraint satisfaction flag, and the solution timestamp, for subsequent decision-making modules to call.
[0153] At the same time, record metadata such as the constraint relaxation parameters and the number of iterations at the time of convergence for model effect retrospective analysis.
[0154] The method provided in the above embodiments can be executed by an electronic device. The following describes this electronic device in the embodiments of the present invention from the perspective of hardware processing. Please refer to Figure 3 , which is a schematic structural diagram of an entity device of the electronic device in the embodiments of the present invention.
[0155] It should be noted that Figure 3 The structure of the electronic device shown is only an example and should not bring any limitations to the functions and usage scopes of the embodiments of the present invention.
[0156] Such as Figure 3As shown, the electronic device includes a Central Processing Unit (CPU) 401, which can perform various appropriate actions and processes according to the program stored in the Read-Only Memory (ROM) 402 or the program loaded from the storage section 408 into the Random Access Memory (RAM) 403, such as executing the method described in the above embodiments. In the Random Access Memory (RAM) 403, various programs and data required for system operation are also stored. The Central Processing Unit (CPU) 401, the Read-Only Memory (ROM) 402, and the Random Access Memory (RAM) 403 are connected to each other via a bus 404. The Input / Output (I / O) interface 405 is also connected to the bus 404.
[0157] The following components are connected to the Input / Output (I / O) interface 405: an input section 406 including an audio input device, a button switch, etc.; an output section 407 including a display, an audio output device, an indicator light, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the Input / Output (I / O) interface 405 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 410 as needed so that a computer program read from it can be installed into the storage section 408 as needed.
[0158] Specifically, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 409, and / or installed from the removable medium 411. When the computer program is executed by the Central Processing Unit (CPU) 401, various functions defined in the present invention are executed.
[0159] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0160] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the block may occur in a different order from that marked in the accompanying drawings.
[0161] Specifically, the electronic device in this embodiment includes a processor and a memory. The memory is coupled to one or more processors. The memory is used to store computer program code, and the computer program code includes computer instructions. One or more processors call the computer instructions to enable the electronic device to execute the method provided in the above embodiment.
[0162] On the other hand, the present invention also provides a computer-readable storage medium. This storage medium may be included in the electronic device described in the above embodiment; or it may exist separately and not be assembled into the electronic device. The above storage medium carries one or more computer programs. When the above one or more computer programs are executed by a processor of the electronic device, the electronic device is enabled to implement the method provided in the above embodiment.
[0163] As mentioned above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention.
[0164] As used in the foregoing embodiments, depending on the context, the term "when" may be construed to mean "if", "after", "in response to determining", or "in response to detecting". Similarly, depending on the context, the phrase "when determining" or "if (the stated condition or event) is detected" may be construed to mean "if determined", "in response to determining", "when (the stated condition or event) is detected", or "in response to detecting (the stated condition or event)".
[0165] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the foregoing embodiments can be implemented by a computer program instructing relevant hardware. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the foregoing method embodiments. The foregoing storage media include various media that can store program codes, such as ROM, random access memory (RAM), magnetic disks, or optical discs.
Claims
1. A multi-objective economic dispatch optimization method for thermal power plants, characterized in that, include: Construct an initial data set based on the power grid operation parameters, generator set status parameters and environmental monitoring data of the thermal power plant; Determine characteristic parameters of the current operation and dispatching demand according to the initial data set, wherein the characteristic parameters include load demand fluctuation rate, unit adjustable margin and environmental protection limit compliance rate; based on the characteristic parameters, determine the real-time target weight coefficients of the economic index, environmental protection index and dispatching flexibility index through a preset dynamic weight allocation model; Generating an initial candidate scheduling solution set that satisfies the safety constraints of the generator set according to the real-time target weight coefficient; Performing main target oriented optimization on the initial candidate scheduling solution set to obtain a Pareto optimal solution set that meets a preset convergence condition; A scheduling recommendation scheme is generated based on the Pareto optimal solution set.
2. The method according to claim 1, wherein The construction and training process of the dynamic weight allocation model includes: constructing a multimodal hybrid model structure, which includes an input layer for integrating power grid, generator set data and environmental data, a constraint encoding layer embedded in physical mechanism equations, and a hierarchical strategy network based on a multi-head attention mechanism; Generate a simulation data set containing multiple working conditions through simulation technology, wherein the simulation data set covers fuel shortage, environmental protection equipment failure and power grid frequency disturbance scenarios; Annotating the Pareto optimal weight vector of each scene in the simulation data set to form a training label set; Using the training label set and the physical mechanism equation as network constraint rules, pre-training the initial parameters of the hierarchical strategy network to obtain a pre-training model; The hierarchical strategy network of the pre-trained model is subjected to phased reinforcement learning training in single-unit steady-state, multi-unit coupling and random failure scenarios respectively to obtain the dynamic weight allocation model.
3. The method according to claim 2, wherein After performing phased reinforcement learning training on the hierarchical strategy network of the pre-trained model in the single unit steady state, multi-unit coupling and random failure scenarios to obtain the dynamic weight allocation model, the method further includes: Introduce adversarial generative network to construct perturbation sample set; Performing robustness reinforcement training on the dynamic weight allocation model according to the disturbance sample set; Through the online transfer learning framework, the output layer parameters of the dynamic weight allocation model are dynamically fine-tuned based on the deviation between the actual scheduling results and the simulation prediction.
4. The method according to claim 1, characterized in that The performing of main target oriented optimization on the initial candidate scheduling solution set to obtain a Pareto optimal solution set that meets a preset convergence condition includes: Determine the primary optimization target parameter and the secondary constraint target threshold based on the real-time target weight coefficient; According to the primary optimization objective parameter, a subset of the to-be-optimized solutions that meet the secondary constraint objective threshold are selected from the initial candidate scheduling solution set; Optimizing and solving the subset of solutions to be optimized by a mixed integer programming model to obtain a Pareto solution set; Verify whether the Pareto solution set meets a preset convergence condition, wherein the convergence condition includes an objective function variance threshold and a solution set distribution uniformity index; The Pareto solution set that meets the preset convergence conditions is determined as the Pareto optimal solution set.
5. The method according to claim 4, characterized in that, The mixed integer programming model is used to optimize and solve the subset of solutions to be optimized to obtain a Pareto solution set, including: Construct a mixed-integer programming model that includes a main objective function, secondary constraints, and equipment operation rules; Based on the mixed-integer programming model, call a solver to perform multi-threshold relaxation solving to obtain a set of candidate solutions; Extract the Pareto solution set from the set of candidate solutions through non-dominated sorting and uniform screening strategies.
6. The method according to claim 1, wherein The generating an initial set of candidate scheduling plans that satisfy the safety constraints of the generator set according to the real-time objective weight coefficient includes: Construct a multi-objective weighted aggregation function based on the real-time objective weight coefficient; According to the multi-objective weighted aggregation function, call a fast heuristic algorithm to generate a basic feasible solution; Expand the basic feasible solution by the parameter perturbation method to obtain multiple different candidate solutions; Perform hard filtering of safety constraints on the multiple different candidate solutions to obtain an initial set of candidate scheduling plans.
7. The method according to claim 6, characterized in that The expanding the basic feasible solution by the parameter perturbation method to obtain multiple different candidate solutions includes: Based on the real-time objective weight coefficient and the basic feasible solution, determine the perturbation parameter and the perturbation direction; Generate multiple sets of parameter combinations according to the perturbation parameter and the perturbation direction; Adjust the weight coefficients and output baselines of the multiple sets of parameter combinations to obtain multiple sets of directional parameter combinations; Use each set of the directional parameter combinations as an input for independent solving to obtain different candidate solutions.
8. An electronic device, characterized in that, Comprising one or more processors and a memory; The memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the electronic device to execute the method according to any one of claims 1-7.
9. A computer-readable storage medium storing computer instructions, characterized in that, When the computer instructions are run on an electronic device, the electronic device is caused to execute the method according to any one of claims 1-7.
10. A computer program product, characterized in that, When the computer program product is run on an electronic device, the electronic device is caused to execute the method according to any one of claims 1-7.
Citation Information
Patent Citations
Environmental economy power generation dispatching method
CN104009494A
Thermal power plant environment economic dispatching method based on multi-target differential evolution algorithm
CN105809297A
Multi-optimization target weighting method and system in construction of integrated energy system
CN112949177A
Intelligent charging pile scheduling method and system based on dynamic adjustment of energy storage battery pack
CN118966580A
Multi-target flexible workshop adaptive scheduling method
CN119575910A
Cited By
Multi-agent optimal compromise scheduling method for multi-agent hydrogen-containing comprehensive energy system
CN120450394A
Energy storage operation optimization method and system based on hierarchical multi-agent reinforcement learning
CN120952277A
Network source collaborative primary frequency modulation optimization method and system based on reinforcement learning
CN121192745A
Intelligent optimization method and system for multi-process production energy efficiency of lithium battery diaphragm
CN121209433A
Method for evaluating coupling frequency modulation control capability of super capacitor and thermal power generating unit
CN121216518A