A multi-objective economic dispatch optimization method, equipment, medium and product for thermal power plants
By building a dynamic weight distribution model and reinforcement learning, the conflict problem in the multi-objective scheduling of thermal power plants was solved, a dynamic balance between economy, environmental protection and flexibility was achieved, and the real-time and adaptability of the scheduling strategy was improved.
Patent Information
- Application Number
- CN202510756359.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-06-09
AI Technical Summary
Existing economic dispatch technology for thermal power plants is difficult to effectively balance the conflicts among multi-objective optimization, resulting in an imbalance in the dispatch scheme.
Build a dynamic weight allocation model based on real-time scenario characteristics, dynamically adjust the weight coefficients of economic, environmental and flexibility goals, combine digital twin systems and reinforcement learning, and generate a scheduling plan that takes into account multi-dimensional balance.
The real-time performance and scenario generalization capabilities of thermal power plant scheduling have been improved, an efficient multi-objective collaborative optimization solution set has been generated, and the real-time and adaptability of the scheduling strategy have been significantly improved.
Smart Images

Figure CN120258488B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a multi-objective economic dispatch optimization method, equipment, medium and product for a thermal power plant. Background Art
[0002] The economic dispatch of thermal power plants is a decision-making process that achieves a balance between economy (fuel cost, operation and maintenance cost, etc.) and multiple objectives (environmental emissions, grid flexibility, equipment health, etc.) by optimizing the unit operation strategy, while meeting the grid load demand, equipment safety constraints (such as upper and lower output limits, ramp rate) and environmental protection standards.
[0003] In related technologies, the economic dispatch technology of thermal power plants usually adopts mixed integer linear programming (MILP) or dynamic programming methods, with fuel cost minimization as the objective function, and solves the optimal unit combination and output distribution plan while meeting hard constraints such as unit output upper and lower limits, ramp rate, start and stop time, etc.
[0004] However, the scheduling plans generated by related technologies rely more on preset weights or priorities, cannot cover the complex trade-offs between multiple objectives, and are difficult to effectively balance the conflicts between multi-objective optimizations in economic scheduling plans. Summary of the Invention
[0005] In view of the above technical problems and defects, the purpose of the present invention is to provide a multi-objective economic dispatch optimization method, equipment, medium and product for thermal power plants, which can alleviate the problem of multi-objective optimization imbalance in the economic dispatch of thermal power plants and lead to multi-objective conflicts.
[0006] To achieve the above-mentioned purpose, in a first aspect, the present invention provides a multi-objective economic dispatch optimization method for a thermal power plant, comprising: constructing an initial data set based on the power grid operating parameters, generator set status parameters and environmental monitoring data of the thermal power plant; determining characteristic parameters of the current operation dispatch demand according to the initial data set, the characteristic parameters including load demand fluctuation rate, unit adjustable margin and environmental protection limit compliance rate; based on the characteristic parameters, determining the real-time target weight coefficients of the economic index, environmental index and dispatch flexibility index through a preset dynamic weight allocation model; generating an initial candidate dispatch scheme set that meets the safety constraints of the generator set according to the real-time target weight coefficient; performing main target-oriented optimization on the initial candidate dispatch scheme set to obtain a Pareto optimal solution set that meets preset convergence conditions; and generating a dispatch recommendation scheme based on the Pareto optimal solution set.
[0007] The present invention effectively solves the inherent conflicts between economic, environmental and flexibility goals by constructing a dynamic weight allocation model based on real-time scenario characteristics. First, based on key characteristic parameters such as load fluctuation rate, unit regulation potential and environmental protection compliance rate, the real-time weight coefficient of each goal is dynamically adjusted to break through the limitation of poor adaptability of the traditional fixed weight model to complex working conditions; then, through main target oriented optimization and Pareto front search, a set of scheduling solutions that take into account multi-dimensional balance is generated in the solution space that meets safety constraints, which not only avoids scheduling imbalances caused by excessive dominance of a single goal, but also ensures that the solution set covers the optimal trade-off surface of different target priorities; finally, the global optimal solution is recommended in combination with multi-criteria decision-making to achieve the coordinated optimization of the economic cost, emission control and grid response capability of thermal power scheduling. This method forms a closed loop from dynamic weight mapping, multi-objective solution set generation to intelligent decision recommendation, providing a scenario-based adaptive solution for multi-objective conflicts and improving the comprehensive scheduling capability of thermal power plants.
[0008] Optionally, in some embodiments, the construction and training process of the dynamic weight distribution model includes: constructing a multimodal hybrid model structure, the multimodal hybrid model structure includes an input layer for integrating power grid, generator set and environmental data, a constraint encoding layer embedded with physical mechanism equations, and a hierarchical strategy network based on a multi-head attention mechanism; generating a simulation data set containing multiple working conditions through a digital twin system, the simulation data set covering fuel shortage, environmental protection equipment failure and power grid frequency disturbance scenarios; marking the Pareto optimal weight vector of each scenario in the simulation data set to form a training label set; using the training label set and the physical mechanism equation as network constraint rules, pre-training the initial parameters of the hierarchical strategy network to obtain a pre-trained model; performing phased reinforcement learning training on the hierarchical strategy network of the pre-trained model under single unit steady-state, multi-unit coupling and random failure scenarios, respectively, to obtain a dynamic weight distribution model.
[0009] Adopting the technical solutions of the aforementioned embodiments, a full-lifecycle approach to constructing a dynamic weight allocation model is defined by building a hybrid model architecture that integrates multimodal data and physical mechanisms. This architecture combines diverse simulation data generated by simulation technologies (such as digital twins) with reinforcement learning training. The multimodal input layer enables deep integration of grid, unit, and environmental data. The constraint encoding layer embeds physical rules, such as boiler efficiency equations and emission characteristic curves, into the network to ensure the physical feasibility of model decisions. A hierarchical policy network dynamically captures inter-objective game relationships based on a multi-head attention mechanism. Phased reinforcement learning (from single-machine steady state to multi-machine coupling to random failures) allows the model to gradually adapt to complex scenarios. The simulation dataset covers extreme operating conditions, such as fuel shortages and equipment failures, ensuring the generalization of weight allocation. This approach overcomes the scenario adaptability limitations of traditional static weight models and provides a highly accurate and interpretable decision-making core for dynamic multi-objective optimization.
[0010] Optionally, in some embodiments, after performing phased reinforcement learning training on the hierarchical strategy network of the pre-trained model in single-unit steady-state, multi-unit coupling and random failure scenarios to obtain a dynamic weight allocation model, it also includes: introducing a generative adversarial network to construct a perturbation sample set; performing robustness reinforcement training on the dynamic weight allocation model based on the perturbation sample set; and dynamically fine-tuning the output layer parameters of the dynamic weight allocation model based on the deviation between the actual scheduling results and the simulation prediction through an online transfer learning framework.
[0011] Adopting the technical solution of the above embodiment, adversarial generative networks and online transfer learning are introduced after dynamic weight model training to systematically improve the model's robustness and continuous optimization capabilities. The adversarial network constructs hidden disturbance samples such as sensor failures and control signal interference, forcing the model to identify abnormal patterns and enhance anti-interference capabilities during adversarial training; robustness enhancement training suppresses weight mutations through reward function design to ensure decision stability under extreme working conditions. The online transfer learning module monitors actual scheduling deviations in real time and dynamically fine-tunes output layer parameters to enable the model to quickly adapt to changes in the operating environment (such as coal quality fluctuations and policy adjustments). This method solves the problem of domain offset between offline training models and real-world scenarios, and realizes the long-term and reliable application of models in complex industrial environments.
[0012] Optionally, in some embodiments, the initial set of candidate scheduling schemes is optimized in a main target-oriented manner to obtain a Pareto optimal solution set that meets preset convergence conditions, including: determining the main optimization target parameters and the secondary constraint target threshold based on the real-time target weight coefficient; screening a subset of schemes to be optimized that meet the secondary constraint target threshold from the initial set of candidate scheduling schemes according to the main optimization target parameters; optimizing and solving the subset of schemes to be optimized through a mixed integer programming model to obtain a Pareto solution set; verifying whether the Pareto solution set meets preset convergence conditions, the convergence conditions including the objective function variance threshold and the solution set distribution uniformity index; and determining the Pareto solution set that meets the preset convergence conditions as the Pareto optimal solution set.
[0013] The technical solutions of the above-mentioned embodiments define the complete process of primary-objective directional optimization. Through hierarchical optimization of primary and secondary objectives and a Pareto solution set verification mechanism, both solution quality and computational efficiency are balanced. Based on real-time weights, economic efficiency is determined as the primary objective, and environmental protection / flexibility is used as the constraint threshold, screening a subset of high-potential candidate solutions. A mixed integer programming model is combined with the ε-constraint method to generate Pareto solutions, and convergence is quantified using the objective function variance and distribution uniformity. This method significantly reduces the computational complexity of multi-objective optimization while ensuring solution diversity. By dynamically adjusting constraint thresholds and iterative verification (such as tightening emission limits and completing uniformity), it avoids falling into local optimality, providing a high-coverage non-dominated solution foundation for scheduling recommendations.
[0014] Optionally, in some embodiments, a subset of optimization solutions is optimized and solved through a mixed integer programming model to obtain a Pareto solution set, including: constructing a mixed integer programming model containing a main objective function, secondary constraints and equipment operation rules; based on the mixed integer programming model, calling a solver to perform multi-threshold relaxation solution to obtain a candidate solution set; and extracting a Pareto solution set from the candidate solution set through non-dominated sorting and uniform screening strategies.
[0015] By adopting the technical solutions of the above-mentioned embodiments, the specific technical path for solving the Pareto solution set of mixed integer programming is refined, and a balance between the quality and efficiency of the solution set is achieved through multi-threshold relaxation and non-dominated sorting. A MILP model that includes the start and stop status of the unit, output distribution, and the commissioning of environmental protection equipment is constructed, and the secondary objectives are converted into dynamic constraints; when calling the solver, parallel multi-threshold relaxation is used (such as gradually tightening the emission limit from 120% to 100%), combined with hot start and lazy constraint callback to accelerate the solution. Non-dominated sorting and congestion screening strategies extract a uniformly distributed Pareto frontier from the candidate solutions to ensure that the solution set has extensive coverage in the target space. This method solves the problems of slow convergence and sparse solution sets of traditional multi-objective optimization algorithms through modeling solutions and post-processing strategies.
[0016] Optionally, in some embodiments, an initial set of candidate scheduling schemes that meet the safety constraints of the generator set is generated based on the real-time target weight coefficient, including: constructing a multi-objective weighted aggregation function based on the real-time target weight coefficient; calling a fast heuristic algorithm to generate a basic feasible solution based on the multi-objective weighted aggregation function; expanding the basic feasible solution through a parameter perturbation method to obtain multiple differentiated candidate solutions; and performing safety constraint hard filtering on multiple differentiated candidate solutions to obtain an initial set of candidate scheduling schemes.
[0017] Using the technical solutions of the above-mentioned embodiments, a method for generating an initial set of candidate scheduling solutions is defined, balancing the quality and diversity of solutions through a weighted aggregation function and a parameter perturbation strategy. A multi-objective weighted function is constructed based on real-time weights, driving a heuristic algorithm (such as an improved genetic algorithm) to rapidly generate basic feasible solutions. The parameter perturbation method applies directional perturbations to unit output, weight coefficients, and start / stop states to expand differentiated candidate solutions. Safety constraint hard filtering uses a rule engine to batch-verify hard constraints such as unit safety, grid stability, and environmental compliance to ensure the feasibility of the solution set. This method generates a high-quality initial solution set that is both economical, environmentally friendly, and flexible within seconds, providing ample search space for subsequent optimization.
[0018] Optionally, in some embodiments, the basic feasible solution is expanded by the parameter perturbation method to obtain multiple differentiated candidate solutions, including: determining the disturbance parameters and disturbance directions based on the real-time target weight coefficients and the basic feasible solution; generating multiple groups of parameter combinations according to the disturbance parameters and disturbance directions; adjusting the weight coefficients and output baselines of the multiple groups of parameter combinations to obtain multiple groups of directional parameter combinations; and independently solving each group of directional parameter combinations as input to obtain differentiated candidate solutions.
[0019] The technical solutions of the above-mentioned embodiments clarify the implementation details of parameter perturbation, and enhance the diversity of the solution set through systematic perturbation parameter generation and directional adjustment. Based on weight sensitivity analysis and unit priority, the perturbation direction is determined (e.g., economic weight ±5%, critical unit output ±3%), and Monte Carlo sampling is used to generate multiple parameter combinations. Physical feasibility adjustments to parameter combinations are achieved through weight normalization, output baseline compensation, and start-stop state synchronization, and a lightweight solver is called to quickly generate candidate solutions. This method uses a controllable perturbation strategy to efficiently explore the target trade-off space within the neighborhood of the basic solution, avoiding the computational waste caused by blind search and providing a standardized process for generating differentiated, high-quality solutions.
[0020] In a second aspect, an embodiment of the present invention provides an electronic device, comprising: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the electronic device to execute the method described in the first aspect or the second aspect, and any possible implementation of the first aspect or the second aspect.
[0021] In a third aspect, the present invention provides a computer-readable storage medium comprising instructions, which, when executed on the electronic device, enables the electronic device to execute the method described in the first aspect or the second aspect, and any possible implementation of the first aspect or the second aspect.
[0022] In a fourth aspect, the present invention provides a computer program product comprising instructions, which, when the computer program product is run on the electronic device, enables the electronic device to execute the method described in the first aspect or the second aspect, and any possible implementation of the first aspect or the second aspect.
[0023] It is understood that the electronic device provided in the second aspect, the storage medium provided in the third aspect, and the computer program product provided in the fourth aspect are all used to execute the method provided by the present invention. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding methods and will not be repeated here.
[0024] The one or more technical solutions provided by the present invention have at least the following technical effects or advantages:
[0025] 1. Dynamic multi-objective weight adaptation breaks through the rigid constraints of traditional optimization models: By constructing a dynamic weight allocation model that integrates physical mechanisms and deep learning, this model addresses the difficulty of adapting fixed weights in traditional multi-objective scheduling to complex operating conditions. The model automatically adjusts the weights for economy, environmental protection, and flexibility based on real-time characteristics (such as compliance with environmental limits). Combining digital twin simulation with staged reinforcement learning, it precisely matches weighting decisions with operational scenarios. For example, in the event of fuel shortages, the weight for economy is automatically reduced to ensure unit safety, while in the event of a surge in grid frequency regulation demand, the weight for flexibility is increased to prioritize response. This adaptability elevates multi-objective optimization from a static trade-off to a dynamic game, significantly improving the real-time performance and scenario-wide generalization capabilities of the scheduling strategy.
[0026] 2. Efficient Pareto solution generation mechanism to ensure multi-objective coordinated optimization: Based on mixed integer programming and multi-threshold relaxation solution technology, a primary objective-oriented optimization and solution convergence verification method is proposed to effectively balance solution quality and computational efficiency. The secondary objective thresholds are dynamically tightened using the ε-constraint method. Combined with non-dominated sorting and uniform screening strategies, a high-density Pareto solution set is generated that covers the three-dimensional trade-off surface of economy, environmental protection, and flexibility. Compared with traditional multi-objective algorithms, this method improves the solution hypervolume (HV) by over 20% and shortens convergence time by 50%. This provides a rich and reliable library of non-dominated solutions for scheduling decisions, avoiding decision imbalances caused by sparse solutions or local optimality.
[0027] 3. Enhanced robustness across the entire chain, enabling reliable industrial-grade applications: Through adversarial training, online transfer learning, and hard filtering of safety constraints, a closed-loop reliability assurance system is constructed from model training to solution generation. A generative adversarial network is introduced to construct complex disturbance samples, such as sensor failures and control interference, enhancing the model's anti-interference capabilities. The online transfer framework enables real-time fine-tuning based on actual scheduling deviations, addressing domain offsets between the digital twin simulation and the physical system. Parameter perturbation expansion and safety constraint verification ensure that candidate solutions are 100% compliant with unit safety and grid stability requirements. This system enables the dispatch system to output compliant solutions even in extreme scenarios such as fuel fluctuations and equipment anomalies, increasing mean time between failures (MTBF) by three times and meeting the stringent requirements for continuous and stable operation of thermal power plants. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present invention, and together with the specification, are used to explain the principles of the present invention. Obviously, the drawings described below are only some embodiments of the present invention, and those skilled in the art can derive other drawings based on these drawings without inventive effort. In the drawings:
[0029] Figure 1 This is a flow chart of a multi-objective economic dispatch optimization method for a thermal power plant according to an embodiment of the present invention;
[0030] Figure 2 This is a schematic diagram of a process for constructing and training a dynamic weight allocation model in an embodiment of the present invention;
[0031] Figure 3 It is a schematic diagram of the architecture of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0032] The terms used in the following embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. As used in the specification of the present invention, the singular expressions "a," "an," "above," "the," and "this" are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used in the present invention refers to any and all possible combinations of one or more of the listed items.
[0033] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying relative importance or implicitly indicating the quantity of the technical features indicated. Thus, a feature designated "first" or "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present invention, unless otherwise specified, "plurality" means two or more.
[0034] It should also be noted that, unless otherwise clearly specified and limited, in the embodiments of the present invention, terms such as "setting" and "connection" should be understood in a broad sense. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, or it can be the internal connection of two components; it can be a wired communication connection or a wireless communication connection. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances. The embodiments of the present invention are described in detail below.
[0035] The embodiment of the present invention provides a multi-objective economic dispatch optimization method for a thermal power plant, such as Figure 1 As shown, the following steps are included:
[0036] Step 201 : constructing an initial data set based on the power grid operation parameters, generator set status parameters and environmental monitoring data of the thermal power plant.
[0037] Specifically, first, three categories of core data are obtained in real time through various sensors, monitoring systems and data interfaces deployed in thermal power plants and power grids.
[0038] Grid operating parameters include real-time grid load power, peak and valley period division, electricity price fluctuation curve, real-time output of renewable energy (wind power, photovoltaic) and grid dispatching instructions (such as ramp rate requirements and backup capacity requirements). These data are synchronized to the dispatching center in real time through the power dispatching data network (such as the SCADA system).
[0039] The status parameters of the generator sets include the active / reactive output of each unit, main steam temperature / pressure, turbine vibration amplitude, boiler efficiency, unit start / stop status, cumulative operating time, temperature field and stress monitoring data of key components (such as superheaters and turbine blades), and fuel inventory (coal calorific value, sulfur content, and blending ratio), etc., which are mainly collected through the unit DCS (distributed control system) and equipment Internet of Things sensors.
[0040] Environmental monitoring data includes real-time monitoring values of SOx, NOx, and CO2 concentrations at the chimney outlet, the operational status and efficiency of environmental protection equipment (such as denitrification equipment and dust removal equipment), wastewater / waste residue emission indicators, etc., which are derived from online pollutant monitors and environmental protection ledger systems.
[0041] After data collection, preprocessing is required to ensure quality. First, data cleaning is performed to remove outliers (such as jump data caused by sensor failures), and high-frequency data is smoothed using sliding average or Kalman filtering algorithms. Second, multi-source heterogeneous data is time-aligned and formatted uniformly. For example, second- and minute-level data from different systems are unified into a 5-minute resolution time series. Finally, data normalization (such as Z-score standardization) is performed to eliminate dimensional differences and construct a structured initial dataset containing timestamps, unit numbers, parameter categories, and values. This dataset not only covers the current real-time status but also includes historical data from the previous hour to the previous day to capture short-term trends (such as load fluctuation inertia and equipment thermal inertia), providing a comprehensive data foundation for subsequent feature parameter extraction.
[0042] Step 202 : determining characteristic parameters of the current operation scheduling demand based on the initial data set, wherein the characteristic parameters include load demand fluctuation rate, unit adjustable margin, and environmental protection limit compliance rate.
[0043] Among them, load demand fluctuation rate refers to the percentage change of the real-time load of the power grid relative to the benchmark load (such as the predicted value or historical average), which is used to quantify the urgency of the imbalance between electricity supply and demand. It is usually calculated by dividing the load change per unit time by the maximum regulation rate of the unit.
[0044] The unit's adjustable margin represents the range of power generation capacity that can be flexibly adjusted at the current output of the generator set. The calculation formula is (the difference between the current output and the maximum / minimum technical output) divided by the unit's rated capacity, reflecting the unit's potential to participate in frequency and peak regulation.
[0045] The environmental protection limit compliance rate is the percentage ratio of the real-time monitored pollutant emission concentration (such as SO2, NOx) to the environmental protection regulatory limit. If the compliance rate is lower than 100%, it means that the emission exceeds the standard and the environmental protection weight enhancement mechanism needs to be triggered immediately.
[0046] Specifically, the calculation of load demand fluctuation requires combining real-time grid load with historical data. First, the load forecast and actual values within the current dispatch period (e.g., 15 minutes) are obtained. The ratio of the absolute deviation to the rated load is calculated, or the standard deviation of the load variation over the past 30 minutes is calculated using a sliding window to reflect the severity of the load fluctuation. If the grid is experiencing peak load periods and renewable energy output plummets, the load demand fluctuation rate will increase significantly, indicating an urgent need for rapid response from thermal power units and the need to increase the weight of flexibility in dispatch.
[0047] The assessment of unit turnaround margin requires a comprehensive consideration of both individual unit and plant-wide constraints. For a single unit, turnaround margin = min (rated output - current output, current output - minimum technical output) / rated output, reflecting the unit's capacity for adjustment under the current load. Ramping rate constraints are also considered: the maximum adjustable capacity within the next 15 minutes = current output + ramp rate × time interval. The plant-wide turnaround margin is the weighted sum of each unit's turnaround margin (weighted by the percentage of the unit's rated capacity). If the plant-wide turnaround margin is less than 10%, the overall turnaround capacity of the units is nearing its limit, and scheduling must prioritize avoiding the risk of equipment overload caused by excessive peak load regulation.
[0048] The environmental limit compliance rate is calculated based on real-time emissions data and environmental standards. For NOx, the compliance rate = real-time emission concentration / environmental limit (e.g., the ultra-low emission limit is 50 mg / m³). If the compliance rate exceeds 90% (i.e., the concentration is approaching the limit), it indicates tightening environmental constraints and the need to increase the weight of environmental indicators. Emissions should be reduced by adjusting combustion parameters or implementing deep purification equipment. This parameter can also be dynamically adjusted based on the operating status of environmental protection equipment (e.g., denitrification catalyst activity and ammonia reserves). For example, if the denitrification system fails, the compliance rate calculation needs to include the response capacity of backup environmental protection facilities.
[0049] Step 203 : Based on the characteristic parameters, the real-time target weight coefficients of the economic index, the environmental index and the scheduling flexibility index are determined through a preset dynamic weight allocation model.
[0050] The dynamic weight allocation model is a rule-driven algorithm that adaptively adjusts target weights based on multi-dimensional scenario characteristics. Taking the real-time operating status of the thermal power plant and external environmental parameters as input, the model dynamically calculates weight coefficients for economic, environmental, and flexibility indicators through a predefined rule base for resolving conflicting objectives and a fuzzy logic inference mechanism. Its core lies in constructing a mapping table between target weights and characteristic parameters. For example, when the environmental limit compliance rate falls below a preset threshold, an environmental weight increase rule is triggered; when grid frequency regulation demand surges, the flexibility weight is automatically increased.
[0051] The dynamic weight allocation model utilizes a hierarchical decision-making architecture. The first layer uses thresholds to screen key constraints (e.g., whether environmental limits are on the verge of being exceeded). The second layer calculates the weights of non-critical objectives using linear interpolation or a lookup table. Finally, normalization is performed to ensure that all weight coefficients sum to 1. This model avoids complex mathematical optimization and relies on expert experience and historical data statistics, completing weighting decisions within 10 milliseconds.
[0052] The specific implementation process of this step is as follows: First, three characteristic parameters are extracted from real-time data: load demand fluctuation rate (the ratio of the current load change rate to the unit's maximum ramp rate), unit adjustable margin (the difference between the current output and the maximum adjustable output), and environmental protection limit compliance rate (the ratio of actual emission concentration to the limit). These characteristic parameters are then input into a dynamic weight allocation model: 1) If the environmental protection limit compliance rate is greater than 95%, the environmental weighting is set to a base value of 0.2; otherwise, the weighting increases by 0.05 for every 1% decrease in load demand fluctuation; 2) When the load demand fluctuation rate is greater than 5%, the flexibility weighting increases linearly from 0.1 to 0.3, an increase of twice the excess fluctuation; 3) When the unit adjustable margin is less than 20%, the economic weighting is reduced by 0.1 to maintain regulatory margin. Finally, the three benchmark weights are normalized. For example, the initial benchmark values are 0.6 for economy, 0.2 for environmental protection, and 0.2 for flexibility. After adjustment, they become 0.55 for economy, 0.25 for environmental protection, and 0.2 for flexibility. The sum is then verified to be 1 and output. This process is implemented using precompiled rule scripts, with weight updates executed every 15 seconds to ensure that the scheduling strategy adapts to changes in the operating scenario in real time.
[0053] In this embodiment, the construction and training process of the dynamic weight allocation model is divided into four stages:
[0054] First, an initial rule base was constructed based on historical scheduling data and expert experience. For example, the optimal weight distribution under different load demands and environmental protection limit scenarios in the past year was collected, and regular parameters such as the average economic weight during the peak-shaving period was 0.55 and the environmental protection weight was increased to 0.4 when the emission exceeded the standard warning was issued were statistically analyzed.
[0055] Secondly, a fuzzy logic system is used to design weight adjustment rules, and the characteristic parameters are divided into three fuzzy sets of "low, medium, and high". A fuzzy inference rule table is defined (for example, when the environmental limit compliance rate is "low" and the load fluctuation rate is "high", the adjustment action of "environmental protection weight +0.15, flexibility weight +0.05, and economic weight -0.2" is triggered). The matching degree between parameters and rules is quantified through the membership function.
[0056] Next, offline historical data is used to optimize the rule parameters, and the grid search method is used to traverse the candidate values of each threshold in the rule base (such as the environmental protection compliance rate trigger threshold is tested from 90% to 98% with a step size of 2%). The economic and environmental comprehensive cost of the actual scheduling results is used as the evaluation indicator to select the optimal parameter combination.
[0057] Finally, an online adaptive calibration module is deployed to compare the deviation between the weights output by the model and the scheduling effect in real time. When the target achievement degree in five consecutive scheduling plans is lower than the expected value, the rule base iteration is automatically triggered (such as relaxing the trigger sensitivity of the environmental protection weight increase). At the same time, the latest data is reloaded every month to batch update the rule base to ensure that the model continues to adapt to changes in the power plant operating environment.
[0058] Step 204 : generating an initial candidate scheduling solution set that satisfies the safety constraints of the generator set according to the real-time target weight coefficient.
[0059] First, a weighted single objective function is constructed based on real-time weight coefficients: for example, total cost = economic weight × fuel cost + environmental weight × emission penalty + flexibility weight × frequency regulation cost. Subsequently, a greedy algorithm or relaxed linear programming is used to quickly generate a basic feasible solution: Based on the current unit status, units with low coal consumption and excellent emissions performance are prioritized for output increase while also satisfying power balance constraints. For example, if a unit's output limit is 300 MW and its current output is 200 MW, the output is increased in 10 MW steps until the ramp rate limit (e.g., 5 MW / min) is reached. During this process, hard safety constraints (e.g., minimum unit start / stop times and vibration zone entry restrictions) are verified in real time through logical judgment, and violating solutions are directly eliminated. To increase the diversity of the solution set, the weight coefficients can be perturbed within a small range (e.g., ±5%) to generate multiple sets of differentiated objective functions, which are then calculated using the solver separately. For example, the original weights are 0.6 for economy, 0.3 for environmental protection, and 0.1 for flexibility. After perturbation, three sets of weights are generated (0.63 / 0.28 / 0.09, 0.57 / 0.31 / 0.12, and 0.60 / 0.30 / 0.10), each corresponding to a candidate solution. Finally, all feasible solutions are collected through parallel computing or iterative search, forming an initial candidate set of 10-20 solutions for subsequent optimization. This process typically completes within 5 seconds, meeting the timeliness requirements of real-time scheduling.
[0060] Step 205 : performing main target-oriented optimization on the initial candidate scheduling solution set to obtain a Pareto optimal solution set that meets a preset convergence condition.
[0061] The implementation of this step is based on a hierarchical optimization strategy. First, the target with the highest weight from the weight coefficients output by the dynamic weight allocation model is selected as the main optimization target (for example, the economic weight is 0.6), and the remaining targets are used as constraints.
[0062] In specific implementation, a mixed integer linear programming (MILP) model centered around the primary objective is constructed: 1) The primary objective function is an economic indicator (e.g., minimizing fuel costs), including parameters such as coal consumption rate and startup and shutdown costs. 2) Secondary objectives are converted into constraints. For example, the environmental objective requires total emissions to not exceed 110% of the real-time limit, and the flexibility objective requires the total ramp rate of the units to be no less than 90% of the grid frequency regulation demand. 3) Hard safety constraints include upper and lower limits on unit output, minimum startup and shutdown times, and vibration zone entry prohibition rules. A branch-and-bound approach is used for precise solution, dynamically adjusting constraint boundaries during the calculation process to explore the solution space. For example, if the optimal economic solution results in an environmental performance violation, the emission constraint is gradually tightened (e.g., from 110% to 105%) and the solution is repeated until all constraints are satisfied. To generate the Pareto frontier, the relaxation range of the secondary objective constraints is iteratively adjusted using the ε-constraint method. After the initial solution, the optimal environmental performance value is recorded, relaxed by 5% as a new constraint, and the economic objective is solved again. This process is repeated until all non-inferior solutions are found. The convergence condition is set as the improvement of the objective function is less than 0.5% in three consecutive iterations or the maximum number of iterations (e.g., 50) is reached, and the final output is a Pareto optimal solution set containing 10 to 20 groups of non-dominated solutions.
[0063] Step 206: Generate a scheduling recommendation plan based on the Pareto optimal solution set.
[0064] Specifically, a multi-criteria decision analysis can be used to select the optimal solution from the Pareto solution set as the recommended scheduling solution. The recommended scheduling solution is an optimal unit output allocation and start-stop combination strategy that takes into account economy, environmental protection, flexibility, and no hidden risks.
[0065] In this step, we first construct an evaluation index system, which includes four dimensions: economy (unit power supply cost), environmental protection (emission intensity), flexibility (frequency regulation response time), and equipment health (turbine stress coefficient). Each dimension is assigned a standardized scoring weight (e.g., economy 40%, environmental protection 30%, flexibility 20%, and health 10%).
[0066] Quantify the indicators of each solution in the Pareto solution set:
[0067] 1) Economic efficiency score = (historical lowest cost - current cost) / (historical highest cost - historical lowest cost) × 100;
[0068] 2) Environmental performance score = 1-(actual emissions / limit) × 100 (0 points if exceeding the limit);
[0069] 3) Flexibility score = (theoretical maximum frequency modulation rate - actual rate) / theoretical value × 100;
[0070] 4) Equipment health score = 1-(stress accumulation / safety threshold) × 100.
[0071] Subsequently, TOPSIS (top-up ideal solution ranking method) was used to calculate the closeness of each solution to the ideal solution: the positive ideal solution (highest score in each dimension) and the negative ideal solution (lowest score in each dimension) of each indicator were determined, the Euclidean distance of each solution to the positive and negative ideal solutions was calculated, and finally the solutions were sorted from high to low according to the closeness.
[0072] For example, a solution scored 90 for economy, 85 for environmental performance, 70 for flexibility, and 80 for health, with a closeness of 0.82, ranking first. To enhance decision reliability, a digital twin verification step was added: the top three solutions were input into the power plant's digital twin system, which simulated operating conditions over the next two hours to detect any implicit constraint violations (such as delayed NOx concentration exceeding the standard). If the simulation results showed that a solution would exceed the emission standard after 30 minutes, the solution was automatically downgraded and the suboptimal solution was selected. The solution with the highest overall score and verified by simulation was ultimately output as the recommended scheduling solution, and a list of alternative solutions was provided for manual decision-making.
[0073] This embodiment significantly improves the comprehensive efficiency and real-time decision-making ability of economic dispatch of thermal power plants through the collaborative design of a dynamic weight allocation model and a multi-stage optimization strategy, and effectively balances the conflicts between multi-objective optimizations in the economic dispatch scheme.
[0074] This embodiment constructs multi-dimensional characteristic parameters based on real-time collected power grid, unit and environmental data to accurately depict the core contradictions of the current scheduling scenario. For example, it quantifies the power grid stability pressure through load demand volatility, evaluates the peak-shaving capacity through the adjustable margin of the unit, and reflects the intensity of emission constraints through the environmental protection limit compliance rate. It provides high-information-density input for dynamic weight allocation and solves the problem of poor scenario adaptability caused by traditional methods relying on manual experience or static rules. Secondly, the dynamic weight allocation model adopts a hybrid architecture that integrates reinforcement learning and physical mechanisms. It can automatically adjust the weight coefficients of economic, environmental and flexibility goals according to real-time scenario characteristics. For example, when the power grid frequency regulation instructions are frequently issued, the flexibility weight is increased to more than 70%, and in the stage of strengthening environmental monitoring, the environmental protection weight is dynamically increased to a dominant position. This breaks through the limitations of the traditional fixed weight model in the target conflict scenario, which has a single solution set and cannot cover multi-dimensional trade-offs.
[0075] Furthermore, a two-stage optimization strategy (initial candidate solution generation and main target-oriented optimization) is used to significantly reduce computational complexity. The initial candidate solution set is generated using heuristic rules and parameter perturbation methods, which can provide 20 to 30 feasible solutions that meet safety constraints within 3 seconds, more than 5 times faster than traditional multi-objective optimization algorithms. The main target-oriented optimization stage focuses on the target with the highest weight for mixed integer programming solution, while converting suboptimal targets into constraints, which not only ensures the diversity of the Pareto front, but also controls the overall optimization time to within 10 seconds, meeting the time sensitivity requirements of real-time scheduling.
[0076] In addition, the generation process of the Pareto optimal solution set introduces an equipment life loss prediction model and a digital twin verification mechanism, which can avoid the hidden costs caused by traditional methods that ignore long-term equipment health. For example, the accumulated stress of the turbine is used as a constraint item, which can reduce the life loss caused by frequent start-up and shutdown by more than 40%.
[0077] While ensuring real-time computing, this embodiment can reduce overall dispatch costs by 15% to 22%, reduce environmental violations by 90%, increase unit frequency regulation response speed by 35%, and improve solution diversity indicators (such as Hypervolume) by over 50%. This provides dispatchers with multi-dimensional decision support that is both economical, environmentally friendly, and reliable, effectively resolving the three core challenges and contradictions of traditional technologies: rigid optimization objectives, low computing efficiency, and uncontrollable long-term equipment loss.
[0078] In some embodiments, a method for constructing and training a dynamic weight distribution model is also provided, such as Figure 2 As shown, the specific process includes the following:
[0079] S301, construct a multimodal hybrid model structure.
[0080] The multimodal hybrid model structure includes an input layer for integrating grid, generator, and environmental data, a constraint encoding layer for embedding physical mechanism equations, and a hierarchical strategy network based on a multi-head attention mechanism. Physical mechanism equations are mathematical expressions based on physical laws such as thermodynamics and fluid mechanics. They quantitatively describe the inherent laws of energy conversion, equipment efficiency, and pollutant generation during the operation of thermal power plants. Examples include the quadratic function relationship between boiler combustion efficiency and load, and the differential equation for the relationship between turbine heat rate and steam parameters.
[0081] Specifically, the input layer is designed as a multi-channel data fusion module. The power grid data (frequency deviation, frequency regulation demand) is standardized using a sliding window, the generator set data (coal calorific value, output curve) is normalized using the Z-score, and the environmental data (emission concentration, meteorological parameters) are time series aligned.
[0082] The constraint coding layer integrates physical mechanism equations such as the thermodynamic equations of thermal power units. For example, it converts the boiler efficiency-load relationship formula (η=aP²+bP+c) into an inequality constraint matrix and embeds it into the network forward propagation process through automatic differentiation technology.
[0083] The hierarchical strategy network adopts a multi-head attention mechanism, in which the economic attention head calculates the correlation weight between fuel cost and load demand, the environmental attention head focuses on the temporal dependency between emission concentration and environmental protection equipment status, and the flexibility attention head analyzes the dynamic matching between frequency modulation instructions and unit ramp rate. The outputs of each attention head are fused into the final weight vector after residual connection and layer normalization.
[0084] S302 , generating a simulation data set including multiple operating conditions through simulation technology, where the simulation data set covers fuel shortage, environmental protection equipment failure, and power grid frequency disturbance scenarios.
[0085] Specifically, first, a high-fidelity simulation model is constructed based on the actual power plant design parameters (such as boiler capacity, turbine model, and environmental protection equipment specifications). The core equipment of the unit adopts a mechanism model (such as the boiler combustion process is modeled by the mass-energy conservation partial differential equation, and the denitrification reactor simulates the dynamic relationship between ammonia injection and NOx conversion rate based on the chemical reaction kinetic equation). The grid interaction module integrates the IEEE standard node model to simulate the frequency-power coupling characteristics.
[0086] Then, a variety of operating conditions were generated through the parameter perturbation method: 1) The fuel shortage scenario was achieved by adjusting the coal quality parameters. For example, the received base low calorific value was randomly reduced from the design value of 20MJ / kg to 15~18MJ / kg, while the ash content was increased (from 15% to 25%), and the fuel delivery delay caused by coal blockage in the coal feeder was simulated (random perturbation of 0~30 seconds); 2) Environmental protection equipment failure scenarios included abnormal ammonia escape rate of the selective catalytic reduction (SCR) system (suddenly increased from the normal 3ppm to 15ppm), dust removal The short circuit of the plate causes the dust removal efficiency to drop (from 99.9% to 85%), which is achieved by modifying the equipment state variable parameters and injecting a step signal or a ramp failure curve; 3) The grid frequency disturbance scenario uses a random signal generator to simulate the grid frequency fluctuation, superimposing normal distribution noise (standard deviation 0.1Hz) and periodic disturbance (amplitude ±0.5Hz, random period 10-60 seconds) on the reference frequency of 50Hz, and at the same time associating the frequency modulation command issuance logic (when the frequency deviation exceeds 0.2Hz, the AGC command is triggered).
[0087] All scenarios are combined, including, for example, the simultaneous triggering of a compound abnormal operating condition characterized by "decreased coal calorific value, SCR ammonia pump failure, and a sudden drop in grid frequency of 0.3 Hz." During data generation, the simulation model runs continuously at a sampling interval of one second, recording over 300 parameters, including grid frequency, unit output, emission concentration, and equipment status. Each scenario lasts for at least six hours to capture both transient and steady-state characteristics. Ultimately, a simulation dataset containing over 100,000 samples is generated, each annotated with the corresponding operating condition type (e.g., fuel shortage level, equipment fault code), the true value of the real-time target weight (pre-calculated through offline multi-objective optimization), and key constraint violation flags for subsequent model training and validation.
[0088] S303: Label the Pareto optimal weight vector of each scene in the simulation data set to form a training label set.
[0089] Among them, the training label set refers to the set of optimal target weight vectors that match each operating scenario in the simulation data set, which is pre-calculated by the multi-objective optimization algorithm. It contains the Pareto optimal solution of the economic, environmental and flexibility weights under different working conditions, and is used to supervise the model learning the mapping relationship between scenario features and dynamic weights.
[0090] Specifically, for each simulation scenario, the improved NSGA-III algorithm is used for multi-objective optimization. The objective function is to solve the three-dimensional Pareto front of economy (fuel cost), environmental performance (total emissions), and flexibility (frequency regulation response error). Constraints include unit output limits, start-up and shutdown times, and a hard upper limit on instantaneous NOx concentration.
[0091] After the optimization is completed, the weight vectors of all non-dominated solutions on the Pareto front are extracted, and their distribution consistency is evaluated by KL divergence to eliminate abnormal solutions that deviate from the main cluster.
[0092] The final labeled data contains scenario feature vectors (load demand, coal quality parameters, equipment status) and the corresponding Pareto optimal weight set, forming a large number of training label sets.
[0093] S304 , using the training label set and the physical mechanism equation as network constraint rules, pre-training the initial parameters of the hierarchical strategy network to obtain a pre-trained model.
[0094] Specifically, a composite loss function is first constructed, such as L_total=α·L_label+β·L_phy, where L_label is the cosine similarity loss between the predicted weight and the Pareto label, calculated as 1-(y_pred·y_true) / (||y_pred||·||y_true||), which is used to drive the network to approach the optimal weight distribution; L_phy is the physical mechanism constraint loss, which includes three types of penalty terms: 1) Based on the constraint of the boiler efficiency-load curve, the mean square error between the theoretical coal consumption rate and the actual value corresponding to the network output economic weight is calculated; 2) Environmental weight constraint, the predicted environmental weight is substituted into the SCR denitrification efficiency equation (η=k·C_nox·T_reactor) to verify whether the emission concentration meets the hard requirement of η≥85%, otherwise a gradient penalty is imposed; 3) Flexibility weight constraint, check whether the unit frequency regulation rate meets ΔP / Δt≥(weight×theoretical maximum value), if violated, the square of the slack variable is calculated as the loss term.
[0095] Training utilizes projected gradient descent, forcing weight coefficients to be projected into the feasible region after each parameter update (e.g., Σweight = 1, all weights ≥ 0.1). During data preprocessing, input features undergo sliding window normalization (window length 60 seconds), and physical equation parameters (e.g., turbine heat rate and NOx generation coefficient) are encoded as 32-dimensional vectors and fed into the constraint encoding layer. Pretraining is performed for 200 epochs with a batch size of 128. The initial learning rate is 0.01, decaying to 1 / 10 every 50 epochs. The ratio of α to β is dynamically adjusted (initial α:β = 1:1, adjusted after each epoch based on the validation set loss; if L_label decreases slowly, α is increased by 20%).
[0096] The obtained pre-trained model must meet the following requirements when passing the validation set test: cosine similarity ≥ 0.92, physical constraint violation rate < 3%, and weight normalization error < 0.001.
[0097] S305 , respectively, performs phased reinforcement learning training on the hierarchical strategy network of the pre-trained model in single-unit steady-state, multi-unit coupling, and random failure scenarios to obtain a dynamic weight allocation model.
[0098] This step can be implemented in the following stages:
[0099] During the single-unit steady-state training phase, with a fixed grid load (e.g., 500 MW) and no equipment faults, environmental parameters are initialized (coal calorific value 20 MJ / kg, SCR denitrification efficiency 90%), and training is performed using the Deep Deterministic Policy Gradient (DDPG) algorithm. The action space consists of continuous adjustments (Δw∈[-0.1,+0.1]) to the weights of economy, environmental protection, and flexibility. The state space contains 20-dimensional features such as the unit's current output, coal consumption rate, and NOx concentration. The reward function is designed as R = 0.7 × (baseline fuel cost - actual cost) / baseline cost + 0.2 × (1 - emission limit violation flag) + 0.1 × frequency modulation response rate. The exploration noise is set to an O(n) process (θ = 0.15, σ = 0.2). Training is performed for 500,000 steps until the average reward stabilizes above 0.85.
[0100] During the multi-unit coupling training phase, the coordinated scheduling of six units was expanded, and the state space increased to 120 dimensions (20 parameters per unit). The action space consisted of a global weight vector plus the output allocation ratio of each unit. A multi-agent proximal policy optimization (MAPPO) framework was employed, with a central critic network evaluating the global reward, and each unit's actor network independently outputting local actions. The reward function was upgraded to R = 0.5 × total economic benefit + 0.3 × environmental compliance rate + 0.2 × Σ (the inverse of the frequency synchronization error of unit i). A coupling penalty term was also introduced: if the output deviation between any two units exceeded 20% for 5 minutes, a reward of 0.1 was deducted. Training employed a curriculum learning strategy, gradually increasing the number of units (from 2 to 6) and linearly decaying the exploration rate from 0.3 to 0.1, for a total of 800,000 training steps.
[0101] During the random fault training phase, dynamic environmental perturbations are introduced. Each training round has a 30% probability of triggering one of the following events: 1) coal quality mutation (calorific value step change within ±15%); 2) SCR ammonia pump failure (denitrification efficiency drops to 50% within 10 seconds); or 3) grid frequency perturbation (superimposed ±0.5Hz sinusoidal fluctuations). Using an adversarial reinforcement learning framework, an additional adversarial network is deployed to generate deceptive environmental monitoring signals (such as false NOx compliance data). The policy network is required to identify anomalies and adjust weights within 0.5 seconds. The reward function reinforces a long-term penalty term: R = R_base - 0.5 × Σ (predicted equipment life loss rate for the next hour). Training utilizes prioritized experience replay, prioritizing sampling of fault scenario data. The exploration rate is reset to 0.2, and training is conducted for 1.2 million steps until the average reward in the fault scenario is ≥ 0.75. After each stage of training, the model performance is tested on an independent validation set (such as 10,000 sets of unseen disturbance scenarios), requiring the weight decision delay to be less than 50ms and the Pareto solution coverage to be more than 95%. Finally, through three stages of progressive training, the model is made multi-scenario robust to obtain a dynamic weight allocation model.
[0102] S306: Introduce a generative adversarial network to construct a perturbation sample set.
[0103] Specifically, a conditional generative adversarial network (CGAN) is constructed. The generator receives the real environmental status (such as unit output and emission concentration) and a random noise vector as input, and outputs forged samples with disturbance characteristics, including: 1) sensor failure disturbance (such as coal calorific value data drift of ±20%); 2) control signal interference (such as false AGC frequency modulation instructions superimposed on Gaussian noise with an amplitude of ±5%); 3) covert environmental data tampering (injecting a slow rising trend when the NOx concentration meets the standard to make it appear to exceed the standard).
[0104] The discriminator uses a hybrid convolutional neural network (CNN)-long short-term memory (LSTM) architecture, determining whether input data is a generated sample through temporal pattern recognition. During adversarial training, the generator's objective function minimizes the policy network's weight fluctuation (Δw < 0.05) on forged samples while maximizing the discriminator's misclassification probability. The discriminator maximizes the accuracy of classifying genuine and fake samples. After each round of training, perturbed samples that cause the policy network's weights to shift by more than 10% are retained, resulting in a perturbed data set containing a large number of highly offensive samples.
[0105] S307: Perform robustness enhancement training on the dynamic weight allocation model according to the perturbation sample set.
[0106] Specifically, within a reinforcement learning environment, a dynamic weight allocation model (the agent) interacts with a simulated environment that is injected with perturbation samples. During each interaction, the environmental state is sampled with a 60% probability from a real dataset and a 40% probability from an adversarially generated perturbation dataset. For example: 1) The sensor perturbation sample modifies the actual unit output of 300MW to a random value between 280 and 320MW; 2) The control signal perturbation sample superimposes a sinusoidal perturbation with an amplitude of ±5% and a frequency of 0.1 to 1Hz on the actual AGC command; and 3) The environmental data perturbation sample causes the NOx concentration to fluctuate around the limit (95% to 105%).
[0107] The agent receives the perturbed state observations, outputs economic, environmental, and flexibility weight vectors, and obtains feedback based on the improved robustness reward function: R = 0.6 × R_base (base reward, including economic benefits and environmental penalties) + 0.3 × R_stability (stability reward, calculating the cosine similarity between the current weight and the average of the previous 10 steps to suppress mutations) - 0.1 × R_attack (attack detection penalty, a 0.1x penalty is imposed when the discriminator network determines that the current state is an adversarial example).
[0108] Training uses the TD3 reinforcement learning algorithm. The actor network is updated at each step, while the critic network uses dual Q-learning to reduce overestimation bias. The target network update coefficient τ = 0.005. To improve anti-interference capabilities, a state recovery mechanism is designed: when the variance of the critic's predicted Q value exceeds a threshold for five consecutive steps, the historical state is rolled back (replacing the last 10 steps with the unperturbed version), forcing the agent to make a new decision.
[0109] During training, robustness stress testing is conducted every 50,000 steps. The model is evaluated on an independent test set (containing 2,000 extreme perturbation scenarios, such as a 30% drop in coal calorific value and a 40% drop in SCR efficiency). The objective function volatility (maximum - minimum) / mean) is required to be less than 10%, and the confidence level of weight decisions (calculated using softmax probabilistic entropy) is required to be greater than 0.7. The final model must pass 300 adversarial attack tests, maintaining a weight adjustment amplitude Δw less than 0.15 and a Pareto solution coverage of ≥85% under perturbation.
[0110] S308, through the online transfer learning framework, dynamically fine-tune the output layer parameters of the dynamic weight allocation model based on the deviation between the actual scheduling results and the simulation prediction.
[0111] Specifically, during the deployment phase, actual scheduling data (such as unit output, fuel consumption, and emission concentration per minute) can be collected in real time and compared with the model prediction value to calculate multi-dimensional deviation indicators: economic deviation ΔC = (actual cost - predicted cost) / predicted cost, environmental deviation ΔE = MAX (0, actual emissions - predicted emissions) / emission limit, flexibility deviation ΔF = 1 - (actual frequency regulation response rate / predicted value).
[0112] When ΔC>10% or ΔE>5% in 5 consecutive sets of data, the parameter update process is triggered:
[0113] 1) Freeze the parameters of the feature extraction layer and attention mechanism layer of the policy network, and only open the fully connected weight matrix of the output layer (dimension 128×3) for fine-tuning;
[0114] 2) Construct an online training set and extract the 500 sets of samples with the largest deviations within the past 24 hours from the ring buffer. Each sample contains the environment state (a normalized 120-dimensional vector), the original model weight output, and the actual scheduling result.
[0115] 3) Design a transfer loss function, such as:
[0116] L = 0.5 × MSE (prediction weight, correction weight) + 0.3 × |ΔC + ΔE| + 0.2 × L2 regularization term;
[0117] MSE (prediction weight, correction weight) refers to the mean squared error between the target weight coefficient output by the calculation model and the target weight adjusted by the expert rule. It is used to constrain the direction of weight adjustment during the model fine-tuning phase so that the prediction results approach the ideal value after domain knowledge correction. Corrected weight = original weight + K·Δ (Δ is the increment given by the expert rule library, and K=0.1 is the learning rate decay coefficient).
[0118] 4) An online momentum SGD optimizer is used with a learning rate of 0.0001 and a momentum of 0.9. The model is iterated 50 times in small batches (16 samples / batch). The output weights are projected and normalized after each batch update (Σwi = 1 and w_i ≥ 0.1).
[0119] During the fine-tuning process, gradient truncation (threshold 0.01) was introduced to prevent overtuning, and validation was achieved through A / B testing: the updated model was piloted on 10% of real-time traffic. If the overall deviation decreased by more than 15% within two hours, full deployment was initiated. A rollback mechanism was also established: if ΔE > 10% was observed three consecutive times after fine-tuning, the model would automatically revert to the previous stable version, triggering expert intervention and diagnostics.
[0120] In some embodiments, step 204 may specifically include the following steps:
[0121] S2041, constructing a multi-objective weighted aggregation function based on the real-time target weight coefficients.
[0122] Specifically, a weighted aggregate objective function can be constructed based on real-time objective weight coefficients (e.g., 0.6 for economy, 0.3 for environmental performance, and 0.1 for flexibility): Total cost = Economic weight × Fuel cost (Σ Coal consumption rate × Output) + Environmental weight × Emission penalty (Σ Exceeding emission standards × Penalty coefficient) + Flexibility weight × Frequency modulation performance loss (1 / Frequency modulation response rate). Fuel cost and emission penalty must be normalized to the [0, 1] range to avoid imbalance in the objective function due to dimensional differences. For example, the fuel cost normalization formula is (Current value - Historical minimum) / (Historical maximum - Historical minimum). The penalty coefficient is dynamically adjusted based on real-time environmental policies (e.g., a penalty of 200 yuan per minute for each milligram of NOx exceeding the standard).
[0123] S2042: Based on the multi-objective weighted aggregation function, a fast heuristic algorithm is called to generate a basic feasible solution.
[0124] Specifically, a greedy algorithm or an improved genetic algorithm can be used to quickly generate a basic solution.
[0125] Taking the greedy algorithm as an example: 1) Sort the unit list from low to high by coal consumption rate, and prioritize starting the unit with the lowest coal consumption rate to its technical output lower limit; 2) Gradually increase the output according to load demand, and each time select the unit with the smallest marginal cost increment (Δcost / Δoutput) until load balance is achieved; 3) Introduce a relaxation mechanism, allowing the ramp rate constraint to be temporarily relaxed by 10% (for example, temporarily increasing the unit's maximum ramp rate from 5MW / min to 5.5MW / min) to accelerate the search.
[0126] The genetic algorithm uses binary coding (unit start and stop status) + real number coding (output value). The initial population is generated based on random perturbations of the current operating status. The fitness function is the weighted total cost. After 50 generations of iteration, the top 10 individuals are selected as the basic solution.
[0127] S2043, expanding the basic feasible solution through the parameter perturbation method to obtain multiple differentiated candidate solutions.
[0128] Specifically, the parameter perturbation method generates differentiated candidate solutions in the neighborhood space of the basic feasible solution through a controllable random perturbation mechanism. The specific implementation method is as follows:
[0129] Unit output perturbation: Randomly select one to three non-critical units from the base solution (e.g., units with an output ratio of less than 30% or in continuous operation) and apply a random perturbation of ±3% to 5% of the rated capacity to their current output. For example, if a unit has a rated capacity of 300MW and a current output of 180MW (60%), the perturbation amplitude is ±5% (i.e., ±15MW). This generates a new output of 185MW or 175MW, while ensuring that the adjusted output does not exceed the limit (e.g., not less than the minimum technical output of 120MW). If the perturbation causes the output to enter the vibration forbidden zone (e.g., the 150-160MW range), it is automatically shifted to the nearest safe point (e.g., 161MW).
[0130] Perturbation of target weight coefficients: A small random offset is applied to the real-time weight vector to generate multiple sets of perturbation weights. For example, if the original weights are 0.6 for economy, 0.3 for environmental protection, and 0.1 for flexibility, the perturbation adjusts them within a ±5% range: the economy weight is randomly increased by 0.03 (to 0.63) or decreased by 0.03 (to 0.57), the environmental weight is adjusted in the opposite direction (to 0.27 or 0.33), and the flexibility weight remains unchanged at 0.1. Finally, the perturbations are normalized to sum to 1. Each set of perturbation weights recalculates the weighted objective function and generates a new solution. For example, weights (0.63, 0.27, 0.1) may favor a more economically aggressive scheduling solution.
[0131] Start / Stop State Perturbation: The start / stop state of any standby unit that meets the minimum downtime requirement (e.g., downtime ≥ 4 hours) in the base solution is reversed. For example, if a unit is down and its downtime meets the requirement, an attempt is made to start it and allocate its minimum technical output (e.g., 120 MW), while simultaneously shutting down another less efficient unit to maintain power balance. Only one unit's state is allowed to change per perturbation, generating complementary start / stop combinations.
[0132] Time series perturbations: To address frequency demand fluctuations, we apply time-shift perturbations to the output plan in the base solution. For example, we can shift the output curve of a particular unit in time period t forward or backward by 5 minutes and recheck the ramp rate constraint (e.g., 5 MW / min). This type of perturbation can explore the flexibility potential at different time granularities.
[0133] Through this multi-dimensional perturbation strategy, each basic feasible solution can generate 5–10 or even more candidate solutions. For example, basic solution A generates A1 (output + 5% of unit X), A2 (weights 0.63 / 0.27 / 0.1), A3 (activating standby unit Y), and so on. All perturbed solutions must pass through a fast verification module (such as precalculating the feasible region boundaries) to filter out obvious violations. Ultimately, a differentiated set of candidate solutions is formed, covering different objective trade-offs and ensuring a uniform distribution of solutions across the target space.
[0134] In some embodiments, this step S2043 may further specifically include:
[0135] (1) Based on the real-time target weight coefficient and the basic feasible solution, the disturbance parameters and disturbance direction are determined.
[0136] Specifically, the main perturbation parameters (economic weight, key unit output) and perturbation direction can be selected based on the real-time target weight coefficient and the characteristics of the basic feasible solution (e.g., the economic weight accounts for more than 60%). For example, if the economic weight is 0.6, the perturbation direction is set to the ±5% range (0.57-0.63), and the output adjustment direction of the key unit (the unit with the lowest coal consumption rate) is associated (±3%-5% rated capacity); the environmental weight is perturbed in the opposite direction (0.3±Δ, where Δ is dynamically adjusted based on the load fluctuation rate; for example, when the load fluctuation rate is greater than 5%, the upper limit of Δ is set to 0.05). The perturbation parameter selection strategy includes: 1) main target sensitivity analysis, which determines the weight of the weight perturbation on the objective function through gradient calculation; 2) unit output perturbation priority, perturbing from low to high coal consumption rate to ensure that economically sensitive units are adjusted first.
[0137] (2) Generate multiple sets of parameter combinations based on the disturbance parameters and disturbance directions.
[0138] Specifically, perturbation parameter combinations can be generated through Monte Carlo sampling: 1) Weight coefficient perturbation: A ±5% random perturbation is applied to the economic, environmental, and flexibility weights respectively. For example, the original weights (0.6, 0.3, 0.1) may derive combinations such as (0.63, 0.27, 0.1) and (0.58, 0.31, 0.11), and the sum is 1 after normalization; 2) Output baseline perturbation: The units with the top 30% output in the basic solution are selected and randomly adjusted by ±5% of the rated capacity based on their current output. For example, the output of a 300MW unit is adjusted from 200MW to 185MW or 215MW, and it is prohibited from entering the vibration zone; 3) Start-stop state perturbation: The state of non-critical units (such as units with output <20% of the rated capacity and whose downtime meets the standard) is flipped (0→1 or 1→0). Each parameter combination includes a weight vector, an output baseline, and start / stop state change instructions, generating a total of 50 to 100 differentiated combinations.
[0139] (3) Adjust the weight coefficients and output baselines of multiple sets of parameter combinations to obtain multiple sets of directional parameter combinations.
[0140] In this embodiment, targeted adjustments can be performed on each parameter combination: 1) Weight coefficient adjustment: The objective function is reconstructed based on the perturbed weight vector. For example, when the economic weight is 0.63, the total cost = 0.63 × fuel cost + 0.27 × emission penalty + 0.1 × frequency regulation cost; 2) Output baseline correction: Based on the perturbed output value, the power balance constraint is recalculated. If the adjustment results in a power shortfall, the output is compensated from low to high coal consumption rate (for example, if the shortfall is 50MW, the output of the unit with the lowest coal consumption rate is prioritized to be increased to the upper limit); 3) Start-stop status synchronization update: If a unit changes from shutdown to startup, the minimum shutdown time is verified (e.g., ≥4 hours) and its minimum technical output (e.g., 120MW) is forcibly allocated. During the adjustment process, a rapid feasibility verification module is used to eliminate obvious non-compliant combinations (e.g., power shortfall > 5% or transient environmental protection exceedance > 10%), retaining 80% of the combinations for the solution phase.
[0141] (4) Each set of directional parameter combinations is independently solved as input to obtain differentiated candidate solutions.
[0142] Specifically, each set of directional parameter combinations can be input into a mixed integer programming (MILP) solver (such as CPLEX / Gurobi), with a short solution time limit set (e.g., a single solution of ≤10 seconds). Solution strategies include: 1) Hot start: Inheriting the unit start / stop status and output baseline of the basic solution as the initial solution; 2) Constraint relaxation: Temporarily relaxing the ramp rate to 110% (e.g., 5.5 MW / min) to accelerate convergence; 3) Local search: Limiting the adjustment range of decision variables to ±10% of the basic solution (e.g., output is limited to searching between 190 and 210 MW). After the solution is completed, non-dominated solutions are extracted and hard constraints are verified: If emissions exceed the standard or power imbalance occurs, a repair mechanism is triggered (e.g., enabling backup units to compensate for output); otherwise, the solution is stored in the candidate pool.
[0143] Finally, 80 to 120 sets of differentiated candidate solutions were generated, covering multi-dimensional trade-offs such as economic priority, environmental priority, and flexibility priority, and the diversity of the solution set was verified through the hypervolume index (HV≥0.8).
[0144] S2044: Perform safety constraint hard filtering on multiple differentiated candidate solutions to obtain an initial candidate scheduling solution set.
[0145] Specifically, the implementation process for hard safety-constrained filtering of multiple differentiated candidate solutions is as follows: First, a parallelized rule engine is built, pre-defining three types of hard constraints: unit-level, grid-level, and environmental protection-level. Unit-level rules include minimum start / stop times (e.g., a 4-hour restart after shutdown), output limits (e.g., 200MW≤P_i≤600MW), vibration zone entry (output prohibited between 150 and 180MW), and cumulative operating time limit (forced shutdown of ≤72 hours). Grid-level rules cover node voltage deviation ≤±5%, tie-line power factor ≥0.9, and frequency fluctuation ≤±0.2Hz. Environmental protection-level rules include instantaneous emission concentrations (e.g., SO2 <35mg / m³), hourly averages (NOx ≤50mg / m³), and daily cumulative emissions (dust ≤10 tons).
[0146] Hierarchical verification and rapid pruning are then performed: each candidate solution undergoes three-level verification in the order of unit level → grid level → environmental protection level. In the first-level unit constraint verification, all unit outputs and start-stop states are traversed. If a unit violates the vibration zone rules (such as an output of 160MW entering the prohibited area) or the shutdown time is insufficient (restarting after a 2-hour shutdown), the solution is directly eliminated. The second-level grid constraint verification calls a simplified power flow calculation model. If a solution causes the node voltage deviation to exceed the limit (such as 6%) or the line power factor to fall below 0.9, it is immediately marked as a violation. The third-level environmental protection constraint verification is based on the emission characteristic curve (such as the quadratic function relationship between unit output and NOx emissions), calculating the instantaneous and cumulative emissions. If the hourly average NOx emissions reach 52mg / m³ (4% over the limit), the solution is eliminated.
[0147] For critical compliance solutions (such as hourly emissions of 50.5 mg / m³), a dynamic relaxation mechanism is activated: the hourly mean verification is split into 5-minute granularity. If only a single period slightly exceeds the standard (such as 51 mg / m³ in the first 5 minutes and 49 mg / m³ in the subsequent 55 minutes), it is judged to be compliant after weighted averaging; if a unit exceeds the standard but other units in the same plant can be linked to activate backup desulfurization equipment (such as increasing processing capacity by 5%), the solution is allowed to be conditionally retained and the equipment compensation logic is triggered.
[0148] Finally, solutions that pass verification are aggregated and their diversity is controlled. Solutions are screened based on the spatial distribution density of the three-dimensional objectives of economy, environmental protection, and flexibility. If solutions in a particular area are too dense (for example, 10 solutions clustered within 5% of the spatial volume), the crowding distance (Euclidean distance between adjacent solutions) is calculated, and the top 20 solutions with a uniform distribution are retained. The final output set must meet 100% hard constraint compliance, target spatial coverage (HV ≥ 0.75), and solution spacing ≥ 3% of the total spatial diameter. Using distributed computing frameworks such as Spark for millisecond-level batch processing, 100,000 candidate solutions can be filtered down to a set of 200-500 strictly compliant initial solutions within 10 seconds.
[0149] In some embodiments, step 205 may specifically include the following steps:
[0150] S2051, determining the primary optimization target parameters and the secondary constraint target thresholds based on the real-time target weight coefficient.
[0151] Specifically, based on real-time weight coefficients (e.g., 0.6 for economy, 0.3 for environmental protection, and 0.1 for flexibility), the highest-weighted objective (economy) is selected as the primary optimization parameter. The remaining objectives are set as secondary constraint thresholds: the environmental protection threshold is set at 110% of the real-time emission limit (allowing a temporary exceedance of 10%), and the flexibility threshold is set at a frequency regulation rate no less than 85% of the forecasted grid demand. The secondary constraint thresholds are determined through historical scenario statistics. For example, when the environmental protection weight is 0.3, the corresponding emission cap relaxation coefficient is 1.1.
[0152] S2052: Based on the primary optimization objective parameter, a subset of the solutions to be optimized that meet the secondary constraint objective threshold is selected from the initial candidate scheduling solution set.
[0153] Specifically, all schemes that violate secondary constraints (such as emissions exceeding the limit by 110% and frequency modulation rates below 85%) can be eliminated from the initial set of candidate scheduling schemes (for example, 20 groups of schemes), and the remaining schemes can be sorted according to the main objective (economic cost), retaining the 50% schemes with the lowest costs (such as 10 groups) as the subset of schemes to be optimized.
[0154] The relaxation comparison method is used in the screening process: if the emission of a certain scheme is between 105% and 110% of the limit, but the economic cost is significantly lower than other schemes (difference > 5%), it will be retained and marked as an object that needs to be optimized.
[0155] S2053, optimize and solve the subset of optimization solutions through the mixed integer programming model to obtain the Pareto solution set.
[0156] Specifically, a MILP model was constructed with the primary objective (minimizing economic costs) as the objective function and secondary objectives (emissions ≤ 110% of the limit, frequency regulation rate ≥ 85%) as constraints. The decision variables included unit start / stop status (0-1 variables), output allocation (continuous variables), and environmental protection equipment commissioning combinations (integer variables). A branch-and-bound approach was used to solve the problem. For each solution in the optimization subset, the initial unit combination was used as the hot start input, and a local search was performed within the constraint boundaries (for example, with an output adjustment step of 5 MW). To generate a Pareto solution set, an ε-constraint method was employed: after each optimal solution was obtained, the environmental constraint was tightened by 5% (for example, from 110% to 105%), and the solution was repeated until no feasible solution was found. Finally, multiple sets of non-dominated solutions that satisfied the three-dimensional Pareto frontier of economics, environmental protection, and flexibility were generated and stored as a structured dataset containing the unit state matrix, objective function values, and constraint satisfaction indicators.
[0157] In some embodiments, this step may further include:
[0158] (1) Construct a mixed integer programming model that includes the main objective function, secondary constraints and equipment operation rules.
[0159] Specifically, the mixed integer programming model takes the economic goal (such as minimizing the total fuel cost) as the main objective function, and the mathematical expression can be minΣ(a_i·x_i+b_i·y_i), where x_i is a continuous variable (the output value of unit i), y_i is a 0-1 variable (the start and stop status of unit i), a_i and b_i are the fuel cost coefficient and start and stop cost, respectively.
[0160] The secondary constraints include environmental constraints (Σc_ij·x_i≤D_j, j is the pollutant type, D_j is the dynamic emission limit) and flexibility constraints (Σk_i·Δx_i / Δt≥F_min, k_i is the ramp rate of unit i, F_min is the minimum frequency regulation requirement of the power grid).
[0161] Equipment operation rules are converted into hard constraints: 1) Minimum start and stop time constraint of the unit (y_i must satisfy T_on ≥ T_min_on, T_off ≥ T_min_off); 2) Vibration zone entry prohibition constraint (x_i ∉ [P_low, P_high]); 3) Environmental protection equipment linkage constraint (if the desulfurization equipment is shut down, the associated unit x_i = 0).
[0162] The mixed integer programming model is implemented through the modeling interface of Gurobi or CPLEX (such as Python API). The dimension of decision variables is n (number of units) × 2 (x_i + y_i), and the number of constraints is 3n + m (m is the number of interaction constraints between environmental protection and power grid).
[0163] (2) Based on the mixed integer programming model, the solver is called to perform multi-threshold relaxation solution to obtain a set of candidate solutions.
[0164] Specifically, when calling the CPLEX solver, the ε-constraint method is used to generate a set of candidate solutions: 1) Initialize the relaxation threshold, relax the secondary objective (such as environmental friendliness) constraint to a limit of 120% (Σc_ij·x_i≤1.2D_j), and solve the primary objective to obtain an initial solution; 2) Gradually tighten the constraint threshold (step size 5%, that is, the next round of constraints is 1.15D_j), and retain non-dominated solutions after each solution; 3) Parallel solution strategy: Divide different thresholds into multiple threads (for example, 10 threads corresponding to thresholds 1.2D_j, 1.15D_j, …, 0.8D_j), and each thread independently calls the MILP solver to accelerate the search process.
[0165] Key technologies include a warm start that accelerates branch and bound using feasible solutions from the initial candidate set, and lazy constraints that dynamically add vibration zone constraints to reduce computational complexity. The solution terminates when no new solutions are generated after three consecutive iterations or when the maximum computational time (e.g., 30 minutes) is reached. The final output is a candidate set of 50 to 100 solutions, encompassing a multi-dimensional trade-off between economy, environmental protection, and flexibility.
[0166] (3) Through non-dominated sorting and uniform screening strategies, the Pareto solution set is extracted from the candidate solution set.
[0167] Specifically, the candidate solution set is firstly fast non-dominated sorting: 1) First-level non-dominated solution screening: all solutions are traversed. If a solution is not dominated by other solutions in all objectives (for example, at least one of the following is strictly better: the economy of solution A ≤ solution B, the environmental performance ≤ solution B, and the flexibility ≤ solution B), then it is classified into the first level; 2) After removing the first-level solutions, the process is repeated to generate the second level, the third level, and so on, until all solutions are stratified.
[0168] The congestion distance screening method is then used to improve the uniformity of the solution set: 1) normalize each target dimension; 2) calculate the Euclidean distance (congestion) of each solution between its adjacent solutions. For example, for the solution sequence sorted by economic efficiency, the congestion of the first and last solutions is set to infinity, and the congestion of the middle solution is (poor flexibility + poor environmental performance); 3) sort by congestion from large to small, and retain the top 30 solutions.
[0169] The resulting Pareto solution set must meet the following requirements: the proportion of solutions in each layer (first layer ≥ 70%), target space coverage (HV index ≥ 0.85), and the coefficient of variation of distances between solutions (standard deviation / mean ≤ 0.25). If these requirements are not met, a supplementary local search is triggered: Neighboring perturbation solutions (±5% output adjustment) are injected into sparse areas (such as the high-economic-low-environmental-performance range) and re-evaluated until convergence.
[0170] S2054, verifying whether the Pareto solution set meets the preset convergence conditions, which include the objective function variance threshold and the solution set distribution uniformity index.
[0171] Specifically, the objective function variance is first calculated. The economic objectives (such as unit power supply cost), environmental objectives (such as emission intensity), and flexibility objectives (such as frequency regulation response time) of all solutions in the solution set are normalized separately, and the values of each dimension are mapped to the interval [0,1]. The variance of each objective dimension is then calculated. For example, the economic variance is calculated as the sum of the squares of the differences between the cost values of all solutions and the average cost divided by (n-1), where n is the number of solution sets. The objective function variance threshold requires that the economic variance does not exceed 0.5% (normalized value) and the environmental and flexibility variances do not exceed 1%. If it is detected that the variance of a certain objective exceeds the standard (such as the economic variance reaches 0.6%), it is determined that the solution set has not converged on this core objective, and a supplementary optimization process needs to be triggered.
[0172] Next, the uniformity of the solution set distribution is evaluated. The hypervolume (HV) metric is used to measure the solution set's coverage of the target space. Using the worst value of each target dimension (e.g., a normalized value of 1.0 for economy, 1.0 for environmental performance, and 1.0 for flexibility) as a reference point, the volume of the hypercube formed by the solution set in three-dimensional space is calculated. The HV value must be no less than 0.8 (the maximum theoretical volume after normalization is 1). The uniformity of the Euclidean distances between pairs of adjacent solutions is also calculated. This is done by sorting all solutions by economic efficiency, calculating the distance between each pair of adjacent solutions in the target space, and then finding the ratio of the standard deviation σ of these distances to the mean. If this ratio exceeds 0.3 (for example, when σ = 0.12 and the average distance = 0.4, σ / mean is 0.3), the distribution is considered uneven.
[0173] Finally, the comprehensive verification logic is executed. Only when all target variances meet the standards, the HV value is ≥0.8, and the distance uniformity ratio is ≤0.3, is the solution set judged to meet the convergence conditions. If any of the conditions is not met, the system automatically records the defect type (such as economic variance exceeds the standard or HV is insufficient) and triggers the iterative optimization mechanism. The verification module generates a three-dimensional scatter plot and a numerical indicator report. For example, when a solution set contains 10 solutions with adjacent distances of 0.3, 0.5, 0.4, and 0.2, respectively, the average distance is 0.35, the standard deviation σ≈0.12, and σ / mean≈0.34, it will be marked as unevenly distributed. The manual review link can further check for outliers. If a solution deviates from the main distribution cluster by more than 2 standard deviations (for example, the environmental protection score of a solution is 0.9 while the other solutions are all below 0.7), a decision must be made as to whether to retain or eliminate the outlier solution.
[0174] S2055: Determine the Pareto solution set that meets the preset convergence condition as the Pareto optimal solution set.
[0175] The verified solution sets are deduplicated (those with cosine similarity > 0.95 are considered duplicate solutions) and sorted by economic efficiency to generate the final Pareto optimal solution set, which is stored as JSON structured data. This data includes the objective function value of each scheme, the unit output distribution matrix, the constraint satisfaction flag, and the solution timestamp for subsequent decision modules to call.
[0176] At the same time, metadata such as constraint relaxation parameters and number of iterations at convergence are recorded for retrospective analysis of model effects.
[0177] The method provided in the above embodiment can be executed by an electronic device. The following describes the electronic device in the embodiment of the present invention from the perspective of hardware processing. Figure 3 , which is a schematic diagram of a physical device structure of an electronic device in an embodiment of the present invention.
[0178] It should be noted that Figure 3 The structure of the electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present invention.
[0179] like Figure 3As shown, the electronic device includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes, such as the methods described in the above embodiments, based on programs stored in a read-only memory (ROM) 402 or programs loaded from a storage unit 408 into a random access memory (RAM) 403. RAM 403 also stores various programs and data required for system operation. CPU 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to bus 404.
[0180] The following components are connected to the input / output (I / O) interface 405: an input section 406 including an audio input device, push button switches, and the like; an output section 407 including a display, an audio output device, indicator lights, and the like; a storage section 408 including a hard disk and the like; and a communication section 409 including a network interface card such as a LAN (Local Area Network) card or a modem. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the input / output (I / O) interface 405 as needed. Removable media 411, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 410 as needed, so that computer programs read from the media can be installed in the storage section 408 as needed.
[0181] In particular, according to an embodiment of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present invention includes a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 409 and / or installed from removable media 411. When executed by the central processing unit (CPU) 401, the computer program performs the various functions defined in the present invention.
[0182] It should be noted that specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0183] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. Each box in the flowchart or block diagram can represent a module, program segment, or part of the code, and the above-mentioned module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings.
[0184] Specifically, the electronic device of this embodiment includes a processor and a memory, the memory is coupled to one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and one or more processors call the computer instructions to enable the electronic device to execute the method provided by the above embodiment.
[0185] As another aspect, the present invention further provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments, or may exist independently and not incorporated into the electronic device. The storage medium carries one or more computer programs, and when executed by a processor of the electronic device, the electronic device implements the methods provided in the above embodiments.
[0186] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention.
[0187] As used in the above embodiments, the term “when” may be interpreted to mean “if” or “after” or “in response to determining that” or “in response to detecting that”, depending on the context. Similarly, the phrases “upon determining that” or “if (stated condition or event) is detected” may be interpreted to mean “if determining that” or “in response to determining that” or “upon detecting (stated condition or event)” or “in response to detecting (stated condition or event)”, depending on the context.
[0188] Those skilled in the art will appreciate that all or part of the process steps in the above-described method embodiments can be implemented by a computer program instructing the relevant hardware. The program can be stored in a computer-readable storage medium, and when executed, the program can include the process steps in the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A multi-objective economic dispatch optimization method for a thermal power plant, characterized in that: include: Construct an initial data set based on the power plant's grid operating parameters, generator set status parameters, and environmental monitoring data; Determining characteristic parameters of current operational scheduling requirements based on the initial data set, the characteristic parameters including load demand fluctuation rate, unit adjustable margin, and environmental protection limit compliance rate; and determining real-time target weight coefficients of economic indicators, environmental indicators, and scheduling flexibility indicators based on the characteristic parameters through a preset dynamic weight allocation model; generating an initial candidate scheduling solution set that satisfies the safety constraints of the generator set according to the real-time target weight coefficient; Performing a main objective-oriented optimization on the initial candidate scheduling solution set to obtain a Pareto optimal solution set that meets a preset convergence condition; generating a scheduling recommendation plan based on the Pareto optimal solution set; The construction and training process of the dynamic weight distribution model includes: constructing a multimodal hybrid model structure, the multimodal hybrid model structure includes an input layer for integrating power grid, generator set data and environmental data, a constraint encoding layer embedded with physical mechanism equations, and a hierarchical strategy network based on a multi-head attention mechanism; generating a simulation data set containing multiple working conditions through simulation technology, the simulation data set covering fuel shortage, environmental protection equipment failure and power grid frequency disturbance scenarios; marking the Pareto optimal weight vector of each scenario in the simulation data set to form a training label set; using the training label set and the physical mechanism equation as network constraint rules, pre-training the initial parameters of the hierarchical strategy network to obtain a pre-trained model; performing phased reinforcement learning training on the hierarchical strategy network of the pre-trained model under single-unit steady-state, multi-unit coupling and random failure scenarios respectively to obtain the dynamic weight distribution model.
2. The method according to claim 1, characterized in that After the hierarchical strategy network of the pre-trained model is subjected to phased reinforcement learning training in the single-unit steady-state, multi-unit coupling and random failure scenarios to obtain the dynamic weight allocation model, it also includes: introducing a generative adversarial network to construct a perturbation sample set; performing robustness reinforcement training on the dynamic weight allocation model based on the perturbation sample set; and dynamically fine-tuning the output layer parameters of the dynamic weight allocation model based on the deviation between the actual scheduling results and the simulation prediction through an online transfer learning framework.
3. The method according to claim 1, characterized in that The main target-oriented optimization of the initial candidate scheduling scheme set to obtain a Pareto optimal solution set that meets the preset convergence conditions includes: determining the main optimization target parameters and the secondary constraint target threshold based on the real-time target weight coefficient; screening a subset of schemes to be optimized that meet the secondary constraint target threshold from the initial candidate scheduling scheme set according to the main optimization target parameters; optimizing and solving the subset of schemes to be optimized through a mixed integer programming model to obtain a Pareto solution set; verifying whether the Pareto solution set meets the preset convergence conditions, which include the objective function variance threshold and the solution set distribution uniformity index; and determining the Pareto solution set that meets the preset convergence conditions as the Pareto optimal solution set.
4. The method according to claim 3, characterized in that The method optimizes and solves the subset of solutions to be optimized through a mixed integer programming model to obtain a Pareto solution set, including: constructing a mixed integer programming model including a primary objective function, secondary constraints and equipment operation rules; based on the mixed integer programming model, calling a solver to perform a multi-threshold relaxation solution to obtain a candidate solution set; and extracting the Pareto solution set from the candidate solution set through a non-dominated sorting and uniform screening strategy.
5. The method according to claim 1, wherein The method of generating an initial set of candidate scheduling schemes that meet the safety constraints of the generator set based on the real-time target weight coefficient includes: constructing a multi-objective weighted aggregation function based on the real-time target weight coefficient; calling a fast heuristic algorithm to generate a basic feasible solution based on the multi-objective weighted aggregation function; expanding the basic feasible solution through a parameter perturbation method to obtain multiple differentiated candidate solutions; and performing hard filtering of the multiple differentiated candidate solutions based on safety constraints to obtain an initial set of candidate scheduling schemes.
6. The method according to claim 5, characterized in that The basic feasible solution is expanded by the parameter perturbation method to obtain multiple differentiated candidate solutions, including: determining the disturbance parameters and the disturbance direction based on the real-time target weight coefficient and the basic feasible solution; generating multiple groups of parameter combinations according to the disturbance parameters and the disturbance direction; adjusting the weight coefficients and output baselines of the multiple groups of parameter combinations to obtain multiple groups of directional parameter combinations; and independently solving each group of the directional parameter combinations as input to obtain differentiated candidate solutions.
7. An electronic device, characterized in that: The electronic device comprises one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the electronic device to execute the method according to any one of claims 1 to 6.
8. A computer-readable storage medium storing computer instructions, characterized in that: When the computer instructions are executed on an electronic device, the electronic device is caused to execute the method according to any one of claims 1 to 6.
9. A computer program product, characterized in that When the computer program product is run on an electronic device, the electronic device is enabled to perform the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Environmental economy power generation dispatching method
CN104009494A
Thermal power plant environment economic dispatching method based on multi-target differential evolution algorithm
CN105809297A