Low-carbon park control method and system based on digital twinning and dynamic carbon footprint

By combining digital twin technology and dynamic carbon footprint, the problems of static carbon accounting and rigid control strategies in the optimized operation of low-carbon parks have been solved. Dynamic carbon management and adaptive optimization of the low-carbon park system have been realized, improving the system's operating efficiency and carbon emission reduction capabilities under uncertain environments.

CN122092388BActive Publication Date: 2026-07-03STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO
Filing Date
2026-04-22
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

The existing low-carbon industrial park optimization operation suffers from problems such as static carbon accounting, disconnect between day-ahead planning and real-time control, and rigid control strategy parameters. As a result, the optimization model cannot identify the temporal differences in carbon emissions, is unable to cope with the fluctuations of renewable energy and load, and the system operates rigidly and lacks adaptability.

Method used

A low-carbon park control method based on digital twins and dynamic carbon footprint is adopted. By constructing a digital twin simulation model and combining dynamic carbon cost and reinforcement learning, a two-level closed-loop optimization of day-ahead optimization and real-time control is achieved. Carbon accounting is performed using dynamic marginal carbon emission factors, and carbon weights are dynamically adjusted through reinforcement learning agents. A hybrid intelligent optimization algorithm is constructed to improve the system's adaptability.

Benefits of technology

It enables dynamic sensing and optimization of carbon accounting, improves the system's robustness and carbon efficiency stability under uncertain environments, can proactively avoid high-carbon periods, adaptively adjust operating strategies, and enhance the intelligence level and long-term carbon efficiency performance of the park's energy system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122092388B_ABST
    Figure CN122092388B_ABST
Patent Text Reader

Abstract

The application relates to a low-carbon park control method and system based on digital twinning and dynamic carbon footprint, a digital twinning simulation model mapped with a physical entity of a low-carbon park is constructed; in a day-ahead optimization stage, a day-ahead operation plan of a park energy system is formulated with the objective of minimizing total day operation cost in a scheduling period, serving as a baseline and operation boundary for real-time control rolling optimization; in a real-time control stage, a model predictive control is adopted to perform rolling optimization on the day-ahead operation plan based on a predicted curve of system state in each control period, the first control instruction in an optimized sequence obtained in the current control period is applied to an actual system, and a dynamic carbon cost weight in a model predictive control objective function is dynamically adjusted by using a reinforcement learning intelligent agent. Compared with the prior art, the application realizes accurate perception, forward deduction and intelligent decision of a whole process of the park energy system, and improves overall optimization performance and operation robustness of the system in an uncertain environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of low-carbon industrial park optimization control technology, and in particular to a low-carbon industrial park optimization control method based on digital twins and dynamic carbon footprint. Background Technology

[0002] Constructing integrated low-carbon / zero-carbon industrial park energy systems that combine distributed renewable energy sources (such as photovoltaics and wind power), energy storage systems, and flexible loads, and improving their operational efficiency through optimized scheduling, is a current research hotspot. Although significant progress has been made in the research on optimized operation of low-carbon industrial parks, their practical application still faces three core challenges:

[0003] First, there is a contradiction between the static nature of carbon accounting methods and the extensive nature of carbon reduction decision-making. Most existing studies use periodically averaged carbon emission factors for carbon accounting and optimization. This method severely ignores the essential characteristic of real-time fluctuations in the marginal carbon emission factor of the power grid. This static perspective prevents optimization models from identifying and utilizing the temporal differences in carbon emissions, resulting in scheduling outcomes that only minimize energy costs, not carbon costs. It fails to guide industrial parks to proactively reduce grid purchases, increase energy storage discharge, or reduce load during periods of high marginal carbon intensity, thus missing the enormous potential for refined carbon management over time.

[0004] Second, there is a contradiction between the fragmentation of day-ahead planning and real-time control and system uncertainty. Current research often treats day-ahead optimization and real-time control as two independent processes. Day-ahead planning generates rigid scheduling schemes based on deterministic predictions, while real-time control often uses simple rule-based or proportional control to track the plan. This open-loop architecture is ill-suited to effectively address the inherent volatility and uncertainty of renewable energy and loads. Once prediction deviations occur, the real-time layer can only respond passively, lacking the ability to proactively re-optimize, which can easily lead to the system deviating from its optimal trajectory and even causing reliability issues.

[0005] Third, there is a contradiction between the fixed control strategy parameters and the complex and ever-changing operating environment. The weight parameters in the objective function of existing optimization models are usually fixed values ​​pre-set based on experience. This setting cannot adapt to the changing scenarios in actual park operation. Parameters optimal in one scenario may lead to a significant performance decline or even conflict in another. The system lacks an intelligent mechanism capable of online evaluation of operational performance and autonomously adjusting optimization objectives accordingly. This deficiency makes the entire energy system rigid and sluggish, making it difficult to maintain optimal carbon efficiency stably throughout the entire operating cycle, thus limiting its long-term applicability and effectiveness. Summary of the Invention

[0006] This invention addresses the problems existing in the current optimized operation of low-carbon industrial parks, such as static carbon accounting, disconnect between day-ahead planning and real-time control, and fixed control strategy parameters. It provides a low-carbon industrial park control method and system based on digital twins and dynamic carbon footprint.

[0007] The objective of this invention can be achieved through the following technical solutions:

[0008] As a first aspect of the present invention, a low-carbon industrial park control method based on digital twins and dynamic carbon footprints is provided, comprising the following steps:

[0009] Construct a digital twin simulation model that maps to the physical entities of the low-carbon park;

[0010] In the day-ahead optimization phase, with the goal of minimizing the total daily operating cost within the scheduling cycle, a day-ahead operation plan for the park's energy system is formulated as the baseline and operating boundary for real-time control and rolling optimization. The total daily operating cost includes the dynamic carbon cost, which converts carbon emissions into economic costs based on the dynamic marginal carbon emission factor of the power grid.

[0011] In the real-time control phase, during each control cycle, the model predictive control is used to perform rolling optimization of the day-ahead operating plan based on the prediction curve of the system state, and the first control command in the optimized sequence obtained in the current control cycle is applied to the actual system. The objective function of the model predictive control introduces dynamic carbon weights on the basis of the day-ahead optimized objective function, and a reinforcement learning agent is used to dynamically adjust the dynamic carbon cost weights in the objective function of the model predictive control.

[0012] As a preferred technical solution, the total system operating cost in the objective function of the current day optimization includes: electricity trading cost, which is the economic expenditure and income generated by the park through electricity trading with the grid; operation and maintenance cost, which is the loss and maintenance cost of distributed generation equipment and energy storage system in the park during operation; dynamic carbon cost, which introduces the dynamic marginal carbon emission factor of the grid to convert carbon emissions into economic costs; and penalty cost, which considers the total load power, energy storage charging and discharging power, and photovoltaic and wind power output to penalize the power imbalance of the park's energy system.

[0013] As a preferred technical solution, the current-day optimized constraints include: power balance constraints, ensuring that the power generation and power consumption of the park's energy system are balanced in real time; energy storage system operation constraints, including state of charge evolution constraints, upper and lower limits of SOC constraints, charging and discharging power constraints, and charging and discharging mutual exclusion constraints; grid interaction power constraints, constraining the power exchange between the park and the upper-level grid at the connection point to ensure operational safety and compliance with the grid connection protocol; distributed generator output constraints, ensuring that the generator output is between the lower and upper technical limits; and renewable energy curtailment constraints, ensuring that the curtailed power does not exceed the available output of renewable energy.

[0014] As a preferred technical solution, the real-time control, based on the day-ahead optimization constraints, introduces an initial state constraint into the energy storage system operation constraints, constraining the initial state of the park's energy system to be equal to the actual measured value at the corresponding time; the optimal control sequence is obtained through model predictive control, and the control command for the first time period is issued to the physical system for execution. The new measured value in the next control cycle is used as the initial state for re-prediction and optimization.

[0015] As a preferred technical solution, the reinforcement learning model for the dynamic carbon cost weight adjustment problem is as follows:

[0016] Reinforcement learning's state-space representation of system operational characteristics and resource status:

[0017]

[0018] In the formula: Set time for the past Average carbon intensity of the inner park; Set time for the past Average wind and solar penetration rate within the area; for t The actual state of charge of stored energy at all times.

[0019] The actions of a reinforcement learning agent are decisions about adjusting dynamic carbon cost weights:

[0020]

[0021] In the formula: This is the carbon cost weighting adjustment amount; The preset adjustment step size;

[0022] The action space is designed with three discrete options: increase weights, decrease weights, or keep the weights constant.

[0023] The reward function of reinforcement learning balances carbon efficiency goals with operational economics:

[0024]

[0025] In the formula: for t The average carbon intensity of the park during the actual operation of the time period; The target carbon intensity value that the park hopes to achieve; for t Total operating cost of the time-slot system; is the regularization coefficient.

[0026] As a preferred technical solution, the reinforcement learning agent interacts with the digital twin simulation model in the early stage of deployment and with the actual park energy system in the later stage to learn the optimal strategy.

[0027] As a second aspect of the present invention, a low-carbon park optimization control system based on digital twin and dynamic carbon footprint is provided, the system being composed of a digital twin layer, a day-ahead optimization layer, a real-time control layer and a physical system layer.

[0028] The digital twin layer integrates a simulation model, a multi-timescale prediction module, a dynamic carbon footprint database, and an optimization algorithm library, providing full-domain data perception and decision support for the day-ahead optimization layer and the real-time control layer.

[0029] The day-ahead optimization layer is based on high-precision wind and solar load forecasting, and constructs a mixed-integer linear programming model with multiple objectives of economy, low carbon emissions, and reliability to generate day-ahead scheduling plans.

[0030] The real-time control layer receives the day-ahead scheduling plan through the edge computing device and uses model predictive control method to continuously correct the scheduling instructions; the instructions are then sent to the physical system layer for execution.

[0031] The physical system layer state data is uploaded to the digital twin layer in real time to update the virtual model state. The real-time control layer uses the updated data to optimize and provide feedback on the execution effect. The digital twin layer uses the accumulated data to train and optimize the algorithm.

[0032] As a preferred technical solution, the day-ahead optimization layer formulates the next-day scheduling plan for the energy system, and its functions include:

[0033] Multi-timescale forecast data fusion integrates meteorological and production planning information to generate high-precision forecast curves for wind and solar power output, load demand, time-of-use electricity prices, and dynamic marginal carbon emission factors of the power grid for the next day.

[0034] A multi-objective optimization scheduling model is solved, and a planning model is constructed and solved with the goal of minimizing the total daily operating cost. The total daily operating cost incorporates the dynamic carbon cost, which converts carbon emissions into economic costs based on the dynamic marginal carbon emission factor of the power grid.

[0035] The plan is issued and the boundary is set. The day-ahead scheduling plan obtained from the solution is issued to the real-time control layer as the baseline and operating boundary for rolling optimization.

[0036] As a preferred technical solution, the real-time control layer is deployed on the edge computing nodes of the park, using high-frequency rolling optimization and adaptive adjustment to compensate for the deviation between the current plan and the actual operating status caused by prediction errors. The functional modules include:

[0037] High-frequency data sensing and analysis: Decompose the total load of the park's energy system and identify the operating status of equipment;

[0038] Ultra-short-term forecasting module: performs rolling minute-level forecasts of future ultra-short-term wind and solar power output and load demand;

[0039] The dynamic carbon footprint real-time tracking engine accesses the dynamic marginal carbon emission factor of the power grid in real time via API, and combines the interaction power between the park and the power grid to calculate the total carbon flow of the park and the carbon footprint of key equipment at a frequency of minutes.

[0040] The core of the hybrid intelligent optimization control system is model predictive control, which acts as the executor of rolling optimization. It executes rolling optimization on a minute-by-minute cycle, receiving the daily plan as a benchmark and combining the latest ultra-short-term forecasts, the real-time status of the park's energy system, and carbon weight coefficients dynamically provided by a reinforcement learning agent. It solves the optimization problem within a shortened prediction time domain, generates real-time control commands, and issues them for execution. The reinforcement learning agent operates on a timescale longer than the rolling cycle of model predictive control. Through interaction with the environment, it autonomously learns and dynamically adjusts the carbon weight coefficients in the objective function of model predictive control.

[0041] As a preferred technical solution, the digital twin layer, as a driving engine for perception and closed-loop optimization, includes the following functions:

[0042] High-fidelity multi-dimensional modeling is used to construct a virtual model that maps to the physical entities of the park's energy system, including a refined model at the equipment level and a network topology model at the system level.

[0043] The data fusion and governance center integrates and manages real-time data, historical data, and external data from the physical park's energy system.

[0044] The algorithm model library and simulation sandbox encapsulate and manage prediction, optimization and reinforcement learning training algorithms, and provide a simulation environment for testing new policies, extrapolating extreme scenarios and offline training of reinforcement learning agents.

[0045] Visual analysis and decision support enable panoramic visualization and dynamic tracking of energy flow, carbon flow, and economic indicators.

[0046] Compared with the prior art, the present invention has the following beneficial effects:

[0047] 1) This invention proposes a method for modeling and quantifying the carbon footprint of industrial parks based on dynamic marginal carbon emission factors, and embeds it into the entire process of day-ahead optimization and real-time control. It abandons the traditional periodic average carbon emission factor and instead adopts a time-varying sequence of dynamic marginal carbon emission factors from an authoritative database. In the day-ahead optimization stage, this dynamic factor sequence is used as a key input parameter to construct an optimization objective function that considers dynamic carbon costs. This allows optimization decisions to proactively avoid purchasing electricity from the grid during periods of high marginal carbon intensity, guiding energy consumption habits towards lower-carbon periods. In the real-time control stage, real-time updated grid carbon intensity data is accessed, and combined with ultra-short-term forecasts, carbon costs are calculated in real-time during the rolling optimization of model predictive control, achieving a minute-level response to the dynamic carbon footprint. This method essentially upgrades carbon accounting from static accounting to dynamic perception and optimization, fully exploring the carbon reduction potential over time.

[0048] 2) This invention designs a hybrid intelligent optimization algorithm integrating reinforcement learning and model predictive control for adaptive decision-making in the real-time control layer. A reinforcement learning agent is designed with the system's recent average carbon intensity, wind and solar penetration rate, and energy storage state of charge as the state space, carbon cost weighting actions as the action space, and the deviation between actual carbon efficiency and the target, as well as the total operating cost, as the reward function. This agent operates through offline training and online application, capable of autonomously learning and dynamically adjusting the carbon weight coefficients in the objective function of the model predictive controller. The long-term reward mechanism of reinforcement learning compensates for the limitations of finite-time-domain optimization in model predictive control, enabling the system not only to cope with short-term fluctuations but also to adaptively adjust its operating strategy to approach the long-term carbon efficiency target, effectively improving the system's intelligence level and carbon efficiency stability under different operating scenarios.

[0049] 3) This invention constructs a digital twin-driven, two-stage prediction-control dual-layer closed-loop optimization framework for low-carbon industrial parks, encompassing day-ahead and real-time prediction. It consists of a digital twin layer, a day-ahead optimization layer, a real-time control layer, and a physical system layer, forming an organic whole. The digital twin layer integrates a high-fidelity model, multi-timescale prediction modules, a dynamic carbon footprint database, and an optimization algorithm library, providing comprehensive data perception and decision support for both upper and lower layers, serving as the digital brain for closed-loop optimization. The day-ahead optimization layer, based on high-precision wind and solar load prediction, constructs a mixed-integer linear programming model with multiple objectives of economy, low carbon emissions, and reliability to generate day-ahead scheduling plans. The real-time control layer, through edge computing devices, integrates non-intrusive load monitoring and equipment-level real-time carbon flow metering technology, frequently sensing system status and employing model predictive control methods to continuously correct scheduling instructions. This addresses the disconnect between day-ahead planning and real-time control, significantly improving the robustness and overall optimization performance of the industrial park system under uncertain environments. Attached Figure Description

[0050] Figure 1This forms the core structure and information interaction process of the overall framework of the present invention.

[0051] Figure 2 The following is an optimized solution flowchart for this invention.

[0052] Figure 3 This is a flowchart of the real-time control solution method of the present invention.

[0053] Figure 4 The input data in this embodiment of the invention are: a) load and renewable energy output, b) time-of-use electricity price curve, c) dynamic change of marginal carbon intensity of the power grid, and d) change of renewable energy penetration rate.

[0054] Figure 5 The following are the day-ahead optimization results in the embodiments of the present invention: a) day-ahead optimization - power balance analysis, b) day-ahead optimization - energy storage operation strategy, c) day-ahead optimization - cost structure analysis, and d) day-ahead optimization - renewable energy curtailment analysis.

[0055] Figure 6 The real-time control results in the embodiments of the present invention are as follows: a) Real-time control - power tracking with grid interaction, b) Real-time control - adaptive adjustment of energy storage operation, c) Real-time control - adaptive adjustment of carbon weight, and d) Real-time control - carbon intensity tracking effect.

[0056] Figure 7 The following is a carbon efficiency analysis diagram in an embodiment of the present invention: a) carbon intensity comparison analysis, b) real-time carbon intensity change trend, c) carbon emission reduction contribution decomposition, and d) carbon efficiency target achievement status.

[0057] Figure 8 For the comparative analysis of current optimization and actual control in the embodiments of the present invention, a) is a comparison of total cost, b) is a comparison of carbon emissions, c) is a comparison of renewable energy utilization, and d) is a comparison of system operation performance indicators.

[0058] Figure 9 This is a schematic diagram of the reinforcement learning process of the reinforcement learning agent in an embodiment of the present invention. a) is the reinforcement learning carbon weight adaptive adjustment process, and b) is the reinforcement learning action distribution statistics. Detailed Implementation

[0059] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.

[0060] Example 1

[0061] This invention addresses the three core challenges faced by low-carbon industrial parks in their operation: static carbon accounting, fragmented optimization control, and rigid strategy parameters. It innovatively proposes a digital twin-driven, day-ahead / real-time two-stage closed-loop optimization control method. This method deeply integrates cyber-physical systems, advanced forecasting, multi-timescale optimization, and artificial intelligence into a cohesive whole. Using dynamic carbon footprint as the core optimization guide and evaluation benchmark, it breaks down the barriers between the digital virtual space and the physical entity system through a closed-loop data flow of perception-prediction-decision-control-feedback, achieving coordinated optimal and adaptive operation of the park's energy system under multiple objectives of economy, low carbon emissions, and reliability.

[0062] The core concept of this invention lies in utilizing digital twin technology to construct a high-fidelity virtual model that synchronously maps, interactively verifies, and operates in parallel with the physical park. This makes it not merely a digital mirror, but a digital brain capable of panoramic perception, deep insight, forward-looking projection, and intelligent decision-making. Under this concept, the framework logically constructs four distinct and functionally coupled layers from top to bottom (from macro-strategy to micro-execution): the digital twin layer, the day-ahead optimization layer, the real-time control layer, and the physical system layer. Data, instructions, and feedback information form an efficient, bidirectional, and closed-loop information and control flow among these four layers, jointly constructing an energy system organism capable of continuous perception, dynamic optimization, autonomous decision-making, and continuous intelligent evolution. The core structure and information interaction process of the overall framework are as follows: Figure 1 As shown below, the four core levels and their collaborative closed-loop mechanism are explained.

[0063] 1.1 Physical System Layer: As the physical foundation and holographic sensing source of the energy system, the physical system layer is the physical entity on which the entire optimization framework operates and relies, and is the carrier and source of energy flow, material flow and information flow.

[0064] The system's composition and functional positioning encompass the energy supply side, energy storage side, energy consumption side, and interconnection side. The energy supply side is the foundation for achieving low-carbon development, including distributed renewable energy sources with intermittent and fluctuating output (such as rooftop photovoltaics (PV) and distributed wind power (WT)), as well as conventional distributed generator sets (such as gas turbines) that offer strong controllability and serve as important supplementary and regulating units. The energy storage side, represented by electrochemical energy storage systems (ESS), is a key flexible resource for solving the problem of energy mismatch in time and space. It achieves energy transport and buffering through charging and discharging, and its state of charge (SOC) is a crucial state variable connecting short-term and long-term optimization. The energy consumption side is divided into flexible loads (such as air conditioning, electric vehicle charging stations, and adjustable industrial loads) and rigid loads (critical loads that must be met immediately) based on their controllability. Flexible loads are important resources for demand-side response and mitigating fluctuations. The interconnection side, or connection point (PCC) with the upper-level power grid, serves as a bridge for interconnecting and sharing energy resources between the park and regional energy systems, participating in broader energy optimization through electricity purchase and sale operations.

[0065] The core functions include an energy flow execution terminal and a holographic data sensing terminal. The energy flow execution terminal faithfully executes the control commands issued by the upper-level optimization, completing a series of physical processes such as energy production, storage, consumption, and exchange. The holographic data sensing terminal, through the deployment of a wide range of sensor networks (smart meters, PMUs, environmental sensors, etc.), collects real-time, high-frequency panoramic operation data of the system (power, voltage, SOC, temperature, etc.), providing a real and up-to-date data foundation for upper-level optimization decisions.

[0066] 1.2 Real-time Control Layer: As the edge intelligence and adaptive tactical execution unit, the real-time control layer is deployed on the edge computing nodes of the campus. Its core mission is to compensate for deviations between the current plan and the actual operating state caused by prediction errors. Through high-frequency rolling optimization and intelligent adaptive adjustments, it ensures that the system dynamically approaches the optimal trajectory during actual operation. In terms of deployment and performance, it adopts edge deployment, close to physical devices to ensure low latency and high reliability of control; it possesses high-performance computing capabilities to meet the timeliness requirements of solving complex optimization problems within minute-level time windows. Its core functional modules include:

[0067] (1) High-frequency data sensing and intelligent analysis. Based on technologies such as non-intrusive load monitoring (NILM), the total load is finely decomposed and the operating status of key equipment is identified, providing data support for precise control.

[0068] (2) Ultra-short-term forecast module. It performs rolling forecasts of wind and solar power output and load demand for the next 15-30 minutes. Its accuracy is significantly higher than that of day-ahead forecasts, which is the key to the effectiveness of model predictive control (MPC).

[0069] (3) Dynamic carbon footprint real-time tracking engine. By accessing the dynamic marginal carbon emission factor of the power grid in real time through API, and combining the interaction power between the park and the power grid, the total carbon flow of the park and the carbon footprint of key equipment can be calculated at the minute level.

[0070] (4) Hybrid Intelligent Optimization Control Core (MPC+RL). Model Predictive Control (MPC) is the executor of rolling optimization, executing on a 5-15 minute cycle. It receives the day-ahead plan as a benchmark, combines the latest ultra-short-term forecasts, real-time system status (such as measured SOC values), and carbon weight coefficients dynamically provided by the RL agent, solves the optimization problem within a shortened prediction time domain, generates real-time control commands, and issues them for execution. Its rolling optimization and feedback correction mechanisms are the core of its ability to cope with uncertainty. The Reinforcement Learning (RL) agent is the adaptive adjuster of policy parameters, operating on a slower time scale (such as hourly). It evaluates the system's recent macroscopic carbon efficiency performance, autonomously learns and dynamically adjusts the carbon weight coefficients in the MPC objective function through interaction with the environment, enabling the system behavior to transcend the finite time domain of MPC and continuously tend towards the long-term optimal carbon efficiency.

[0071] 1.3 Day-ahead Optimization Layer: As the center for overall coordination and forward-looking strategic planning, the day-ahead optimization layer is responsible for developing a comprehensive and forward-looking scheduling plan for the park for the next 24 hours based on relatively complete medium-term forecast information. Its core functions include:

[0072] (1) Multi-timescale forecast data fusion. Integrate meteorological, production planning and other information to generate high-precision forecast curves for wind and solar power output, load demand, time-of-use electricity price and dynamic marginal carbon emission factor of the power grid for the next day.

[0073] (2) Solving the multi-objective optimization scheduling model. A mixed integer linear programming (MILP) model is constructed and solved. The model aims to minimize the total daily operating cost. Its innovation lies in the deep integration of dynamic carbon cost into the cost function. It uses the time-varying marginal carbon emission factor of the power grid as a key input to guide the optimization decision to actively avoid purchasing electricity during high-carbon periods, and strategically tap the carbon emission reduction potential in the time dimension.

[0074] (3) Plan issuance and boundary setting. The day-ahead scheduling plan (energy storage SOC plan, grid interaction plan, etc.) obtained from the solution is issued to the real-time control layer as the baseline and operating boundary for its rolling optimization, so as to ensure the consistency between real-time control and the overall strategic objectives.

[0075] 1.4 Digital Twin Layer: As the driving engine for comprehensive perception and closed-loop optimization, the digital twin layer is not an independent execution layer, but a systematic platform that provides comprehensive, high-fidelity, and immersive decision support for both day-ahead and real-time optimization layers. Its core functions include:

[0076] (1) High-fidelity multi-dimensional modeling. Construct virtual models that are mapped one-to-one with physical entities, including device-level refined models (such as photovoltaic efficiency curves and energy storage degradation models) and system-level network topology models.

[0077] (2) Data Fusion and Governance Center. Integrates and manages real-time data, historical data and external data (weather, carbon intensity API) from physical systems to form a unified, high-quality data lake.

[0078] (3) Algorithm model library and simulation sandbox. Encapsulates and manages various prediction, optimization and RL training algorithms. Provides powerful simulation capabilities for new policy testing, extreme scenario simulation and safe and efficient offline training of RL agents.

[0079] (4) Visualization analysis and decision support. Enable panoramic visualization and dynamic traceability of energy flow, carbon flow and economic indicators, providing operators with in-depth insights.

[0080] 1.5 Closed-loop information flow: As the source of system intelligence and vitality, the superior performance of this framework stems from its bidirectional closed-loop information flow that runs through four layers, which endows the system with the vitality of self-perception, decision-making, execution and evolution.

[0081] Forward instruction flow (top-down decision execution): The path is as follows: the day-ahead optimization layer generates a scheduling plan → the real-time control layer receives the plan and continuously generates fine-grained instructions → the instructions are issued to the physical system layer for execution. This process achieves the connection from macro strategy to micro tactics, ensuring that every action of the system serves the global optimization goal.

[0082] The reverse feedback flow (bottom-up state awareness and learning evolution) follows this path: real-time uploading of physical system state data → updating of the virtual model state by the digital twin layer → optimization using the latest data and feedback of execution results by the real-time control layer → training and optimization of algorithms (such as RL) using accumulated data by the digital twin layer → improvement of decision-making quality by the optimized algorithms at both the top and bottom layers. This process achieves a precise mapping from the physical world to the digital world, and through a data-driven approach, enables the system's prediction accuracy, optimization strategies, and control intelligence to continuously iterate and evolve, forming a positive cycle that becomes smarter with use.

[0083] The overall framework constructed in this invention achieves deep integration and bidirectional driving of virtual and physical spaces through digital twin technology. It addresses the disconnect between planning and control through day-ahead and real-time two-stage collaborative optimization, achieves refined and guided carbon management through the embedding of dynamic carbon footprints, and finally realizes rapid response and long-term self-adaptation of the system under uncertain environments through hybrid intelligence of MPC and RL. This framework not only provides a systematic and implementable technical solution for the optimized operation of low-carbon industrial parks, but also lays a solid theoretical and engineering foundation for building future smart energy systems with self-sensing, self-decision-making, and self-evolution capabilities.

[0084] 2. The current optimization layer modeling.

[0085] The optimization layer, serving as the strategic decision-making center of the entire framework, bears the crucial responsibility of formulating the 24-hour operation plan for the park's energy system. Its core is to construct and solve a multi-objective mixed-integer linear programming model, aiming to balance economy, low-carbon performance, and reliability. Compared to traditional models, the core innovation of this invention lies in introducing a carbon cost term based on a dynamic marginal carbon emission factor. This transforms the optimization decision-making from simply pursuing the lowest energy consumption cost to pursuing the lowest overall carbon cost, thereby guiding energy consumption behavior towards low-carbon transformation over time.

[0086] 2.1 Objective Function: The objective of the day-ahead optimization model is to minimize the scheduling cycle. T (usually 24 hours, in order to) This refers to the total system operating cost over a time interval (such as 15 minutes or 1 hour). This total cost is a composite cost, consisting of four parts: electricity trading costs, equipment maintenance costs, dynamic carbon costs, and penalty costs for exceeding system operating limits. Its mathematical expression is as follows:

[0087] (1)

[0088] In the formula: t Indexed by time period; T This represents the total number of time periods within a scheduling cycle. for t Cost of electricity traded with the grid during a given time period (in yuan); for t Operation and maintenance costs (in yuan) of each power generation device in the park during the specified period. for t The carbon cost (in yuan) incurred from purchasing electricity from the grid during a given period is the core of reflecting the dynamic carbon footprint; for t The penalty cost (in yuan) for system power imbalance and other issues during certain periods is used to ensure system reliability.

[0089] 2.1.1 Electricity trading costs: This cost represents the economic expenditures and revenues generated by the park through electricity trading with the power grid.

[0090] (2)

[0091] In the formula: for t The time period is based on the grid's electricity purchase price (yuan / kWh); for t Power purchased from the grid during the specified time period (kW); for t Electricity price sold to the grid during specific time periods (RMB / kWh); for t Electricity sold to the grid during a given time period (kW); The duration (h) of each time period.

[0092] 2.1.2 Operation and Maintenance Costs: This cost describes the losses and maintenance expenses incurred during the operation of distributed power generation equipment and energy storage systems within the park. It is usually simplified to be proportional to the power generation or usage.

[0093] (3)

[0094] In the formula: This represents the total number of distributed generation units (such as gas turbines); For the first i Operation and maintenance coefficient of each distributed generation unit (yuan / kWh); For the first i Distributed generation units in t Output power (kW) for a given time period; The operation and maintenance coefficient of the energy storage system (yuan / kWh); For energy storage systems in t The charging and discharging power (kW) of the segment is negative during charging and positive during discharging. Taking the absolute value indicates that losses occur during both charging and discharging.

[0095] 2.1.3 Dynamic Carbon Cost: This is the key innovation of the model, transforming traditional static carbon emission accounting into dynamic carbon footprint optimization based on time series. It enables the model to detect and avoid periods of high marginal carbon intensity in the power grid.

[0096] (4)

[0097] In the formula: for t The dynamic marginal carbon emission factor (kgCO2 / kWh) of the power grid over a given period is derived from forecast data from an authoritative database and varies over time. t Change is the core variable that reflects dynamic attributes; The carbon price (yuan / kgCO2) converts carbon emissions into economic costs, thereby achieving a balance between economic efficiency and low carbon emissions in the optimization process.

[0098] 2.1.4 Penalty Cost: To ensure the reliability of system operation and the feasibility of the mathematical model, a penalty term is introduced to apply soft constraints to behaviors such as power imbalance.

[0099] (5)

[0100] In the formula: Power imbalance penalty factor (yuan / kW) 2 ); for t Total load power (kW) for the time period; , They are respectively t For ease of modeling, the charging and discharging power (kW) of time-period energy storage is usually split into two non-negative variables; , They are respectively t Forecast output (kW) of photovoltaic and wind power during the time period.

[0101] 2.2 Constraints. Constraints define the physical boundaries and operating rules for the safe and stable operation of the system, and are the fundamental guarantee for the feasibility of the optimization results.

[0102] 2.2.1 Power balance constraint: The power generation and power consumption (including losses) of the system must be kept in real time balance at any time. This is the most basic operating constraint.

[0103] (6)

[0104] In the formula: for t The curtailment power of renewable energy during the time period (kW) is a slack variable introduced to ensure that the model can still be solved under extreme conditions.

[0105] 2.2.2 Energy Storage System Operation Constraints: Energy storage systems (ESS) are key to realizing energy time shift, and their models need to accurately reflect their energy dynamics and operational limitations.

[0106] 1. Constraints on State of Charge (SOC) Evolution:

[0107] (7)

[0108] In the formula: for t State of charge of energy storage at the end of the time period; , These are the charging and discharging efficiencies of energy storage, respectively. The rated capacity of the energy storage (kWh).

[0109] 2. SOC upper and lower limit constraints:

[0110] (8)

[0111] In the formula: , The lower and upper limits of the energy storage SOC are usually determined by the battery characteristics to prevent overcharging and over-discharging.

[0112] 3. Charge and discharge power constraints:

[0113] (9)

[0114] (10)

[0115] In the formula: , The maximum charging and discharging power (kW) of the energy storage; , It is a 0-1 binary variable representing the charging and discharging states.

[0116] 4. Charge / discharge mutual exclusion constraint:

[0117] (11)

[0118] This constraint ensures that energy storage cannot be in a charging and discharging state simultaneously within the same time period, protecting the safety of the equipment.

[0119] 2.2.3 Power Constraints Between Power Grids:

[0120] Constrain the power exchange between the industrial park and the upstream power grid at the point of connection (PCC) to ensure operational safety and compliance with grid connection protocols.

[0121] (12)

[0122] (13)

[0123] In the formula: The maximum allowed switching power (kW) at the connection point (PCC); A binary variable of 0-1 representing the power purchase status, used to prevent unreasonable operations such as purchasing electricity from the grid and selling electricity to the grid at the same time.

[0124] 2.2.4 Output constraints of distributed generator sets:

[0125] (14)

[0126] In the formula: , For the first i The technical lower limit and upper limit (kW) of the output of the distributed generator set.

[0127] 2.2.5. Constraints on curtailment of renewable energy: Although it is undesirable, curtailment is permitted under specific operating scenarios (such as when there is excess power generation and energy storage is full) to ensure system safety.

[0128] (15)

[0129] This constraint limits the amount of abandoned electricity to no more than the available output of new energy sources.

[0130] The optimization layer recently developed a rigorous MILP problem through the aforementioned refined mathematical modeling. This model, while satisfying all physical and operational constraints, aims to minimize the total cost, including dynamic carbon costs, and generates a forward-looking scheduling plan, laying a strategic foundation for the park's economical and low-carbon operation. The model is solved using a high-speed mixed-integer programming solver to ensure that the global optimum is obtained within an acceptable timeframe.

[0131] 3. Real-time control layer modeling.

[0132] The real-time control layer is the core of the entire optimization framework, tasked with addressing the disconnect between day-ahead planning and actual operating conditions caused by forecast errors. Deployed at the edge, this layer possesses high-frequency data sensing, rapid computation, and low-latency response capabilities. It employs Model Predictive Control (MPC) to achieve rolling optimization based on the latest system state, addressing ultra-short-term fluctuations in renewable energy and load. Simultaneously, a reinforcement learning (RL) agent is introduced to dynamically adjust the carbon cost weights in the MPC objective function, overcoming the drawbacks of fixed control strategy parameters and forming an adaptive closed-loop control system capable of autonomous learning and adapting to complex operating environments.

[0133] 3.1 MPC Rolling Optimization Model: MPC is the execution engine of the real-time control layer. Its core idea is: at the beginning of each control cycle, based on the current measured state of the system, an ultra-short-term prediction model is used to optimize the system behavior over a finite future time domain (prediction domain), and the first control command in the optimized sequence is applied to the actual system. This process is then repeated in the next cycle. This rolling optimization and feedback correction mechanism makes it extremely robust to uncertainties.

[0134] 3.1.1 Objective Function: At each rolling optimization time step t MPC controller solves the future The objective function is to minimize the total cost over a given time period (prediction time domain). It inherits from the day-ahead optimization layer but introduces a key adaptive parameter—dynamic carbon weights. And calculations are performed based on ultra-short-term forecast data.

[0135] (16)

[0136] In the formula: t This is the current moment for scrolling optimization; k To predict the time period index within the time domain, ; To predict the length of the time domain; k | t Represents the moment t For the future k The predicted or planned value of a time-matter variable; For the current moment t The carbon cost weighting coefficient is a key parameter dynamically output by the RL agent, used to adjust the priority of low-carbon goals and economic goals in real time. For prediction k Time-of-use electricity trading cost (RMB); For prediction k Time-based maintenance costs (RMB); For prediction k Carbon cost per time period (RMB); For prediction k Time-based penalty cost (RMB).

[0137] The specific calculation models for each cost are as follows:

[0138] 1. Electricity trading costs:

[0139] (17)

[0140] In the formula: , In order to be in t Time prediction k The periodic electricity purchase / sale price (yuan / kWh) is usually based on the day-ahead planned value or short-term electricity price forecast; , In order to be in t Optimized in real time k Interaction power between the time period and the power grid (kW).

[0141] 2. Operation and maintenance costs:

[0142] (18)

[0143] In the formula: In order to be in tThe first time to optimize i A distributed power source in Power output (kW) during a given time period; In order to be in t Optimized energy storage at all times The charging and discharging power (kW) during the time period.

[0144] 3. Dynamic carbon cost:

[0145] (19)

[0146] In the formula: In order to be in t Real-time acquisition or prediction Real-time marginal carbon emission factor of the power grid (kgCO2 / kWh). This data comes from an authoritative database that is updated in real time and is a direct reflection of the dynamic carbon footprint at the real-time level.

[0147] 4. Cost of punishment:

[0148] (20)

[0149] In the formula: , , In order to be in t Time-to-time prediction obtained using ultra-short-term prediction models Time-of-use load, photovoltaic power, and wind power (kW).

[0150] 3.1.2 Constraints: The constraint form of MPC is basically the same as that of the day-ahead optimization layer, but there are two fundamental differences, which reflect its rolling implementation characteristics.

[0151] 1. Power balance constraints:

[0152] (twenty one)

[0153] 2. Energy storage system operation constraints:

[0154] SOC evolution constraints:

[0155] (twenty two)

[0156] SOC upper and lower bound constraints:

[0157] (twenty three)

[0158] Charging and discharging power and mutual exclusion constraints:

[0159] (twenty four)

[0160] (25)

[0161] (26)

[0162] Key difference 1: Initial state constraints

[0163] (27)

[0164] In the formula: for t The actual measured value of the stored energy state of charge at any given time. This constraint anchors the optimization problem to the actual state of the system and is the basis for implementing feedback correction.

[0165] 3. Grid interaction power, DG output and curtailment constraints: Same as the day-ahead layer, but based on real-time optimization variables and ultra-short-term forecasts.

[0166] 4. Key Difference Two: Rolling Implementation. After solving the above optimization problem, the entire prediction time domain is not executed. Instead of all instructions, it is:

[0167] (28)

[0168] In the formula: - In order to be in t In the optimal control sequence obtained by solving for each time interval, only the first time interval... The corresponding control commands.

[0169] The command is then sent to the physical system for execution. The next control cycle will then begin. With new measured values Using this as the initial state, we re-predict and optimize, and repeat this process continuously.

[0170] 3.2 Reinforcement Learning-Assisted Dynamic Carbon Weight Adjustment Strategy: Fixed Carbon Weights This invention addresses the complex and ever-changing operating environment. It introduces an RL agent whose core function is to act as a decision-maker beyond the MPC prediction time domain, adaptively adjusting the MPC objective function by evaluating the system's macroscopic operational performance. This enables real-time control to not only cope with short-term fluctuations, but also to approach long-term carbon efficiency targets.

[0171] 3.2.1 Model Construction: The weight adjustment problem is modeled as a Markov Decision Process (MDP).

[0172] 1. State State space design is a set of key indicators that can characterize the recent operating characteristics and resource status of the system.

[0173] (29)

[0174] In the formula: For the past period of time The average carbon intensity (kgCO2 / kWh) of the park within 4 hours reflects recent carbon efficiency performance. For the past period of time The average wind and solar penetration rate within the region characterizes the supply of renewable energy. for t The actual state of charge of the stored energy at any given time reflects the energy regulation potential of the system.

[0175] 2. Action Actions are the agent's decisions to adjust dynamic carbon cost weights.

[0176] (30)

[0177] In the formula: This is the carbon cost weighting adjustment amount; This is the preset adjustment step size.

[0178] The action space is designed with three discrete options: increase weights, decrease weights, or keep weights unchanged. The weight update formula is:

[0179] (31)

[0180] 3. Reward function The reward function is used to evaluate the quality of actions and guide the agent to learn the optimal strategy. Its design balances carbon efficiency goals with operational economy.

[0181] (32)

[0182] In the formula: for t Average carbon intensity of actual operation in the park during the period ; The carbon intensity target value that the park expects to achieve ; for t Total operating cost of the time-segment system (RMB); It is a small regularization coefficient used to avoid excessive operating costs while prioritizing carbon efficiency.

[0183] The significance of the reward function: the greater the deviation of the actual carbon intensity from the target value, the greater the penalty (negative reward); the higher the total operating cost, the greater the penalty. The goal of the agent is to maximize the long-term cumulative reward through learning, that is, to find a set of weight adjustment strategies that can stabilize the system's carbon intensity near the target value while keeping the operating cost within a reasonable range.

[0184] 3.2.2 Learning and Decision-Making Mechanism. The intelligent agent learns the optimal strategy by continuously interacting with the operating environment (initially a high-fidelity digital twin simulation model, and later the actual system). .

[0185] Offline training: Before system deployment, the RL agent is trained offline using historical running data or a large amount of simulation experience generated by digital twin models, so that it can obtain a preliminary effective weight adjustment strategy and ensure basic performance and security after going online.

[0186] Online decision-making and fine-tuning: After deployment, the trained agent makes decisions based on real-time status. Query strategy, output optimal action (i.e., weight adjustment amount), thereby dynamically setting the MPC. At the same time, it can be based on real operational feedback (rewards). Continue to fine-tune its strategies online to enable lifelong learning and adapt to the slow changes in system characteristics.

[0187] In summary, the real-time control layer, through the hybrid intelligence of MPC and RL, constitutes an advanced control architecture that combines rapid response capabilities with long-term strategic adaptability. MPC is responsible for high-frequency rolling optimization, accurately tracking day-ahead plans and suppressing short-term fluctuations; RL evaluates macro performance at a slower pace, dynamically adjusting the optimization direction of MPC. Working together, they ensure that the park's energy system, in actual operation and facing uncertainties, can still robustly and intelligently approach its long-term optimal carbon efficiency target.

[0188] 4. Solution method.

[0189] The two-stage optimization framework proposed in this invention focuses on solving two types of mathematical problems: mixed-integer linear programming problems at the current-day level and collaborative solutions of mixed-integer linear programming and reinforcement learning at the real-time level. The solution algorithms, input data requirements, and specific implementation procedures for these two types of problems will be described in detail below.

[0190] 4.1 Day-ahead Optimization Solution Method: The mathematical model of the day-ahead optimization layer aims to formulate a global scheduling plan for the park over the next 24 hours. The model includes continuous decision variables (such as equipment power and energy storage SOC) and binary integer decision variables (such as equipment start / stop status), and the objective function and all constraints are linear in form. Therefore, this optimization problem is reduced to a mixed-integer linear programming problem.

[0191] 4.1.1 Algorithm Selection

[0192] Problem type: Mixed Integer Linear Programming (MILP).

[0193] Mathematical form:

[0194] (33)

[0195] In the formula: A continuous variable vector, including ; A vector of integer variables, including ; These are the objective function coefficient vectors for continuous variables and integer variables, respectively; These are constraint matrices and vectors, corresponding to all linear constraints such as power balance and equipment operation limits.

[0196] Solution Algorithm: MILP problems are NP-hard in terms of computational complexity. However, modern commercial mathematical optimization solvers, based on advanced algorithms such as branch and bound and cutting plane methods, and incorporating heuristic strategies and large-scale preprocessing techniques, exhibit extremely high solution efficiency and robustness for such problems, typically finding a globally optimal solution or a high-quality feasible solution near the globally optimal solution within an acceptable timeframe. Therefore, this invention employs such a mature high-speed mixed-integer programming solver to solve the current-day optimization model.

[0197] 4.1.2 Input data and solution process.

[0198] Solving current MILP models requires high-precision prediction data as input. The complete solution process and data requirements are as follows: Figure 2 As shown.

[0199] Data preparation phase: Obtain the following prediction curves from the digital twin platform, with a time resolution typically of 15 minutes or 1 hour, covering the entire scheduling cycle. T (e.g., 24 hours): Photovoltaic output forecast curve Wind power output prediction curve Load demand forecast curve Time-of-use electricity price curve Grid dynamic marginal carbon emission factor prediction curve (Core innovation input)

[0200] Model instantiation phase: In a programming environment (such as Python + Gurobi interface, MATLAB + YALMIP), construct the MILP model as described above based on the input data. This includes: defining decision variables (continuous variables and 0-1 integer variables); setting the coefficients of each cost term in the objective function based on the input data; and adding all constraints, including power balance constraints, energy storage operation constraints, and grid interaction constraints.

[0201] Model Solving and Result Output Stage: The API interface of the selected mixed-integer programming high-speed solver is called to solve the constructed MILP model. After the solution is completed, the optimal solution is extracted and stored, and a day-ahead scheduling plan is generated, including: the optimal power purchase and sale plan. Energy storage system scheduling plan: Distributed generator set output plan; Flexible load start-up and shutdown plans, etc. The daily plan is issued to the real-time control layer as a benchmark for its rolling optimization.

[0202] 4.2 Real-time control solution method: such as Figure 3 As shown, solving the real-time control layer involves two parallel and coordinated processes: fast solving of the MPC rolling optimization problem and learning and decision-making by the RL agent. The former requires millisecond to second-level computation speed, while the latter is a continuous learning process.

[0203] 4.2.1 Solving the MPC Optimization Problem. The MPC model in the real-time control layer is also a MILP problem in mathematical essence, but its scale is much smaller than that of the current problem.

[0204] Problem type: Mixed Integer Linear Programming (MILP).

[0205] Scale characteristics: prediction time domain Np With shorter intervals (e.g., 12 five-minute intervals, or one hour), the number of decision variables and constraints is significantly reduced to meet the timeliness requirements of rapid solutions at the minute level.

[0206] Solution Algorithm: A high-speed mixed-integer programming solver is also employed. These solvers provide APIs and parameter configurations optimized for embedded systems and fast solution scenarios. During implementation, the solver kernel is embedded into a real-time control program deployed on edge computing nodes. By pre-setting optimal algorithm parameters (such as MIPGap and TimeLimit), the solution quality is guaranteed while the solution time is strictly limited, ensuring real-time control.

[0207] Input data: at each scroll optimization time point The MPC solver requires the following real-time and predicted data: measured system state values: (State of charge of energy storage) (Total load), etc.; Ultra-short-term forecast curve: Future Rainfall forecast for different time periods and load forecasting Real-time grid data: Real-time / ultra-short-term grid carbon emission factors And electricity price information; adaptive parameters: current carbon weight coefficients provided by the RL agent. .

[0208] 4.2.2 Reinforcement Learning Agent Learning and Decision Making: The goal of a reinforcement learning agent is to learn an optimal policy. π ∗ ( s ), to determine based on system status s t Intelligent adjustment of carbon weight .

[0209] 1. Algorithm Selection: Based on the complexity of the state and action space, different levels of RL algorithms can be selected:

[0210] Tabular approach (suitable for problems with a small state-action space).

[0211] Q-learning: A classic offline strategy algorithm that updates iteratively. Q Value table To learn the optimal action value function. Its update rule is:

[0212] (34)

[0213] In the formula: The learning rate controls how new information affects old information. Q The degree of influence of the value; The discount factor is used to weigh the importance of current rewards against future rewards. To perform the action The instant reward obtained afterward; For the next state The maximum expected future cumulative reward; SARSA: an online policy algorithm whose updates depend on the actual actions taken.

[0214] Deep reinforcement learning is suitable for problems with large or continuous state-action spaces.

[0215] Deep Q-Network (DQN): Uses deep neural networks to approximate the Q-value function. The parameters are This overcomes the curse of dimensionality in tabular methods. Its core techniques include experience replay and target networks.

[0216] Policy gradient methods (such as Actor-Critic): directly learn a parameterized policy function. It is suitable for continuous motion spaces.

[0217] 2. Training and Deployment Method: The training and deployment of RL agents adopt a phased strategy to ensure the security and stability of system operation.

[0218] Offline training phase: A high-fidelity simulation environment constructed using historical running data or a digital twin layer is used as the training environment. The agent explores and learns in this virtual environment through numerous rounds, continuously updating its Q-table or neural network weights through repeated trial and error, ultimately converging to a preliminarily effective policy. This avoids high-risk, inefficient exploration on real physical systems and accelerates the learning process.

[0219] Online application and fine-tuning phase: Applying the offline-trained strategy Deployed into the real-time control system.

[0220] Online decision-making: At each decision point (e.g., every hour), the agent makes decisions based on the current actual system state. Choose actions based on their strategy (i.e., weight adjustment amount) ).

[0221] Online learning: Rewards for agents based on feedback from real systems and new status They continue to update their strategies online at a relatively low learning rate, enabling lifelong learning and adaptation to slow environmental changes.

[0222] The constructed solution system combines the advantages of two core branches of artificial intelligence: mathematical programming and reinforcement learning. For deterministic optimization problems, the high efficiency of a mixed-integer programming high-speed solver is utilized to obtain an exact solution; for policy parameter adaptation problems under uncertain environments, the perception-learning-decision-making capabilities of reinforcement learning are leveraged. The two work together to ensure the advanced computational efficiency and decision-making intelligence of the entire two-stage closed-loop optimization framework.

[0223] Example 2

[0224] As one specific embodiment of the present invention, this embodiment constructs a typical low-carbon industrial park integrated energy system. The system configuration includes 500kW photovoltaic power generation, 300kW wind power generation, 400kW distributed gas turbine, and a 1000kWh electrochemical energy storage system. The load characteristics exhibit typical bi-peak features, with a base load of 400kW and a daily load fluctuation range of 150-650kW. The operating environment adopts a time-of-use pricing mechanism, with a peak-hour price of 0.9 yuan / kWh, a normal-hour price of 0.7 yuan / kWh, and an off-peak price of 0.4 yuan / kWh. The electricity sales price is 80% of the purchase price. In terms of carbon management, the target carbon intensity is set at 0.3 kgCO2 / kWh, the marginal carbon intensity of the grid fluctuates within the range of 0.3-0.8 kgCO2 / kWh, and the carbon price is 0.08 yuan / kgCO2. The optimization algorithm adopts a hybrid intelligent architecture combining day-ahead mixed integer linear programming and real-time model predictive control with reinforcement learning. The prediction time domain is set to 1 hour, and the RL state update frequency is once per hour.

[0225] like Figure 4 This embodiment demonstrates the input data configuration, which provides a real-world basis for optimization. The combination of highly fluctuating loads and renewable energy tests the system's regulation capabilities, while the introduction of dynamic carbon intensity and time-of-use pricing makes optimization more challenging but also creates conditions for tapping into carbon reduction potential. This configuration is typical and reasonable, ensuring the validity and comparability of subsequent optimization results.

[0226] Analysis of recent optimization results: Figure 5 The system's day-ahead optimization scheduling plan is demonstrated. During off-peak hours at night, the system purchases 150-200kW of electricity to fully utilize low-priced power. During morning and evening peak hours, the purchased power is reduced to 50-100kW to avoid high electricity prices. A small amount of electricity is sold during the midday photovoltaic peak, reflecting an economic optimization orientation. The energy storage system exhibits typical "valley charging, peak discharging" characteristics, with charging concentrated during off-peak hours at night and midday photovoltaic peaks, and discharging concentrated during peak electricity price periods. The State of Charge (SOC) safely cycles within the range of 20%-85%. Cost structure analysis shows a total cost of 4825 yuan, of which electricity purchase cost accounts for 65.3%, operation and maintenance cost 21.1%, and carbon cost 13.6%. The significant proportion of carbon cost highlights the crucial role of dynamic carbon cost in the objective function. These results validate the model's effective balance between economic efficiency and low carbon emissions. Through fine-grained scheduling in the time dimension, it not only reduces operating costs but also proactively avoids purchasing electricity during high-carbon periods, laying the foundation for deep carbon reduction.

[0227] Real-time control effect analysis: Figure 6The results demonstrate the operational effectiveness of the real-time control layer. Grid-interactive power tracking shows a correlation coefficient of 0.92 between real-time power purchases and day-ahead plans. Power purchases are reduced by 15-25 kW when photovoltaic output exceeds forecasts, and appropriately increased when load exceeds expectations. Real-time power sales increase by 10-20 kW compared to the plan during peak photovoltaic periods, showcasing the rolling optimization capability of the MPC. Energy storage SOC dynamic adjustment is excellent, with a root mean square tracking error of 0.032. Adaptive charging and discharging based on actual demand ensures energy storage safety. Regarding adaptive carbon weight adjustment, the RL agent increases the weight to 1.2-1.5 during peak carbon intensity and decreases it to 0.8-1.0 during troughs, with actions distributed as 8 increases, 6 decreases, and 10 holds, demonstrating the flexibility of intelligent decision-making. These results prove that the real-time layer effectively compensates for prediction errors. Through the collaboration of MPC and RL, robust system operation under uncertainty and long-term carbon efficiency convergence are achieved.

[0228] Carbon emission reduction effect assessment: Figure 7 A detailed assessment of carbon emission reduction effectiveness was provided. The average carbon intensity of the power grid was 0.52 kg CO2 / kWh, while the actual carbon intensity of the industrial park decreased to 0.28 kg CO2 / kWh, a reduction of 46.2%, which is lower than the target value of 0.3 kg CO2 / kWh, achieving a carbon efficiency of 106.7%, exceeding the target. Regarding total carbon emissions, the current optimized level is 1245 kg CO2, and real-time control further reduces it to 1085 kg CO2, achieving a carbon emission reduction of 160 kg CO2, a reduction of 12.9%, highlighting the added value of real-time carbon footprint tracking. The carbon emission reduction contribution decomposition shows that energy storage contributes the most to time-shifting (32.5%), followed by renewable energy utilization (28.5%), DG substitution (22.5%), and load adjustment (16.5%), indicating that the carbon time-shifting effect of the energy storage system is a key emission reduction factor. This analysis confirms the effectiveness of the dynamic carbon footprint method, achieving significant carbon emission reduction performance through multi-resource synergy.

[0229] Economic analysis: such as Figure 8As shown, comparative analysis reveals the improved economic performance of the system. The total cost of optimization was 4825 yuan, which was reduced to 4310 yuan by the real-time control layer, achieving a cost saving of 515 yuan, or 10.7%. This is attributed to the rolling optimization of MPC and the adaptive adjustment of RL. In terms of cost structure, real-time control reduced the proportion of electricity purchase cost from 65.3% to 61.2%, and the proportion of carbon cost from 13.6% to 11.8%, while maintaining stable operation and maintenance costs. This indicates that the system has been highly effective in reducing dependence on external electricity purchases and carbon costs. The renewable energy utilization rate increased from 94.2% to 98.5%, an increase of 4.3 percentage points. Precise scheduling reduced wind and solar curtailment and improved the efficiency of clean energy utilization. This analysis confirms the economic advantages of the two-stage closed-loop optimization, which not only reduces carbon costs through dynamic carbon management but also optimizes energy procurement and allocation through real-time control, maximizing overall benefits.

[0230] Comprehensive evaluation of system performance indicators: Figure 9 The learning process of the reinforcement learning agent is demonstrated. RL converges within 8-12 hours, maintaining good policy stability and adaptability. The learning curve shows that as training progresses, the cumulative reward gradually increases and stabilizes, indicating that the agent has successfully learned to adjust carbon weights based on system states (such as average carbon intensity, wind and solar penetration, and SOC). After deployment, the agent can fine-tune online to adapt to changes in the operating environment, such as increasing weights during peak carbon intensity to enhance carbon reduction and decreasing weights during troughs to balance economic efficiency. This process verifies the effectiveness of the RL-MPC hybrid architecture. RL, acting as a "strategist," compensates for the limited time domain of MPC, enabling the system to not only cope with short-term fluctuations but also approach long-term carbon efficiency goals, thus improving the overall intelligence level and robustness.

[0231] Compared with traditional static carbon accounting methods, the carbon cost calculation based on dynamic marginal carbon intensity (0.3-0.8 kg CO2 / kWh) proposed in this invention reduces the error by 35-40% compared with the traditional average carbon intensity (0.5 kg CO2 / kWh) method, and can actively avoid purchasing electricity during high-carbon periods, improving the carbon reduction effect by more than 25%. Compared with single optimization control methods, the robustness of two-stage closed-loop optimization is significantly improved, overcoming the defect of single day-ahead optimization causing a 15-20% performance drop due to prediction errors. At the same time, it solves the problem of single real-time control lacking a long-term optimization perspective, improving the overall performance by 10-15%. Adaptive control can maintain stable performance under different operating scenarios, which is better than fixed parameter control.

[0232] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.

Claims

1. A low-carbon park control method based on digital twin and dynamic carbon footprint, characterized by the steps of include: Construct a digital twin simulation model that maps to the physical entities of the low-carbon park; In the day-ahead optimization phase, with the goal of minimizing the total daily operating cost within the scheduling cycle, a day-ahead operation plan for the park's energy system is formulated as the baseline and operating boundary for real-time control and rolling optimization. The total daily operating cost includes the dynamic carbon cost, which converts carbon emissions into economic costs based on the dynamic marginal carbon emission factor of the power grid. In the real-time control phase, during each control cycle, model predictive control is used to perform rolling optimization of the day-ahead operating plan based on the predicted curve of the system state. The first control command in the optimized sequence obtained in the current control cycle is applied to the actual system. The objective function of model predictive control introduces dynamic carbon weights based on the day-ahead optimized objective function, and a reinforcement learning agent is used to dynamically adjust the dynamic carbon cost weights in the objective function of model predictive control. The problem of adjusting the dynamic carbon cost weights by reinforcement learning is modeled as follows: Reinforcement learning's state-space representation of system operational characteristics and resource status: In the formula: Set time for the past Average carbon intensity of the inner park; Set time for the past Average wind and solar penetration rate within the area; for t The actual state of charge of stored energy at all times. The actions of a reinforcement learning agent are decisions about adjusting dynamic carbon cost weights: In the formula: This is the carbon cost weighting adjustment amount; The preset adjustment step size; The action space is designed with three discrete options: increase weights, decrease weights, or keep the weights constant. The reward function of reinforcement learning balances carbon efficiency goals with operational economics: In the formula: for t The average carbon intensity of the park during the actual operation of the time period; The target carbon intensity value that the park hopes to achieve; for t Total operating cost of the time-slot system; This is the regularization coefficient.

2. The low-carbon industrial park control method based on digital twin and dynamic carbon footprint according to claim 1, characterized in that, The total system operating cost in the objective function of the current optimization phase includes: Electricity trading costs: the economic expenditures and revenues generated by the park through electricity trading with the power grid; Operation and maintenance costs, including losses and maintenance expenses incurred during the operation of distributed power generation equipment and energy storage systems within the park; Dynamic carbon cost, which introduces the dynamic marginal carbon emission factor of the power grid, transforms carbon emissions into economic costs; The penalty cost takes into account the total load power, energy storage charging and discharging power, and photovoltaic and wind power output, and penalizes the power imbalance of the park's energy system.

3. The low-carbon industrial park control method based on digital twin and dynamic carbon footprint according to claim 1, characterized in that, The constraints of the day-ahead optimization phase include: Power balance constraints ensure that the power generation and power consumption of the park's energy system are kept in real time. Energy storage system operation constraints include state of charge evolution constraints, SOC upper and lower limit constraints, charging and discharging power constraints, and charging and discharging mutual exclusion constraints; Power exchange constraints between the power grid and the upper-level power grid at the connection point limit the exchange power between the park and the upper-level power grid to ensure operational safety and compliance with the grid connection agreement; Distributed generator set output constraints: the generator set output is located between the lower and upper technical limits. The curtailment of renewable energy must not exceed the available output of renewable energy.

4. The low-carbon industrial park control method based on digital twin and dynamic carbon footprint according to claim 1, characterized in that, The real-time control, based on the day-ahead optimization constraints, introduces an initial state constraint into the energy storage system operation constraints, constraining the initial state of the park's energy system to be equal to the actual measured value at the corresponding time. The optimal control sequence is obtained by solving the model predictive control, and the control command for the first time period is issued to the physical system for execution. In the next control cycle, the new measured values ​​are used as the initial state to re-predict and optimize.

5. The low-carbon industrial park control method based on digital twin and dynamic carbon footprint according to claim 1, characterized in that, The reinforcement learning agent interacts with the digital twin simulation model in the early stages of deployment and with the actual park energy system in the later stages to learn the optimal strategy.

6. A low-carbon industrial park control system based on digital twin and dynamic carbon footprint, characterized in that, The system implements the low-carbon park control method based on digital twin and dynamic carbon footprint as described in any one of claims 1-5. The system is composed of a digital twin layer, a day-ahead optimization layer, a real-time control layer and a physical system layer. The digital twin layer integrates a simulation model, a multi-timescale prediction module, a dynamic carbon footprint database, and an optimization algorithm library, providing full-domain data perception and decision support for the day-ahead optimization layer and the real-time control layer; The day-ahead optimization layer is based on high-precision wind and solar load forecasting, and constructs a mixed-integer linear programming model with multiple objectives of economy, low carbon emissions, and reliability to generate day-ahead scheduling plans. The real-time control layer receives the day-ahead scheduling plan through the edge computing device and uses model predictive control method to continuously correct the scheduling instructions; the instructions are then sent to the physical system layer for execution. The physical system layer uploads state data to the digital twin layer in real time to update the virtual model state. The real-time control layer uses the updated data to optimize and provides feedback on the execution results. The digital twin layer uses accumulated data to train and optimize algorithms.

7. A low-carbon industrial park control system based on digital twin and dynamic carbon footprint as described in claim 6, characterized in that, The day-ahead optimization layer is responsible for formulating the next day's scheduling plan for the park's energy system based on medium-term forecast information. Its functions include: Multi-timescale forecast data fusion integrates meteorological and production planning information to generate high-precision forecast curves for wind and solar power output, load demand, time-of-use electricity prices, and dynamic marginal carbon emission factors of the power grid for the next day. A multi-objective optimization scheduling model is solved, and a planning model is constructed and solved with the goal of minimizing the total daily operating cost. The total daily operating cost incorporates the dynamic carbon cost, which converts carbon emissions into economic costs based on the dynamic marginal carbon emission factor of the power grid. The plan is issued and the boundary is set. The day-ahead scheduling plan obtained from the solution is issued to the real-time control layer as the baseline and operating boundary for rolling optimization.

8. A low-carbon industrial park control system based on digital twin and dynamic carbon footprint according to claim 6, characterized in that, The real-time control layer is deployed on the edge computing nodes of the park, and uses high-frequency rolling optimization and adaptive adjustment to compensate for the deviation between the current plan and the actual operating status caused by prediction errors. The functional modules include: High-frequency data sensing and analysis: Decompose the total load of the park's energy system and identify the operating status of equipment; Ultra-short-term forecasting module: performs rolling minute-level forecasts of future ultra-short-term wind and solar power output and load demand; The dynamic carbon footprint real-time tracking engine accesses the dynamic marginal carbon emission factor of the power grid in real time via API, and combines the interaction power between the park and the power grid to calculate the total carbon flow of the park and the carbon footprint of key equipment at a frequency of minutes. The core of the hybrid intelligent optimization control system is model predictive control, which acts as the executor of rolling optimization. It executes rolling optimization on a minute-by-minute cycle, receiving the daily plan as a benchmark and combining the latest ultra-short-term forecasts, the real-time status of the park's energy system, and carbon weight coefficients dynamically provided by a reinforcement learning agent. It solves the optimization problem within a shortened prediction time domain, generates real-time control commands, and issues them for execution. The reinforcement learning agent operates on a timescale longer than the rolling cycle of model predictive control. Through interaction with the environment, it autonomously learns and dynamically adjusts the carbon weight coefficients in the objective function of model predictive control.

9. A low-carbon industrial park control system based on digital twin and dynamic carbon footprint according to claim 6, characterized in that, The digital twin layer, as the driving engine for global perception and closed-loop optimization, includes the following functions: High-fidelity multi-dimensional modeling is used to construct a virtual model that maps to the physical entities of the park's energy system, including a refined model at the equipment level and a network topology model at the system level. The data fusion and governance center integrates and manages real-time data, historical data, and external data from the physical park's energy system. The algorithm model library and simulation sandbox encapsulate and manage prediction, optimization and reinforcement learning training algorithms, and provide a simulation environment for testing new policies, extrapolating extreme scenarios and offline training of reinforcement learning agents. Visual analysis and decision support enable panoramic visualization and dynamic tracking of energy flow, carbon flow, and economic indicators.

Citation Information

Patent Citations

  • CN120255459A

  • CN121414236A