Industrial park multi-load electricity-carbon collaborative management decision-making method based on hierarchical safety reinforcement learning
By constructing a dynamic power-carbon flow coupling model using a hierarchical security reinforcement learning method, the problems of coordinated scheduling of multiple types of loads and power-carbon flow coupling management in industrial parks were solved, enabling economical, safe, and low-carbon operation of industrial parks and improving energy utilization efficiency and carbon emission control capabilities.
Patent Information
- Application Number
- CN202610211047.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-13
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies are insufficient to address the challenges of coordinated scheduling of various load types in industrial parks, power-carbon flow coupling management, and real-time performance and security of scheduling under dynamic operating conditions. They also lack systematic collaborative optimization schemes, resulting in low energy utilization efficiency, difficulties in carbon emission traceability and control, and high safety risks.
A hierarchical security reinforcement learning approach is adopted to construct a dynamic power-carbon flow coupling model. Combining power flow calculation and carbon emission flow theory, a carbon flow entropy index is introduced, and a dual-timescale hierarchical reinforcement learning decision architecture is designed. An enhanced priority experience replay mechanism and a constraint violation cost embedding reward function are introduced to form a closed-loop control system of perception-modeling-decision-execution-feedback.
It has enabled the economical, safe, and low-carbon coordinated operation of multiple loads and energy systems in industrial parks, improved the coordinated scheduling capability of multiple types of loads, enhanced the pertinence of carbon emission reduction measures and the dynamic adaptability of the system, and ensured the real-time and security of scheduling decisions.
Smart Images

Figure CN122022368A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the technical fields of industrial energy optimization, artificial intelligence and low-carbon environmental protection. Specifically, it relates to a decision-making method for multi-load electricity-carbon collaborative management in industrial parks based on hierarchical security reinforcement learning. It includes technical directions such as multi-entity energy collaborative scheduling in industrial parks, coupled management of electricity-carbon emission flows, and hierarchical security reinforcement learning algorithm design. It is particularly suitable for complex industrial park energy management scenarios that integrate intelligent manufacturing units, data centers, vehicle-to-grid (V2G) charging stations, photovoltaic power generation systems and energy storage systems. Background Technology
[0002] Against the backdrop of the comprehensive advancement of the "dual-carbon" strategy, industrial parks, as the core carriers of manufacturing clusters, are a crucial link in achieving carbon peaking and carbon neutrality in the industrial sector. Industrial parks integrate technologies such as the Internet of Things, big data, and artificial intelligence, forming a complex cyber-physical energy system encompassing "source-grid-load-storage." Their energy structure covers diverse components, including renewable energy sources such as photovoltaic power generation, mains grid power supply, energy storage systems, manufacturing loads, data center computing loads, and electric vehicle charging loads, playing an irreplaceable role in promoting the green transformation of manufacturing.
[0003] However, the low-carbon operation and energy management of industrial parks currently face multi-dimensional technical challenges. Traditional scheduling methods and control systems are no longer suitable for the complex energy ecosystem and dynamic operating conditions. Specific problems are reflected in the following aspects:
[0004] The lack of a multi-type load coordination and dispatch mechanism results in low energy utilization efficiency.
[0005] Industrial parks exhibit significant heterogeneous characteristics of "slow load" and "fast load": Slow load is centered on the production load of manufacturing units, characterized by stable energy consumption, long scheduling cycles, and strict process constraints. Dynamic adjustments to production plans (such as the insertion of emergency orders, changes in product types, and adjustments to delivery dates) can easily disrupt the original resource scheduling scheme. Fast load includes data center computing load (such as tasks like load forecasting, production optimization, and fault diagnosis) and charging station charging load, characterized by frequent energy consumption fluctuations, fast response speed, and large adjustment potential. Its computing task allocation and charging / discharging power adjustment are easily affected by real-time demand.
[0006] Traditional scheduling methods often focus on single-objective optimization for a single load type, such as minimizing the energy consumption of manufacturing loads or optimizing the completion time of charging loads. They lack a collaborative scheduling mechanism across load types, making it difficult to balance the stability requirements of "slow loads" with the flexibility requirements of "fast loads," resulting in low renewable energy absorption rates, large peak-valley differences in grid loads, and low efficiency in energy resource allocation.
[0007] The inadequate power-carbon flow coupling management system makes carbon emission traceability and control difficult.
[0008] Carbon emissions in industrial parks are deeply intertwined with electricity flow. As a virtual network flow coexisting with active power flow, the distribution of carbon emissions is influenced by multiple factors, including the carbon emission intensity of power generation units, grid topology, line impedance, and load distribution. However, current technologies have not established a dynamic coupling model between electricity flow and carbon emission flow. They can only achieve macroscopic statistics on carbon emissions and cannot quantify the carbon footprint of each node, load, and branch, making it difficult to achieve accurate tracking and dynamic management of carbon emissions.
[0009] Meanwhile, existing studies are mostly based on static or quasi-static assumptions to analyze carbon emission flows, without fully considering the impact of factors such as fluctuations in renewable energy output, dynamic changes in load, and adjustments in energy price signals on the distribution of carbon emission flows. This results in a lack of targeted carbon reduction measures, making it impossible to achieve efficient utilization of carbon allowances while ensuring the economic operation of industrial parks, and making it difficult to balance operating costs and carbon reduction targets.
[0010] The real-time performance and security of scheduling under dynamic operating conditions are insufficient.
[0011] Industrial parks face multiple dynamic disturbances in their energy operations, including the intermittency and volatility of photovoltaic power generation, sudden changes in production tasks, randomness in charging demand, and dynamic changes in the power grid's operating status. These factors place extremely high demands on the real-time performance of scheduling decisions. Traditional optimization scheduling methods (such as integer programming and heuristic algorithms) are computationally complex, time-consuming, and lack timeliness, making it difficult to quickly generate real-time control decisions.
[0012] In recent years, reinforcement learning has been gradually applied to energy dispatching in industrial parks. However, existing multi-agent reinforcement learning methods are prone to parameter drift during training and operation. Erroneous actions and negative reward experiences generated during upper and lower layer policy training cannot be effectively corrected, leading to the breach of grid security constraints (such as node voltage constraints, branch current constraints, power flow constraints, and carbon quota constraints), and even grid operation failures, posing serious safety risks. Furthermore, existing reinforcement learning algorithms do not fully consider decision-making interactions across multiple time scales, making it difficult to achieve synergy between global optimization and local adjustment, resulting in poor policy robustness and adaptability.
[0013] Existing technological research has limitations and lacks systematic collaborative optimization schemes.
[0014] Current research largely focuses on optimizing single energy systems or improving single technological aspects, failing to establish a comprehensive electricity-carbon coordinated control system covering the entire process from "data perception to modeling and analysis, decision optimization, and execution feedback." Some studies only focus on the coordination of electricity generation, grid, load, and storage, without incorporating carbon emission flow control objectives; some studies only achieve static analysis of carbon emission flows without integrating them with real-time dispatch decisions; some studies only optimize single-type load dispatch without integrating the regulation potential of multiple load types such as manufacturing, computing, and charging; and some studies only design basic reinforcement learning frameworks without designing hierarchical decision-making systems and security enhancement mechanisms tailored to the heterogeneous characteristics of industrial parks.
[0015] In summary, existing technologies are insufficient to address core issues in the low-carbon operation of industrial parks, such as multi-load coordination, precise carbon flow control, and dynamic safety scheduling. There is an urgent need to develop an intelligent control method that takes into account multi-load coordinated scheduling, dynamic coupling of power and carbon flow, multi-timescale decision-making, and grid safety constraints, in order to construct a systematic power-carbon coordinated control system for industrial parks and achieve economic, safe, and low-carbon operation of industrial parks. Summary of the Invention
[0016] To address the shortcomings and deficiencies of existing technologies, this invention provides a decision-making method and system for multi-load power-carbon collaborative management in industrial parks based on hierarchical security reinforcement learning. This scheme constructs a dynamic power-carbon flow coupling model, combines power flow calculation and carbon emission flow theory to calculate the carbon potential and carbon flow rate of distribution network nodes, and introduces a carbon flow entropy index to quantify carbon emission distribution characteristics. The carbon flow entropy is used as a component of the reinforcement learning reward function to guide the scheduling strategy towards a more uniform carbon emission distribution. Addressing the heterogeneous characteristics of slow and fast loads in industrial parks, a dual-timescale hierarchical reinforcement learning decision architecture is adopted. The upper-level decision module generates global scheduling strategies on an hourly basis, while the lower-level decision module generates local adjustment strategies on a minute-based basis, achieving multi-timescale collaborative optimization. The decision system introduces an enhanced priority experience replay mechanism, prioritizing experience samples that lead to violations of grid constraints or carbon quota constraints, thus correcting the parameter drift problem in multi-agent reinforcement learning. Simultaneously, grid security constraints and carbon quota constraints are quantified as constraint violation costs and embedded in the reward function. Through constraint violation cost calculation and reward function correction mechanisms, the security of scheduling decisions is ensured. Based on the execution feedback data, the model parameters and training strategies are updated online, forming a closed-loop management and control system of perception-modeling-decision-execution-feedback, so as to realize the economic, safe and low-carbon coordinated control of multiple loads and energy systems in industrial parks.
[0017] The specific technical solution adopted by this invention to solve its technical problem is as follows:
[0018] A decision-making method for multi-load electricity and carbon collaborative management in industrial parks based on hierarchical security reinforcement learning includes:
[0019] Based on multi-source operation data of industrial parks, combined with power flow calculation and carbon emission flow theory, the carbon potential and carbon flow rate of distribution network nodes are calculated. The carbon flow entropy index is introduced to quantify the carbon emission distribution characteristics and is used as a component of the reward function of the lower-level decision module. At the same time, a multi-load collaborative energy consumption model is established, and a dynamic power-carbon flow coupling model is constructed.
[0020] Based on the dynamic power-carbon flow coupling model, a hierarchical reinforcement learning decision-making system is constructed, adopting a dual-timescale architecture with a first decision cycle and a second decision cycle, where the duration of the first decision cycle is longer than that of the second decision cycle. The upper-level decision module generates a global scheduling strategy for the main grid's power purchase and renewable energy generation output, while the lower-level decision module generates local adjustment strategies for slow load, fast load, and energy storage systems in industrial parks. The decision-making system introduces an enhanced priority experience replay mechanism to correct the parameter drift problem in multi-agent reinforcement learning, and embeds grid security constraints and carbon quota constraints into the reward function to quantify the cost of constraint violations.
[0021] The global scheduling strategy and local adjustment strategy are converted into scheduling instructions and issued to the execution unit to realize the coordinated control of multiple loads and energy systems in the industrial park; the operation of the park's power grid and the status of carbon emission flow are monitored in real time, and the calculation parameters of the dynamic power-carbon flow coupling model and the training strategy of the hierarchical reinforcement learning decision system are updated based on the monitored feedback data to complete the iterative optimization of the model and strategy.
[0022] Furthermore, the multi-source operation data of the industrial park includes power parameters, multi-type load operation parameters, environmental meteorological parameters, and economic carbon parameters. The multi-source operation data needs to be time-synchronized, outlier removal, missing value completion, and standardized preprocessing before being input into the dynamic power-carbon flow coupling model. The standardized preprocessing is to normalize the original data of different dimensions to the [0,1] interval.
[0023] Furthermore, the node carbon potential is the carbon emission on the generation side corresponding to a unit of electricity consumption at the node. The carbon potential of a single node is obtained by dividing the active power flux of the node by the carbon flow density and power flow of all branches flowing into the node, the power injected by the node, and the carbon emission intensity of the corresponding power generation unit. The carbon flow rate includes the load carbon flow rate, the branch carbon flow rate, and the total system carbon flow rate. The load carbon flow rate is obtained by multiplying the load distribution parameters by the corresponding node carbon potential. The branch carbon flow rate is calculated based on the power flow distribution of the distribution network and the node carbon potential. The carbon flow entropy index is calculated based on the information entropy theory. It is obtained by multiplying the proportion of each node's carbon emissions to the total system carbon emissions by the natural logarithm of that proportion and then summing the calculation results for all nodes.
[0024] Furthermore, the slow load is the manufacturing production load, and the fast load includes the data center computing load and the vehicle-to-grid charging station charging load; the first decision cycle is on the hourly level, and the second decision cycle is on the minute level; the upper-level decision result serves as the environmental condition for the lower-level decision to achieve state transmission, and the upper-level reward function integrates all lower-level reward values within the corresponding time period to achieve reward fusion; the operation of the energy storage system must meet the upper and lower limits of the state of charge and the range of charging and discharging power.
[0025] Furthermore, the enhanced priority experience replay mechanism is implemented as follows: the sample priority is calculated based on the temporal difference error of the experience samples, and the larger the temporal difference error, the higher the sample priority; samples are selected from the experience replay buffer using a weighted random sampling method based on the priority, while importance sampling weights are introduced to correct the training bias caused by non-uniform sampling; the erroneous actions—negative reward experiences—that are generated during the training process and cause violations of grid constraints or carbon quota constraints are set to the highest priority, and the reinforcement learning network is forced to repeatedly learn such experiences to correct the policy network parameters.
[0026] Furthermore, the power grid security constraints include node voltage constraints, branch current constraints, and power flow constraints; the constraint violation cost is the cumulative value of the degree of violation of node voltage constraints, branch current constraints, and power flow constraints; wherein the degree of violation of node voltage constraints is the sum of the value by which the actual node voltage exceeds the lower voltage limit and the value by which the upper voltage limit exceeds the actual node voltage, and only the non-negative part is calculated; the degree of violation of branch current constraints is the result of the ratio of the actual branch current to the rated branch current minus 1, and only the non-negative part is calculated; the degree of violation of power flow constraints is the result of the ratio of the actual branch apparent power to the rated branch apparent power minus 1, and only the non-negative part is calculated.
[0027] Furthermore, the specific method of embedding the power grid security constraints and carbon quota constraints into the reward function is as follows: the constraint violation cost is introduced into the original reward function to obtain the modified reward function. The modified reward function is obtained by subtracting the product of the constraint violation cost and the preset penalty coefficient from the original reward value, and then adding the constraint satisfaction reward. When the constraint violation cost is zero, the preset constraint satisfaction reward is obtained; when the constraint violation cost is not zero, there is no such reward.
[0028] Furthermore, the hierarchical reinforcement learning decision-making system is constructed based on a multi-agent, dual-delay deep deterministic policy gradient framework combined with a hierarchical reinforcement learning algorithm, comprising upper and lower layers of actor-commentator networks; the target network parameters are updated through a soft update mechanism.
[0029] Furthermore, the feedback data from the monitoring includes the actual operating power of each execution unit, the actual voltage of the distribution network nodes, the actual power flow of the branches, the actual state of charge of the energy storage system, the actual monitoring parameters of the carbon emission flow, and the actual operating cost and carbon emissions of the industrial park; the calculation parameters of the dynamic power-carbon flow coupling model are updated online, and the training strategy of the hierarchical reinforcement learning decision system is updated to add the real-time running state-action-reward experience to the experience replay buffer and train the reinforcement learning network online.
[0030] Furthermore, a multi-load electric-carbon collaborative management decision-making system for industrial parks based on hierarchical security reinforcement learning is provided to implement the method described above. The system includes a multi-source information sensing module, an electric-carbon flow coupling modeling module, a hierarchical security reinforcement learning decision-making module, and an execution and feedback closed-loop module. These modules interact and transmit instructions through standardized data interfaces. The multi-source information sensing module collects and preprocesses multi-source operational data from the industrial park. The electric-carbon flow coupling modeling module constructs a dynamic electric-carbon flow coupling model and calculates node carbon potential, carbon flow rate, and carbon flow entropy, using carbon flow entropy as a component of the reward function for the lower-level decision-making module. The hierarchical security reinforcement learning decision-making module constructs a hierarchical reinforcement learning decision-making system and generates global scheduling strategies and local adjustment strategies. The execution and feedback closed-loop module issues scheduling instructions, monitors operational status, collects feedback data, and performs iterative optimization of the model and strategies.
[0031] And a computer device including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the method described above.
[0032] A non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described above.
[0033] Compared to existing technologies, this invention and its preferred solution, by constructing a dynamic power-carbon flow coupling model, deeply binds carbon emission flows with power flow, achieving precise traceability and dynamic control of carbon emissions in industrial parks. This solves the problems of imperfect power-carbon flow coupling management systems and difficulties in carbon emission traceability and control in existing technologies. Addressing the heterogeneous characteristics of slow and fast loads in industrial parks, a dual-timescale hierarchical reinforcement learning decision architecture is adopted, achieving coordinated optimization of global scheduling and local adjustment. This balances the stability requirements of slow loads with the flexibility requirements of fast loads, improving the collaborative scheduling capability of multiple load types. By incorporating carbon flow entropy as a component of the reinforcement learning reward function, the scheduling strategy is guided towards a more uniform carbon emission distribution, enhancing the targeting of carbon reduction measures. An enhanced priority experience replay mechanism is introduced, prioritizing experience samples that violate grid constraints or carbon quota constraints, effectively correcting the parameter drift problem in multi-agent reinforcement learning and improving the safety and robustness of the decision system. By quantifying grid security constraints and carbon quota constraints into violation costs and embedding them into a reward function, a soft embedding of security constraints is achieved. This avoids the computational complexity and feasibility issues of traditional hard constraint processing methods, ensuring the real-time performance and security of scheduling decisions. A closed-loop iterative optimization mechanism based on execution feedback data enables online updates of model parameters and training strategies, enhancing the system's adaptability to dynamic operating conditions. Ultimately, a grid-carbon coordinated control system covering the entire process of data perception, modeling and analysis, decision optimization, and execution feedback is formed, enabling economical, safe, and low-carbon coordinated operation of multiple loads and energy systems in industrial parks. Attached Figure Description
[0034] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:
[0035] Figure 1 This is a schematic diagram of the test environment for an embodiment of the present invention;
[0036] Figure 2 This is a diagram of a hierarchical security reinforcement learning decision-making framework according to an embodiment of the present invention;
[0037] Figure 3 This is a time-series diagram showing the dynamic changes in the 24-hour load power and corresponding total carbon flow of the industrial park manufacturing unit and data center, as described in an embodiment of the present invention.
[0038] Figure 4 This is a timing diagram of multi-load collaborative scheduling under typical operating conditions in an embodiment of the present invention;
[0039] Figure 5 This is a comparison chart of the training performance of the algorithms in embodiments of the present invention;
[0040] Figure 6 This is a general framework diagram of an embodiment of the present invention. Detailed Implementation
[0041] To make the features and advantages of the present invention more apparent and understandable, specific embodiments are described below in detail:
[0042] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.
[0043] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this application. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0044] Under the "dual-carbon" strategy, industrial parks, as core carriers of manufacturing clusters, face key technological challenges such as the lack of multi-type load coordination and dispatch mechanisms, an imperfect power-carbon flow coupling management system, and insufficient real-time performance and security of dispatch under dynamic operating conditions. Traditional dispatch methods are difficult to adapt to complex energy ecosystems and dynamic operating conditions, and existing technologies lack systematic collaborative optimization solutions.
[0045] To address the challenges posed by the diverse loads (manufacturing, computing, and charging) within complex industrial parks, including significant differences in spatiotemporal distribution, lack of collaborative scheduling mechanisms, difficulty in tracing heterogeneous carbon emission flows, and weak grid security constraints under dynamic operating conditions, this invention constructs a closed-loop management and control system encompassing perception, modeling, decision-making, execution, and feedback.
[0046] This solution relies on a collaborative control system comprising a multi-source information sensing module, a power-carbon flow coupling modeling module, a hierarchical security reinforcement learning decision-making module, and an execution and feedback closed-loop module, specifically including:
[0047] 1. Multi-source heterogeneous data end-to-end perception: Collect power, load and environmental parameters and perform spatiotemporal synchronization and standardized preprocessing;
[0048] 2. Dynamic power-carbon flow coupling modeling: Based on carbon emission flow theory and power flow analysis, the carbon potential and carbon flow rate of nodes are calculated, and the carbon emission flow entropy index is introduced to quantify the carbon emission distribution characteristics.
[0049] 3. Hierarchical Security Reinforcement Learning Decision-Making: Addressing the energy consumption characteristics and spatiotemporal distribution differences between "slow loads" (manufacturing production loads) and "fast loads" (computing loads, charging loads) within industrial parks, a mathematical model deeply coupled with the electricity-carbon flow is constructed. A hierarchical security reinforcement learning decision-making system based on Multi-Agent Dual-Delay Deep Deterministic Policy Gradient (MATD3) is designed to achieve coordinated optimization scheduling of electricity and carbon at multiple time scales. As a preferred embodiment, the specific method of combining the HIRO hierarchical reinforcement learning algorithm with the MATD3 framework is as follows: The hierarchical architecture of HIRO (upper-layer global decision-lower-layer local decision-making) is deeply integrated with the MATD3 multi-agent actor-critic network. Independent MATD3 actor-critic networks are constructed for the upper and lower layers respectively. The upper-layer network generates global scheduling actions, and the lower-layer network generates local adjustment actions within the lower-layer state space. Coordinated training and decision-making of the two-layer network are achieved through state transfer and reward fusion.
[0050] 4. Execution and feedback closed loop: The scheduling manufacturing unit, data center and energy storage equipment operate in coordination, and online correction and model update are achieved based on operational performance feedback.
[0051] The core innovation of this invention lies in:
[0052] Construct a dynamic power-carbon flow coupling model to achieve precise traceability and dynamic control of carbon emissions;
[0053] Design a multi-timescale collaborative scheduling mechanism to balance the stability of slow loads and the flexibility of fast loads;
[0054] An improved hierarchical security reinforcement learning algorithm is proposed, which introduces an enhanced priority experience replay mechanism to solve the parameter drift problem, and embeds grid security and carbon quota constraints into the reward function.
[0055] Establish a multi-objective optimization system that takes into account the constraints of power grid safety operation, the minimization of operating costs, and the carbon emission reduction target.
[0056] This invention enables multi-load coordination and precise carbon flow control, effectively solving the safety risks caused by reinforcement learning parameter drift, significantly improving the energy utilization efficiency, low-carbon operation level and power grid operation safety of industrial parks, providing core technical support for the digital, intelligent and green transformation of industrial parks, and can be widely applied to energy management systems in various industrial cluster scenarios such as machinery manufacturing, new energy industrial parks and smart industrial parks.
[0057] To address this, this invention provides a method and system for coordinated power-carbon control in industrial parks based on hierarchical security reinforcement learning. It constructs a four-level core architecture consisting of a multi-source information sensing module, a power-carbon flow coupling modeling module, a hierarchical security reinforcement learning decision-making module, and an execution and feedback closed-loop module. This achieves intelligent and integrated management and control of the entire process from data acquisition to control execution. The specific technical solution is as follows:
[0058] 1. Overall Architecture Design
[0059] The overall system architecture of this invention is a closed-loop management and control system. Each module interacts with the other through standardized data interfaces to exchange data and transmit commands. The overall process is as follows: a multi-source information sensing module collects multi-dimensional data on electricity, load, environment, and economy from the industrial park and performs preprocessing; a power-carbon flow coupling modeling module analyzes the power flow distribution based on the preprocessed data, calculates node carbon potential, carbon flow rate, and carbon emission flow entropy, and establishes a multi-load collaborative power-carbon coupling model; a hierarchical safety reinforcement learning decision-making module generates upper-level global optimization decisions and lower-level local real-time adjustment decisions based on the power-carbon coupling model and multi-timescale data; and an execution and feedback closed-loop module issues decision commands to each execution unit and collects the operating status and execution results in real time, feeding them back to the modeling and decision-making modules to achieve online optimization of the model and strategy.
[0060] The functions of each module are interconnected and mutually supportive, forming a complete closed loop of "perception-modeling-decision-execution-feedback" to ensure the dynamic adaptability and global optimization of the system.
[0061] 2. Multi-source information sensing module
[0062] This module serves as the data foundation for the entire system. It is deployed at various nodes, load units, renewable energy devices, and energy storage systems in the industrial park's power distribution network. Through various sensors, smart meters, data acquisition units (DTUs), and monitoring systems, it enables real-time acquisition, time synchronization, and preprocessing of multi-source heterogeneous data, providing high-quality, standardized datasets for subsequent modeling and decision-making.
[0063] 2.1 Data Collection Dimensions
[0064] The data collected in this module covers four main categories: power parameters, load data, environmental parameters, and economic parameters, specifically including:
[0065] 1. Power parameters: voltage, current, active power, reactive power of each node in the distribution network, power flow of each branch, real-time output and predicted output of photovoltaic power generation units, charging and discharging power and state of charge (SOC) of the energy storage system, and power purchased from the main grid interface.
[0066] 2. Load data: production plan, real-time operating power, process start / stop status, and load adjustment range of manufacturing units; computing workload, server operating power, task unloading status, and energy consumption characteristics of data centers; electric vehicle charging demand, charging and discharging power, V2G operating status, and charging service revenue of charging stations.
[0067] 3. Environmental parameters: meteorological data such as light intensity, temperature, and wind speed (used for photovoltaic power output forecasting), and weather coefficients (to quantify the impact of meteorological conditions on the energy system);
[0068] 4. Economic and carbon parameters: time-of-use electricity price, carbon price, carbon emission intensity of each power generation unit, total carbon quota of industrial park, carbon flow rate of each load, and carbon potential of each node.
[0069] 2.2 Data Preprocessing Methods
[0070] The collected multi-source data suffers from issues such as time asynchrony, outliers, and missing values. This module preprocesses the data using the following methods:
[0071] 1. Time Synchronization: Based on BeiDou / GPS time synchronization technology, the timestamps of all collected data are aligned to a unified time scale (minute level) to ensure the spatiotemporal consistency of the data;
[0072] 2. Outlier Removal: The 3σ criterion and the Isolation Forest algorithm are used to identify and remove outlier data caused by sensor malfunctions and communication interference, ensuring data accuracy.
[0073] 3. Missing value imputation: For short-term missing data, linear interpolation is used for imputation; for long-term missing data, an LSTM-based time series prediction model is used for imputation to ensure data integrity.
[0074] 4. Data standardization: Normalize the original data of different dimensions to the [0,1] interval to eliminate the influence of dimensions and facilitate subsequent model calculation and algorithm training;
[0075] 5. Data Classification and Storage: The preprocessed datasets are classified and stored in the industrial big data platform, supporting rapid access and analysis by multiple modules.
[0076] 3. Electricity-Carbon Fluid Coupling Modeling Module
[0077] This module is the core analysis unit of the system. Based on carbon emission flow theory and power flow analysis methods, it constructs a dynamic coupling model of electricity and carbon flow to achieve synchronous analysis of electricity flow and carbon emission flow. Simultaneously, it establishes energy consumption models and constraints for multiple entities, including manufacturing, computing, charging, photovoltaics, and energy storage, providing quantitative analysis basis for the decision-making module. This module includes three sub-models: a carbon emission flow model, a carbon emission flow entropy model, and a multi-load coordinated energy consumption model.
[0078] 3.1 Carbon Emission Flow Model
[0079] Carbon emission flow is a virtual network flow that coexists with active power flow, used to characterize electricity-related carbon emissions. Its distribution is influenced by factors such as grid topology, line impedance, carbon emission intensity of generating units, and load distribution. Assuming an industrial park distribution network contains N nodes, K generating units, and M load nodes, this model uses power flow calculation and matrix analysis to achieve real-time calculation of node carbon potential, load carbon flow rate, branch carbon flow rate, and total system carbon flow rate.
[0080] 1. Basic calculations of power flow
[0081] First, the Nth-order branch power flow distribution matrix is obtained through power flow calculation. This matrix contains information about the topology of the power network and the dynamic active power flow distribution of the system. For any node i, its power balance equation is:
[0082] , ,
[0083] in, and Let represent the active and reactive power of node i.
[0084] Node voltage is expressed through the relationship between node current and impedance:
[0085] in, Let be the current from node i to j. Let be the line impedance from node i to j.
[0086] 2. Core Parameter Definitions
[0087] Power generation injection distribution matrix : A K×N matrix that describes the active power injected into the system by all power generation units and specifies the boundary conditions for the carbon emission flow;
[0088] Load distribution matrix A K×N matrix that describes the active power consumption of all load nodes;
[0089] Node active flux matrix : An N-order diagonal matrix, for node i, ;
[0090] Carbon emission intensity vector of power generation unit : Let be the carbon emission intensity (kgCO2 / kWh) of the k-th power generation unit.
[0091] Nodal carbon potential vector : Let be the carbon potential of node i, which is the carbon emission on the generation side corresponding to a unit of electricity consumption at node i, and also the carbon emission intensity of the node load.
[0092] 3. Calculation of nodal carbon potential
[0093] Nodal carbon potential is a core indicator in carbon emission flow models, and its calculation formula is as follows:
[0094] .
[0095] For a single node i, its carbon potential can be expressed as:
[0096] ,
[0097] Among them, I + Let i be the set of branches flowing into node i. The carbon flux density (the ratio of carbon flux rate to active power flow) is the carbon flux density of branch ij.
[0098] 4. Carbon Flow Rate Calculation
[0099] Load carbon flow rate: the amount of carbon emissions per unit time at a load node; load carbon flow rate vector. The calculation formula is:
[0100]
[0101] Branch carbon flow rate: the amount of carbon emissions per unit time from each branch; Nth-order branch carbon flow rate matrix. The calculation formula is:
[0102] .
[0103] Total carbon flow rate of the system: The total carbon emissions of the industrial park's power distribution network, calculated using the following formula:
[0104] ,
[0105] in, Let be the carbon flux rate at node i.
[0106] 3.2 Carbon Emission Flow Entropy Model
[0107] The carbon emission flow entropy (CEFE) index is introduced, which measures the degree of disorder in the distribution of carbon emission flows in industrial parks based on information entropy theory, providing a quantitative basis for carbon flow optimization. The formula for calculating carbon emission flow entropy is:
[0108] ,
[0109] in, Let represent the proportion of carbon emissions from the i-th node to the total carbon emissions of the system.
[0110] The numerical characteristics of carbon emission flow entropy are as follows: the higher the entropy value, the more dispersed the carbon emission distribution; the lower the entropy value, the more concentrated the carbon emission distribution; the theoretical maximum entropy is... When the entropy value is close to At this time, carbon emissions are relatively evenly distributed. This model uses carbon emission flow entropy as an important indicator for carbon flow optimization, and adjusts the carbon emission flow entropy to a reasonable range through scheduling decisions, thereby achieving uniform distribution and efficient management of carbon emissions.
[0111] 3.3 Multi-load Coordinated Energy Consumption Model
[0112] This embodiment establishes energy consumption models and constraints for manufacturing units, data centers, charging stations, photovoltaic power generation, and energy storage systems, clarifying the power regulation range and operating characteristics of each entity, thus providing a theoretical basis for multi-load coordinated scheduling. The power of each entity at time t is defined as: manufacturing unit power. Data center power Charging station power Photovoltaic power generation Energy storage system power .
[0113] 1. System power balance constraints
[0114] The total power supply of the industrial park is equal to the total load power, that is:
[0115]
[0116] in, Power purchased from the main power grid The charging power of the charging station (positive value for charging, negative value for discharging).
[0117] 2. Energy consumption models and constraints for each entity
[0118] Manufacturing unit (slow load): power constraint is It must meet the constraints of production process and delivery time, and the adjustment cycle is on the order of hours;
[0119] Data Center (High Load): Power Constraints are Energy consumption is affected by energy consumption per CPU cycle, the size of the computing model, and the number of CPU cycles required to process a task. It supports task unloading and server power adjustment, with an adjustment cycle of minutes.
[0120] Charging station (fast load): Supports V2G mode, power constraint is (Negative values represent discharging, positive values represent charging), charging service revenue ( (Revenue from unit charging)
[0121] Photovoltaic power generation unit: power constraint is The output is affected by weather conditions, which can be predicted using a time-series machine learning model;
[0122] Energy storage system: power constraint is (Negative values indicate discharging, positive values indicate charging);
[0123] Main grid generation unit: ramping constraint is The output power constraint is Carbon quota constraints are ( (Total carbon allowance).
[0124] 3. Electricity Carbon Cost Model
[0125] An operating cost model for the industrial park is established, including electricity costs and carbon costs, while also considering revenue from computing services and charging services, providing a quantitative basis for multi-objective optimization. The comprehensive cost at time t is:
[0126] in, For time-of-use electricity pricing, For carbon price, To calculate service revenue.
[0127] 4. Layered Security Reinforcement Learning Decision Module
[0128] This module is the core decision-making unit of the system and is key to solving the multi-timescale, multi-agent, and multi-constraint coordinated scheduling problem of electricity and carbon in industrial parks. This module formulates the multi-agent resource scheduling problem in industrial parks as a dual-timescale partially observable Markov decision process (POMDP). Based on the multi-agent dual-delay deep deterministic policy gradient (MATD3) framework and combined with the HIRO hierarchical reinforcement learning algorithm, it designs a two-layer collaborative decision-making architecture with upper-layer global decision-making and lower-layer local decision-making. An enhanced priority experience replay mechanism is introduced to address parameter drift issues, and grid security constraints are embedded into the reward function to achieve safe, efficient, and globally optimized scheduling decision generation.
[0129] 4.1 Problem Statement of POMDP
[0130] This embodiment models the multi-agent resource scheduling problem in industrial parks as a POMDP. The elements are defined as follows:
[0131] 1. Subject set i: All nodes i=1,2,...,N of the industrial park power distribution network. Each subject is an independent learning unit, and they communicate and coordinate by exchanging observations.
[0132] 2. Global state space S: It consists of all information of the industrial park, such as time, weather, power grid topology, power parameters, carbon flow parameters, and the operating status of each load. It is divided into upper state and lower state.
[0133] 3. Action Space A: Includes upper-level global actions and lower-level local actions, involving power allocation and discrete events (process start / stop, task unloading, charge / discharge selection);
[0134] 4. Transfer function σ: σ: S×A×S→[0,1], defines the Markov transition probability distribution, realizes state transition based on power flow calculation, and includes updates of physical quantities such as voltage, power flow, and SOC;
[0135] 5. Reward function r: Divided into upper-level reward and lower-level reward, it quantifies the optimization effect of decision-making and includes indicators such as cost, carbon emission reduction, and safety constraints;
[0136] 6. Discount factor γ: γ∈[0,1], used to calculate long-term cumulative rewards. In this invention, γ=0.99 is set.
[0137] Each agent obtains partially observable state information from the global state s, executes action a, and generates a reward r and a new global state s' through the transition function σ. The goal of the algorithm is to find the optimal joint policy. To maximize the expected cumulative rewards for the entire park.
[0138] 4.2 Two-tier decision-making architecture design
[0139] Based on the characteristics of multiple time scales, this embodiment designs a two-layer architecture with hourly-level upper-layer global decision-making and minute-level lower-layer local decision-making, realizing an interactive mechanism for state transmission and reward fusion, and ensuring the coordination of global optimization and local adjustment.
[0140] 1. Top-level global decision-making (hourly level)
[0141] The upper-level decision-making focuses on grid security, carbon emission reduction, and minimizing overall costs. It optimizes the main grid's power purchase, photovoltaic output, and energy storage system charging and discharging plans, and outputs global control targets (such as the power allocation upper limit for each load cluster) to set constraints for lower-level decision-making.
[0142] upper state ,in For weather coefficients, } represents the energy storage state of charge. The set of injected power for all load nodes. Let be the nodal carbon potential vector. The load carbon flow rate vector;
[0143] Upper-level actions This refers to the optimized allocation of power purchased by the main power grid and photovoltaic power output;
[0144] Upper-level rewards The calculation primarily considers electricity costs, carbon costs, and penalty constraints, while also incorporating the sum of all lower-level rewards during that period. The formula is as follows:
[0145]
[0146] in, Photovoltaic carbon emission intensity ( =0).
[0147] 2. Lower-level local decision-making (minute-level)
[0148] The lower-level decision-making receives the control objectives and constraints from the upper-level decision-making. With local load optimization, real-time response, and constraint satisfaction as the core objectives, it optimizes the start and stop of manufacturing unit processes, data center task scheduling, charging and discharging power of charging stations, and real-time adjustment of energy storage systems to achieve refined management and control of each load.
[0149] Lower state That is, the combination of upper-level state and upper-level action to ensure that lower-level decisions are constrained by the upper-level global goal;
[0150] Lower-level actions This refers to real-time power regulation in manufacturing, data centers, charging stations, and energy storage systems.
[0151] Lower-level rewards Considering rewards for completing manufacturing tasks, data center satisfaction, charging service revenue, and carbon emission entropy, while also introducing a power imbalance penalty, the calculation formula is as follows:
[0152] ,
[0153] in, - The weighting coefficients are used to balance the relative importance of each reward item. They are determined as follows: based on the optimization objectives of the industrial park (such as economic benefits, carbon emission reduction, grid security, etc.), a normalization process is used to ensure their sum equals 1, or through multiple reinforcement learning simulations and adjustments, the overall benefit of the park is maximized, with H... t The optimal weight combination is determined with the carbon emission flow entropy (CEFE) as the objective. To earn rewards for completing the manufacturing task. This represents the power imbalance value.
[0154] 3. Integration of State Transfer and Rewards
[0155] State transmission: The results of upper-level decisions (power purchased by the main grid and photovoltaic output) are transmitted to the lower-level decision model as part of the lower-level environmental conditions, so that the lower-level decision can determine the available power range and ensure that local regulation meets the global optimization goal;
[0156] Reward integration: The upper-level reward not only includes its own cost and carbon emission reduction reward, but also integrates the sum of all lower-level rewards during the period, so that the learning of the upper-level strategy can comprehensively consider the effect of local adjustment, ensuring the consistency of the two-level decision-making and the global optimization.
[0157] 4.3 Layered Secure Reinforcement Learning Network Structure
[0158] This embodiment is based on the MATD3 framework, and designs the Actor-Critic network structure for each of the upper and lower layers. LeakyReLU is used as the activation function, and a dual-Critic network is introduced to reduce overestimation bias. The target network parameters are updated through a soft update mechanism to ensure the stability of network training.
[0159] This module contains a total of 12 neural networks: 6 networks in the upper layer (2 Actor networks and 4 Critic networks) and 6 networks in the lower layer (2 Actor networks and 4 Critic networks). The input and output dimensions of each network are designed according to the state and action space.
[0160] Upper-layer Actor network: Input dimension 104 (number of upper-layer state features), output dimension 2 (upper-layer actions: main grid power purchase, photovoltaic power output);
[0161] Upper-layer Critic network: Input dimension 106 (number of upper-layer state features + upper-layer action dimension), output dimension 1 (upper-layer reward value);
[0162] Lower-layer Actor network: Input dimension 106 (number of lower-layer state features), output dimension 9 (lower-layer actions: power of 3 manufacturing units, 2 data centers, 3 charging stations, and 1 energy storage system);
[0163] Lower-layer Critic network: Input dimension 115 (lower-layer state feature count + lower-layer action dimension), output dimension 1 (lower-layer reward value).
[0164] like Figure 2 As shown, the network structure and interaction mechanism of upper-layer global decision-making and lower-layer local decision-making are illustrated. The upper layer obtains state features through state extraction, generates actions through the Actor network, and evaluates rewards through the Critic network; the lower layer receives the upper layer's state and actions and completes local decisions; the priority experience replay module implements priority management of experience samples, the safety constraint embedding module completes constraint quantification and reward correction, and finally achieves two-layer decision-making collaboration through state transmission and reward fusion.
[0165] 2. Network training parameters
[0166] This embodiment ensures the efficiency and convergence of network training by setting appropriate training parameters:
[0167] Learning rate: 0.0004 for upper Actor network and 0.004 for upper Critic network; 0.0002 for lower Actor network and 0.002 for lower Critic network. The learning rate is dynamically adjusted during training.
[0168] Soft update parameter: τ=0.005, target network parameter update formula is θ'=τθ+(1-τ)θ';
[0169] Experience replay buffer: Capacity 200~500 samples, used to store training experience;
[0170] Batch size: 64, 64 samples are randomly sampled from the experience buffer for each training session;
[0171] Training epochs: 800 epochs, to ensure the network fully converges.
[0172] 3. Network Training Formula
[0173] Lower-level objective Q-value and lower-level Critic network loss function:
[0174]
[0175] The upper-layer objective Q-value, Critic loss function, and Actor gradient formula are similar to those of the lower layer, except that the state, action, and reward are replaced with relevant parameters from the upper layer.
[0176] 4.4 Enhanced Priority Experience Replay Mechanism
[0177] To address the parameter drift problem during multi-agent reinforcement learning training and correct erroneous actions and negative reward experiences generated during upper and lower layer policy training, this invention proposes an enhanced priority experience replay mechanism. This mechanism prioritizes experience samples, training high-priority experience samples first, thereby improving training efficiency and policy robustness.
[0178] 1. Empirical Priority Calculation: Priority is calculated based on the TD error (temporal difference error) of empirical samples. The larger the TD error, the higher the priority of the sample, indicating that the sample contains more learning information. The calculation formula is: TD = |y - Q(s,a)| (y is the target Q value, and Q(s,a) is the Q value predicted by the current network).
[0179] 2. Empirical Sample Sampling: Based on priority, weighted random sampling is used to select samples from the empirical buffer. High-priority samples have a higher probability of being sampled. At the same time, importance sampling weights are introduced to correct the training bias caused by non-uniform sampling and ensure the unbiasedness of training.
[0180] 3. Error correction: For erroneous actions generated during training—negative reward experiences (such as actions that lead to violations of power grid constraints)—their priority is set to the highest, forcing the network to repeatedly learn such experiences, correcting the parameters of the policy network, avoiding the generation of similar erroneous actions in the future, and ensuring the safety of decision-making.
[0181] 4.5 Security Constraint Embedding Mechanism
[0182] To ensure the safe and stable operation of the power grid, power grid security constraints and carbon quota constraints are embedded into the training process and reward function of reinforcement learning. By using a strategy network guided by constraint violation penalties and constraint satisfaction rewards, decisions that meet the constraints are generated, thus achieving a balance between safety and optimization.
[0183] 1. Quantification of constraint violation degree
[0184] Cost of defining constraints The degree of violation of node voltage, branch current, and power flow is quantified by the following formula:
[0185]
[0186] in:
[0187] The degree of violation of node voltage constraints;
[0188] The degree of violation of branch current constraints;
[0189] : Degree of violation of power flow constraints.
[0190] 2. Constraint Embedded Reward Function
[0191] By incorporating the cost of constraint violation into the reward function, decisions that violate constraints are penalized, while decisions that satisfy constraints are rewarded. The modified reward function is as follows:
[0192] r'=r - λ·c t + μ·I( =0)
[0193] Where λ is the penalty coefficient, μ is the reward coefficient, and I( =0) is the indicator function (when the constraint is violated, the cost is 0). When =0 (i.e., both grid security constraints and carbon quota constraints are satisfied), I( =0) is 1, when the constraint violates the cost When I > 0 (i.e., any constraint is violated), =0) is 0). Among them, the penalty coefficient λ ranges from [0.1,10], and the reward coefficient μ ranges from [0.01,1]. μ is usually less than λ. The two are determined based on the relative importance of constraint violation cost and original reward, combined with the simulation debugging of the strictness of industrial park power grid security constraints.
[0194] 3. Training constraints and limitations
[0195] During training, an upper limit d for the constraint violation cost is set (the preferred method for determining this is: based on the power grid safety standards of the industrial park distribution network (such as node voltage deviation controlled within ±5%, branch current not exceeding the rated value, and power flow not exceeding the limit) and engineering experience, the maximum permissible degree of violation of power grid safety constraints is determined by quantifying them; in this embodiment, the value range of d is [0.01, 0.1], representing the maximum permissible degree of constraint violation), requiring the training strategy π to satisfy... ( For the expected discount return based on the cost of constraint violation, If the strategy violates this constraint, training is stopped and the network parameters are readjusted to ensure that the trained strategy complies with the power grid security constraints.
[0196] 5. Execution and Feedback Closed-Loop Module
[0197] This module is the system's execution unit, responsible for the precise issuance of scheduling decisions, execution monitoring, effect evaluation, and model feedback optimization. It establishes a closed-loop control system across the entire process, ensuring the effective execution of decisions and the system's dynamic adaptability. This module comprises two sub-units: the execution unit and the feedback unit.
[0198] 5.1 Execution Unit
[0199] Based on Industrial Internet of Things (IIoT) and fieldbus technology, the execution unit accurately sends the scheduling instructions (upper-level global power allocation instructions and lower-level local power adjustment instructions) generated by the hierarchical security reinforcement learning decision module to each field execution device, realizing intelligent control of manufacturing units, data centers, charging stations, photovoltaic inverters, energy storage converters, and main grid interfaces.
[0200] 1. Command issuance method: Standardized industrial communication protocols (Modbus TCP / IP, Profinet, OPCUA) are used to issue scheduling commands to the controllers of each device. The commands include information such as power values, execution time, and adjustment cycle to ensure coordinated execution of each device;
[0201] 2. Perform equipment control:
[0202] Manufacturing unit controller: Adjusts the start / stop status and operating power of processes according to instructions to achieve orderly regulation of production load;
[0203] Data center controller: Schedules computing tasks according to instructions, adjusts server operating power, offloads tasks, and optimizes computing load;
[0204] Charging station controller: Adjusts charging and discharging power and guides V2G operation according to instructions to realize the interaction between charging load and power grid;
[0205] Photovoltaic inverters: Adjust photovoltaic output according to instructions to improve the renewable energy consumption rate;
[0206] Energy storage converter: It realizes charging and discharging control according to instructions, and smooths the peak-valley difference of grid load;
[0207] Main grid interface controller: Adjusts the purchased power according to instructions to achieve coordinated operation with the main grid.
[0208] 5.2 Feedback Unit
[0209] The feedback unit collects the operating status of each execution device, the power parameters of the power grid, and the carbon emission flow parameters in real time, evaluates the execution effect of the scheduling decision, calculates the operation efficiency index, and feeds back the monitoring data and efficiency index to the power-carbon flow coupled modeling module and the hierarchical security reinforcement learning decision module to realize the online updating and optimization of the model and strategy.
[0210] 1. Operational Status Monitoring: Real-time monitoring of the actual operating power of each device, node voltage, branch power flow, energy storage SOC, actual output of photovoltaic power generation, carbon potential of each node, carbon flow rate of each load, total carbon emissions of the system, operating costs, etc. The time scale of the monitoring data is in the minute range to ensure real-time performance.
[0211] 2. Operational performance evaluation: Set multi-dimensional operational performance indicators to quantify the execution effect of dispatching decisions, including economic indicators (operating cost reduction rate, renewable energy consumption rate), environmental indicators (carbon emission reduction rate, carbon emission flow entropy, carbon quota utilization rate), safety indicators (voltage constraint satisfaction rate, power flow constraint satisfaction rate), and dispatching indicators (power balance, load peak-valley difference rate, decision response time).
[0212] 3. Model and Strategy Feedback Optimization: Monitoring data and performance indicators are fed back to the upstream module to update the parameters of the carbon flow model online (such as the node carbon potential calculation coefficient and carbon flow rate correction coefficient). The real-time running state-action-reward experience is added to the experience replay buffer and the reinforcement learning network is trained online to optimize the scheduling strategy so that the strategy can adapt to the dynamic operating conditions of the industrial park.
[0213] like Figure 6The diagram illustrates the four-level core architecture of this invention and the sub-units of each module. Arrows indicate the direction of data and instruction transmission. The multi-source information sensing module collects multi-dimensional data, preprocesses it, and then transmits it to the industrial big data platform for storage. The power-carbon flow coupling modeling module retrieves data, completes power-carbon coupling analysis, and then transmits it to the decision-making module. The decision-making module generates scheduling decisions and transmits them to the execution unit. The execution unit issues instructions to field equipment, the feedback unit collects the operating status and evaluates the effect, and finally feeds the data back to the big data platform to achieve closed-loop control.
[0214] 6. Control Method Implementation Steps
[0215] Based on the above system architecture, the specific implementation steps of the power-carbon coordinated control method for industrial parks based on hierarchical security reinforcement learning provided in this embodiment of the invention are as follows:
[0216] S1: Multi-source data acquisition and preprocessing
[0217] The power parameters, load data, environmental parameters, economic and carbon parameters of the industrial park are collected by a multi-source information sensing module. The data is preprocessed using methods such as time synchronization, outlier removal, missing value completion, and data standardization to generate a standardized dataset, which is then classified and stored.
[0218] S2: Construction and Calculation of Electricity-Carbon Flow Coupling Model
[0219] The preprocessed dataset is input into the power-carbon flow coupling modeling module. First, the branch power flow distribution matrix is obtained through power flow calculation. Based on the carbon emission flow theory, the node carbon potential, load carbon flow rate, branch carbon flow rate, and total system carbon flow rate are calculated. The carbon emission flow entropy index is introduced to analyze the carbon emission distribution characteristics. At the same time, a multi-load coordinated energy consumption model and constraints for manufacturing, computing, charging, photovoltaic, and energy storage are established to calculate the electricity carbon cost and comprehensive benefits, and generate an electricity-carbon coupling analysis report.
[0220] S3: Layered Secure Reinforcement Learning Decision Initialization
[0221] Input the electricity-carbon coupling analysis report into the hierarchical security reinforcement learning decision module, initialize the state, action, and reward space of the POMDP model, set the structure and training parameters (learning rate, discount factor, soft update parameters, etc.) of the hierarchical reinforcement learning network, initialize the experience replay buffer, and set the thresholds for power grid security constraints and carbon quota constraints.
[0222] S4: Upper-level global decision generation (hourly level)
[0223] Based on daily / hourly forecast data (PV output, load demand, electricity price, and carbon price), the upper-level Actor network generates upper-level actions for main grid power purchase and PV output according to the current upper-level state. The reward value is evaluated by the upper-level Critic network. If the training convergence condition is met, the upper-level global control objective (the upper limit of power allocation for each load cluster) is output. If convergence is not achieved, the upper-level network continues to be trained until convergence.
[0224] S5: Lower-level local decision generation (minute-level)
[0225] The lower-level decision-making module receives the global control target from the upper level, and uses the upper-level state and actions as the lower-level state. The lower-level actor network generates lower-level actions for manufacturing, data centers, charging stations, and energy storage systems based on the current lower-level state. The lower-level critic network evaluates the reward value and introduces an enhanced priority experience replay mechanism to train the lower-level network, ensuring that the lower-level actions comply with the upper-level constraints and grid security constraints.
[0226] S6: Execution of scheduling instructions
[0227] The execution unit of the execution and feedback closed-loop module sends the scheduling instructions of the upper-level global decision and the lower-level local decision to each field execution device (manufacturing unit controller, data center controller, charging station controller, photovoltaic inverter, energy storage converter, etc.) through the industrial communication protocol to realize the coordinated control of each device.
[0228] S7: Real-time monitoring and performance evaluation of operational status
[0229] The feedback unit collects the operating status of each device, power grid parameters, and carbon emission flow parameters in real time, calculates multi-dimensional operational efficiency indicators in terms of economy, environment, safety, and scheduling, and evaluates the effectiveness of scheduling decisions.
[0230] S8: Feedback Optimization and Closed-Loop Iteration
[0231] The monitoring data and performance indicators are fed back to the power-carbon flow coupling modeling module and the hierarchical safety reinforcement learning decision module to update the carbon flow model parameters online, add real-time experience to the experience replay buffer and train the reinforcement learning network online to optimize the scheduling strategy; return to step S2 to perform the power-carbon coupling calculation and decision generation for the next time period, and realize closed-loop iterative control of the whole process.
[0232] The above-described solutions in this invention construct a closed-loop, end-to-end control system encompassing multi-source information perception, power-carbon flow coupled modeling, hierarchical security reinforcement learning decision-making, execution, and feedback. This system integrates carbon emission flow theory with hierarchical security reinforcement learning algorithms, solving the core technical challenges of power-carbon coordinated control in industrial parks. It achieves multi-timescale coordinated scheduling of slow and fast loads, dynamic coupled control of power-carbon flows, and effective fulfillment of grid security constraints, yielding significant results in terms of technology, economy, environment, and application. Specifically:
[0233] The constructed dynamic power-carbon flow coupling model can calculate and accurately trace the carbon emissions of each node, load, and branch in the park in real time, quantify the core indicators of node carbon potential, carbon flow rate, and carbon emission flow entropy, and achieve a carbon emission traceability accuracy of over 95%.
[0234] The multi-timescale hierarchical decision-making architecture enables the coordination of hourly global optimization and minute-level local adjustment, with a decision response time of minutes, which can quickly adapt to dynamic operating conditions such as sudden changes in production tasks and fluctuations in photovoltaic output.
[0235] The enhanced priority experience replay mechanism effectively solves the parameter drift problem in multi-agent reinforcement learning. Combined with the safety constraint embedding mechanism, it controls the grid voltage deviation within ±5%, and the power flow and branch current constraint satisfaction rate is 100%, thus preventing grid safety accidents caused by parameter drift.
[0236] The closed-loop system enables integrated collection of multi-source data, online model updates, and iterative optimization of strategies, significantly improving the system's dynamic adaptability and robustness, and making it suitable for the operational needs of industrial parks of different sizes and business types.
[0237] To make the technical solution and implementation effects of the present invention clearer and more explicit, the technical solution of the present invention will be described in detail below with reference to the IEEE 33-node power distribution network test system through specific embodiments. The scope of protection of the present invention is not limited to the following embodiments.
[0238] like Figure 1 As shown, this embodiment constructs an industrial park test environment based on an improved IEEE 33-node distribution network system. This system is a medium-voltage radial distribution network, containing 33 nodes and 32 branches, adapted to the network topology, line parameters, and load curve design of the industrial park. Specific parameter settings are as follows:
[0239] 1. Power supply side: Main grid interface capacity 1000kW, photovoltaic (PV) system rated capacity 500kW (output is subject to weather conditions), energy storage system (ESS) capacity 1000kWh, maximum charging / discharging power 300kW / 200kW, SOC limit 0.1-0.9;
[0240] 2. Load side: 3 manufacturing units (each rated 100kW), 2 data centers (each rated 150kW), 3 charging stations (each rated 50kW, supporting V2G bidirectional adjustment);
[0241] 3. Economic and Carbon Parameters: Time-of-use electricity price (off-peak 00:00-07:00: 0.35 yuan / kWh; peak 08:00-19:00: 0.65 yuan / kWh; normal 20:00-23:00: 0.45 yuan / kWh); carbon price 0.08 yuan / kgCO2; main grid carbon emission intensity 0.85 kgCO2 / kWh, photovoltaic carbon emission intensity 0 kgCO2 / kWh; charging station customer charging fee 1.2 yuan / kWh, grid electricity sales price 0.9 yuan / kWh, V2G discharge compensation 0.3 yuan / kWh;
[0242] 4. Algorithm Implementation: Implemented in Python, using PyTorch to build a hierarchical security reinforcement learning network. The PandaPower library is used to calculate system-level power flow, node voltage, and line load. The Actor / Critic network uses the LeakyReLU activation function, and the experience replay buffer capacity is 200-500.
[0243] like Figure 3 As shown, in the power-carbon flow coupling model of this embodiment, the carbon flow change is highly correlated with the load power change, and the differences in power and carbon flow between slow load (manufacturing) and fast load (data center) are clearly distinguished. At the same time, it provides data support for hierarchical security reinforcement learning decision-making: the hourly upper-level decision can formulate global power allocation based on the stability of manufacturing load, and the minute-level lower-level decision can make local adjustments to the volatility of carbon flow in data center, ultimately realizing dynamic control of multi-load power and carbon flow coordination.
[0244] like Figure 4 As shown, using a 24-hour time axis, the power variation curves of photovoltaic output, manufacturing load, data center load, and charging station load in this embodiment are displayed, along with the corresponding dynamic variation curves of carbon emissions. Photovoltaic output exhibits a characteristic of being high during the day and low at night, while the manufacturing load remains stable (slow load characteristic). The data center and charging station loads dynamically adjust with photovoltaic output and electricity prices (fast load characteristic). The carbon emission curve is highly correlated with the power curve, and carbon emissions are significantly reduced during periods of high photovoltaic output.
[0245] like Figure 5 As shown, with the number of training rounds (0-800) on the horizontal axis and cumulative reward on the vertical axis, the training performance of the HSRL algorithm of this invention is compared with that of TD3, HRL-only, Greedy, and Random algorithms. The cumulative reward of the HSRL algorithm is significantly higher than that of other algorithms, and the reward fluctuation is small and the convergence speed is fast during training, which reflects the excellent learning efficiency and robustness of the algorithm provided by this invention.
[0246] In the above test examples, multi-load collaborative scheduling improves the renewable energy absorption rate, reducing industrial park operating costs by more than 18%. In the IEEE 33-node test system, the system reward is 21.9% higher than the traditional baseline model, with an average total training reward of 2205.12. Continuous operation over multiple days saves more than 120,000 yuan in costs compared to traditional scheduling methods. The solution provided by this invention can tap the flexible adjustment potential of data centers and charging stations, combining V2G technology and computing task offloading to increase charging and computing service revenue, improving the return on investment of industrial parks by 10%-15%. Through time-of-use pricing response and energy storage system charging and discharging optimization, the peak-valley difference of the power grid load is smoothed, reducing the demand for electricity purchases during peak hours of the main grid and lowering electricity procurement costs. Based on precise carbon flow management, carbon quotas are used efficiently, reducing the cost of carbon over-emission. At the same time, carbon emission flow entropy optimization improves the uniformity of carbon emission distribution, reducing the marginal cost of carbon emission reduction. Through power-carbon flow coupling management and multi-load coordinated dispatch, carbon emissions in the industrial park have been reduced by more than 15%, with a reduction of up to 25% under high-output photovoltaic scenarios. Cumulative carbon emission reductions exceeded 3.2 tons during multi-day continuous operation verification. The carbon emission flow entropy index effectively quantifies the distribution characteristics of carbon emissions. After dispatch decision optimization, the carbon emission flow entropy has been reduced by more than 30%, resulting in a more uniform carbon emission distribution and improved overall carbon emission reduction effectiveness. The utilization rate of renewable energy sources such as photovoltaic power generation has increased to over 95%, reducing fossil fuel consumption, lowering carbon emission intensity on the power generation side, and promoting the green transformation of the industrial park's energy structure. Accurate carbon emission traceability and clear carbon responsibility delineation have been achieved, providing quantitative basis for the industrial park to formulate targeted carbon emission reduction measures and helping the park achieve its carbon peak and carbon neutrality goals.
[0247] In practical applications, this solution adopts a modular design, with modules interacting through standardized data interfaces. It boasts strong scalability, facilitating the integration of new load types or energy devices such as hydrogen fuel cells and industrial waste heat utilization. It is adaptable to various industrial clusters, including machinery manufacturing, new energy industrial parks, and smart industrial parks. Quantitative indicators such as carbon potential, carbon flow rate, and constraint violation degree at each node in the carbon flow model and decision-making process enhance system interpretability, solving the "black box decision-making" problem of traditional reinforcement learning and making it easier for engineers to understand and operate. The system can be developed based on mature software frameworks such as Python, PyTorch, and PandaPower, combined with general-purpose hardware facilities such as industrial IoT and big data platforms. This reduces engineering implementation difficulty and allows for rapid deployment. Furthermore, it can seamlessly integrate with existing energy management systems (EMS) and intelligent manufacturing systems (MES) in industrial parks without requiring large-scale modifications to existing equipment, reducing application costs and demonstrating broad engineering promotion value.
[0248] Based on the same inventive concept, this invention also provides a computer device, comprising: one or more processors, and a memory for storing one or more computer programs; the programs include program instructions, and the processor executes the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, used to implement one or more instructions, specifically for loading and executing one or more instructions stored in a computer storage medium to implement the above-described method.
[0249] It should be further explained that, based on the same inventive concept, the present invention also provides a computer storage medium storing a computer program, which, when executed by a processor, performs the above-described method. This storage medium can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0250] It should be noted that, unless otherwise defined, the technical or scientific terms used in this invention should have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0251] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
[0252] This invention is not limited to the preferred embodiment described above. Anyone inspired by this invention can derive other forms of a multi-load electric carbon collaborative management decision-making method for industrial parks based on hierarchical security reinforcement learning. All equivalent variations and modifications made within the scope of the patent applications of this invention shall fall within the scope of this invention.
Claims
1. A decision-making method for multi-load electricity-carbon collaborative management in industrial parks based on hierarchical security reinforcement learning, characterized in that, include: Based on multi-source operation data of industrial parks, combined with power flow calculation and carbon emission flow theory, the carbon potential and carbon flow rate of distribution network nodes are calculated. The carbon flow entropy index is introduced to quantify the carbon emission distribution characteristics and is used as a component of the reward function of the lower-level decision module. At the same time, a multi-load collaborative energy consumption model is established, and a dynamic power-carbon flow coupling model is constructed. Based on the dynamic power-carbon flow coupling model, a hierarchical reinforcement learning decision-making system is constructed, adopting a dual-timescale architecture with a first decision cycle and a second decision cycle, where the duration of the first decision cycle is longer than that of the second decision cycle. The upper-level decision module generates a global scheduling strategy for the main grid's power purchase and renewable energy generation output, while the lower-level decision module generates local adjustment strategies for slow load, fast load, and energy storage systems in industrial parks. The decision-making system introduces an enhanced priority experience replay mechanism to correct the parameter drift problem in multi-agent reinforcement learning, and embeds grid security constraints and carbon quota constraints into the reward function to quantify the cost of constraint violations. The global scheduling strategy and local adjustment strategy are converted into scheduling instructions and sent to the execution unit to realize the coordinated control of multiple loads and energy systems in the industrial park. The system monitors the operation of the park's power grid and the status of carbon emission flows in real time. Based on the feedback data from the monitoring, it updates the calculation parameters of the dynamic power-carbon flow coupling model and the training strategy of the hierarchical reinforcement learning decision system, thereby completing the iterative optimization of the model and strategy.
2. The decision-making method for multi-load electricity and carbon collaborative management in industrial parks based on hierarchical security reinforcement learning as described in claim 1, characterized in that: The multi-source operation data of the industrial park includes power parameters, operation parameters of various types of loads, environmental meteorological parameters, and economic carbon parameters. The multi-source operation data needs to be time-synchronized, outlier removal, missing value completion, and standardized preprocessing before being input into the dynamic power-carbon flow coupling model. The standardized preprocessing is to normalize the original data of different dimensions to the [0,1] interval.
3. The decision-making method for multi-load electricity and carbon collaborative management in industrial parks based on hierarchical security reinforcement learning as described in claim 1, characterized in that: The node carbon potential is the carbon emission on the generation side corresponding to a unit of electricity consumption at the node. The carbon potential of a single node is obtained by dividing the active power flux of the node by the carbon flow density and power flow of all branches flowing into the node, the power injected by the node, and the carbon emission intensity of the corresponding power generation unit. The carbon flow rate includes the load carbon flow rate, the branch carbon flow rate, and the total system carbon flow rate. The load carbon flow rate is obtained by multiplying the load distribution parameters by the corresponding node carbon potential. The branch carbon flow rate is calculated based on the power flow distribution of the distribution network and the node carbon potential. The carbon flow entropy index is calculated based on the information entropy theory. It is obtained by multiplying the proportion of each node's carbon emissions to the total system carbon emissions by the natural logarithm of that proportion, and then summing the calculation results for all nodes.
4. The decision-making method for multi-load electricity and carbon collaborative management in industrial parks based on hierarchical security reinforcement learning as described in claim 1, characterized in that: The slow load is the manufacturing load, and the fast load includes the data center computing load and the vehicle-to-grid charging station charging load; the first decision cycle is on the hour level, and the second decision cycle is on the minute level; the upper-level decision result serves as the environmental condition for the lower-level decision to achieve state transmission, and the upper-level reward function integrates all lower-level reward values within the corresponding time period to achieve reward fusion; the operation of the energy storage system must meet the upper and lower limits of the state of charge and the range of charging and discharging power.
5. The decision-making method for multi-load electricity and carbon collaborative management in industrial parks based on hierarchical security reinforcement learning as described in claim 1, characterized in that: The enhanced priority experience replay mechanism is implemented as follows: the sample priority is calculated based on the temporal difference error of the experience samples, and the larger the temporal difference error, the higher the sample priority; samples are selected from the experience replay buffer using a weighted random sampling method based on the priority, and importance sampling weights are introduced to correct the training bias caused by non-uniform sampling. The negative reward experience, which is an erroneous action that leads to a violation of grid constraints or carbon quota constraints during training, is given the highest priority. The reinforcement learning network is forced to repeatedly learn this type of experience to correct the policy network parameters.
6. The decision-making method for multi-load electricity and carbon collaborative management in industrial parks based on hierarchical security reinforcement learning as described in claim 1, characterized in that: The power grid security constraints include node voltage constraints, branch current constraints, and power flow constraints; the constraint violation cost is the sum of the degree of violation of node voltage constraints, branch current constraints, and power flow constraints. The degree of violation of node voltage constraints is the sum of the difference between the actual node voltage exceeding the lower voltage limit and the difference between the upper voltage limit and the actual node voltage, with only the non-negative part being calculated. The degree of violation of branch current constraints is the result of the ratio of the actual branch current to the rated branch current minus 1, with only the non-negative part being calculated. The degree of violation of power flow constraints is the result of the ratio of the actual branch apparent power to the rated branch apparent power minus 1, with only the non-negative part being calculated.
7. The decision-making method for multi-load electricity and carbon collaborative management in industrial parks based on hierarchical security reinforcement learning as described in claim 1, characterized in that: The specific method for embedding power grid security constraints and carbon quota constraints into the reward function is as follows: the constraint violation cost is introduced into the original reward function to obtain the modified reward function. The modified reward function is obtained by subtracting the product of the constraint violation cost and the preset penalty coefficient from the original reward value, and then adding the constraint satisfaction reward. When the constraint violation cost is zero, the preset constraint satisfaction reward is obtained; when the constraint violation cost is not zero, there is no such reward.
8. The decision-making method for multi-load electricity and carbon collaborative management in industrial parks based on hierarchical security reinforcement learning as described in claim 1, characterized in that: The hierarchical reinforcement learning decision-making system is constructed based on a multi-agent, dual-delay deep deterministic policy gradient framework combined with a hierarchical reinforcement learning algorithm, and includes actor-commentator networks in the upper and lower layers respectively. Update the target network parameters using a soft update mechanism.
9. The decision-making method for multi-load electricity and carbon collaborative management in industrial parks based on hierarchical security reinforcement learning as described in claim 1, characterized in that: The monitored feedback data includes the actual operating power of each execution unit, the actual voltage of the distribution network nodes, the actual power flow of the branches, the actual state of charge of the energy storage system, the actual monitoring parameters of the carbon emission flow, and the actual operating cost and carbon emissions of the industrial park. The calculation parameters of the dynamic power-carbon flow coupling model are updated online, and the training strategy of the hierarchical reinforcement learning decision system is updated to add the real-time running state-action-reward experience to the experience replay buffer and train the reinforcement learning network online.
10. A multi-load electric carbon collaborative management decision-making system for industrial parks based on hierarchical security reinforcement learning, characterized in that, The system for implementing the method of any one of claims 1 to 9 comprises a multi-source information sensing module, a power-carbon flow coupling modeling module, a hierarchical security reinforcement learning decision-making module, and an execution and feedback closed-loop module. The modules interact and transmit instructions through a standardized data interface. The multi-source information sensing module collects and preprocesses multi-source operational data from the industrial park. The power-carbon flow coupling modeling module constructs a dynamic power-carbon flow coupling model and calculates node carbon potential, carbon flow rate, and carbon flow entropy, using carbon flow entropy as a component of the reward function of the lower-level decision-making module. The hierarchical security reinforcement learning decision-making module constructs a hierarchical reinforcement learning decision-making system and generates global scheduling strategies and local adjustment strategies. The execution and feedback closed-loop module issues scheduling instructions, monitors operational status, collects feedback data, and performs iterative optimization of the model and strategies.