Industrial park carbon emission prediction method, equipment and medium
By combining digital twin models and multi-agent reinforcement learning with calibration models, high-precision carbon emission predictions and confidence intervals are generated, solving the problems of "black box" and poor timeliness in carbon emission prediction for industrial parks, and realizing high-precision, short-term carbon emission prediction and decision support.
Patent Information
- Application Number
- CN202511722126.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-21
- Publication Date
- 2026-02-13
AI Technical Summary
Existing carbon emission prediction technologies for industrial parks suffer from problems such as "black box" prediction, lack of foresight, coarse prediction granularity, and poor timeliness, which cannot meet the park's needs for refined energy dispatch and carbon quota trading.
By employing a digital twin model combined with multi-agent reinforcement learning and a calibration model, a set of carbon emission evolution trajectories is generated based on the randomness of the policy by inputting the production plan into the decision layer of the digital twin model. The residuals are then calibrated using the calibration model to obtain high-precision carbon emission predictions and confidence intervals.
It achieves high-precision, short-term carbon emission forecasting, has decision support capabilities, significantly improves forecast timeliness and accuracy, can detect carbon emission exceedance risks in advance, and supports proactive scheduling and optimization of the park.
Smart Images

Figure CN121525979A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of carbon management, and particularly relates to an industrial park carbon emission prediction method, device and medium. BACKGROUND
[0002] An industrial park is a concentrated area of carbon emissions, and accurate prediction of carbon emissions in the industrial park is a key to achieving fine carbon management and the "double carbon" goal. Existing industrial park carbon emission prediction technologies mainly have the following limitations: (1) "black box" prediction, ignoring internal mechanisms: Most methods (such as LSTM, ARIMA) regard the entire park as a black box and only perform time series modeling based on historical carbon emission data. This method cannot reveal the complex dynamic coupling relationship between enterprises and production units within the park, for example, how the production fluctuations of upstream enterprises affect the energy consumption of downstream enterprises through the supply chain.
[0003] (2) Passive response, lack of foresight: Traditional prediction is "backward-looking", that is, it infers the future from the past. It cannot answer "if" type questions, for example: "If enterprise A adjusts the production shift tomorrow, how will the carbon emissions of the entire park change?" The lack of this ability makes it difficult for the prediction result to directly guide the park managers to conduct active and forward-looking scheduling optimization.
[0004] (3) Coarse prediction granularity, poor timeliness: Existing predictions are mostly daily or longer, which cannot meet the fine needs of the park for hour-level energy scheduling and carbon quota trading. Moreover, the prediction model is usually disconnected from the actual production and operation system (such as MES, ERP) of the park, resulting in a mismatch between the prediction result and the actual production plan. SUMMARY
[0005] The embodiments of the present application provide an industrial park carbon emission prediction method, device and medium, to solve the technical problem of how to achieve high-precision, short-term prediction and decision support capability of carbon emission prediction.
[0006] In a first aspect, an embodiment of the present application provides an industrial park carbon emission prediction method, which comprises: setting a physical layer state of a digital twin model of an industrial park to be consistent with a current real world, and inputting a production plan in a preset time period into a decision layer of the digital twin model, wherein a node in the digital twin model is a trained intelligent agent, and one enterprise or key production unit in the industrial park corresponds to one intelligent agent; based on the randomness of a strategy, driving the digital twin model to deduce at least once based on the production plan to generate a carbon emission evolution trajectory set, wherein the carbon emission evolution trajectory set comprises at least one trajectory, the trajectory records a time sequence change of total carbon emission of the industrial park from a current time to a future target time, and a length of the preset time period is greater than or equal to a length between the current time and the future target time; inputting the carbon emission evolution trajectory set into a trained calibration model to obtain a residual output by the calibration model; obtaining a carbon emission prediction value of the industrial park at the future target time according to a sum of an average value corresponding to the carbon emission evolution trajectory set and the residual, wherein the average value is used to indicate an average value of total carbon emission of at least one trajectory at the future target time; and obtaining a confidence interval corresponding to the industrial park at the future target time based on a distribution of at least one trajectory in the carbon emission evolution trajectory set.
[0007] In a second aspect, an embodiment of the present application further provides an industrial park carbon emission prediction device, which comprises: at least one processor; and a memory in communication connection with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the industrial park carbon emission prediction method in the first aspect.
[0008] In a third aspect, an embodiment of the present application further provides a computer storage medium storing computer executable instructions, and the computer executable instructions are executed to implement the industrial park carbon emission prediction method in the first aspect.
[0009] The industrial park carbon emission prediction method, device and medium provided by the embodiment of the present application have the following beneficial effects: In the embodiment of the present application, the physical layer state of the digital twin model of the industrial park can be set to be consistent with the current real world, and the production plan in the preset time period is input into the decision layer of the digital twin model, and then based on the randomness of the strategy, the digital twin model is driven to deduce at least once based on the production plan to generate a set of carbon emission evolution trajectories, and then the set of carbon emission evolution trajectories is input into the trained calibration model to obtain the residual output by the calibration model. Finally, according to the sum of the average value corresponding to the set of carbon emission evolution trajectories and the residual, the carbon emission prediction value of the industrial park at the future target time is obtained, and based on the distribution of at least one of the trajectories in the set of carbon emission evolution trajectories, the confidence interval corresponding to the industrial park at the future target time is obtained. In this way, the prediction is changed from passive data analysis to active scenario deduction, the accuracy of the obtained carbon emission prediction value is high, the timeliness is significantly improved, and the carbon emission prediction value and the confidence interval have decision support capability. BRIEF DESCRIPTION OF DRAWINGS
[0010] The accompanying drawings used to provide further understanding of the present application and form a part of the present application, and the illustrative embodiments of the present application and the description thereof are used to explain the present application and do not constitute improper limitations on the present application. In the drawings: Figure 1 A flowchart of an industrial park carbon emission prediction method provided by an embodiment of the present application; Figure 2 An internal structure schematic diagram of an industrial park carbon emission prediction device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0011] To make the purpose, technical scheme and advantages of the present application clearer, the technical scheme of the present application will be described in detail below with reference to the embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0012] The embodiment of the present application provides an industrial park carbon emission prediction scheme, and the technical scheme proposed by the embodiment of the present application will be described in detail below with reference to the drawings.
[0013] Figure 1 A flowchart of an industrial park carbon emission prediction method provided by an embodiment of the present application. As shown in Figure 1 The industrial park carbon emission prediction method provided by the embodiment of the present application specifically includes the following steps: Step 101, set the physical layer state of the digital twin model of the industrial park to be consistent with the current real world, and input the production plan in the preset time period into the decision layer of the digital twin model.
[0014] Among them, the nodes in the digital twin model are trained intelligent agents, and one enterprise or key production unit in the industrial park corresponds to one intelligent agent.
[0015] In actual application, the digital twin model is a data-driven virtual model, which is a mirror / copy of the physical entity (real world). Using the digital twin model, the invisible state of the physical world (such as internal wear and tear of equipment, energy flow direction) can be converted into visual and readable data, achieving "one-stop" global control. Intelligent agent refers to an entity that can perceive the environment and take autonomous action to achieve a specific goal.
[0016] In the embodiments of the present application, the physical layer state of the digital twin model of the industrial park can be set to be consistent with the current real world, which can ensure that the digital twin model can reflect various data of the industrial park in real time. And input the production plan in the preset time period into the decision layer of the digital twin model, so that the digital twin model can simulate the industrial process based on the production plan based on the trained intelligent agent, and then perform carbon emission prediction.
[0017] Step 102, based on the randomness of the strategy, drive the digital twin model to evolve at least once based on the production plan to generate a set of carbon emission evolution trajectories.
[0018] Among them, the set of carbon emission evolution trajectories includes at least one trajectory, which records the time sequence change of the total carbon emission of the industrial park from the current time to the future target time, and the length of the preset time period is greater than or equal to the length between the current time and the future target time.
[0019] In actual application, the intelligent agent strategy has randomness, that is, when the intelligent agent selects an action, it does not fixedly select a certain determined action, but selects different actions according to a certain probability distribution. Therefore, In the embodiments of the present application, based on the randomness of the above strategy, the digital twin model can be repeatedly driven to evolve based on the above production plan. Each time the evolution is performed, each intelligent agent decides to adjust the production power, switch the energy use, etc. according to its strategy and current state, and then obtains the total carbon emission of the industrial park. In this way, a set of carbon emission evolution trajectories can be obtained.
[0020] It should be noted that the above production plan can be an initial production plan, or a plan that is regenerated after the plan is evolved and the carbon emission exceeds the standard, which needs to be determined according to the actual situation.
[0021] In a conventional manner, the total carbon emission of next week or next month is generally predicted based on the historical total monthly electricity fee and total output. In the embodiment of the present application, the carbon emission trajectory of the entire industrial park is synthesized from bottom to top by simulating the dynamic behavior and interaction of each enterprise or key production unit in the industrial park, so that the timeliness of the prediction can be improved.
[0022] In step 103, the set of carbon emission evolution trajectories is input into the trained calibration model to obtain the residual output by the calibration model.
[0023] In actual application, digital twinning simulation is mechanism-driven, but there may be systematic deviation from the physical world, which can be calibrated by a data-driven calibration model. In the embodiment of the present application, the set of carbon emission evolution trajectories is input into the trained calibration model M calibrate , to obtain the residual output by the calibration model, which can indicate the difference between the predicted value and the true value.
[0024] The mechanism model of digital twinning is scientific, but it is always based on simplification and assumption. The calibration model can learn and correct these inherent systematic errors. Moreover, there are a large number of factors in the real industrial environment that are difficult to completely describe by mechanism formula (such as equipment aging, unmeasured loss, operator habit, etc.), and the data-driven calibration model can capture these complex patterns. In this way, the prediction accuracy can be significantly improved. Moreover, the calibration model changes the digital twinning model from a theoretically perfect but possibly "out of touch with reality" laboratory model to a reliable, credible and usable decision support system in the real industrial environment. It is a key technical link to obtain maximum accuracy improvement at minimum cost without giving up the strong deduction ability of the mechanism model, and fully plays the synergistic advantages of mechanism and data-driven.
[0025] In step 104, the carbon emission prediction value of the industrial park at the future target time is obtained according to the sum of the average value corresponding to the set of carbon emission evolution trajectories and the residual.
[0026] The average value is used to indicate the average value of the total carbon emission of at least one of the trajectories at the future target time.
[0027] In the embodiment of the present application, the carbon emission prediction value of the industrial park at the future target time is determined by the sum of the average value of the total carbon emission of at least one of the trajectories at the future target time (the average value C sim (tp) at the future target time tp=t0+T) and the above residual, and the formula is as follows:
[0028] The carbon emission prediction value thus obtained can more accurately describe the carbon emission of the industrial park at the future target time.
[0029] Step 105: Obtain a confidence interval corresponding to the industrial park at the future target time based on the distribution of at least one of the trajectories in the set of carbon emission evolution trajectories.
[0030] In the embodiments of the present application, the confidence interval corresponding to the industrial park at the future target time can be obtained based on the distribution of at least one of the trajectories in the set of carbon emission evolution trajectories. The confidence interval corresponds to the carbon emission prediction value, and the confidence interval is constructed by the standard error (or quantile) of the set of carbon emission evolution trajectories, directly quantifying the uncertainty of the prediction. For example, the 95% confidence interval can be expressed as [θ-1.96SE, θ+1.96SE]. - 1.96 SE, + 1.96 SE], which can provide a risk assessment basis for decision-making.
[0031] The confidence interval is an interval constructed based on a sample statistic (such as a sample mean or a sample proportion), and has the form [θ-margin of error, θ+margin of error], where θ is the sample statistic, and the margin of error is determined by the confidence level, the standard error, and the distribution characteristics. If a 95% confidence level is set, it means that in 100 repeated samplings, 95 times of the constructed confidence intervals will contain the true population parameter (such as the population mean μ).
[0032] In the above carbon emission prediction scheme, the confidence interval quantifies the uncertainty of the prediction by the standard error (such as SE=standard deviation of the set of trajectories at tp, and N is the number of trajectories) of the set of trajectories. For example, the 95% confidence interval [θ-1.96SE, θ+1.96SE] can be expressed as [θ-1.96SE, θ+1.96SE]. -1.96 SE, + 1.96 SE] can represent that the prediction value at the future target time tp has a 95% probability of falling within the interval. Therefore, the performance of the current production plan can be determined based on the carbon emission prediction value and the confidence interval, and the user can be assisted in determining the production plan.
[0033] In the traditional way, the user only knows that the carbon emission has exceeded the limit when the carbon emission monitor alarms, which is too late. However, in the method provided in the application, the prediction result can show that there is a risk of exceeding the carbon emission limit several hours in advance, which can foresee the future risk and realize "early warning". Instead of passively responding to the alarm, the user can actively issue a warning. For example, "According to the current operating situation and production plan, the total carbon emission of the industrial park will exceed the limit in 4 hours with a confidence of 95%." This wins a valuable time window for adjusting the dispatch. Moreover, the consequences of different dispatch strategies can be evaluated and compared. Assuming that the prediction shows that the carbon emission will exceed the limit in 4 hours, the user can test multiple intervention measures, such as asking high-energy-consuming enterprise A to reduce its production power by 10%, or adjusting the energy center to increase the proportion of clean energy (such as photovoltaic) and starting the standby energy storage. Then in the digital twin model, by modifying the production plan of the decision layer and re-running the forward-looking scenario deduction, the prediction result of carbon emission is obtained.
[0034] In the embodiment of the application, the physical layer state of the digital twin model of the industrial park can be set to be consistent with the current real world, and the production plan in the preset time period can be input into the decision layer of the digital twin model, and then based on the randomness of the strategy, the digital twin model is driven to deduce at least once based on the production plan to generate a set of carbon emission evolution trajectories, and then the set of carbon emission evolution trajectories is input into the calibrated model to obtain the residual output by the calibrated model. Finally, according to the sum of the average value corresponding to the set of carbon emission evolution trajectories and the residual, the carbon emission prediction value of the industrial park at the future target time is obtained, and based on the distribution of at least one of the trajectories in the set of carbon emission evolution trajectories, the confidence interval corresponding to the industrial park at the future target time is obtained. In this way, the prediction is changed from passive data analysis to active scenario deduction, the accuracy of the obtained carbon emission prediction value is high, the timeliness is significantly improved, and the carbon emission prediction value and the confidence interval have decision support capability.
[0035] In actual application, before carbon emission prediction is performed in the industrial park, a digital twin model of the industrial park needs to be obtained, and the nodes of the digital twin model are trained intelligent agents. It should be noted that if the industrial park is performing carbon emission prediction for the first time, a digital twin model and intelligent agents need to be constructed, or a digital twin model of an industrial park similar to the industrial park is found for application. If the industrial park is not performing carbon emission prediction for the first time, the digital twin model of the last time can be directly used, and if the industrial park has changed, fine tuning can be performed, so that resources can be saved and the speed of carbon emission prediction can be improved.
[0036] In a possible implementation, before the step of setting the physical layer state of the digital twin model of the industrial park to be consistent with the current real world, the method can further include the following steps: Step 1, constructing a digital twin model of the industrial park; Step 2, modeling at least one enterprise or key production unit in the industrial park as at least one agent; Step 3, constructing a multi-agent reinforcement learning environment; Step 4, performing multi-agent reinforcement learning training so that the agent makes decisions independently using its own policy network and local observation.
[0037] In the above embodiment, a digital twin model of the industrial park can be constructed, and at least one enterprise or key production unit in the industrial park can be modeled as at least one agent, where the enterprise or key production unit modeling is in a one-to-one relationship with the agent. Then, a multi-agent reinforcement learning environment is constructed, and based on the above learning environment, multi-agent reinforcement learning training is performed so that the agent can make decisions independently using its own policy network and local observation. This one-to-one modeling and multi-agent reinforcement learning architecture can find a feasible path to achieve global optimization of the system while respecting the distributed, privacy, and complex coupling constraints of the real world.
[0038] In a possible implementation, the step of constructing a digital twin model of the industrial park can include: performing fine-grained modeling on key energy-consuming equipment in the industrial park to construct a physical layer of the digital twin model; constructing a directed graph based on the enterprises or key production units in the industrial park to construct a network layer of the digital twin model, where the directed graph is used to describe the coupling relationship between the enterprises or key production units, and the nodes in the directed graph are used to represent the enterprises or key production units in the industrial park; constructing a decision layer of the digital twin model by integrating or simulating the production plans of the enterprises or the key production units.
[0039] In the above embodiment, the digital twin model includes three layers: a physical layer, a network layer, and a decision layer.
[0040] When constructing the physical layer, fine-grained modeling can be performed on the key energy-consuming equipment in the park. For example, for a boiler, its energy consumption model can be represented as:
[0041] wherein, is the energy consumption (such as coal consumption) of the boiler at time t, is the steam output power at time t, is the real-time operation efficiency of the boiler, is the ambient temperature. This function f can be a mechanism formula based on the equipment manual, or a regression model fitted by historical data. This layer communicates with the SCADA, EMS system in real time through OPC UA, MQTT, etc. to synchronize , key parameters.
[0042] In this way, the key energy-consuming equipment (such as boilers) that affect global energy consumption and carbon emissions are modeled in detail, ensuring the accuracy of core link predictions.
[0043] When building the network layer, a directed graph can be constructed to describe the coupling relationship between enterprises or key production units.
[0044] The node represents an enterprise or a key energy center (such as a thermal power plant).
[0045] The directed edge (i, j) ∈ represents the energy or material flow from node i to node j. The weight of the edge represents the flow at time t, for example, the steam flow from thermal power plant i to chemical plant j. The dynamic evolution of this layer is driven by the device output of the physical layer and the scheduling decisions of the agent.
[0046] Traditional models are difficult to quantify the impact of one enterprise's decisions on another. The above directed graph can clearly indicate how downstream enterprises depend on upstream energy centers. When the state of a node (enterprise) changes (such as adjusting production in the physical layer), the model can accurately simulate how this change spreads like a ripple to the entire network along the directed edges, affecting the external state of other nodes.
[0047] When building the decision layer, the production plan of an enterprise or key production unit can be integrated or simulated. This layer can provide macro constraints such as the planned production of enterprise i at time t , which is an important input for agent decision-making. The decision layer ensures that the digital twin is synchronized with real business goals, making it no longer a physical simulation that is detached from reality. Instead of simulating physical processes in isolation, the model takes the most important business driving factor, the production plan, as the top-level input. This means that the operation of the digital twin is always centered around the core goal of "how to efficiently and low-carbon complete the production task". This makes the simulation deduction and optimization results have direct business guidance significance. All analyses (such as energy consumption analysis, carbon emission prediction) are carried out under the given production plan, and the conclusions and scheduling recommendations can be directly understood and applied by the business department.
[0048] In the above embodiments, the physical layer provides the device-level real dynamics and is the basis for calculating energy consumption costs. The network layer defines the interaction environment between agents (enterprises), and the coupling state of the agents directly originates from this. The decision layer provides agents with macro-level constraints that they must adhere to (such as planned output), ensuring the feasibility of the learning strategy. This structured state representation and clear system boundaries are prerequisites for subsequent multi-agent reinforcement learning.
[0049] In one possible implementation, constructing the multi-agent reinforcement learning environment may include: Define a Markov decision process tuple for each agent; A multi-agent reinforcement learning environment is constructed based on the Markov decision process tuples of the agents.
[0050] In the above embodiments, a Markov Decision Process (MDP) tuple can be defined for each agent to construct a multi-agent reinforcement learning environment. Specifically, in multi-agent reinforcement learning, each agent can define its own Markov Decision Process (MDP) tuple (such as a state set, action set, transition function, reward function, and discount factor). Then, the MDP tuples of all agents are merged or coupled to form a unified multi-agent environment framework for simulating and training interactions and collaborations between agents. This design is based on the Markov property (future states depend only on the current state and actions) and extends to multi-agent scenarios, enabling each agent to learn autonomously and adapt to environmental changes.
[0051] In practical applications, in addition to methods based on MDP tuples, constructing multi-agent reinforcement learning environments can also employ methods based on environment models, interaction mechanisms, evolutionary strategies, self-supervised learning, graph neural networks, federated learning, etc., without specific limitations.
[0052] In practical applications, a Markov decision process consists of the following quintuple: State space (S): The set of all possible states that a system can be in; Action space (A): The set of all possible actions that can be taken in a given state; Transition Probability (P): Given the current state s and action a, the probability of transitioning to the next state (s'): P(s'|s, a); Reward Function (R): the immediate reward obtained from transitioning from state s to state s' after taking action a: R(s, a, s') ; Discount Factor (γ): controls the importance of future rewards, with a value range of (0-1).
[0053] In one possible implementation, the defining the Markov Decision Process tuple of the agent for each agent includes: defining the state space of the agent based on the internal state vector, the external state vector, and the coupling state vector; defining the action space of the agent based on the actions executable by the agent; constructing a reward function based on profit and carbon emission cost, wherein the reward function is used to balance economy and environment.
[0054] In the above embodiment, for each agent a Markov Decision Process tuple is defined .
[0055] wherein, for the state space the complete state of the agent at time t is a multi-dimensional vector, defined as:
[0056] wherein, is the internal state vector, such as , representing actual production, inventory level, and equipment status, respectively. is the external environment state vector, such as , representing park energy price, upstream material supply status, and downstream product demand, respectively. is the coupling state vector, from the network layer, such as , representing the key medium flow in and out of the node.
[0057] the action space of the agent at time t is the action executable by the agent. It can be discrete, such as ∈{production line 开启 , production line 关闭}; or continuous, such as , representing the adjustment amount of production power.
[0058] A typical composite action can be: where, is a Boolean variable, determining whether to use park steam or self-provided boiler.
[0059] Reward function is the core of guiding the agent to learn. A balanced economic and environmental reward function is designed as follows: where, is the profit at time t. is the carbon emission cost at time t, is the carbon emission, is the real-time carbon price (which can be a simulated internal carbon price). and are weight coefficients, used to adjust the preferences of the agent. For example, in the initial stage can be set to a small value, and then gradually increased after the agent learns the basic production optimization, guiding it to explore low-carbon strategies.
[0060] In one possible implementation, the performing multi-agent reinforcement learning training can include: initializing a policy network and a value network for each agent; In the training process, the value network of each agent accesses the actions and states of all agents, and the policy network of each agent outputs actions according to its own local observation, wherein the objective function of the agent is to maximize the expected cumulative reward, and the objective function corresponding to the policy network is to improve the policy using the gradient of the value network.
[0061] In the above embodiment, a “centralized training, decentralized execution” (CTDE) framework can be used for multi-agent reinforcement learning training. An Actor-Critic architecture can be constructed, i.e., each agent i has its own Actor network (policy network) and Critic network (value network).
[0062] In training, centralized training is performed, and the Critic network of each agent can access the actions and states of all agents. Its objective function is to maximize the expected cumulative reward: where (s, a) is the current joint state and joint action, and (s', a') is the joint state and joint action at the next time.
[0063] The Actor network of each agent then outputs actions according to its own part of the local observation ( ). = The objective function is to improve the policy using the gradient of the Critic network:
[0064] After training, each agent i only uses its own Actor network and local observation to make decisions independently (t) without a central coordinator, which is consistent with the real-world scenario of independent decision-making by enterprises.
[0065] In one possible implementation, the policy-based randomness drives the digital twin model to evolve at least once based on the production plan to generate a set of carbon emission trajectories, including: At each time step, each agent makes a move according to its local observation and policy network, triggering changes in the state and energy consumption of the key energy-consuming equipment, where the length of the time step is less than the length of the preset time period. Through the network layer, the external state of other agents is affected, the environment state is transferred to the next time step, and the reward of each agent and the total carbon emission are calculated until the future target time is reached. Record the sequence of the total carbon emissions of the industrial park in the preset time period to form a trajectory. Repeat the above evolution process at least once to obtain a set of carbon emission evolution trajectories.
[0066] In the above embodiment, starting from the current time , advance T (such as 4 hours) with a step size of Δt (such as 1 minute), so that the agent makes a move every minute, and the prediction of carbon emissions can be accurate to the hour level.
[0067] At each time step t, each agent can make a move according to its observation and policy . The action is input to the physical layer of the digital twin model, triggering changes in the state and energy consumption of the key energy-consuming equipment. The changes in the physical layer are propagated through the network layer, affecting the external state of other agents . The environment state is transferred to s(t+1), and the reward of each agent is calculated . In practical applications, carbon emissions can be calculated based on changes in the energy consumption of key energy-consuming equipment.
[0068] During the above advancing process, the total carbon emissions of the park sequences, forming a complete evolutionary trajectory .
[0069] Due to the inherent randomness in reinforcement learning strategies (e.g. Ornstein-Uhlenbeck process for exploration), a Monte Carlo simulation can be performed, repeating the above deduction process K times (e.g. K = 50) to obtain a set of trajectories { , ,..., }. In reinforcement learning training, randomness can be intentionally introduced into the strategy (e.g. Ornstein-Uhlenbeck process) to encourage exploration. Even in the execution phase, this randomness can be retained to simulate the slight decision fluctuations in the real world. Monte Carlo simulation fully captures this inherent randomness in the strategy through multiple runs. Each time the agent makes slightly different decisions due to random noise (e.g. the exact time point for shutting down the production line is different, the amplitude of power adjustment is slightly different), resulting in different final system states (e.g. total carbon emissions). This reflects the fact that real-world decision makers are not completely predictable machines, and their decisions are influenced by various factors. Monte Carlo simulation faithfully reflects how this microscopic decision uncertainty converges into macroscopic systemic risk. Managers see a more realistic, uncertain future rather than an idealized, deterministic future.
[0070] In one possible implementation, the calibration model is trained by: Collecting historical data of the industrial park over a period of time; For each historical time point, use the environmental state at that time to perform retrospective deduction in the digital twin model to obtain a simulated carbon emission sequence; Use the simulated carbon emission sequence and the real monitoring sequence as input-output pairs to train the calibration model.
[0071] In practical applications, a lightweight Gated Recurrent Unit (GRU) network can be used as the calibration model .
[0072] During model training, historical data over a period of H can be collected. For each historical time point , use the environmental state at that time to perform retrospective deduction in the digital twin to obtain a simulated carbon emission sequence . Use and the real monitoring sequence as input-output pairs to train the GRU model: That is, let the GRU learn the residual error between the simulation value and the true value .
[0073] In practical applications, the above calibration model can be used for the calibration process of predicting carbon emissions in the industrial park, which can be updated regularly or once per prediction of a future target time, without limitation. In order to save resources, if park A is similar to park B, they can share a calibration model, without limitation.
[0074] The application of the above industrial park carbon emission prediction method in practice can be illustrated by examples, for example, predicting the carbon emissions of a chemical industrial park in the next 4 hours, which can include the following steps: Step 1, build a digital twin model for the chemical industrial park.
[0075] Wherein, when building the physical layer, the boiler efficiency curve of the park thermal power plant and the energy consumption-yield relationship of each chemical plant main reaction kettle can be modeled.
[0076] When building the network layer, a steam pipe network diagram can be built to determine which enterprise supplies steam to which enterprise and the transmission loss of the pipe network.
[0077] When building the decision layer, the future 24-hour production plan of each enterprise can be accessed.
[0078] Steps 2-4, multi-agent reinforcement learning modeling and training.
[0079] The thermal power plant and three core chemical plants are set as four agents.
[0080] State: The state of the thermal power plant includes boiler load, coal warehouse inventory, and power grid price; the state of the chemical plant includes current order quantity, reaction kettle temperature, and steam receiving pressure.
[0081] Action: The action of the thermal power plant is to adjust the generator set load; the action of the chemical plant is to adjust the reaction rate or apply for steam increment.
[0082] Reward: The reward of the thermal power plant is (electricity sales revenue + steam sales revenue - fuel cost - carbon emission cost); the reward of the chemical plant is (product value - energy cost - carbon emission cost).
[0083] In the training process, a Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm can be used for training until the policy of each agent converges. MADDPG is a reinforcement learning algorithm specially designed for multi-agent environments. It solves the problems of multi-agent learning instability and credit assignment difficulty through "centralized training", and ensures the practicability of the algorithm in real distributed systems through "decentralized execution".
[0084] Step 5, forward-looking deduction is performed.
[0085] Suppose the current time is 8 am, and it is known that a chemical plant will have equipment maintenance at 2 pm. Input this information into the digital twin system, and let the above four agents run autonomously from 8 am to 12 noon.
[0086] Record the total carbon emissions of the entire park every minute to obtain an evolutionary trajectory. Repeat the deduction 50 times to obtain 50 trajectories.
[0087] Step 6, prediction calibration and output.
[0088] Use the GRU model to learn the deviation between the "simulation trajectory" and the "measured data" in the past week.
[0089] Calibrate the 50 trajectories obtained in step 5, take their mean and confidence interval as the final prediction result output: "The total carbon emissions of the park in the next 4 hours are XX tons, and the 95% confidence interval is [XX-Δ, XX+Δ]".
[0090] In the case where the above obtained result does not meet the expectation, the production plan can be adjusted and prediction can be performed again. Meanwhile, in the case where the adjusted production plan does not meet the expectation, the production plan can be adjusted again and prediction can be performed again.
[0091] The above is a method embodiment proposed by the present application. Based on the same inventive concept, the present application embodiment also provides an industrial park carbon emission prediction device, the structure of which is shown in Figure 2 .
[0092] Figure 2 An internal structure diagram of an industrial park carbon emission prediction device provided by the present application embodiment is shown in Figure 2 . As shown in the figure, the device comprises: at least one processor 201; and a memory 202 in communication connection with the at least one processor; The memory 202 stores instructions executable by the at least one processor 201, and the at least one processor 201 executes the instructions to enable the at least one processor 201 to perform the industrial park carbon emission prediction method described above.
[0093] Some embodiments of the present application provide a non-transitory computer storage medium corresponding to Figure 1 The computer executable instructions are configured to perform the industrial park carbon emission prediction method described above.
[0094] Each of the embodiments in the present application is described in a progressive manner, and the same or similar parts of each of the embodiments can be referred to each other. Each of the embodiments mainly describes the difference from other embodiments. In particular, the IoT device and medium embodiments are basically similar to the method embodiments, and thus are described simply. The relevant parts can be referred to the description of the method embodiments.
[0095] The system and medium provided by the embodiments of the present application are one-to-one corresponding to the method, and thus the system and medium have similar beneficial technical effects to the method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the system and medium will not be described here.
[0096] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can be in the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0097] The present application is described with reference to flowcharts and / or block diagrams according to the method, device (system), and computer program product of the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to produce a machine, so that the instructions executed by the computer or other programmable data processing devices produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The device that implements the functions specified in one or more flows and / or blocks in the flowcharts and / or block diagrams. Figure 1 The device that implements the functions specified in one or more flows and / or blocks in the flowcharts and / or block diagrams.
[0098] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the Figure 1 function specified in the flow or flows and / or blocks Figure 1 of the block or blocks.
[0099] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions that are executed on the computer or other programmable apparatus provide steps for implementing the Figure 1 function specified in the flow or flows and / or blocks Figure 1 of the block or blocks.
[0100] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0101] The memory can include non-persistent memory and / or volatile memory, such as a random access memory (RAM) including a cache area for the temporary storage of data. The memory can also include non-volatile memory, such as a read only memory (ROM), EPROM, EEPROM, or flash memory. The memory can be another form of computer-readable media.
[0102] Computer-readable media includes permanent and non-permanent, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to a computing device. According to the definition herein, computer-readable media does not include transitory media, such as modulated data signals and carrier waves.
[0103] It is also to be noted that the terms "comprising", "including", and any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises a... " does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0104] The above description is merely illustrative of the application, and not restrictive. Various modifications and changes can become apparent to those skilled in the art. Incorporating any modification, equivalent substitution, improvement, etc. within the spirit and principle of the application, shall be included in the scope of the claims of the application.
Claims
1. An industrial park carbon emission prediction method, characterized in that, The method comprises the following steps: setting the physical layer state of the digital twin model of the industrial park to be consistent with the current real world, and inputting the production plan in a preset time period into the decision layer of the digital twin model, wherein the nodes in the digital twin model are trained intelligent agents, and one enterprise or key production unit in the industrial park corresponds to one intelligent agent; based on the randomness of the strategy, driving the digital twin model to evolve at least once based on the production plan to generate a set of carbon emission evolution trajectories, wherein the set of carbon emission evolution trajectories includes at least one trajectory, which records the time sequence change of the total carbon emission of the industrial park from the current time to the future target time, and the length of the preset time period is greater than or equal to the length between the current time and the future target time; inputting the set of carbon emission evolution trajectories into the trained calibration model to obtain the residual output by the calibration model; obtaining the carbon emission prediction value of the industrial park at the future target time according to the sum of the average value corresponding to the set of carbon emission evolution trajectories and the residual, wherein the average value is used to indicate the average value of the total carbon emission of at least one trajectory at the future target time; based on the distribution of at least one trajectory in the set of carbon emission evolution trajectories, obtaining the confidence interval corresponding to the industrial park at the future target time.
2. The method of claim 1, wherein, Before the step of setting the physical layer state of the digital twin model of the industrial park to be consistent with the current real world, the method further comprises: constructing a digital twin model of the industrial park; modeling at least one enterprise or key production unit in the industrial park as at least one intelligent agent; constructing a multi-agent reinforcement learning environment; performing multi-agent reinforcement learning training to enable the intelligent agent to make decisions independently using its own strategy network and local observation.
3. The method of claim 2, wherein, The step of constructing a digital twin model of the industrial park comprises: performing fine modeling on key energy-consuming equipment in the industrial park to construct a physical layer of the digital twin model; based on the enterprises or key production units in the industrial park, constructing a directed graph to construct a network layer of the digital twin model, wherein the directed graph is used to describe the coupling relationship between enterprises or key production units, and the nodes in the directed graph represent the enterprises or key production units in the industrial park; by integrating or simulating the production plans of the enterprises or the key production units, constructing a decision layer of the digital twin model.
4. The method of claim 2, wherein, The step of constructing a multi-agent reinforcement learning environment comprises: defining a Markov decision process tuple of the intelligent agent for each intelligent agent; based on the Markov decision process tuple of the intelligent agent, constructing a multi-agent reinforcement learning environment.
5. The method of claim 4, wherein, The step of defining a Markov decision process tuple of the intelligent agent for each intelligent agent comprises: based on the internal state vector, external state vector and coupling state vector of the intelligent agent, defining the state space of the intelligent agent; based on the actions executable by the intelligent agent, defining the action space of the intelligent agent; based on the profit and carbon emission cost, constructing a reward function, wherein the reward function is used to balance the economy and the environment.
6. The method of claim 2, wherein, The performing multi-agent reinforcement learning training comprises: initializing a policy network and a value network for each agent; During the training process, the value network of each agent accesses the actions and states of all agents, and the policy network of each agent outputs actions according to its own local observation, wherein the objective function of the agent is to maximize the expected cumulative reward, and the objective function corresponding to the policy network is to improve the policy by using the gradient of the value network.
7. The method of claim 3, wherein, Once deduction is performed, the digital twin model is driven to advance the preset time period in steps of target length from the current time, and the process is as follows: At each time step, each agent is driven to make actions according to its local observation and policy network; The actions are input to the physical layer of the digital twin model, triggering changes in the state and energy consumption of the key energy-consuming equipment; The changes in the physical layer are propagated through the network layer to affect the external states of other agents; The state of the digital twin model is transferred to the next time step, and the rewards of each agent and the total carbon emissions are calculated.
8. The method of claim 1, wherein, The calibration model is trained in the following way: Collecting historical data of the industrial park over a period of time; At each historical time point, the environmental state at that time is used to perform backtracking deduction in the digital twin model to obtain a simulated carbon emission sequence; The simulated carbon emission sequence and the real monitoring sequence are used as input-output pairs to train the calibration model.
9. An industrial park carbon emission prediction device characterized by, The device comprises: at least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform an industrial park carbon emission prediction method according to any one of claims 1-8.
10. A computer storage medium storing computer-executable instructions, which, when executed by a processor, cause the processor to perform acts comprising: The computer executable instructions, when executed, implement an industrial park carbon emission prediction method according to any one of claims 1-8. The computer executable instructions, when executed, implement an industrial park carbon emission prediction method according to any one of claims 1-8.