Trusted interpretation method and system for intelligent scheduling decision-making in heterogeneous energy systems
By building a trusted reinforcement learning model, the problem of insufficient security and interpretability of DRL models in heterogeneous energy system scheduling is solved, providing a clear explanation of the decision-making process, and enhancing the credibility and scheduling efficiency of the model.
Patent Information
- Application Number
- CN202410408471.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-07
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2044-04-07
AI Technical Summary
Existing DRL models have problems with insufficient security and interpretability in heterogeneous energy system scheduling, especially when dealing with complexity and multidimensional data. Traditional methods cannot effectively balance security and efficiency, and humans do not have sufficient trust in DRL models.
The trusted reinforcement learning method is adopted to build a heterogeneous energy system model for trusted intelligent scheduling, perform CMDP modeling, and solve it using the S-DRL model, and combine the S-DRL model interpretation method of AP-CSI to provide clear explanation of the decision process, enhancing the credibility and decision quality of the model.
It realizes that while ensuring the safety of system operation, it provides clearer and more intuitive explanation of decision-making process, enhances the credibility and decision-making quality of the model, and improves human trust and scheduling efficiency in the DRL model.
Smart Images

Figure CN118690632B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of energy management technology based on artificial intelligence, and in particular to a credible interpretation method and system for intelligent scheduling decision-making in heterogeneous energy systems. Background Art
[0002] The rapid development of measurement and communication infrastructure for heterogeneous energy systems (HES) has generated massive amounts of data from multidimensional and heterogeneous sources, transforming the system into a complex, high-dimensional, heterogeneous, time-varying, and nonlinear system. Against this backdrop, the application of artificial intelligence (AI) in HES scheduling can significantly improve the decision-making performance and intelligence level of the scheduling system. By assisting with analysis and autonomous learning during the decision-making process, AI is expected to meet safe operation requirements, improve the speed and efficiency of decision-making in response to abnormal and unexpected situations, and thus reduce the workload of manual scheduling.
[0003] Currently, DRL models have the potential to solve complex scheduling and control problems in HES. Numerous AI models have been applied to heterogeneous energy systems. However, despite significant progress in improving scheduling efficiency, the application of DRL models in optimizing the scheduling of heterogeneous energy systems is currently limited to decision support and edge technologies, and their support for the scope and degree of control remains far from sufficient. In the field of HES scheduling, human trust in DRL models remains insufficient. This lack of trust is primarily due to issues with the security and interpretability of DRL models.
[0004] In heterogeneous energy systems, credible DRL models are crucial for establishing a mutually trusting human-machine collaborative decision-making process. The introduction of safety can enhance human trust in DRL model results and improve their reliability. Simultaneously, the introduction of explainability improves the quality of human-machine interaction and enhances operator understanding; both are indispensable.
[0005] Overall, the DRL models currently used in HES have some shortcomings in ensuring safe operation and providing explainable decisions. In particular, the free exploration mechanism of traditional DRL models ignores safety constraints, introducing potential risks, while existing solutions fail to effectively balance safety and efficiency. Furthermore, in terms of explainability, although some methods have attempted to improve model transparency through feature analysis, they remain insufficient when dealing with the complexity and multidimensional data in heterogeneous energy systems. Therefore, there is an urgent need for a method that can explain the decision-making process of these complex AI models while ensuring safety. Summary of the Invention
[0006] The purpose of the present invention is to provide a credible interpretation method and system for intelligent scheduling decision-making of heterogeneous energy systems, so as to solve at least one technical problem existing in the above-mentioned background technology.
[0007] In order to achieve the above object, the present invention adopts the following technical solutions:
[0008] In a first aspect, the present invention provides a trusted interpretation method for intelligent scheduling decision-making in heterogeneous energy systems, comprising:
[0009] Construct a heterogeneous energy system model for trusted intelligent dispatch, including: determining power balance constraints, determining HES operation constraints, determining PV equipment reactive output constraints, determining the steady-state model of the thermal network, determining GB operation constraints, determining CHP operation constraints, and determining EB operation constraints;
[0010] Conducting CMDP modeling on the heterogeneous energy system model for trusted intelligent scheduling;
[0011] The S-DRL model is used to solve the CMDP model;
[0012] The S-DRL model interpretation method of AP-CSI is used to interpret the solution results.
[0013] In a second aspect, the present invention provides a trusted interpretation system for intelligent scheduling decisions in heterogeneous energy systems, comprising:
[0014] Model building device, used to build a heterogeneous energy system model for trusted intelligent scheduling, including: determining power balance constraints, HES operation constraints, photovoltaic equipment reactive output constraints, thermal network steady-state model, GB operation constraints, CHP operation constraints, and EB operation constraints;
[0015] A model conversion device, used for performing CMDP modeling on the heterogeneous energy system model for trusted intelligent scheduling;
[0016] A model solving device, used for solving the CMDP model using the S-DRL model;
[0017] The result interpretation device is used to interpret the solution results using the S-DRL model interpretation method of AP-CSI.
[0018] In a third aspect, the present invention provides a non-transitory computer-readable storage medium, which is used to store computer instructions. When the computer instructions are executed by a processor, the trusted interpretation method for intelligent scheduling decisions of heterogeneous energy systems as described in the first aspect is implemented.
[0019] In a fourth aspect, the present invention provides a computer device comprising a memory and a processor, wherein the processor and the memory communicate with each other, the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the trusted interpretation method for intelligent scheduling decision-making of heterogeneous energy systems as described in the first aspect.
[0020] In a fifth aspect, the present invention provides an electronic device comprising: a processor, a memory, and a computer program; wherein the processor is connected to the memory, and the computer program is stored in the memory. When the electronic device is running, the processor executes the computer program stored in the memory so that the electronic device executes instructions for implementing the trusted interpretation method for intelligent scheduling decisions of heterogeneous energy systems as described in the first aspect.
[0021] Explanation of terms:
[0022] (1) Heterogeneous Energy Systems (HES) are systems that utilize multiple energy forms and technologies. These systems are designed to optimize energy efficiency, improve energy security, and reduce environmental impact by combining different energy resources (e.g., renewable energy sources such as thermal energy, solar energy, wind energy, hydropower, geothermal energy, and traditional fossil fuels).
[0023] (2) Trusted reinforcement learning is an advanced technique in machine learning that focuses on ensuring that scheduling behavior meets safety standards while pursuing maximum performance. This may include avoiding harm to the environment, ensuring the agent's own safety, or following specific ethical and legal guidelines. In short, trusted reinforcement learning aims to balance performance and safety, ensuring that the agent does not produce adverse consequences when pursuing its goals, and improving human trustworthiness. Trusted reinforcement learning provides an ideal solution for the management of heterogeneous energy systems because it can safely predict, schedule, and optimize strategies from historical data while also adapting to new situations and constraints. This safety, adaptability, and optimization capabilities make trusted reinforcement learning very suitable for energy management in HES.
[0024] (3) Dynamic Time Warping (DTW) is an algorithm that is used to measure the similarity between two time series, even if they are of different lengths or offset in time. This algorithm is very useful when processing time series data in fields such as speech recognition, data mining, and pattern recognition. The core idea of DTW is to find the best match between two time series. This match does not require the time series to be strictly aligned in time, making it very suitable for processing time series that may have different speeds. DTW allows time series to be stretched in time, which means that even if the events in the two sequences occur at different time points or have different durations, DTW can find the best match between them. In HES scheduling, this can help analyze and compare operation patterns on different days, even if they differ in specific time.
[0025] (4) K-means is a widely used clustering algorithm that is mainly used to divide data into a pre-specified number (k) of different groups or "clusters". The core idea of this algorithm is to find the center point of each cluster (called the centroid) so that the sum of the distances from the data points in the cluster to its centroid is minimized. The K-means algorithm is favored for its simplicity and efficiency, and is particularly suitable for processing large amounts of data. K-means groups data points so that the similarity between members in a group is high, while the similarity between members in different groups is low. This distance-based grouping method is easy to understand, making the results of the algorithm easy to interpret and verify. Each cluster has a centroid, which can be regarded as the "average" representative of all the points in the cluster. The location and characteristics of the centroid provide an intuitive understanding of the cluster and help explain the characteristics of the cluster. The results of the K-means algorithm are easy to present through visualization, especially in two-dimensional or three-dimensional space. Through graphical display, the distribution of different clusters and the location of their centroids can be intuitively seen, which increases the interpretability of the algorithm.
[0026] (5) The Shapley Value Additive Explanation Algorithm is a method for explaining machine learning models that uses the Shapley Value in game theory to assess the contribution of specific features to model predictions. In the context of heterogeneous energy systems, this means that it is possible to more clearly understand which factors (such as energy prices, supply or demand changes) have a significant impact on the choice of energy management strategies, thereby increasing decision makers' trust in the model and improving transparency.
[0027] Beneficial effects of the present invention: The trusted reinforcement learning method for heterogeneous energy systems provided by the present invention is based on the S-DRL model interpretation method of AP-CSI, which is used to achieve the interpretability of the S-DRL model, thereby optimizing the heterogeneous energy system. The SHAP method quantifies the contribution of each input variable to the ML model by analyzing the feature importance, thereby solving the interpretability problem of the DRL model in the HES environment. At the same time, the method adopts CMDP with safety constraints to describe the interaction between the agent and the environment, and combines dynamic time warping (DTW) and Kmeans clustering analysis to improve the interpretability and credibility of the model. Through this method, not only can the impact of the HES environment on the energy management strategy be quantified, but the decision-making process of the DRL model can also be explained through detailed probability analysis and feature classification, thereby achieving more accurate and safe energy management decisions. The present invention understands the internal mechanism of the proposed S-DRL model through AP-CSI, and provides a clearer and more intuitive explanation of the decision-making process while ensuring the safety of system operation, so as to enhance the credibility of the model and the quality of decision-making.
[0028] Additional advantages of the present invention will be more clearly given in the following description or learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0030] Figure 1 This is a flow chart of a trusted interpretation method for intelligent scheduling decision-making in heterogeneous energy systems according to an embodiment of the present invention.
[0031] Figure 2 This is a diagram of the trusted reinforcement learning framework described in an embodiment of the present invention. DETAILED DESCRIPTION
[0032] The embodiments of the present invention are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention and are not to be construed as limiting the present invention.
[0033] Those skilled in the art will understand that unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art to which this invention belongs.
[0034] It should also be understood that terms, such as those defined in commonly used dictionaries, should be understood to have a meaning consistent with their meaning in the context of the prior art and will not be interpreted in an idealized or overly formal sense unless as defined herein.
[0035] Those skilled in the art will appreciate that, unless otherwise stated, the singular forms "a," "an," "said," and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of the stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, and / or groups thereof.
[0036] In the description of this specification, reference to the terms "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples. Those skilled in the art may combine and integrate different embodiments or examples described in this specification, as well as features of different embodiments or examples, unless otherwise contradictory.
[0037] To facilitate understanding of the present invention, the present invention is further explained below with reference to specific embodiments in conjunction with the accompanying drawings. However, the specific embodiments do not constitute a limitation on the embodiments of the present invention.
[0038] Those skilled in the art should understand that the drawings are merely schematic diagrams of embodiments, and the components in the drawings are not necessarily necessary for implementing the present invention.
[0039] The application of artificial intelligence in the scheduling of heterogeneous energy systems can significantly improve decision-making performance and intelligence level. As a potential solution to this problem, the deep reinforcement learning (DRL) model has the potential to achieve intelligent scheduling. However, the application of DRL models in heterogeneous energy systems is mainly limited to auxiliary decision-making and edge technology, and its trust and control range need to be improved. In order to establish trust in artificial intelligence decisions and achieve efficient "human-machine" collaborative decision-making, it is necessary to focus on the credibility of scheduling decisions. To address this problem, the present invention proposes a trusted reinforcement learning method for intelligent scheduling of heterogeneous energy systems, which will provide a clearer and more intuitive explanation of the decision-making process while ensuring the safety of system operation, so as to enhance the credibility of the scheduling model and the quality of decision-making.
[0040] Example 1
[0041] In this embodiment 1, a trusted interpretation system for intelligent scheduling decision-making of heterogeneous energy systems is first provided, including: a model building device for building a heterogeneous energy system model for trusted intelligent scheduling, including: determining power balance constraints, HES operation constraints, photovoltaic equipment reactive output constraints, thermal network steady-state model, GB operation constraints, CHP operation constraints and EB operation constraints; a model conversion device for performing CMDP modeling on the heterogeneous energy system model for trusted intelligent scheduling; a model solving device for solving the CMDP model using the S-DRL model; and a result interpretation device for interpreting the solution results using the S-DRL model interpretation method of AP-CSI.
[0042] In this embodiment, the above-mentioned system is used to implement a trusted interpretation method for intelligent scheduling decisions of heterogeneous energy systems, including: constructing a heterogeneous energy system model for trusted intelligent scheduling; using the S-DRL model to solve the CMDP model; and using the AP-CSI S-DRL model interpretation method to interpret the solution results.
[0043] The CMDP modeling of the heterogeneous energy system model for trusted intelligent scheduling is carried out, including: defining the observation state and action space of the intelligent agent, determining the objective function as the minimum system operation cost and safety cost constraint, determining the reward function, defining the unit output safety constraint function and the voltage offset safety constraint function, and determining the HES scheduling model based on CMDP.
[0044] The S-DRL model is used to solve the CMDP model, including:
[0045] The goal of the S-DRL model is to meet the long-term cost-benefit C i (π)≤d i In this case, we maximize the reward Q(π), Q(π) is:
[0046]
[0047]
[0048] The process can be restructured as:
[0049]
[0050]
[0051] Among them, β i is the Lagrange multiplier;
[0052] For the Lagrange multiplier β i To update:
[0053]
[0054] Among them, η k is the learning rate of the Lagrange multiplier;
[0055] Parameters of the policy network Update formula:
[0056]
[0057] Parameters θ of the target policy network π′ Update formula:
[0058] θ π′ ←τθ π +(1-τ)θ π′
[0059] Critic network loss function L(φ Q )for:
[0060]
[0061] Where, y is the action value function output by Critic network 1; t is the target Q value at time t, which is the smaller value output by the two target value networks;
[0062] Parameters of the value network Update formula:
[0063]
[0064] Where, α φ is the value network learning rate;
[0065] Safety constraint loss function output by the Safety network for:
[0066]
[0067] Where, is the safety constraint function corresponding to different constraint conditions;
[0068] Update the security constraint parameters:
[0069]
[0070] Target value network parameter soft update formula:
[0071]
[0072]
[0073] The results of the solution are explained using AP-CSI's S-DRL model interpretation method. The method includes: in the S-DRL model, the scheduling agent generates a constrained behavior trajectory τ through interaction with the environment at each moment; after the trajectory distance is calculated, the training scenario is obtained by combining the clustering method; after the number of clusters is determined, for each data point, its distance to each cluster center is calculated and assigned to the nearest cluster center.
[0074] SHAP is used to analyze the characteristics of the clustering results. For a given prediction, the SHAP value measures the average contribution of each state feature to the agent's action. For the model f and input x, the SHAP value is defined as:
[0075]
[0076] Among them, φ i is the SHAP value of state feature i, f x (S) is the contribution of a subset of state features S to the agent's action, and n is the total number of state features.
[0077] Kmeans clustering is used: first, k data points are randomly selected as the initial cluster centers, and the Calinski-Harabasz index and Davies-Bouldin index are used to comprehensively evaluate the number of clusters. The solution formulas are:
[0078]
[0079]
[0080] Where T is the total dispersion between all points, B k is the class discreteness; σ i is the distance from all points in the i-th cluster to the cluster center c i The average distance, d(c i ,c j ) are different cluster centers c i and c j The distance between them.
[0081] Example 2
[0082] In this embodiment 2, a trusted reinforcement learning system for a heterogeneous energy system is first provided, including: a model building device for building a heterogeneous energy system model for trusted intelligent scheduling; a model conversion device for performing CMDP modeling on the heterogeneous energy system model for trusted intelligent scheduling; a model solving device for solving the CMDP model using the S-DRL model; and a result interpretation device for interpreting the solution result using the S-DRL model interpretation method of AP-CSI.
[0083] In this embodiment, the aforementioned system was used to implement a trusted reinforcement learning method for heterogeneous energy system scheduling, which was applied to intelligent HES scheduling, overcoming the shortcomings of existing HES intelligent scheduling technologies. While ensuring system security, this method provides clearer and more intuitive explanations of the decision-making process, enhancing the model's credibility and decision quality. The method includes: constructing a heterogeneous energy system model for trusted intelligent scheduling; performing CMDP modeling on the heterogeneous energy system model for trusted intelligent scheduling; solving the CMDP model using the S-DRL model; and interpreting the solution results using the S-DRL model interpretation method of AP-CSI.
[0084] The construction of a heterogeneous energy system model for trusted intelligent scheduling includes:
[0085] The power balance constraint is determined based on the active and reactive power flowing into node i at time t, the active and reactive power output of the upper power grid at node i at time t, the active and reactive output of the CHP at node i at time t, the charging and discharging power of the ES at node i at time t, the active and reactive output of the CHP at node i at time t, the active power consumed by the EB at node i at time t, the active and reactive loads at node i at time t, the conductance and susceptance between nodes i and node j, the phase angle difference between nodes i and node j at time t, and the lower and upper voltage limits of node i.
[0086] The operation constraints of the BES are determined based on the charging and discharging flags of the BES, the lower and upper limits of the ES charging power, the lower and upper limits of the ES discharging power, the state of charge of the ES at node i at time t, the lower and upper limits of the ES state of charge, the charging and discharging efficiency of the ES, and the capacity of the ES.
[0087] Determine the reactive power output constraint of photovoltaic equipment based on PV apparent power
[0088] The steady-state model of the thermal network is determined based on the specific heat capacity of water, the upper and lower correlation matrices of the thermal network nodes and pipeline branches, the transposed matrix of the lower correlation matrix, the pipeline flow diagonal matrix, the temperature attenuation coefficient diagonal matrix, the thermal power matrix flowing into the node at time t, the node temperature matrix at time t, the pipeline end temperature matrix and the pipeline external environment temperature matrix, the flow rate of the i-th pipeline, the thermal conductivity per unit length, and the length.
[0089] The operation constraints of GB are determined according to the thermal power output of GB at node i at time t, the conversion efficiency of GB, and the lower and upper limits of thermal power output.
[0090] The CHP operation constraints are determined based on the calorific value of natural gas, the conversion efficiency of CHP active output, the lower and upper limits of CHP active output, the CHP ramp rate, the CHP apparent power, and the CHP power factor.
[0091] The EB operation constraints are determined based on the thermal power output of the EB at node i at time t, the conversion efficiency of the EB, the lower and upper limits of the EB thermal power output, and the ramp rate of the EB.
[0092] The power balance constraint is determined according to the active and reactive power flowing into the node i at time t, the active and reactive power output by the upper power grid at the node i at time t, the active and reactive output of the CHP at the node i at time t, the charge and discharge power of the ES at the node i at time t, the active and reactive output of the CHP at the node i at time t, the active power consumed by the EB at the node i at time t, the active and reactive loads at the node i at time t, the conductance and susceptance between the node i and the node j, the phase angle difference between the node i and the node j at time t, and the lower and upper voltage limits of the node i, including:
[0093]
[0094]
[0095]
[0096]
[0097] Where, P i t and are the active and reactive powers flowing into node i at time t, respectively; and are the active and reactive power output by the upper power grid at node i at time t; and are the active and reactive outputs of the CHP at node i at time t, respectively; and are the charge and discharge power of ES at node i at time t; and are the active and reactive outputs of CHP at node i at time t, respectively; is the active power consumed by EB at node i at time t; and are the active and reactive loads at node i at time t; G ij and B ij are the conductance and susceptance between node i and node j respectively; is the phase angle difference between node i and node j at time t; and are the lower and upper limits of the voltage at node i, respectively.
[0098] The operation constraints of the BES are determined based on the charge and discharge flag of the BES, the lower limit and upper limit of the ES charging power, the lower limit and upper limit of the ES discharging power, the state of charge of the ES at node i at time t, the lower limit and upper limit of the ES state of charge, the charging and discharging efficiency of the ES, and the capacity of the ES, including:
[0099]
[0100]
[0101]
[0102]
[0103]
[0104] Where, and They are respectively the charge and discharge flags of BES, which are 0-1 variables; and are the lower and upper limits of ES charging power respectively; and are the lower and upper limits of ES discharge power respectively; is the state of charge of the ES at node i at time t; and are the lower and upper limits of the ES state of charge respectively; η ESc and η ESd are the charging and discharging efficiencies of ES, respectively; Q ES is the capacity of ES.
[0105] Determining the reactive output constraint of the photovoltaic device according to the apparent power of the PV includes:
[0106]
[0107] Where S PV is the apparent power of PV.
[0108] The steady-state model of the thermal network is determined based on the specific heat capacity of water, the upper and lower correlation matrices of the thermal network nodes and pipeline branches, the transposed matrix of the lower correlation matrix, the pipeline flow diagonal matrix, the temperature attenuation coefficient diagonal matrix, the thermal power matrix flowing into the node at time t, the node temperature matrix at time t, the pipeline end temperature matrix and the pipeline external environment temperature matrix, the flow rate of the i-th pipeline, the thermal conductivity per unit length, and the length, including:
[0109]
[0110]
[0111]
[0112] M=diag(m1,m2…m i )
[0113]
[0114] Where C p is the specific heat capacity of water; A up With A down are the upper and lower correlation matrices of the heating network nodes and pipeline branches respectively; is the transposed matrix of the lower correlation matrix; M is the pipeline flow diagonal matrix; E is the temperature attenuation coefficient diagonal matrix; is the thermal power matrix flowing into the node at time t; and are the node temperature matrix, the pipe end temperature matrix and the pipe external environment temperature matrix at time t; m i ,λ i and L i are the flow rate, thermal conductivity per unit length and length of the i-th pipeline respectively.
[0115] The operation constraints of the GB are determined based on the thermal power output of the GB at node i at time t, the conversion efficiency of the GB, and the lower and upper limits of the thermal power output, including:
[0116]
[0117]
[0118] Where, is the thermal power output of GB at node i at time t; η GB is the conversion efficiency of GB; and They are the lower and upper limits of thermal power output respectively.
[0119] The CHP operation constraints are determined based on the calorific value of natural gas, the conversion efficiency of CHP active output, the lower limit and upper limit of CHP active output, the ramp rate of CHP, the apparent power of CHP, and the power factor of CHP, including:
[0120]
[0121]
[0122]
[0123]
[0124]
[0125]
[0126] Where H NG is the calorific value of natural gas; η CHP The conversion efficiency of CHP active power; and are the lower and upper limits of CHP active output respectively; R CHP is the ramp rate of CHP; S CHP is the apparent power of CHP; is the power factor of CHP. is the thermal power output of CHP at node i at time t; b CHP is the thermoelectric ratio of CHP.
[0127] The EB operation constraints are determined based on the thermal power output of the EB at node i at time t, the conversion efficiency of the EB, the lower limit and upper limit of the EB thermal power output, and the ramp rate of the EB, including:
[0128]
[0129]
[0130]
[0131] Where, is the thermal power output of EB at node i at time t; η EB is the transformation efficiency of EB; and are the lower and upper limits of EB thermal power output respectively; R EB is the climbing rate of EB.
[0132] The CMDP modeling of the heterogeneous energy system model for trusted intelligent scheduling is carried out, including:
[0133] Define the agent observation state s t :
[0134]
[0135] Where, is the PV reactive output power of node i at time t-1, are the active output power of CHP of node i at time t-1, is the thermal output power of EB at node i at time t-1, is the thermal output power of GB at node i at time t-1.
[0136] Define the agent action space a t :
[0137]
[0138]
[0139]
[0140]
[0141]
[0142] Where: The action values of CHP and EB output by credible reinforcement learning.
[0143] Determine the objective function F:
[0144]
[0145] Where: F is the system operation cost; T is the number of time periods in the operation cycle; Ω NG ,Ω grid and Ω ES These are the node sets for gas purchase, electricity purchase and electric energy storage respectively; is the gas purchase price at time t; is the electricity purchase price at time t; α ES is the depreciation cost coefficient of electric energy storage; and are the total amount of natural gas consumed by GB and CHP at node i at time t; is the total amount of electricity purchased by node i from the upper power grid at time t; and are the ES charging and discharging powers of node i at time t; Δt is the time interval between adjacent time periods.
[0146] Determine the reward function R t (s t ,a t ):
[0147] R t (s t ,a t )=-ξ0F t
[0148] Where: ξ0 is the scaling factor of the immediate reward.
[0149] Define the unit output safety constraint function C1 and voltage offset safety constraint function C2:
[0150]
[0151]
[0152] Where: is the active power output of the unit at node i at time t; and are the maximum and minimum output values of the unit at any moment respectively; N is the number of nodes in the power grid; is the actual voltage value of grid node i at time t; is the upper limit of the reference voltage of grid node i at any moment; is the lower limit of the reference voltage of grid node i at any moment.
[0153] Define the security cost constraint C i (π):
[0154]
[0155]
[0156] Where: d i is the given constraint tolerance.
[0157] Define the HES scheduling model based on CMDP:
[0158]
[0159]
[0160] The use of the S-DRL model to solve the CMDP model includes:
[0161] The goal of the S-DRL model is to meet the long-term cost-benefit C i (π)≤d i In this case, we maximize the reward Q(π), Q(π) is:
[0162]
[0163]
[0164] The process can be restructured as:
[0165]
[0166]
[0167] Among them, β i is the Lagrange multiplier.
[0168] For the Lagrange multiplier β i To update:
[0169]
[0170] Among them, η k is the learning rate of the Lagrange multiplier;
[0171] Parameters of the policy network Update formula:
[0172]
[0173] Parameters θ of the target policy network π′ Update formula:
[0174] θ π′ ←τθ π +(1-τ)θ π′
[0175] Critic network loss function L(φ Q )for:
[0176]
[0177] Where, y is the action value function output by Critic network 1; t is the target Q value at time t, which is the smaller value output by the two target value networks.
[0178] Parameters of the value network Update formula:
[0179]
[0180] Where, α φ is the value network learning rate.
[0181] Safety constraint loss function output by the Safety network for:
[0182]
[0183] Where, is the safety constraint function corresponding to different constraint conditions.
[0184] Update the security constraint parameters:
[0185]
[0186] Target value network parameter soft update formula:
[0187]
[0188]
[0189] The AP-CSI S-DRL model interpretation method is used to interpret the solution results, including:
[0190] In this step, the scheduling agent in the S-DRL model generates a constrained behavior trajectory τ by interacting with the environment at each moment. Each constraint trajectory uses a sequence of "state-action" (τ i =(o0,a0),(o1,a1),...,(o k-1 ,a k-1 )) represents the learning behavior of the scheduling agent during each day-ahead scheduling cycle. To understand the learning behavior of the scheduling agent during each day-ahead scheduling cycle, it is necessary to embed the physical knowledge of the scheduling process and transform these disordered sequences into trajectories suitable for the scheduling cycle. Therefore, a groupby method is used to group the data according to the length of the day-ahead schedule. For each group, a custom day-ahead schedule is filtered for 24 hours to remove erroneous or missing data. This is done to extract trajectory data with a specific structure from the raw data for subsequent analysis or model training.
[0191] After the trajectory distance calculation is completed, combined with the clustering method, similar trajectories can be clustered to obtain training scenarios, thereby gaining a deeper understanding of the behavioral preferences of the scheduling agent under different environmental conditions. The classic Kmeans clustering is adopted. Kmeans clustering is a widely used unsupervised learning algorithm, which is mainly used to divide data points into a predefined number of clusters or classes. Its purpose is to find inherent groups in the data. The working steps of Kmeans clustering are as follows: First, k data points are randomly selected as the initial cluster centers. The Calinski-Harabasz indicator and the Davies-Bouldin indicator are used to comprehensively evaluate the number of clusters, and the solution formulas are:
[0192]
[0193] Where T is the total dispersion between all points, B k It is the class dispersion. The higher the value of this indicator, the better the clustering effect is. It means that the variance within the class is small and the variance between classes is large.
[0194]
[0195] Among them, σ i is the distance from all points in the i-th cluster to the cluster center c i The average distance, d(c i ,c j ) are different cluster centers c i and c jA lower Davies-Bouldin index value usually indicates a better clustering effect, because it indicates that the compactness within the cluster is high and the separation between clusters is also high.
[0196] After determining the number of clusters, for each data point, calculate its distance to each cluster center and assign it to the closest cluster center. Distances are typically calculated using the distances obtained using dynamic time warping in the previous section. Simultaneously, update each cluster center to the average position of all points within the cluster. Repeat this step until the cluster centers no longer change significantly or the predefined number of iterations is reached. The algorithm outputs the cluster assignments and cluster centers.
[0197] To further explain the characteristics of each scheduling scenario, SHAP is used to analyze the characteristics of the clustering results to more clearly explain the key characteristics of each scheduling scenario to humans. SHAP is a method for explaining machine learning model predictions based on the Shapley value in game theory. This method decomposes the model's output and divides it into the contribution of each input feature, thereby providing a detailed explanation of the model. For a given prediction, such as (s t ,a t ), the SHAP value measures the average contribution of each state feature to the agent's action. For model f and input x, the SHAP value is defined as:
[0198]
[0199] Among them, φ i is the SHAP value of state feature i, f x (S) is the contribution of a subset of state features S to the agent's actions, and n is the total number of state features. By calculating the SHAP value of each state feature, we can intuitively demonstrate the impact of each state feature on the agent's actions, helping us understand the decision-making process of complex models.
[0200] Example 3
[0201] This embodiment 3 provides a non-transitory computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the trusted interpretation method for intelligent scheduling decision-making of a heterogeneous energy system as described above is implemented. The method includes:
[0202] Construct a heterogeneous energy system model for trusted intelligent dispatch, including: determining power balance constraints, determining HES operation constraints, determining PV equipment reactive output constraints, determining the steady-state model of the thermal network, determining GB operation constraints, determining CHP operation constraints, and determining EB operation constraints;
[0203] Conducting CMDP modeling on the heterogeneous energy system model for trusted intelligent scheduling;
[0204] The S-DRL model is used to solve the CMDP model;
[0205] The S-DRL model interpretation method of AP-CSI is used to interpret the solution results.
[0206] Example 4
[0207] This embodiment 4 provides a computer device, including a memory and a processor, wherein the processor and the memory communicate with each other, the memory stores program instructions executable by the processor, and the processor invokes the program instructions to execute the above-mentioned trusted interpretation method for intelligent scheduling decision-making of a heterogeneous energy system, the method comprising:
[0208] Construct a heterogeneous energy system model for trusted intelligent dispatch, including: determining power balance constraints, determining HES operation constraints, determining PV equipment reactive output constraints, determining the steady-state model of the thermal network, determining GB operation constraints, determining CHP operation constraints, and determining EB operation constraints;
[0209] Conducting CMDP modeling on the heterogeneous energy system model for trusted intelligent scheduling;
[0210] The S-DRL model is used to solve the CMDP model;
[0211] The S-DRL model interpretation method of AP-CSI is used to interpret the solution results.
[0212] Example 5
[0213] This embodiment 5 provides an electronic device, including: a processor, a memory, and a computer program; wherein the processor is connected to the memory, and the computer program is stored in the memory. When the electronic device is running, the processor executes the computer program stored in the memory to cause the electronic device to execute instructions for implementing the above-mentioned trusted interpretation method for intelligent scheduling decision-making of heterogeneous energy systems. The method includes:
[0214] Construct a heterogeneous energy system model for trusted intelligent dispatch, including: determining power balance constraints, determining HES operation constraints, determining PV equipment reactive output constraints, determining the steady-state model of the thermal network, determining GB operation constraints, determining CHP operation constraints, and determining EB operation constraints;
[0215] Conducting CMDP modeling on the heterogeneous energy system model for trusted intelligent scheduling;
[0216] The S-DRL model is used to solve the CMDP model;
[0217] The S-DRL model interpretation method of AP-CSI is used to interpret the solution results.
[0218] In summary, the trusted interpretation method for intelligent scheduling decision-making of heterogeneous energy systems described in the embodiment of the present invention includes: constructing a heterogeneous energy system model for trusted intelligent scheduling; performing CMDP modeling on the heterogeneous energy system model for trusted intelligent scheduling; using the Safe Reinforcement Learning (S-DRL) model to solve the CMDP model; proposing an agent path cluster game interpretation method (Agent Path Cluster-SHAPInsight, AP-CSI) and interpreting the decision path of the S-DRL model. While utilizing the intelligent scheduling S-DRL model to ensure the safety of system operation, the present invention provides humans with a clearer and more intuitive explanation of the decision-making process through the internal mechanism understood by the AP-CSI interpretation method, thereby enhancing humans' trust in the intelligent scheduling method.
[0219] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0220] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0221] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0222] These computer program instructions can also be loaded onto a computer or other programmable data processing device, and a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide the functions for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0223] Although the above describes the specific embodiments of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solutions disclosed in the present invention without the need for creative work should be included in the scope of protection of the present invention.
Claims
1. A credible interpretation method for intelligent scheduling decision-making in heterogeneous energy systems, characterized by: include: Construct a heterogeneous energy system model for trusted intelligent dispatch, including: determining power balance constraints, determining HES operation constraints, determining PV equipment reactive output constraints, determining the steady-state model of the thermal network, determining GB operation constraints, determining CHP operation constraints, and determining EB operation constraints; Conducting CMDP modeling on the heterogeneous energy system model for trusted intelligent scheduling; The S-DRL model is used to solve the CMDP model; The S-DRL model interpretation method of AP-CSI is used to interpret the solution results.
2. The credible interpretation method for intelligent scheduling decision-making of heterogeneous energy systems according to claim 1 is characterized in that: The CMDP modeling of the heterogeneous energy system model for trusted intelligent scheduling is carried out, including: defining the observation state and action space of the intelligent agent, determining the objective function as the minimum system operation cost and safety cost constraint, determining the reward function, defining the unit output safety constraint function and the voltage offset safety constraint function, and determining the HES scheduling model based on CMDP.
3. The credible interpretation method for intelligent scheduling decision-making of heterogeneous energy systems according to claim 1 is characterized in that: The S-DRL model is used to solve the CMDP model, including: The goal of the S-DRL model is to meet the long-term cost-benefit C i (π)≤d i In this case, we maximize the reward Q(π), Q(π) is: The process can be restructured as: Among them, β i is the Lagrange multiplier; For the Lagrange multiplier β i To update: Among them, η k is the learning rate of the Lagrange multiplier; Parameters of the policy network Update formula: Parameters θ of the target policy network π′ Update formula: i π′ ←tth π +(1-τ)θ π′ Critic network loss function L(φ Q )for: Where, y is the action value function output by Critic network 1; t is the target Q value at time t, which is the smaller value output by the two target value networks; Parameters of the value network Update formula: Where, α φ is the value network learning rate; Safety constraint loss function output by the Safety network for: Where, is the safety constraint function corresponding to different constraint conditions; Update the security constraint parameters: Target value network parameter soft update formula:
4. The credible interpretation method for intelligent scheduling decision-making of heterogeneous energy systems according to claim 1 is characterized in that: The results of the solution are explained using AP-CSI's S-DRL model interpretation method. The method includes: in the S-DRL model, the scheduling agent generates a constrained behavior trajectory τ through interaction with the environment at each moment; after the trajectory distance is calculated, the training scenario is obtained by combining the clustering method; after the number of clusters is determined, for each data point, its distance to each cluster center is calculated and assigned to the nearest cluster center.
5. The credible interpretation method for intelligent scheduling decision-making of heterogeneous energy systems according to claim 4 is characterized in that: SHAP is used to analyze the characteristics of the clustering results. For a given prediction, the SHAP value measures the average contribution of each state feature to the agent's action. For the model f and input x, the SHAP value is defined as: Among them, φ i is the SHAP value of state feature i, f x (S) is the contribution of a subset of state features S to the agent's action, and n is the total number of state features.
6. The credible interpretation method for intelligent scheduling decision-making of heterogeneous energy systems according to claim 4 is characterized in that: Kmeans clustering is used: first, k data points are randomly selected as the initial cluster centers, and the Calinski-Harabasz index and Davies-Bouldin index are used to comprehensively evaluate the number of clusters. The solution formulas are: Where T is the total dispersion between all points, B k is the class discreteness; σ i is the distance from all points in the i-th cluster to the cluster center c i The average distance, d(c i ,c j ) are different cluster centers c i and c j The distance between them.
7. A trusted interpretation system for intelligent scheduling decisions in heterogeneous energy systems, characterized by: include: Model building device, used to build a heterogeneous energy system model for trusted intelligent scheduling, including: determining power balance constraints, HES operation constraints, photovoltaic equipment reactive output constraints, thermal network steady-state model, GB operation constraints, CHP operation constraints, and EB operation constraints; A model conversion device, used for performing CMDP modeling on the heterogeneous energy system model for trusted intelligent scheduling; A model solving device, used for solving the CMDP model using the S-DRL model; The result interpretation device is used to interpret the solution results using the S-DRL model interpretation method of AP-CSI.
8. A non-transitory computer-readable storage medium, characterized in that The non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by the processor, the trusted interpretation method for intelligent scheduling decision-making of heterogeneous energy systems as described in any one of claims 1 to 6 is implemented.
9. A computer device, characterized in that: It includes a memory and a processor, the processor and the memory communicate with each other, the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the trusted interpretation method for intelligent scheduling decision-making of heterogeneous energy systems as described in any one of claims 1 to 6.
10. An electronic device, characterized in that: include: A processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to execute instructions for implementing the trusted interpretation method for intelligent scheduling decision-making of heterogeneous energy systems as described in any one of claims 1 to 6.