An Optimization Decision Method and Device for a Low-Carbon Transformation Potential Model of a Power System Based on Interpretability Reinforcement Learning
Explainable reinforcement learning enhances electric power system optimization by integrating lifecycle assessment and optimal transport mapping to generate interpretable rules, addressing the lack of reliable low-carbon strategies in existing systems.
Patent Information
- Application Number
- CN202510611787.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-05-13
AI Technical Summary
Existing power system optimization technologies are difficult to provide reliable low-carbon transformation strategy support, and the low-carbon optimization decision-making strategies are relatively low.
Using an interpretability reinforcement learning method, carbon emission accounting and status representation are performed by obtaining multi-source data for the entire life cycle of the power system, and using optimal transmission interpretation mapping and causal analysis to generate symbolic rule sets, and output optimization decision report.
It has realized carbon emission management throughout the life cycle of the power system, improved the scientificity and reliability of low-carbon optimization decisions, enhanced the transparency and credibility of decision-making results, and promoted the transparency and compliance of low-carbon transformation.
Smart Images

Figure CN120146632B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of power technology, and in particular to an optimized decision-making method, device, computer equipment, computer-readable storage medium and computer program product for a power system low-carbon transformation potential model based on interpretable reinforcement learning. Background Art
[0002] With the increasingly serious global climate change problem, countries around the world have introduced relevant policies and measures to control greenhouse gas emissions. As the core of the energy supply system, the carbon emission level of the power system is directly related to the realization of energy structure adjustment and emission reduction goals.
[0003] The power system optimization technologies in the related art mostly focus on the efficiency improvement and economic objectives of a single link, such as minimizing the generation cost or maximizing the power supply reliability, and it is difficult to provide reliable decision-making support for the long-term low-carbon transformation strategy.
[0004] Therefore, there is a problem of low reliability in the low-carbon optimization decision-making strategy of the power system in the related art. Summary of the Invention
[0005] Based on this, it is necessary to provide an optimized decision-making method, device, computer equipment, computer-readable storage medium and computer program product for a power system low-carbon transformation potential model based on interpretable reinforcement learning, which can improve the reliability of the low-carbon optimization decision-making strategy of the power system, aiming at the above technical problems.
[0006] In a first aspect, the present application provides an optimized decision-making method for a power system low-carbon transformation potential model based on interpretable reinforcement learning, which is applied to the power system low-carbon transformation potential model, and includes:
[0007] Obtain multi-source data of the power system during its entire life cycle; the multi-source data includes power system data corresponding to each link during the entire life cycle;
[0008] Through the life cycle assessment method, quantify the carbon emissions of the power system during its entire life cycle according to the multi-source data, and obtain carbon emission accounting data; the carbon emission accounting data includes the carbon emission characteristics of each link and the total carbon emissions of the power system during the entire life cycle;
[0009] According to the carbon emission accounting data, convert the multi-source data into a state representation set in the reinforcement learning decision-making problem; the state representation set is used to instruct the reinforcement learning agent to determine a low-carbon optimization decision-making strategy for the power system;
[0010] Map the high-dimensional state representation set to a low-dimensional explanatory feature space through optimal transport interpretation mapping to generate a symbolic rule set; the symbolic rule set is used to characterize the key causal logic and decision-making path of the low-carbon optimization decision-making strategy;
[0011] Output an optimization decision report for the power system according to the low-carbon optimization decision-making strategy and the symbolic rule set.
[0012] In one embodiment, the mapping the high-dimensional state representation set to a low-dimensional explanatory feature space through optimal transport interpretation mapping to generate a symbolic rule set includes:
[0013] Map the high-dimensional state representation set to a low-dimensional explanatory feature space to obtain a low-dimensional explanatory feature set;
[0014] Use a causal discovery algorithm to construct a causal graph based on the low-dimensional explanatory feature set; the causal graph consists of an explanatory feature node set and a causal edge set of a directed acyclic graph;
[0015] Extract the key factors and paths leading to high carbon emissions or low-efficiency decisions from the causal graph to generate the symbolic rule set.
[0016] In one embodiment, the extracting the key factors and paths leading to high carbon emissions or low-efficiency decisions from the causal graph to generate the symbolic rule set includes:
[0017] According to the causal graph, screen out high-contribution feature combinations in the low-dimensional explanatory feature set; the high-contribution feature combinations include multiple high-contribution features;
[0018] According to the preset threshold sets corresponding to each high-contribution feature, determine the judgment conditions corresponding to each strategy action in the preset strategy action set to obtain the symbolic rule set.
[0019] In one embodiment, the outputting an optimization decision report for the power system according to the low-carbon optimization decision-making strategy and the symbolic rule set includes:
[0020] Perform counterfactual analysis based on the low-dimensional explanatory feature set to generate counterfactual explanation suggestions; the counterfactual explanation suggestions are used to indicate the impact of changes in the power system data on the low-carbon optimization decision-making strategy;
[0021] Generate the optimization decision report according to the low-carbon optimization decision-making strategy, the symbolic rule set and the counterfactual explanation suggestions, and output the optimization decision report.
[0022] In one embodiment, the method further includes:
[0023] Obtain the scoring results input by the user account for the symbol rule set;
[0024] According to the scoring results, determine the preference weight of the user account for the symbol rule set and the deviation between the symbol rule set and the user's expectation;
[0025] Update the symbol rule set and the low-carbon optimization decision strategy according to the preference weight and the deviation.
[0026] In one embodiment, the power system data includes the usage amounts of at least one material or energy. By using the life cycle assessment method, the carbon emissions in the whole life cycle of the power system are quantitatively accounted for based on the multi-source data, and carbon emission accounting data is obtained, including:
[0027] Multiply the usage amount of at least one material or energy corresponding to each link by its corresponding carbon emission factor, and sum up the product results to obtain the carbon footprint of each link; the carbon footprint is used to characterize the carbon emission characteristics;
[0028] Sum up the carbon footprints of each link to obtain the total carbon emissions.
[0029] In a second aspect, the present application further provides an optimization decision device for a power system low-carbon transformation potential model based on interpretable reinforcement learning, which is applied to the power system low-carbon transformation potential model and includes:
[0030] A data acquisition module, configured to acquire multi-source data in the whole life cycle of the power system; the multi-source data includes power system data corresponding to each link in the whole life cycle;
[0031] An accounting module, configured to quantitatively account for the carbon emissions in the whole life cycle of the power system based on the multi-source data by using the life cycle assessment method, and obtain carbon emission accounting data; the carbon emission accounting data includes the carbon emission characteristics of each link and the total carbon emissions of the power system in the whole life cycle;
[0032] A conversion module, configured to convert the multi-source data into a set of state representations in the reinforcement learning decision problem according to the carbon emission accounting data; the set of state representations is used to instruct the reinforcement learning agent to determine a low-carbon optimization decision strategy for the power system;
[0033] A mapping module, configured to map the high-dimensional set of state representations to a low-dimensional interpretable feature space through an optimal transport interpretation mapping to generate a symbol rule set; the symbol rule set is used to characterize the key causal logic and decision path of the low-carbon optimization decision strategy;
[0034] An output module, configured to output an optimization decision report for the power system according to the low-carbon optimization decision strategy and the symbolic rule set.
[0035] In a third aspect, the present application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the steps of the above-mentioned method are implemented.
[0036] In a fourth aspect, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the steps of the above-mentioned method are implemented.
[0037] In a fifth aspect, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by the processor, the steps of the above-mentioned method are implemented.
[0038] The above method, device, computer device, computer-readable storage medium, and computer program product for the low-carbon transformation potential model of the power system based on interpretable reinforcement learning are applied to the low-carbon transformation potential model of the power system. By acquiring multi-source data of the power system during its entire life cycle; the multi-source data includes power system data corresponding to each link during the entire life cycle; through the life cycle assessment method, the carbon emissions of the power system during its entire life cycle are quantitatively accounted for based on the multi-source data to obtain carbon emission accounting data; the carbon emission accounting data includes the carbon emission characteristics of each link and the total carbon emissions of the power system during its entire life cycle; according to the carbon emission accounting data, the multi-source data is transformed into a set of state representations in the reinforcement learning decision problem; the set of state representations is used to instruct the reinforcement learning agent to determine a low-carbon optimization decision strategy for the power system; through the optimal transport interpretation mapping, the high-dimensional set of state representations is mapped to a low-dimensional interpretive feature space to generate a symbolic rule set; the symbolic rule set is used to characterize the key causal logic and decision-making path of the low-carbon optimization decision strategy; according to the low-carbon optimization decision strategy and the symbolic rule set, an optimization decision report for the power system is output.
[0039] Thus, through the full-life-cycle carbon emission accounting method, this application obtains the comprehensive carbon emission accounting data of the power system, breaking through the limitations of only focusing on a single link or short-term goals, and expanding to cover the carbon emission management of the entire life cycle of the power system. This comprehensive analysis can comprehensively reveal the carbon emission sources of the power system at each stage of equipment manufacturing, installation, operation, maintenance, and retirement, thereby providing a more accurate and reliable carbon footprint assessment. This not only improves the scientific nature of optimization decisions but also avoids the risk of overestimating or underestimating the low-carbon transformation effect due to neglecting certain links. Subsequently, based on the carbon emission accounting data, using feature engineering and state representation methods, multi-source data is mapped into high-quality state information that can be processed by reinforcement learning to indicate the low-carbon optimization decision-making strategy for the power system by the reinforcement learning agent. Based on this, interpretable reinforcement learning is introduced, and the optimal transport interpretation mapping and causal analysis method are used to map the high-dimensional state to a low-dimensional interpretation space and generate a symbolic rule set for characterizing the key causal logic and decision-making path of the low-carbon optimization decision-making strategy, realizing a highly transparent and interpretable decision-making process, which can clearly present the key influencing factors and causal relationships behind the decision, enhancing the credibility and reviewability of the decision result. This helps users better understand and accept the optimization strategy, promoting the transparency and compliance of the decision-making process. At the same time, according to the low-carbon optimization decision-making strategy and the symbolic rule set, an optimization decision report for the power system is output, providing effective information reference and review basis for decision-makers and operation and maintenance personnel.
[0040] This solution provides a full-life-cycle and interpretable intelligent decision-making tool for the low-carbon transformation of the power system, thus effectively solving the problems of failure to globally optimize and lack of interpretability in existing research, and effectively improving the reliability of the low-carbon optimization decision-making strategy of the power system. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of this application or related technologies. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0042] Figure 1 It is a schematic flowchart of an optimization decision-making method for a low-carbon transformation potential model of a power system based on interpretable reinforcement learning in one embodiment;
[0043] Figure 2 It is a schematic flowchart of another optimization decision-making method for a low-carbon transformation potential model of a power system based on interpretable reinforcement learning in one embodiment;
[0044] Figure 3Schematic flowchart of an optimization decision-making method for a low-carbon transformation potential model of a power system based on interpretable reinforcement learning in another embodiment;
[0045] Figure 4 Structural block diagram of a device for optimizing a decision-making of a low-carbon transformation potential model of a power system based on interpretable reinforcement learning in an embodiment;
[0046] Figure 5 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0047] In order to make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0048] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such used data may be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0049] In one embodiment, as Figure 1 shown, an optimization decision-making method for a low-carbon transformation potential model of a power system based on interpretable reinforcement learning is provided, which is applied to a low-carbon transformation potential model of a power system. The low-carbon transformation potential model of a power system can be deployed in a computer device. It can be understood that the computer device can be a terminal, a server, or a system including a terminal and a server. In this embodiment, the method includes the following steps:
[0050] Step S110, obtaining multi-source data of the power system during its entire life cycle.
[0051] Among them, the multi-source data includes power system data corresponding to each link during the entire life cycle.
[0052] Among them, each link during the entire life cycle may include links such as equipment manufacturing, installation, operation, maintenance, and retirement.
[0053] Among them, the power system data corresponding to each link may include, but is not limited to, the following content:
[0054] Material and energy consumption data (material, energy consumption and corresponding carbon emission records) during the manufacturing and installation of power generation equipment.
[0055] Input-output data for the construction and maintenance of transmission lines (such as line load, loss, and maintenance frequency).
[0056] Operating status of the distribution network and related loss data (such as node voltage, load level, and energy storage configuration).
[0057] Power consumption characteristics and demand response data on the user side (such as power demand curves, price elasticity, and participation in demand response).
[0058] Data on the retirement and recycling of power equipment (waste treatment methods, recovery rates, and related emissions).
[0059] In a specific implementation, based on building a power system low-carbon transformation potential model based on interpretable reinforcement learning, first, comprehensive collection and preprocessing of power system data for each link in the entire life cycle are carried out. Specifically, it is necessary to obtain material and energy consumption data during the manufacturing and installation of power generation equipment, input-output data for the construction and maintenance of transmission lines, operating status of the distribution network and related loss data, power consumption characteristics and demand response data on the user side, and data on the retirement and recycling of power equipment. At the same time, to consider external influencing factors (such as weather changes and economic indicators), the multi-source data should also include relevant meteorological, macroeconomic, and industrial information.
[0060] For the convenience of those skilled in the art to understand, Figure 2 A flowchart of an optimization decision method for a power system low-carbon transformation potential model based on interpretable reinforcement learning is provided. As Figure 2 shown, it is necessary to perform data cleaning (processing missing values, outliers, and inconsistent data points) and standardization processing (normalizing features with different dimensions) on the multi-source data, and perform denoising, stationarity, and feature extraction on time series features (extracting short-term dynamic features through a sliding window) to obtain preprocessed multi-source data.
[0061] Specifically, in the data preprocessing link, it is necessary to clean and screen the above-mentioned collected multi-source data, including processing missing values, outliers, and inconsistent data points. Subsequently, standardization and normalization processing are performed on features with different dimensions, scales, and units to ensure the stability and effectiveness of the data. In addition, feature extraction and noise reduction operations are performed on time series data (such as generating time series features using a sliding window, and performing stationarity and detrending processing) to lay a high-quality data foundation for subsequent feature engineering and modeling, and obtain preprocessed multi-source data for subsequent processing.
[0062] Step S120: Quantitatively calculate the carbon emissions throughout the life cycle of the power system based on multi-source data through the life cycle assessment method to obtain carbon emission accounting data.
[0063] Among them, the carbon emission accounting data includes the carbon emission characteristics of each link and the total carbon emissions of the power system throughout the life cycle.
[0064] In specific implementation, referring to Figure 2 , after obtaining the pre-processed multi-source data that has been cleaned and standardized, the life cycle assessment (LCA) method is used to quantitatively calculate the carbon emissions throughout the life cycle of the power system to obtain carbon emission accounting data.
[0065] Specifically, the power system data includes the usage amounts of at least one material or energy. Multiply the usage amounts of at least one material or energy corresponding to each link by their respective corresponding carbon emission factors, and sum up the product results to obtain the carbon footprint of each link; the carbon footprint is used to characterize the carbon emission characteristics; then sum up the carbon footprints of each link to obtain the total carbon emissions. Through this process, users can obtain quantitative indicators reflecting the environmental impact of the entire process operation of the power system. This result will be used by the subsequent optimization module to set environmental target constraints or optimization directions.
[0066] To accurately quantify the carbon emissions of each life cycle link, in this embodiment, the link carbon footprint is defined as:
[0067]
[0068] Among them, is the carbon footprint of the th link throughout the life cycle, represents the consumption amount of the material or energy in link , is the consumption amount, is the carbon emission factor corresponding to the material or energy . By multiplying and summing up the usage amounts of materials and energy in each link with their corresponding carbon emission factors, the carbon footprint of each link can be obtained. Summing up the carbon footprints of each link can obtain the total carbon emissions throughout the life cycle of the power system:
[0069]
[0070] Among them, is the total carbon emissions, and N is the total number of links throughout the life cycle. This indicator ( ) is a comprehensive measure of the environmental impact of the power system throughout the entire life cycle and is an important target reference quantity for subsequent optimization decisions.
[0071] In this way, through the life cycle assessment method, based on the multi-source data of the power system throughout its life cycle, the carbon emissions of the power system throughout its life cycle are quantitatively accounted for, making up for the deficiencies of existing power system optimization methods, breaking through the limitations of only focusing on a single link or short-term goals, and extending to the carbon emission management covering the entire life cycle of the power system. Such an all-round analysis can comprehensively reveal the carbon emission sources of the power system at each stage of equipment manufacturing, installation, operation, maintenance, and decommissioning, thus providing a more accurate and reliable carbon footprint assessment. This not only improves the scientific nature of optimization decisions but also avoids the risk of overestimating or underestimating the low-carbon transformation effect due to neglecting certain links.
[0072] Step S130, according to the carbon emission accounting data, convert the multi-source data into a set of state representations in the reinforcement learning decision problem.
[0073] Among them, the set of state representations is used to instruct the reinforcement learning agent to determine the low-carbon optimization decision strategy for the power system.
[0074] In specific implementation, after clarifying the carbon emission characteristics of each link and the total carbon emissions throughout the life cycle, it is necessary to reasonably model and represent the characteristics of the power system state to convert the multi-source data into the state representation in the reinforcement learning decision problem and obtain the set of state representations. This set of state representations is used to instruct the reinforcement learning agent to determine the low-carbon optimization decision strategy for the power system. The set of state representations includes the key operating parameters, equipment life cycle information, user-side response characteristics, and external environment impact factors of the power system in each link of production, transmission, distribution, use, and abandonment. To reduce data redundancy and improve modeling efficiency, feature selection, dimensionality reduction, and feature combination can be performed on the features (such as automatically constructing cross features and extracting key principal components). The goal is to obtain a compact but information-rich state description on which the reinforcement learning agent can make policy decisions.
[0075] Reference Figure 2 , in the feature screening process, define the state space of the power system as a set containing multi-dimensional features. These features include but are not limited to:
[0076] Generation link features: unit output, proportion of renewable energy, unit start-stop status, maintenance plan, etc.
[0077] Transmission link features: line load rate, transmission loss, power flow distribution, section constraint, etc.
[0078] Distribution link features: distribution network topology structure, node voltage and load level, energy storage configuration and scheduling strategy.
[0079] Demand-side management features: user electricity consumption curve, price elasticity, demand response participation rate, proportion of shiftable load.
[0080] Full - life - cycle characteristics of the equipment: equipment construction and retirement time, maintenance cycle and cost, component replacement plan.
[0081] External influencing factors: weather conditions (wind speed, light intensity, temperature), economic indicators (electricity price, GDP growth rate), changes in market rules, etc.
[0082] When performing feature engineering on the above - mentioned characteristics, this application makes the state representation more discriminative and compact through steps of feature selection, dimensionality reduction, reduction, and interaction feature construction. Considering the time - series correlation, the sliding window method and time - series decomposition technology can be used to extract and fuse short - term and long - term features. To highlight key factors in high - dimensional features, methods such as statistical correlation, principal component analysis (PCA), and independent component analysis (ICA) can be used to screen out the feature subsets most relevant to carbon emissions, economic benefits, and system security. After completing the above data processing, carbon emission accounting, and feature engineering, a high - quality state representation set for reinforcement learning modeling can be obtained. .
[0083] At this time, to optimize the decision - making for low - carbon transformation, this application applies reinforcement learning to the policy search process under this state representation set to determine the low - carbon optimization decision - making strategy for the power system. For example, for the state at each time step, the agent can adjust the power generation structure (increase the proportion of clean energy, reduce the use of high - carbon fuels), optimize the power transmission path and maintenance plan, adjust the distribution network configuration, formulate demand response strategies, and determine the equipment replacement time. The reinforcement learning agent evaluates and updates the policy according to the immediate reward (comprehensive carbon emissions, economic cost, system stability, resilience, and robustness against disturbances).
[0084] Specifically, after obtaining high - quality data, clarifying the full - life - cycle carbon emission target, and constructing an appropriate state representation set, the present invention turns to the decision - making optimization of the reinforcement learning (RL) agent. In this problem, the power system can be represented as an environment with time - series dynamics. , the reinforcement learning agent faces the state at each time step and selects a decision - making strategy (low - carbon optimization decision - making strategy) from the set of feasible actions . After the agent executes the action , the environment gives an immediate reward and transfers to a new state
[0085] State space It includes multi-dimensional characteristics in the whole life cycle of the power system, such as power generation links (unit output, proportion of renewable energy, etc.), transmission links (transmission path selection, line operation mode, etc.), distribution links (network topology, load dispatching, etc.), demand-side management (user load response, energy storage strategy, etc.), and comprehensive descriptions of the whole life cycle characteristics of equipment (construction, maintenance, retirement, etc.). Considering the carbon emission accounting method for the whole life cycle, it is defined as:
[0086]
[0087] Among them, is the carbon emission of the th link in the whole life cycle, is the consumption of materials or energy in link , is the material or energy 's carbon emission factor.
[0088] Even if the reward function comprehensively considers carbon emissions, economic costs, system stability (and additional expansion goals), let:
[0089]
[0090] Among them is the weighting coefficient, is the current carbon emission, is the economic cost, is the system stability index, represents the system resilience reward, represents the robustness reward of the strategy under random perturbations.
[0091] The goal of the agent is to maximize the cumulative expected return:
[0092]
[0093] Among them is the decision-making strategy, is the discount factor.
[0094] Step S140, through the optimal transmission interpretation mapping, map the high-dimensional state representation set to the low-dimensional interpretation feature space to generate a symbolic rule set.
[0095] Among them, the symbolic rule set is used to characterize the key causal logic and decision-making path of the low-carbon optimization decision-making strategy.
[0096] Reinforcement learning strategies in the related art are difficult to intuitively understand. In this application, through optimal transport mapping and causal analysis, a set of high-dimensional state representation sets are mapped to a lower-dimensional and easily understandable interpretive feature space. In this interpretive space, causal discovery algorithms and rule extraction techniques can be used to generate a set of symbolic IF-THEN rules (symbolic rule sets) to express the key causal logic and decision-making paths behind the strategies.
[0097] Specifically, in the process of solving the optimization of low-carbon transformation strategies for power systems, reinforcement learning agents search for optimal strategies in high-dimensional and complex state spaces, often resulting in decision-making logics that are difficult for human experts to intuitively understand. This "black box" characteristic not only reduces the trust and usability of strategy results but also brings difficulties to the review and compliance of strategies in practical applications. To improve the transparency and interpretability of strategy decisions, this application introduces an interpretability mechanism in strategy optimization, enabling key influencing factors, logical paths, and causal chains in the decision-making process to be clearly identified and displayed.
[0098] Reference Figure 2 , in order to obtain feature representations that are easy to understand and analyze from complex, high-dimensional, and multi-source power system state data, through an optimal transport (OT) interpretive mapping, the complex high-dimensional state distribution is mapped to a lower-dimensional and interpretable distribution .
[0099] Let represent the high-dimensional state feature set, represent the low-dimensional interpretive feature set. Define the mapping , with the goal:
[0100]
[0101] Among them, , is the transport cost function, typically taking the square of the Euclidean distance:
[0102]
[0103] To adapt to the multi-stage characteristics of the entire life cycle of the power system, the state features can be decomposed into several subsets , corresponding to different links or time periods in the life cycle, and the OT problem is solved separately for each subset:
[0104]
[0105] Then, through quadratic OT solution, multiple sub-mappings are combined into global interpretive features . This process can be recorded as:
[0106]
[0107] where is the composite mapping for the second - stage merging. When introducing domain - knowledge constraints, a regularization term can be added to the objective function, such as
[0108]
[0109] where penalizes the transmission mapping through prior knowledge (such as power - grid topology constraints and upper and lower bounds of unit output), is the regularization coefficient.
[0110] Through causality identification, the key factors and paths leading to high - carbon emissions or low - efficiency decisions are extracted. Then, continuous features are divided into meaningful intervals to provide clear decision - making conditions for choosing which strategic actions in different situations. The finally generated symbolic rule set enables users to clearly understand the reasons for choosing a certain strategic action under specific conditions.
[0111] Specifically, after obtaining the low - dimensional explanatory features perform causal analysis and rule extraction on the low - dimensional explanatory features. Causal analysis: Use a causal discovery algorithm (such as the PC algorithm) to construct a causal graph where is the set of explanatory - feature nodes, is the set of causal edges of the directed acyclic graph. In the process of extracting the key factors and paths leading to high - carbon emissions or low - efficiency decisions based on the causal graph to generate a symbolic rule set, if there exists causal relationships, then give priority to considering the influence path on in the decision - making.
[0112] The causal graph helps to screen high - contribution feature combinations during rule generation and ensures that the decision - making logic is based on robust causal links rather than accidental correlations. Symbolic rule - set generation: To obtain directly understandable decision - making logic, divide the continuous - feature space into decision regions. For example, select several high - contribution features and set a threshold set for each feature to generate IF - THEN rules:
[0113]
[0114] The symbolic rule set is represented as:
[0115]
[0116] To reduce redundancy and complexity, the rule complexity can be set as (such as the number of features, logical length) and the rule accuracy (the matching degree of the rule to the action prediction under the current policy). Introduce the optimization objective:
[0117]
[0118] By minimizing this loss, a set of concise, highly accurate, and causally robust symbolic rule sets is obtained.
[0119] In this way, by introducing interpretable reinforcement learning, it not only has efficient intelligent optimization capabilities but also realizes a highly transparent and interpretable decision-making process. Using the optimal transport mapping and causal analysis methods, the key influencing factors and causal relationships behind the decision can be clearly presented, enhancing the credibility and scrutability of the decision results. This helps managers and regulatory agencies better understand and accept the optimization strategy, promoting the transparency and compliance of the decision-making process.
[0120] Step S150, according to the low-carbon optimization decision-making strategy and the symbolic rule set, output an optimization decision-making report for the power system.
[0121] In specific implementation, an optimization decision-making report for the power system can be output according to the low-carbon optimization decision-making strategy and the symbolic rule set, allowing users to review and file it offline.
[0122] In the above method for optimizing the decision-making of the low-carbon transformation potential model of the power system based on interpretable reinforcement learning, it is applied to the low-carbon transformation potential model of the power system. By obtaining multi-source data of the power system throughout the life cycle; the multi-source data includes power system data corresponding to each link throughout the life cycle; through the life cycle assessment method, the carbon emissions of the power system throughout the life cycle are quantitatively accounted for based on the multi-source data to obtain carbon emission accounting data; the carbon emission accounting data includes the carbon emission characteristics of each link and the total carbon emissions of the power system throughout the life cycle; according to the carbon emission accounting data, the multi-source data is converted into a set of state representations in the reinforcement learning decision-making problem; the set of state representations is used to instruct the reinforcement learning agent to determine the low-carbon optimization decision-making strategy for the power system; through the optimal transport interpretation mapping, the high-dimensional set of state representations is mapped to a low-dimensional interpretive feature space to generate a symbolic rule set; the symbolic rule set is used to characterize the key causal logic and decision-making path of the low-carbon optimization decision-making strategy; according to the low-carbon optimization decision-making strategy and the symbolic rule set, an optimization decision-making report for the power system is output.
[0123] Thus, through the full-life-cycle carbon emission accounting method, this application obtains the comprehensive carbon emission accounting data of the power system, breaking through the limitations of only focusing on a single link or short-term goals and expanding to cover the carbon emission management throughout the life cycle of the power system. This all-round analysis can comprehensively reveal the carbon emission sources at each stage of equipment manufacturing, installation, operation, maintenance, and decommissioning of the power system, thus providing a more accurate and reliable carbon footprint assessment. This not only improves the scientific nature of optimization decisions but also avoids the risk of overestimating or underestimating the low-carbon transformation effect due to neglecting some links. Subsequently, based on the carbon emission accounting data, using feature engineering and state representation methods, multi-source data is mapped into high-quality state information that can be processed by reinforcement learning to indicate the low-carbon optimization decision-making strategy for the power system by the reinforcement learning agent. Based on this, interpretable reinforcement learning is introduced, and the optimal transport interpretation mapping and causal analysis method are used to map the high-dimensional state to a low-dimensional interpretation space and generate a symbolic rule set for characterizing the key causal logic and decision-making path of the low-carbon optimization decision-making strategy, realizing a highly transparent and interpretable decision-making process, being able to clearly present the key influencing factors and causal relationships behind the decision, enhancing the credibility and scrutability of the decision result, which helps users better understand and accept the optimization strategy and promotes the transparency and compliance of the decision-making process. At the same time, according to the low-carbon optimization decision-making strategy and the symbolic rule set, an optimization decision report for the power system is output, providing an effective information reference and review basis for the decision-making layer and operation and maintenance personnel.
[0124] This solution provides a full-life-cycle and interpretable intelligent decision-making tool for the low-carbon transformation of the power system, thus effectively solving the problems of failure to globally optimize and insufficient interpretability in existing research and effectively improving the reliability of the low-carbon optimization decision-making strategy of the power system.
[0125] In some embodiments, according to the low-carbon optimization decision-making strategy and the symbolic rule set, an optimization decision report for the power system is output, including: performing counterfactual analysis based on the low-dimensional interpretation feature set to generate counterfactual interpretation suggestions; the counterfactual interpretation suggestions are used to indicate the impact of changes in the power system data on the low-carbon optimization decision-making strategy; according to the low-carbon optimization decision-making strategy, the symbolic rule set, and the counterfactual interpretation suggestions, an optimization decision report is generated and the optimization decision report is output.
[0126] In specific implementation, referring to Figure 2 , in the process of outputting an optimization decision report for the power system according to the low-carbon optimization decision-making strategy and the symbolic rule set, counterfactual analysis can be performed based on the low-dimensional interpretation feature set to generate counterfactual interpretation suggestions; the counterfactual interpretation suggestions are used to indicate the impact of changes in the power system data on the low-carbon optimization decision-making strategy; then, according to the low-carbon optimization decision-making strategy, the symbolic rule set, and the counterfactual interpretation suggestions, an optimization decision report is generated and the optimization decision report is output.
[0127] Among them, counterfactual analysis, that is, answering "if a certain condition is changed, can the decision result be optimized?", by providing counterfactual suggestions, the operation and maintenance personnel can obtain intervention strategies (such as slightly increasing the energy storage capacity or optimizing the maintenance frequency of specific lines) to improve the system performance and emission reduction effect.
[0128] Specifically, during the counterfactual analysis process, let the current state be and the corresponding explanatory features , and the decision is . The counterfactual explanation problem is to find a point close to , so that the strategy produces a different decision, such as a lower carbon emission action:
[0129]
[0130] Where is a similarity metric (such as L1 or L2 distance). Counterfactual explanations provide insights into how the decision changes if a certain feature changes slightly, assisting the operation and maintenance personnel to adjust the corresponding features in practice to improve the decision-making performance.
[0131] Furthermore, a multi-level explanation mechanism can also be adopted. The multi-level explanation mechanism enables users to view the overall strategy logic and main causal chains at the global level, and also to examine the specific reasons and key features for a decision at a specific moment in detail at the local level.
[0132] In practical applications, the multi-level explanation structure includes:
[0133] Global level: Output the causal link diagram and rule set of the entire strategy, showing the global decision-making logic framework.
[0134] Local level: For the state at a specific moment, give a readable subset of decision rules, causal paths, and counterfactual suggestions.
[0135] In this way, the multi-level explanation mechanism can enable users to have a clear understanding from the high-level overview (overall optimization strategy for the life cycle) to the micro-local level (decision-making motivation and improvement path at a specific time point).
[0136] In some embodiments, during the training process of the low-carbon transformation potential model of the power system, multi-objective optimization is achieved by adding an explanation loss term to the proximal policy optimization (PPO):
[0137]
[0138] Where: is the loss of the reinforcement learning itself, and the CLIP loss of PPO is used here:
[0139]
[0140] is the estimated value of the advantage function, is the truncation range.
[0141] Ensure the stability of the interpretation mapping:
[0142]
[0143] Optimize the symbol rule set:
[0144]
[0145] During the policy optimization process, the interpretation weight can be dynamically adjusted according to the current decision-making performance and interpretation quality. If the rules are too complex or the accuracy is not ideal, the optimization of the interpretation rules will tend to be strengthened; if the policy performance is poor, the focus will be on enhancing the performance of the reinforcement learning itself. This adaptive adjustment ensures that the final policy can achieve the low-carbon and efficient goals without becoming an incomprehensible black box. Specifically:
[0146] is the weight of the interpretation-related loss, and a dynamic adjustment strategy can be adopted, such as defining a metric function:
[0147]
[0148] where represents the performance related to interpretation (such as rule simplicity and accuracy), represents the performance related to reinforcement learning. If the interpretation performance ratio is lower than the expectation, increase , otherwise decrease . This can be achieved through an adaptive scheme, such as:
[0149]
[0150] where is the adjustment rate, is the expected interpretation-performance ratio.
[0151] In addition, the bi-level optimization technique can be used, taking interpretability as the upper-level constraint and RL performance as the lower-level problem, and solving the KKT conditions through Lagrange multipliers to ensure the optimal trade-off.
[0152] In this way, the model has the ability of multi-objective optimization and can achieve balanced optimization among multiple objectives such as carbon emissions, economic costs, and system stability. By comprehensively considering various indicators, it ensures that the low-carbon transformation strategy can achieve environmental goals without sacrificing the economic benefits and operational safety of the power system. This multi-dimensional optimization ability makes the low-carbon transformation more balanced and feasible, promoting the development of the power system towards a more efficient and environmentally friendly direction.
[0153] In some other embodiments, referring to Figure 2 , this application can update data and models online regularly. Specifically, as time goes by and external conditions change (such as sudden weather changes and market price fluctuations), data and models can be updated online regularly to make the strategy and interpretation structure adapt to the latest conditions. Users can view low-carbon optimization decision-making strategies, symbolic rule sets, causal paths, and counterfactual suggestions in the visualization interface, and provide feedback, rate the rules, or edit them. Update the interpretation rules and decision-making strategies according to user feedback, so as to achieve human-machine collaborative optimization and co-creation of interpretations.
[0154] In the actual operation of the power system, data distribution and conditions change over time. For this reason, this application supports online iterative updates and human-computer interaction: (1) Online iteration: After a certain period of time , use new data to re-estimate the state distribution , transmission mapping , and causal graph, and dynamically adjust the symbolic rules and interpretation weights to ensure the long-term adaptability of the strategy. (2) Human-computer interaction: Build a visualization interface to display low-carbon optimization decision-making strategies, symbolic rule sets, causal chains, and counterfactual suggestions. The user account rates the symbolic rule set, and the computer device can obtain the rating results input for the symbolic rule set. According to the rating results, determine the preference weight of the user account for the symbolic rule set and the deviation between the symbolic rule set and the user's expectations; update the symbolic rule set and the low-carbon optimization decision-making strategy according to the preference weight and the deviation.
[0155] The formula is as follows:
[0156]
[0157] Among them, is the feedback information, represents the preference weight of the user account feedback for the symbolic rule , is the deviation between the symbolic rule and the user's expectations. Incorporate into the overall goal to complete the closed-loop optimization of the "agent-human" co-constructed interpretation.
[0158] In this way, the model has a high degree of self - adaptability and real - time response ability. Through online iterative updates and human - machine interaction mechanisms, it can dynamically adapt to changes in external conditions (such as weather, market prices), and adjust and optimize strategies and explanation rules in real time. This flexibility and self - adaptability ensure that the power system always maintains an efficient and low - carbon operation state in a changing environment, while enhancing users' trust and support for system decisions, promoting the wide application of low - carbon technologies and the realization of sustainable development goals.
[0159] Furthermore, this application can evaluate the stability and explanation quality of decision - making strategies during the actual operation of the power system or through historical scenario playback. By comparing with the benchmark results of the non - explanatory model, it examines the performance of the current decision - making strategy when facing extreme load shocks and renewable energy fluctuations. At the same time, through the automatic report generation function, it regularly outputs a comprehensive report containing decision - making strategies, symbolic rules, causal links, and counterfactual explanation suggestions, providing effective information reference and review basis for decision - making levels and operation and maintenance personnel.
[0160] Specifically, referring to Figure 2 , the evaluation mechanism includes two parts: report generation and strategy playback:
[0161] (1) Strategy playback: To ensure the effectiveness and practicality of explanations, this application adds a verification link: replay the strategy decision - making process and explanation output for known historical extreme scenarios (such as large load shocks, extreme weather), compare the causal logic and rule applicability of the strategy in dealing with extreme situations, and set a confidence measure:
[0162]
[0163] If the confidence in the extreme scenario is high, it indicates that the explanation maintains consistency and robustness in real challenges.
[0164] (2) Automatic report generation. Build an automated reporting function that regularly outputs a comprehensive report containing decision - making strategies, symbolic rules, causal paths, counterfactual cases, and the processing results of user feedback, allowing users to review offline and file. Mathematical statistical indicators, such as rule distribution entropy (measuring rule diversity), can be added to the report:
[0165]
[0166] is the probability of a rule being triggered to meet the requirements. If the entropy is too large, it means that the rules that meet the requirements are too scattered, and redundancy can be reduced by further optimization
[0167] In summary, the low-carbon transformation potential model of the power system based on interpretable reinforcement learning not only realizes the comprehensive management and optimal decision-making of carbon emissions throughout the life cycle at the technical level, but also significantly promotes the low-carbon transformation process of the power system by improving decision-making transparency and enhancing system adaptability in practical applications, and has broad application prospects.
[0168] In another embodiment, as Figure 3 shown, an optimal decision-making method for the low-carbon transformation potential model of the power system based on interpretable reinforcement learning is provided, which is applied to the low-carbon transformation potential model of the power system and includes the following steps:
[0169] Step S302, obtain multi-source data of the power system throughout the life cycle.
[0170] Step S304, multiply the usage amount of at least one material or energy corresponding to each link by their respective corresponding carbon emission factors, and sum the product results to obtain the carbon footprint of each link.
[0171] Step S306, add up the carbon footprints of each link to obtain the total carbon emissions.
[0172] Step S308, according to the carbon emission accounting data, convert the multi-source data into a set of state representations in the reinforcement learning decision problem.
[0173] Step S310, map the high-dimensional set of state representations to a low-dimensional interpretable feature space to obtain a low-dimensional interpretable feature set.
[0174] Step S312, use the causal discovery algorithm to construct a causal graph based on the low-dimensional interpretable feature set.
[0175] Step S314, according to the causal graph, screen out high-contribution feature combinations in the low-dimensional interpretable feature set.
[0176] Step S316, according to the preset threshold set corresponding to each high-contribution feature, determine the judgment conditions corresponding to each policy action in the preset policy action set to obtain a symbolic rule set.
[0177] Step S318, perform counterfactual analysis based on the low-dimensional interpretable feature set to generate counterfactual explanation suggestions.
[0178] Step S320, generate an optimal decision-making report according to the low-carbon optimal decision-making strategy, the symbolic rule set and the counterfactual explanation suggestions, and output the optimal decision-making report.
[0179] It should be noted that the specific limitations of the above steps can refer to the specific limitations of an optimal decision-making method for the low-carbon transformation potential model of the power system based on interpretable reinforcement learning described above.
[0180] This solution provides a full - life - cycle, interpretable, and self - adaptive intelligent decision - making tool for the low - carbon transformation of the power system, thus effectively solving the problems of failure to globally optimize, insufficient interpretability, and lack of online collaborative regulation in existing research.
[0181] It should be understood that although the steps in the flowcharts involved in the above - mentioned embodiments are shown in sequence according to the arrows, these steps do not necessarily execute in the order indicated by the arrows. Unless there is a clear description in this article, there is no strict order limit for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above - mentioned embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily execute at the same moment, but can execute at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0182] Based on the same inventive concept, the embodiments of the present application also provide a device for optimizing and making decisions on the low - carbon transformation potential model of the power system based on interpretable reinforcement learning, which is used to implement the method for optimizing and making decisions on the low - carbon transformation potential model of the power system based on interpretable reinforcement learning mentioned above. The implementation solutions for solving problems provided by this device are similar to those recorded in the above - mentioned method. Therefore, the specific limitations in one or more embodiments of the device for optimizing and making decisions on the low - carbon transformation potential model of the power system based on interpretable reinforcement learning provided below can refer to the limitations on the method for optimizing and making decisions on the low - carbon transformation potential model of the power system based on interpretable reinforcement learning in the above text, and will not be elaborated here.
[0183] In an exemplary embodiment, as Figure 4 shown, a device for optimizing and making decisions on the low - carbon transformation potential model of the power system based on interpretable reinforcement learning is provided. It is applied to the low - carbon transformation potential model of the power system and includes: a data acquisition module 410, an accounting module 420, a transformation module 430, a mapping module 440, and an output module 450, where:
[0184] The data acquisition module 410 is used to acquire multi - source data of the power system during its full life cycle; the multi - source data includes power system data corresponding to each link during the full life cycle.
[0185] The accounting module 420 is used to quantitatively account for the carbon emissions of the power system during its full life cycle according to the multi - source data by using the life - cycle assessment method, and obtain carbon emission accounting data; the carbon emission accounting data includes the carbon emission characteristics of each link and the total carbon emissions of the power system during the full life cycle.
[0186] A conversion module 430, configured to convert the multi-source data into a set of state representations in a reinforcement learning decision-making problem according to the carbon emission accounting data; the set of state representations is used to instruct a reinforcement learning agent to determine a low-carbon optimization decision-making strategy for the power system.
[0187] A mapping module 440, configured to map the high-dimensional set of state representations to a low-dimensional explanatory feature space through optimal transport interpretation mapping to generate a symbolic rule set; the symbolic rule set is used to characterize the key causal logic and decision-making path of the low-carbon optimization decision-making strategy.
[0188] An output module 450, configured to output an optimization decision report for the power system according to the low-carbon optimization decision-making strategy and the symbolic rule set.
[0189] In one embodiment, the mapping module 440 is specifically configured to map the high-dimensional set of state representations to a low-dimensional explanatory feature space to obtain a low-dimensional explanatory feature set; use a causal discovery algorithm to construct a causal graph based on the low-dimensional explanatory feature set; the causal graph consists of an explanatory feature node set and a causal edge set of a directed acyclic graph; extract key factors and paths leading to high carbon emissions or low-efficiency decisions from the causal graph to generate the symbolic rule set.
[0190] In one embodiment, the mapping module 440 is specifically configured to screen out high-contribution feature combinations in the low-dimensional explanatory feature set according to the causal graph; the high-contribution feature combinations include multiple high-contribution features; determine the determination conditions corresponding to each policy action in a preset policy action set according to a preset threshold set corresponding to each high-contribution feature to obtain the symbolic rule set.
[0191] In one embodiment, the output module 450 is specifically configured to perform counterfactual analysis according to the low-dimensional explanatory feature set to generate counterfactual explanation suggestions; the counterfactual explanation suggestions are used to indicate the impact of changes in the power system data on the low-carbon optimization decision-making strategy; generate the optimization decision report according to the low-carbon optimization decision-making strategy, the symbolic rule set and the counterfactual explanation suggestions, and output the optimization decision report.
[0192] In one embodiment, the device further includes: an update module, configured to obtain a scoring result input by a user account for the symbolic rule set; determine a preference weight of the user account for the symbolic rule set and a deviation between the symbolic rule set and the user's expectation according to the scoring result; update the symbolic rule set and the low-carbon optimization decision-making strategy according to the preference weight and the deviation.
[0193] In one embodiment, the power system data includes the usage amount of at least one material or energy. The accounting module 420 is specifically configured to multiply the usage amount of at least one material or energy corresponding to each link by the respective corresponding carbon emission factor, and sum up the product results to obtain the carbon footprint of each link; the carbon footprint is used to characterize the carbon emission characteristics; and add up the carbon footprints of each link to obtain the total carbon emissions.
[0194] Each module in the above power system low-carbon transformation potential model optimization decision-making device based on interpretable reinforcement learning can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or be independent of it, or stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0195] In an exemplary embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 5 shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a method for optimizing the decision-making of the power system low-carbon transformation potential model based on interpretable reinforcement learning. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.
[0196] Those skilled in the art can understand, Figure 5The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0197] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0198] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0199] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0200] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0201] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include Read-Only Memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, Resistive Random Access Memory (ReRAM), Magnetoresistive Random Access Memory (MRAM), Ferroelectric Random Access Memory (FRAM), Phase Change Memory (PCM), graphene memory, etc. Volatile memory can include Random Access Memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, Artificial Intelligence (AI) processors, etc., and are not limited thereto.
[0202] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.
[0203] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. An optimization decision-making method for the low-carbon transformation potential model of a power system based on interpretable reinforcement learning, characterized in that Applied to the low-carbon transformation potential model of the power system, the method includes: Obtain multi-source data of the power system throughout its life cycle; the multi-source data includes power system data corresponding to each link in the life cycle; Through the life cycle assessment method, quantitatively calculate the carbon emissions of the power system throughout its life cycle based on the multi-source data to obtain carbon emission accounting data; the carbon emission accounting data includes the carbon emission characteristics of each link and the total carbon emissions of the power system throughout the life cycle; According to the carbon emission accounting data, transform the multi-source data into a set of state representations in the reinforcement learning decision-making problem; the set of state representations is used to instruct the reinforcement learning agent to determine the low-carbon optimization decision-making strategy for the power system; Through the optimal transport interpretation mapping, map the high-dimensional set of state representations to a low-dimensional interpretation feature space to generate a symbolic rule set; the symbolic rule set is used to characterize the key causal logic and decision-making path of the low-carbon optimization decision-making strategy; including: mapping the high-dimensional set of state representations to a low-dimensional interpretation feature space to obtain a low-dimensional interpretation feature set; using a causal discovery algorithm to construct a causal graph based on the low-dimensional interpretation feature set; the causal graph consists of an interpretation feature node set and a causal edge set of a directed acyclic graph; extract the key factors and paths leading to high carbon emissions or low-efficiency decisions from the causal graph to generate the symbolic rule set; the extracting the key factors and paths leading to high carbon emissions or low-efficiency decisions from the causal graph to generate the symbolic rule set includes: screening out high-contribution feature combinations from the low-dimensional interpretation feature set according to the causal graph; the high-contribution feature combinations include multiple high-contribution features; according to the preset threshold sets corresponding to each high-contribution feature, determine the judgment conditions corresponding to each policy action in the preset policy action set to obtain the symbolic rule set; According to the low-carbon optimization decision-making strategy and the symbolic rule set, output an optimization decision report for the power system, including: generating counterfactual explanation suggestions through counterfactual analysis based on the low-dimensional interpretation feature set; the counterfactual explanation suggestions are used to indicate the impact of changes in the power system data on the low-carbon optimization decision-making strategy; generate the optimization decision report according to the low-carbon optimization decision-making strategy, the symbolic rule set and the counterfactual explanation suggestions, and output the optimization decision report.
2. The method according to claim 1, wherein The method further includes: Obtain the scoring result input by the user account for the symbolic rule set; According to the scoring result, determine the preference weight of the user account for the symbolic rule set and the deviation between the symbolic rule set and the user's expectation; Update the symbolic rule set and the low-carbon optimization decision-making strategy according to the preference weight and the deviation.
3. The method according to claim 1, wherein The power system data includes the usage amount of at least one material or energy. The quantitatively calculating the carbon emissions of the power system throughout its life cycle based on the multi-source data through the life cycle assessment method to obtain carbon emission accounting data includes: Multiply the usage amount of at least one material or energy corresponding to each of the said links by their respective corresponding carbon emission factors, and sum up the product results to obtain the carbon footprint of each of the said links; the carbon footprint is used to characterize the carbon emission characteristics; Sum up the carbon footprints of each of the said links to obtain the total carbon emissions.
4. An optimization decision-making device for a low-carbon transformation potential model of a power system based on interpretable reinforcement learning, characterized in that, Applied to the low-carbon transformation potential model of the power system, the device includes: A data acquisition module, configured to acquire multi-source data of the power system during its entire life cycle; the multi-source data includes the power system data corresponding to each link during the entire life cycle; An accounting module, configured to quantitatively account for the carbon emissions of the power system during its entire life cycle based on the multi-source data through the life cycle assessment method, to obtain carbon emission accounting data; the carbon emission accounting data includes the carbon emission characteristics of each of the said links and the total carbon emissions of the power system during the entire life cycle; A transformation module, configured to transform the multi-source data into a set of state representations in the reinforcement learning decision-making problem according to the carbon emission accounting data; the set of state representations is used to instruct the reinforcement learning agent to determine the low-carbon optimization decision-making strategy for the power system; A mapping module, configured to map the high-dimensional set of state representations to a low-dimensional interpretive feature space through optimal transport interpretation mapping, to generate a symbolic rule set; the symbolic rule set is used to characterize the key causal logic and decision-making path of the low-carbon optimization decision-making strategy; The mapping module is specifically configured to map the high-dimensional set of state representations to a low-dimensional interpretive feature space to obtain a low-dimensional interpretive feature set; construct a causal graph based on the low-dimensional interpretive feature set using a causal discovery algorithm; the causal graph consists of a set of interpretive feature nodes and a set of causal edges of a directed acyclic graph; extract the key factors and paths leading to high carbon emissions or low-efficiency decisions from the causal graph to generate the symbolic rule set; The mapping module is specifically configured to screen out high-contribution feature combinations in the low-dimensional interpretive feature set according to the causal graph; the high-contribution feature combinations include multiple high-contribution features; determine the judgment conditions corresponding to each policy action in the preset policy action set according to the preset threshold set corresponding to each of the high-contribution features to obtain the symbolic rule set; An output module, configured to output an optimization decision report for the power system according to the low-carbon optimization decision-making strategy and the symbolic rule set; The output module is specifically configured to perform counterfactual analysis based on the low-dimensional interpretive feature set to generate counterfactual interpretation suggestions; the counterfactual interpretation suggestions are used to indicate the impact of the change in the power system data on the low-carbon optimization decision-making strategy; generate the optimization decision report according to the low-carbon optimization decision-making strategy, the symbolic rule set and the counterfactual interpretation suggestions, and output the optimization decision report.
5. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 3.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 3.
7. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Urban water delivery system water volume prediction method and system based on low carbon emission
CN117422165A