Multi-objective regulation method and device for power distribution network based on potential game and dynamic reward

CN122533112APending Publication Date: 2026-08-07STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO
Filing Date
2026-01-15
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

然而,现有方法多基于非合作博弈或简单共享奖励机制,各智能体以自身奖励最大化为目标,缺乏有效的全局协同引导,易产生策略冲突与振荡,导致算法收敛速度慢、稳定性差,且最终达成的纳什均衡点常与系统帕累托最优解相去甚远

Benefits of technology

1、收敛速度快:相较于传统MARL方法,训练轮次减少50%,10kV节点配电网达到纳什均衡的时间缩短至3.2小时,优化效率大幅提升。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122533112A_ABST
    Figure CN122533112A_ABST
Patent Text Reader

Abstract

The application discloses a power distribution network multi-objective regulation method based on potential game and dynamic reward, comprising the following steps: constructing a first intelligent agent including power grid architecture features and a second intelligent agent including power grid economic features, and designing a multi-objective potential function based on the two intelligent agents; acquiring a differential reward function by constructing a counterfactual baseline in an inner layer based on the first intelligent agent and the second intelligent agent, and updating a reward weight vector by constructing a regret matching algorithm in an outer layer; constructing a state space and an action space of a power distribution network; updating a loss function of a strategy network by adding an N-1 safety criterion opportunity constraint penalty term based on the loss function; updating network parameters by training intelligent agents of the power distribution network; and inputting real-time working data to output a Pareto optimal solution set and a corresponding optimization scheme through preprocessing, the intelligent agents of the power distribution network, safety checking and convergence judgment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power distribution network optimization operation technology, and particularly relates to a multi-objective control method and device for power distribution networks based on potential game theory and dynamic rewards. Background Technology

[0003] With the advancement of energy transition, distribution networks are developing towards high penetration rates of renewable energy integration, high levels of power electronics, and high levels of user interaction and participation. Against this backdrop, the optimized operation of distribution networks needs to simultaneously coordinate multiple conflicting objectives such as power supply reliability, operational economy, voltage quality, and investment costs. Traditional optimization methods face significant challenges. First, existing methods mostly employ static weighted linear scaling or hierarchical optimization strategies, transforming multi-objective problems into single-objective problems or sequential problems with pre-defined priorities. These methods cannot dynamically adapt to changes in system state, and the setting of weights or priorities relies heavily on expert experience, resulting in strong subjectivity. This often leads to optimization results getting trapped in local optima, making it difficult to achieve a globally optimal balance among the objectives at the Pareto front.

[0004] Second, distributed optimization using multi-agent reinforcement learning is a current research hotspot. However, existing methods are mostly based on non-cooperative game theory or simple shared reward mechanisms, where each agent aims to maximize its own reward, lacking effective global coordination guidance, which easily leads to policy conflicts and oscillations. This results in slow convergence speed, poor stability, and the final Nash equilibrium point often deviates significantly from the Pareto optimal solution of the system.

[0005] Third, most studies use simplified discrete or continuous action spaces, failing to fully consider the characteristics of discrete-continuous hybrid decision-making in actual distribution network scheduling and planning. State representations also often neglect key operational characteristics such as node voltage sensitivity and line load balance, making it difficult to directly apply the trained strategies to real-world engineering scenarios.

[0006] Fourth, existing optimization models typically treat safety constraints as hard constraints or ex-post verification terms, failing to deeply embed them into the agent's learning process. This may result in optimization schemes that perform well in terms of economy or performance indicators, but fail to meet the safety and reliability requirements of actual operation, or require extensive manual adjustments to be feasible, thus weakening the practical value of the method.

[0007] Fifth, the reward function is crucial for guiding the agent's learning. Existing methods often employ simple weighted sums with fixed weights, lacking a clear normalization process and dynamic adjustment mechanism. Because the dimensions and orders of magnitude of the various optimization objectives differ significantly, fixed weights can easily lead the agent to overemphasize one objective while neglecting others, resulting in an unstable learning process and unclear decision-making direction. Summary of the Invention

[0008] To address the aforementioned issues, this invention proposes a multi-objective control method and device for distribution networks based on potential game theory and dynamic rewards. By constructing a cooperative architecture through potential game theory and utilizing a two-layer dynamic reward shaping mechanism, it resolves the policy conflict and guidance ambiguity problems in traditional multi-agent reinforcement learning. This enables the algorithm to converge quickly and stably to Nash equilibrium. Through a designed potential function and dynamic weight mechanism, it automatically coordinates the inherent conflicts among multiple objectives such as reliability, economy, and voltage quality. This invention is the first to systematically introduce potential game theory into the field of multi-agent optimization for distribution networks, mathematically ensuring the unity of Nash equilibrium and the pursuit of Pareto optimality, avoiding the disconnect between equilibrium solutions and optimal solutions in non-cooperative game theory, and providing a model foundation for the optimality of the solution. A dedicated state-action space for distribution networks is defined, fully considering the characteristics of discrete-continuous mixed decision-making in actual scheduling, enabling seamless integration between the algorithm's output actions and actual engineering operation instructions, avoiding decision biases caused by universal models, and significantly improving practicality. By deeply embedding the N-1 safety criterion as an opportunity constraint penalty term into the training loss function of the policy network, rather than verifying it afterward, it is ensured that all output optimization schemes meet the strict safety operation requirements during the training process.

[0009] The first aspect of this invention provides a multi-objective control method for distribution networks based on potential game theory and dynamic rewards, comprising: Construct a first intelligent agent that includes the characteristics of the power grid architecture and a second intelligent agent that includes the economic characteristics of the power grid. Design a multi-objective potential function based on the above two intelligent agents. Based on the first intelligent agent and the second intelligent agent, the differential reward function is obtained by constructing a counterfactual baseline in the inner layer, and the reward weight vector is updated by constructing a regret matching algorithm in the outer layer; Construct the state space and action space of the power distribution network; The loss function based on the policy network is updated by adding an N-1 security criterion chance constraint penalty term; Network parameters are updated by training agents in the distribution network; The input real-time working data is preprocessed and processed by the intelligent agent of the power distribution network, security verification, and convergence determination to output the Pareto optimal solution set and the corresponding optimization scheme.

[0010] Preferably, the topology adjustment action is output by constructing a first reward function for the first agent, and the calculation expression of the first reward function is: In the formula, To improve reliability, For voltage quality improvement, To cover renovation costs, , , Weights are assigned to reliability, quality improvement, and modification. The second reward function of the second agent is used to output investment timing actions. The calculation expression of the second reward function is as follows: In the formula, To improve operational economic efficiency, Conditional risk value, This refers to the risk weighting coefficient. The expression for calculating the potential function is: In the formula, , , These are the potential energy term of the distribution network structure, the economic potential energy term, and the strategy distribution divergence penalty term, respectively. For environmental conditions, , These are topology adjustment actions and investment timing actions, respectively. , Let be the action probability distribution of the first agent and the action probability distribution of the second agent. This is the penalty coefficient.

[0011] Preferably, the calculation expression for the differential reward function is: In the formula, Let i be the differential reward for the first agent and the second agent. Let i be the initial reward for the first and second agents. , , These represent the actions and action probability distributions of agents with indices -i, i, and -i, respectively. The expected action value of the agent with index -i; The calculation expression for the reward weight vector is as follows: In the formula, , The weight vector with time series t+1 and t, For learning rate, , This is the regret value matrix with time series t and k+1. This is the projection function.

[0012] Preferably, the state space of the distribution network includes at least the voltage sensitivity matrix of the distribution network nodes, line load balance, distributed power output, load time sequence data, and equipment operating status.

[0013] Preferably, the action space is obtained through a continuous hybrid decision-making mode, including discrete actions such as grid topology switching operations and equipment switching status, and continuous actions such as equipment capacity reactive power compensation and investment scale adjustment.

[0014] Preferably, the expression for calculating the loss function is: In the formula, For the loss function based on reinforcement learning, The penalty coefficient is... The N-1 safety criterion for the occurrence of a fault. As a safety threshold, The confidence level.

[0015] Preferably, the real-time working data includes at least distribution network topology data, SCADA measured data stream, distributed power generation output data, and load time series data, and the preprocessing includes at least filtering and normalization processing.

[0016] Preferably, the convergence determination rule is as follows: The inequality for ε-Nash equilibrium is: ε≤0.01, where ε is the accuracy parameter of ε-Nash equilibrium; the convergence time of nodes in the distribution network in the first and second agents is ≤3.2 hours.

[0017] A second aspect of the present invention provides a multi-objective control device for a distribution network based on potential game theory and dynamic rewards, comprising: A dual-agent potential game construction module is used to construct a first agent that includes the characteristics of the power grid architecture and a second agent that includes the economic characteristics of the power grid. Based on the above two agents, a multi-objective potential function is designed. A two-layer dynamic reward shaping module is used to obtain a differential reward function by constructing a counterfactual baseline in the inner layer based on the first agent and the second agent, and to update the reward weight vector by constructing a regret matching algorithm in the outer layer. The space definition module is used to construct the state space and action space of the power distribution network; A security constraint embedding module is used to update the loss function based on the policy network by adding an N-1 security criterion chance constraint penalty term; The agent training module is used to update network parameters by training agents in the distribution network. The scheme generation module is used to input real-time working data, and through preprocessing and intelligent agents, security verification and convergence determination of the power distribution network, output Pareto optimal solution set and corresponding optimization scheme.

[0018] A third aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the multi-objective control method for a distribution network based on potential game theory and dynamic rewards as described in any of the preceding claims.

[0019] Because the present invention adopts the above technical solution, it has the following advantages and positive effects compared with the prior art: 1. Fast convergence speed: Compared with the traditional MARL method, the number of training rounds is reduced by 50%, achieving faster convergence speed in 10kV node distribution networks. The time to Nash equilibrium is reduced to 3.2 hours, significantly improving optimization efficiency.

[0020] 2. Multi-objective balance: The coverage of Pareto optimal solution set is improved by 35%, and the power supply reliability index is improved by 8.7% while the economic cost is reduced by 12%, achieving the optimal balance between reliability, economy and voltage quality.

[0021] 3. Solid theoretical support: For the first time, the theory of potential game theory is introduced into the field of multi-agent reinforcement learning in power distribution networks, which theoretically ensures that the Nash equilibrium coincides with the Pareto optimal solution and solves the problem of multi-objective conflict.

[0022] 4. High practicality: The dedicated state-action space is closely aligned with the actual operation requirements of the distribution network, and the optimized scheme is closely bound to the specific distribution network topology data and measured data stream, avoiding universality deviation.

[0023] 5. Embedding N-1 security criterion opportunity constraints ensures the safety and feasibility of the optimization scheme and reduces the operation risk of the distribution network. Attached Figure Description

[0024] The specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings, wherein: Figure 1 This is the main flowchart of a multi-objective control method for distribution networks based on potential game theory and dynamic rewards in this invention; Figure 2 This is a diagram of the dual-agent potential game collaborative optimization model architecture in this invention. Figure 3 This is a flowchart for obtaining the optimization scheme in this invention. Detailed Implementation

[0025] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. The advantages and features of the present invention will become clearer from the following description and claims. It should be noted that the drawings are all in a very simplified form and use non-precise ratios, and are only used to facilitate and clarify the illustration of the embodiments of the present invention.

[0026] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present invention are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indication will also change accordingly.

[0027] First Embodiment See Figure 1 The first aspect of the present invention provides a multi-objective control method for distribution networks based on potential game theory and dynamic rewards, comprising: S100: Construct a first intelligent agent that includes the characteristics of the power grid architecture and a second intelligent agent that includes the economic characteristics of the power grid. Design a multi-objective potential function based on the above two intelligent agents. S200: Based on the first and second intelligent agents, the differential reward function is obtained by constructing a counterfactual baseline in the inner layer, and the reward weight vector is updated by constructing a regret matching algorithm in the outer layer; S300: Constructs the state space and action space of the power distribution network; S400: The loss function based on policy networks is updated by adding an N-1 security criterion chance constraint penalty term; S500: Updates network parameters by training agents in the distribution network; S600: Input real-time working data, through preprocessing and distribution network intelligent agent, security verification and convergence determination, output Pareto optimal solution set and corresponding optimization scheme.

[0028] This application, for the first time, deeply integrates game theory into a multi-agent reinforcement learning framework for power distribution networks, mathematically guaranteeing the automatic alignment of agent selfish behavior with the overall interests of the system. It solves the theoretical problems of agent policy conflict and Pareto optimality uncoupling in traditional methods, providing a new methodology with a solid game theory foundation for the collaborative optimization of complex systems. Through two-layer reward shaping and dedicated state-action space design, the algorithm can dynamically weigh and accurately balance multiple conflicting objectives. It is not only an intelligent decision-making engine that outputs a reliable, comparable, and executable set of optimal solutions, but also elevates mathematically feasible solutions to engineering-optimal solutions. Its empirical performance (efficiency improvement of over 50%, cost reduction of 12% while reliability improvement of 8.7%, and 100% satisfaction of the N-1 criterion) signifies that power distribution network optimization has moved from conservative and extensive control relying on experience to lean collaborative autonomy driven by data and algorithms, providing technical support for the digital transformation and intelligent upgrading of the power grid.

[0029] In the above process, S500 and S600 are optimized processes, specifically: Data acquisition and preprocessing: Collect distribution network topology data, SCADA measured data streams, distributed power generation output data, load time series data, etc. Preprocessing includes at least filtering and normalization.

[0030] Initialize the parameters of the agents: network structure of the first and second agents, and penalty coefficients of the potential function. Initial value of the reward function ; Iterative training: The first and second agents perform actions in the state space, calculate rewards through a two-layer dynamic reward shaping mechanism, and update the policy network and potential function parameters.

[0031] Dynamic iteration: In each iteration, the reward weight vector is updated based on the regret value matrix. Optimize the decision-making direction of the intelligent agent.

[0032] Safety constraint verification: The solution is ensured to meet the opportunity constraint penalty term in the loss function. Safety guidelines.

[0033] Convergence criterion: When convergence is reached The Nash equilibrium ε ≤ 0.01, where ε is the accuracy parameter of the ε-Nash equilibrium; the convergence time of the nodes in the distribution network for the first and second agents is ≤ 3.2 hours. This application takes a large 10kV distribution network as an example to verify the effectiveness of the method. The specific parameters are as follows: Distribution network parameters: voltage level 110kV / 10kV, total line length 1200km, distributed power penetration rate 25% (photovoltaic + wind power), peak load 1500MW; Optimization goals: Improve power supply reliability, reduce economic costs, and improve voltage quality (node ​​voltage deviation ≤ ±5%). Comparison methods: traditional multi-agent reinforcement learning (MARL) method and static weight multi-objective optimization method.

[0034] The real-time data collection during the above process includes: 1-year scale basic data, specifically including: 1) Topology data, including line resistance / reactance, node connection relationships, and equipment parameters (circuit breakers, transformers, etc.); 2) Measured data, including real-time data such as node voltage, line current, and power flow collected by the SCADA system. In this embodiment, the sampling interval for measured data is 5 minutes. 3) Time series data, including distributed power generation output fluctuation data and user load time series data.

[0035] The above real-time working data is filtered to remove outliers and normalized to obtain a training dataset or an actual dataset.

[0036] The initialization and parameter setting process for the two agents is as follows: Initialize the network structure of the first and second agents, and the penalty coefficient of the potential function. The initial values ​​of the reward weights; the network structures of the first and second agents adopt the Deep Deterministic Policy Gradient (DDPG) network, with 3 hidden layers (256, 128, and 64 neurons respectively). Potential function parameter initialization: penalty coefficient The risk weighting coefficient is 0.05. It is 0.1. Reward weight vector initialization: These are respectively the reliability weight, the quality improvement weight, and the modification weight; Convergence threshold initialization: Confidence level It is 0.95.

[0037] The first and second intelligent agents execute the actions: The first intelligent agent outputs actions such as line switching and topology reconfiguration, while the second intelligent agent outputs the investment nodes (nodes 32 / 67 / 105), investment scale (total capacity 1200kVar), and timing (months 3 / 6 / 9) of the reactive power compensation equipment. The differential reward function and reward weight vector are iterated and updated: the contribution of the first agent and the second agent are decoupled through the differential reward function. The contributions of the two agents are evaluated, and the weight vector is updated each round based on the regret value matrix to eventually stabilize the weights. Based on the first and second intelligent agents, a differential reward function is obtained by constructing a counterfactual baseline in the inner layer, and the reward weight vector is updated by constructing a regret matching algorithm in the outer layer. ; Safety constraint verification: Ensure that the line load rate is ≤80% and the node voltage deviation is ≤±5% under N-1 fault conditions through opportunity constraint penalty terms; Convergence criterion: Iteration reached after 8000 rounds of training. Nash equilibrium was reached, taking 3.2 hours. The experimental results are shown in Table 1. The above operations show that in 10kV node distribution networks, the convergence time of this method is 50% shorter than that of the traditional MARL method, the economic cost reduction rate is increased by 4.2 percentage points, the power supply reliability improvement rate is increased by 4.5 percentage points, and the N-1 safety criterion is met 100%, which verifies its efficiency, safety and practicality in multi-objective collaborative optimization of large-scale distribution networks.

[0038] This application selects traditional deterministic programming (worst-case scenario) and conventional Monte Carlo method (10,000 samplings) as comparison benchmarks. From four core dimensions, namely economic cost, operational performance, safety level and computational efficiency, the technical advantages of the method of this invention are quantitatively analyzed. The data are all from the actual test verification of the IEEE 33-node system (35kV voltage level, 30% photovoltaic penetration rate scenario), which provides objective data support for patent application and technical presentation. For specific comparison results, please refer to Table 2.

[0039] The advantages of this application compared to other methods are: 1) Economic and cost advantages: Precise cost control and avoidance of over-investment. Traditional deterministic planning uses worst-case scenario design, resulting in redundant investment in reactive power compensation equipment (total investment of 2.102 million yuan); this invention accurately characterizes uncertainty through an improved point estimation method, and combines opportunity constraints to balance security and cost, reducing total investment by 34.1%, while reducing the average annual network loss cost by 22.9%, and reducing the overall operating cost to 63% of the traditional method, achieving the optimal balance between economy and practicality.

[0040] 2) Safety and Performance Synergy: Strict Risk Control and Enhanced Operational Stability. In a high-fluctuation scenario with 30% photovoltaic penetration, this invention controls the voltage over-limit probability to 2.8%, which not only surpasses the national standard requirement of 5% but also reduces it by 39.1% compared to traditional deterministic planning. The node voltage qualification rate is increased to 97.2%, effectively solving the voltage fluctuation problem caused by distributed power source integration. Simultaneously, the branch power overload probability is only 1.5%, ensuring the implementation of the N-1 safety principle in the distribution network.

[0041] 3) Breakthrough in computational efficiency: Highly efficient solution, adapted to engineering applications. Conventional Monte Carlo methods require 10,000 samples and take up to 36.5 hours to compute, which is difficult to meet the real-time planning needs of engineering projects. This invention adopts a genetic algorithm that combines quantum behavior. Through quantum state encoding and adaptive crossover mutation mechanism, the computation time is shortened to 0.8 hours, the efficiency is improved by more than 40 times, and the number of iterations is reduced by 84%, which can quickly respond to the dynamic planning needs of power distribution networks.

[0042] Economically, the overall operating cost is reduced by 37.0%, avoiding over-investment; in terms of performance, the voltage qualification rate is improved by 1.8 percentage points, and network losses are reduced by 22.9%; in terms of safety, the probability of voltage exceeding limits is better than the national standard, and the risk is reduced by more than 39.1%; in terms of efficiency, the calculation time is shortened by 40 times, making it suitable for practical engineering applications. The data fully verify the practical value of this method in high-fluctuation distributed power source access scenarios, providing a more scientific and efficient solution for reactive power planning in distribution networks.

[0043] See Figure 2 Preferably, the topology adjustment action is output by constructing a first reward function for the first intelligent agent. The calculation expression of the first reward function is as follows: In the formula, To improve reliability, For voltage quality improvement, To cover renovation costs, , , Weights are assigned to reliability, quality improvement, and modification. By constructing a second reward function for the second agent, the investment timing actions are output. The calculation expression for the second reward function is as follows: In the formula, To improve operational economic efficiency, Conditional risk value, This refers to the risk weighting coefficient. The expression for calculating the potential function is: In the formula, , , These are the potential energy term of the distribution network structure, the economic potential energy term, and the strategy distribution divergence penalty term, respectively. For environmental conditions, , These are topology adjustment actions and investment timing actions, respectively. , Let be the action probability distribution of the first agent and the action probability distribution of the second agent. This is the penalty coefficient.

[0044] The first reward function integrates three objectives with different dimensions and directions—reliability, voltage quality, and modification cost—into a unified, optimizable signal through reliability weights, quality improvement weights, and modification weights. This allows the first agent A to adaptively adjust its optimization focus based on the system state and training phase, avoiding the rigidity and subjectivity of traditional static weighted methods. The second reward function not only pursues improved operational economy but also innovatively incorporates Conditional Value at Risk (CVaR) as a penalty term. This imbues the second agent's decision-making with risk aversion, enabling it to proactively manage tail risks arising from uncertainty while pursuing returns, thus outputting a more robust and reliable investment plan. The potential function is the theoretical core of this scheme. It unifies the objectives of the two agents into a global situational potential function and adds a policy distribution divergence penalty term. This ensures that any action by an agent to increase its own reward will inevitably lead to an increase in global situational potential. The policy distribution divergence penalty term forces the decision-making approaches of the two agents to converge at the policy level, preventing strategic conflicts. The mathematical model ensures that the equilibrium point of system convergence (Nash equilibrium) is also the optimal point of the overall system (Pareto optimality), thus solving the fundamental problem of multiple agents acting independently.

[0045] Preferably, the calculation expression for the differential reward function is as follows: In the formula, Let i be the differential reward for the first agent and the second agent. Let i be the initial reward for the first and second agents. , , These represent the actions and action probability distributions of agents with indices -i, i, and -i, respectively. The expected action value of the agent with index -i; The expression for calculating the reward weight vector is: In the formula, , The weight vector with time series t+1 and t, For learning rate, , This is the regret value matrix with time series t and k+1. This is the projection function.

[0046] The differential reward received by each agent reflects the marginal contribution of its actions. This makes the learning objective extremely clear, improving the accuracy and efficiency of individual learning, which is the direct mathematical reason for the increased convergence speed of the algorithm. This encourages agents to explore and learn actions that can proactively create synergies and compensate for high-value actions, thereby driving the entire system to evolve towards a higher-performance Pareto front, rather than simply reaching an equilibrium.

[0047] See Figure 3 Preferably, the state space of the distribution network includes at least the voltage sensitivity matrix of the distribution network nodes, the line load balance, the output of distributed power sources, load time series data, and equipment operating status.

[0048] It accurately depicts the power grid's operational status, providing intelligent agents with comprehensive and practical decision-making information. It supports precise voltage regulation, directly improving voltage qualification rates. Embedded safety margin sensing prevents overload risks at the source. It endows strategies with adaptability to fluctuations, enhancing the integration and economic efficiency of renewable energy. It ensures all output actions are physically feasible, eliminating invalid operations.

[0049] Preferably, the action space is obtained through a continuous hybrid decision-making mode, including discrete actions such as grid topology switching operations and equipment switching status, as well as continuous actions such as equipment capacity reactive power compensation and investment scale adjustment.

[0050] This hybrid action space design enables the intelligent agent to simultaneously output discrete commands for switching and continuous commands for capacity adjustment, ensuring that optimized solutions can be directly converted into control commands without secondary conversion. This improves the precision and coordination of decision-making and significantly enhances the direct executability of the solution, making it a key design feature that guarantees the practicality of this invention.

[0051] Preferably, the expression for calculating the loss function is: In the formula, For the loss function based on reinforcement learning, The penalty coefficient is... The N-1 safety criterion for the occurrence of a fault. As a safety threshold, The confidence level.

[0052] By actively penalizing strategies that violate safety confidence levels during training, the agent is guided to automatically learn optimized solutions that strictly meet engineering safety standards, rather than through post-hoc verification. This ensures that the final solution possesses both theoretical optimality and engineering feasibility, directly supporting the empirical results of 100% satisfaction of the N-1 safety criterion and fundamentally eliminating operational risks.

[0053] See Figure 3Preferably, the real-time working data includes at least distribution network topology data, SCADA measured data stream, distributed power generation output data, and load time series data, and the preprocessing includes at least filtering and normalization processing.

[0054] Topological data defines the optimization boundary, SCADA real-time data reflects the instantaneous state, and distributed power source and load time-series data characterize the core fluctuation sources, together forming a spatiotemporally comprehensive input to ensure that optimization is based on real-world scenarios. Filtering removes measurement noise and outliers, avoiding interference from dirty data on policy learning; normalization unifies the dimensions and scales of multi-source data, significantly improving the numerical stability and convergence efficiency of model training. The preprocessed data can be directly and efficiently input into the defined dedicated state space, supporting the agent in accurate situational awareness and rapid decision-making, serving as a crucial data bridge for the entire method to move from theory to engineering application.

[0055] See Figure 3 The preferred convergence determination rule is: The inequality for ε-Nash equilibrium is: ε≤0.01, where ε is the accuracy parameter of ε-Nash equilibrium; the convergence time of nodes in the distribution network in the first and second agents is ≤3.2 hours.

[0056] Setting ε≤0.01 as the accuracy standard provides a rigorous mathematical criterion for ε-Nash equilibrium. This ensures that the final output strategy combination is a highly stable game equilibrium point, theoretically guaranteeing the stability and optimality of the solved strategy. The requirement of a convergence event ≤3.2 hours transforms the abstract concept of convergence into a clear engineering time constraint, directly verifying the effectiveness of the aforementioned two-layer dynamic reward shaping mechanism in accelerating convergence, and meeting the rigid timeliness requirements of practical engineering applications.

[0057] Second Embodiment A second aspect of the present invention provides a multi-objective control device for a distribution network based on potential game theory and dynamic rewards, comprising: A dual-agent potential game construction module is used to construct a first agent that includes the characteristics of the power grid architecture and a second agent that includes the economic characteristics of the power grid. Based on the above two agents, a multi-objective potential function is designed. A two-layer dynamic reward shaping module is used to obtain the differential reward function by constructing a counterfactual baseline in the inner layer based on the first and second intelligent agents, and to update the reward weight vector by constructing a regret matching algorithm in the outer layer. The space definition module is used to construct the state space and action space of the power distribution network; A security constraint embedding module is used to update the loss function of the policy network by adding an N-1 security criterion chance constraint penalty term; The agent training module is used to update network parameters by training agents in the distribution network. The scheme generation module is used to input real-time working data, which is preprocessed and processed by the intelligent agent of the distribution network, security verification and convergence determination to output Pareto optimal solution set and corresponding optimization scheme.

[0058] The dual-agent potential game construction module and the dual-layer dynamic reward shaping module correspond to the core algorithm engine, ensuring the theoretical advantages of collaborative optimization and efficient convergence. The space definition module and the safety constraint embedding module correspond to the engineering adaptation and safety kernel, ensuring the practicality and security of the decision-making. The agent training module and the scheme generation module correspond to the system operation and output interface, completing the end-to-end closed loop from data to decision. This design ensures that every theoretical innovation of the patent can be faithfully and efficiently executed in the device.

[0059] Third Embodiment A third aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of a multi-objective control method for a power distribution network based on potential game theory and dynamic rewards as described above.

[0060] Through modular packaging, this device transforms the innovative potential game collaborative theory and dynamic reward shaping algorithm into a standardized system with a clear structure that can be independently deployed and maintained, achieving a key leap from methodology to deliverable product.

[0061] In the description of this application, it should be noted that the terms "inner" and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship commonly used when the product is in use. They are used only for the convenience of describing this application and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application. Furthermore, the terms "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0062] It should also be noted that, unless otherwise explicitly specified and limited, the terms "setup" and "connection" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.

[0063] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific identification content executed by the system and device described above can be referred to the corresponding process in the foregoing method embodiments.

[0064] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings, but the present invention is not limited to the above embodiments. Even if various changes are made to the present invention, if these changes fall within the scope of the claims of the present invention and their equivalents, they shall still fall within the protection scope of the present invention.

Claims

1. A multi-objective control method for distribution networks based on potential game theory and dynamic rewards, characterized in that, include: Construct a first intelligent agent that includes the characteristics of the power grid architecture and a second intelligent agent that includes the economic characteristics of the power grid. Design a multi-objective potential function based on the above two intelligent agents. Based on the first intelligent agent and the second intelligent agent, the differential reward function is obtained by constructing a counterfactual baseline in the inner layer, and the reward weight vector is updated by constructing a regret matching algorithm in the outer layer; Construct the state space and action space of the power distribution network; The loss function based on the policy network is updated by adding an N-1 security criterion chance constraint penalty term; Network parameters are updated by training agents in the distribution network; The input real-time working data is preprocessed and processed by the intelligent agent of the power distribution network, security verification, and convergence determination to output the Pareto optimal solution set and the corresponding optimization scheme.

2. The multi-objective control method for distribution networks based on potential game theory and dynamic rewards according to claim 1, characterized in that, The topology adjustment action is output by constructing a first reward function for the first agent. The calculation expression for the first reward function is as follows: In the formula, To improve reliability, For voltage quality improvement, To cover renovation costs, , , Weights are assigned to reliability, quality improvement, and modification. The second reward function of the second agent is used to output investment timing actions. The calculation expression of the second reward function is as follows: In the formula, To improve operational economic efficiency, Conditional risk value, This refers to the risk weighting coefficient. The expression for calculating the potential function is: In the formula, , , These are the potential energy term of the distribution network structure, the economic potential energy term, and the strategy distribution divergence penalty term, respectively. For environmental conditions, , These are topology adjustment actions and investment timing actions, respectively. , Let be the action probability distribution of the first agent and the action probability distribution of the second agent. This is the penalty coefficient.

3. The multi-objective control method for distribution networks based on potential game theory and dynamic rewards according to claim 1, characterized in that, The calculation expression for the differential reward function is as follows: In the formula, Let i be the differential reward for the first agent and the second agent. Let i be the initial reward for the first and second agents. , , These represent the actions and action probability distributions of agents with indices -i, i, and -i, respectively. The expected action value of the agent with index -i; The calculation expression for the reward weight vector is as follows: In the formula, , The weight vector with time series t+1 and t, For learning rate, , This is the regret value matrix with time series t and k+1. This is the projection function.

4. The multi-objective control method for distribution networks based on potential game theory and dynamic rewards according to claim 1, characterized in that, The state space of the distribution network includes at least the voltage sensitivity matrix of the distribution network nodes, line load balance, distributed power output, load time series data, and equipment operating status.

5. The multi-objective control method for distribution networks based on potential game theory and dynamic rewards according to claim 1, characterized in that, The action space acquires discrete actions, including grid topology switching operations and equipment switching status, as well as continuous actions, including equipment capacity reactive power compensation and investment scale adjustment, through a continuous hybrid decision-making mode.

6. The multi-objective control method for distribution networks based on potential game theory and dynamic rewards according to claim 1, characterized in that, The expression for calculating the loss function is as follows: In the formula, For the loss function based on reinforcement learning, The penalty coefficient is... The N-1 safety criterion for the occurrence of a fault. As a safety threshold, The confidence level.

7. The multi-objective control method for distribution networks based on potential game theory and dynamic rewards according to claim 1, characterized in that, The real-time working data includes at least distribution network topology data, SCADA measured data stream, distributed power output data, and load time series data. Preprocessing includes at least filtering and normalization.

8. The multi-objective control method for distribution networks based on potential game theory and dynamic rewards according to claim 1, characterized in that, The convergence determination rule is as follows: The inequality for ε-Nash equilibrium is: ε≤0.01, where ε is the accuracy parameter of ε-Nash equilibrium; the convergence time of nodes in the distribution network in the first and second agents is ≤3.2 hours.

9. A multi-objective control device for distribution networks based on potential game theory and dynamic rewards, characterized in that, include: A dual-agent potential game construction module is used to construct a first agent that includes the characteristics of the power grid architecture and a second agent that includes the economic characteristics of the power grid. Based on the above two agents, a multi-objective potential function is designed. A two-layer dynamic reward shaping module is used to obtain a differential reward function by constructing a counterfactual baseline in the inner layer based on the first agent and the second agent, and to update the reward weight vector by constructing a regret matching algorithm in the outer layer. The space definition module is used to construct the state space and action space of the power distribution network; A security constraint embedding module is used to update the loss function based on the policy network by adding an N-1 security criterion chance constraint penalty term; The agent training module is used to update network parameters by training agents in the distribution network. The scheme generation module is used to input real-time working data, and through preprocessing and intelligent agents, security verification and convergence determination of the power distribution network, output Pareto optimal solution set and corresponding optimization scheme.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the computer program is executed by the processor, it implements the steps of a multi-objective control method for a distribution network based on potential game theory and dynamic rewards as described in any one of claims 1-8.