Power system load dynamic optimization method based on reinforcement learning

Through the combination of layered reinforcement learning and evolutionary optimization, the strategy instability and adaptability problems in multi-level scheduling scenarios in power system load scheduling are solved, efficient and accurate load scheduling control is achieved, and the response speed and stability of the power system are improved.

CN120258254AActive Publication Date: 2025-07-04ZHEJIANG JUHUA THERMAL POWER CO LTD

Patent Information

Application Number
CN202510741404.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-07-04
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

The existing reinforcement learning methods have problems such as slow convergence in multi-level scheduling scenarios, easy to fall into local optimality, unclear control particle size, insufficient generalization ability of strategy and poor adaptability in power system load scheduling, especially in cross-regional power grids and volatile load environments, which are difficult to deploy quickly.

Method used

The hierarchical reinforcement learning model is adopted to divide the load scheduling tasks of the power system into two subtasks: high-level decision making and low-level execution. Combining evolutionary optimization algorithm and improved Soft Actor-Critic strategy gradient method, scheduling task target category instructions are generated through the high-level decision model, the low-level execution model outputs specific control actions, and a state distribution gating mechanism, node attention mechanism and dynamic entropy adjustment strategy are introduced to achieve efficient scheduling.

Benefits of technology

It realizes efficient, accurate and stable load scheduling of power system, improves clarity of scheduling processes and system adaptability, and improves scheduling efficiency and strategy robustness in multi-region power grid scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258254A_ABST
    Figure CN120258254A_ABST
Patent Text Reader

Abstract

The invention discloses a power system load dynamic optimization method based on reinforcement learning. The method comprises the following steps: S1, collecting power system data to construct a state space; s2, constructing a hierarchical reinforcement learning model based on the state space, and dividing a high-level decision and a low-level execution task; s3, training a high-level decision model, and outputting a scheduling task target category instruction in a high-level state; s4, training a low-layer execution model, and outputting a control action in combination with a current node state and a high-layer instruction; s5, introducing an evolutionary mechanism to generate a strategy population and optimizing a low-layer execution model; s6, fusing an evolutionary mechanism and a strategy gradient to synchronously optimize individuals with excellent performance; s7, deploying the trained model to a power dispatching system; s8, performing model parameter fine tuning based on scheduling feedback; and S9, continuously applying the fine-tuned model to load scheduling control. According to the invention, power load accurate scheduling and strategy efficient adaptive optimization are realized, and system responsiveness and operation stability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power load optimization, and particularly to a dynamic optimization method for power system load based on reinforcement learning. Background Art

[0002] Currently, with the continuous increase in the penetration rate of renewable energy in the power system and the increasing complexity of user-side electricity consumption behaviors, traditional static load scheduling methods are difficult to cope with the increasingly dynamic and variable operating environment. Especially in scenarios such as demand response, energy storage control, and load migration, scheduling tasks face problems such as high-dimensional state space, continuous action output, and system uncertainty. To address the above challenges, researchers have widely introduced reinforcement learning methods to optimize the scheduling of power system loads. Among them, policy gradient algorithms have received attention for their performance in continuous control problems, and representative methods include Advantage Actor-Critic (A2C), Proximal Policy Optimization (PPO), and Soft Actor-Critic (SAC), etc.

[0003] In the prior art, most load optimization methods adopt a single-layer reinforcement learning structure, that is, a mapping is directly established from the system state to the scheduling action in a unified policy model. When facing multi-level and multi-objective scheduling scenarios of the power grid, such methods often show problems such as slow convergence, easy to fall into local optima, and unclear control granularity. In addition, some studies attempt to improve the model performance through policy gradient methods, but they are often limited to the standard forms of A2C or PPO, and there are still problems of poor generalization ability and insufficient stability when dealing with the complex state structure and high-dimensional decision-making actions of the power system.

[0004] In the research on evolutionary methods for power load optimization tasks, traditional genetic algorithms are used to optimize control strategies, but most only stay in fixed fitness evaluation and parameter mutation methods, lacking deep integration with the policy network structure, and it is difficult to capture the structural differences or dynamic characteristics between individuals. At the same time, current reinforcement learning methods often ignore the adaptability limitations of policy models in terms of migration and generalization capabilities, and it is difficult to be quickly deployed especially in cross-regional power grids or volatile load environments.

[0005] Based on this, the existing technologies still face the following problems: First, the reinforcement learning model fails to effectively combine multi-layer structures with task division and lacks a scheduling process with strong controllability and clear responses; Second, the two optimization paths of evolutionary optimization and policy gradient do not form a collaborative mechanism, suffering from defects such as information fragmentation and isolated optimization processes; Third, although policy gradient methods such as SAC are robust, their specific structures in power scenarios have not been improved for state clustering, key node perception, and policy distribution control, restricting their performance upper limits. Therefore, there is an urgent need for a power system load dynamic optimization scheme that integrates hierarchical modeling, evolutionary optimization, and improved policy gradient methods to improve scheduling efficiency, model stability, and adaptability.

[0006] Therefore, how to provide a power system load dynamic optimization method based on reinforcement learning is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0007] An object of the present invention is to propose a power system load dynamic optimization method based on reinforcement learning. The present invention integrates hierarchical reinforcement learning, evolutionary optimization algorithm, and improved Soft Actor-Critic policy gradient method, details the hierarchical decision-making structure from high-level scheduling policy generation to low-level control action output in the power system, and introduces a structural optimization mechanism to enhance the policy network, having the advantages of high control accuracy, fast policy convergence, and strong system adaptability.

[0008] The power system load dynamic optimization method based on reinforcement learning according to an embodiment of the present invention includes the following steps:

[0009] S1. Collect data of the power system and construct a state space;

[0010] S2. Based on the state space, construct a hierarchical reinforcement learning model, and divide the hierarchical reinforcement learning model into a high-level decision-making model and a low-level execution model;

[0011] S3. Train the high-level decision-making model, where the high-level decision-making model takes the high-level state in the state space as input and outputs a scheduling task target category instruction;

[0012] S4. Train the low-level execution model, where the low-level execution model takes the current node-level state and the scheduling task target category instruction output by the high-level decision-making model as input and outputs a control action sequence including the load migration amount of each node, energy storage charge and discharge instructions, and user response incentive coefficients;

[0013] S5. When training the low-level execution model, introduce an evolutionary reinforcement learning mechanism, construct a number of policy individuals, calculate the reward value for each policy individual through simulated environment interaction, and perform selection, crossover, and mutation operations on the population composed of policy individuals using an improved genetic algorithm to obtain a new generation of policy individuals;

[0014] S6. Combine the evolutionary reinforcement learning mechanism with the optimization method of policy gradient. While evolving and updating the policy individuals, use the optimization method based on policy gradient to further update the parameters of the relatively better-performing policy individuals.

[0015] S7. Deploy the hierarchical reinforcement learning model to the power system load scheduling module.

[0016] S8. During the deployment and operation, continuously collect scheduling feedback data to fine-tune the parameters of the deployed hierarchical reinforcement learning model.

[0017] S9. Continuously apply the high-level decision-making model and the low-level execution model after transfer and fine-tuning to the load scheduling process of the power system.

[0018] Optionally, the S2 specifically includes:

[0019] S21. Divide the state space into a high-level state space and a low-level state space. The high-level state space includes the global load trend of the power system, the electricity price fluctuation range, and the renewable energy output prediction data. The low-level state space includes the real-time load values of each node, the energy storage state information, and the controllable load regulation ability parameters.

[0020] S22. Build a high-level decision-making model based on the high-level state space. The high-level decision-making model takes the high-level state as the input and outputs the scheduling task target category instructions. The scheduling task target category instructions include peak shaving and valley filling scheduling instructions, energy storage priority regulation instructions, or demand response control instructions.

[0021] S23. Build a low-level execution model based on the low-level state space and the scheduling task target category instructions output by the high-level decision-making model. The low-level execution model outputs the control action sequences of the load migration amount of each power grid node, the energy storage charge and discharge control instructions, and the user response incentive coefficient.

[0022] S24. Establish the mapping relationship between the high-level decision-making model and the low-level execution model, so that the scheduling task target category instructions serve as the input conditions of the low-level execution model, constituting a hierarchical and linked scheduling control structure.

[0023] Optionally, the S3 specifically includes:

[0024] S31. Obtain the input data of the high-level state space as the training sample.

[0025] S32. Build the high-level decision-making policy function , where represents the high-level decision-making policy function, is the input high-level state input, is the output scheduling task target category instruction;

[0026] S33. Define the immediate reward function of the high-level decision-making model , where represents the time step, and evaluates the performance of the low-level execution results guided by the output of the high-level decision-making policy function in terms of load balancing, operating cost control, and scheduling volatility;

[0027] S34. Use the reinforcement learning method based on policy gradient to train the high-level decision-making policy function , and update the parameters of the high-level decision-making model by maximizing the expected value of the immediate reward function . When facing the input high-level state , output the optimal scheduling task target category instruction .

[0028] Optionally, the reinforcement learning method based on policy gradient in S34 is the SoftActor-Critic method based on maximum entropy policy optimization, which specifically includes:

[0029] S341. Construct two high-level Q-value function networks and , take the high-level state and the corresponding output scheduling task target category instruction as inputs, and respectively output Q-value estimates and , and use the smaller value to participate in policy optimization;

[0030] S342. Introduce an objective function containing a policy entropy term as the optimization objective for training the high-level decision-making policy function , which is used to guide the parameter update process under the given high-level state ;

[0031] S343. Alternately update the high-level decision-making policy function and the high-level Q-value function networks , , and adopt the following structural improvement methods: add a residual connection structure to the high-level decision-making policy function, use a shared state encoder in the Q-value function network, and introduce a dynamic entropy coefficient adjustment mechanism based on the fluctuation amplitude of the target Q network to replace the static or predefined entropy coefficient update method.

[0032] Optionally, S4 specifically includes:

[0033] S41. Define the low-level state input as , including the current load value of each node in the power grid, the current state of charge of the energy storage device corresponding to the node, and the adjustable load capacity:

[0034] S42. Combine the scheduling task target category instruction output by the high-level decision-making model with the low-level state input as the combined input of the low-level execution model to construct a low-level policy function;

[0035] S43. Construct a low-level immediate reward function , and evaluate the system performance generated by the low-level control action under the low-level state input and the scheduling task target category instruction output by the high-level decision-making model;

[0036] S44. Using the low-level immediate reward function as the optimization objective, train the low-level policy function , and output the optimal control action under the given low-level state input and the scheduling task target category instruction output by the high-level decision-making model; .

[0037] Optionally, the specific steps of S5 include:

[0038] S51. Initialize the low-level policy function and construct individuals of the low-level policy function , where each individual of the low-level policy function is a set of independent low-level policy network parameters, corresponding to the control action , being the control action of the th individual;

[0039] S52. Record the cumulative reward value of each policy individual in consecutive training cycles, and calculate the comprehensive fitness value of the policy individual with historical smoothed weights;

[0040] S53. In the selection operation of the genetic algorithm, adopt the selection method of the elite retention strategy, and preferentially retain the first policy individuals with the largest comprehensive fitness values, and at the same time perform roulette selection based on the ranking from the remaining individuals to generate a candidate parent set;

[0041] S54. Improve the offspring crossover operation, judge the performance similarity of two parent policy individuals in the policy parameter distribution space. If their Euclidean distance or KL divergence is less than the set threshold, then allow crossover, otherwise this combination is discarded;

[0042] S55. Perform a mutation operation on the selected offspring parameter set, introduce a target guidance mechanism, and calculate the mutation probability according to the current target residual of the policy individual;​

[0043] S56. Take the new generation of strategy individuals after screening, crossover, and target-guided mutation as the current strategy population, repeat steps S52 to S55, and iterate and update until convergence. Select strategy individuals as the output strategy structure of the low-level execution model according to the final comprehensive fitness value from high to low.

[0044] Optionally, step S6 specifically includes:

[0045] S61. After each generation of evolutionary operation, select the top strategy individuals from the current strategy population as the candidate subset for policy gradient optimization, where represents the th low-level strategy individual, ;

[0046] S62. For each strategy individual , construct a set of parameter-sharing double Q-value function networks , , with the low-level state input and the scheduling task target category instruction output by the high-level decision-making model and the corresponding control action as the input;

[0047] S63. Construct the low-level policy objective function ;

[0048] S64. Take as the optimization objective, and update the parameters of the low-level strategy individual and the of the corresponding Q-value function network respectively. The update adopts an asynchronous alternating training method, and adjusts the parameters of the target Q network based on the soft update rule;

[0049] S65. While evolving and updating the strategy individuals, further update the parameters of the relatively better-performing strategy individuals using an optimization method based on policy gradient. Among them, the optimization method based on policy gradient is the improved SoftActor-Critic method, and the SoftActor-Critic method has the following structural improvements: introduce a feature gating mechanism based on the clustering label of the input state distribution in the low-level double Q-value function network to adjust the response intensity of the Q-value channel, introduce a collaborative attention mechanism in the policy network, and adopt a dynamic target KL divergence adjustment strategy in the entropy coefficient update process to adaptively adjust according to the deviation amplitude of the policy distribution during the training process;

[0050] S66. Use the policy individual after being jointly trained by the genetic algorithm and the improved SoftActor-Critic method as the low-level execution model as the target output policy.

[0051] Optionally, the structural improvement of the low-level SoftActor-Critic method by S65 specifically includes:

[0052] S651. Introduce a state distribution gating mechanism in the low-level double Q-value function network and to cluster the low-level state input and obtain the state category label . Use the state category label as the input to the gating module, and activate the corresponding channel control factor in the feature extraction layer of the low-level double Q-value function network according to to perform gating adjustment on the intermediate features to generate a category-sensitive Q-value estimate;

[0053] S652. Introduce a cross-node state collaborative attention mechanism in the low-level policy individual . Use the state feature of each node that makes up the low-level state as the attention encoding input respectively, construct an attention weight distribution based on the inter-node dependency relationship, and perform weighted aggregation on each node feature as the input of the policy distribution;

[0054] S653. Introduce a dynamic adjustment mechanism for the entropy coefficient in the low-level policy network, set the target policy distribution offset value . Calculate the KL divergence between the current low-level policy individual and the previous low-level policy individual in each round of training. Adjust the entropy coefficient based on the deviation between . The adjustment rule of the deviation-adjusted entropy coefficient is as follows: If

[0055] If , increase ;

[0056] If , decrease ;

[0057] If , keep the current unchanged;

[0058] Control the exploration degree and training stability of the policy in a dynamic manner.

[0059] The beneficial effects of the present invention are as follows:

[0060] (1) By constructing a hierarchical reinforcement learning model, the present invention divides the power load scheduling task into two subtasks: high-level scheduling target decision-making and low-level specific execution control, solving the problems of task aliasing and fuzzy control granularity existing in existing reinforcement learning methods when facing complex multi-level scheduling tasks, realizing hierarchical response from global policy planning to local control execution, and effectively improving the clarity and decision-making efficiency of the scheduling process.

[0061] (2) The present invention introduces an evolutionary reinforcement learning mechanism in the training of the low-level execution model, and combines a policy gradient optimization method based on the Soft Actor-Critic method to perform two-stage hybrid training on high-quality individuals in the policy population, retaining both the diversity advantage of the evolutionary algorithm in global search and integrating the high-precision local update ability of the policy gradient, solving the problems of slow convergence and discontinuous policy update in existing evolutionary optimization methods, and significantly enhancing the policy performance and robustness.

[0062] (3) Based on the Soft Actor-Critic method, the present invention proposes structural improvements at the low level, specifically including a state distribution gating mechanism, a node attention mechanism, and a dynamic entropy adjustment strategy based on KL divergence, enhancing the model's ability to model the response difference of different types of node states and the ability to regulate the policy distribution, effectively solving the problems of insufficient generalization ability and unstable policy in the traditional SAC algorithm in the power dispatching scenario, and improving the adaptability and deployment efficiency of the model in multi-region and multi-type power grid scenarios. Description of the Drawings

[0063] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention. In the drawings:

[0064] Figure 1 is the overall flowchart of the power system load dynamic optimization method based on reinforcement learning proposed by the present invention. Detailed Embodiments

[0065] Now, the present invention will be further described in detail with reference to the drawings. These drawings are all simplified schematic diagrams, only showing the basic structure of the present invention in a schematic way, so they only show the components related to the present invention.

[0066] Refer to Figure 1 , the power system load dynamic optimization method based on reinforcement learning includes the following steps:

[0067] S1. Collect data of the power system and construct a state space, specifically including collecting key operation data such as real-time load, energy storage device status, user response ability, renewable energy output, electricity price information, and historical dispatching records of each node from the power dispatching system in the target area. The state space includes a high-level state space and a low-level state space. The high-level state space is used to describe the overall operation situation of the system, including the total network load, energy storage margin, electricity price trend, etc. The global features are uniformly input into the high-level decision-making model in vector form. The low-level state space is constructed as a set of local states of each node, including features such as the current load of the node, the state of charge of the energy storage, the adjustable load capacity, and the node category;

[0068] S2. Construct a hierarchical reinforcement learning model based on the state space, and divide the hierarchical reinforcement learning model into a high-level decision-making model and a low-level execution model;

[0069] S3. Train the high-level decision-making model. The high-level decision-making model takes the high-level state in the state space as input and outputs the scheduling task target category instruction;

[0070] S4. Train the low-level execution model. The low-level execution model takes the current node-level state and the scheduling task target category instruction output by the high-level decision-making model as input and outputs a control action sequence including the load transfer amount of each node, the energy storage charge and discharge instruction, and the user response incentive coefficient;

[0071] S5. When training the low-level execution model, introduce an evolutionary reinforcement learning mechanism, construct several policy individuals, calculate the reward value for each policy individual through simulated environment interaction, and perform selection, crossover, and mutation operations on the population composed of policy individuals using an improved genetic algorithm to obtain a new generation of policy individuals;

[0072] S6. Combine the evolutionary reinforcement learning mechanism with the optimization method of policy gradient. While the policy individuals evolve and update, use the optimization method based on policy gradient to further update the parameters of the relatively better-performing policy individuals;

[0073] S7. Deploy the hierarchical reinforcement learning model to the power system load dispatching module. Specifically: after completing the training of the high-level decision-making model and the low-level execution model, integrate the two into the power system load dispatching platform as the core for generating dispatching instructions. During the deployment process, first configure a model inference module in the master control center to receive the real-time collected state data as input. First, the high-level model generates the current scheduling task target category instruction, and then the low-level model outputs specific control actions based on the high-level instruction and the node state, including load transfer, energy storage charge and discharge, and user response incentive. The model sends the actions to each execution unit through the edge computing node or the control terminal to achieve distributed execution. The deployment scheme supports modular docking with the existing dispatching platform and can be loaded hierarchically by region to achieve fast response and scalable operation;

[0074] S8. During the deployment and operation period, continuously collect scheduling feedback data and fine-tune the parameters of the deployed hierarchical reinforcement learning model. Specifically: after the hierarchical reinforcement learning model goes online, the system continuously collects feedback information during the real-time scheduling process, including the execution results of scheduling instructions, actual load changes, energy storage response behaviors, user participation, and system volatility indicators, etc. Compare the above feedback information with the expected output results of the model, evaluate the performance deviation of the strategy in the real environment. If there are obvious deviations or environmental changes, such as changes in user behavior patterns, adjustments to the electricity price mechanism, or changes in the load structure, then start the fine-tuning mechanism. During the fine-tuning process, an incremental training method is adopted. On the basis of retaining the original model structure, new collected data is introduced to update the strategy parameters, especially dynamically adjust the strategy weights related to key nodes in the lower-level model;

[0075] S9. Continuously apply the migrated and fine-tuned high-level decision-making model and low-level execution model to the load scheduling process of the power system. Specifically: after completing the fine-tuning optimization, the updated high-level decision-making model and low-level execution model are reloaded into the scheduling system and run continuously as the main control strategy. The high-level model dynamically generates scheduling objectives according to the overall state of the system, such as peak shaving, energy storage priority, load shifting and other task instructions. The low-level model combines the local states of each node and high-level instructions to output refined control actions to achieve intelligent scheduling control of node loads, energy storage devices and user behaviors. The scheduling process runs in the online environment in real time, and the model can adaptively adjust the strategy output according to system load changes, grid topology adjustments and user behavior differences.

[0076] The present invention proposes a power system load dynamic optimization method integrating hierarchical reinforcement learning and evolutionary optimization, which realizes hierarchical response of scheduling tasks based on high-level and low-level strategy division, effectively solves the problems of unstable strategy and unclear control existing in the existing single-layer reinforcement learning structure when dealing with multi-objective and multi-granularity scheduling problems, and improves the optimization efficiency and controllability of strategy decision-making.

[0077] In this embodiment, the specific content of S2 includes:

[0078] S21. Divide the state space into a high-level state space and a low-level state space. The high-level state space includes the global load trend of the power system, the electricity price fluctuation range, and the predicted data of renewable energy output. The low-level state space includes the real-time load values of each node, the energy storage state information, and the controllable load regulation ability parameters;

[0079] S22. Construct a high-level decision-making model based on the high-level state space. The high-level decision-making model is responsible for capturing the overall operation trend of the system, formulating the macro scheduling direction, and serving as the guiding signal for low-level control. The high-level decision-making model takes the high-level state as the input and outputs the scheduling task target category instruction, where the scheduling task target category instruction includes the peak shaving and valley filling scheduling instruction, the energy storage priority regulation instruction, or the demand response control instruction;

[0080] S23. Construct a low-level execution model based on the low-level state space and the scheduling task target category instruction output by the high-level decision-making model. The low-level execution model takes the local state information of each node and the scheduling task target instruction output by the high-level as the input, generates the specific control action sequence, including the node load adjustment amount, the energy storage charge and discharge plan, and the user incentive strategy, and is responsible for realizing the refined execution of the high-level strategy at specific nodes. The low-level execution model outputs the control action sequence of the load migration amount, the energy storage charge and discharge control instruction, and the user response incentive coefficient of each power grid node;

[0081] S24. Establish the mapping relationship between the high-level decision-making model and the low-level execution model, so that the scheduling task target category instruction serves as the input condition of the low-level execution model, forming a hierarchical and linked scheduling control structure. Specifically: Take the scheduling task target category instruction output by the high-level decision-making model as an auxiliary input and directly input it into the low-level execution model, so that the low-level can not only consider the local state when generating the control actions of each node, but also combine the high-level instructions for strategy guidance. This mapping relationship realizes the effective connection from the global scheduling target to the local control execution, enables the two-layer model to form a unified input-output chain in structure, and ensures the consistency of the scheduling strategy and the accuracy of execution.

[0082] In view of the complexity of the power load scheduling task, the present invention proposes to divide the state space into high and low levels, and construct the corresponding high-level decision-making model and low-level execution model, realizing a task modeling process with a clear structure. Compared with the traditional integrated strategy model, it has stronger interpretability and scalability, and is suitable for large-scale scheduling system deployment.

[0083] In this embodiment, the specific content of S3 includes:

[0084] S31. Obtain the high-level state space as the input data of the training sample;

[0085] S32. Construct the high-level decision-making policy function , where represents the high-level decision-making policy function, is the input high-level state, is the output scheduling task target category instruction;

[0086] S33. Define the immediate reward function of the high-level decision-making model ,in, Represents the time step, evaluating the output of the high-level decision strategy function The performance of the guided low-level execution results in terms of load balancing, operating cost control and scheduling volatility, the reward function is: ;

[0087] in, represents the high-level strategy at time step The immediate reward value, represents the system load deviation cost under the current scheduling instruction, represents the scheduling resource cost, represents the volatility cost of scheduling actions, , , is the weighting coefficient;

[0088] S34, using policy gradient-based reinforcement learning methods to improve high-level decision-making policy functions Training is performed by maximizing the immediate reward function The expected value of the high-level decision model is updated, and the high-level state input is faced with the input Output the optimal scheduling task target category instruction .

[0089] The present invention establishes a high-level strategy training process by introducing advantage functions and instant reward construction mechanisms, and combines the policy gradient method to achieve fine optimization of target category instructions. This is different from the existing methods that have vague or non-discriminatory designs for strategy training objectives, and improves the rationality and stability of the scheduling strategy.

[0090] In this implementation, the policy gradient-based reinforcement learning method in S34 is a SoftActor-Critic method based on maximum entropy strategy optimization, specifically including:

[0091] S341. Construct two high-level Q-value function networks and , in high-level status The corresponding output scheduling task target category instruction As input, the output Q value estimation and , use the smaller value to participate in strategy optimization;

[0092] S342, introduce the objective function containing the policy entropy term as the training high-level decision-making policy function The optimization goal is to guide the optimization of the given high-level state The parameter update process under

[0093] S343, high-level decision strategy function and high-level Q value function network , Perform alternating updates, specifically: during model training, do not update the parameters of the policy network and the Q-value network at the same time, but first fix one, update the other, and then exchange them alternately. The following structural improvements are adopted: Add a residual connection structure to the high-level decision strategy function to improve training stability, use a shared state encoder in the Q-value function network to reduce parameter redundancy, and introduce a dynamic entropy coefficient adjustment mechanism based on the fluctuation amplitude of the target Q network to replace the static or predefined entropy coefficient update method.

[0094] The present invention explicitly defines the high-level strategy training method as Soft Actor-Critic, and introduces structural improvements such as residual connection and dynamic entropy adjustment mechanism, so that the high-level model exhibits stronger convergence stability and strategy distribution control ability when dealing with electricity price fluctuations and load uncertainties, which is better than traditional A2C or PPO methods.

[0095] In this implementation manner, the S4 specifically includes:

[0096] S41, low-level state input is defined as , including the current load value of each node in the power grid, the current charge state of the energy storage device corresponding to the node, and the adjustable load capacity:

[0097] S42, the scheduling task target category instruction output by the high-level decision model With low-level state input The combination is used as the joint input of the low-level execution model to construct the low-level policy function: ;

[0098] in, Represents the control action output by the low-level execution model, wherein the control action includes the load migration instruction of each node, the charge and discharge operation instruction of the energy storage device, and the user response incentive coefficient;

[0099] S43. Constructing low-level immediate reward function , evaluates the low-level state input and the scheduling task target category instructions output by the high-level decision model Lower level control actions The resulting system performance, the reward function is defined as: ;

[0100] in, represents the scheduling loss cost caused by load migration, represents the operating cost of the energy storage equipment at this moment, represents the incentive cost generated by motivating user responses is the weighted coefficient for each cost item;

[0101] S44. Using the low-level immediate reward function as the optimization objective, train the low-level policy function , and the training specifically aims to optimize the low-level policy function so that the control actions output by the low-level policy function can achieve tasks such as load adjustment, energy storage operation, and user response at each node, thereby maximizing the overall scheduling effect; under the condition of given low-level state input and the scheduling task objective category instruction output by the high-level decision-making model , output the optimal control action .

[0102] The present invention details the input structure and control action output form of the low-level execution model, enabling the model to simultaneously perceive the node load status, energy storage information, and user response capabilities, and generate multi-category schedulable instructions that can be implemented. Compared with the existing control models with a single load as the optimization objective, it has more practical scheduling and deployment value.

[0103] In this embodiment, the specific steps of S5 are as follows:

[0104] S51. Initialize the low-level policy function and construct individuals of the low-level policy function , and each individual of the low-level policy function is a set of independent low-level policy network parameters, corresponding to the control action , is the control action of the th individual;

[0105] S52. Record the cumulative reward value of each policy individual in consecutive training cycles , and calculate the comprehensive fitness value of the policy individual with historical smoothed weights: ;

[0106] wherein, is the comprehensive fitness value, is the time step of each training cycle, is the low-level immediate reward function of the th individual at the th cycle and the th step;

[0107] S53. In the selection operation of the genetic algorithm, adopt the selection method of the elite retention strategy and preferentially retain the top The strategy individual with the largest comprehensive fitness value, and at the same time, roulette selection is performed based on the ranking from the remaining individuals to generate a candidate parent set;

[0108] S54. Improve the offspring crossover operation. Specifically, on the basis of the traditional strategy parameter crossover, a strategy similarity discrimination mechanism is introduced to avoid invalid combinations. That is, before crossover, first calculate the similarity between the two parent strategy individuals in the parameter space. If the similarity is higher than the set threshold, it is considered that the strategy difference is insufficient, and no crossover is performed. Instead, the individual is directly copied or rematched. If the similarity is within a reasonable range, the crossover operation is executed. Judge the performance similarity between the two parent strategy individuals in the strategy parameter distribution space. If their Euclidean distance or KL divergence is less than the set threshold, crossover is allowed; otherwise, this combination is discarded to avoid invalid or degenerate strategy fusion;

[0109] S55. Perform a mutation operation on the selected offspring parameter set. Specifically, after the strategy parameter crossover is completed, a certain probability of perturbation is applied to the parameters of the generated offspring strategy individuals to introduce new strategy features and enhance the diversity of the strategy population. Different from the traditional method with a fixed mutation rate, the present invention adopts a guidance mechanism based on the target residual. Calculate the strategy target residual between each offspring individual and the current optimal strategy, and dynamically determine the mutation probability according to the size of the residual. The larger the residual, the higher the mutation probability; otherwise, it decreases; introduce a target guidance mechanism to calculate the mutation probability according to the current target residual of the strategy individual , specifically defined as: ;

[0110] where, is the mutation intensity coefficient, represents the gap index between the current target output of the strategy individual and the optimal individual;

[0111] S56. Use the new generation of strategy individual set after screening, crossover, and target-guided mutation as the current strategy population, and repeat S52 to S55 until convergence. Select the strategy individuals as the output strategy structure of the low-level execution model according to the final comprehensive fitness value from high to low.

[0112] The present invention introduces an evolutionary reinforcement learning mechanism, generates strategy individuals through evolutionary processes such as population evaluation and selection, crossover, and mutation, and deeply integrates with the reinforcement learning process, making up for the defects of strong dependence on the initial strategy and insufficient global exploration ability in traditional reinforcement learning, and improving the diversity and search ability of the scheduling strategy.

[0113] In this embodiment, the S6 specifically includes:

[0114] S61. After each generation of evolutionary operation, select the top A strategic individual As a candidate subset for policy gradient optimization, where represents the th low-level strategic individual, ;

[0115] S62. For each strategic individual , construct a set of dual Q-value function networks with shared parameters , , with the low-level state input and the scheduling task target category instruction output by the high-level decision-making model as well as the corresponding control action as the input;

[0116] S63. Construct the low-level policy objective function as: ;

[0117] Among them, is to take the expectation of the joint distribution of the low-level state and the scheduling task target category instruction output by the high-level decision-making model , represents the minimum value function, is the entropy weight coefficient, is the logarithm of the probability generated by the policy network for the current action;

[0118] S64. Take as the optimization objective, and update the parameters of the low-level strategic individual and the corresponding of the Q-value function network respectively. The update adopts an asynchronous alternating training method, and adjusts the parameters of the target Q network based on the soft update rule. Specifically, during the training process, the policy network and the Q-value function network are not updated simultaneously, but alternately: first fix the Q network, update the policy network, and then fix the policy network and update the Q network parameters, which is called the asynchronous alternating training method. In addition, to improve the training stability, the Q network adopts a soft update mechanism, that is, the parameters of the target Q network are not completely replaced by the current network parameters, but are smoothly updated according to a certain proportion;

[0119] S65. While evolving and updating the policy individuals, an optimization method based on policy gradient is used to further update the parameters of the better-performing policy individuals. Among them, the optimization method based on policy gradient is an improved SoftActor-Critic method, and the SoftActor-Critic method is structurally improved as follows: A feature gating mechanism based on the clustering labels of the input state distribution is introduced into the low-level double Q-value function network to adjust the response intensity of the Q-value channels. A collaborative attention mechanism is introduced into the policy network. During the update process of the entropy coefficient, a dynamic target KL divergence adjustment strategy is adopted, and it is adaptively adjusted according to the deviation amplitude of the policy distribution during the training process. ;

[0120] S66. The policy individuals jointly trained by the genetic algorithm and the improved SoftActor-Critic method are used as the target output policies of the low-level execution model.

[0121] The present invention combines evolutionary optimization with Soft Actor-Critic policy gradient training to further optimize the excellent low-level policy individuals, constructs an evolutionary-gradient fusion learning framework, and takes into account both policy accuracy and exploration efficiency compared with the separate evolutionary or separate gradient optimization schemes, achieving better performance.

[0122] In this embodiment, the structural improvement of the low-level SoftActor-Critic method in S65 specifically includes:

[0123] S651. A state distribution gating mechanism is introduced into the low-level double Q-value function network and to cluster the low-level state input to obtain state category labels . The state category labels are used as inputs to embed the gating module, and corresponding channel control factors are activated according to in the feature extraction layer of the low-level double Q-value function network to perform gating adjustment on the intermediate features to generate category-sensitive Q-value estimates.

[0124] S652. A cross-node state collaborative attention mechanism is introduced into the low-level policy individuals . The state feature of each node that composes the low-level state is used as the attention coding input respectively to construct an attention weight distribution based on the inter-node dependency relationship, and the features of each node are weighted and aggregated as the input of the policy distribution.

[0125] S653. A dynamic adjustment mechanism is introduced for the entropy coefficient in the low-level policy network, and a target policy distribution deviation value is set., calculate the current low-level policy individual in each round of training and the low-level policy individual in the previous round of KL divergence , based on and to adjust the entropy coefficient of the deviation between , the deviation-adjusted entropy coefficient The adjustment rule is:

[0126] If , then increase ;

[0127] If , then decrease ;

[0128] If , then keep the current unchanged;

[0129] Control the exploration degree and training stability of the policy in a dynamic way.

[0130] The present invention structurally improves the low-level Soft Actor-Critic method, introduces a state gating mechanism, a node attention mechanism and an entropy adjustment strategy based on KL divergence, forms a policy optimization path customized for the characteristics of power grid dispatching, is different from the traditional SAC structure, and improves the model's expression ability and output stability for complex states.

[0131] Example 1:

[0132] In order to verify the feasibility of the present invention in implementation, the present invention is applied to the real-time load optimal dispatching task of a regional distribution system within the jurisdiction of a provincial power grid control center. This region has 21 power grid dispatching nodes, covering urban residential areas, commercial areas and some light industrial areas, with typical characteristics of strong power consumption volatility, uneven user load response capabilities, and unstable renewable energy output. Moreover, the current dispatching mode combines static priority rules with empirical models, resulting in problems such as response delay, low dispatching efficiency, and poor load reallocation accuracy.

[0133] The main problems faced by the power system dispatching task include: during the two peak periods of 12:00 - 14:00 and 18:00 - 20:00 every day, the electricity load fluctuates violently, and local nodes are prone to overloading alarms; after a large number of household energy storage devices and user air conditioning devices are connected to the system, it causes a high degree of uncertainty in load forecasting and dispatching response; when the system responds to a sudden drop in new energy output (such as photovoltaic shading or wind speed mutation), the dispatching reaction is lagged, and the system stability cannot be controlled in time. The method proposed in the present invention takes reinforcement learning as the core and constructs a hierarchical dispatching structure. The high-level model is responsible for identifying the global state and formulating the dispatching target type, such as peak shaving, energy storage priority, or load reduction; the low-level model generates specific dispatching behaviors for each node according to the node state and high-level instructions, including energy storage instructions, load migration plans, and strategies to encourage users to participate in demand response.

[0134] During the implementation of this project, first, the electricity load data, electricity price signals, operation status of user energy storage, and photovoltaic output data in this area from July 2024 to January 2025 were collected to construct a state space and perform clustering modeling on different nodes. The high-level decision-making model was trained using an improved Soft Actor-Critic algorithm, and the classification accuracy of the dispatching task was improved by about 14.2% compared with the traditional K-means priority assignment model. The low-level control model combines evolutionary optimization and the improved SAC method, dynamically adapts to the response capabilities of different nodes during policy iteration, and introduces a gating mechanism and an attention mechanism to strengthen the control of the policy weights of key nodes. After actual deployment, the response speed to load mutations was shortened from the original average of 68 seconds to 24 seconds, and the load redistribution error of each node in the peak shaving scenario was reduced by 38%.

[0135] During the dispatching operation period, November 16th to November 30th, 2024 was selected as a continuous verification window. The system automatically collected node data and executed the trained model for dispatching control. During the peak period of 18:00 - 20:00 on November 22nd, when the load fluctuated violently, the system processed a total of 17 energy storage call events and 43 load migration operations, and no node over-limit phenomenon occurred. Compared with the original benchmark model (linear programming + expert rules), the energy-saving effect was improved by 7.9%, the user response rate was increased by about 12.6%, and the volatility of the daily average load balance index decreased by 23%. Further observing the event of a significant reduction in photovoltaic production during extreme weather at the beginning of December, this model can complete the policy adaptive fine-tuning and restore dispatching stability within 4 minutes, while the traditional model requires more than 10 minutes.

[0136] Table 1: Performance comparison table between the present invention and traditional methods on a typical dispatching day

[0137] Analysis of the data results in Table 1, "Performance Comparison Table of the Present Invention and Traditional Methods on a Typical Scheduling Day", clearly shows that the present invention has significant advantages over traditional scheduling systems in multiple key performance indicators, demonstrating its comprehensive improvement effects in aspects such as response efficiency, control accuracy, user participation, and system stability.

[0138] In terms of load scheduling response, the average policy execution response time of the system of the present invention is only 24 seconds, which is 64.7% shorter than the 68 seconds of the traditional scheduling system. The maximum single-node response delay is shortened from 112 seconds to 39 seconds, with a reduction amplitude as high as 65.2%, indicating that the present invention has an obvious improvement in the execution efficiency of policy triggering and action distribution. This high responsiveness is particularly crucial when dealing with sudden load fluctuations during peak grid hours and can effectively reduce the overload risk.

[0139] In terms of load balancing performance, the optimized policy of the present invention performs more precisely in peak shaving and valley filling control. The residual deviation after regulation is reduced from the original 11.4% to 6.2%, and the average value of the reallocation error is also reduced from 9.3 kW to 5.7 kW, with an error reduction of 38.7%. This shows that the low-level execution model in the present invention can more reasonably allocate load tasks and improve the system balance and regulation effect.

[0140] In terms of user behavior response, the present invention improves the enthusiasm of users on the user side to cooperate with scheduling through the user incentive mechanism and node attention strategy in the policy. The user participation rate is increased from 54.1% to 60.9%, an increase of 12.6%. At the same time, the average adoption time of the incentive signal is shortened from 36 seconds to 22 seconds, indicating that the timeliness of user response has also been optimized.

[0141] In terms of energy storage control effect, the policy of the present invention can more efficiently call energy storage resources. The daily energy storage utilization rate is increased from 73.5% to 82.4%, and the coverage ratio of energy storage for peak shaving load is increased from 48.2% to 56.8%. This means that the present invention can more fully mobilize energy storage units to participate in power optimization scheduling and reduce the dependence on external power supply to the grid.

[0142] In terms of energy economy, the electricity procurement cost during peak hours is reduced from 468,000 yuan to 421,000 yuan, with a cost reduction of 10%, indicating that the present invention better realizes the coordination of power consumption-side resources and cuts peak procurement electricity during peak hours, with significant economic benefits.

[0143] Finally, in terms of system stability, the present invention reduces the number of abnormal fluctuations during the switching process of the scheduling policy from 8 times per day to 2 times, with a 75% reduction in fluctuations. This shows that the present invention has more advantages in the continuity during the policy adjustment process and the stability of system control, and can effectively prevent local instability caused by frequent switching.

[0144] In summary, through the hierarchical reinforcement learning model with clear structure, the evolution-gradient fusion optimization mechanism, and the improved SAC policy structure, the present invention significantly improves the real-time performance, accuracy, robustness, and economy of the dynamic optimization process of the power system load, and has system-level performance superior to existing scheduling strategies.

[0145] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.

Claims

1. A dynamic optimization method for power system load based on reinforcement learning, characterized in that, It includes the following steps: S1. Collect data of the power system and construct a state space; S2. Based on the state space, construct a hierarchical reinforcement learning model, and divide the hierarchical reinforcement learning model into a high-level decision-making model and a low-level execution model; S3. Train the high-level decision-making model. The high-level decision-making model takes the high-level state in the state space as input and outputs a scheduling task target category instruction; S4. Train the low-level execution model. The low-level execution model takes the current node-level state and the scheduling task target category instruction output by the high-level decision-making model as input and outputs a control action sequence including the load migration amount of each node, energy storage charge and discharge instructions, and user response incentive coefficients; S5. When training the low-level execution model, introduce an evolutionary reinforcement learning mechanism, construct several policy individuals, calculate the reward value for each policy individual through simulated environment interaction, and perform selection, crossover, and mutation operations on the population composed of policy individuals using an improved genetic algorithm to obtain a new generation of policy individuals; S6. Combine the evolutionary reinforcement learning mechanism with the optimization method of policy gradient. While evolving and updating the policy individuals, further update the parameters of the relatively better-performing policy individuals using the optimization method based on policy gradient; S7. Deploy the hierarchical reinforcement learning model to the power system load scheduling module; S8. During the deployment and operation period, continuously collect scheduling feedback data and fine-tune the parameters of the deployed hierarchical reinforcement learning model; S9. Continuously apply the high-level decision-making model and the low-level execution model after migration and fine-tuning to the load scheduling process of the power system.

2. The dynamic optimization method for power system load based on reinforcement learning according to claim 1, characterized in that The specific content of S2 includes: S21. Divide the state space into a high-level state space and a low-level state space. The high-level state space includes the global load trend of the power system, the electricity price fluctuation range, and renewable energy output prediction data. The low-level state space includes the real-time load value of each node, energy storage state information, and controllable load regulation ability parameters; S22. Based on the high-level state space, construct a high-level decision-making model. The high-level decision-making model takes the high-level state as input and outputs a scheduling task target category instruction. The scheduling task target category instruction includes a peak shaving and valley filling scheduling instruction, an energy storage priority regulation instruction, or a demand response control instruction; S23. Based on the low-level state space and the scheduling task target category instruction output by the high-level decision-making model, construct a low-level execution model. The low-level execution model outputs a control action sequence including the load migration amount of each power grid node, energy storage charge and discharge control instructions, and user response incentive coefficients; S24. Establish a mapping relationship between the high-level decision-making model and the low-level execution model, and use the scheduling task target category instruction as the input condition of the low-level execution model to form a hierarchical and linked scheduling control structure.

3. The method for dynamically optimizing the load of a power system based on reinforcement learning according to claim 2, wherein, The specific content of S3 includes: S31. Obtain the high-level state space as the input data of the training sample; S32. Construct a high-level decision-making policy function , where represents the high-level decision-making policy function,[ is the input high-level state,[ is the output scheduling task target category instruction; S33. Define the immediate reward function of the high-level decision-making model , where represents the time step, and evaluates the performance of the low-level execution results guided by output by the high-level decision-making policy function in terms of load balancing, operating cost control, and scheduling volatility; S34. Use a reinforcement learning method based on policy gradients to train the high-level decision-making policy function and update the parameters of the high-level decision-making model by maximizing the expected value of the immediate reward function . When facing the input high-level state , output the optimal scheduling task target category instruction .

4. The load dynamic optimization method for a power system based on reinforcement learning according to claim 3, wherein The reinforcement learning method based on policy gradient in S34 is the SoftActor-Critic method based on maximum entropy policy optimization, which specifically includes: S341. Construct two high-level Q-value function networks and , taking the high-level state and the scheduling task target category instruction corresponding to the output as inputs, and respectively outputting Q-value estimates and . Use the smaller value among them to participate in policy optimization; S342. Introduce an objective function containing a policy entropy term as the optimization objective for training the high-level decision-making policy function to guide the parameter update process under a given high-level state ; S343. Alternately update the high-level decision-making policy function and the high-level Q-value function network , using the following structural improvement method: add a residual connection structure to the high-level decision-making policy function, adopt a shared state encoder in the Q-value function network, introduce a dynamic entropy coefficient adjustment mechanism based on the fluctuation amplitude of the target Q network, and replace the static or predefined entropy coefficient update method.

5. The dynamic optimization method for power system load based on reinforcement learning according to claim 4, wherein The specific content of S4 includes: S41. The low-level state input is defined as , including the current load value of each node in the power grid, the current state of charge of the energy storage device corresponding to the node, and the adjustable load capacity: S42. Combine the scheduling task target category instruction output by the high-level decision-making model with the low-level state input as the combined input of the low-level execution model to construct a low-level policy function; S43. Construct a low-level immediate reward function , evaluate the low-level state input and the scheduling task target category instruction output by the high-level decision-making model for the low-level control actions generated by the system performance; S44. Using the low-level immediate reward function as the optimization objective, train the low-level policy function , and under the condition of the given low-level state input and the scheduling task target category instruction output by the high-level decision-making model , output the optimal control action .

6. The dynamic optimization method for power system load based on reinforcement learning according to claim 5, characterized in that The specific content of S5 includes: S51. Initialize the low-level policy function and construct individuals of the low-level policy function , and each individual of the low-level policy function is a set of independent low-level policy network parameters, corresponding to the control action , is the control action of the th individual; S52. Record the cumulative reward value of each policy individual in consecutive training cycles, and calculate the comprehensive fitness value of the policy individual with historical smoothed weights; ​ S53. In the screening operation of the genetic algorithm, a selection method with an elite retention strategy is adopted to preferentially retain the first strategy individuals with the largest comprehensive fitness values, and at the same time, roulette selection is performed based on the ranking among the remaining individuals to generate a candidate parent set; S54. Improve the offspring crossover operation, judge the performance similarity between two parent policy individuals in the policy parameter distribution space. If their Euclidean distance or KL divergence is less than the set threshold, crossover is allowed; otherwise, this combination is discarded. S55. Perform a mutation operation on the selected set of offspring parameters, introduce a target guidance mechanism, and calculate the mutation probability based on the current target residual of the policy individual ; S56. Use the new generation of policy individual set after screening, crossover, and target-guided mutation as the current policy population, repeat S52 to S55, and iterate and update until convergence. Select policy individuals as the output policy structure of the low-level execution model according to the final comprehensive fitness value from high to low.

7. The load dynamic optimization method for a power system based on reinforcement learning according to claim 6, wherein The specific content of S6 includes: S61. After each generation of evolutionary operation, select the first policy individuals from the current policy population as the candidate subset for policy gradient optimization, where represents the th low-level policy individual, ; S62. For each policy individual , construct a set of double Q-value function networks with shared parameters , , with the low-level state input and the scheduling task target category instruction output by the high-level decision-making model and the corresponding control action as inputs; S63. Construct the low-level policy objective function as ; S64. Take as the optimization objective, and update the low-level policy individuals and the parameters of the corresponding Q-value function network respectively. The update adopts an asynchronous alternating training method, and adjusts the parameters of the target Q network based on the soft update rule; While evolving and updating the policy individuals, an optimization method based on policy gradient is used to further update the parameters of the better-performing policy individuals. Among them, the optimization method based on policy gradient is an improved SoftActor-Critic method, and the SoftActor-Critic method has the following structural improvements: a feature gating mechanism based on the clustering labels of the input state distribution is introduced into the lower-layer double Q-value function network to adjust the response intensity of the Q-value channels, a collaborative attention mechanism is introduced into the policy network, and a dynamic target KL divergence adjustment strategy is adopted in the entropy coefficient update process, which is adaptively adjusted according to the deviation amplitude of the policy distribution during the training process ; S66. Use the policy individual after being jointly trained by the genetic algorithm and the improved SoftActor-Critic method as the target output policy of the low-level execution model. ​ 8. The method for dynamically optimizing the load of a power system based on reinforcement learning according to claim 7, characterized in that, The specific structural improvement of the low-level SoftActor-Critic method in S65 includes: S651. Introduce a state distribution gating mechanism in the low-level double Q-value function network and cluster the low-level state input to obtain state category labels . Use the state category labels as input to embed into the gating module, and activate the corresponding channel control factors in the feature extraction layer of the low-level double Q-value function network to perform gating adjustment on the intermediate features to generate category-sensitive Q-value estimates; S652. Introduce a cross-node state collaborative attention mechanism in the low-level policy individuals to form the low-level state Take the state features of each node that makes up the low-level state as the input of the attention encoding respectively, construct the attention weight distribution based on the inter-node dependency relationship, and perform weighted aggregation on each node feature as the input of the policy distribution; S653. Entropy coefficient in the low-level policy network Introduce a dynamic adjustment mechanism and set the target policy distribution offset value . Calculate the current low-level policy individual in each round of training and the low-level policy individual in the previous round to obtain the KL divergence . Based on the deviation between and , adjust the entropy coefficient . The adjustment rule of the deviation-adjusted entropy coefficient is as follows: If , then increase ; If , then reduce ; If , then keep the current unchanged; Control the exploration degree and training stability of the policy in a dynamic manner.

Citation Information

Patent Citations

  • Electric power system source-load look-ahead scheduling method and device based on deep reinforcement learning

    CN113902176A

  • Double-layer agent decision control method based on reinforcement learning

    CN117555229A

  • Satellite exploration control system and method based on deep reinforcement learning

    CN119239995A

  • Load management and control method and system based on reinforcement learning

    CN119674994A

  • Method for robotic multi-peg-in-hole assembly based on hierarchical reinforcement learning and distributed learning and system thereof

    US20240361732A1

Cited By

  • Energy storage income maximization method based on strategy gradient algorithm

    CN120896202A

  • Network source collaborative primary frequency modulation optimization method and system based on reinforcement learning

    CN121192745A

  • Water affair system operation data model optimization scheduling method based on reinforcement learning

    CN121212713A

  • Water system operation data model optimization scheduling method based on reinforcement learning

    CN121212713B

  • Virtual reality photography composition optimization method and system based on hierarchical reinforcement learning

    CN121259674A