Power grid multi-stage planning intelligent auxiliary decision-making method and system based on reinforcement learning fusion

CN122844082APending Publication Date: 2026-09-29南方电网能源发展研究院有限责任公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610985605.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-03
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0006]为克服现有技术的不足,本发明旨在提供一种融合强化学习的电网多阶段规划智能辅助决策方法及系统,其能够深度融合精准环境建模、高效策略学习与内在探索机制,以自主、适应性地解决复杂环境下电网多阶段协同规划问题

Benefits of technology

[0028]本发明的有益效果在于:1)通过引入条件生成对抗网络构建高精度环境动力学模型,显著提升了样本利用效率和规划仿真的可信度,降低了对真实系统试错的依赖与风险;2)创新性地设计了策略与内在奖励交替优化的双层梯度机制,有效解决了稀疏奖励下的探索难题,引导智能体学习更优的协同策略;3)将平均场理论与基于模型的强化学习深度结合,构建了概率测度空间上的决策框架,简化了大规模多资源协同规划的复杂性;4)形成了完整的虚拟-现实闭环迭代学习范式,使系统具备持续适应环境变化与进化的能力,最终输出的规划方案在自主性、适应性与经济协同性方面均得到显著提升。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122844082A_ABST
    Figure CN122844082A_ABST
Patent Text Reader

Abstract

This invention discloses an intelligent auxiliary decision-making method and system for multi-stage power grid planning that integrates reinforcement learning. The method models multi-stage power grid planning as a Markov decision process in a probability measure space, constructing a decision framework through three neural networks: initializing the policy, environment, and intrinsic reward. It then uses historical data to train a conditional generative adversarial network to construct an initial environmental dynamics model. At each planning stage, forward simulation is performed using this model and the policy network to generate a virtual planning trajectory. Based on this trajectory, an alternating optimization two-layer gradient mechanism is used to collaboratively update the parameters of the policy and intrinsic reward networks. The optimized policy is applied to actual power grid interactions to collect new data, which is then used to iteratively update the environment model. The output is a multi-stage planning scheme that satisfies multi-priority load demands, adapts to environmental uncertainties, and achieves coordinated optimization of power generation and energy storage. This invention significantly improves the autonomy, adaptability, and collaborative optimization capabilities of planning decisions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the application of artificial intelligence in the field of power system planning, specifically to an intelligent auxiliary decision-making method and system for multi-stage power grid planning that integrates reinforcement learning. Background Technology

[0002] With the accelerated transformation of the energy structure, new power systems characterized by a high proportion of renewable energy, diversified loads, and energy storage systems are constantly developing. Power grid planning is becoming increasingly complex, requiring long-term, multi-stage optimization decisions under multiple constraints, including meeting multi-priority load demands, addressing environmental uncertainties (such as the intermittency of renewable energy output and load fluctuations), and achieving economic synergy between power generation and energy storage resources. Traditional planning methods typically rely on precise mathematical models and deterministic scenario assumptions, making it difficult to dynamically adapt to complex and changing environments, and they face the curse of dimensionality when dealing with high-dimensional, continuous decision spaces.

[0003] Reinforcement learning, as an artificial intelligence method that learns optimal decisions through trial and error and interaction with the environment, offers a new approach to addressing the aforementioned challenges. Existing research attempts to introduce reinforcement learning into power grid planning, but significant shortcomings remain: Firstly, classic model-free reinforcement learning methods suffer from low sample efficiency, making extensive trial-and-error in high-risk real-world systems like power grids costly and risky. Secondly, while model-based reinforcement learning methods can improve sample efficiency by establishing environmental models, existing models often struggle to accurately characterize the highly nonlinear and uncertain nature of the power grid environment, leading to accumulated model errors and impacting the reliability of planning schemes. Furthermore, traditional reward designs are ineffective in guiding agents to explore sparse reward problems like power grid planning, easily leading to local optima.

[0004] In particular, the complexity of collaborative decision-making increases dramatically in multi-agent or large-scale system scenarios. While mean-field theory can simplify multi-agent interactions, deeply integrating it with model-based reinforcement learning and designing effective exploration mechanisms to address uncertainty and sparse reward problems in planning remains a pressing technical challenge.

[0005] Therefore, there is an urgent need for an intelligent auxiliary decision-making method that can deeply integrate accurate environmental modeling, efficient strategy learning and intelligent exploration mechanisms, and adapt to the characteristics of multi-stage power grid planning. Summary of the Invention

[0006] To overcome the shortcomings of existing technologies, this invention aims to provide an intelligent auxiliary decision-making method and system for multi-stage planning of power grids that integrates reinforcement learning. This method can deeply integrate accurate environment modeling, efficient policy learning and intrinsic exploration mechanisms to autonomously and adaptively solve the multi-stage collaborative planning problem of power grids in complex environments.

[0007] In a first aspect, embodiments of this application provide a smart auxiliary decision-making method for multi-stage planning of power grids that integrates reinforcement learning, the method comprising:

[0008] S1. Based on the distribution of power grid resources, load priority grouping and environmental state parameter set, the multi-stage planning problem of power grid is modeled as a Markov decision process on the probability measure space, and the policy network, environmental model network and intrinsic reward network are initialized.

[0009] S2. Based on historical power grid operation data, an initial environmental dynamics model is constructed by training a conditional generative adversarial network. This model takes load fluctuation rate, renewable energy output characteristics and fault probability as input conditions, and predicts the future power grid state based on the current state and decision actions.

[0010] S3. In each planning stage, forward simulation is performed using the current environmental dynamics model and policy network to generate virtual planning trajectory samples containing state transitions.

[0011] S4. Based on the virtual planning trajectory sample, update the parameters of the policy network and the intrinsic reward network; wherein, first update the parameters of the policy network according to the composite gradient of the external reward and the intrinsic reward, and then update the parameters of the intrinsic reward network according to the gradient of the external reward.

[0012] S5. Apply the updated policy network to the actual power grid environment to perform resource allocation decisions and collect the generated state and decision interaction data.

[0013] S6. Add the new data collected in S5 to the training set to update the environmental dynamics model; after the model update is completed, return to S3 to start a new round of iteration; repeat S3 to S6 until the system reaches the preset convergence criterion;

[0014] S7. Output a multi-stage power grid planning scheme that meets the needs of multiple priority loads, adapts to environmental uncertainties, and achieves coordinated optimization of power generation and energy storage.

[0015] In a second aspect, embodiments of this application provide an intelligent auxiliary decision-making system for multi-stage power grid planning that integrates reinforcement learning, applied to the intelligent auxiliary decision-making method for multi-stage power grid planning that integrates reinforcement learning as described in the first aspect, the system comprising:

[0016] The modeling and initialization module is used to model the multi-stage planning problem of the power grid as a Markov decision process in the probability measure space based on the distribution of power grid resources, load priority grouping and environmental state parameter set, and to initialize the policy network, environmental model network and intrinsic reward network.

[0017] The initial environment model building module is used to build an initial environment dynamics model based on historical power grid operation data and by training a conditional generative adversarial network. The model takes load fluctuation rate, renewable energy output characteristics and fault probability as conditional inputs and predicts the future power grid state based on the current state and decision actions.

[0018] The trajectory generation and simulation module is used to perform forward simulation using the current environmental dynamics model and policy network at each planning stage, generating virtual planning trajectory samples that include state transitions.

[0019] The strategy and reward optimization module is used to update the parameters of the strategy network and the intrinsic reward network based on the virtual planning trajectory sample; wherein, the strategy network parameters are first updated according to the composite gradient of the external reward and the intrinsic reward, and then the intrinsic reward network parameters are updated according to the gradient of the external reward.

[0020] The actual interaction and data collection module is used to apply the updated policy network to the actual environment of the power grid, execute resource allocation decisions, and collect the generated status and decision interaction data.

[0021] The iterative control and model update module is used to add the collected new data to the training set to update the environmental dynamics model, and after the model update is completed, it controls the return to the trajectory generation and simulation module to start a new round of iteration until the system reaches the preset convergence criterion.

[0022] The scheme output module is used to output a multi-stage power grid planning scheme that meets the multi-priority load demand, adapts to environmental uncertainties, and achieves coordinated optimization of power generation and energy storage when the system reaches the convergence criterion.

[0023] Thirdly, embodiments of this application provide an electronic device, including:

[0024] processor;

[0025] Memory used to store processor-executable instructions;

[0026] The processor is configured to implement the intelligent auxiliary decision-making method for multi-stage planning of power grids that incorporates reinforcement learning as described in the first aspect when executing the instructions.

[0027] Fourthly, embodiments of this application provide a computer-readable storage medium storing a program that instructs a device to execute the intelligent auxiliary decision-making method for multi-stage planning of power grids incorporating reinforcement learning as described in the first aspect.

[0028] The beneficial effects of this invention are as follows: 1) By introducing conditional generative adversarial networks to construct a high-precision environmental dynamics model, the efficiency of sample utilization and the credibility of planning simulation are significantly improved, and the dependence on and risk of trial and error in real systems are reduced; 2) An innovative two-layer gradient mechanism for alternating optimization of policy and intrinsic reward is designed, which effectively solves the exploration problem under sparse rewards and guides the agent to learn better collaborative strategies; 3) The mean field theory is deeply combined with model-based reinforcement learning to construct a decision-making framework in the probability measure space, which simplifies the complexity of large-scale multi-resource collaborative planning; 4) A complete virtual-reality closed-loop iterative learning paradigm is formed, which enables the system to continuously adapt to environmental changes and evolution, and the final output planning scheme is significantly improved in terms of autonomy, adaptability and economic synergy. Attached Figure Description

[0029] Figure 1 This is a schematic diagram of a method for intelligent auxiliary decision-making in multi-stage planning of power grids that incorporates reinforcement learning, provided as an embodiment of this application.

[0030] Figure 2 The architecture diagram of the intelligent auxiliary decision-making system for multi-stage planning of power grids that incorporates reinforcement learning is provided for this application.

[0031] Figure 3 A schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0032] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them.

[0033] It should be noted that in the embodiments of this application, "at least one" refers to one or more, and "more than one" refers to two or more. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the specification of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application.

[0034] Based on the embodiments described in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0035] Example 1

[0036] Figure 1 This is a schematic flowchart illustrating a multi-stage intelligent auxiliary decision-making method for power grid planning that incorporates reinforcement learning, provided as an embodiment of this application. Figure 1As shown, a smart auxiliary decision-making method for multi-stage power grid planning integrating reinforcement learning includes:

[0037] S1. Based on the distribution of power grid resources, load priority grouping, and environmental state parameter set, the multi-stage planning problem of the power grid is modeled as a Markov decision process in the probability measure space, and the policy network, environment model network, and intrinsic reward network are initialized. This step defines the task and builds the algorithm skeleton through framework construction and initialization. The complex multi-stage planning problem of the power grid is transformed into a mathematical model (Markov decision process) that can be handled by reinforcement learning, and three core neural networks (policy, environment model, and intrinsic reward) are prepared for subsequent learning, completing the initialization.

[0038] Specifically, in this embodiment, the multi-stage power grid planning problem in step S1 is modeled as a Markov decision process in the probability measure space, which is constructed in the following way:

[0039] (1) The state space is defined as a joint state vector containing generator output levels, energy storage charging and discharging status, real-time demand of each priority load group, load fluctuation statistics, deviation between predicted and actual renewable energy output, and key node failure probability indicators. The state space is the set of all key information that an agent (decision-maker) can observe at each planning moment, reflecting the current operating status of the system. It provides information input for decision-making. It determines what the agent sees and is its window to the environment. A comprehensive and accurate state space is a prerequisite for making correct decisions. For example, a virtual power grid dispatch center screen displays: Generator A outputs 500 MW, generator B outputs 300 MW; the current energy storage station has 60% power and is charging at 50 MW; high-priority load demand is 800 MW, medium-priority demand is 400 MW; the wind farm's current actual output is 20% lower than predicted; the failure risk assessment of a key transmission line is 0.5% (low risk). The set of all these real-time data constitutes the state vector at the current moment.

[0040] (2) The action space is defined as a multi-dimensional continuous action vector containing the power allocation coefficients of each priority load group, generator start-up and shutdown and output adjustment commands, energy storage charging and discharging power commands, and reserve capacity allocation schemes. The action space is a multi-dimensional continuous action vector, defining the set of all control commands that the agent can execute. The action space is the set of all possible operations or decisions that the agent can execute at each planning moment. It defines the output form and feasible range of the decisions. It tells the agent what it can do. Because grid control is a fine-grained continuous adjustment, the action space is designed as a multi-dimensional continuous vector, allowing for precise output commands such as increasing the power of generator A by 3.5 MW. Based on the above state, the agent needs to issue commands. These commands may include: proportionally allocating available power, for example, 55% to high-priority loads, 30% to medium-priority loads, and 15% to low-priority loads; commanding generator A to increase its output by 10 MW and generator B to decrease its output by 5 MW; commanding the energy storage station to switch to discharging at 30 MW; and reserving a certain amount of reserve capacity for each generator to cope with emergencies. This set of commands constitutes an action vector.

[0041] (3) The state transition function is learned through the environmental model network and is used to approximate the probability distribution of transitioning to the next state given the current state and decision action. The state transition function is a physical law simulator of the power grid, describing how the system will change to the next state after performing a certain action in the current state. It defines the dynamic model of the environment. It answers the question of what will happen next if I do this now. In this invention, this complex law, which is difficult to describe with precise mathematical equations, is not defined manually, but is learned and approximated by the environmental model network (a conditional generative adversarial network) from historical data. For example, when the agent issues the above action command, the state transition function (i.e., the learned environmental model) will predict: due to the command to discharge energy storage, its power will drop to 58% at the next moment; due to changes in load demand and power generation output, the system frequency may fluctuate slightly; the randomness of wind power may cause changes in the output deviation of renewable energy. This prediction result is probabilistic and represents the most likely future scenario.

[0042] (4) The reward function adopts a hierarchical structure including basic external rewards and intrinsic exploration rewards. The basic external rewards reflect the generation cost and penalties for violating service quality constraints, while the intrinsic exploration rewards are dynamically generated based on state prediction errors. Specifically, the reward function refers to the scoring criteria for the planning task, used to evaluate the quality of each decision action in real time and guide its strategy to improve in the right direction. This invention adopts a unique hierarchical structure: Basic external rewards: correspond to the final business objectives, such as low generation costs and never cutting off power to high-priority loads. Good performance is rewarded with negative points (cost) or fewer deductions, while violations of constraints are severely penalized. Intrinsic exploration rewards: correspond to curiosity during the learning process. Generation based on prediction errors of the environment model: If the agent explores a region and the environment model's prediction of its next state is very inaccurate (large error), a reward is given to encourage the agent to learn about this unknown region and thus discover potentially better strategies. For example, after performing the above actions, the system calculates the rewards: the generation cost was 1000 units (deduction); all loads were satisfied, with no penalty (no deduction); however, because the decision entered a region where the model's prediction was inaccurate, an exploration reward of +5 was obtained. The total reward is -1000 + 5 = -995.

[0043] (5) The value function adopts an approximation based on mean-field theory. It measures the long-term comprehensive value of different decision-making strategies by evaluating the expected cumulative discounted reward in the probability measure space. The value function is defined as a long-term investment evaluation report of a decision-making strategy. It does not only look at the reward of the current step, but evaluates the sum of all future rewards (considering discounts) that can be obtained by following a certain strategy from the current state. It is the ultimate optimization goal of reinforcement learning. The purpose of strategy optimization is to find a strategy that maximizes this long-term total return. This invention adopts an approximation based on mean-field theory. This is a mathematical tool for dealing with large-scale systems such as power grids that contain a large number of interacting individuals (numerous load points, power generation units). It simplifies the joint state of all individuals, which is difficult to handle, to the distribution characteristics of the state, thus making it feasible to calculate the long-term value. For example, strategy A may have a slightly higher cost in the first step (-1020), but it creates favorable conditions for subsequent steps, and the long-term total cost is expected to be -100,000. Strategy B has a low cost in the first step (-990), but it makes subsequent adjustments difficult, and the long-term total cost is expected to be -120,000. Therefore, strategy A has higher value and is the better choice. The value function is used to make this kind of long-term judgment.

[0044] Furthermore, an approximate value function based on mean-field theory is employed to measure the long-term comprehensive value of different decision-making strategies by evaluating the expected cumulative discount reward in the probability measure space; the value function is specifically expressed as:

[0045]

[0046] in, Indicates time State distribution measure, Representation strategy In distribution The expected value below Represents the composite reward function. As a discount factor, Represents the state transition probability. Indicates the state distribution Next state Seeking expectations, Indicates the state transition probability Next to the next state Seeking expectations.

[0047] These five components collectively translate the real-world power grid planning problem into a machine-understandable and optimizeable language: the state is the machine's eyes, actions are its hands and feet, state transitions are its understanding of the world's operating rules, rewards are the immediate feedback guiding its learning, and value is the ultimate goal it pursues. This complete mathematical model forms the foundation for the operation and optimization of all subsequent intelligent algorithms (policy networks, environment model networks, and intrinsic reward networks).

[0048] Furthermore, the policy network adopts a deep deterministic policy gradient framework, in which the Actor network outputs continuous resource allocation decisions based on the power grid state, and the Critic network evaluates the long-term value of different decisions in a given environment.

[0049] The Actor network, acting as the executor of the policy network, has the core function of directly generating specific resource allocation decisions based on the current power grid state. This is a deep neural network that receives a complete state vector containing information on generation, energy storage, and load, and outputs a set of continuous decision instructions. The specific implementation steps are as follows: 1. State Feature Extraction: The network first processes the input power grid state vector, extracting high-order features through a multi-layer fully connected network. For example, identifying the urgency of the current load pattern, the available capacity of energy storage devices, and the output trend of renewable energy. 2. Decision Instruction Generation: Based on the extracted features, the network outputs a multi-dimensional continuous vector, which is divided into several key parts: Load Allocation Instruction: Outputs power allocation ratios (e.g., [0.6, 0.3, 0.1]) for high, medium, and low priority load groups; Generation Scheduling Instruction: Outputs power adjustment amounts for each generator unit (positive values ​​increase generation, negative values ​​decrease generation); Energy Storage Control Instruction: Outputs charging and discharging power values ​​for each energy storage device; Reserve Allocation Instruction: Reserves capacity reserved for key units; 3. Output Normalization: Through activation functions (e.g., tanh) and scaling operations, ensures that all output instructions are within the range allowed by the physical devices.

[0050] The Critic network, as an evaluator in a policy network, primarily evaluates the long-term value of a specific state-action combination. It receives the current state and the action proposed by the Actor, outputting a single value representing the long-term reward of this decision. The specific implementation steps can be: 1. State-Action Feature Fusion: Concatenating the state vector and action vector into a new feature vector. 2. Deep Value Evaluation: Learning the deep value features of the state-action combination through multi-layered cascaded neural networks. 3. Long-Term Return Prediction: Outputting a numerical value that comprehensively considers: the immediate effects of the current decision (e.g., cost, reliability), the potential impact on the future state of the system, and the cascading effects across multiple planning stages.

[0051] The environment model network adopts a conditional generative adversarial network architecture. Its generator takes the current state, decision action, and environmental parameters as input conditions and introduces random noise to predict the power grid state in the next stage. The discriminator judges the authenticity of the state through adversarial training, thereby learning high-precision environmental transition dynamics. The intrinsic reward network is designed as an exploration incentive mechanism based on the state prediction error. Its output intrinsic reward signal is positively correlated with the prediction error of the environment model, which is used to drive the policy network to explore effectively in a sparse reward environment.

[0052] The generator is the core predictive component of the environmental model, learning the environmental dynamics of the power grid. Unlike traditional models, it takes the current decision as input and predicts the next state the power grid will enter after executing that decision. The specific implementation steps are: 1. Condition information construction: Concatenate the current power grid state, the planned decision action, and environmental uncertainty parameters (such as load volatility and renewable energy output characteristics) into a condition vector. 2. Random noise injection: Introduce a random noise vector to ensure the generator can learn the probability distribution of state transitions rather than deterministic mappings. 3. Next state prediction: Based on the conditions and noise, the generator outputs a complete prediction of the power grid state at the next moment, including: the new output level of each unit, the new state of charge of energy storage devices, possible voltage and frequency changes, and the operating status of key equipment.

[0053] The discriminator acts as the quality inspector of the environment model, responsible for determining whether a state prediction is realistic and credible. It improves alongside the generator through adversarial training. Specific implementation steps include: 1. Real vs. Fake Sample Comparison: The discriminator receives two types of input: real power grid operation data (real samples) and the state predicted by the generator (fake samples). 2. Deep Feature Recognition: Deep statistical features of the state vector are identified through convolutional layers or self-attention mechanisms. 3. Real vs. Fake Probability Output: A probability value between 0 and 1 is output, representing the degree to which the input state resembles the real state.

[0054] For example, when an agent tries a novel power generation-storage combination strategy, the environmental model's prediction of its consequences may be inaccurate (with large errors) due to the lack of similar cases in historical data. In this case, the intrinsic reward network will give a high reward for this attempt, encouraging the agent to continue exploring this new path, even if the current external reward (cost) may not be ideal.

[0055] S2. Based on historical power grid operation data, an initial environmental dynamics model is constructed by training a conditional generative adversarial network (GAN). This model takes load fluctuation rate, renewable energy output characteristics, and fault probability as input conditions, and predicts the future power grid state based on the current state and hypothetical decision actions. This step establishes a virtual cognition or digital twin of the power grid environment. Using historical data, a high-precision GAN model is trained, enabling it to predict the future power grid state based on the current state and hypothetical decision actions. This lays the foundation for safe and efficient trial and error in the virtual space.

[0056] Specifically, in this embodiment, step S2, which involves constructing an initial environmental dynamics model by training a conditional generative adversarial network, specifically includes:

[0057] (1) A conditional generative adversarial network (GAN) framework is constructed. The generator takes the current grid state, resource allocation decisions, load volatility, renewable energy output characteristics, and fault probability as joint conditional inputs, and introduces random noise to output a prediction of the grid state in the next stage. The discriminator takes the same conditional inputs and the actual or generated next state as inputs, and outputs the probability that the state is the actual state. The generator is designed as a deep neural network, and its core function is to predict the future grid state based on given conditional information. The generator network processes conditional and noise information through multi-layer neural networks, ultimately outputting a prediction vector with the same dimension as the input state vector, representing the most likely next grid state after executing a given decision. Through learning, the discriminator gradually becomes able to accurately distinguish the subtle differences between the generator's predicted state and the actual observed state. When the discriminator struggles to distinguish between the two, it indicates that the generator's prediction quality is already high.

[0058] Construct a conditional generative adversarial network framework, whose generator is based on the current power grid state. Resource allocation decision Load fluctuation rate, renewable energy output characteristics, and fault probability are used as joint input conditions, and random noise is introduced. The output predicts the state of the power grid in the next stage, specifically expressed as follows: ,in, Represents a vector of parameters indicating environmental uncertainty. Represents a generator network;

[0059] Its discriminator takes the same input conditions and the true or generated next state as input, and outputs the probability that the state is the true state. This represents random noise.

[0060] (2) Phased adversarial training: First, pre-training is performed based on the historical power grid operation data to allow the generator to initially learn the state transition rules. Subsequently, during the closed-loop iteration process from S3 to S6, new data collected in S5 is continuously added for online fine-tuning and adversarial optimization. Phase 1: Historical data pre-training. In the initial stage of system deployment, offline training is performed using historical power grid operation data. The pre-training phase is completed when the generator's predictions reach a predetermined accuracy threshold on the historical data test set. Phase 2: Online iterative fine-tuning. After the system is put into operation, continuous online learning and optimization are performed: newly collected real-time interactive data (real state transition records) are added to the training set. Each time a certain amount of new data is collected, model fine-tuning is initiated. During fine-tuning, new and old data are mixed in a certain proportion, and the learning rate is appropriately reduced to avoid forgetting historical learning results. The new data reflects the latest operating characteristics of the power grid (such as new generating units and changes in load patterns). Through continuous fine-tuning, the environmental model can adapt to these changes and maintain prediction accuracy.

[0061] (3) Model Application: After training convergence, the stable generator is used as the environmental dynamics model to simulate and predict the future multi-stage power grid state evolution trajectory based on the current state and candidate decisions during the planning phase. Specifically, after training convergence, the generator is extracted as an independent environmental dynamics model: Model Solidification: The network structure and weight parameters of the generator are saved as an independently runnable model file. Prediction Interface Encapsulation: A standardized prediction interface is provided, with the current state, candidate decisions, and environmental parameters as inputs, and the predicted state for the next stage as output.

[0062] The real-time implementation of multi-stage trajectory simulation is as follows: The application process in the planning phase is as follows: 1. Single-step prediction call: Given the current state and candidate decisions, the environmental model performs a single-step prediction, outputting the possible next state. 2. Recursive rolling prediction: The predicted next state is used as the new current state. Combined with the subsequent decisions generated by the policy network for this new state, the environmental model is called again for the next prediction. This process is repeated to generate a complete N-stage predicted trajectory. 3. Uncertainty handling: For critical planning, multiple predictions are performed, each injected with different random noise, resulting in a set of possible future trajectories for risk assessment. 4. Application of prediction results: The generated predicted trajectories are used to: evaluate the long-term effects of candidate decisions, identify potential operational risks, and provide virtual training samples for policy optimization. After each actual decision is executed, the real results are compared with the predictions to monitor model accuracy. When the prediction error continuously exceeds a threshold, model retraining is triggered. Simultaneously, multiple versions of the model are retained, allowing for regression to a historically stable version when predictions become unstable. Through this implementation, the environmental dynamics model not only provides high-quality predictions but also continuously improves itself as the system operates, adapting to the dynamic changes in the power grid.

[0063] S3. In each planning phase, forward simulation is performed using the current environmental dynamics model and policy network to generate virtual planning trajectory samples that include state transitions. This step involves safe exploration and planning deduction in a virtual sandbox through virtual trial and error and sample generation. Combining the pre-trained environmental model (S2) and the current policy (initialized in S1 or updated in S4), a large number of possible planning trajectories and results (virtual samples) are simulated and generated. This avoids high-risk, high-cost trial and error directly in the real power grid.

[0064] Specifically, in this embodiment, generating virtual planning trajectory samples in step S3 includes:

[0065] (1) Based on the current actual state and state distribution of the power grid, set the initial conditions for the simulation. Obtain accurate operating data at the current moment from the actual power grid monitoring system, including the output of all generators, the state of charge of energy storage and the demand of each priority load, and construct a complete initial state vector; at the same time, based on the load type distribution and historical statistical data, calculate the initial state probability distribution as the starting point of the simulation.

[0066] (2) Perform multi-step forward inference: Repeat the following operations until the preset simulation step size or termination condition is reached: a. Input the current simulation state and distribution into the policy network to obtain the decision action; b. Input the current state, decision action, and environmental parameters into the environmental dynamics model to obtain the prediction of the next simulation state; c. Calculate the immediate reward according to the reward function; d. Update the state distribution based on the predicted next state; e. Record the current state, decision action, reward, next state, and updated distribution as a transition sample, and update the current state. Establish an iterative loop in the simulation environment: In each step, first call the policy network to generate a decision action, then input the action into the environmental dynamics model to predict the next state, then calculate the benefit or cost of the step according to the preset reward rule, then update the probability distribution of the system state based on the prediction result, and finally package and store all the information of this transition as a training sample.

[0067] (3) Construct a complete virtual planning trajectory from the ordered transition sample sequence generated in the loop. Connect all the transition samples generated in the deduction loop in chronological order to form a complete decision sequence; timestamp and verify the integrity of this sequence to ensure that each decision node has corresponding state, action, reward and next state information.

[0068] (4) Multiple independent trajectories are generated in parallel from different initial states or by introducing different random noises to construct a batch sample set for training. Parallel computing technology is used to start multiple simulation instances simultaneously, each instance using slightly different initial conditions or injecting different random noises; the independent trajectories generated by all instances are aggregated into a large-capacity training dataset, and then standardized and quality-screened to remove abnormal trajectories.

[0069] S4. Based on the virtual planning trajectory samples, update the parameters of the policy network and the intrinsic reward network. First, update the policy network parameters based on the composite gradient of external and intrinsic rewards, then update the intrinsic reward network parameters based on the gradient of external rewards. This step optimizes the learning and adjustment of the exploration direction from virtual experience through policy and exploration mechanisms. Based on the virtual samples (S3), a two-layer gradient mechanism is used to optimize core decision-making capabilities: Optimize the policy network: learn how to make better resource allocation decisions. Optimize the intrinsic reward network: dynamically adjust curiosity or exploration motivation to make exploration more efficiently serve the final task objective. This is key to solving the sparse reward problem.

[0070] Specifically, in this embodiment, the parameter updates for the policy network and the intrinsic reward network in step S4 are performed in the following manner:

[0071] (1) Calculation of Composite Reward Gradient: For the virtual planning trajectory, the accumulated composite reward is calculated, which is the weighted sum of the external reward and the intrinsic reward. Based on the expected value of the composite reward, the update gradient of the policy network parameters is calculated using the policy gradient algorithm. For each virtual trajectory, the sum of the external reward and the intrinsic reward at each step is calculated, and then the reward sequence of the entire trajectory is accumulated with time discount to obtain the total composite reward. Based on this total reward, the advantage function estimation method in the policy gradient algorithm is used to backpropagate and calculate the gradient direction of each parameter of the policy network.

[0072] (2) Policy network parameter update: Using the gradient, the policy network parameters are optimized using the gradient ascent method. The calculated gradient is input into the optimizer (e.g., Adam), and the policy network weights are slightly adjusted along the gradient ascent direction; the adjustment range is controlled by the preset learning rate, and gradient pruning is used to prevent the update step size from being too large, which would lead to training instability.

[0073] (3) Setting the intrinsic reward objective: The learning objective of the intrinsic reward network parameters is set to maximize the expected cumulative external reward obtained from the same batch of virtual trajectories. The current policy network is fixed, and the same batch of virtual trajectories is re-evaluated, but this time only the cumulative value of pure external reward is calculated; the maximization of this cumulative external reward is set as the direct optimization objective of the intrinsic reward network in this round of update.

[0074] (4) Intrinsic Reward Gradient Calculation and Update: Using the chain rule, the gradient of the expected cumulative external reward with respect to the intrinsic reward network parameters is calculated, and the intrinsic reward network parameters are updated accordingly. Using an automatic differentiation tool, the gradient chain of the cumulative external reward value with respect to the intrinsic reward network parameters is calculated, i.e., how the external reward is affected by the intrinsic reward signal through the policy network decision-making process. The parameters of the intrinsic reward network are updated along this gradient direction so that the generated intrinsic reward can better guide the policy to improve the external reward.

[0075] Furthermore, the parameter update process employs an alternating optimization two-layer gradient mechanism:

[0076] (1) Outer Layer Optimization: Based on the current intrinsic reward network parameters, calculate the composite reward and update the policy network parameters with the goal of maximizing its expected value. The output of the current intrinsic reward network is added to the external task reward to form a composite reward signal. The policy gradient algorithm is then used to update the parameter weights of the policy network along the direction that increases this composite reward. Specifically, this is achieved through the following gradient equation:

[0077] ,

[0078] in, Represents the policy network parameter vector gradient operator, This represents the composite reward objective function. This represents the composite advantage function. Indicates the trajectory The mathematical expectation, Represents a virtual planned trajectory. These represent the time index and the total number of times, respectively.

[0079] (2) Inner Layer Optimization: With the updated policy network parameters fixed, and maximizing the expected cumulative external reward as the independent objective, the intrinsic reward network parameters are calculated and updated. The newly updated policy network is temporarily fixed, using only external task rewards as feedback signals. The parameters of the intrinsic reward network are calculated and updated through backpropagation, making its output exploration signals more conducive to the policy network obtaining higher external rewards in the future. This is specifically achieved through the following gradient equation:

[0080] ,

[0081] in, Represents the policy network parameter vector gradient operator, This represents the objective function with purely external rewards. This represents the external dominance function.

[0082] (3) Alternately execute steps (1) and (2) to form an optimization loop that alternates between inner and outer layers, so that the internal reward signal can be adaptively adjusted to drive the policy network to explore in the direction of improving the final task performance. Repeat the outer and inner layer optimization steps in sequence, with each round of outer layer optimization followed by a round of inner layer optimization, and so on until the policy performance is stable; save the intermediate model state at each switch between inner and outer layers to ensure that the optimization process is continuous and converges stably.

[0083] S5. Apply the updated policy network to the actual power grid environment to execute resource allocation decisions and collect the resulting state and decision interaction data. This step applies the learned policy to reality through real-world application and data collection, and gathers real feedback. The optimized policy network is deployed to the real power grid to execute decisions, while simultaneously acting as a sensor to collect data on the real environment's response to new decisions. This serves as a bridge connecting virtual learning and the real world.

[0084] Specifically, in this embodiment, step S5 applies the policy network to real-world environmental interactions and collects data, specifically including:

[0085] (1) Online deployment: The updated strategy network is deployed to the actual power grid control system, and decisions are generated and executed based on real-time status. The trained strategy network is compiled into a real-time control module and deployed to the power grid energy management system. Real-time telemetry data is received through a standard communication interface, and resource allocation instructions are automatically generated and sent to the generation, energy storage and load control units for execution.

[0086] (2) Synchronous Recording: After each decision is executed, the actual state at the time of the decision, the decision action output by the strategy network, the immediate external reward from the environmental feedback, and the observed actual state of the next stage are recorded. Within each control cycle, the system automatically collects and timestamps and stores four types of data: the real-time operating status of the power grid before execution, the decision instructions output by the strategy network, the real-time operating cost or reliability indicators fed back by the power grid after execution, and the new power grid state after execution.

[0087] (3) Data Validation: Validate the collected interactive data and filter out invalid data. Design data validation rules to automatically check the integrity and physical rationality of the collected data, and remove data records containing outliers, violating physical constraints, or exceeding communication timeouts to ensure the quality of the training dataset.

[0088] (4) Dataset expansion: Add valid data to the training set used for model and policy updates. Merge validated valid data samples with historical datasets to build a dynamically expanding training sample pool and establish a data version management mechanism to ensure that newly added data can be used in an orderly manner for subsequent model updates and policy iteration training.

[0089] S6. Add the new data collected in S5 to the training set to update the environmental dynamics model. After the model update is complete, return to S3 to start a new round of iteration. Repeat S3 to S6 until the system reaches the preset convergence criterion. This step achieves a closed loop of continuous learning from real-world feedback through closed-loop iteration and model evolution. Real-world interaction data (S5) is fed back to the environmental model (S2) for updating, making its cognition closer to reality. Then, based on the updated and more accurate model, a new round of virtual simulation (S3) and policy optimization (S4) begins. This step is a key loop that enables the method to possess continuous adaptability and evolutionary capabilities.

[0090] Specifically, in this embodiment, the preset convergence criterion includes the following quantifiable criteria:

[0091] (1) Strategy Convergence Criterion: The resource allocation decision output by the strategy network remains stable over multiple consecutive planning periods, wherein the rate of change of the power ratio allocated to each priority load group is lower than the first preset threshold, and the fluctuation range of the dispatch instructions for power generation and energy storage equipment is lower than the second preset threshold. At the end of each planning period, the difference between the key decision parameters output by the current strategy and the historical strategies of the most recent periods is automatically calculated. If the rate of change of the load allocation ratio and the fluctuation range of the equipment dispatch instructions are lower than their respective preset thresholds for multiple consecutive periods, the strategy is determined to have converged. The resource allocation decision output by the strategy network remains stable over M consecutive planning periods, satisfying the following mathematical conditions:

[0092] ,

[0093] in, This represents the policy function at time t. Describes the function space norm. This is the first preset threshold.

[0094] (2) Model accuracy criterion: The average single-step prediction error of the environmental dynamics model for the power grid state remains below the third preset threshold over multiple consecutive planning periods, and the growth rate of its multi-step cumulative prediction error is lower than the fourth preset threshold. The prediction results of the environmental model are periodically compared with actual power grid operation data. The growth trends of the average error of single-step prediction and the error of multi-step cumulative prediction are statistically analyzed. If both indicators remain stable within the preset accuracy requirement range for multiple consecutive periods, the model accuracy is deemed to meet the standard. The prediction accuracy of the environmental dynamics model for the power grid state remains at the required level over M consecutive planning periods, satisfying the following mathematical conditions:

[0095] ,

[0096] in Indicates the Kullback-Leibler divergence. Represents the true state transition distribution. Indicates the model's predicted distribution. This is the second preset threshold.

[0097] (3) Reward Effectiveness Criterion: The strength of the intrinsic reward signal generated by the intrinsic reward network and its correlation with external rewards tend to stabilize over multiple consecutive planning periods, wherein the variance of the intrinsic reward is lower than the fifth preset threshold, and the absolute value of its correlation coefficient with the external reward obtained in the same decision period is lower than the sixth preset threshold. The changing patterns of the intrinsic reward signal and its correlation with external rewards are monitored in real time, and the fluctuation level of the intrinsic reward and their correlation coefficient are calculated. If the fluctuation of the intrinsic reward slows down over multiple consecutive periods and the correlation with external rewards remains at a low level, then the reward mechanism is considered stable and effective. The intrinsic reward signal generated by the intrinsic reward network tends to stabilize over M consecutive planning periods, satisfying the following mathematical conditions:

[0098] ,

[0099] in, Represents variance. Expressing expectations, This is the third preset threshold.

[0100] When each of the criteria (1), (2), and (3) is continuously satisfied within multiple consecutive planning periods, the strategy network, environmental dynamics model, and intrinsic reward network are determined to have reached a state of coordinated convergence. A real-time monitoring dashboard is established to track the compliance status of the above three criteria in parallel. When all three indicators simultaneously meet the threshold requirements within a preset number of consecutive periods, the system automatically issues a convergence completion signal and initiates the final planning scheme solidification output process.

[0101] S7. Output a multi-stage power grid planning scheme that meets multi-priority load demands, adapts to environmental uncertainties, and achieves coordinated optimization of generation and energy storage. This step terminates the learning process and delivers the final result through scheme generation and output. When the entire system (strategy, environment model, intrinsic reward) reaches a stable and coordinated convergence state after multiple rounds of iterative training, learning stops, and the final multi-stage power grid planning scheme is solidified and output.

[0102] Specifically, in this embodiment, the multi-stage power grid planning scheme disclosed in this invention is a complete and detailed intelligent decision-making output document. This scheme not only provides specific hierarchical scheduling instructions, but also covers adaptive strategies for various uncertainties, quantitative evaluation of collaborative optimization effects, identification and control of potential risks, and dynamic monitoring and adjustment guidelines during implementation, forming a systematic planning outcome integrating strategy, scenario, evaluation, risk control, and monitoring.

[0103] Specifically, at the strategy level, the solution outputs a phased, fine-grained sequence of dispatch instructions, clarifying the power supply guarantee schemes for each priority load, the start-up and shutdown and output plans of generator units, and the charging and discharging dispatch strategies for energy storage devices, providing a direct and actionable guide for the implementation of the plan. At the adaptability level, the solution pre-sets differentiated response strategies and contingency measures for typical uncertainty scenarios such as load fluctuations, renewable energy output deviations, and key equipment failures, enhancing the robustness of the plan in complex and volatile environments. At the effectiveness evaluation level, the solution provides a series of quantitative evaluation results, including multi-stage total cost predictions, power supply reliability indicators, and generation-storage synergy efficiency parameters, systematically evaluating the effectiveness of the plan from multiple dimensions such as economy, reliability, and synergy. At the risk control level, the solution conducts a forward-looking analysis of potential risks that may arise during the implementation of the plan, including an analysis of the satisfaction of key constraints, a sensitivity assessment of the impact of uncertainties, and provides corresponding early warning thresholds and risk response priorities. Finally, at the monitoring and adjustment level, the plan clarifies the key monitoring variables during implementation, the decision-making and adjustment trigger conditions when deviations from expectations occur, and the scheme switching logic when environmental parameters change significantly, ensuring that the planning scheme can be fed back and optimized in a timely and effective manner during dynamic execution.

[0104] Example 2

[0105] like Figure 2 As shown, this application provides an architecture diagram of a power grid multi-stage planning intelligent auxiliary decision-making system that integrates reinforcement learning. It is applied to the power grid multi-stage planning intelligent auxiliary decision-making system that integrates reinforcement learning as described in Embodiment 1, including: a modeling and initialization module 210, an initial environment model construction module 220, a trajectory generation and simulation module 230, a strategy and reward optimization module 240, an actual interaction and data collection module 250, an iterative control and model update module 260, and a scheme output module 270.

[0106] The modeling and initialization module 210 is used to model the multi-stage planning problem of the power grid as a Markov decision process in the probability measure space based on the distribution of power grid resources, load priority grouping and environmental state parameter set, and to initialize the policy network, environmental model network and intrinsic reward network.

[0107] The initial environment model construction module 220 is used to construct an initial environment dynamics model based on historical power grid operation data and by training a conditional generative adversarial network. The model takes load fluctuation rate, renewable energy output characteristics and fault probability as input conditions, and predicts the future power grid state based on the current state and decision actions.

[0108] The trajectory generation and simulation module 230 is used to generate virtual planning trajectory samples containing state transitions by performing forward simulation using the current environmental dynamics model and policy network at each planning stage.

[0109] The strategy and reward optimization module 240 is used to update the parameters of the strategy network and the intrinsic reward network based on the virtual planning trajectory sample; wherein, the strategy network parameters are first updated according to the composite gradient of the external reward and the intrinsic reward, and then the intrinsic reward network parameters are updated according to the gradient of the external reward.

[0110] The actual interaction and data collection module 250 is used to apply the updated policy network to the actual environment of the power grid, execute resource allocation decisions, and collect the generated status and decision interaction data.

[0111] The iterative control and model update module 260 is used to add the collected new data to the training set to update the environmental dynamics model, and after the model update is completed, it controls the return to the trajectory generation and simulation module to start a new round of iteration until the system reaches the preset convergence criterion.

[0112] The scheme output module 270 is used to output a multi-stage power grid planning scheme that meets the multi-priority load demand, adapts to environmental uncertainties, and achieves coordinated optimization of power generation and energy storage when the system reaches the convergence criterion.

[0113] Figure 3 This is an electronic device provided in one embodiment of this application. For example... Figure 3 As shown, the electronic device includes at least the following components: processor 301 and memory 300, communication interface 303, and bus 302.

[0114] In this embodiment of the application, memory 300 is used to store executable instructions of processor 301, which, when configured to execute instructions, implements the method as described in the first aspect.

[0115] In embodiments of this application, a computer-readable storage medium includes instructions that instruct a device to perform the method as described in the first aspect. For example, the instructions instruct the device to perform... Figure 1 The method is shown in the process steps.

[0116] In one embodiment of this application, the program operating in the electronic device may be a program that controls a central processing unit (CPU) or similar device to achieve the functions of the above-described embodiments of the present invention (a program that enables the computer to function). Information processed by these systems is then temporarily stored in random access memory (RAM) during processing, and subsequently stored in various ROMs such as read-only memory (FlashROM) and hard disk drives (HDDs), and read, corrected, and written by the CPU as needed.

[0117] It should be noted that a portion of the electronic device described in the above embodiments can also be implemented using a computer. In this case, the program for implementing the control function can be recorded on a computer-readable recording medium, and the program recorded on the recording medium can be read into the computer and executed.

[0118] It should be noted that the computer mentioned here refers to a computer built into an electronic device, employing hardware including an operating system and peripheral devices. Furthermore, computer-readable recording media refers to removable media such as floppy disks, magneto-optical disks, ROMs, and CD-ROMs, as well as storage systems such as hard drives built into the computer.

[0119] Furthermore, computer-readable recording media can include: media that dynamically stores programs for short periods of time, such as communication lines used when transmitting programs via networks like the Internet or communication lines like telephone lines; and media that store programs for fixed periods of time, such as volatile memory inside a computer that serves as a server or client in this case. In addition, the aforementioned program can be a program used to implement the above-mentioned functions, or it can be a program that can implement the above-mentioned functions by combining them with programs already recorded in the computer.

[0120] Furthermore, the electronic device in the above embodiments can also be implemented as an assembly (system group) composed of multiple systems. Each system constituting the system group can possess some or all of the functions or functional blocks of the electronic device in the above embodiments. As a system group, it is sufficient to have all the functions or functional blocks of the electronic device.

[0121] Those skilled in the art should recognize that the above embodiments are only used to illustrate this application and are not intended to limit this application. Any appropriate changes and variations made to the above embodiments within the essential spirit and scope of this application fall within the scope of protection claimed in this application.

Claims

1. A method for intelligent auxiliary decision-making in multi-stage planning of power grids that integrates reinforcement learning, characterized in that, Includes the following steps: S1. Based on the distribution of power grid resources, load priority grouping and environmental state parameter set, the multi-stage planning problem of power grid is modeled as a Markov decision process in the probability measure space, and the policy network, environmental model network and intrinsic reward network are initialized. S2. Based on historical power grid operation data, an initial environmental dynamics model is constructed by training a conditional generative adversarial network. This model takes load fluctuation rate, renewable energy output characteristics and fault probability as input conditions, and predicts the future power grid state based on the current state and decision actions. S3. In each planning stage, forward simulation is performed using the current environmental dynamics model and policy network to generate virtual planning trajectory samples containing state transitions. S4. Based on the virtual planning trajectory sample, update the parameters of the policy network and the intrinsic reward network; wherein, first update the parameters of the policy network according to the composite gradient of the external reward and the intrinsic reward, and then update the parameters of the intrinsic reward network according to the gradient of the external reward. S5. Apply the updated policy network to the actual power grid environment to perform resource allocation decisions and collect the generated state and decision interaction data. S6. Add the new data collected in S5 to the training set to update the environmental dynamics model; after the model update is completed, return to S3 to start a new round of iteration; repeat S3 to S6 until the system reaches the preset convergence criterion; S7. Output a multi-stage power grid planning scheme that meets the needs of multiple priority loads, adapts to environmental uncertainties, and achieves coordinated optimization of power generation and energy storage.

2. The method according to claim 1, characterized in that, In S1, the multi-stage power grid planning problem is modeled as a Markov decision process in the probability measure space, specifically constructed in the following way: (1) The state space is defined as a joint state vector that includes the output level of generator sets, the charging and discharging status of energy storage, the real-time demand of each priority load group, the statistical characteristics of load fluctuation rate, the deviation between the predicted output and the actual output of renewable energy, and the failure probability index of key nodes. (2) The action space is defined as a multi-dimensional continuous action vector that includes the power allocation coefficient of each priority load group, generator start-up and shutdown and output adjustment commands, energy storage charging and discharging power commands, and standby capacity allocation scheme. (3) The state transition function is learned through the environment model network and is used to approximate the probability distribution of transitioning to the next state given the current state and decision action; (4) The reward function adopts a hierarchical structure that includes basic external rewards and intrinsic exploration rewards. The basic external rewards reflect the generation cost and the penalty for violating service quality constraints, while the intrinsic exploration rewards are dynamically generated based on state prediction errors. (5) The value function adopts an approximation based on mean field theory, and measures the long-term comprehensive value of different decision-making strategies by evaluating the expected cumulative discount reward on the probability measure space.

3. The method according to claim 1 or 2, characterized in that: The policy network adopts a deep deterministic policy gradient framework, in which the Actor network outputs continuous resource allocation decisions based on the power grid state, and the Critic network evaluates the long-term value of different decisions in a given environment. The environment model network adopts a conditional generative adversarial network architecture. Its generator takes the current state, decision action, and environmental parameters as input conditions and introduces random noise to predict the power grid state in the next stage. The discriminator judges the authenticity of the state through adversarial training, thereby learning high-precision environmental transition dynamics. The intrinsic reward network is designed as an exploration incentive mechanism based on the state prediction error. Its output intrinsic reward signal is positively correlated with the prediction error of the environment model, which is used to drive the policy network to explore effectively in a sparse reward environment.

4. The method according to claim 1, characterized in that, The initial environmental dynamics model is constructed by training a conditional generative adversarial network in S2, specifically including: (1) Construct a conditional generative adversarial network framework. Its generator takes the current power grid state, resource allocation decision, load fluctuation rate, renewable energy output characteristics and fault probability as joint conditional inputs, and introduces random noise to output a prediction of the power grid state in the next stage. Its discriminator takes the same conditional inputs and the real or generated next state as inputs and outputs the probability that the state is the real state. (2) Phased adversarial training: First, pre-training is performed based on the historical power grid operation data to enable the generator to initially learn the state transition rules; then, in the closed-loop iteration process from S3 to S6, new data collected in S5 is continuously added for online fine-tuning and adversarial optimization. (3) Model application: After training convergence, the stable generator is used as the environmental dynamics model to simulate and predict the future multi-stage power grid state evolution trajectory based on the current state and candidate decisions during the planning stage.

5. The method according to claim 1, characterized in that, The generation of virtual planning trajectory samples in S3 specifically includes: (1) Based on the current actual state and state distribution of the power grid, set the initial conditions for the simulation; (2) Perform multi-step forward simulation: Repeat the following operations until the preset simulation step size or termination condition is reached: a. Input the current simulation state and distribution into the policy network to obtain the decision action; b. Input the current state, decision-making actions, and environmental parameters into the environmental dynamics model to obtain a prediction of the next simulation state; c. Calculate the immediate reward based on the reward function; d. Update the state distribution based on the predicted next state; e. Record the current state, decision action, reward, next state, and updated distribution as a transition sample, and update the current state; (3) Construct a complete virtual planning trajectory from the ordered transition sample sequence generated in the loop; (4) By generating multiple independent trajectories in parallel from different initial states or by introducing different random noise, a batch sample set for training is constructed.

6. The method according to claim 1, characterized in that, In step S4, the parameters of the policy network and the intrinsic reward network are updated, specifically through the following method: (1) Calculation of composite reward gradient: For the virtual planning trajectory, calculate its accumulated composite reward, which is the weighted sum of external reward and internal reward; based on the expected value of the composite reward, calculate the update gradient of the policy network parameters through the policy gradient algorithm. (2) Policy network parameter update: Using the gradient, the gradient ascent method is used to optimize the policy network parameters; (3) Setting the intrinsic reward objective: The learning objective of the intrinsic reward network parameters is to maximize the expected cumulative external reward obtained from the same batch of virtual trajectories; (4) Calculation and update of intrinsic reward gradient: The gradient of the expected cumulative external reward relative to the intrinsic reward network parameters is calculated by using the chain rule, and the intrinsic reward network parameters are updated accordingly.

7. The method according to claim 6, characterized in that, The parameter update process employs an alternating optimization two-layer gradient mechanism: (1) Outer layer optimization: Based on the current intrinsic reward network parameters, calculate the compound reward and update the policy network parameters with the goal of maximizing its expected value; (2) Inner layer optimization: Fix the updated policy network parameters and calculate and update the inner reward network parameters with the goal of maximizing the expected cumulative external reward. (3) Alternately execute steps (1) and (2) to form an optimization loop that alternates between inner and outer layers, so that the internal reward signal can be adaptively adjusted to drive the policy network to explore in the direction of improving the performance of the final task.

8. The method according to claim 1, characterized in that, In S5, the policy network is applied to real-world environmental interactions and data collection, specifically including: (1) Online deployment: Deploy the updated strategy network to the actual power grid control system, generate and execute decisions based on real-time status; (2) Synchronous recording: After each decision is executed, the actual state at the moment of decision, the decision action output by the policy network, the immediate external reward from the environment, and the observed actual state of the next stage are recorded. (3) Data validation: Validate the collected interactive data and filter out invalid data; (4) Dataset expansion: Add effective data to the training set for model and policy updates.

9. The method according to claim 1, characterized in that, The preset convergence criteria specifically include the following quantifiable criteria: (1) Strategy convergence criterion: The resource allocation decision output by the strategy network remains stable over multiple consecutive planning periods, wherein the rate of change of the power ratio allocated to each priority load group is lower than the first preset threshold, and the change amplitude of the scheduling instructions for power generation and energy storage equipment is lower than the second preset threshold. (2) Model accuracy criteria: The mean single-step prediction error of the environmental dynamics model for the power grid state remains below the third preset threshold in multiple consecutive planning periods, and the growth rate of its multi-step cumulative prediction error is lower than the fourth preset threshold. (3) Reward effectiveness criterion: The strength of the intrinsic reward signal generated by the intrinsic reward network and its correlation with the external reward tend to be stable over multiple consecutive planning periods, wherein the variance of the intrinsic reward is lower than the fifth preset threshold, and the absolute value of its correlation coefficient with the external reward obtained in the same decision period is lower than the sixth preset threshold. When each of the criteria (1), (2), and (3) is continuously satisfied over multiple consecutive planning periods, it is determined that the policy network, the environmental dynamics model, and the intrinsic reward network have reached a state of coordinated convergence.

10. A smart auxiliary decision-making system for multi-stage planning of power grids integrating reinforcement learning, applied to the method described in any one of claims 1 to 9, characterized in that, The system includes: The modeling and initialization module is used to model the multi-stage planning problem of the power grid as a Markov decision process in the probability measure space based on the distribution of power grid resources, load priority grouping and environmental state parameter set, and to initialize the policy network, environmental model network and intrinsic reward network. The initial environment model building module is used to build an initial environment dynamics model based on historical power grid operation data and by training a conditional generative adversarial network. The model takes load fluctuation rate, renewable energy output characteristics and fault probability as conditional inputs and predicts the future power grid state based on the current state and decision actions. The trajectory generation and simulation module is used to perform forward simulation using the current environmental dynamics model and policy network at each planning stage, generating virtual planning trajectory samples that include state transitions. The strategy and reward optimization module is used to update the parameters of the strategy network and the intrinsic reward network based on the virtual planning trajectory sample; wherein, the strategy network parameters are first updated according to the composite gradient of the external reward and the intrinsic reward, and then the intrinsic reward network parameters are updated according to the gradient of the external reward. The actual interaction and data collection module is used to apply the updated policy network to the actual environment of the power grid, execute resource allocation decisions, and collect the generated status and decision interaction data. The iterative control and model update module is used to add the collected new data to the training set to update the environmental dynamics model, and after the model update is completed, it controls the return to the trajectory generation and simulation module to start a new round of iteration until the system reaches the preset convergence criterion. The scheme output module is used to output a multi-stage power grid planning scheme that meets the multi-priority load demand, adapts to environmental uncertainties, and achieves coordinated optimization of power generation and energy storage when the system reaches the convergence criterion.