Load aggregator flexible resource regulation and control method and system based on DDPG algorithm

By adopting a differentiated user incentive and multi-objective optimization control model based on the DDPG algorithm, the problems of lack of differentiation in user incentive strategies and low scheduling optimization efficiency are solved. This achieves the maximization of total profit for load aggregators and scheduling compliance, thereby improving the effectiveness and economy of flexible load resource scheduling.

CN122052044APending Publication Date: 2026-05-15CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD +3
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-23
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies lack differentiation in user incentive strategies, have low scheduling optimization efficiency, and suffer from wasteful incentive costs and scheduling conflicts due to decentralized constraint processing, making it difficult to maximize the total profit of load aggregators.

Method used

A flexible resource control method based on the DDPG algorithm is adopted. By constructing a differentiated user incentive cost model and a multi-objective optimization control model, and combining Markov decision process, the electricity price and dispatch strategy are dynamically adjusted to optimize the flexible load resource dispatch.

Benefits of technology

It improved user willingness to respond, reduced incentive costs, enhanced the real-time performance and compliance of scheduling strategies, increased the total profit of load aggregators, and achieved efficient and flexible load resource scheduling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122052044A_ABST
    Figure CN122052044A_ABST
Patent Text Reader

Abstract

The invention discloses a load aggregator flexible resource regulation and control method and system based on a DDPG algorithm. The method comprises the steps of obtaining historical operation data of a load aggregator; based on the user comfort loss, dynamically adjusting the electricity price through a pre-constructed user incentive cost model; substituting the historical operation data of the load aggregator and the dynamically adjusted electricity price into a pre-constructed multi-target optimization regulation and control model, and solving through a Markov decision process to obtain a flexible load resource scheduling scheme; regulation and control are carried out based on a flexible load resource scheduling scheme; wherein the pre-constructed multi-target optimization regulation and control model is constructed by taking the maximum total profit of the load aggregator as a target and combining constraint conditions. According to the method, the differentiated user incentive cost model is designed, the user comfort loss is matched by dynamically adjusting the electricity price, the user response willingness is improved, the invalid incentive cost is reduced, the multi-target optimization regulation and control model is constructed, the DDPG reinforcement learning algorithm is introduced, and efficient and dynamic solving of the scheduling strategy in the multi-constraint scene is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power system automation control technology, specifically to a load aggregator flexible resource regulation method and system based on the DDPG algorithm. Background Technology

[0002] As the penetration rate of new energy sources (wind power, solar power, etc.) in the power system continues to increase, the volatility and intermittency of their output significantly increase the difficulty of balancing power grid supply and demand. Flexible load resources (such as industrial production equipment, electric vehicles, commercial building air conditioning, etc.), as the core carrier of demand-side response, have become a key force in smoothing power grid load fluctuations and improving system flexibility. Against this backdrop, load aggregators have emerged, whose core function is to integrate dispersed flexible load resources and formulate control strategies based on power grid dispatching needs (such as peak-valley regulation and emergency power support) to achieve precise matching between resources and power grid demand.

[0003] Currently, key technologies in the field of flexible load resource regulation include: 1) load resource aggregation technology, which integrates adjustable loads of different types and with different spatiotemporal characteristics through a unified platform; 2) dispatch strategy optimization technology, which needs to meet multi-dimensional constraints such as grid demand, upper and lower limits of resource dispatch volume, and dispatch duration / interval; 3) user incentive technology, which guides users to cooperate with load reduction or transfer through economic compensation to ensure that the regulation target is achieved.

[0004] However, existing technical solutions have the following significant drawbacks: User incentive strategies lack differentiation: Traditional incentive models are centered on "the scalability of all users" and adopt uniform electricity prices or fixed subsidy standards. They ignore the differences in comfort loss of different users (or different resources) when load is reduced (e.g., the cost of reducing load when industrial equipment is at full load is higher than when it is at low load, and the demand sensitivity of electric vehicle charging is higher in the later stages of reduction). This leads to wasted incentive costs or insufficient user response, making it difficult to achieve the expected control effect. Low efficiency in scheduling optimization: Existing scheduling models mostly adopt traditional optimization methods such as linear programming and integer programming. When faced with multiple constraints such as "upper and lower limits of resource scheduling, total scheduling volume per hour not exceeding grid demand, single scheduling duration / interval, zeroing of industrial equipment load transfer, and satisfaction of electric vehicle charging demand", they have poor dynamic adaptability and real-time performance, cannot quickly respond to changes in the dynamic demand of the grid, and are difficult to achieve the core objective of "maximizing the total profit (compensation revenue - incentive cost) of load aggregators". The constraints are handled in a decentralized manner: Traditional solutions often use the method of "embedding optimization equations one by one" to handle various scheduling constraints, which leads to high model complexity, slow solution convergence, and easy constraint conflicts (such as the conflict between industrial equipment load transfer constraints and the real-time demand of the power grid), affecting scheduling compliance and efficiency. Summary of the Invention

[0005] To address the problems of traditional user incentive models, such as ignoring individual differences, poor real-time performance, difficulty in achieving profit targets, dispersion and complexity, and susceptibility to conflicts, this invention proposes a flexible resource regulation method for load aggregators based on the DDPG algorithm, including: Obtain historical operational data from load aggregators; Based on the loss of user comfort, the electricity price is dynamically adjusted through a pre-built user incentive cost model. By substituting the historical operating data of the load aggregator and the dynamically adjusted electricity price into the pre-constructed multi-objective optimization control model, and solving it through the Markov decision process, a flexible load resource scheduling scheme is obtained. Regulation is carried out based on the aforementioned flexible load resource scheduling scheme; The pre-constructed multi-objective optimization control model is built with the goal of maximizing the total profit of the load aggregator, combined with constraints.

[0006] Preferably, the construction of the multi-objective optimization control model includes: Construct an objective function with the goal of maximizing the total profit of the load aggregator; Set constraints for the objective function; The constraints include: upper and lower limits of the schedulable quantity of each resource, the schedulable quantity of all resources not exceeding the grid demand for each set time period, single scheduling duration constraints, scheduling interval constraints, maximum schedulable times during the control period constraints, total schedulable quantity constraints of industrial equipment, and electric vehicle charging quantity constraints.

[0007] Preferably, the objective function is as follows:

[0008] In the formula, To the total profit of the load aggregator, The load aggregator's electricity price at time t; This represents the actual response of all adjustable resources at time t; The user incentive electricity price represents the price at time t. This represents the actual response amount of resource i at time t; N represents the ratio of the current scheduling quantity to the available scheduling quantity of the i-th resource at time t; N represents the number of resources; and T represents the control period.

[0009] Preferably, the pre-built user incentive cost model is shown in the following formula:

[0010] In the formula, This represents the user incentive electricity price for the i-th resource; This represents the ratio of the current schedulable quantity to the available schedulable quantity of the i-th resource. Represents the current scheduling amount of the i-th resource; This represents the maximum schedulable quantity of the i-th resource; The coefficient represents the quadratic function of electricity price, and a > 0. Preferably, the step of substituting historical operating data of load aggregators and dynamically adjusted electricity prices into a pre-constructed multi-objective optimization control model, and solving it through a Markov decision process to obtain a flexible load resource scheduling scheme, includes: Set the constraints in the multi-objective optimization control model to the environment; The scheduling amount of all resources for each set period of time in the historical operation data of the load aggregator is used as the status. The current scheduling amount of each resource per set time period is taken as an action; the process of executing an action, updating the scheduling table, and advancing to the next moment is taken as a transition; the scheduling profit for this set time period is taken as a reward. The state is input into the Actor network, which outputs the action. The state and action are then input into the Critic network, which outputs the Q value. This process is repeated until convergence or a predetermined number of rounds is reached, resulting in a flexible load resource scheduling scheme.

[0011] Preferably, the step of inputting the state into the Actor network and outputting the action, inputting the state and action into the Critic network and outputting the Q value, and repeating this process continuously until convergence or a predetermined number of rounds is reached to obtain a flexible load resource scheduling scheme, includes: Step 1: Initialize the Actor network and Critic network, as well as the target network corresponding to the Actor network and Critic network, and initialize the experience replay pool; Step 2: In each step, use an Actor network based on the current state s t Select action a, increase exploration noise, and you get action a with increased exploration noise. t Execute a t Observe the new state s returned by the environment. t+1 and reward r t ,Will Add to the experience replay pool; Step 3: Randomly sample N experience samples from the experience replay pool. The Q-value y of the next state is estimated using a target network consisting of Critic and Actor. i The Critic network is updated using mean squared error loss; the Actor network is updated using policy gradient; and the target networks corresponding to the Actor and Critic networks are updated using a small step size τ of the parameters of the Actor and Critic networks. Step 4: Determine whether convergence has been achieved or the predetermined number of rounds has been reached. If convergence has not been achieved or the predetermined number of rounds has not been reached, return to Step 2. Otherwise, obtain the action that maximizes the output Q, and use the current hourly scheduling amount of each resource corresponding to the action as the flexible load resource scheduling scheme.

[0012] Furthermore, this invention also provides a load aggregator flexible resource regulation system based on the DDPG algorithm, comprising: The parameter acquisition module is used to acquire historical operating data of the load aggregator; The electricity price adjustment module is used to dynamically adjust the electricity price based on the loss of user comfort using a pre-built user incentive cost model. The scheme generation module is used to input the historical operating data of the load aggregator and the dynamically adjusted electricity price into the pre-built multi-objective optimization control model, and solve it through the Markov decision process to obtain a flexible load resource scheduling scheme. The control module is used to perform control based on the flexible load resource scheduling scheme; The pre-constructed multi-objective optimization control model is built with the goal of maximizing the total profit of the load aggregator, combined with constraints.

[0013] Preferably, it further includes a model building module, the model building module being used for: Construct an objective function with the goal of maximizing the total profit of the load aggregator; Set constraints for the objective function; The constraints include: upper and lower limits of the schedulable quantity of each resource, the schedulable quantity of all resources not exceeding the grid demand for each set time period, single scheduling duration constraints, scheduling interval constraints, maximum schedulable times during the control period constraints, total schedulable quantity constraints of industrial equipment, and electric vehicle charging quantity constraints.

[0014] Preferably, the objective function is as follows:

[0015] In the formula, The load aggregator's electricity price at time t; This represents the actual response of all adjustable resources at time t; The user incentive electricity price represents the price at time t. This represents the actual response amount of resource i at time t; N represents the ratio of the current scheduling quantity to the available scheduling quantity of the i-th resource at time t; N represents the number of resources; and T represents the control period.

[0016] Preferably, the pre-built user incentive cost model is shown in the following formula:

[0017] In the formula, This represents the user incentive electricity price for the i-th resource; This represents the ratio of the current schedulable quantity to the available schedulable quantity of the i-th resource. Represents the current scheduling amount of the i-th resource; This represents the maximum schedulable quantity of the i-th resource; The coefficient represents the quadratic function of electricity price, and a > 0.

[0018] Preferably, the scheme generation module is specifically used for: Set the constraints in the multi-objective optimization control model to the environment; The scheduling amount of all resources for each set period of time in the historical operation data of the load aggregator is used as the status. The current scheduling amount of each resource per set time period is taken as an action; the process of executing an action, updating the scheduling table, and advancing to the next moment is taken as a transition; the scheduling profit for this set time period is taken as a reward. The state is input into the Actor network, which outputs the action. The state and action are then input into the Critic network, which outputs the Q value. This process is repeated until convergence or a predetermined number of rounds is reached, resulting in a flexible load resource scheduling scheme.

[0019] Preferably, in the scheme generation module, the state is input into the Actor network, which outputs actions; the state and actions are input into the Critic network, which outputs Q-values; this process is repeated until convergence or a predetermined number of rounds is reached to obtain a flexible load resource scheduling scheme. Specific implementation steps include: Step 1: Initialize the Actor network and Critic network, as well as the target network corresponding to the Actor network and Critic network, and initialize the experience replay pool; Step 2: In each step, use an Actor network based on the current state s t Select action a, increase exploration noise, and you get action a with increased exploration noise. t Execute a t Observe the new state s returned by the environment. t+1 and reward r t ,Will Add to the experience replay pool; Step 3: Randomly sample N experience samples from the experience replay pool. The Q-value y of the next state is estimated using a target network consisting of Critic and Actor. iThe Critic network is updated using mean squared error loss; the Actor network is updated using policy gradient; and the target networks corresponding to the Actor and Critic networks are updated using a small step size τ of the parameters of the Actor and Critic networks. Step 4: Determine whether convergence has been achieved or the predetermined number of rounds has been reached. If convergence has not been achieved or the predetermined number of rounds has not been reached, return to Step 2. Otherwise, obtain the action that maximizes the output Q, and use the current hourly scheduling amount of each resource corresponding to the action as the flexible load resource scheduling scheme.

[0020] In another aspect, this application also provides an electronic device, comprising: at least one processor and a memory; the memory and the processor are connected via a bus; The memory is used to store one or more programs; When the one or more programs are executed by the at least one processor, the load aggregator flexible resource regulation method based on the DDPG algorithm described above is implemented.

[0021] In another aspect, this application also provides a computer-readable storage medium having an executable program stored thereon, which, when executed, implements the load aggregator flexible resource regulation method based on the DDPG algorithm as described above.

[0022] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention provides a flexible resource regulation method for load aggregators based on the DDPG algorithm. The method includes acquiring historical operating data of load aggregators; dynamically adjusting electricity prices based on user comfort loss using a pre-constructed user incentive cost model; substituting the historical operating data and dynamically adjusted electricity prices into a pre-constructed multi-objective optimization regulation model, solving it through a Markov decision process to obtain a flexible load resource scheduling scheme; and performing regulation based on the flexible load resource scheduling scheme. The pre-constructed multi-objective optimization regulation model is built with the goal of maximizing the total profit of the load aggregator, combined with constraints. This invention designs a differentiated user incentive cost model that fits the characteristics of resource scheduling. By dynamically adjusting electricity prices to match user comfort loss, it improves user response willingness and reduces ineffective incentive costs. It constructs a multi-objective optimization regulation model with the goal of "maximizing the total profit of the load aggregator," and introduces the DDPG reinforcement learning algorithm to achieve efficient and dynamic solution of scheduling strategies under multiple constraints. This solves the problems of traditional user incentive models, such as ignoring individual differences, poor real-time performance, difficulty in achieving profit targets, dispersion and complexity, and susceptibility to conflicts. Attached Figure Description

[0023] Figure 1 This is a flowchart of the load aggregator flexible resource regulation method based on the DDPG algorithm of the present invention; Figure 2 This is a schematic diagram of the DDPG network structure of the present invention; Figure 3 This is a flowchart of the DDPG algorithm of the present invention; Figure 4 This is a schematic diagram of an electronic device structure according to the present invention. Detailed Implementation

[0024] This invention provides a flexible load resource regulation method based on the DDPG algorithm, aiming to solve the following technical problems to improve the effectiveness, economy, and compliance of flexible load resource regulation: To address the problem of traditional user incentive models ignoring individual differences, a differentiated user incentive cost model tailored to resource scheduling characteristics is designed. This model dynamically adjusts electricity prices to match user comfort losses, thereby increasing user willingness to respond and reducing ineffective incentive costs. To address the issues of poor real-time performance and difficulty in achieving profit targets in scheduling optimization under multiple constraints, a multi-objective optimization and control model is constructed with the objective of maximizing the total profit of load aggregators. The DDPG reinforcement learning algorithm is introduced to achieve efficient and dynamic solution of scheduling strategies under multiple constraints. To address the issues of "dispersed, complex, and conflict-prone" constraint handling, scheduling constraints are uniformly incorporated into the "environment" of a Markov Decision Process (MDP). Through a mechanism of "action violation - environment adjustment - reward and punishment," the constraint handling logic is simplified, scheduling compliance is ensured, and solution efficiency is improved.

[0025] To better understand the present invention, the following description, in conjunction with the accompanying drawings and embodiments, will further illustrate the content of the present invention.

[0026] Example 1: A flexible resource regulation method based on the load aggregator algorithm, such as... Figure 1 As shown, it includes: Step 1: Obtain historical operating data from the load aggregator; Step 2: Based on the loss of user comfort, dynamically adjust the electricity price using a pre-built user incentive cost model; Step 3: Substitute the historical operating data of the load aggregator and the dynamically adjusted electricity price into the pre-built multi-objective optimization control model, and solve it through the Markov decision process to obtain the flexible load resource scheduling scheme; Step 4: Perform regulation based on the aforementioned flexible load resource scheduling scheme; The pre-constructed multi-objective optimization control model is built with the goal of maximizing the total profit of the load aggregator, combined with constraints.

[0027] The following is a further description of the various steps involved in this invention: Step 1: Obtain historical operational data from the load aggregator, including: Obtain the load aggregator's electricity price at time t, the actual response of all adjustable resources at time t, the user incentive electricity price at time t, the actual response of resource i at time t, the current scheduling quantity of the i-th resource, the scheduling quantity of the i-th resource, the number of air conditioners in the shopping mall, the number of resources, the response of the i-th air conditioner at time t, the number of industrial equipment, the response of the i-th industrial equipment at time t, the number of electric vehicles, and the response of the i-th electric vehicle at time t, etc.

[0028] Step 2: Based on the loss of user comfort, dynamically adjust the electricity price using a pre-built user incentive cost model, including: The pre-built user incentive cost model is shown in the following equation:

[0029] In the formula, This represents the user incentive electricity price for the i-th resource; This represents the ratio of the current schedulable quantity to the available schedulable quantity of the i-th resource. Represents the current scheduling amount of the i-th resource; represents the maximum schedulable quantity of the i-th resource; a, b, and c all represent the coefficients of the quadratic function of electricity price, and a > 0.

[0030] Step 2 specifically includes: S2 User Incentive Cost Model Construction.

[0031] During demand-side response, load aggregators need to provide incentives to users based on their reduction amounts. Traditional user incentive models generally take the reduction feasibility of all users as the starting point and formulate a uniform incentive strategy. The resulting user incentive models ignore the differences in individual users' reduction efforts, and often fail to achieve the expected results in practical applications. A user incentive cost model is formulated based on the principle that the closer the resource scheduling volume is to its upper limit, the higher the electricity price.

[0032] For an individual user, initially, the power reduction has a small impact on user comfort, and the incentive cost is also small. However, as the power reduction increases, the impact on user comfort also increases, and the incentive cost becomes larger. Assume the electricity price C and the resource scheduling quantity x... i The relationship between them can be represented by a quadratic function, where the closer the resource allocation is to its upper limit, the higher the electricity price. The user cost incentive model is shown in formula (10).

[0033]

[0034] In the formula: This represents the user incentive electricity price for the i-th resource; This represents the ratio of the current schedulable quantity to the available schedulable quantity of the i-th resource. Represents the current scheduling amount of the i-th resource; This represents the maximum schedulable quantity of the i-th resource; The coefficients represent the quadratic function of electricity price, and a>0 ensures that the electricity price increases with the increase of dispatch volume.

[0035] Before step 3, a multi-objective optimization scheduling model is constructed. The construction process of this multi-objective optimization scheduling model includes: Construct an objective function with the goal of maximizing the total profit of the load aggregator; Set constraints for the objective function; The constraints include: upper and lower limits of the schedulable quantity of each resource, the schedulable quantity of all resources per hour not exceeding the grid demand, single scheduling duration constraints, scheduling interval constraints, maximum number of schedulable times per day constraints, total schedulable quantity constraints of industrial equipment, and electric vehicle charging quantity constraints.

[0036] Furthermore, the objective function is shown in the following equation:

[0037] In the formula, The load aggregator's electricity price at time t; This represents the actual response of all adjustable resources at time t; The user incentive electricity price represents the price at time t. This represents the actual response amount of resource i at time t; The ratio of the current scheduling quantity to the available scheduling quantity of the i-th resource at time t is represented by N; N represents the number of resources; T represents the control period. In this embodiment, the control period is set to 24 hours and the interval setting time is 1 hour.

[0038] The process of constructing a multi-objective optimization control model specifically includes: The objective function is to maximize the total profit (compensation revenue minus incentive cost) of the load aggregator. Constraints include upper and lower limits on the dispatchable quantities of each resource, a maximum hourly dispatchable quantity of all resources not exceeding grid demand, constraints on single dispatch duration, dispatch intervals, the maximum number of dispatchable times per day, the total dispatchable quantity of industrial equipment, and constraints on electric vehicle charging volume.

[0039] (1) Objective function The profit of a load aggregator is the difference between the economic benefits generated by grid regulation and the user incentive costs. The objective function of this paper is to maximize the total profit of the load aggregator.

[0040]

[0041] In the formula: The load aggregator's electricity price at time t; This represents the actual response of all adjustable resources at time t; The user incentive electricity price represents the price at time t. The electricity price coefficient representing resource i; This represents the actual response amount of resource i at time t; N represents the ratio of the current scheduling quantity to the available scheduling quantity of the i-th resource; N represents the number of resources; T represents 24 hours. The number of air conditioners in the shopping mall. Let be the response quantity of the i-th air conditioner at time t. For the number of industrial equipment, Let be the response of the i-th industrial device at time t. For the number of electric vehicles, Let be the response of the i-th electric vehicle at time t.

[0042] (2) Constraints The constraints include upper and lower limits on the dispatchable amount of each resource, the dispatchable amount of all resources per hour not exceeding the grid demand, single dispatch duration constraints, dispatch interval constraints, maximum number of dispatchable times per day constraints, total dispatchable amount of industrial equipment constraints, and electric vehicle charging amount constraints.

[0043] ① Upper and lower limits of schedulable quantities for each resource For each resource type, the scheduling amount must be within the schedulable range:

[0044] In the formula, Let i be the minimum scheduling amount for resource i. Let be the resource scheduling amount at time t. Let i be the maximum scheduling amount for resource i.

[0045] ② The hourly dispatch volume of all resources shall not exceed the grid demand constraint.

[0046] In the formula: D(t) is the power grid demand at time t.

[0047] ③ Single scheduling duration constraint For each scheduling event k of each resource, its duration must satisfy minimum and maximum duration constraints:

[0048] In the formula: and These are the start and end times of the k-th scheduling of resource i; This represents the total number of times resource i was scheduled that day. and These are the minimum and maximum allowed single scheduling durations, where k is the number of scheduling attempts.

[0049] ④ Scheduling interval constraints For each resource, the minimum interval time must be satisfied between two consecutive schedulings:

[0050] In the formula, Let be the end time of the (k-1)th scheduling of resource i. The interval between two consecutive resource scheduling operations.

[0051] ⑤ Maximum number of schedulable events per day constraint The total number of times each resource can be scheduled in a day cannot exceed the maximum allowed number of times:

[0052] In the formula: This is the maximum number of scheduling attempts allowed (the value varies depending on the resource type).

[0053] ⑥ Total scheduling constraints for industrial equipment The total scheduling volume of industrial equipment within a day must be zero (load transfer):

[0054] And each time a scheduling occurs:

[0055] In the formula, Let be the maximum schedulable quantity of the i-th industrial equipment.

[0056] ⑦ Electric vehicle charging quantity constraint Electric vehicles must meet charging requirements at the end of the scheduling cycle:

[0057] In the formula: This represents the total amount of charging required for electric vehicle i; This represents the initial charge level of electric vehicle i; This represents the time window during which electric vehicles can be scheduled.

[0058] Step 3: Substitute the historical operating data of the load aggregator and the dynamically adjusted electricity price into the pre-constructed multi-objective optimization control model, and solve it through a Markov decision process to obtain a flexible load resource scheduling scheme, including: Set the constraints in the multi-objective optimization control model to the environment; The scheduling amount of all resources for each set period of time in the historical operation data of the load aggregator is used as the status. The current scheduling amount of each resource per set time period is taken as an action; the process of executing an action, updating the scheduling table, and advancing to the next moment is taken as a transition; the scheduling profit for this set time period is taken as a reward. The state is input into the Actor network, which outputs the action. The state and action are then input into the Critic network, which outputs the Q value. This process is repeated until convergence or a predetermined number of rounds is reached, resulting in a flexible load resource scheduling scheme.

[0059] Step 3 specifically includes: The solution for flexible load resource scheduling based on the DDPG algorithm includes: First, the problem is modeled as a Markov Decision Process (MDP), which includes four elements: state, action, transition, and reward. Constraints are not directly processed in the "state" or "action" space, but are uniformly processed in the "environment." If an action is violated, the environment will adjust the action and impose a penalty in the reward.

[0060] Table 1. Correspondence between MDP elements and power dispatching

[0061] The network structure of the DDPG algorithm is as follows: Figure 2 As shown in the diagram, Actor is the Actor network, Critic is the Critic network, T_actor is the target network of Actor, and T_critic is the target network of Critic. The input state s of the Actor network is , and the output action is . The parameter is θ. The output action; the input state and action of the Critic network. Output Q value The parameters are The target network of the Actor network is represented as follows; the target network of the Critic network is represented as follows. .

[0062] The DDPG algorithm flowchart is as follows: Figure 3 As shown, S represents the current state, A represents the current action, Q represents the score under the current state and action, S' represents the state at the next moment, A' represents the action at the next moment, TD-Error is the difference between the two scores, Q' represents the score under the state and action at the next moment, R is the reward, Minimize is the minimization function, and Target is the target value. The specific process steps are as follows.

[0063] ① Initialize the experience pool: Initialize the Actor network and Critic Network and their corresponding target networks and .

[0064] ② Interaction with the environment: At each step, the Actor network selects action a based on the current state s, usually with some exploration noise (N). t ),Right now Execute action a t Observe the new state s returned by the environment. t+1 and reward r t .Bundle This experience is added to the experience pool.

[0065] ③ Sample from the experience pool and update the network.

[0066] Randomly sample N experience samples from the experience pool. .

[0067] Calculate the target Q value y i The Q-value of the next state is estimated using the target Critic and the target Actor.

[0068]

[0069] In the formula: It is the target Q-value that the Critic network learns to fit; It is the reward after the AI ​​is executed; It is a discount factor (0 < <1, which determines the importance of future rewards; , It is the target network, and then the target Actor outputs the next action. Then, the target Critic provides an estimate of its future Q value.

[0070] Update the Critic network using mean squared error loss.

[0071]

[0072] The standard mean squared error loss (MSE) aims to make the Q-value output by the Critic network as close as possible to the target Q-value calculated in the previous step. L is the loss function of the Critic network.

[0073] Update the Actor network using policy gradients.

[0074] The core idea is that the Actor wants the output action to maximize the Q-value of the Critic network output in the current state.

[0075] (13) in, Taking the derivative with respect to a represents how the Q value changes if action a changes slightly in a certain state s. The derivative with respect to θ represents how the Actor's output is affected by the network parameters. After multiplication, backpropagation optimizes θ, making the Actor more inclined to produce actions with high Q values. J is the objective function of the Actor network policy.

[0076] Soft update of the target network: The parameters of the target network are updated using a small step size τ of the parameters of the main network; this prevents the target value from changing too quickly, thus ensuring training stability. The parameters of the target Critic network, For the parameters of the Critic network, The parameters of the target Actor network.

[0077]

[0078] ④ Repeat continuously: interact, sample, update until convergence or a predetermined number of rounds are reached. The task of the Actor network is to find A to maximize the output Q.

[0079] Modifications to DDPG: By adding random noise to deterministic actions, the DDPG agent can try more actions during training and find the optimal policy more quickly. Some changes were made to the program's constraint handling and noise exploration mechanisms. ①Incorporate a separation constraint handling mechanism: The `calculate_step_reward` method in the program handles immediate feedback, such as constraints that the hourly resource scheduling exceeds the grid demand, and updates it once every time step. The `calculate_final_reward` method in the program handles global constraints, such as timing constraints, industrial equipment scheduling mismatch constraints, and electric vehicle not fully charged constraints, and updates it once every 24 time steps.

[0080] ② Add Ornstein-Uhlenbeck noise:

[0081] In the formula: This represents the noise value at the current moment; This represents the long-term mean of the noise. For random noise; σ is the tiny increment of the Wiener process; θ is the mean regression velocity parameter, the larger θ is, the faster the noise returns to the mean; σ is the volatility, which controls the intensity of random fluctuations, the larger σ is, the more violent the noise fluctuations.

[0082] First, compared to simple Gaussian noise, OU noise produces a smoother and more continuous exploration trajectory, which is crucial for problems like virtual power plant dispatching. Second, in power dispatching, actions typically cannot change abruptly. The exploration trajectory generated by OU noise better reflects the inertial characteristics of real systems. Finally, by controlling the parameters θ and σ, the exploration level can be dynamically adjusted during training to achieve a balance between exploration and utilization. In the early stages of training, a high σ value promotes extensive exploration. In the later stages of training, a lower σ value allows for greater utilization of the learned strategy.

[0083] The multi-objective optimization control model is modeled as a Markov decision process: Directly solving multi-objective optimization control models using dynamic programming results in high time complexity, large space complexity, and difficulty in handling multi-dimensional constraints. Therefore, we first decompose the multi-objective optimization control model into a Markov Decision Process (MDP), and then use the DDPG algorithm to solve the model. In this section, we will introduce the four basic elements of an MDP: state, action, transition, and reward. These are specifically expressed as follows:

[0084] In the formula: Representative Resources exist The amount of time-based scheduling; This represents the state and action of the DDPG model at time t; Represents the state transition probability; Represents resource prices; T represents resource cost; T represents 24 hours; N represents the quantity of resources. Represents the length of the time period; These represent penalties for exceeding demand, timing constraints, mismatched industrial equipment scheduling, and electric vehicles not being fully charged, respectively. It is a scheduling duration penalty; It is a scheduling interval penalty; It's a penalty for the number of scheduling attempts; These represent the penalty coefficients for response exceeding demand, timing constraints, mismatch in industrial equipment scheduling, and incomplete charging of electric vehicles, respectively. It is the response quantity of resource i at time t; It represents the demand of the distribution network at time t; It is the scheduling duration of resource i; It is the maximum scheduling duration; It is the minimum scheduling time; It is the interval between two consecutive schedulings of resource i; It is the minimum scheduling interval; It represents the number of times resource i is scheduled; It is the maximum number of times resource i is scheduled. It represents the amount of charge applied to electric vehicle i at time t; K represents the maximum electricity demand of electric vehicle i, and K is the sequence number of the scheduling task. For the reward function, For the number of scheduling, Let n be the scheduling amount of n resources in the first t time steps, where n is the total number of resources. Let n be the scheduling amount of n resources at time t. The time when the k-th scheduled task ends. The start time of the k-th scheduled task. The number of scheduling intervals, The number of industrial equipment.

[0085] This invention, through a technical solution combining a "multi-objective optimization control model + differentiated user incentive model + DDPG algorithm solution," achieves the following significant advantages compared to existing technologies: (1) Improved effectiveness of user incentives: The user incentive cost model (Formula 10) constructed using a quadratic function links the electricity price to the resource scheduling volume (the closer the scheduling volume is to the upper limit, the higher the electricity price), accurately matching the individual differences of different resources (such as high electricity price when industrial equipment is under high load, and high electricity price in the later stages of electric vehicle charging). On the one hand, it can improve users' willingness to respond proactively (higher compensation is obtained in high-sensitivity scenarios), and on the other hand, it can avoid the cost waste caused by uniform incentives. According to theoretical calculations, under the same control effect, the incentive cost can be reduced by 15%-25%; (2) Optimization of scheduling improves both efficiency and economy: The DDPG algorithm uses Markov decision process (MDP) modeling to link "system state (historical scheduling quantity) - action (current scheduling decision) - transition (state update) - reward (profit - penalty)" in a closed loop. It can dynamically adapt to changes in grid demand. Compared with traditional linear programming methods, the convergence speed of scheduling strategy solution is improved by more than 30%, meeting the real-time control requirements of the power grid. The multi-objective optimization model (S1) explicitly aims to maximize the total profit of load aggregators. Combined with seven types of constraints (including zeroing out industrial equipment load transfer and meeting electric vehicle charging demand), it can maximize the difference between compensation revenue and incentive costs under the premise of compliance. In practical applications, the total profit of load aggregators can be increased by 10%-20%. (3) Significantly enhanced dispatch compliance: All constraints (upper and lower limits of dispatch volume, duration, interval, etc.) are handled in the MDP “environment”. Through the “adjustment of violations + reward and punishment” mechanism (such as deducting rewards for violations), the problem of “violations caused by constraint conflicts” in traditional schemes can be effectively avoided. The dispatch compliance rate has increased from about 85% of the existing technology to more than 98%, ensuring the reliability of power grid response. (4) Wide adaptability to scenarios: The technical solution is compatible with different types of flexible load resources such as industrial equipment, electric vehicles, and commercial buildings. It can also flexibly adjust the weight of the objective function according to different grid needs (peak-valley regulation, emergency power support) (e.g., in emergency scenarios, the proportion of "response speed" in the reward can be increased), and is suitable for demand-side response scenarios of power systems of different scales such as provincial and municipal levels.

[0086] The core technical solution of this invention revolves around the "three-in-one" approach of "precise incentives + efficient solution + compliant scheduling," and its key components include: (1) Construction logic of multi-objective optimization and control model: with the core objective of "maximizing the total profit (compensation revenue - incentive cost) of load aggregator", and embedding 7 types of targeted constraints (upper and lower limits of resource dispatch volume, total dispatch volume per hour does not exceed grid demand, single dispatch duration / interval, maximum number of dispatches per day, zeroing of industrial equipment load transfer, and meeting the charging demand of electric vehicles), forming a constraint system covering the needs of the three parties of "resources-grid-users"; (2) Design method of differentiated user incentive cost model: Based on the rule that "the closer the scheduling volume is to the upper limit, the greater the loss of user comfort", a quadratic function (Formula 10) is used to establish the relationship between electricity price and scheduling volume. The coefficient a (a>0) is used to ensure that the electricity price increases with the increase of scheduling volume, so as to achieve quantitative adaptation of individual differences. (3) Scheduling solution framework based on DDPG algorithm: The precise correspondence between the four elements of a Markov Decision Process (MDP) and power dispatch (State = Historical dispatch quantity, Action = Current dispatch decision, Transfer = State update, Reward = Profit - Constraint / Penalty); A unified mechanism for handling constraints: Instead of directly embedding constraints in the "state / action" space, compliance management is achieved through "environmental adjustment of non-compliant actions + rewards and penalties"; (4) Design logic of reward function: integrate "subsidy income, incentive cost, and constraint penalty" into a single reward signal to guide the DDPG algorithm to find the optimal balance between "pursuing high profits" and "ensuring compliance".

[0087] A differentiated user incentive cost model for flexible load resources is proposed, which uses a quadratic function to establish the relationship between electricity price and resource scheduling quantity. The electricity price increases as the resource scheduling quantity approaches its upper limit. A flexible load resource scheduling solution based on the DDPG algorithm is proposed. The scheduling problem is modeled as a Markov decision process (MDP). Multiple types of scheduling constraints are handled uniformly in the "environment". Compliance is ensured through "adjustment of violations + reward and punishment". Dynamic optimization of scheduling strategy is achieved through Actor / Critic dual network and soft update of target network. A multi-objective optimization control model that incorporates the above-mentioned incentive model and solution method aims to maximize the total profit of the load aggregator, and includes seven types of constraints such as upper and lower limits of resource scheduling, zeroing of industrial equipment load transfer, and meeting the charging demand of electric vehicles.

[0088] Example 2 Based on the same inventive concept, this invention also provides a load aggregator flexible resource control system based on the DDPG algorithm, comprising: The parameter acquisition module is used to acquire historical operating data of the load aggregator; The electricity price adjustment module is used to dynamically adjust the electricity price based on the loss of user comfort using a pre-built user incentive cost model. The scheme generation module is used to input the historical operating data of the load aggregator and the dynamically adjusted electricity price into the pre-built multi-objective optimization control model, and solve it through the Markov decision process to obtain a flexible load resource scheduling scheme. The control module is used to perform control based on the flexible load resource scheduling scheme; The pre-constructed multi-objective optimization control model is built with the goal of maximizing the total profit of the load aggregator, combined with constraints.

[0089] Preferably, it further includes a model building module, the model building module being used for: Construct an objective function with the goal of maximizing the total profit of the load aggregator; Set constraints for the objective function; The constraints include: upper and lower limits of the schedulable quantity of each resource, the schedulable quantity of all resources not exceeding the grid demand for each set time period, single scheduling duration constraints, scheduling interval constraints, maximum schedulable times during the control period constraints, total schedulable quantity constraints of industrial equipment, and electric vehicle charging quantity constraints.

[0090] Preferably, the objective function is as follows:

[0091] In the formula, The load aggregator's electricity price at time t; This represents the actual response of all adjustable resources at time t; The user incentive electricity price represents the price at time t. This represents the actual response amount of resource i at time t; N represents the ratio of the current scheduling quantity to the available scheduling quantity of the i-th resource at time t; N represents the number of resources; and T represents the control period.

[0092] Preferably, the pre-built user incentive cost model is shown in the following formula:

[0093] In the formula, This represents the user incentive electricity price for the i-th resource; This represents the ratio of the current schedulable quantity to the available schedulable quantity of the i-th resource. Represents the current scheduling amount of the i-th resource; This represents the maximum schedulable quantity of the i-th resource; The coefficient represents the quadratic function of electricity price, and a > 0.

[0094] Preferably, the scheme generation module is specifically used for: Set the constraints in the multi-objective optimization control model to the environment; The scheduling amount of all resources for each set period of time in the historical operation data of the load aggregator is used as the status. The current scheduling amount of each resource per set time period is taken as an action; the process of executing an action, updating the scheduling table, and advancing to the next moment is taken as a transition; the scheduling profit for this set time period is taken as a reward. The state is input into the Actor network, which outputs the action. The state and action are then input into the Critic network, which outputs the Q value. This process is repeated until convergence or a predetermined number of rounds is reached, resulting in a flexible load resource scheduling scheme.

[0095] Preferably, in the scheme generation module, the state is input into the Actor network, which outputs actions; the state and actions are input into the Critic network, which outputs Q-values; this process is repeated until convergence or a predetermined number of rounds is reached to obtain a flexible load resource scheduling scheme. Specific implementation steps include: Step 1: Initialize the Actor network and Critic network, as well as the target network corresponding to the Actor network and Critic network, and initialize the experience replay pool; Step 2: In each step, use an Actor network based on the current state s t Select action a, increase exploration noise, and you get action a with increased exploration noise. t Execute a t Observe the new state s returned by the environment. t+1 and reward r t ,Will Add to the experience replay pool; Step 3: Randomly sample N experience samples from the experience replay pool. The Q-value y of the next state is estimated using a target network consisting of Critic and Actor. i The Critic network is updated using mean squared error loss; the Actor network is updated using policy gradient; and the target networks corresponding to the Actor and Critic networks are updated using a small step size τ of the parameters of the Actor and Critic networks. Step 4: Determine whether convergence has been achieved or the predetermined number of rounds has been reached. If convergence has not been achieved or the predetermined number of rounds has not been reached, return to Step 2. Otherwise, obtain the action that maximizes the output Q, and use the current hourly scheduling amount of each resource corresponding to the action as the flexible load resource scheduling scheme.

[0096] Example 3 like Figure 4 As shown, the present invention also provides an electronic device, which may be a computer device, a microcontroller device, a smart mobile device, etc. The electronic device in this embodiment may include a processor, a memory, a transceiver component, etc. The memory, processor, and transceiver component are connected via a bus; the memory can be used to store executable programs, and an exemplary executable program may include instructions; the processor is used to execute the instructions stored in the memory. The memory can also be used to store data, which can be accessed and / or modified when instructions are executed.

[0097] The processor may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, and it is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the storage medium to realize the corresponding method flow or corresponding function, so as to realize the steps of the load aggregator flexible resource regulation method based on the DDPG algorithm in the above embodiments.

[0098] Example 4 Based on the same inventive concept, this invention also provides a readable storage medium, specifically an electronic device readable storage medium (Memory). This readable storage medium is a memory device within an electronic device used to store programs and data. It is understood that the storage medium here can include both built-in storage media within the electronic device and extended storage media supported by the electronic device. The storage medium provides storage space, which stores the terminal's operating system. Furthermore, this storage space also stores one or more instructions suitable for loading and execution by a processor. These instructions can be one or more executable programs (including program code). It should be noted that the storage medium here can be high-speed RAM or non-volatile memory, such as at least one disk storage device. Loading and executing one or more instructions stored in the storage medium by the processor can implement the steps of the load aggregator flexible resource control method based on the DDPG algorithm in the above embodiments.

[0099] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0100] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0101] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1The function specified in one or more boxes.

[0102] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0103] The above are merely embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of the claims of the present invention pending approval.

Claims

1. A flexible resource regulation method based on the DDPG algorithm for load aggregator, characterized in that, include: Obtain historical operational data from load aggregators; Based on the loss of user comfort, the electricity price is dynamically adjusted through a pre-built user incentive cost model. By substituting the historical operating data of the load aggregator and the dynamically adjusted electricity price into the pre-constructed multi-objective optimization control model, and solving it through the Markov decision process, a flexible load resource scheduling scheme is obtained. Regulation is carried out based on the aforementioned flexible load resource scheduling scheme; The pre-constructed multi-objective optimization control model is built with the goal of maximizing the total profit of the load aggregator, combined with constraints.

2. The method as described in claim 1, characterized in that, The construction of the multi-objective optimization control model includes: Construct an objective function with the goal of maximizing the total profit of the load aggregator; Set constraints for the objective function; The constraints include: upper and lower limits of the schedulable quantity of each resource, the schedulable quantity of all resources not exceeding the grid demand for each set time period, single scheduling duration constraints, scheduling interval constraints, maximum schedulable times during the control period constraints, total schedulable quantity constraints of industrial equipment, and electric vehicle charging quantity constraints.

3. The method as described in claim 1, characterized in that, The objective function is shown in the following equation: In the formula, To the total profit of the load aggregator, The load aggregator's electricity price at time t; This represents the actual response of all adjustable resources at time t; The user incentive electricity price represents the price at time t. This represents the actual response amount of resource i at time t; The ratio of the current scheduled quantity to the available scheduled quantity of the i-th resource at time t; N represents the number of resources; T represents the period of regulation.

4. The method as described in claim 1, characterized in that, The pre-built user incentive cost model is shown in the following equation: In the formula, This represents the user incentive electricity price for the i-th resource; This represents the ratio of the current schedulable quantity to the available schedulable quantity of the i-th resource. Represents the current scheduling amount of the i-th resource; This represents the maximum schedulable quantity of the i-th resource; The coefficient represents the quadratic function of electricity price, and a >

0.

5. The method as described in claim 1, characterized in that, The process involves substituting historical operating data of load aggregators and dynamically adjusted electricity prices into a pre-constructed multi-objective optimization control model, and solving it through a Markov decision process to obtain a flexible load resource scheduling scheme, including: Set the constraints in the multi-objective optimization control model to the environment; The scheduling amount of all resources for each set period of time in the historical operation data of the load aggregator is used as the status. The current scheduling amount of each resource per set time period is taken as an action; the process of executing an action, updating the scheduling table, and advancing to the next moment is taken as a transition; the scheduling profit for this set time period is taken as a reward. The state is input into the Actor network, which outputs the action. The state and action are then input into the Critic network, which outputs the Q value. This process is repeated until convergence or a predetermined number of rounds is reached, resulting in a flexible load resource scheduling scheme.

6. The method as described in claim 5, characterized in that, The process of inputting the state into the Actor network and outputting actions, and then inputting the state and actions into the Critic network and outputting Q-values, is repeated until convergence or a predetermined number of rounds is reached to obtain a flexible load resource scheduling scheme, including: Step 1: Initialize the Actor network and Critic network, as well as the target network corresponding to the Actor network and Critic network, and initialize the experience replay pool; Step 2: In each step, use an Actor network based on the current state s t Select action a, increase exploration noise, and you get action a with increased exploration noise. t Execute a t Observe the new state s returned by the environment. t+1 and reward r t ,Will Add to the experience replay pool; Step 3: Randomly sample N experience samples from the experience replay pool. The Q-value y of the next state is estimated using a target network consisting of Critic and Actor. i The Critic network is updated using mean squared error loss; the Actor network is updated using policy gradient; the target networks corresponding to the Actor and Critic networks are updated using a small step size τ of the parameters of the Actor and Critic networks, where N represents the number of resources. Step 4: Determine whether convergence has been achieved or the predetermined number of rounds has been reached. If convergence has not been achieved or the predetermined number of rounds has not been reached, return to Step 2. Otherwise, obtain the action that maximizes the output Q, and use the current scheduling amount of each resource per set duration corresponding to the action as the flexible load resource scheduling scheme.

7. A load aggregator flexible resource control system based on the DDPG algorithm, characterized in that, include: The parameter acquisition module is used to acquire historical operating data of the load aggregator; The electricity price adjustment module is used to dynamically adjust the electricity price based on the loss of user comfort using a pre-built user incentive cost model. The scheme generation module is used to input the historical operating data of the load aggregator and the dynamically adjusted electricity price into the pre-built multi-objective optimization control model, and solve it through the Markov decision process to obtain a flexible load resource scheduling scheme. The control module is used to perform control based on the flexible load resource scheduling scheme; The pre-constructed multi-objective optimization control model is built with the goal of maximizing the total profit of the load aggregator, combined with constraints.

8. The system as described in claim 7, characterized in that, It also includes a model building module, which is used for: Construct an objective function with the goal of maximizing the total profit of the load aggregator; Set constraints for the objective function; The constraints include: upper and lower limits of the schedulable quantity of each resource, the schedulable quantity of all resources not exceeding the grid demand for each set time period, single scheduling duration constraints, scheduling interval constraints, maximum schedulable times during the control period constraints, total schedulable quantity constraints of industrial equipment, and electric vehicle charging quantity constraints.

9. The system as described in claim 8, characterized in that, The objective function is shown in the following equation: In the formula, To the total profit of the load aggregator, The load aggregator's electricity price at time t; This represents the actual response of all adjustable resources at time t; The user incentive electricity price represents the price at time t. This represents the actual response amount of resource i at time t; The ratio of the current scheduled quantity to the available scheduled quantity of the i-th resource at time t; N represents the number of resources; T represents the period of regulation.

10. The system as described in claim 7, characterized in that, The scheme generation module is specifically used for: Set the constraints in the multi-objective optimization control model to the environment; The scheduling amount of all resources for each set period of time in the historical operation data of the load aggregator is used as the status. The current scheduling amount of each resource per set duration is taken as an action; the process of executing an action, updating the scheduling table, and advancing to the next moment is taken as a transition. The scheduling profit within this set duration will be used as a reward. The state is input into the Actor network, which outputs the action. The state and action are then input into the Critic network, which outputs the Q value. This process is repeated until convergence or a predetermined number of rounds is reached, resulting in a flexible load resource scheduling scheme.

11. The system as described in claim 10, wherein the scheme generation module inputs the state into an Actor network and outputs an action, inputs the state and action into a Critic network and outputs a Q-value, repeating this process until convergence or a predetermined number of rounds is reached to obtain a flexible load resource scheduling scheme, the specific implementation steps of which include: Step 1: Initialize the Actor network and Critic network, as well as the target network corresponding to the Actor network and Critic network, and initialize the experience replay pool; Step 2: In each step, use an Actor network based on the current state s t Select action a, increase exploration noise, and you get action a with increased exploration noise. t Execute a t Observe the new state s returned by the environment. t+1 and reward r t ,Will Add to the experience replay pool; Step 3: Randomly sample N experience samples from the experience replay pool. The Q-value y of the next state is estimated using a target network consisting of Critic and Actor. i The Critic network is updated using mean squared error loss; the Actor network is updated using policy gradient; and the target networks corresponding to the Actor and Critic networks are updated using a small step size τ of the parameters of the Actor and Critic networks. Step 4: Determine whether convergence has been achieved or the predetermined number of rounds has been reached. If convergence has not been achieved or the predetermined number of rounds has not been reached, return to Step 2. Otherwise, obtain the action that maximizes the output Q, and use the current hourly scheduling amount of each resource corresponding to the action as the flexible load resource scheduling scheme.

12. An electronic device, characterized in that, include: At least one processor and memory; The memory and processor are connected via a bus; The memory is used to store one or more programs; When the one or more programs are executed by the at least one processor, the load aggregator flexible resource regulation method based on the DDPG algorithm as described in any one of claims 1 to 6 is implemented.

13. A readable storage medium, characterized in that, It contains an execution program, which, when executed, implements the load aggregator flexible resource regulation method based on the DDPG algorithm as described in any one of claims 1 to 6.