Expert-data dual-driven demand response optimization method considering user self-behavior coupling

By employing an expert-data dual-driven demand response optimization method, and utilizing DQN networks and expert participation evaluators to generate load dispatch instructions, the problems of resource instability and insufficient user response in the electricity market are solved, thereby achieving precise control of flexible loads and maximizing long-term benefits.

CN120633946BActive Publication Date: 2026-04-17CHINA THREE GORGES UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA THREE GORGES UNIV
Filing Date
2025-07-16
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

The instability of supply-side resources and insufficient awareness of demand-side user response in the electricity market make it difficult to balance supply and demand, and existing technologies are insufficient to achieve effective flexible load control and maximize long-term benefits.

Method used

We adopt an expert-data dual-driven demand response optimization method that considers user self-behavior coupling. We generate load scheduling instructions through DQN network and expert participation evaluator, and combine expert-driven and data-driven collaborative training to optimize load reduction strategies, thereby achieving precise control of flexible load resources and maximizing long-term benefits.

Benefits of technology

It improved the robustness and stability of power supply and demand, enhanced the accuracy and reliability of load regulation, reduced trial-and-error costs and manpower burden, and achieved the goal of maximizing load reduction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633946B_ABST
    Figure CN120633946B_ABST
Patent Text Reader

Abstract

The application discloses a method for optimizing demand response of experts-data double driving considering user self-behavior coupling, and relates to the technical field of power system demand response optimization. The application determines the expert intervention or exit state by constructing an expert participation degree evaluator. First, the expert intervention is used to learn the behavior characteristics of each user to demand response, and an expert strategy is constructed. After the expert exits, a neural network is autonomously grown to construct a strategy. The generated strategy is in the form of a signal of "load scheduling instruction value" and is issued to each user. After considering the influence of user behavior self-coupling and expert preference, the expert strategy is appropriately improved. Finally, the data after the user executes the strategy is returned to the "experience replay pool" for the expert strategy constructor / neural network training sampling, so as to form a collaborative cycle training system of expert driving and data driving. The network is trained according to the set round, and the strategy is optimized, that is, the most suitable "load scheduling instruction value" for each user is set, and the maximum total load reduction is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power system demand response optimization technology, and particularly relates to an expert-data dual-drive demand response optimization method that considers user self-behavior coupling. Background Technology

[0002] With increasing global focus on environmental protection and sustainable energy development, the power system is undergoing profound changes. The advancement of smart grid construction provides technical support for the refined management and efficient operation of the power system, but it also places higher demands on demand response optimization methods. In this context, demand response, as a means of regulating flexible loads from the demand side and promoting supply-demand balance, incentivizes users to adjust their electricity consumption behavior, such as reducing electricity consumption during peak hours and increasing it during off-peak hours, thereby achieving peak shaving and valley filling of loads, alleviating power supply shortages, and maximizing long-term benefits. However, the current electricity market still faces many challenges: On the one hand, as a natural resource to be converted, it is difficult to regulate the supply and demand balance on the supply side of the electricity market through technical means. This is because the supply of electricity resources is unstable, resulting in poor controllability and flexibility. On the other hand, regulating the supply and demand balance from the demand side of electricity resources is more flexible. Adjusting electricity consumption from the user's perspective can effectively adapt to the supply and demand balance. However, the majority of users lack awareness and participation in the demand side of the electricity market and have a weak market consciousness, making it difficult for suppliers to issue appropriate load control measures to users from the demand side, thus failing to effectively regulate the supply and demand balance of the electricity market.

[0003] In summary, the current electricity market suffers from unstable energy inputs, limited power resources, and poor controllability and flexibility on the supply side. Furthermore, demand-side user awareness and participation are insufficient, and market consciousness is weak. Therefore, the controllability of demand response in the current electricity market presents significant problems. Summary of the Invention

[0004] To address this issue, this invention proposes an expert-data dual-driven demand response optimization method that considers user self-behavior coupling. It introduces a DQN network with "load dispatch command values" as actions and constructs an expert participation evaluator. By evaluating expert participation, the expert state is determined, and a strategy is generated and sent to users as a load reduction signal. Simultaneously, an expert-driven and data-driven collaborative training network is introduced to continuously optimize the strategy, maximizing load reduction. This method can also accurately regulate flexible loads on the demand side of the electricity market, effectively improving the robustness and stability of power resource supply and demand, thereby optimizing the efficiency of power resource allocation and supporting the efficient operation of the electricity market.

[0005] To solve the above technical problems, the present invention adopts the following technical solution:

[0006] This invention proposes an expert-data dual-driven demand response optimization method that considers user self-behavior coupling. Based on a DQN neural network algorithm with "load dispatch command value" as the action, the method determines expert intervention or withdrawal by evaluating expert participation. First, experts learn the behavioral characteristics of each user in response to demand response and construct expert strategies. After experts withdraw, the neural network autonomously trains and constructs its own strategies. The generated strategies are then issued to each user in the form of "load dispatch command value" signals. The expert strategies are appropriately improved after considering user behavior self-coupling and the influence of expert preferences. Finally, the results of user strategy execution are returned to an "experience replay pool" for expert strategy builder / neural network training sampling, forming a collaborative cyclical training system of expert-driven and data-driven approaches. This system trains the network according to a set number of rounds to optimize strategies, i.e., setting the most suitable "load dispatch command value" for each user, to maximize total load reduction and achieve precise control of flexible load resources and maximize long-term benefits on the demand side of the electricity market.

[0007] Furthermore, the expert-data dual-drive demand response optimization method proposed in this invention, which considers user self-behavior coupling, introduces the "load scheduling instruction value" for each user. This serves as the action layer A (i.e., the policy) for training the neural network. It represents the user. i exist t The load reduction target that the supplier wants to cut is determined by a strategy and then input as a load reduction signal to each user to generate a response. Based on this response, a DQN neural network is constructed to iteratively train the optimal allocation scheme. The output layer of this neural network takes indoor temperature, excitation, user self-influence, and their respective weights as state inputs. The second fully connected layer contains 256 neurons, and the third fully connected layer contains 128 neurons. The output actions are discretized into five levels: 20, 40, 60, 80, and 100, corresponding to five output layer neurons. The user then executes one of these five actions. Q The maximum value is executed, and the model is iteratively optimized through training. Q The value (i.e., the maximum expected total load reduction) is used to plan the most suitable "load dispatch command value" for each user. This aims to maximize the long-term expected load reduction within a limited budget. Due to the different needs of various users... Since the responses differ, the expected value is used to represent the total load reduction of users within a certain period of time. The optimization model constructed is shown below:

[0008]

[0009] in, Expressing expectations; express t Always i Users are requesting a reduction in load; Indicates the user at time t i Past response history. Constraint (a2) states that the "required load reduction" should not exceed the total load reduction.

[0010] The iterative formula for its training is:

[0011]

[0012] in, In order to be in t Time (parameter is) w Given a "load dispatch instruction value" ,from t Total load reduction from the current time to the deadline Given a "load dispatch command value" at time t. The load reduction immediately As a discount factor, To enter t The optimal value for total load reduction from time +1 to the deadline.

[0013] Furthermore, the user behavior self-coupling proposed in this invention, when issuing the generated strategy to users, fully considers the self-herding effect and fatigue effect of users regarding load reduction. It constructs a user behavior self-coupling system framework, continuously incorporating the user's current action into the F-factor simulation, and introducing the concept of user self-influence as the state input of the neural network. This allows the demand response to adapt to user behavior, improving optimization effectiveness. The formula is:

[0014]

[0015] in, User i "Self-influence" is based on users' past response patterns. ( Construct a sequence (i = 1, 2, 3, ..., t). The user's response... It is a random variable that satisfies:

[0016]

[0017] in (·) indicates the Bernoulli distribution. Indicates user i The response probability can be obtained by learning from a neural network.

[0018] Furthermore, the expert-data dual-driven demand response optimization method proposed in this invention, which considers user self-behavior coupling, evaluates expert participation through an expert simulator to construct an expert participation evaluator to determine the expert's involvement / exit status. This evaluator generates participation process parameters in each round of demand response. This leads to the generation of the discrete distribution function of participation. The distribution is sampled to determine whether an expert has intervened or withdrawn; the formula is as follows:

[0019]

[0020] in, Indicates the initial level of expert participation. express t Expert involvement at all times In response to i Users t The sampling time in the distribution function is 0. If the value is 0, the expert intervenes; otherwise, the expert withdraws.

[0021] Furthermore, the cyclic training system proposed in this invention, which returns the results of the user's execution of the strategy for training sampling, incorporates the data after the user executes the strategy into the "experience replay pool" for neural network training sampling each time a strategy is generated and distributed to each user.

[0022] Furthermore, the expert-data dual-driven demand response optimization method proposed in this invention, which considers user self-behavior coupling, first starts with expert experience driving the model's initial development, and then allows the neural network to learn and grow independently. The network training method is divided into expert-driven training and data-driven training based on the expert's intervention and exit states. First, when an expert intervenes, expert strategies are accurately generated by capturing the response characteristics of each user. After the data from user strategy execution is included in the replay pool, the expert-driven training network is started, continuously optimizing the expert strategy by mimicking errors, and the above steps are repeated. When an expert exits, the trained neural network autonomously generates strategies. After the data from user strategy execution is included in the regression pool, the data-driven training network is started, calling historical interaction data stored in the replay pool, achieving self-optimization by minimizing sampling errors, and repeating the above steps to plan the most suitable "load scheduling instruction value" for each user. From an effectiveness standpoint, this dual-driven method ensures efficient convergence of network learning in the early stages due to expert intervention, reducing trial-and-error costs, avoiding resource waste, and achieving rapid strategy optimization. The subsequent self-growth of the neural network effectively reduces human costs and the cognitive burden on experts, showing significant advantages compared to full expert intervention or network self-growth. When generating expert strategies, it incorporates user weights based on prior knowledge. Based on extensive surveys conducted before the demand response project commences and the expert's own experience, it generates expert predicted values ​​for each user's weights regarding different influencing factors.

[0023]

[0024] in, For expert prediction of response characteristic functions. , , and For users respectively i The expert-predicted values ​​of temperature weight, excitation weight, self-influence weight, and load reduction command value weight. This refers to the user response probability predicted by experts. , , , These are the temperature, excitation, self-influence, and load reduction command values, respectively. i Users t The actual value at that moment.

[0025] Simultaneously, capture expert preference factors in demand response actions. , A positive value indicates that the expert is aggressive, while a negative value indicates that the expert is conservative.

[0026] Furthermore, this invention considers the impact of expert preferences on the generation strategy by removing expert-preference action factors from the generated expert actions to reduce errors caused by expert preferences. This is because human individuals are relatively short-sighted, and expert policy builders only focus on immediate rewards; therefore, to generate expert policies without expert preferences, [further measures are needed].

[0027]

[0028] in, This represents the final generated expert action. For expert preference factor action This refers to the ideal expert action after eliminating expert preference factors.

[0029] The present invention adopts the above technical solution, and its progress compared with the prior art is as follows:

[0030] This invention, based on existing technologies, fully considers the self-herding effect and fatigue benefits of user behavior trends, making it more practical. By quantifying the user's self-influence to construct a self-feedback equation, it introduces a user self-coupling mechanism to form a closed-loop feedback system, deeply exploring the potential for flexible adjustment on the demand side. Compared with existing technologies, it has higher accuracy and reliability in controlling the load.

[0031] Meanwhile, given the high trial-and-error costs and low efficiency of rote learning in existing technologies for learning user behavior trends, this invention constructs an expert strategy generation framework based on dynamic expert participation evaluation. Through time-sharing expert intervention and withdrawal, it achieves adaptive collaborative training of expert experience and data-driven approaches, accurately generating strategies. Early expert intervention ensures efficient convergence of network learning, reduces trial-and-error costs, avoids resource waste, and enables rapid strategy optimization. Later, the neural network grows independently, effectively reducing human resources costs and the cognitive burden on experts. Conversely, full expert intervention may lead to serious errors due to expert bias, poor sensitivity to changes in user behavior patterns, and high human resources costs; full network self-growth, on the other hand, leads to prolonged trial-and-error cycles and low training efficiency due to random exploration. Therefore, the dual-drive method of this invention has significant advantages. Attached Figure Description

[0032] Figure 1 This is an illustration of a self-coupling expert-data dual-driven demand response optimization method.

[0033] Figure 2 This is a flowchart of the method of the present invention.

[0034] Figure 3 These are simulation comparison diagrams of the present invention. Detailed Implementation

[0035] The present invention will now be further described with reference to the accompanying drawings. The following embodiments are only used to illustrate the invention more clearly.

[0036] The technical solutions described herein should not be used to limit the scope of protection of this invention.

[0037] The present invention describes a self-coupling expert-data dual-driven demand response optimization method as follows: Figure 1 As shown, it specifically includes the following:

[0038] like Figure 1 As shown, the expert participation evaluator first determines the expert status. When an expert intervenes or withdraws, the expert strategy builder and the neural network generate strategies respectively. These strategies are then sent to the user and executed with a load reduction signal. At this time, the user behavior self-coupling and the influence of expert preferences are taken into account. The results after the strategy execution are returned to the experience pool for training and optimization of the strategy.

[0039] Combination Figure 1 As shown, the following details specific embodiments of the expert-data dual-drive demand response optimization method of the present invention, which considers user self-behavior coupling:

[0040] This invention designs an expert-data dual-driven demand response optimization method that considers user self-behavior coupling, aiming to maximize total load reduction and thereby achieve precise regulation of flexible load resources and maximize long-term benefits on the demand side of the electricity market. The method uses an expert participation evaluator to determine the expert's state, initially generating an expert strategy by the expert's intervention. After considering user behavior self-coupling and the influence of expert preferences, the data after the user executes the strategy is returned to an "experience replay pool" for the expert strategy builder to optimize the strategy. After the expert exits, the neural network autonomously trains to generate a strategy and repeats the above steps, continuously incorporating user responses into the experience pool and optimizing the strategy through training. The DQN neural network optimization model constructed in this invention is as follows:

[0041]

[0042] in, Expressing expectations; This indicates a request to reduce the load. This indicates the user's past response history. Constraint (a2) states that the "required load reduction" should not exceed the total load reduction.

[0043] The iterative formula for its training is:

[0044]

[0045] in, In order to be in t Time (parameter is) w Given a "load dispatch instruction value" ,from t Total load reduction from the current time to the deadline Given a "load dispatch command value" at time t. The load reduction immediately As a discount factor, To enter t The optimal value for total load reduction from time +1 to the deadline.

[0046] The user behavior self-coupling parameter model of the present invention is as follows:

[0047]

[0048] in, User i "Self-influence" is based on users' past response patterns. ( Construct a sequence (i = 1, 2, 3, ..., t). The user's response... It is a random variable that satisfies:

[0049]

[0050] in (·) indicates the Bernoulli distribution. Indicates user i The response probability at time t can be learned by the neural network. During the simulation, the solution... The sigmoid function is shown below:

[0051]

[0052] in, , and They are respectively t Time users i Indoor temperature, motivation and self-influence, This refers to the amount of load that users are required to reduce. This is the maximum amount of load that can be reduced for this user; , , and These are the weights of the corresponding factors. A negative value indicates that the greater the amount of load reduction required from the user, the lower the probability of the user responding.

[0053] The expert participation evaluation model of this invention is as follows:

[0054]

[0055] in, Indicates the initial level of expert participation. express t Expert involvement at all times In response to i Users t The sampling time in the distribution function is 0. If the value is 0, the expert intervenes; otherwise, the expert withdraws.

[0056] The model in this invention for capturing user behavior features and then generating expert strategies is as follows:

[0057]

[0058] in, For expert prediction of response characteristic functions. , , and For users respectively i The expert-predicted values ​​of temperature weight, excitation weight, self-influence weight, and load reduction command value weight. This refers to the user response probability predicted by experts. , , , These are the temperature, excitation, self-influence, and load reduction command values, respectively. i Users t The actual value at that moment.

[0059] The strategy optimization model for expert preferences in this invention is as follows:

[0060]

[0061] in, This represents the final generated expert action. For expert preference factor action This refers to the ideal expert action after eliminating expert preference factors.

[0062] This invention uses a solver to solve the above model, obtaining a solution that maximizes the total load reduction. The relevant parameters in the solver are set as follows: total community members N = 100, demand response rounds T = 200, and load reduction values ​​for each user range from 0% to 100%. ,in The initial self-influence of each user is 3kW. =0; temperature weight, excitation weight User weights are taken from uniformly distributed interval values. The value range is [-1.5, -0.5]. The neural network consists of four layers. The number of neurons in the input layer is equal to the state space dimension. The first hidden layer (fully connected layer) contains 256 neurons, the second hidden layer (fully connected layer) contains 128 neurons, and the output layer contains five neurons, corresponding to five action levels: 20%. 40% 60% 80% 100% The user then executes the five given Q-values ​​with the highest value; reinforcement learning discount factor. The value is 0.9; the expert preference factor action... -20% 20% A positive value indicates aggressive experts, while a negative value indicates conservative experts; this represents the initial level of expert participation. It is 0.2.

[0063] The simulation results of this invention are as follows:

[0064] As shown in Table a1, this represents the average aggregated load value per round when actions are randomly given within 200 rounds, experts are completely withdrawn, experts are fully involved, and the process is driven by both experts and data.

[0065] Table a1

[0066]

[0067]

[0068]

[0069]

[0070]

[0071]

[0072]

[0073] Appendix Figure 3 This is a line graph showing the average load aggregation value over 200 rounds under four scenarios: random action, full expert intervention, full expert withdrawal, and expert-data dual-drive.

[0074] As shown in Table a2, it represents the average aggregated load value over 200 rounds under four scenarios: random action, full expert intervention, full expert withdrawal, and expert-data dual-drive.

[0075] Table a2

[0076]

[0077] When actions are randomly assigned, the line chart shows significant fluctuations. When experts are involved throughout the process, the line chart data tends to stabilize with smaller fluctuations due to the precise control provided by the experts' experience. However, long-term expert involvement can lead to a series of problems, such as expert bias, automatic outlier removal, and poor sensitivity to changes in user habits. When experts are not involved throughout the process, the line chart data shows a clear upward trend. This is because the blindness of data-driven approaches makes the data more susceptible to noise. When a dual expert-data approach is adopted, the line chart data shows a high degree of fit with the data from the period of full expert involvement in the early stages and a high degree of fit with the data from the period of full expert withdrawal in the later stages, showing an overall trend towards stability. This is not only due to the precise fitting of the optimal curve by expert involvement but also avoids problems such as expert bias, achieving the goal of optimizing the average load aggregation value.

[0078] The results show that when a load reduction command value is randomly given, the average load aggregation value is 84.77 kW; when the entire process is driven by data and experts are not involved, the average load aggregation value is 105.42 kW; when experts are involved throughout the process, the average load aggregation value is 120.36 kW; and when both experts and data are involved, the average load aggregation value is 154.96 kW, which is 147.0% of the data-driven approach and 128.7% of the expert-involved approach. This demonstrates that the self-coupling expert-data dual-driven demand response optimization method of this invention has significant advantages.

[0079] Therefore, the expert-data dual-driven demand response optimization method proposed in this invention, which considers self-coupling, achieves the goal of maximizing the total load reduction.

[0080] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. An expert-data dual-driven demand response optimization method considering user self-behavior coupling, characterized in that, The process includes the following steps: building an expert strategy builder, which generates load reduction instructions for each user based on preset rules and expert experience, while taking into account the influence of expert preferences; firstly, experts intervene to learn the behavioral characteristics of each user in responding to demand and build expert strategies for allocating load reduction instructions to each user. After the experts leave, the DQN neural network will be used to construct the load reduction and allocation strategy. Load reduction and allocation strategies are based on load scheduling command values. The signal is sent to each user. After considering the influence of user behavior self-coupling and expert preferences, the load reduction and allocation strategy is optimized. Finally, the response results of users after implementing the load reduction and allocation strategy are returned to the "experience replay pool" for DQN neural network training sampling. This is integrated to form a collaborative cyclical training system of expert-driven and data-driven approaches. According to the set training rounds, the DQN neural network is trained to optimize the strategy, that is, to set the load scheduling command value that is most suitable for each user. The load dispatch command value This represents the amount of load that the supplier wants to reduce for user i at time t. The user behavior self-coupling refers to the self-coupling effect of the user due to historical response behavior. The consideration of user behavior self-coupling is specifically achieved by constructing a user behavior self-coupling system framework. Based on the user's past response data, the user's self-influence is dynamically calculated through an iterative update formula, and this calculation is used as the state input of the DQN neural network to further optimize the load reduction and allocation strategy. The iterative update formula is as follows: in, This represents the self-influence of user i at time t+1, which is based on the user's past response history. to Build; The consideration of the influence of expert preferences refers to removing expert preference action factors from the generated load reduction and allocation strategy to reduce the error caused by expert preferences. The expert strategy builder only focuses on immediate rewards and generates expert strategies without expert preferences. in, This represents the expert strategy that ultimately allocates load reduction command values ​​to each user. For expert preference factor actions, This refers to the expert strategy for assigning load reduction command values ​​to each user after removing expert preference factors.

2. The expert-data dual-driven demand response optimization method considering user self-behavior coupling as described in claim 1, characterized in that: Introduce load scheduling command values ​​for each user As the action layer A of the neural network, after the strategy is constructed, it is used as a load reduction signal input to each user to generate a response. Based on the response, a DQN neural network is constructed. The input layer takes indoor temperature, excitation, and user self-influence as state inputs, and the output layer predicts different... The value corresponds to the estimated total load reduction. The value is discretized into five levels, specifically 20. ,in , The upper limit for load reduction is set; users execute the load reduction allocation strategy that maximizes the estimated total load reduction value, and generate responses for network training. Through continuous training, the most suitable load scheduling command values ​​for each user are planned. Because different users have different preferences Given different responses, the expected total load reduction for users within a certain period of time is used to achieve the following overall optimization objective: maximizing the expected total load reduction. in, Expressing expectations; express t Always i Users are requesting a reduction in load; Indicates the user at time t i Historical response data shows that constraint (a4) indicates that the "required load reduction" does not exceed the total load reduction; where T represents the total number of demand response rounds; to approximate the overall optimization objective, the training objective of the DQN neural network is to maximize the action value function, i.e., to maximize the estimated total load reduction. ; The training iteration formula is: in, In order to be in t Time and parameters are w、 Given a load dispatch instruction value At that time, from t The estimated total load reduction from the current time to the deadline. Given a load dispatch instruction value at time t. The load reduction immediately afterwards This is the discount factor.

3. The expert-data dual-driven demand response optimization method considering user self-behavior coupling as described in claim 2, characterized in that, An expert participation evaluator is constructed by assessing expert engagement through an expert simulator to determine expert involvement / exit status. This evaluator generates participation process parameters in each round of demand response. This leads to the generation of the discrete distribution function of participation. The distribution is sampled to determine whether an expert has intervened or withdrawn; the formula is as follows: in, (·) represents the Bernoulli distribution. Indicates the initial level of expert participation. express t Expert involvement at all times In response to i Users t The sampling time in the distribution function is 0. If the value is 0, the expert intervenes; otherwise, the expert withdraws.

4. The expert-data dual-driven demand response optimization method considering user self-behavior coupling as described in claim 3, characterized in that, The response results after users execute the load reduction and allocation strategy are returned to the cyclic training system for training sampling. Each time a load reduction and allocation strategy is generated and distributed to each user, the data after users execute the load reduction and allocation strategy is included in the "experience replay pool" for neural network training sampling.

5. The expert-data dual-driven demand response optimization method considering user self-behavior coupling as described in claim 4, characterized in that, Based on the expert involvement and exit status, the training methods of DQN neural networks are divided into expert-driven training and data-driven training. When experts are involved, an expert policy builder first generates expert policies based on preset rules and expert experience. Specifically, this includes: by incorporating user weights from prior knowledge, and based on extensive surveys conducted before the demand response project and the expert's own experience, generating expert prediction values ​​for the weights of each user regarding different influencing factors. in, For expert prediction of response characteristic function, , , and For users respectively i The expert-predicted values ​​of temperature weight, excitation weight, self-influence weight, and load reduction command value weight. This refers to the user response probability predicted by experts. , , , These are the temperature, excitation, self-influence, and load reduction command values, respectively. i Users t The actual value at any given moment, while simultaneously capturing expert preference factors in the demand response. , A positive value indicates that the expert is aggressive, while a negative value indicates that the expert is conservative; When the expert exits, the DQN neural network autonomously generates a strategy to allocate load reduction instruction values ​​to each user. After the response data of the users after executing the load reduction allocation strategy is included in the regression pool, the data-driven training network is started. The historical interaction data stored in the replay pool is called to achieve self-optimization by minimizing the sampling error. The above steps are repeated to plan the most suitable load scheduling instruction value for each user.

Citation Information

Patent Citations

  • Strategy protection defense method for deep reinforcement learning

    CN113392396A

  • Intelligent optimization method for power grid safe operation strategy based on deep reinforcement learning

    CN114048903A