Expert-data dual-drive demand response optimization method considering self-coupling of user

By introducing a collaborative training system of DQN network and expert participation evaluator, the demand response strategy coupled with user behavior is optimized, which solves the problem of supply and demand balance regulation in the power market, maximizes load reduction and improves resource allocation efficiency.

CN120633946AActive Publication Date: 2025-09-12CHINA THREE GORGES UNIV
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510978980.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-09-12
Estimated Expiration
2045-07-16

AI Technical Summary

Technical Problem

In the electricity market, the supply-side energy input is unstable, the supply flexibility is poor, and the demand-side users lack awareness and participation in response, which makes it difficult to regulate the supply and demand balance. Existing technologies make it difficult to achieve effective demand response optimization.

Method used

An expert-data dual-driven demand response optimization method considering user behavior coupling is adopted. By introducing the DQN network and expert participation evaluator, load scheduling instruction values ​​are generated. Combining the user behavior self-coupling system and expert preferences, a collaborative training system is constructed to optimize the load reduction strategy.

Benefits of technology

It has achieved precise regulation of flexible loads on the demand side of the electricity market and maximized long-term benefits, improved the robustness and stability of power resource allocation, and reduced trial and error costs and manpower burden.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633946A_ABST
    Figure CN120633946A_ABST
Patent Text Reader

Abstract

The invention discloses an expert-data dual-drive demand response optimization method considering user self-coupling, and relates to the technical field of power system demand response optimization. According to the method, an expert participation degree evaluator is constructed to judge an expert intervention or exit state, firstly, an expert intervenes to learn behavior characteristics of each user responding to a demand, and an expert strategy is constructed; and after the expert exits, the neural network autonomously grows to construct a strategy. The generated strategy is issued to each user in a'load scheduling instruction value 'signal form, the expert strategy is properly improved after the influence of user behavior self-coupling and expert preference is considered, and finally data after the user executes the strategy is returned to an'experience playback pool' for training and sampling by an expert strategy builder / neural network. The expert-driven and data-driven collaborative cyclic training system is formed through fusion, and a strategy is optimized according to a set round training network, that is, a load scheduling instruction value which is most adaptive to each user is set, so that the purpose of maximizing total load reduction is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of power system demand response optimization, and in particular relates to an expert-data dual-driven demand response optimization method considering user behavior coupling. Background Art

[0002] With growing global attention to environmental protection and sustainable energy development, the power system is undergoing profound changes. The advancement of smart grid construction provides technical support for refined management and efficient operation of power systems, but it also places higher demands on demand response optimization methods. In this context, demand response, as a means of responding to and regulating flexible loads from the demand side and promoting supply and demand balance, incentivizes users to adjust their electricity consumption, such as reducing electricity consumption during peak hours and increasing it during off-peak hours, to achieve load shaving, alleviate power supply constraints, and maximize long-term benefits. However, today's electricity market still faces many challenges: on the one hand, as a natural resource to be converted, it is difficult to regulate the supply and demand balance through technical means on the supply side of the electricity market. This is because the supply of electricity resources is unstable, and therefore there are practical problems such as poor controllable flexibility; on the other hand, the controllable flexibility of regulating the supply and demand balance from the demand side of electricity resources is relatively strong, and adjusting electricity consumption from the user's perspective can effectively adapt to the supply and demand balance. However, the majority of users lack awareness and participation in the response of the demand side of the electricity market and have weak market awareness, which makes it difficult for suppliers to issue appropriate load control quantities to users from the demand side, and unable to effectively regulate the supply and demand balance of the electricity market.

[0003] Overall, the current power market faces unstable energy inputs on the supply side, limited power resources, and poor controllability. Furthermore, on the demand side, users lack awareness and participation in demand response, and market awareness is weak. Consequently, significant controllability issues exist in the current power market regarding demand response. Summary of the Invention

[0004] In order to solve this problem, the present invention proposes an expert-data dual-driven demand response optimization method that considers the coupling of user behavior. It introduces a DQN network with "load dispatch instruction value" as the action and constructs an expert participation evaluator. The expert status is determined by evaluating the expert participation, and then a strategy is generated. The strategy is sent to the user as a load reduction signal. At the same time, the expert-driven and data-driven collaborative training networks are introduced to continuously optimize the strategy to achieve the goal of maximizing load reduction. It can also accurately regulate the flexible load on the demand side of the power market, effectively improve the robustness and stability of the supply and demand of power resources, thereby optimizing the efficiency of power resource allocation and supporting the efficient operation of the power market.

[0005] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0006] This paper proposes an expert-data dual-driven demand response optimization method that considers user behavior coupling. Based on a DQN neural network algorithm with "load dispatch command values" as its action, the method determines whether experts are engaged or disengaged by assessing their participation. First, experts engage to learn the behavioral characteristics of each user in demand response and construct an expert strategy. After the experts withdraw, the neural network autonomously grows and constructs a strategy. The generated strategy is then issued to each user in the form of a "load dispatch command value" signal. The expert strategy is then appropriately refined by considering the influence of user behavior coupling and expert preferences. Finally, the results of user strategy execution are returned to an "experience replay pool" for training and sampling by the expert strategy builder / neural network. This system integrates expert-driven and data-driven collaborative cyclic training. The network is trained according to a set number of rounds to optimize the strategy, specifically setting the "load dispatch command value" that best suits each user, achieving the goal of maximizing total load reduction. This enables precise regulation of flexible load resources and maximizes long-term benefits on the demand side of the power market.

[0007] Furthermore, the expert-data dual-driven demand response optimization method proposed in this invention, which considers the coupling of user behavior, introduces the “load dispatch instruction value” ΔP of each user. i,t As the action layer A (i.e., strategy) for neural network training. It represents the amount of load that the supplier wants to reduce for user i at time t. After constructing the strategy, it is input as a load reduction signal to each user to generate a response. Based on this response, a DQN neural network for iterative training of the optimal allocation plan is constructed: the output layer of the neural network uses indoor temperature, incentives, user self-influence and their respective weights as state inputs. The second fully connected layer contains 256 neurons, and the third fully connected layer contains 128 neurons. The output action of the output layer is discretized into five levels of 20, 40, 60, 80, and 100, corresponding to the five output layer neurons. The user executes the five given Q values ​​with the largest execution, and iteratively optimizes the Q value (i.e., the expected maximum value of total load reduction) through the training model to plan the most suitable "load scheduling instruction value" ΔP for each user. i,t , to achieve the goal of maximizing the long-term expected load reduction within a limited budget. i,t The responses are different, so the expectation is used to represent the total load reduction of users in a certain period of time. The constructed optimization model is as follows:

[0008]

[0009]

[0010] in, represents expectation; ΔP i,t represents the load reduction required for user i at time t; x i,trepresents the past response of user i at time t. Constraint (a2) indicates that the “required load reduction” should not exceed the total load reduction.

[0011] The iterative formula for its training is:

[0012]

[0013] in, Given the “load dispatch instruction value” a at time t (parameter is w) t , the total load reduction from time t to the deadline, r t At time t, given the “load dispatch instruction value” a t The load reduction immediately, γ is the discount factor, It is the optimal value of total load reduction from time t+1 to the deadline.

[0014] Furthermore, the proposed method of incorporating user behavior autocoupling into the generated strategy when it is delivered to users fully considers the user's self-herding and fatigue effects in response to load reduction. This system framework builds a user behavior autocoupling system, continuously incorporating the user's current actions into the F factor for simulation, and introducing the concept of user self-influence as the state input of the neural network, allowing demand response to adapt to user behavior and improve optimization results. Its formula is:

[0015]

[0016] Among them, F i,t is the “self-influence” of user i, which is based on the user’s past responses x i,τ (τ=1,2,3,…,t) is constructed. And the user’s response x t is a random variable that satisfies:

[0017]

[0018] in represents Bernoulli distribution. p i,t It represents the response probability of user i, and its value can be learned by the neural network.

[0019] Furthermore, the expert-data dual-driven demand response optimization method proposed in this invention, which considers the coupling of user behavior, evaluates the expert participation through an expert simulator, thereby constructing an expert participation evaluator to determine the expert's involvement / exit status. The evaluator generates a participation construction process parameter ε in each round of demand response. t , and then generate the participation discrete distribution function Sampling from this distribution and then determining the expert's intervention or exit status is done using the following formula:

[0020]

[0021] Among them, ε0 represents the initial participation of experts, ε t represents the expert participation at time t, z i,t is the sampling of user i in the distribution function at time t. If the value is 0, the expert intervenes, otherwise the expert exits.

[0022] Furthermore, the present invention proposes a cyclic training system that returns the results after the user executes the strategy for training sampling. Every time a strategy is generated and distributed to each user, the data after the user executes the strategy is included in the "experience replay pool" for neural network training sampling.

[0023] Furthermore, the proposed expert-data dual-driven demand response optimization method, which considers user behavior coupling, begins with expert experience driving the model, followed by self-learning and growth of the neural network. Network training is divided into expert-driven and data-driven training, depending on the status of expert intervention and withdrawal. First, when an expert intervenes, the network captures the response characteristics of each user to accurately generate an expert strategy. After the data after the user executes the strategy is added to the replay pool, the expert-driven training network is activated, continuously optimizing the expert strategy by imitating the error, and repeating these steps. When the expert withdraws, the trained neural network autonomously generates a strategy. After the data after the user executes the strategy is added to the regression pool, the data-driven training network is activated, using the historical interaction data stored in the replay pool to achieve self-optimization by minimizing sampling error. These steps are repeated to determine the optimal "load dispatch command value" for each user. This dual-driven approach effectively ensures efficient convergence of network learning, reduces trial-and-error costs, avoids resource waste, and achieves rapid strategy optimization. Later, the neural network self-grows, effectively reducing labor costs and the cognitive burden of experts, offering significant advantages over full expert intervention or self-growth of the network. When generating expert strategies, it incorporates user weights based on prior knowledge, and based on extensive surveys before the demand response project and the experts' own experience and knowledge, generates expert prediction values ​​for each user's weights on different influencing factors, namely:

[0024]

[0025] Among them, G(x) is the expert prediction response characteristic function. and are the expert prediction values ​​of temperature weight, incentive weight, self-influence weight and load reduction instruction value weight for user i, is the user response probability predicted by the expert, μ i,t , χ i,t 、F i,t , ΔP i,tare the actual values ​​of temperature, incentive, self-influence, and load reduction instruction value for user i at time t.

[0026] At the same time, capturing expert preference factor actions in demand response A positive value indicates that the expert is aggressive, and a negative value indicates that the expert is conservative.

[0027] Furthermore, the present invention considers the impact of expert bias on the generated strategy. It removes the expert bias action factor from the generated expert actions to reduce the error caused by expert bias. This is because human individuals are relatively short-sighted. The expert strategy builder only focuses on immediate rewards. In order to generate an expert strategy without expert bias,

[0028]

[0029] in, represents the final generated expert action, is the expert preference factor action is the ideal expert action after removing the expert preference factor.

[0030] The present invention adopts the above technical solution, and compared with the prior art, the improvements are:

[0031] Based on the existing technology, the present invention fully considers the self-herding effect and fatigue effect of user behavior trends, is more practical, constructs a self-feedback equation by quantifying the user's self-influence, introduces a user self-coupling mechanism to form a closed-loop feedback system, deeply explores the flexible adjustment potential on the demand side, and has higher accuracy and reliability in load regulation than the existing technology.

[0032] At the same time, in view of the high trial and error cost of learning user behavior trends and the low efficiency of mechanical learning in the existing technology, the present invention constructs an expert strategy generation framework based on dynamic expert participation evaluation, and realizes adaptive collaborative training driven by expert experience and data through time-sharing expert intervention and exit, and accurately generates strategies. Early expert intervention can ensure the efficient convergence of network learning, reduce trial and error costs, avoid waste of resources, and achieve rapid optimization of strategies; in the later stage, the neural network will self-grow, which effectively reduces labor costs and expert cognitive burden. In contrast, full expert intervention may lead to serious errors caused by expert preferences, poor sensitivity to changes in user behavior patterns, and high labor costs; full network self-growth will extend the trial and error cycle due to random exploration, and the training efficiency is low. Therefore, there are significant advantages in adopting the dual-drive method of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 This diagram illustrates the expert-data dual-driven demand response optimization method considering self-coupling.

[0034] Figure 2 It is a flow chart of the method of the present invention.

[0035] Figure 3 It is a simulation comparison diagram of the present invention. DETAILED DESCRIPTION

[0036] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention.

[0037] The expert-data dual-driven demand response optimization method designed by the present invention considering self-coupling is described as follows: Figure 1 As shown, specifically including the following:

[0038] like Figure 1 As shown in the figure, the expert participation evaluator first determines the expert status. When the expert intervenes and exits, the expert strategy builder and neural network generate strategies respectively, which are sent to the user as load reduction signals for execution. At this time, the influence of user behavior self-coupling and expert preference is taken into account, and the results after executing the strategy are returned to the experience pool to train the optimization strategy.

[0039] Combine Figure 1 As shown, the specific embodiment of the expert-data dual-driven demand response optimization method considering user behavior coupling of the present invention is described in detail below:

[0040] The present invention designs an expert-data dual-driven demand response optimization method that takes into account the coupling of user behavior, in order to maximize the total load reduction, and then achieve precise regulation of flexible load resources and maximize long-term benefits on the demand side of the power market. The expert status is determined by the expert participation evaluator, and the expert intervenes to generate an expert strategy. After considering the influence of user behavior self-coupling and expert preferences, the data after the user executes the strategy is returned to the "experience replay pool" for the expert strategy builder to optimize the strategy; after the expert exits, the neural network self-grows to generate a strategy and repeat the above steps, continuously incorporating the user's response into the experience pool and optimizing the strategy through training. Among them, the DQN neural network optimization model constructed by the present invention is:

[0041]

[0042] in, represents expectation; ΔP i,t Indicates the requirement to reduce load; x i,t represents the user's past response. Constraint (a2) indicates that the "required load reduction" should not exceed the total load reduction.

[0043] The iterative formula for its training is:

[0044]

[0045] in, Given the “load dispatch instruction value” a at time t (parameter is w) t , the total load reduction from time t to the deadline, r t At time t, given the “load dispatch instruction value” a t The load reduction immediately, γ is the discount factor, It is the optimal value of total load reduction from time t+1 to the deadline.

[0046] The user behavior self-coupling parameter model of the present invention is:

[0047]

[0048] Among them, F i,t is the “self-influence” of user i, which is based on the user’s past responses x i,τ (τ=1,2,3,…,t) is constructed. And the user’s response x t is a random variable that satisfies:

[0049]

[0050] in represents Bernoulli distribution. p i,t represents the response probability of user i at time t, and its value can be learned by the neural network. In the simulation process, solve p i,t The sigmoid function is as follows:

[0051]

[0052] Among them, μ i,t , χ i,t and F i,t are the indoor temperature, incentive and self-influence of user i at time t, ΔP i,t is the amount of load that users are required to reduce. is the maximum value of the load that can be reduced for this user; α i , β i ,ω i and θ i are the weights of the corresponding factors. i A negative value indicates that the greater the load reduction required of the user, the smaller the user response probability.

[0053] The expert participation evaluator model of the present invention is:

[0054]

[0055] Among them, ε0 represents the initial participation of experts, ε t represents the expert participation at time t, z i,t is the sampling of user i in the distribution function at time t. If the value is 0, the expert intervenes, otherwise the expert exits.

[0056] The model used in this invention to capture user behavior characteristics and generate expert strategies is:

[0057]

[0058] Among them, G(x) is the expert prediction response characteristic function. and are the expert prediction values ​​of temperature weight, incentive weight, self-influence weight and load reduction instruction value weight for user i, is the user response probability predicted by the expert, μ i,t , χ i,t 、F i,t , ΔP i,t are the actual values ​​of temperature, incentive, self-influence, and load reduction instruction value for user i at time t.

[0059] The strategy optimization model for expert preference in this invention is:

[0060]

[0061] in, represents the final generated expert action, is the expert preference factor action is the ideal expert action after removing the expert preference factor.

[0062] The present invention uses a solver to solve the above model and obtains a solution that achieves the maximum total load reduction. The relevant parameters in the solver are set as follows: the total number of community members N is 100, the demand response round T is 200, and the load reduction value of each user is in is 3kW, and the initial self-influence of each user is F i,0 is 0; temperature weight, incentive weight α i , β i The user weights are taken from the uniformly distributed interval values, α i The value range is [-1.5,-0.5], β i The value range is [0.4, 0.9]. For the neural network, there are four layers in total. The number of neurons in the input layer is the dimension of the state space. The first hidden layer (fully connected layer) contains 256 neurons, the second hidden layer (fully connected layer) contains 128 neurons, and the output layer contains five neurons, corresponding to the five levels of action: The user executes the action with the largest Q value of the five given actions; the reinforcement learning discount factor γ is 0.9; the expert preference factor action A positive value indicates that the expert is aggressive, and a negative value indicates that the expert is conservative. The initial expert participation ε0 is 0.2.

[0063] The simulation results of the present invention are as follows:

[0064] As shown in Table a1, the average load aggregation value in each round is shown when actions are randomly given, experts are completely absent, experts are fully involved, and experts and data are dual-driven within 200 rounds:

[0065] Table a1

[0066]

[0067]

[0068]

[0069]

[0070]

[0071]

[0072]

[0073]

[0074] Attachment Figure 3 A line graph showing the average load aggregation value within 200 rounds under four conditions: random action, full expert intervention, full expert withdrawal, and expert-data dual drive.

[0075] Table a2 shows the average load aggregation values ​​within 200 rounds for four scenarios: random action, full expert intervention, full expert withdrawal, and expert-data dual drive.

[0076] Table a2

[0077] Randomly given actions Data-driven throughout the process Expert-driven throughout the process Expert-Data Dual Drive Expert participation round / round 0 0 200 78 Average aggregate load / kW 84.77 105.42 120.36 154.96

[0078] When actions are given randomly, the line graph fluctuates greatly; when experts are introduced to intervene throughout the process, the line graph data tends to be stable due to the precise regulation of expert experience, and the fluctuation range is small, but long-term expert intervention will lead to a series of problems such as expert preference, automatic elimination of outliers, and poor sensitivity to changes in user habits; when experts withdraw from the process, the line graph data has an obvious upward trend. This is because the blindness of data-driven causes the data to be greatly affected by noise; when expert-data dual drive is introduced, the line graph data has a higher degree of fit with the data of full expert intervention in the early stage, and a higher degree of fit with the data of full expert withdrawal in the later stage, and tends to be stable overall. Not only does the expert intervention accurately fit the optimal curve, but it also avoids problems such as expert preference, and achieves the goal of optimizing the average load aggregation value.

[0079] The results show that when the load reduction instruction value is randomly given, the average load aggregation value is 84.77kW; when the full expert exits and the data is driven by self, the average load aggregation value is 105.42kW; when the full expert intervenes, the average load aggregation value is 120.36kW; when the expert and data are dual-driven, the average load aggregation value is 154.96kW, which is 147.0% of the full data drive and 128.7% of the full expert intervention. It can be seen that the expert-data dual-driven demand response optimization method considering self-coupling of the present invention has significant advantages.

[0080] It can be seen that the expert-data dual-driven demand response optimization method considering self-coupling proposed in the present invention achieves the goal of maximizing the total load reduction.

[0081] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. An expert-data dual-driven demand response optimization method considering user behavior coupling is characterized by: The method includes the following steps: first, experts intervene to learn the behavioral characteristics of each user in response to demand, and construct an expert strategy for assigning load reduction instructions to each user; after the experts exit, the neural network grows autonomously and constructs a load reduction allocation strategy. The load reduction allocation strategy is based on the "load dispatch instruction value" ΔP i,t The system sends a signal to each user in the form of a signal. After considering the influence of user behavior autocoupling and expert preference, it optimizes the load reduction allocation strategy. Finally, the response results after the user executes the load reduction allocation strategy are returned to the "experience replay pool" for expert strategy builder / neural network training sampling, and a collaborative cycle training system driven by experts and data is formed. The network is trained according to the set rounds to optimize the strategy, that is, the "load scheduling instruction value" ΔP that is most suitable for each user is set. i,t , "load dispatch instruction value" ΔP i,t represents the load that the supplier wants to reduce for user i at time t.

2. The expert-data dual-driven demand response optimization method considering user behavior coupling according to claim 1 is characterized by: Introducing the "load dispatch instruction value" ΔP of each user i,t As the action layer A for neural network training, after constructing the strategy, it is input as a load reduction signal to each user to generate a response. Based on the response, a DQN neural network is constructed to iteratively train the optimal allocation plan. The input layer uses indoor temperature, incentives, user self-influence and their respective weights as state inputs. The output layer action is discretized into five levels of 20*n (n=1,2…5). The user executes the load reduction allocation strategy with the largest Q value, and a response is generated again for network training. Through continuous training and optimization of the Q value, the most suitable "load scheduling instruction value" ΔP for each user is planned. i,t , because each user has different ΔP i,t The response is different. The expectation is used to represent the total load reduction of users in a certain period of time. The constructed optimization model is as follows: in, represents expectation; ΔP i,t represents the load reduction requirement for user i at time t; x i,t represents the past response of user i at time t, and constraint (a2) indicates that the "required load reduction" does not exceed the total load reduction; The iterative formula for training is: in, At time t, the parameter w is given as the "load dispatch instruction value" a t , the total load reduction from time t to the deadline, r t At time t, given the "load dispatch instruction value" a t The load reduction immediately, γ is the discount factor, It is the optimal value of total load reduction from time t+1 to the deadline.

3. The expert-data dual-driven demand response optimization method considering user behavior coupling according to claim 2 is characterized in that: When the generated load reduction allocation strategy is delivered to users, the user behavior autocoupling is incorporated. Considering the self-herding effect and fatigue effect of users on the load reduction amount, a user behavior autocoupling system framework is constructed. Based on the user's previous response data, the concept of user self-influence is introduced as the state input of the neural network to further optimize the load reduction allocation strategy. The formula is: Among them, F i,t is the "self-influence" of user i, which is based on the user's past responses x i,τ (τ=1,2,3,…,t) is constructed, and the user’s response x t is a random variable that satisfies: in represents Bernoulli distribution, p i,t Represents the response probability of user i, and its value is learned by the neural network.

4. The expert-data dual-driven demand response optimization method considering user behavior coupling according to claim 3 is characterized in that: The expert participation is evaluated by the expert simulator, and an expert participation evaluator is constructed to determine the expert involvement / exit status. The expert participation evaluator generates a participation construction process parameter ε in each round of demand response. t , and then generate the participation discrete distribution function Sampling from this distribution and then determining the expert's intervention or exit status is done using the following formula: Among them, ε0 represents the initial participation of experts, ε t represents the expert participation at time t, z i,t is the sampling of user i in the distribution function at time t. If the value is 0, the expert intervenes, otherwise the expert exits.

5. The expert-data dual-driven demand response optimization method considering user behavior coupling according to claim 4 is characterized in that: The response results after the user executes the load reduction allocation strategy are returned to the cyclic training system for training sampling. Each time the load reduction allocation strategy is generated and distributed to each user, the data after the user executes the load reduction allocation strategy is included in the "experience replay pool" for neural network training sampling.

6. The expert-data dual-driven demand response optimization method considering user behavior coupling according to claim 5 is characterized in that: Based on the status of expert intervention and exit, the network training method is divided into expert-driven training and data-driven training. First, when the expert intervenes, the expert strategy for assigning load reduction instructions to each user is generated by capturing the response characteristics of each user. After the response data after the user executes the load reduction allocation strategy is added to the replay pool, the expert-driven training network is started. That is, the expert strategy is continuously optimized by the imitation error method, and the above steps are repeated. When the expert exits, the trained neural network autonomously generates a strategy for assigning load reduction instructions to each user. After the response data after the user executes the load reduction allocation strategy is added to the regression pool, the data-driven training network is started, calling the historical interaction data stored in the replay pool, that is, the method of minimizing sampling error to achieve self-optimization, and repeating the above steps to plan the most suitable "load dispatch instruction value" for each user. When generating the load reduction allocation strategy, by incorporating user weights based on prior knowledge, based on extensive surveys before the demand response project and the expert's own experience and knowledge, expert prediction values ​​are generated for each user's weights of different influencing factors, namely: Among them, G(x) is the expert prediction response characteristic function, and are the expert prediction values ​​of temperature weight, incentive weight, self-influence weight and load reduction instruction value weight for user i, is the user response probability predicted by the expert, μ i,t , χ i,t 、F i,t , ΔP i,t are the actual values ​​of temperature, incentive, self-influence, and load reduction instruction value for user i at time t; At the same time, capturing expert preference factor actions in demand response A positive value indicates that the expert is aggressive, while a negative value indicates that the expert is conservative.

7. The expert-data dual-driven demand response optimization method considering user behavior coupling according to claim 6 is characterized in that: The expert preference action factor is removed from the generated load reduction allocation strategy to reduce the error caused by expert preference. The expert strategy builder only focuses on immediate rewards and generates an expert strategy without expert preference: in, represents the expert strategy for the load reduction instruction value finally assigned to each user, is the expert preference factor action, It is the expert strategy for assigning load reduction instruction values ​​to each user after removing the expert preference factor.

Citation Information

Patent Citations

  • Battery energy storage system peak clipping and valley filling real-time control method based on load prediction

    CN102624017A

  • Demand side response electrical load regulation and control method and system for virtual power plant

    CN110675275A

  • Expert system and method for power equipment failure analysis

    CN112116108A

  • Strategy protection defense method for deep reinforcement learning

    CN113392396A

  • Intelligent optimization method for power grid safe operation strategy based on deep reinforcement learning

    CN114048903A