Electric energy demand response resource allocation method based on load division and Q-learning

Through load partitioning and Q-learning-based energy demand response resource allocation method, based on historical load data of smart meters, the total energy cost problem is solved without relying on user equipment information, and the optimization of electricity cost and peak-to-average power ratio is achieved.

CN120638334AActive Publication Date: 2025-09-12WUHAN UNIV OF SCI & TECH +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511119915.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-09-12
Estimated Expiration
2045-08-12

AI Technical Summary

Technical Problem

Existing technologies are unable to minimize the total energy cost of a residential community based on historical load data from smart meters without relying on user device information.

Method used

Load partitioning and Q-learning methods are used to obtain historical load data from smart meters, user types are clustered through SOM, and Q-learning is used to allocate electricity, and electrical appliances are scheduled in combination with real-time electricity prices.

Benefits of technology

It has achieved the goal of reducing the community’s total electricity cost and peak-to-average ratio and optimizing electricity distribution without affecting users’ electricity usage habits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120638334A_ABST
    Figure CN120638334A_ABST
Patent Text Reader

Abstract

The invention discloses an electric energy demand response resource allocation method based on load division and Q-learning, and relates to the technical field of demand response scheduling. Comprising the following steps: S1, obtaining historical load data from an intelligent electric meter, normalizing a daily load curve of each user through a daily peak load, and obtaining a user type by using SOM based on relative load curve clustering; s2, using Q-learning to distribute energy to each residential user according to the daily peak load of the user; s3, collecting the electric energy demand of each residence user, and adjusting the residence user distribution electric energy according to the electric energy demand; and S4, according to the distributed electric energy and the real-time electricity price, scheduling the operation of the electric appliance. The invention aims to solve the technical problem that the existing method cannot minimize the total energy cost of the residential community under the condition of not depending on user equipment information and only based on historical load data of an intelligent electric meter.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of demand response scheduling, and more particularly to an electric energy demand response resource allocation method based on load partitioning and Q-learning. Background Art

[0002] Demand response (DR) refers to the autonomous adjustment of electricity users' existing electricity usage behavior by reducing or shifting peak electricity consumption when wholesale electricity prices rise or grid reliability is threatened, after receiving a notice of load reduction compensation or a price increase signal from the electricity supplier. This response is achieved by maintaining grid stability and curbing short-term sharp increases in electricity prices. In a smart grid environment, DR enables the controlled and intelligent transmission of electricity from the power generation side to active users on the demand side. Given that residential electricity consumption accounts for 30% to 40% of total global energy consumption, research on demand response in residential communities has both profound theoretical value and significant practical significance.

[0003] In recent years, users have become increasingly concerned about privacy protection. Traditional methods often rely on users proactively providing detailed information such as device operating status, appliance type, and energy usage preferences to achieve refined modeling and personalized scheduling. However, some users, out of concern for data security and privacy, refuse to share information about their household electrical appliances. This results in incomplete user-side information obtained by the scheduling model, thus affecting the effectiveness of the scheduling strategy.

[0004] To address these issues, some current research attempts to introduce privacy-preserving mechanisms, such as federated learning and differential privacy algorithms, or employ information desensitization techniques for model training. While these approaches can reduce reliance on raw device data to a certain extent, they often increase model complexity and computational resource consumption. Currently, no method exists to minimize the total energy cost of a residential community based solely on historical smart meter load data, independent of user device information. Summary of the Invention

[0005] In view of this, the present invention provides an energy demand response resource allocation method based on load partitioning and Q-learning, aiming to solve the technical problem that existing methods cannot minimize the total energy cost of residential communities without relying on user device information and only based on historical load data of smart meters.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions: A method for allocating electric energy demand response resources based on load partitioning and Q-learning includes the following steps: S1. Obtain historical load data from smart meters, normalize the daily load curve of each user by the daily peak load, and use SOM to cluster the user types based on the relative load curves. S2, using Q-learning to allocate energy to each residential user based on the user's daily peak load; S3. Collect the power demand of each residential user and adjust the power allocation to the residential user according to the power demand; S4. Dispatch electrical appliances to operate according to the allocated electricity and real-time electricity prices.

[0007] Optionally, in S1, the specific process is as follows: S11. Obtain a historical load curve from the smart meter and normalize the user's daily load curve using the daily load peak value; S12, best matching unit It refers to the mapping unit with the smallest distance to the current input vector x during SOM training. The distance is calculated using the Euclidean distance, as shown below: ; Among them, x represents the input vector; j represents the mapping unit, and each mapping unit j can be represented as a prototype vector , where d is the input dimension; S13. In each iteration, the prototype vector is updated. This process is repeated for a predetermined number of iterations. The adaptive coefficient and neighborhood radius The value of decreases over time and is defined as follows: ; ; Where k represents the number of iterations of SOM; represents the neighborhood kernel centered on the best matching unit; f represents the location of neurons in the SOM grid; S14. After obtaining the cluster type, the central curve expression of the relative load curve of different types of users is as follows: ; Where u represents the user category; Represents the relative load curve of user type u.

[0008] Optional, for each residential user i According to the following formula, t The historical load data is normalized, and the normalized value Between 0 and 1, the calculation formula is as follows: ; represents the historical load data of residential user i at time slot t; represents the daily load peak of residential user i at time slot t.

[0009] Optionally, in S2, after receiving the user type, the DR is used to calculate the energy distribution of each type u and each user i in the community; different agents adopt different allocation strategies for different types of users. In time slot t, the agent observes the state , and select actions according to the strategy π , get the current reward according to the reward function , and the agent observes the new state .

[0010] Optionally, the specific modeling method of the agent is as follows: S21. State modeling: Take the real-time price and the central value of this type of user as the current state, and the expression is as follows: ; in, Represents the central value of the daily normalized load of this type of users; S22, Action Modeling: Actions represent the relative load assigned to the corresponding type of user; the normalized load distribution takes discrete values, i.e., [0.1, 0.2, ..., 0.9, 1], and the corresponding indexes are [0, 1, ..., 8, 9]; the minimum normalized load is 0.1; the current state is selected using the ε-greedy strategy. Next action , the ε-greedy strategy is defined as follows: ; in, is the Q value of the state-action pair, and ε is the probability of selecting the action strategy; S23. Reward modeling: Reward modeling consists of three parts: electricity cost reduction, comfort distance reduction, and load distribution rationality, which are user electricity cost rewards. , User Comfort Reward and reasonable rewards for load distribution ;User electricity cost reward The definition is as follows: ; ; ; in, represents the relative load assigned to this type of user at time t; Indicates the minimum relative load assigned to this type of user at time t; Indicates the maximum relative load assigned to this type of user at time t; User Comfort Rewards The definition is as follows: ; ; The reasonable reward for load distribution is calculated by cosine similarity, so that the energy distributed is similar to the normalized load center curve of this type of user. is the value of the normalized cosine phase, load distribution reasonable reward The definition is as follows: ; ; ; in, for The maximum value of for The minimum value of award Rewarded by user electricity costs , User Comfort Reward and reasonable rewards for load distribution The composition is expressed as follows: ; in, 、 、 is the normalization parameter; S24, Q-learning algorithm: In the context of user energy consumption management, Q-learning simulates user electricity consumption behavior, allocates electricity to users, and optimizes electricity consumption strategies; The goal of energy consumption scheduling is to find the optimal strategy , that is, to determine the optimal distribution value of relative load to maximize the action value function; the basic mechanism of the Q-learning algorithm is to construct a Q value table, in which the Q value of each state-action pair is It is updated in each iteration until the convergence condition is met; in this way, the best action with the best Q value in each state is selected; the best Q value The definition is as follows: ; The Q value is updated based on the reward, learning rate, and discount factor, and its expression is as follows: ; Among them, θ represents the learning rate, When , the agent only uses prior information, and Indicates that the agent only considers the current estimate and ignores the prior information; according to the power allocation results of different agents Q-learning and the user's daily load peak, the power is allocated to each user in each type .

[0011] The above technical solution shows that, compared with the existing technology, the present invention provides a method for allocating electricity demand response resources based on load partitioning and Q-learning. First, the historical load data from the user's smart meter is sent to the load aggregator. The load aggregator normalizes the load data by the user's daily load peak and then uses the SOM to partition the users, distinguishing users with different electricity usage habits. The load aggregator then allocates electricity to different types of users in the community based on the user partitioning results and real-time electricity prices. Finally, the user's electricity consumption behavior after scheduling forms new energy consumption data, which is then fed back to the load aggregator by the smart meter. The method of the present invention is a complete residential community demand response resource allocation method that has the characteristics of reducing the community's total electricity cost and peak-to-average ratio. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0013] Figure 1 This is a schematic diagram of the overall framework of power demand response resource allocation based on load partitioning and Q-learning; Figure 2 It is a flow chart of the method for allocating electric energy demand response resources based on load partitioning and Q-learning provided by the present invention; Figure 3 It is a comparison chart of the loads assigned to 6 types of users and the original loads; Figure 4 This is a comparison chart of the electricity consumption of residential user 1 before and after participating in demand response; Figure 5 This is a comparison of electricity consumption before and after residential communities participated in demand response. DETAILED DESCRIPTION

[0014] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0015] An embodiment of the present invention discloses a method for allocating power demand response resources based on load partitioning and Q-learning, comprising the following steps: S1. Obtain historical load data from smart meters, normalize the daily load curve of each user by the daily peak load, and use SOM to cluster the user types based on the relative load curves. S2, using Q-learning to allocate energy to each residential user based on the user's daily peak load; S3. Collect the power demand of each residential user and adjust the power allocation to the residential user according to the power demand; S4. Dispatch electrical appliances to operate according to the allocated electricity and real-time electricity prices.

[0016] like Figure 1 As shown in FIG, this embodiment provides an energy demand response resource allocation method based on load partitioning and Q-learning, which is applied to demand response scheduling in residential communities. It includes two key parts. One part is user partitioning, where the load aggregator divides user types using SOM based on historical load curves (historical energy consumption data). The other part is energy distribution, where the load aggregator uses Q-learning to calculate the energy distribution for each user in the community based on user type and real-time electricity prices. Figure 2 As shown, the method of the present invention comprises the following steps: Step S1: The load aggregator obtains historical load data from smart meters, normalizes the daily load curve of each user by the daily peak load, and uses SOM to cluster the relative load curves to obtain the user type.

[0017] The specific process of this step is as follows: S11. The load aggregator first obtains the historical load curve (historical energy consumption data) from the smart meter and normalizes the user's daily load curve using the daily load peak value to ensure that the daily load curve size of different users is in the range of 0 to 1, so as to eliminate the clustering between load curves caused by differences in electricity consumption scale between different users and fail to reflect the user's electricity consumption habits.

[0018] Residential users For example, for each residential user According to the following formula, The historical load data is normalized, and the normalized value Between 0-1.

[0019] ; S12, SOM is an unsupervised learning algorithm, usually represented as a two-dimensional grid mapping unit, used to map high-dimensional data to low-dimensional space while maintaining the topological structure of the data. It refers to the mapping unit with the smallest distance to the current input vector x during SOM training. The distance is calculated using the Euclidean distance, as shown below: ; Among them, x represents the input vector; j represents the mapping unit, and each mapping unit j can be represented as a prototype vector , where d is the input dimension.

[0020] S11 is equivalent to data preprocessing, which involves normalizing the dataset and then clustering it using the SOM. The reason for normalization is mentioned in S11—it eliminates clustering between load curves caused by differences in electricity consumption between different users, which can prevent them from reflecting their electricity usage habits.

[0021] S12 and S13 are introductions to the basic principles of SOM. S12 explains that the SOM in the present invention is to find the minimum matching unit (the mapping unit with the minimum distance to the current input vector x) based on the Euclidean distance - the SOM clustering process can be understood as the process of continuously assigning the input vector to the minimum matching unit as the algorithm iterates. S13 explains the process of SOM learning data features and updating clustering results. The classification process does not end immediately after the input vector finds the minimum mapping unit. SOM also continues to learn and update the unit and its neighborhood to gradually optimize the clustering structure. As the iteration proceeds, the adaptive coefficient and neighborhood radius It will gradually decrease, and the algorithm will gradually converge.

[0022] S13. In each iteration, the prototype vector will be updated. This process will be repeated for a predetermined number of iterations. The adaptive coefficient and neighborhood radius The value of decreases over time and is defined as follows: ; ; Where k represents the number of iterations of SOM; represents the neighborhood kernel centered on the best matching unit; f Represents the location of neurons in the SOM grid.

[0023] S14. After obtaining the cluster type, the central curve expression of the relative load curve of different types of users is as follows: ; Where u represents the user category; Represents the relative load curve of user type u.

[0024] Step S2: The load aggregator uses Q-learning to allocate energy to each residential user according to the user's daily peak load.

[0025] After receiving the user type, the load aggregator calculates the energy allocation of each type u and each user i in the community through DR. Different agents adopt different allocation strategies for different types of users. In time slot t, the agent observes the state , and select actions according to the strategy π , get the current reward according to the reward function After taking this action, the agent observes the new state This process continues until all services are assigned. The following example shows the specific modeling of an agent. The other agents are modeled using the same method.

[0026] S21. State modeling: Take the real-time price and the central value of this type of user as the current state, and the expression is as follows: ; in, Represents the central value of the daily normalized load of this type of users.

[0027] S22. Action modeling: Actions represent the relative load assigned to this type of user. The normalized load distribution takes discrete values, i.e., [0.1, 0.2, ..., 0.9, 1], and the corresponding indexes are [0, 1, ..., 8, 9]. The minimum value of the normalized load is 0.1. This is because each user has unschedulable devices, so the normalized load cannot be 0 at any time. In the strategy selection process, the ε-greedy strategy is used to select the current state. Next action , the ε-greedy strategy is defined as follows: ; in, is the Q value of the state-action pair, and ε is the probability of choosing the action strategy.

[0028] S23. Reward modeling: Reward modeling consists of three parts: electricity cost reduction, comfort distance reduction, and load distribution rationality, which are user electricity cost rewards. , User Comfort Reward and reasonable rewards for load distribution ;User electricity cost reward The definition is as follows: ; ; ; in, represents the relative load assigned to this type of user at time t; Indicates the minimum relative load assigned to this type of user at time t; Indicates the maximum relative load assigned to this type of user at time t.

[0029] User Comfort Rewards The definition is as follows: ; ; The reasonable reward for load distribution is calculated by cosine similarity, so that the energy distributed is similar to the normalized load center curve of this type of user. is the value of the normalized cosine phase, load distribution reasonable reward The definition is as follows: ; ; ; in, for The maximum value of for The minimum value of .

[0030] award Rewarded by user electricity costs , User Comfort Reward and reasonable rewards for load distribution The composition is expressed as follows: ; in, 、 、 is the normalization parameter, .

[0031] S24. Q-learning algorithm: Q-learning is a reinforcement learning algorithm that learns optimal policies through interaction with the environment, specifically selecting actions that maximize cumulative rewards in a given state. In the context of user energy management, Q-learning can be used to simulate user electricity usage behavior, allocate energy to users, and optimize electricity usage strategies.

[0032] The goal of energy consumption scheduling is to find the optimal strategy , that is, to determine the optimal distribution value of relative load to maximize the action value function. The basic mechanism of this algorithm is to construct a Q value table, in which the Q value of each state-action pair is It is updated in each iteration until the convergence condition is met. In this way, the best action with the best Q value in each state can be selected. The definition is as follows: ; The Q value can be updated based on the reward, learning rate and discount factor, and its expression is as follows: ; Among them, θ represents the learning rate, which ranges from 0 to 1. When , the agent only uses prior information, and Indicates that the agent only considers the current estimate and ignores the prior information. Based on the energy allocation results of different agents Q-learning and the user's daily load peak, energy is allocated to each user in each type. .

[0033] Step S3: The load aggregator receives the power demand of each residential user from the home energy management system. If the power allocated to the residential user is less than the power demand, the load aggregator adjusts the power allocated to the residential user.

[0034] The electricity demand of each residential user is different, which depends on the user's previous energy consumption data. The load aggregator allocates more energy to users who use more electricity and less energy to users who use less electricity.

[0035] Step S4: The home energy management system schedules the operation of electrical appliances according to the electric energy allocated by the load aggregator and the real-time electricity price.

[0036] The home energy management system compares the real-time electricity price with the price the user is willing to pay. If the current price is greater than the user's WTP, the current price is considered high; if the current price is less than or equal to the user's WTP, the current price is considered low.

[0037] Residential users i For example, residential users The home energy management system determines the residential users based on the high price / low price information received i Start and stop of home appliances. i Time slots tIf the electricity price is high, the home energy management system may suspend the operation of some dispatchable devices and reduce the power of some adjustable devices, and retain non-dispatchable devices. Time slots t If the electricity price is low, the home energy management system may enable some dispatchable devices and increase the power of some adjustable devices.

[0038] The following is an analysis through specific actual cases to prove the effect of the present invention.

[0039] like Figure 3 The figure shows the loads allocated to six user categories and their original loads, as determined by SOM clustering. It shows that each user category achieves optimal energy allocation: when electricity prices are above average, allocated energy decreases; when prices are below average, allocated energy increases; and allocated energy is roughly the same as when no DR strategy is used. The proposed method, referred to as the DR-SOM-Q algorithm, demonstrates that the DR-SOM-Q algorithm can help users reduce their electricity costs while minimizing changes to their electricity usage habits.

[0040] Figure 4 This figure compares the electricity consumption of user 1 using the DR-SOM-Q algorithm with that of user 1 without a demand response strategy. Using the DR-SOM-Q algorithm increases electricity consumption when the price is below the average and decreases it when the price is above the average. The electricity consumption trends for both the DR-SOM-Q and non-DR strategies are similar, with minimal changes, indicating that the DR-SOM-Q algorithm barely alters the user's electricity usage habits.

[0041] Figure 5 This is a comparison of loads in communities using the DR-SOM-Q algorithm and those not using a demand response strategy. The DR-SOM-Q algorithm increases electricity consumption when electricity prices are below average and reduces it when prices are above average. The load trends for both the DR-SOM-Q and non-DR strategies are similar, with minimal variation. This demonstrates that the DR-SOM-Q algorithm helps communities reduce electricity costs while minimizing user comfort.

[0042] To further verify the effectiveness of the algorithm, the peak average ratio (PAR) is introduced. ; Represents the sum of energy consumption of all residential community users, that is, the electricity purchased by the residential community from the grid.

[0043] Table 1 shows the performance of the DR-SOM-Q algorithm and the DR-SOM-MDP algorithm (demand response scheduling algorithms based on SOM and Markov decision-making) compared to a non-DR strategy. The difference between the DR-SOM-Q and DR-SOM-MDP algorithms stems from their different uncertainty resolution strategies. In terms of power consumption, both algorithms are roughly equivalent to a non-DR strategy. In terms of power cost, the DR-SOM-Q algorithm achieves a greater reduction than the DR-SOM-MDP algorithm. In terms of peak-to-average power ratio, the DR-SOM-Q algorithm achieves a slightly greater reduction than the DR-SOM-MDP algorithm. In terms of user comfort distance, the DR-SOM-Q algorithm slightly outperforms the DR-SOM-MDP algorithm, but this remains within an acceptable range and does not significantly impact the user's electricity experience.

[0044] Table 1 Performance changes of DR-SOM-Q algorithm and DR-SOM-MDP algorithm The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0045] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for allocating power demand response resources based on load partitioning and Q-learning, characterized in that: The following steps are involved: S1. Obtain historical load data from smart meters, normalize the daily load curve of each user by the daily peak load, and use SOM to cluster the user types based on the relative load curves. S2, using Q-learning to allocate energy to each residential user based on the user's daily peak load; S3. Collect the power demand of each residential user and adjust the power allocation to the residential user according to the power demand; S4. Dispatch electrical appliances to operate according to the allocated electricity and real-time electricity prices.

2. The method for allocating power demand response resources based on load partitioning and Q-learning according to claim 1, characterized in that: In S1, the specific process is as follows: S11. Obtain a historical load curve from the smart meter and normalize the user's daily load curve using the daily load peak value; S12, best matching unit It refers to the mapping unit with the smallest distance to the current input vector x during SOM training. The distance is calculated using the Euclidean distance, as shown below: ; Among them, x represents the input vector; j represents the mapping unit, and each mapping unit j can be represented as a prototype vector , where d is the input dimension; S13. In each iteration, the prototype vector is updated. This process is repeated for a predetermined number of iterations. The adaptive coefficient and neighborhood radius The value of decreases over time and is defined as follows: ; ; Where k represents the number of iterations of SOM; represents the neighborhood kernel centered on the best matching unit; f represents the location of neurons in the SOM grid; S14. After obtaining the cluster type, the central curve expression of the relative load curve of different types of users is as follows: ; Where u represents the user category; Represents the relative load curve of user type u.

3. The method for allocating power demand response resources based on load partitioning and Q-learning according to claim 2, characterized in that: For each residential user i According to the following formula, t The historical load data is normalized, and the normalized value Between 0 and 1, the calculation formula is as follows: ; represents the historical load data of residential user i at time slot t; represents the daily load peak of residential user i at time slot t.

4. The method for allocating power demand response resources based on load partitioning and Q-learning according to claim 1, characterized in that: In S2, after receiving the user type, the DR calculates the energy allocation of each type u and each user i in the community; different agents adopt different allocation strategies for different types of users. In time slot t, the agent observes the state , and select actions according to the strategy π , get the current reward according to the reward function , and the agent observes the new state .

5. The method for allocating power demand response resources based on load partitioning and Q-learning according to claim 4, characterized in that: The specific modeling method of the agent is as follows: S21. State modeling: Take the real-time price and the central value of this type of user as the current state, and the expression is as follows: ; in, Represents the central value of the daily normalized load of this type of users; S22, Action Modeling: Actions represent the relative load assigned to the corresponding type of user; the normalized load distribution takes discrete values, i.e., [0.1, 0.2, ..., 0.9, 1], and the corresponding indexes are [0, 1, ..., 8, 9]; the minimum normalized load is 0.1; the current state is selected using the ε-greedy strategy. Next action , the ε-greedy strategy is defined as follows: ; in, is the Q value of the state-action pair, and ε is the probability of selecting the action strategy; S23. Reward modeling: Reward modeling consists of three parts: electricity cost reduction, comfort distance reduction, and load distribution rationality, which are user electricity cost rewards. , User Comfort Reward and reasonable rewards for load distribution ;User electricity cost reward The definition is as follows: ; ; ; in, represents the relative load assigned to this type of user at time t; Indicates the minimum relative load assigned to this type of user at time t; Indicates the maximum relative load assigned to this type of user at time t; User Comfort Rewards The definition is as follows: ; ; The reasonable reward for load distribution is calculated by cosine similarity, so that the energy distributed is similar to the normalized load center curve of this type of user. is the value of the normalized cosine phase, load distribution reasonable reward The definition is as follows: ; ; ; in, for The maximum value of for The minimum value of award Rewarded by user electricity costs , User Comfort Reward and reasonable rewards for load distribution The composition is expressed as follows: ; in, 、 、 is the normalization parameter; S24, Q-learning algorithm: In the context of user energy consumption management, Q-learning simulates user electricity consumption behavior, allocates electricity to users, and optimizes electricity consumption strategies; The goal of energy consumption scheduling is to find the optimal strategy , that is, to determine the optimal distribution value of relative load to maximize the action value function; the basic mechanism of the Q-learning algorithm is to construct a Q value table, in which the Q value of each state-action pair is It is updated in each iteration until the convergence condition is met; in this way, the best action with the best Q value in each state is selected; the best Q value The definition is as follows: ; The Q value is updated based on the reward, learning rate, and discount factor, and its expression is as follows: ; Among them, θ represents the learning rate, When , the agent only uses prior information, and Indicates that the agent only considers the current estimate and ignores the prior information; according to the power allocation results of different agents Q-learning and the user's daily load peak, the power is allocated to each user in each type .

Citation Information

Patent Citations

  • Load prediction-based load identification method and system capable of participating in demand response

    CN114662761A

  • Park energy distribution method and system based on multi-element load clustering

    CN116957265A

  • Power distribution network agent group cooperative regulation and control method and system based on deep reinforcement learning

    CN118611038A

  • Multi-load equipment energy consumption management method and device and computer equipment

    CN119539444A

  • Multi-cloud deployment scheduling method, device, equipment, medium and program product

    CN119690603A