Cluster power consumer load reference line response method based on time sequence deep reinforcement learning

Through the cluster power user load reference line response method based on deep reinforcement learning in time series, the problem that traditional demand response mechanisms are difficult to accurately reflect user load characteristics is solved, accurate estimation of load reference lines and dynamic correction of line lines is achieved, the accuracy and intelligence of demand response are improved, and the stability and resource utilization efficiency of power systems are enhanced.

CN120146510APending Publication Date: 2025-06-13SHANGHAI UNIVERSITY OF ELECTRIC POWER
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510262839.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The traditional demand response mechanism is difficult to accurately reflect user load characteristics, resulting in poor response effects, large deviations in baseline estimation, failure to fully consider spatial and temporal differences, low efficiency of optimization algorithms, and difficult to meet the power system's demand for accurate and intelligent response strategies.

Method used

The cluster power user load baseline response method based on time-series deep reinforcement learning is adopted. By cleaning and standardizing massive data, the KMedoids+SoftDTW clustering algorithm is used to build a typical load baseline, dynamically correct the load baseline, and combining multi-factor optimization model and reinforcement learning algorithm to realize the intelligence and precision of the cluster power user demand response.

Benefits of technology

Accurate estimation of user load baseline and dynamic correction of load line are achieved, the accuracy and intelligence of demand response are improved, the stability of the power system and resource utilization efficiency are enhanced, and the phenomenon of wind and light are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146510A_ABST
    Figure CN120146510A_ABST
Patent Text Reader

Abstract

According to the time sequence deep reinforcement learning-based cluster power consumer load reference line response method, intelligence and precision of cluster power consumer demand response are realized by accurately estimating a user load reference line and dynamically correcting a load reference line in combination with a multi-factor optimization model and a reinforcement learning algorithm. The method comprises the following steps: (1) carrying out cleaning and standardization processing on mass power consumer load time series data, and removing noise and abnormal values in the data; (2) determining an optimal clustering number by using an Elbow rule, classifying the power consumers by using a KMedoids + SoftDTW clustering algorithm, and constructing a typical load baseline of the power consumers; and (3) preliminarily determining a cluster power consumer load alignment line CDL, and adjusting the alignment line by considering the response capability. According to the method, the user load baseline is accurately estimated, the load directrix is dynamically corrected, and the multi-factor optimization model and the reinforcement learning algorithm are combined, so that the intelligence and precision of the cluster power user demand response are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power system demand response, and specifically to a method for responding to the load baseline of cluster power users based on time series deep reinforcement learning. Background Technique

[0002] With the rapid growth of the installed capacity of new energy, the load characteristics of the power system have become increasingly complex, and the peak-valley difference has further increased. Traditional demand response mechanisms mainly rely on price signals or general incentive standards, making it difficult to accurately reflect the user load characteristics and resulting in poor response effects. Although the baseline type (CBL) and quasi-linear type (CDL) response mechanisms can partially solve the above problems, they still have the following deficiencies:

[0003] First, the baseline estimation deviation is relatively large, mainly relying on historical averages and making it difficult to adapt to dynamic load changes; second, when evaluating the user response ability, the existing methods fail to fully consider its spatio-temporal differences, resulting in insufficient evaluation of the response potential; finally, traditional optimization algorithms are less efficient in dealing with complex decision-making problems with multi-factor coupling and are difficult to meet the requirements of the power system for precise and intelligent response strategies.

[0004] Therefore, the present invention proposes an optimal strategy system and method for intelligent response of the load baseline of cluster power users based on time series deep reinforcement learning, aiming to solve the above technical problems. Summary of the Invention

[0005] In view of the above problems, the present invention proposes a method for responding to the load baseline of cluster power users based on time series deep reinforcement learning. By accurately estimating the user load baseline, dynamically correcting the load baseline, and combining a multi-factor optimization model and a reinforcement learning algorithm, the intelligent and precise demand response of cluster power users is realized.

[0006] To achieve the above object, the technical solution adopted by the present invention is:

[0007] The method for responding to the load baseline of cluster power users based on time series deep reinforcement learning includes the following steps:

[0008] (1) Clean and standardize the massive power user load time series data to remove noise and outliers in the data;

[0009] (2) Use the Elbow method to determine the optimal number of clusters, classify power users using the KMedoids+SoftDTW clustering algorithm, and construct a typical load baseline for power users;

[0010] (3) Initially determine the load baseline CDL of cluster power users and adjust the baseline considering the response ability.

[0011] As a further improvement of the present invention, in step (2), the KMedoids + SoftDTW clustering algorithm is used to classify power users, which specifically includes: adopting the improved distance metric Soft-DTW method and automatically adjusting the smoothing parameter through grid search to handle the alignment problem between time series, and obtaining the distance matrix:

[0012]

[0013] where d(X i ,Y j ) represents the Euclidean distance between points X i and Y j . By calculating the silhouette coefficient related error, the optimal number of clusters is determined:

[0014]

[0015] where d[C AV (X i )] represents the average distance from the data point X i to other data points within its affiliated cluster, and d[O AV (X i )] is the average distance from the data point to the data points within the nearest other cluster. Taking the data points as the initial medoids clustering centers, the data points are assigned to the nearest clusters:

[0016]

[0017] And the data point that minimizes the sum of the distances from all data points in each cluster to other data points is selected as the new medoid:

[0018]

[0019] Repeat the steps of assignment and update of medoids until the clustering centers no longer change or reach the predetermined number of iterations.

[0020] As a further improvement of the present invention, in the modeling and optimization method for initially determining the cluster power user load reference line CDL in step (3), the CDL is obtained by subtracting the photovoltaic, wind power, and line loss power from the total load at a certain moment to obtain the net load curve and then flipping it:

[0021] P Net,t =P G,t -(P W,t +P S,t )-P U,t (5)

[0022] Take its contour shape as the reference line, and perform per-unit processing on the cluster power user load curve and the reference line:

[0023]

[0024] Where: l i * (t) is the load curve l of type i power user i (t) is normalized by the function f u (·) after the normalized per unit value, construct the power user response capability index system of each node, adopt the data dispersion degree to dynamically determine the gray resolution coefficient to improve the gray relational analysis method IGRA, and use the Euclidean distance to measure the similarity between the normalized baseline load curve and the CDL issued by the power department:

[0025]

[0026] Where: ε is a given constant, which determines the incentives obtained by each cluster power user after participating in the CDL type response:

[0027]

[0028] Where: and P t * (i) are the normalized load values ​​of the demand response type i power users and the power sector in the tth period; β d is the similarity coefficient; Υ is the excitation coefficient; is the total response amount of cluster power users.

[0029] As a further improvement of the present invention, in the modeling and optimization method for the preliminary determination of the cluster power user load line CDL in step (3), the power user response capability index system of each node is constructed, including considering the historical load level, subjective response information and user type index characteristics; the formula for determining the DR potential coefficient to correct the CDL is (9), and the comprehensive adjustment formula is (10);

[0030]

[0031] Where: For X 0 (k) Minus X i (k), σ is the standard deviation of the indicator feature set, w t is the weight coefficient of the i-th time node.

[0032] As a further improvement of the present invention, the optimal strategy model of the cluster power user response based on the baseline in step (3) includes:

[0033] ①Objective function: A multi-objective function is set that includes response incentives, electricity price costs, curtailment costs of wind and light, and action adjustment costs. The specific formula is (11), where the response incentive is the incentive obtained by users participating in the CDL response. The electricity price cost formulates different electricity prices according to the electricity demand and cost differences in different periods. The curtailment costs of wind and light are used to promote the consumption of new energy, and the action adjustment costs are used to reasonably adjust the electricity consumption behavior of users;

[0034] max[R award -R con -R energy -R action (11)

[0035] In the formula: R award Response incentive; R con Electricity price cost; R energy Curtailment costs of wind and light; R action Action adjustment cost;

[0036] ②Constraints: Include quasi-linear response constraints, system security constraints, power balance constraints, new energy output constraints, and flexibility supply-demand balance constraints.

[0037] As a further improvement of the present invention, in step (3), the quasi-linear response constraint formulas in the optimal strategy model of cluster power user response based on the baseline are (16) and (17):

[0038]

[0039] ΔP F,min,t ≤ΔP F,t ≤ΔP F,max,t (17)

[0040] In the formula: P F,t and P F,base,t are the actual load and the baseline load of the user at time t, respectively; ΔP F,t is the adjustable power of the user; ΔP F,max,t and ΔP F,min,t are the upper and lower limits of the adjustment power, respectively. The system security constraint formulas are (18) and (19):

[0041]

[0042] In the formula: U min and U max are the lower and upper voltage limits at the user node, respectively; I ij,max is the safe current of line (i, j) at time t. The power balance constraint formula is (20):

[0043]

[0044] Where: u(j) is the set of the head nodes of the branches with j as the end node; w(j) is the set of the end nodes of the branches with j as the head node; P ij,t and Q ij,t are the active power and reactive power flowing through the feeder (i, j) during the time period t, respectively; I ij is the current amplitude; R ij and X ij are the resistance and reactance, respectively; P j,t and Q j,t are the active power and reactive power injected into node j, respectively. The new energy output constraint formulas are (21) and (22):

[0045]

[0046] The flexibility supply-demand balance constraint formula is (23):

[0047]

[0048] Where: is the set of flexible resources i; ΔP i,t is the regulation power of flexible resource i; ΔP D,t is the flexibility demand.

[0049] Beneficial effects: Compared with the prior art, the present invention adopts the above technical solution and has the following advantages:

[0050] Demand response is the core link to relieve the pressure of the power system and optimize the power supply-demand balance, and it is also an important carrier to realize the stable operation of the power system with a high proportion of renewable energy access. As the main body of demand response, the cluster power users should take corresponding measures according to their own electricity consumption characteristics and system requirements to create a significant load regulation effect. Therefore, analyzing and studying the logical relationship between user electricity consumption behavior and system load demand and quantifying the user response ability are the problems that need to be considered at present.

[0051] In view of the above problems, the present invention proposes a research method for the optimal intelligent response strategy of the cluster power user load baseline based on time-series deep reinforcement learning, and conducts research from multiple aspects such as power user load baseline estimation, load baseline correction, and response strategy formulation.

[0052] By introducing this method based on time-series deep reinforcement learning, the response ability of user-side resources can be better explored, the precise scheduling of the power system and the optimal allocation of the load can be realized, the stability of the power system and the resource utilization efficiency can be improved, and the phenomena of wind and light abandonment can be reduced; at the same time, the power load management can be strengthened, the power supply can be reasonably scheduled, and the power waste can be reduced. Description of the Drawings

[0053] Figure 1 is the framework of the present invention;

[0054] Figure 2 It is the algorithm flowchart of the present invention. Detailed implementation manners

[0055] The present invention will be further described below in conjunction with numerical examples. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention, and cannot be used to limit the protection scope of this application. A method framework for cluster user load baseline estimation, cluster user load reference line correction, and cluster user response strategy formulation based on temporal deep reinforcement learning, referring to Figure 1 shown in the algorithm flowchart, referring to Figure 2 shown in the figure, includes the following steps:

[0056] Step 1: Data collection and preprocessing;

[0057] Collect the basic residential hourly load dataset in a certain area in 2024, covering multi-dimensional information such as residential electricity load data, weather data, and new energy output data. Clean and preprocess the basic residential load data, and use the formula

[0058]

[0059] In the formula: is the baseline load of the i-th resident; basic residential load; λ is the daily maximum load coefficient; τ is the daily maximum peak-valley difference coefficient, to obtain 936 pieces of multi-dimensional relevant data and load baselines per hour from 1:00 on January 1st to 0:00 on January 2nd, providing a high-quality data basis for subsequent analysis.

[0060] Step 2: Determine the clustering number of residential user clusters;

[0061] Use the Elbow method to determine the optimal clustering number of the basic load curves of residential users. Calculate the relevant errors of the silhouette coefficients under different clustering numbers. By observing the relationship diagram between the clustering number and the relevant errors of the silhouette coefficients, find the inflection point similar to the "elbow" in the curve. The clustering number corresponding to this inflection point is the optimal clustering number, which is determined to be 4 after calculation. Table 1 shows the advantages of the clustering algorithm proposed in the present invention.

[0062] Table 1 Comparison of clustering methods

[0063]

[0064] Step 3: Conduct clustering analysis based on the Kemdiods-SoftDTW algorithm;

[0065] Construct typical load baselines for residential users based on the load characteristics after clustering, showing four different residential user load patterns, providing a basis for formulating personalized demand response strategies in the future.

[0066] Step 4: Construct and correct the Cluster Power Users' CDL;

[0067] The power company dispatching center preliminarily determines the CDL by inverting the net load curve. Considering the spatio-temporal differences in the response capabilities of different residential users, a DR potential coefficient is introduced. According to the index characteristics such as historical load levels, response subjective information, and user types, an index system for the response capabilities of power users at each node is constructed. The grey discrimination coefficient is dynamically determined based on the data dispersion degree to improve the Improved Grey Relational Analysis (IGRA), and the CDL is dynamically adjusted. Taking the residents in Cluster I as an example, the adjusted CDL is shown to be more in line with the actual electricity consumption characteristics and regulation potential.

[0068] Step 5: Build an optimal response strategy model for cluster power users based on the baseline;

[0069] Construct an optimal response strategy model for cluster power users considering multiple factors, and set an objective function that includes response incentives, electricity price costs, curtailment of wind and solar costs, and action adjustment costs.

[0070] Step 6: Solve the model using the GRU-MATD3 algorithm;

[0071] Combine GRU and MATD3 to form the GRU-MATD3 algorithm, and integrate it into the Markov decision process. Referring to the Multi-Agent Deep Deterministic Policy Gradient (MADDPG), collect the observation variables of each cluster power user for centralized training - distributed execution to find the optimal response actions.

[0072] Step 7: Evaluate and compare the training effects;

[0073] Train by combining the GRU time series neural network with the MATD3 and MADDPG algorithms respectively, and observe the change of the reward value with the number of training rounds. It is found that the GRU-MATD3 combination first achieves convergence at about 500 rounds, and the reward value stabilizes in the range of 93.1 - 93.5 at 900 - 1000 rounds. Compared with the range of 91.8 - 93.3 of GRU-MADDPG, it is more stable, has a smaller fluctuation range, and a higher cumulative reward value, verifying the advantages of the GRU-MATD3 algorithm in dealing with this problem.

[0074] Step 8: Develop and analyze the optimal response strategy for residents based on the CDL;

[0075] Select the results generated by the GRU-MATD3 algorithm, and analyze the relevant data of the DR actions and DR strategies of the 4 types of cluster residents clustered by Kemdiods-SoftDTW at different times. Observe the electricity power adjustment situations of different residents at different time periods, such as different actions during low-load periods and peak-load periods, as well as the power change trend of the overall DR strategy.

[0076] Step 9: Economic benefit analysis;

[0077] Conduct economic benefit analysis on different resident clusters, calculate indicators such as the CDL response benefit, electricity price cost, and total benefit of Residents 1, 3, and 4. The specific results are shown in Table 2. It is found that the CDL response benefits of Residents 1, 3, and 4 are all significantly higher than their electricity price costs. Although the total benefit is relatively small after considering cost factors such as curtailment of wind and solar power, it still reflects the role of this algorithm in guiding residents to optimize electricity consumption and create value. Finally, the total benefit value is 93.38×10 4 Yuan.

[0078] Table 2 DR benefit values based on the sequential GRU-MATD3 algorithm

[0079]

[0080] Step 10: Comparative analysis to verify the effectiveness of the algorithm;

[0081] Design 6 scenarios for comparative analysis, covering different situations such as only considering a single electricity price factor, mainly focusing on the load characteristic CDL, comprehensively considering both CDL and electricity price factors, emphasizing the response ability, and fully integrating the response ability, electricity price, and CDL factors. Compare the similarity between the demand response (DR) strategy and the expected key demand load (CDL) target and the total economic benefit value under different scenarios. The results are shown in Table 3, indicating that the average similarity index of the scenarios considering the response ability is higher. The similarity of Scenario 6 using the method of this paper is as high as 95.82%, and in most scenarios, the method of this paper can obtain more considerable economic benefit values, verifying the effectiveness and advantages of the GRU-MATD3 algorithm in complex actual scenarios.

[0082] Table 3 Comparison of response indicators of different methods in each scenario

[0083]

[0084] The above are only the preferred embodiments of the present invention, and it is not any other form of limitation to the present invention. Any modification or equivalent change made according to the technical essence of the present invention still belongs to the scope protected by the present invention.

Claims

1. A load baseline response method for cluster power users based on time series deep reinforcement learning, characterized by: The following steps are involved: (1) Clean and standardize the massive load time series data of power users to remove noise and outliers in the data; (2) The Elbow rule is used to determine the optimal number of clusters, and the KMedoids+SoftDTW clustering algorithm is used to classify power users and construct a typical load baseline for power users; (3) Preliminary determination of the load standard line (CDL) of cluster power users, and adjustment of the standard line considering the response capability.

2. The method for cluster power user load baseline response based on time series deep reinforcement learning according to claim 1 is characterized by: The step (2) uses the KMedoids+SoftDTW clustering algorithm to classify power users, specifically including: using the improved distance metric Soft-DTW method and automatically adjusting the smoothing parameters through grid search to handle the alignment problem between time series, and obtaining the distance matrix: In the formula, d(X i ,Y j ) represents point X i and Y j The Euclidean distance between them. The optimal number of clusters is determined by calculating the error associated with the silhouette coefficient: In the formula, d[C AV (X i )] represents the data point X i The average distance to other data points in the cluster to which it belongs, d[O AV (X i )] is the average distance from the data point to the data points in the other nearest clusters. The data point is used as the initial medoids cluster center and the data point is assigned to the nearest cluster: And select the data point that minimizes the sum of distances from all data points in each cluster to other data points as the new medoid: The steps of assigning and updating medoids are repeated until the cluster center does not change or the predetermined number of iterations is reached.

3. The method for cluster power user load baseline response based on time series deep reinforcement learning according to claim 1 is characterized by: In the modeling and optimization method for the preliminary determination of the cluster power user load criterion CDL in step (3), the CDL is based on the total load at a certain moment, deducting the photovoltaic, wind power and line loss power to obtain the net load curve and then flipping it: P Net,t =P G,t -(P W,t +P S,t )-P U,t (5) Take its contour shape as the criterion, and normalize the load curve and criterion of cluster power users: Where: l i * (t) is the load curve l of type i power user i (t) is normalized by the function f u (·) after the normalized per unit value, construct the power user response capability index system of each node, adopt the data dispersion degree to dynamically determine the gray resolution coefficient to improve the gray relational analysis method IGRA, and use the Euclidean distance to measure the similarity between the normalized baseline load curve and the CDL issued by the power department: Where: ε is a given constant, which determines the incentives obtained by each cluster power user after participating in the CDL type response: Where: and P t * (i) are the normalized load values ​​of the demand response type i power users and the power sector in the tth period; β d is the similarity coefficient; Υ is the excitation coefficient; is the total response amount of cluster power users.

4. The method for cluster power user load baseline response based on time series deep reinforcement learning according to claim 3 is characterized by: The modeling and optimization method for the preliminary determination of the cluster power user load line CDL in step (3) constructs a response capability index system for each node power user, including considering historical load levels, subjective response information and user type index characteristics; the formula for determining the DR potential coefficient to correct the CDL is (9), and the comprehensive adjustment formula is (10); Where: = X0(k) minus X i (k), σ is the standard deviation of the indicator feature set, w t is the weight coefficient of the i-th time node.

5. The method for cluster power user load baseline response based on time series deep reinforcement learning according to claim 1 is characterized by: The optimal strategy model of the cluster power user response based on the baseline in step (3) includes: ① Objective function: A multi-objective function including response incentive, electricity price cost, wind and solar power abandonment cost, and action adjustment cost is set. The specific formula is (11), where the response incentive is the incentive obtained by users participating in CDL response, the electricity price cost is used to set different electricity prices according to the power demand and cost differences in different periods, the wind and solar power abandonment cost is used to promote the consumption of new energy, and the action adjustment cost is used to reasonably adjust the user's electricity consumption behavior; max[R award -R con -R energy -R action ] (11) Where: R award Response stimulus; R con Electricity price cost; R energy Cost of curtailing wind and solar power; R action Action adjustment cost; ② Constraints: including quasi-linear response constraints, system safety constraints, power balance constraints, new energy output constraints, and flexibility supply and demand balance constraints.

6. The method for cluster power user load baseline response based on time series deep reinforcement learning according to claim 5 is characterized by: The quasi-linear response constraint formulas in the optimal strategy model of the cluster power user response based on the baseline in step (3) are (16) and (17): ΔP F,min,t ≤ΔP F,t ≤ΔP F,max,t (17) Where: P F,t and P F,base,t are the actual load and the benchmark load of the user at time t; ΔP F,t User adjustable power; ΔP F,max,t and ΔP F,min,t are the upper and lower power regulation limits respectively, and the system safety constraint formulas are (18) and (19): Where: U min and U max are the lower and upper limits of the voltage at the user node, respectively; I ij,max is the safe current of line (i, j) in period t. The power balance constraint formula is (20): Where: u(j) is the set of the first and last nodes of the branch with j as the last node; w(j) is the set of the last nodes of the branch with j as the first and last node; P ij,t and Q ij,t are the active and reactive powers flowing through feeder (i, j) in time period t, respectively; I ij Current amplitude; R ij and X ij are resistance and reactance respectively; P j,t and Q j,t are the active and reactive power injected into node j respectively, and the new energy output constraint formulas are (21) and (22): The flexible supply and demand balance constraint formula is (23): Where: Flexible resource i set; ΔP i,t The regulation power of flexible resource i; ΔP D,t Flexibility requirements.