Power grid dispatching fast optimization method based on cross-dimension migration of knowledge matrix

By adopting a cross-dimensional transfer method based on knowledge matrices, the problem of the inability to directly utilize historical knowledge in power grid dispatching was solved, enabling rapid optimization and efficient learning of power grid dispatching tasks, and improving learning efficiency and dispatching efficiency.

CN114626670BActive Publication Date: 2025-12-05HEFEI UNIV OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210089720.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-25
Publication Date
2025-12-05
Estimated Expiration
2042-01-25

AI Technical Summary

Technical Problem

With the increasing variety of new energy sources and flexible resources, the scale of power grid dispatching problems has expanded. Existing intelligent learning algorithms suffer from the "curse of dimensionality" and long learning times, and historical dispatching knowledge cannot be directly utilized, resulting in low learning efficiency.

Method used

A cross-dimensional transfer method based on knowledge matrix is ​​adopted. By constructing a Q-learning algorithm model of the source power grid system, Euclidean distance and DTW distance are used to measure task similarity, establish a mapping relationship between source tasks and target tasks, realize knowledge transfer, and optimize the scheduling plan of target tasks.

Benefits of technology

It accelerates the learning convergence speed of power grid dispatching tasks, improves learning efficiency, reduces learning costs, avoids redundant learning, and enhances the optimization efficiency of power grid dispatching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114626670B_ABST
    Figure CN114626670B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of power systems, and more particularly to a power grid dispatching fast optimization method based on knowledge matrix cross-dimension migration. First, a source power grid system without elastic resources and a target power grid system with elastic resources are constructed, and pre-learning is performed on the source power grid system on a typical day to obtain an optimized knowledge matrix as historical knowledge of power grid dispatching and form a source task knowledge base; then, the source task with the minimum distance to the target task is obtained by taking the net load as a similarity measurement feature, and an index matrix of similar states and actions between the source task and the target task is constructed; finally, the dispatching knowledge of the source task is migrated to the target task according to the index matrix to accelerate the optimization solution of the target task. The application proposes a knowledge matrix cross-dimension migration method based on reinforcement learning, solves the problem that the dispatching knowledge cannot be directly migrated when the state and action space dimensions are different, and effectively accelerates the optimization solution speed of the algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of power systems, and more particularly to a power grid dispatching fast optimization method based on knowledge matrix cross-dimension transfer. BACKGROUND

[0002] In recent years, with the rapid development of new energy, energy storage and electric vehicles, the complementary use of different types of energy and power grids is becoming an important way to improve the flexibility of power grids. After various flexible resources participate in power grid dispatching, although they can bring greater flexibility to power grid dispatching, with the increasing types of flexible resources on both sides of the source and load that can participate in power grid dispatching, it undoubtedly brings challenges to power grid optimization dispatching. At present, there are certain limitations in solving the power grid dispatching problem based on intelligent learning algorithms, such as the "dimension disaster" that occurs as the problem size increases, the need to re-solve different tasks, and long training time. In fact, different learning tasks are often related, so how to use existing dispatching knowledge and experience and transfer it to the solution of new problems to improve learning efficiency will be a problem worth studying. SUMMARY

[0003] In view of the deficiencies in the prior art, the application provides a power grid dispatching fast optimization method based on knowledge matrix cross-dimension transfer. The mapping between the states and actions of the source task and the target task is obtained by using the proposed knowledge transfer method, and then the dispatching knowledge of the source power grid system is transferred to the target power grid system according to the mapping, which can solve the problem that the historical dispatching knowledge cannot be directly used due to the difference in state or action space when the scale of the dispatching problem is expanded, so as to speed up the convergence speed of the algorithm, improve the learning efficiency, and reduce the learning cost.

[0004] To achieve the above purpose, the application adopts the following technical solutions:

[0005] The power grid dispatching fast optimization method based on knowledge matrix cross-dimension transfer, characterized in that the method comprises the following steps,

[0006] Step 1, first, a source power grid system dispatching optimization model without flexible resources is constructed, and multiple sets of historical typical wind, light and load data are applied to the source power grid system without considering flexible resources to perform pre-learning by using a Q learning algorithm, to obtain the Q matrix after learning optimization of each typical day as the historical knowledge matrix of power grid dispatching and form a source task library;

[0007] Step 2, the flexible resources include deep peak shaving units and cuttable loads, a deep peak shaving model and a cuttable load response model are established, and a target power grid system dispatching optimization model considering flexible resources is constructed;

[0008] Step 3, in the day-ahead scheduling stage, obtain the load demand, wind power output and photovoltaic output data of the power grid system in the future day as the target task;

[0009] Step 4, measure the similarity between the source task and the target task, calculate the net load prediction curve of the scheduling day using the predicted data of the load, wind power and photovoltaic power, select the net load as the similarity correlation feature between the tasks, measure the similarity between the net load of the target task and the net load of each typical day in the source task library based on the Euclidean distance and DTW distance, and find the source task with the minimum distance to the target task in the source task library;

[0010] Step 5, decompose the Q matrix of the source task with the minimum distance to the target task and the Q matrix of the target task into state space feature matrix and action space feature matrix respectively, find the state and action in the source task with the minimum distance to the state and action in the target task based on PCA dimension reduction and Euclidean distance, that is, construct the mapping relationship between the similar state and action in the source task and the target task;

[0011] Step 6, initialize the Q matrix of the source task with the minimum distance to the target task to the Q matrix of the target task to complete the scheduling knowledge transfer, and then optimize and solve the target task using the Q learning algorithm, finally obtain the optimal scheduling plan of the target task, which effectively improves the learning efficiency of the target task.

[0012] Further optimization of the technical solution, the source power grid system does not contain elastic resources, and the internal resources include thermal power units, wind power units and photovoltaic power units.

[0013] Further optimization of the technical solution, the target power grid system introduces source-load bilateral elastic resources to participate in optimal scheduling, and the internal resources include thermal power units, deep peak shaving units, wind power units, photovoltaic power units and reducible load.

[0014] Further optimization of the technical solution, the deep peak shaving model and the reducible load response model are as follows,

[0015] Deep peak shaving model:

[0016] Deep peak shaving can be divided into non-oil deep peak shaving and oil deep peak shaving according to the peak shaving degree of thermal power unit output power, and the operation cost of thermal power unit in different peak shaving stages can be expressed in segments as:

[0017]

[0018] Where, P g is the output power of the thermal power unit, C coal (P g ), C life (Pg ), C oil (P g ) are the coal consumption cost, life consumption cost and additional oil injection cost of the thermal power unit, respectively, P g,min , P g,max are the minimum technical output and the maximum technical output of the thermal power unit; P a is the non-oil injection deep peak shaving stable combustion load value of the thermal power unit; P b is the oil injection deep peak shaving stable combustion limit load value of the thermal power unit, when the system requires the thermal power unit to output lower than P a , oil injection combustion support is needed.

[0019] Load curtailment response model:

[0020] Through giving the user incentive compensation, the flexible adjustable space of the load side response is excavated. The function relationship between the load curtailment amount P r,t and the incentive compensation price ε r,t at time t can be approximately expressed as:

[0021]

[0022] Assuming that from time T to T+△T, the compensation cost of the curtailed load curtailment amount can be expressed as:

[0023]

[0024] Wherein, P r,max is the maximum flexible adjustable amount of the curtailed load, ε r,min , ε r,max are the minimum incentive compensation price and the maximum incentive compensation price, respectively, and α, β, γ are the relationship curve parameters.

[0025] Further optimization of the technical solution, the similarity measure of the source task and the target task is performed according to the following steps:

[0026] Step 5.1: In the power grid operation, considering the fluctuation of load demand and new energy consumption, the wind power and photovoltaic output are regarded as reverse load, and the similarity of the source task and the target task is represented by the distance between the net load curve P nl (i.e. the difference between the load demand and the new energy output at each time), then in the source task φ and the target task ψ, the net load power of the kth decision cycle is respectively:

[0027]

[0028]

[0029] Wherein, P load,k , P wind,k , P pv,kThe predicted power of the load demand, the predicted power of the wind power generation and the predicted power of the photovoltaic power generation in the kth decision period, respectively.

[0030] Step 5.2: Calculate the net load time series according to formula (6) The Euclidean distance between to represent the numerical similarity of the net load between the source task and the target task:

[0031]

[0032] Step 5.3: Calculate the dynamic time warping distance (DTW) between and to represent the morphological similarity of the net load between the source task and the target task:

[0033] Construct a KxK matrix Γ between and , where each element Γ(i,j) in the matrix is the distance between the two points of the two sequences, generally using the Euclidean distance, that is,

[0034] The path from (1,1) to (K,K) in the matrix Γ is called a bending path, denoted as: W={w1,w2,…,w L},K≤L≤2K-1,W∈Ω,for each bending path W in the bending path set Ω, the following constraints need to be met:

[0035]

[0036] Where L is the total number of elements (i,j) in the path. The purpose of the DTW algorithm is to find an optimal bending path that minimizes the cumulative distance on this path, and the cumulative distance is the DTW distance between and , as shown in formula (8).

[0037]

[0038] Step 5.4: Combine the Euclidean distance and the DTW distance to comprehensively measure the similarity between the source task and the target task, calculate the distance of the net load between the source task and the target task according to formula (9), and the smaller the distance, the more similar the tasks:

[0039]

[0040] Where λ e , λ d are the weight coefficients of the Euclidean distance and the DTW distance, respectively, λe +λ d =1.

[0041] The further optimization of the technical solution further optimizes the cross-dimension migration method of the knowledge matrix between the source task and the target task, and comprises the following steps:

[0042] Step 6.1, for the Q learning algorithm, the value of Q(s, a) is used to guide the selection of the action a in the state s, and the Q matrix can be regarded as the knowledge matrix of the current scheduling task, and it is assumed that the knowledge matrix of the source task with the minimum distance is Q φ , with a size of m*n, the state space feature matrix is S φ =(s φ,1 ,s φ,2 ,…,s φ,m ) T , and the action space feature matrix is A φ =(a φ,1 ,a φ,2 ,…,a φ,n ) T ; similarly, the knowledge matrix of the target task is Q ψ , with a size of p*q, the state space feature matrix is S ψ =(s ψ,1 ,s ψ,2 ,…,s ψ,p ) T , and the action space feature matrix is A ψ =(a ψ,1 ,a ψ,2 ,…,a ψ,q ) T . Wherein, s φ,i , a φ,i , s ψ,j , a ψ,j are feature vectors composed of information contained in each state and action of the source task and the target task;

[0043] Step 6.2, if the state feature vector s φ,i of the source task and the state feature vector s ψ,j of the target task are of the same dimension, go to step 6.3; otherwise, first unify the state feature vectors of the source task and the target task to the same dimension, and the state feature matrix and the action feature matrix of the target task after dimension reduction are S′ ψ =(s′ ψ,1 ,s′ ψ,2 ,…,s′ ψ,p ) T and A′ ψ =(a′ ψ,1 ,a′ ψ,2 ,…,a′ ψ,q )T ;

[0044] Step 6.3: Calculate the Euclidean distance matrix D between the source task and target task state feature vectors. S (s φ,i ,s′ ψ,j ),i=1,2,…,m,j=1,2,…,p;

[0045] Step 6.4: Find the relationship between the state s′ of each target task and the target task state. ψ,j The source task state with the smallest distance s φ,i Save s φ,i In D S (s φ,i ,s′ ψ,j The corresponding row index value in ) s,j =i, and the index value is s. φ,i The state number in the source task. Simultaneously, obtain the state index vector Index. S =(index) s,1 ,index s,2 ,…,index s,p ), representing the state s′ in the Q matrix of the target task. ψ,j That is, s ψ,j The state with the minimum distance is the row index of the source task's Q matrix. s,j The state;

[0046] Step 6.5: Similarly, execute steps 6.2 to 6.4 to obtain the Euclidean distance matrix D between the action feature vectors of the source task and the target task. A (a φ,i ,a′ ψ,j ), i = 1, 2, ..., n, j = 1, 2, ..., q, and the action index vector Index A =(index) a,1 ,index a,2 ,…,index a,q ), representing the action a′ in the target task Q matrix. ψ,j That is, a ψ,j The action with the smallest distance is the column index of the source task's Q matrix. a,j The state;

[0047] Step 6.6: Transfer the scheduling knowledge of the source task to the target task, that is, construct the index matrix Index using the obtained state index vector and action index vector according to equation (10). p×q According to Index p×q Use the Q value corresponding to the Q matrix of the source task as the initial Q value of the target task.

[0048]

[0049] The further optimization of the technical solution is that the PCA dimension reduction method is used to unify the state / action feature vectors of the source task and the target task to the same dimension.

[0050] The further optimization of the technical solution is that the Q learning algorithm is as follows:

[0051] Step a: initialize system model parameters, learning parameters, current learning step step = 0, current decision period k = 0, and Q matrix;

[0052] Step b: determine the state of the system at the decision time t k , select the action by using the greedy strategy according to the current state and the Q matrix, and calculate the total operation cost C total,k of the system in a decision period k according to formula (11), and update the Q matrix of the target task;

[0053] C total,k = C g,k + C r,k + C a,k + C b,k (11)

[0054] Wherein, C g,k , C r,k , C a,k , C b,k are the operation cost of the thermal power unit, the curtailed load compensation cost, the wind and light abandoned penalty, and the load loss penalty in the decision period k respectively;

[0055] Step c: let k = k + 1, if k < K, return to step b; otherwise, let k = 0;

[0056] Step d: let step = step + 1, if step < M, M is the total learning step, update the exploration rate and return to step b; otherwise, end the program.

[0057] Compared with the prior art, the above technical solution has the following beneficial effects: for the day-ahead optimal dispatching problem of the power system, based on the reinforcement learning algorithm, the application provides a mapping method between knowledge matrices in different dimensions, and the dispatching experience of the historical task is utilized. Different from the knowledge transfer of general machine learning, the transfer mechanism of the application solves the problem that the original historical dispatching knowledge cannot be directly utilized when the scale of the dispatching problem is expanded, the state and action space is different, or the dimension is changed. When the dispatching task or the state and action dimension of the power grid is changed, the transfer mechanism of the application can effectively utilize the historical dispatching knowledge, speed up the solution of the target task, avoid repeated learning, and reduce the learning cost. BRIEF DESCRIPTION OF DRAWINGS

[0058] Figure 1 A schematic diagram of the source power grid system's scheduling resources and the target power grid system's internal resources;

[0059] Figure 2 This is a schematic diagram of the power system optimization scheduling algorithm based on transfer reinforcement learning.

[0060] Figure 3 This is a schematic diagram of the knowledge matrix transfer method. Detailed Implementation

[0061] To explain in detail the technical content, structural features, objectives, and effects of the technical solution, the following description is provided in conjunction with specific embodiments and accompanying drawings.

[0062] This invention discloses a fast optimization method for power grid scheduling based on cross-dimensional transfer of knowledge matrix. Considering the inclusion of various elastic resources on both the source and load sides, the state-action space of reinforcement learning algorithms will undoubtedly increase, easily leading to problems such as the "curse of dimensionality," slow solution, and the inability to directly utilize historical scheduling knowledge. Therefore, this invention incorporates elastic resources on both the source and load sides into the power grid scheduling problem and utilizes historical scheduling knowledge through transfer learning to accelerate the solution of the target scheduling optimization problem.

[0063] This invention is specifically illustrated using a regional power grid system as an example. The invention will be further described in detail below with reference to specific embodiments and accompanying drawings. See also... Figure 1 The diagram illustrates the scheduling resources of the source grid system and the internal resources of the target grid system. The internal resources of the source grid system include: thermal power units, photovoltaic units, and wind turbines; the internal resources of the target grid system include: thermal power units, deep peak-shaving units, photovoltaic units, wind turbines, and load shedding. Among these, deep peak-shaving units are thermal power units with deep peak-shaving capabilities, which can be obtained through flexibility modifications to existing thermal power units. See also... Figure 2 The diagram shows a flowchart of a power system optimization scheduling algorithm based on transfer reinforcement learning. Taking the optimization algorithm flow of the target task as an example, the power grid scheduling optimization and knowledge matrix transfer method based on reinforcement learning in this embodiment includes the following steps:

[0064] Step 1: Assume the load demand forecast value at time t within the scheduling day is P. load,t The predicted wind power output is P. wind,i The predicted photovoltaic output value is P pv,t Thermal power units are classified into three categories (I, II, and III) based on their operating costs. In the source grid system, all three categories of thermal power units are thermal power units. In the target grid system, Category III thermal power units are designated as deeply peak-shaving units. At time t, the output power of the three categories of thermal power units is Pi. g1,t P g2,t P g3,t .

[0065] Step 2: Establishing the deep peak shaving model and the cuttable load response model,

[0066] Deep peak shaving model:

[0067] Deep peak shaving can be divided into non-oil deep peak shaving and oil deep peak shaving according to the degree of power reduction of thermal power units. The operation cost of thermal power units in different peak shaving stages can be expressed in segments as follows:

[0068]

[0069] Where, P g is the power output of the thermal power unit, C coal (P g ), C life (P g ), and C oil (P g ) are the coal consumption cost, life consumption cost, and additional oil injection cost of the thermal power unit, respectively, P g,min and P g,max are the minimum technical power and the maximum technical power of the thermal power unit; P a is the non-oil deep peak shaving stable combustion load value of the thermal power unit; P b is the oil deep peak shaving stable combustion limit load value of the thermal power unit, and when the system requires the thermal power unit to peak at a power lower than P a , oil combustion assistance is needed.

[0070] Cuttable load response model:

[0071] By giving users incentive compensation, the flexible adjustable space of load side response is excavated. The function relationship between the cuttable load P r,t at time t and the incentive compensation price ε r,t can be approximately expressed as:

[0072]

[0073] Assuming that from time T to T+ΔT, the compensation cost of the cuttable load reduction can be expressed as:

[0074]

[0075] Where, P r,max is the maximum flexible adjustable amount of the cuttable load, ε r,min and ε r,max are the minimum incentive compensation price and the maximum incentive compensation price, respectively, and α, β, and γ are the relationship curve parameters.

[0076] Step 3: Divide a scheduling day into K decision cycles from 0 to K-1. The decision time for the k-th (0≤k<K) decision cycle is t. k Discretize the allowable output power range of Class I thermal power units into 0 to N. g1 Total N g1 +1 level, discretizing the power variation range of Class I thermal power units within the climbing constraint range -N c1 ~N c1 For a total of 2N c1 +1 level, then the decision time t k The power generation capacity of Class I thermal power units is n g1,k The power adjustment level is n c1,k Similarly, the permissible output power range of Class II thermal power units is discretized into 0 to N. g2 Total N g2 +1 level, discretizing the power variation range of Category II thermal power units within the climbing constraint range -N c2 ~N c2 For a total of 2N c2 +1 level, then the decision time t k The power generation level of Class II thermal power units is n g2,k The power adjustment level is n c2,k The permissible output power range of Class III thermal power units is discretized into 0 to N. g3 Total N g3 +1 level, discretizing the power variation range of Category III thermal power units within the climbing constraint range -N c3 ~N c3 For a total of 2N c3 +1 level, then the decision time t k The power generation capacity of Class III thermal power units is n g3,k The power adjustment level is n c3,k .

[0077] Step 4: To explore the flexible adjustment space for load reduction, the incentive compensation price is set to 0~N. r Total N r +1 levels, where level 0 indicates no incentive at that moment, then the decision time t k The incentive level is n r,k The compensation price is ε r,k Then, determine the load reduction amount as P according to formula (2). r,k .

[0078] Step 5: At decision time t k, the state feature vector of the source grid system can be represented as formula (4), the action feature vector can be represented as formula (5), the state feature vector of the target grid system can be represented as formula (6), and the action feature vector can be represented as formula (7). Wherein, n g3,k 、n c3,k is the power generation level and the adjustment power level of the third type thermal power unit in the source and target grid system; is the power generation level and the adjustment power level of the third type deep peak regulation unit in the target grid system.

[0079]

[0080]

[0081]

[0082]

[0083] Step 6: determine the optimization target of the system, that is, the daily operation cost of the system is the lowest,

[0084]

[0085] Wherein, C total is the total daily operation cost, C total,k is the operation cost within a decision-making period k, C g,k , C r,k , C a,k , C b,k are the operation cost of the thermal power unit, the compensation cost of the reducible load, the penalty of abandoned wind and light, and the penalty of load loss within the decision-making period k respectively.

[0086] Step 7: initialize the system model parameters and learning parameters; take the net load as the similarity correlation feature between tasks, calculate to find the source task with the minimum distance in the source task library, and initialize the Q matrix of the target task by using the historical scheduling knowledge of the source task with the minimum distance according to the knowledge matrix migration method.

[0087] Step 8: initialize the current learning step step=0 and the current decision-making period k=0.

[0088] Step 9: determine the state of the system at the decision-making time t k , select the action by using the greedy strategy from the current state and the Q matrix, calculate the total operation cost C total,k of the system within a decision-making period k according to formula (8), and update the Q matrix of the target task.

[0089] Step 10: let k=k+1, if k

[0090] Step 11: Let step = step + 1, if step < M, M is the total number of learning steps, update the exploration rate and return to step 9; otherwise, end the program.

[0091] The source task selection and knowledge transfer in step 7 above include the following steps, and the knowledge matrix transfer method is shown in the schematic diagram as Figure 3 .

[0092] Step 7.1: In power grid operation, considering the fluctuation of load demand and new energy consumption, the wind power and photovoltaic output are regarded as reverse load, and the similarity between the source task and the target task is represented by the distance between the net load curves P nl (i.e. the difference between the load demand and the new energy output at each time), then in the source task φ and the target task ψ, the net load power of the kth decision cycle is respectively:

[0093]

[0094]

[0095] Where P load,k , P wind,k , P pv,k are the load demand prediction power, wind power prediction power and photovoltaic power prediction power of the kth decision cycle respectively.

[0096] Step 7.2: Calculate the Euclidean distance between the net load time series and to represent the numerical similarity between the source task and the target task:

[0097]

[0098] Step 7.3: Calculate the dynamic time warping distance (Dynamic TimeWarping, DTW) between the net load time series and to represent the similarity in shape between the source task and the target task:

[0099] A KxK matrix Γ between and is constructed, and each element Γ(i, j) in the matrix is the distance between the points of the two sequences, generally using the Euclidean distance, that is,

[0100] The path from (1, 1) to (K, K) in the matrix Γ is called a bending path, denoted as: W = {w1, w2, …, w L}, K≤L≤2K-1, W∈Ω, for each curved path W in the curved path set Ω, the following constraints need to be met:

[0101]

[0102] where L is the total number of elements (i, j) in the path. The purpose of the DTW algorithm is to find an optimal curved path that minimizes the cumulative distance on this path, and the cumulative distance is the DTW distance between and , as shown in equation (13):

[0103]

[0104] Step 7.4: Combine the Euclidean distance and the DTW distance to comprehensively measure the similarity between the source task and the target task, calculate the distance between the source task and the target task according to equation (14), and the smaller the distance, the more similar the tasks:

[0105]

[0106] where λ e and λ d are the weight coefficients of the Euclidean distance and the DTW distance, respectively, and λ e + λ d = 1.

[0107] Step 7.5: Define the Q matrix as a knowledge matrix containing historical scheduling experience, and assume that the knowledge matrix of the source task with the smallest distance is Q φ , with a size of m x n, the state space feature matrix is S φ =(s φ,1 , s φ,2 , …, s φ,m ) T , and the action space feature matrix is A φ =(a φ,1 , a φ,2 , …, a φ,n ) T ; Similarly, the knowledge matrix of the target task is Q ψ , with a size of p x q, the state space feature matrix is S ψ =(s ψ,1 , s ψ,2 , …, s ψ,p ) T , and the action space feature matrix is A ψ =(a ψ,1 , a ψ,2 , …, a ψ,q ) T . Where s φ,i , a φ,is ψ,j a ψ,j These are feature vectors composed of information contained in each state and action of the source task and the target task, respectively.

[0108] Step 7.6: If the state feature vector s of the source task φ,i With the state feature vector s of the target task ψ,j If the dimensions are the same, skip to step 7.7; otherwise, first unify the state feature vectors of the source task and the target task to the same dimension. The state feature matrix and action feature matrix of the target task state feature vector after dimensionality reduction are S′ respectively. ψ =(s′) ψ,1 ,s′ ψ,2 ,…,s′ ψ,p ) T A′ ψ =(a′ ψ,1 ,a′ ψ,2 ,…,a′ ψ,q ) T .

[0109] Step 7.7: Calculate the Euclidean distance matrix D between the source task and target task state feature vectors. s (s φ,i ,s′ ψ,j ), i=1, 2,..., m, j=1, 2,..., p.

[0110] Step 7.8: Find the relationship between the state s′ of each target task and the target task state. ψ,j The source task state with the smallest distance s φ,i Save s φ,i In D s (s φ,i ,s′ ψ,j The corresponding row index value in ) s,j =i, and the index value is s. φ,i The state number in the source task. Simultaneously, obtain the state index vector Index. S =(index) s,1 index s,2 , ..., index s,p ), representing the state s′ in the Q matrix of the target task. ψ,j That is, s ψ,j The state with the minimum distance is the row index of the source task's Q matrix. s,j The state.

[0111] Step 7.9: Similarly, execute steps 7.6 to 7.8 to obtain the Euclidean distance matrix D between the action feature vectors of the source task and the target task. A (a φ,i ,a′ ψ,j), i = 1, 2, …, n, j = 1, 2, …, q, and action index vector Index A = (index a,1 , index a,2 , …, index a,q ), represents the action a' that is closest to the target task Q-matrix in the source task Q-matrix. ψ,j That is, a ψ,j is the state with column index index a,j in the source task Q-matrix.

[0112] Step 7.10: Migrate the scheduling knowledge of the source task to the target task, that is, form the index matrix Index p×q according to the obtained state index vector and action index vector according to formula (15), and take the corresponding Q value in the source task Q-matrix as the initial Q value of the target task according to Index p×q .

[0113]

[0114] The present application considers that the addition of elastic resources results in that the historical scheduling knowledge of the source power grid system cannot be directly utilized, and that the state-action space is increased and re-optimized learning converges slowly. Therefore, the present application proposes a cross-dimension migration method between knowledge matrices, which can effectively utilize the historical scheduling knowledge when the power grid scheduling task or the state-action dimension changes, to a certain extent, avoid the curse of dimensionality problem, effectively accelerate the solution speed of the day-ahead scheduling plan, and save time and space resources.

[0115] It should be noted that, in the knowledge migration method and the attached Figure 3 , for the convenience of expression, the state (action) feature vector dimension of the source task is less than the state (action) feature vector dimension of the target task, which is taken as an example for illustration. In fact, the knowledge migration method in the present application is not limited to this restriction, and only needs to reduce the dimension of the feature vector in the two tasks to the same dimension as the other task.

[0116] It is to be noted that, in the present text, terms such as first and second, and the like, merely serve to identify a difference between one entity or action and another entity or action, and do not necessarily require or imply that there is any such actual relationship or order between these entities or actions. Moreover, the terms "comprising", "including", or any other variant thereof are intended to cover a non-exclusive inclusion, such that processes, methods, articles, or apparatuses that comprise a list of elements are not required to comprise only those elements, but can include other elements not expressly listed or inherent to such processes, methods, articles, or apparatuses. Without further limitation, an element preceded by "comprises a" or "comprises" does not, without more limitations, preclude the existence of further elements of the process, method, article, or apparatus that includes the element. Furthermore, in the present text, "greater than", "less than", "exceed", and the like are understood to exclude the number itself; "and above", "and below", "and within", and the like are understood to include the number itself.

[0117] Although the above-mentioned embodiments have been described, those skilled in the art can make further changes and modifications to these embodiments once they know the basic inventive concept, so the above description is only for the embodiments of the present application, and does not limit the patent protection scope of the present application, and any equivalent structure or equivalent process transformation using the content of the present application specification and drawings, or direct or indirect application in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A fast optimization method for power grid dispatching based on cross-dimensional transfer of knowledge matrix, characterized in that, The method includes the following steps: Step 1: First, construct a source power grid system scheduling optimization model without elastic resources. Apply the Q-learning algorithm to pre-learn multiple sets of historical typical wind, solar and load data in the source power grid system without considering elastic resources. The optimized Q matrix of each typical day is used as the historical knowledge matrix of power grid scheduling and constitutes the source task library. Step 2: Flexible resources include deep peak-shaving units and loads that can be reduced. Establish deep peak-shaving models and load-reducible response models to construct a target power grid system dispatch optimization model that takes into account flexible resources. Step 3: During the day-ahead dispatch phase, the target task is to obtain the load demand, wind power output, and photovoltaic power output data of the power grid system for each time period of the next day. Step 4: Measure the similarity between the source task and the target task. Calculate the net load forecast curve for the scheduling day using the load, wind power, and photovoltaic forecast data. Select the net load as the similarity correlation feature between tasks. Measure the similarity between the net load of the target task and the net load of each typical day in the source task library based on Euclidean distance and DTW distance. Find the source task with the smallest distance from the target task in the source task library. Step 5: Decompose the Q matrix of the source task with the smallest distance from the target task and the Q matrix of the target task into state space feature matrix and action space feature matrix, respectively. Based on PCA dimensionality reduction and Euclidean distance, find the state and action in the source task with the smallest distance from the state and action in the target task, that is, construct the mapping relationship between similar states and actions in the source task and the target task. Step 6: Initialize the Q matrix of the target task with the Q matrix of the source task that is closest to the target task according to the mapping relationship to complete the scheduling knowledge transfer. On this basis, the Q learning algorithm is used to optimize and solve the target task. Finally, the optimal scheduling plan for the target task scheduling day is obtained, which effectively improves the learning efficiency of the target task. The deep peak-shaving model and the load-reducible response model are as follows. Deep peak-shaving model: Deep peak shaving can be divided into deep peak shaving without oil injection and deep peak shaving with oil injection, depending on the degree of output reduction in thermal power units. The operating costs of thermal power units in different peak shaving stages can be expressed in segments as follows: (1) in, The output power of thermal power units. , , These are the coal consumption cost, lifespan loss cost, and additional oil injection cost of thermal power units, respectively. , For the minimum and maximum technical output of thermal power units; This refers to the peak-shaving and combustion-stabilizing load value for thermal power units without oil injection. The oil injection depth of thermal power units is the peak-shaving and combustion-stabilizing limit load value. When the system requires the peak-shaving output of thermal power units to be lower than the limit load value, the limit load value is determined by the limit load. At this time, oil needs to be added to assist combustion; Reduced load response model: By providing users with incentives and compensation, we can explore the flexibility and adjustability of the load-side response. Reduce load at all times With incentive compensation price The functional relationship can be approximated as: (2) Assuming from Time running to The compensation cost for reducing load can be expressed as: (3) in, To reduce the maximum adjustable load, , These are the minimum incentive compensation price and the maximum incentive compensation price, respectively. , , These are the parameters of the relationship curve.

2. The fast optimization method for power grid dispatch based on cross-dimensional transfer of knowledge matrix as described in claim 1, characterized in that, The power grid system does not include flexible resources; its internal resources include thermal power units, wind power units, and photovoltaic power units.

3. The fast optimization method for power grid dispatch based on cross-dimensional transfer of knowledge matrix as described in claim 1, characterized in that, The target power grid system incorporates flexible resources on both the source and load sides for optimized scheduling. Its internal resources include thermal power units, deep peak-shaving units, wind power units, photovoltaic power units, and loads that can be reduced.

4. The fast optimization method for power grid dispatch based on cross-dimensional transfer of knowledge matrix as described in claim 1, characterized in that, The similarity measurement between the source task and the target task is performed according to the following steps: Step 5.1: In grid operation, considering load demand fluctuations and renewable energy absorption, wind power and solar power output are treated as reverse loads, and the net load curve is used to determine the load. The distance between the source and target tasks is used to characterize the similarity between them, and the net workload curve is shown. That is, the difference between load demand and renewable energy output at each moment is the source task. and target tasks In the middle, the first The net load power for each decision cycle is as follows: (4) (5) in, , , The first Forecast power of load demand, wind power generation, and photovoltaic power generation for each decision cycle; Step 5.2: Calculate the net load time series according to equation (6). and The Euclidean distance between the source and target tasks is used to characterize the numerical similarity of their net workloads. (6) Step 5.3: Calculate the net load time series and The dynamic time warping (DTW) distance between the source and target tasks is used to characterize the morphological similarity of the payloads between them. Build a and The size between matrix Each element in the matrix The distance between points in two sequences is calculated using Euclidean distance, i.e. , matrix Zhong Cong arrive The path is called a curved path, denoted as: For the set of curved paths Each curved path The following constraints must be met: (7) in, For elements in the path The total number of curves is given. The goal of the DTW algorithm is to find an optimal curved path that minimizes the cumulative distance along this path. The cumulative distance is... and DTW distance between them (8) Step 5.4: Combine Euclidean distance and DTW distance to comprehensively measure the similarity between the source task and the target task. Calculate the net load distance between the source task and the target task according to equation (9). The smaller the distance, the more similar the tasks are. (9) in , These are the weighting coefficients for the Euclidean distance and the DTW distance, respectively. .

5. The fast optimization method for power grid dispatch based on cross-dimensional transfer of knowledge matrix as described in claim 1, characterized in that, The cross-dimensional transfer method of the knowledge matrix between the source task and the target task includes the following steps: Step 6.1: For the Q-learning algorithm, in state... The following is the use of Use values ​​to guide actions The choice of the Q matrix can be seen as the knowledge matrix of the current scheduled task. Assuming the knowledge matrix of the task with the minimum distance to the source task is... Size is The state-space feature matrix is The action space feature matrix is Similarly, the knowledge matrix of the target task is: Size is The state-space feature matrix is The action space feature matrix is ,in, , , , These are feature vectors composed of information contained in each state and action of the source task and the target task, respectively. Step 6.2: If the state feature vector of the source task... State feature vector of the target task If the dimensions are the same, skip to step 6.3; otherwise, first unify the state feature vectors of the source task and the target task to the same dimension. The dimensionality-reduced state feature matrix and action feature matrix of the target task state feature vector are respectively... , ; Step 6.3: Calculate the Euclidean distance matrix between the source task and target task state feature vectors. ; Step 6.4: Identify the status of each target task. The source task state with the smallest distance ,save exist The corresponding row index value The index value is The state number in the source task is used to obtain the state index vector. , representing the state in the Q matrix of the target task Right now The state with the minimum distance is the row index of the source task's Q matrix. The state; Step 6.5: Similarly, execute steps 6.2 to 6.4 to obtain the Euclidean distance matrix between the action feature vectors of the source task and the target task. and action index vector , representing the action in the target task Q matrix Right now The action with the smallest distance is the column index of the source task's Q matrix. The state; Step 6.6: Transfer the scheduling knowledge of the source task to the target task, that is, construct an index matrix from the obtained state index vector and action index vector according to equation (10). ,according to Use the Q-value corresponding to the source task's Q-matrix as the initial Q-value for the target task. (10)。 6. The fast optimization method for power grid dispatch based on cross-dimensional transfer of knowledge matrix as described in claim 5, characterized in that, The method used to unify the state feature vectors of the source task and the target task to the same dimension is PCA dimensionality reduction.

7. The fast optimization method for power grid dispatch based on cross-dimensional transfer of knowledge matrix as described in claim 1, characterized in that, The Q-learning algorithm is as follows: Step a: Initialize system model parameters, learning parameters, and current learning steps. Current decision-making cycle as well as matrix; Step b: Determine the system's decision-making time. The state is determined by the current state and The matrix employs a greedy strategy to select actions and calculates the system's actions over a decision cycle according to equation (11). Total operating cost generated internally Update the target task matrix; (11) in, , , , Decision cycle The operating costs of thermal power units within the unit, the cost of reducing load compensation, the penalty for wind and solar curtailment, and the penalty for load loss; Step c: Let ,like If yes, return to step b; otherwise, let ; Step d: Let ,like , Update the exploration rate to the total number of learning steps and return to step b; otherwise, terminate the program.