Data-driven main-distribution multi-level power grid operation state dynamic evaluation method and device
By constructing a multi-level power grid operation status evaluation index system for the main and distribution networks and a dynamic time bending algorithm based on reinforcement learning, the problem that existing technologies are unable to adapt to complex power grid systems is solved, and accurate evaluation and classification of the main and distribution network operation status are achieved, adapting to the multi-level evaluation needs of the power grid system.
Patent Information
- Application Number
- CN202510942722.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-09-12
AI Technical Summary
The existing main-distribution multi-level power grid operation status assessment method cannot effectively consider multi-level and multi-dimensional influencing indicators, resulting in an inability to adapt to the complex operating conditions, strong fluctuations, and large amounts of heterogeneous data in actual power grid systems, and unable to guarantee voltage stability and power supply.
A multi-level power grid operation status evaluation index system for the main and distribution networks is constructed. The operation status evaluation data is collected, the weight of each indicator is calculated, and the dynamic time warping algorithm based on reinforcement learning is used to solve it. The classification evaluation algorithm of reinforcement learning and dynamic time warping based on good and bad solutions is integrated to achieve accurate evaluation of the main and distribution network operation status.
With fewer training samples, it can adapt to complex power grid systems, achieve accurate evaluation of the operating status of the main and distribution networks, solve decision-making problems of multi-index sequences, and adapt to the different needs of the power grid system in terms of economy, reliability and flexibility.
Smart Images

Figure CN120634366A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of main-distribution multi-level power grid operating status assessment, and in particular to a data-driven main-distribution multi-level power grid operating status dynamic assessment method and device. Background Art
[0002] With the increasing proportion of renewable energy and the increasingly complex operating environment of power systems, the operation of multi-tiered power grids, including those involving main and distribution networks, faces unprecedented challenges. The coordinated operation of these networks is crucial to the overall security and economic viability of the power system. In particular, with a high proportion of distributed power sources connected, operational uncertainty in the grid increases significantly, making it difficult to guarantee voltage stability and power output. Therefore, establishing a multi-tiered operational status assessment system for main and distribution networks and proposing an assessment method for this status are of great theoretical and practical significance.
[0003] The existing main-distribution operation status assessment method cannot consider the multi-level and multi-dimensional influencing indicators of the assessment object, which makes it unable to adapt to the characteristics of complex operating conditions, strong fluctuations, and large amounts of heterogeneous data in actual power grid systems.
[0004] The information disclosed in this background technology section is only intended to deepen the understanding of the overall background technology of the present invention and should not be regarded as an admission or any form of suggestion that the information constitutes the prior art already known to those skilled in the art. Summary of the Invention
[0005] The present invention provides a data-driven method and device for dynamically evaluating the operating status of a main-distribution multi-level power grid, thereby effectively solving the problems in the background technology.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is: a data-driven method for dynamically evaluating the operating status of a main-distribution multi-level power grid, comprising the following steps:
[0007] Construct a multi-dimensional operating status evaluation index system for the main and distribution networks. The main network level includes the main network operating cost indicators, supply and demand balance indicators, steady-state frequency limit indicators and flexibility indicators; at the distribution network level, it mainly includes the distribution network operating cost indicators, branch power limit indicators, flexibility indicators and new energy absorption indicators.
[0008] and collecting main-distribution network operation status evaluation data according to the evaluation index system to construct an evaluation matrix;
[0009] Calculating the weight of each indicator in the evaluation indicator system, and constructing a weighted decision matrix based on the evaluation matrix and the weight;
[0010] The weighted decision matrix is solved using a dynamic time warping algorithm based on reinforcement learning to obtain a comprehensive score value, and the main-distribution network operating status to be evaluated is classified according to the comprehensive score value.
[0011] Furthermore, the calculation method of the main network layer indicators is as follows:
[0012] Mainnet operating cost indicators:
[0013] E1=C st +C fuel ;
[0014] Where, E1 is the operating cost of the main network; C st is the unit start-up and shutdown cost; C fuel fuel costs for electricity generation;
[0015] Mainnet supply and demand balance indicators:
[0016]
[0017] Where, E fle is the supply and demand balance indicator of the main network, ΔP l (t) is the flexibility demand value of the system at time t, ΔP g (t) is the flexibility supply value of the system at time t;
[0018] Main network steady-state frequency limit exceeding indicator:
[0019]
[0020] Where R f is the main network steady-state frequency limit-exceeding index, |Δf| is the system steady-state frequency deviation, Δf max It is the maximum steady-state frequency deviation that can be accepted during normal system operation.
[0021] Mainnet flexibility indicators:
[0022]
[0023] Where AF m is the flexibility index of the main network at time t; P g (t) is the output of the controllable traditional unit at time t, when M is +, it indicates an upward adjustment characteristic, and when it is -, it indicates a downward adjustment characteristic; Respectively represent the upward and downward climbing rates of the unit; P g,max 、P g.min Indicates the upper and lower limits of the unit's output.
[0024] The calculation methods for the indicators of the distribution network layer are as follows:
[0025] Distribution network operating cost indicators:
[0026]
[0027] Where, E2 is the operating cost of the distribution network; C t P C is the cost of electricity interaction between the main grid and the distribution grid; RE P is the total operating cost of new energy for the distribution network.
[0028] Distribution network branch power over-limit indicators:
[0029]
[0030] Where R L N is the power limit indicator of the distribution network branch, L is the total number of system lines, P l is the active power transmitted by line l, is the maximum active power allowed to be transmitted on line 1.
[0031] And distribution network flexibility indicators:
[0032] AF D =P l (t+τ)-P l (t);
[0033] Where AF D is the flexibility index of the distribution network system at time t, ΔP l (t+τ) is the net load size at time t+τ, and τ is the response time scale;
[0034] New energy consumption indicators:
[0035]
[0036] Where R RE is the main grid new energy consumption index, ∑P RE_i is the total actual output of new energy, ∑P RE_j The total output for the new energy ideal.
[0037] Furthermore, constructing the evaluation matrix includes:
[0038] The collected main-distribution network operation status evaluation data is formed into reference and evaluation sample sequences, and the matrix x is constructed. ij ;
[0039] After normalizing the matrix, we get the evaluation matrix:
[0040]
[0041] Where μj is the mean of the j-th column feature, σ j is the standard deviation of the j-th column feature.
[0042] Furthermore, the calculation of the weight of each indicator in the evaluation indicator system includes:
[0043] The entropy weight method is used to calculate the entropy weight of each indicator:
[0044]
[0045] Where, f ij 、y ij is the normalized value and weight of the i-th sample on the j-th indicator, m and n are the number of samples and the number of indicators respectively; α j 、H j 、w j and σ j is the adjustment factor, information entropy, entropy weight and standard deviation of the j-th indicator.
[0046] Furthermore, constructing a weighted decision matrix based on the evaluation matrix and weights includes:
[0047] z ij =w j x′ ij ;
[0048] Where Z is the weighted decision matrix for main-distribution network operation status evaluation.
[0049] Furthermore, solving the weighted decision matrix using a dynamic time warping algorithm based on reinforcement learning comprises the following steps:
[0050] Extract the maximum and minimum values from each sequence of the weighted decision matrix and construct the optimal sequence Z + and the worst solution sequence Z - ;
[0051] Set up the state space of reinforcement learning, the action space and reward function based on the dynamic time warping algorithm;
[0052] Calculate the Z of the sequence to be evaluated and the reference sequence + and Z - Dynamic time warping distance;
[0053] The comprehensive score value of the sequence to be evaluated is calculated according to the dynamic time warping distance.
[0054] Furthermore, the state space S is:
[0055] S = [(i, j)];
[0056] Where i represents the current position of the sequence to be evaluated U, and j represents the current position of the reference sequence V;
[0057] The action space A is:
[0058] A=[(i+1,j+1),(i+1,j),(i,j+1)];
[0059] The reward function R is:
[0060]
[0061] Where s represents the boundary of u, t represents the boundary of v, ρ∈{ρ d ,ρ v ,ρ h} is used as a weight factor to penalize or support the change of direction, ρ∈{ρ d ,ρ v ,ρ h}={1,1,1}, indicating that all directions have equal probability.
[0062] Furthermore, the calculation of the sequence to be evaluated and the reference sequence Z + and Z - Dynamic time warping distances, including:
[0063]
[0064] Where D i + 、D i - Optimal paths generated for reinforcement learning;
[0065] Calculating a comprehensive score value of the sequence to be evaluated according to the dynamic time warping distance includes:
[0066]
[0067] Where S i It is the comprehensive score of the sequence to be evaluated based on the dynamic time warping of the good and bad solutions.
[0068] The present invention also includes a data-driven device for dynamically evaluating the operating status of a main-distribution multi-level power grid, using the above-mentioned method, the device comprising:
[0069] An evaluation matrix construction unit is used to construct a multi-dimensional operating status evaluation index system for the main power grid, and collect main-distribution network operating status evaluation data according to the evaluation index system to construct an evaluation matrix;
[0070] A weighting unit, configured to calculate a weight for each indicator in the evaluation indicator system and construct a weighted decision matrix based on the evaluation matrix and the weight;
[0071] A solving unit is used to solve the weighted decision matrix using a dynamic time warping algorithm based on reinforcement learning to obtain a comprehensive scoring value, and classify the main-distribution network operating status to be evaluated according to the comprehensive scoring value.
[0072] The present invention also includes a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above method when executing the computer program.
[0073] The present invention also includes a storage medium storing a computer program, which implements the above method when executed by a processor.
[0074] The beneficial effects of the present invention are: by constructing a main-distribution network operation status evaluation index system to target the differences in the main-distribution multi-level power grid's requirements for economy, reliability, and flexibility, and targeting the multi-level and multi-dimensional influencing indicators of the evaluated objects, the weights of each indicator in the evaluation index system are calculated, and a weighted decision matrix is constructed based on the evaluation matrix and the weights. A classification evaluation algorithm that integrates reinforcement learning and dynamic time bending of superior and inferior solutions is proposed. This algorithm can not only solve the decision-making problem of multi-indicator sequences, but also adapt to the characteristics of complex working conditions, strong fluctuations, and a large amount of heterogeneous data in actual power grid systems. It can achieve effective classification with fewer training samples and realize accurate evaluation of the main-distribution network operation status. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0076] Figure 1 is a flow chart of the method in Example 1;
[0077] Figure 2 Schematic diagram of the structure of the device in Example 1;
[0078] Figure 3 This is the main-distribution network topology diagram in Example 2;
[0079] Figure 4 This is the minimum bending distance result diagram of the sample to be evaluated in Example 2;
[0080] Figure 5 A schematic diagram of the structure of a computer device. DETAILED DESCRIPTION
[0081] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0082] Example 1:
[0083] like Figure 1 A data-driven method for dynamically evaluating the operating status of a main-distribution multi-level power grid includes the following steps:
[0084] Construct a multi-dimensional operating status evaluation index system for the main-distribution network, collect main-distribution network operating status evaluation data based on the evaluation index system, and construct an evaluation matrix;
[0085] Calculate the weight of each indicator in the evaluation index system, and construct a weighted decision matrix based on the evaluation matrix and weight;
[0086] The dynamic time warping algorithm based on reinforcement learning is used to solve the weighted decision matrix to obtain a comprehensive score value, and the main-distribution network operating status to be evaluated is classified according to the comprehensive score value.
[0087] By constructing a main-distribution network operation status evaluation index system to target the differences in the economy, reliability, and flexibility requirements of the main-distribution multi-level power grid, and targeting the multi-level and multi-dimensional influencing indicators of the evaluated objects, the weights of each indicator in the evaluation index system are calculated, and a weighted decision matrix is constructed based on the evaluation matrix and weights. A classification evaluation algorithm that integrates reinforcement learning and dynamic time bending of the best and worst solutions is proposed. This algorithm can not only solve the decision-making problem of multi-indicator sequences, but also adapt to the characteristics of complex working conditions, strong fluctuations, and a large amount of heterogeneous data in actual power grid systems. It can achieve effective classification with fewer training samples and realize accurate evaluation of the main-distribution network operation status.
[0088] In this embodiment, a multi-dimensional operating status evaluation index system for the main power grid is constructed. The main grid level includes the main grid operating cost index, supply and demand balance index, steady-state frequency over-limit index and flexibility index; the distribution network level mainly includes the distribution network operating cost index, branch power over-limit index, flexibility index and new energy absorption index.
[0089] Furthermore, the calculation method of the main network layer indicators is as follows:
[0090] Mainnet operating cost indicators:
[0091] E1=C st +C fuel ;
[0092] Where, E1 is the operating cost of the main network; C st is the unit start-up and shutdown cost; Cfuel fuel costs for electricity generation;
[0093] Mainnet supply and demand balance indicators:
[0094]
[0095] Where, E fle is the supply and demand balance indicator of the main network, ΔP l (t) is the flexibility demand value of the system at time t, ΔP g (t) is the flexibility supply value of the system at time t;
[0096] Main network steady-state frequency limit exceeding indicator:
[0097]
[0098] Where R f is the main network steady-state frequency limit-exceeding index, |Δf| is the system steady-state frequency deviation, Δf max It is the maximum steady-state frequency deviation that can be accepted during normal system operation.
[0099] Mainnet flexibility indicators:
[0100]
[0101] Where AF m is the flexibility index of the main network at time t; P g (t) is the output of the controllable traditional unit at time t, when M is +, it indicates an upward adjustment characteristic, and when it is -, it indicates a downward adjustment characteristic; Respectively represent the upward and downward climbing rates of the unit; P g,max 、P g.min Indicates the upper and lower limits of the unit's output.
[0102] The calculation methods for the indicators of the distribution network layer are as follows:
[0103] Distribution network operating cost indicators:
[0104]
[0105] Where, E2 is the operating cost of the distribution network; C t P C is the cost of electricity interaction between the main grid and the distribution grid; RE P is the total operating cost of new energy for the distribution network.
[0106] Distribution network branch power over-limit indicators:
[0107]
[0108] Where R LN is the power limit indicator of the distribution network branch, L is the total number of system lines, P l is the active power transmitted by line l, is the maximum active power allowed to be transmitted on line 1.
[0109] And distribution network flexibility indicators:
[0110] AF D =P l (t+τ)-P l (t);
[0111] Where AF D is the flexibility index of the distribution network system at time t, ΔP l (t+τ) is the net load size at time t+τ, and τ is the response time scale;
[0112] New energy consumption indicators:
[0113]
[0114] Where R RE is the main grid new energy consumption index, ∑P RE_i is the total actual output of new energy, ∑P RE_j The total output for the new energy ideal.
[0115] The construction of the evaluation matrix includes:
[0116] The collected main-distribution network operation status evaluation data is formed into reference and evaluation sample sequences, and the matrix x is constructed. ij ;
[0117] After normalizing the matrix, we get the evaluation matrix:
[0118]
[0119] Where μ j is the mean of the j-th column feature, σ j is the standard deviation of the j-th column feature.
[0120] In this embodiment, the weights of the various indicators in the evaluation indicator system are calculated, including:
[0121] The entropy weight method is used to calculate the entropy weight of each indicator:
[0122]
[0123]
[0124] Where, f ij 、y ijis the normalized value and weight of the i-th sample on the j-th indicator, m and n are the number of samples and the number of indicators respectively; α j 、H j 、w j and σ j is the adjustment factor, information entropy, entropy weight and standard deviation of the j-th indicator.
[0125] Construct a weighted decision matrix based on the evaluation matrix and weights, including:
[0126] z ij =w j x′ ij ;
[0127] Where Z is the weighted decision matrix for main-distribution network operation status evaluation.
[0128] The weighted decision matrix is solved using a dynamic time warping algorithm based on reinforcement learning, which includes the following steps:
[0129] Extract the maximum and minimum values from each sequence of the weighted decision matrix and construct the optimal sequence Z + and the worst solution sequence Z - ;
[0130] Set up the state space of reinforcement learning, the action space and reward function based on the dynamic time warping algorithm;
[0131] Calculate the Z of the sequence to be evaluated and the reference sequence + and Z - Dynamic time warping distance;
[0132] The comprehensive score of the sequence to be evaluated is calculated based on the dynamic time warping distance.
[0133] The state space S is:
[0134] S = [(i, j)];
[0135] Where i represents the current position of the sequence to be evaluated U, and j represents the current position of the reference sequence V;
[0136] The action space A is:
[0137] A=[(i+1,j+1),(i+1,j),(i,j+1)];
[0138] The reward function R is:
[0139]
[0140] Where s represents the boundary of u, t represents the boundary of v, ρ∈{ρ d ,ρv ,ρ h} is used as a weight factor to penalize or support the change of direction, ρ∈{ρ d ,ρ v ,ρ h}={1,1,1}, indicating that all directions have equal probability.
[0141] As a preferred embodiment of the above, the sequence to be evaluated and the reference sequence Z are calculated. + and Z - Dynamic time warping distances, including:
[0142]
[0143] Where D i + 、D i - Optimal paths generated for reinforcement learning;
[0144] The comprehensive score of the sequence to be evaluated is calculated based on the dynamic time warping distance, including:
[0145]
[0146] Where S i It is the comprehensive score of the sequence to be evaluated based on the dynamic time warping of the good and bad solutions.
[0147] like Figure 2 As shown, this embodiment also includes a data-driven main-distribution multi-level power grid operating status dynamic evaluation device, using the above method, the device includes:
[0148] An evaluation matrix construction unit is used to construct a multi-dimensional operating status evaluation index system for the main power grid, and collect main-distribution network operating status evaluation data based on the evaluation index system to construct an evaluation matrix;
[0149] The weighting unit is used to calculate the weight of each indicator in the evaluation indicator system and construct a weighted decision matrix based on the evaluation matrix and weights;
[0150] The solving unit is used to solve the weighted decision matrix using a dynamic time warping algorithm based on reinforcement learning to obtain a comprehensive score value, and classify the main-distribution network operating status to be evaluated according to the comprehensive score value.
[0151] Example 2:
[0152] This embodiment discloses a data-driven method for dynamically evaluating the operating status of a main-distribution multi-level power grid, including the following steps:
[0153] Step 1: Build an evaluation indicator system and standardize data;
[0154] Step 11: Construct a multi-dimensional operating status evaluation index system for the main distribution network. The main network level includes main network operating cost indicators, supply and demand balance indicators, steady-state frequency limit indicators, and flexibility indicators. The distribution network level mainly includes distribution network operating cost indicators, branch power limit indicators, flexibility indicators, and new energy consumption indicators.
[0155] Furthermore, the calculation method of the main network layer indicators is as follows:
[0156] Mainnet operating cost indicators:
[0157] E1=C st +C fuel ;
[0158] Where, E1 is the operating cost of the main network; C st is the unit start-up and shutdown cost; C fuel fuel costs for electricity generation;
[0159] Mainnet supply and demand balance indicators:
[0160]
[0161] Where, E fle is the supply and demand balance indicator of the main network, ΔP l (t) is the flexibility demand value of the system at time t, ΔP g (t) is the flexibility supply value of the system at time t;
[0162] Main network steady-state frequency limit exceeding indicator:
[0163]
[0164] Where R f is the main network steady-state frequency limit-exceeding index, |Δf| is the system steady-state frequency deviation, Δf max It is the maximum steady-state frequency deviation that can be accepted during normal system operation.
[0165] Where R f is the main network steady-state frequency limit-exceeding index, ρ is the probability of the system state, |Δf| is the system steady-state frequency deviation, Δf max is the maximum steady-state frequency deviation that can be accepted during normal system operation, and e is a natural constant;
[0166] Mainnet flexibility indicators:
[0167]
[0168] Where AF m is the flexibility index of the main network at time t; P g(t) is the output of the controllable traditional unit at time t, when M is +, it indicates an upward adjustment characteristic, and when it is -, it indicates a downward adjustment characteristic; Respectively represent the upward and downward climbing rates of the unit; P g,max 、P g.min Indicates the upper and lower limits of the unit's output.
[0169] The calculation methods for the indicators of the distribution network layer are as follows:
[0170] Distribution network operating cost indicators:
[0171]
[0172] Where, E2 is the operating cost of the distribution network; C t P C is the cost of electricity interaction between the main grid and the distribution grid; RE P is the total operating cost of new energy for the distribution network.
[0173] Distribution network branch power over-limit indicators:
[0174]
[0175] Where R L N is the power limit indicator of the distribution network branch, L is the total number of system lines, P l is the active power transmitted by line l, is the maximum active power allowed to be transmitted on line 1.
[0176] And distribution network flexibility indicators:
[0177] AF D =P l (t+τ)-P l (t);
[0178] Where AF D is the flexibility index of the distribution network system at time t, ΔP l (t+τ) is t+ τ The size of the net load at the moment, τ is the response time scale;
[0179] New energy consumption indicators:
[0180]
[0181] Where R RE is the main grid new energy consumption index, ∑P RE_i is the total actual output of new energy, ∑P RE_j The total output for the new energy ideal.
[0182] Step 12: Construct the normalized matrix for main-distribution network operation status evaluation. According to the calculation formula of the indicators in step 11, collect the main-distribution network operation status evaluation data, form the reference and evaluation sample sequences, and construct the evaluation matrix x ij , after normalization, the normalized matrix of running status evaluation is:
[0183]
[0184] Where μ j is the mean of the j-th column feature, σ j is the standard deviation of the j-th column feature.
[0185] Step 2: Calculate indicator weights and construct a weighted decision matrix;
[0186] Step 21: Calculate the weight of each indicator. The entropy weight method is used to calculate the weight of each indicator. The specific method is as follows:
[0187]
[0188] Where, f ij 、y ij is the normalized value and weight of the i-th sample on the j-th indicator, m and n are the number of samples and the number of indicators respectively; α j 、H j 、w j and σ j is the adjustment factor, information entropy, final entropy weight and standard deviation of the j-th indicator.
[0189] Step 22: Construct a weighted decision matrix for the main-distribution network operation status. Construct a weighted decision matrix based on the weights of each indicator:
[0190]
[0191] Where Z is the weighted decision matrix for main-distribution network operation status evaluation.
[0192] Step 23: Extract the optimal and inferior solution sequence of the main-distribution network operation status indicators. Extract the maximum and minimum values from each sequence and construct the optimal solution sequence Z + , the worst solution sequence Z - .
[0193] Step 3: Determine the state quality of the dynamic time warping algorithm based on reinforcement learning;
[0194] Step 31: Dynamic Time Warping Distance Calculation Based on Reinforcement Learning
[0195] The state space of reinforcement learning can be set as the current matching position in the sequence to be evaluated U and the reference sequence V:
[0196] S=[(i,j)] (10)
[0197] Here, i represents the current position of u, and j represents the current position of v.
[0198] Based on the recursive relationship of the dynamic time warping algorithm, the action space is set as:
[0199] A=[(i+1,j+1),(i+1,j),(i,j+1)] (11)
[0200] The reward function rewards the user based on the accuracy of the matching. If the selected path results in a smaller matching error, a higher reward is given; if the error is larger, a penalty is imposed.
[0201]
[0202] Where s represents the boundary of u, t represents the boundary of v, ρ∈{ρ d ,ρ v ,ρ h} is used as a weight factor to penalize or support the change of direction, and ρ∈{ρ d ,ρ v ,ρ h}={1,1,1}, indicating that all directions have equal probability.
[0203] Evaluation sequence U and reference sequence Z + , Z - Dynamic time warping distance The calculation is as follows:
[0204]
[0205] Where, is the optimal path generated by reinforcement learning.
[0206] Step 32: Calculate the comprehensive score of the sequence to be evaluated
[0207] According to the requested Calculate the comprehensive score of the sequence to be evaluated, which can be expressed as:
[0208]
[0209] Where S i It is the comprehensive score value of the sequence to be evaluated based on the dynamic time bending of the superior and inferior solutions, and can be used to classify the operating status of the main-distribution network to be evaluated.
[0210] Implementation cases;
[0211] The main distribution network topology is as follows Figure 3As shown in Figure 1, the overall main-distribution system is taken as the evaluation object to evaluate the main-distribution operation status. The distribution network consists of 69 nodes, of which the root node 1 is connected to the node 8 of the main network.
[0212] First, based on the multi-dimensional operation status evaluation system of the main-distribution network constructed in step 1, specific evaluation indicators are set and the corresponding operation data are obtained to generate the original data matrix and its normalized form.
[0213] The entropy weight method is used to calculate the weight of each indicator, and a weighted decision matrix is constructed based on the weight of each indicator. The weighted decision matrix for the main-distribution network operation status is established as follows:
[0214]
[0215] The extracted optimal inferior solution sequence is as follows:
[0216] Z + =[0.0048,0.0970,0.9151,0.9225,1,1] (27)
[0217] Z - =[0,0,0.2380,0.2062,0.4170,0.2920] (28)
[0218] Finally, the minimum bending distance between the sample to be evaluated and the optimal inferior solution sequence is obtained according to the proposed algorithm. The minimum bending distance of the sample to be evaluated is as follows: Figure 4 and Table 1.
[0219] Table 1 Minimum bending distance of sample sequences
[0220]
[0221] The comprehensive score values are shown in Table 2.
[0222] Table 2 Comprehensive scores of sample sequences
[0223]
[0224]
[0225] By comparing the comprehensive score values of different sequences, it can be found that sequences 4, 6, and 7 are relatively close, and sequences 3 and 5 are relatively close. Therefore, the main-support operation status samples to be evaluated can be classified through subsequent clustering or interval division, as shown in Table 3.
[0226] Table 3. Closeness of sample sequence to the optimal evaluation level
[0227] Assessment level Proximity powerful (0.2713,0.3284] Strong (0.3284,0.4274] Slightly stronger (0.4274,0.4317] generally (0.4317,0.5165] Weaker (0.5165,0.5181] weak (0.5181,0.5186]
[0228] Beneficial effects of this embodiment:
[0229] 1. This embodiment takes into account economy, reliability, and flexibility, and elaborates on the operating status evaluation index system of each level of power grid from the two levels of main power grid and distribution network, and then establishes the main-distribution network operating status evaluation index system;
[0230] 2. This embodiment selects the best-bad sample sequence that can reflect the operating status of the main-distribution network according to the definition of each indicator and the best-bad solution distance method to form a reference sample sequence;
[0231] 3. An innovative classification evaluation algorithm is proposed that integrates reinforcement learning and dynamic time bending of good and bad solutions. This algorithm can not only solve the decision-making problem of multi-index sequences, but also measure the morphological similarity of sequences and achieve effective classification with fewer training samples.
[0232] See Figure 5 The computer device 400 provided in the embodiment of the present application includes a processor 410 and a memory 420, wherein the memory 420 stores a computer program executable by the processor 410, and when the computer program is executed by the processor 410, the method described above is performed.
[0233] The embodiment of the present application further provides a storage medium 430 , on which a computer program is stored. When the computer program is run by the processor 410 , the above method is executed.
[0234] Among them, the storage medium 430 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, disk or optical disk.
[0235] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of such features. "Multiple" means two or more, unless otherwise specifically defined.
[0236] In the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," "connect," "fixed," etc. should be understood broadly. For example, they may refer to fixed connection, detachable connection, or integration; mechanical connection or electrical connection; direct connection or indirect connection through an intermediate medium; internal communication between two components or interaction between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0237] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0238] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.
[0239] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection having one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.
[0240] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0241] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0242] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and are not to be construed as limiting the present invention. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A data-driven method for dynamic evaluation of the operating status of a main-distribution multi-level power grid, characterized in that: The steps include: Constructing a multi-dimensional operating status evaluation index system for the main and distribution networks, and collecting main and distribution network operating status evaluation data based on the evaluation index system to construct an evaluation matrix; Calculating the weight of each indicator in the evaluation indicator system, and constructing a weighted decision matrix based on the evaluation matrix and the weight; The weighted decision matrix is solved using a dynamic time warping algorithm based on reinforcement learning to obtain a comprehensive score value, and the main-distribution network operating status to be evaluated is classified according to the comprehensive score value.
2. The data-driven dynamic evaluation method for the main-distribution multi-level power grid operating status according to claim 1 is characterized in that: In the multi-dimensional operation status evaluation index system for the main and distribution networks, the main network level includes: main network operation cost indicators, supply and demand balance indicators, steady-state frequency limit indicators and flexibility indicators; the distribution network level includes: distribution network operation cost indicators, branch power limit indicators, flexibility indicators and new energy absorption indicators.
3. The data-driven dynamic evaluation method for the main-distribution multi-level power grid operating status according to claim 2 is characterized in that: The calculation method for each indicator at the main network level is as follows: Mainnet operating cost indicators: E1=C st +C fuel ; Where, E1 is the operating cost of the main network; C st is the unit start-up and shutdown cost; C fuel The cost of fuel for power generation. Mainnet supply and demand balance indicators: Where, E fle is the supply and demand balance indicator of the main network, ΔP l (t) is the flexibility demand value of the system at time t, ΔP g (t) is the flexibility supply value of the system at time t; Main network steady-state frequency limit exceeding indicator: Where R f is the main network steady-state frequency limit-exceeding index, |Δf| is the system steady-state frequency deviation, Δf max It is the maximum steady-state frequency deviation that can be accepted during normal system operation. Mainnet flexibility indicators: Where AF m is the flexibility index of the main network at time t; P g (t) is the output of the controllable traditional unit at time t, when M is +, it indicates an upward adjustment characteristic, and when it is -, it indicates a downward adjustment characteristic; Respectively represent the upward and downward climbing rates of the unit; P g,max 、P g.min Indicates the upper and lower limits of the unit's output.
4. The data-driven dynamic evaluation method for the main-distribution multi-level power grid operating status according to claim 2 is characterized in that: The calculation methods for the indicators at the distribution network level are as follows: Where, E2 is the operating cost of the distribution network; C t P C is the cost of electricity interaction between the main grid and the distribution grid; RE P is the total operating cost of new energy for the distribution network. Distribution network branch power over-limit indicators: Where R L N is the power limit indicator of the distribution network branch, L is the total number of system lines, P l is the active power transmitted by line l, is the maximum active power allowed to be transmitted on line 1. And distribution network flexibility indicators: OF D =P l (t+τ)-P l (t); Where AF D is the flexibility index of the distribution network system at time t, ΔP l (t+τ) is the net load size at time t+τ, and τ is the response time scale; New energy consumption indicators: Where R RE is the main grid new energy consumption index, ∑P RE_i is the total actual output of new energy, ∑P RE_j The total output for the new energy ideal.
5. The data-driven dynamic evaluation method for the main-distribution multi-level power grid operating status according to claim 1 is characterized in that: The construction of the assessment matrix includes: The collected main-distribution network operation status evaluation data is formed into reference and evaluation sample sequences, and the matrix x is constructed. ij ; After normalizing the matrix, we get the evaluation matrix: Where μ j is the mean of the j-th column feature, σ j is the standard deviation of the j-th column feature.
6. The data-driven dynamic evaluation method for the main-distribution multi-level power grid operating status according to claim 1 is characterized in that: The calculation of the weight of each indicator in the evaluation indicator system includes: Use the entropy weight method to calculate the entropy weight w of each indicator j : Where, f ij 、y ij is the normalized value and weight of the i-th sample on the j-th indicator, m and n are the number of samples and the number of indicators respectively; α j 、H j 、w j and σ j is the adjustment factor, information entropy, entropy weight and standard deviation of the j-th indicator.
7. The data-driven dynamic evaluation method for the main-distribution multi-level power grid operating status according to claim 6 is characterized in that: The step of constructing a weighted decision matrix based on the evaluation matrix and the weights includes: Where Z is the weighted decision matrix for main-distribution network operation status evaluation.
8. The data-driven dynamic evaluation method for the main-distribution multi-level power grid operating status according to claim 1 is characterized in that: Solving the weighted decision matrix using a dynamic time warping algorithm based on reinforcement learning comprises the following steps: Extract the maximum and minimum values from each sequence of the weighted decision matrix and construct the optimal sequence Z + and the worst solution sequence Z - ; Set up the state space of reinforcement learning, the action space and reward function based on the dynamic time warping algorithm; Calculate the Z of the sequence to be evaluated and the reference sequence + and Z - Dynamic time warping distance; The comprehensive score value of the sequence to be evaluated is calculated according to the dynamic time warping distance.
9. The data-driven dynamic evaluation method for the main-distribution multi-level power grid operating status according to claim 8 is characterized in that: The state space S is: S = [(i, j)]; Where i represents the current position of the sequence to be evaluated U, and j represents the current position of the reference sequence V; The action space A is: A=[(i+1,j+1),(i+1,j),(i,j+1)]; The reward function R is: Where s represents the boundary of u, t represents the boundary of v, ρ∈{ρ d ,ρ v ,ρ h } is used as a weight factor to penalize or support the change of direction, ρ∈{ρ d ,ρ v ,ρ h }={1,1,1}, indicating that all directions have equal probability.
10. The data-driven dynamic evaluation method for the main-distribution multi-level power grid operating status according to claim 8, characterized in that: The calculation of the sequence to be evaluated and the reference sequence Z + and Z - Dynamic time warping distances, including: Where D i + 、D i - is the dynamic time bending distance; Calculating a comprehensive score value of the sequence to be evaluated according to the dynamic time warping distance includes: Where S i It is the comprehensive score of the sequence to be evaluated based on the dynamic time warping of the good and bad solutions.
11. A data-driven dynamic evaluation device for the operation status of a main-distribution multi-level power grid, characterized in that: Using the method according to any one of claims 1 to 10, the device comprises: An evaluation matrix construction unit is used to construct a multi-dimensional operating status evaluation index system for the main and distribution networks, and to collect main and distribution network operating status evaluation data according to the evaluation index system to construct an evaluation matrix; A weighting unit, configured to calculate a weight for each indicator in the evaluation indicator system and construct a weighted decision matrix based on the evaluation matrix and the weight; A solving unit is used to solve the weighted decision matrix using a dynamic time warping algorithm based on reinforcement learning to obtain a comprehensive scoring value, and classify the main-distribution network operating status to be evaluated according to the comprehensive scoring value.
12. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 10 is implemented.
13. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.