Sensitivity quantification and adaptive optimization method and system for voltage control

Through the multi-level sensitivity matrix and Transformer model combined with reinforcement learning optimization control, the problem of insufficient modeling of the nonlinear characteristic of the power grid is solved, high-precision voltage prediction and dynamic power grid control are realized, and the response speed and operating efficiency of the power grid are improved.

CN120280936APending Publication Date: 2025-07-08广西电网有限责任公司来宾供电局
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510347985.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

When the existing voltage control methods face complex load changes and large-scale distributed power access, the sensitivity analysis accuracy decreases. The traditional heuristic algorithm has high computational complexity and is prone to local optimality. Static control strategies are difficult to deal with changes in the power grid state in real time and cannot meet the dynamic control requirements of modern smart grids.

Method used

The multi-stage sensitivity matrix and Transformer model are used for voltage prediction, combined with reinforcement learning optimization control strategy, dynamically adjust the charge and discharge state of energy storage equipment, capture the nonlinear characteristics of the power grid by introducing a multi-stage sensitivity matrix, capture variable dependencies using the Transformer model, and optimize the control strategy through the DDPG algorithm.

Benefits of technology

It improves the voltage prediction accuracy, enhances the response speed and adaptability of voltage regulation, avoids local optimization, optimizes the operating efficiency and stability of the power grid, and realizes global optimization control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120280936A_ABST
    Figure CN120280936A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of optimization control of an intelligent power grid, in particular to a sensitivity quantification and adaptive optimization method and system for voltage control, and the method comprises the steps: constructing a multi-level sensitivity matrix based on obtained target power grid data; on the basis of the target power grid data and the multi-level sensitivity matrix, the trained Transform model is adopted to carry out voltage prediction; a reinforcement learning optimization control strategy is adopted according to the voltage prediction result, and an optimized control strategy is obtained; and according to the optimized control strategy and the real-time data, dynamically adjusting the charging and discharging states of the energy storage equipment. The method has the advantages that by introducing the multi-level sensitivity matrix, the defects of a traditional method in power grid nonlinear characteristic modeling are effectively overcome, and the accuracy and reliability of the model are improved. By adopting the Transform model, the complex dependency relationship among variables in the power grid can be captured, the precision of voltage prediction is remarkably improved, and the control error is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of optimal control of smart grids, and particularly to a sensitivity quantification and adaptive optimization method and system for voltage control. Background Art

[0002] In the field of optimal control of smart grids, voltage control is one of the key tasks to ensure the safe and stable operation of the power grid. Current voltage control methods mainly rely on traditional linear sensitivity analysis methods and heuristic algorithms. Linear sensitivity analysis methods, such as the Newton-Raphson method, although widely used in power systems, have the limitation that they cannot accurately capture the complex non-linear characteristics in the power grid. Especially when facing complex load changes and large-scale distributed power source access, the accuracy of sensitivity analysis will drop significantly, making it difficult to provide reliable voltage control schemes. In addition, heuristic algorithms such as Particle Swarm Optimization (PSO) and Genetic Algorithm (GA), although they can avoid the limitations of linear analysis to a certain extent, they have high computational complexity. Especially when dealing with large power grids or high-dimensional optimization problems, they require a large amount of computing resources and are prone to falling into local optimal solutions, unable to guarantee finding the global optimal solution, resulting in limited application effects in complex power systems. More importantly, the existing control strategies based on static models still dominate, and most optimization algorithms are solved under fixed grid states, making it difficult to respond to the dynamic changes of grid states in real time. The volatility of grid loads and power demands makes static control strategies ineffective in real-time optimization, and the flexibility and response speed of voltage regulation cannot meet the requirements of modern smart grids for dynamic control. Summary of the Invention

[0003] To solve the above technical problems, the present invention provides the following technical solutions:

[0004] In a first aspect, the present invention provides a sensitivity quantification and adaptive optimization method for voltage control, including constructing a multi-level sensitivity matrix based on the obtained target grid data;

[0005] Based on the target grid data and the multi-level sensitivity matrix, using a trained Transformer model for voltage prediction;

[0006] According to the voltage prediction results, adopting a reinforcement learning optimization control strategy to obtain an optimized control strategy;

[0007] According to the optimized control strategy and real-time data, dynamically adjust the charging and discharging states of energy storage devices.

[0008] As a preferred scheme of the sensitivity quantification and adaptive optimization method for voltage control of the present invention, wherein: before constructing the multi-level sensitivity matrix based on the obtained target grid data, it further includes,

[0009] Obtain power grid related data, where the power grid related data includes power grid topology data, operating status data, and distributed energy storage system data;

[0010] Preprocess the obtained power grid related data to obtain target power grid data, where the preprocessing includes data cleaning and standardization and normalization processing.

[0011] As a preferred solution of the sensitivity quantification and adaptive optimization method for voltage control of the present invention, wherein: construct a multi-level sensitivity matrix based on the obtained target power grid data, including,

[0012] Construct a first-order sensitivity matrix based on the relationship between voltage change and control variables;

[0013] Decompose the first-order sensitivity matrix by using higher-order derivatives or numerical difference methods to obtain a multi-level sensitivity matrix.

[0014] As a preferred solution of the sensitivity quantification and adaptive optimization method for voltage control of the present invention, wherein: based on the target power grid data and the multi-level sensitivity matrix, use the trained Transformer model to perform voltage prediction, including,

[0015] Take the target power grid data and the multi-level sensitivity matrix as inputs to construct an input feature matrix;

[0016] Perform standardization processing on the input feature matrix to obtain a target input feature matrix;

[0017] Use the trained Transformer model to perform forward propagation calculation on the target input feature matrix and output the voltage prediction value at the future moment.

[0018] As a preferred solution of the sensitivity quantification and adaptive optimization method for voltage control of the present invention, wherein: according to the voltage prediction result, adopt a reinforcement learning optimization control strategy to obtain an optimized control strategy, including,

[0019] Define the state space, action space, and reward function of the reinforcement learning environment;

[0020] Based on the voltage prediction result, use the Deep Deterministic Policy Gradient (DDPG) algorithm to optimize the control strategy to obtain an optimized control strategy.

[0021] As a preferred solution of the sensitivity quantification and adaptive optimization method for voltage control of the present invention, wherein: the expression of the reward function is:

[0022]

[0023] In the formula: Vi is the voltage of the i-th node, V ref is the target voltage, λ is the penalty factor, C energy is the energy consumption cost.

[0024] As a preferred solution of the sensitivity quantification and adaptive optimization method for voltage control of the present invention, wherein: the training of the Transformer model includes,

[0025] Obtain historical power grid data and corresponding multi-level sensitivity matrix data, and construct a training set;

[0026] Adopt the method of supervised learning, use the mean square error as the loss function, and train the Transformer model until the model converges.

[0027] In a second aspect, the present invention provides a sensitivity quantification and adaptive optimization system for voltage control, including: a construction module for constructing a multi-level sensitivity matrix based on the obtained target power grid data;

[0028] A prediction module for predicting the voltage by using the trained Transformer model based on the target power grid data and the multi-level sensitivity matrix;

[0029] An optimization module for optimizing the control strategy by using reinforcement learning according to the voltage prediction result to obtain an optimized control strategy;

[0030] An adjustment module for dynamically adjusting the charge and discharge states of the energy storage device according to the optimized control strategy and real-time data.

[0031] In a third aspect, the present invention provides a computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.

[0032] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0033] Compared with the prior art, the beneficial effects of the present invention are as follows: By introducing a multi-level sensitivity matrix, the deficiencies of traditional methods in modeling the nonlinear characteristics of the power grid are effectively solved, and the accuracy and reliability of the model are improved. By adopting the Transformer model, the complex dependence relationships between various variables in the power grid can be captured, significantly improving the accuracy of voltage prediction and reducing control errors. Combining reinforcement learning for optimal control, especially the DDPG algorithm, can adjust the power grid control strategy in real time, enhance the response speed and adaptability of voltage regulation, and at the same time avoid the problem that traditional heuristic algorithms are prone to falling into local optima, ensuring global optimal control. In addition, by dynamically adjusting the charge and discharge states of energy storage devices, the operating efficiency and stability of the power grid are further optimized. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0035] Figure 1 It is a schematic flow chart of a sensitivity quantification and adaptive optimization method for voltage control. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will give a detailed description of the specific embodiments of the present invention in conjunction with the drawings of the specification. Obviously, the described embodiments are some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0037] Embodiment 1, referring to Figure 1 , which is the first embodiment of the present invention. This embodiment provides a sensitivity quantification and adaptive optimization method for voltage control, including:

[0038] S1. Construct a multi-level sensitivity matrix based on the obtained target power grid data.

[0039] Furthermore, before constructing a multi-level sensitivity matrix based on the obtained target power grid data, it also includes

[0040] Obtain power grid-related data, where the power grid-related data includes power grid topology data, operating state data, and distributed energy storage system data;

[0041] It should be noted that the power grid topology data includes node numbers, voltage amplitudes, line impedances, etc.; the operation data includes real-time monitored data such as node voltages, active powers, and reactive powers; the distributed energy storage system data includes the capacity, charge and discharge status, and response time of energy storage devices, etc.

[0042] Preprocess the obtained power grid-related data to obtain target power grid data, where the preprocessing includes data cleaning and standardization and normalization processing, and data cleaning further includes repairing missing values and removing outliers.

[0043] It should be noted that in order to avoid comparison problems between data with different dimensions, it is necessary to perform standardization and normalization processing on the power grid-related data. The standardization formula is:

[0044]

[0045] In the formula: X is the original data, μ is the mean value, and σ is the standard deviation.

[0046] Normalization is to map the data between 0 and 1. Specifically,

[0047]

[0048] In the formula: X is the original data, X min is the minimum value in the dataset, X max is the maximum value in the dataset.

[0049] Preferably, by preprocessing the power grid-related data, it is ensured that parameters with different units and dimensions have consistent comparability and computability in the same model.

[0050] Furthermore, based on the obtained target power grid data, construct a multi-level sensitivity matrix, including,

[0051] Based on the relationship between voltage change and control variables, construct a first-order sensitivity matrix;

[0052] It should be noted that define the first-order sensitivity matrix S of the target node voltage V in the power grid system to the control variables X (such as active power P, reactive power Q). Specifically:

[0053] ΔV = S·ΔX

[0054] In the formula: ΔV is the change amount of the target node voltage; S is the first-order sensitivity matrix, indicating the direct influence of voltage on the control variables. ΔX is the change amount of the control variables (such as P, Q).

[0055] Furthermore, the influence relationship of the target node voltage V can be derived from the power flow equation of the power system: P-Q-V relationship:

[0056] Under steady-state conditions, the node voltage V i and the active power P i , reactive power Q i The relationship between them can be described by the power flow equation:

[0057]

[0058] Where: P i , Q i are the active and reactive powers of node i; V i , V j are the voltage magnitudes of nodes i and j respectively. G ij , B ij are the admittance matrix parameters (conductance and susceptance) between nodes i and j respectively. θ ij = θ i - θ j is the phase angle difference between nodes.

[0059] Furthermore, the sensitivity matrix S is obtained by taking the partial derivative of the power flow equation:

[0060]

[0061] Where, S ViPj represents the impact of the change in the active power of node j on the voltage of the target node i, and S ViQj represents the impact of the change in the reactive power of node j on the voltage of the target node i. These sensitivity coefficients can be calculated by numerical differentiation or analytical methods and are used to construct a multi-level sensitivity matrix to capture the direct and indirect changes in voltage.

[0062] The multi-level sensitivity matrix is obtained by decomposing the first-order sensitivity matrix through higher-order derivatives or numerical differentiation methods.

[0063] It should be noted that the calculation formula for the multi-level sensitivity matrix is:

[0064] ΔV = S (1) ·ΔX + S (2) ·(ΔX) 2 +…+ S (n) ·(ΔX) n

[0065] Where: S (1) is the first-order sensitivity matrix, representing the direct impact of the control variable on the voltage; S (2) , S (n) are the second-order and higher-order sensitivity matrices, representing the indirect and non-linear impacts of the control variable on the voltage; (ΔX) n is the high power of the control variable, used to represent the degree of non-linear impact. Among them, the higher-order sensitivity matrix S (n)Obtained through numerical approximation or high-order derivative calculation, used to quantify the complex non-linear dependence relationships within the power grid.

[0066] Preferably, in this method, by introducing a multi-level sensitivity matrix S (n) The non-linear characteristics of the system can be comprehensively captured, and the relationships between various factors in the power grid can be quantified more accurately.

[0067] S2. Based on the target power grid data and the multi-level sensitivity matrix, use the trained Transformer model for voltage prediction.

[0068] Furthermore, based on the target power grid data and the multi-level sensitivity matrix, use the trained Transformer model for voltage prediction, including

[0069] Take the target power grid data and the multi-level sensitivity matrix as inputs to construct an input feature matrix, where the target power grid data includes historical voltage data, power load, reactive power compensation, energy storage device status, etc., and the multi-level sensitivity matrix includes sensitivity matrices S of different orders (n) , used to capture the direct and indirect effects of voltage changes on control variables;

[0070] It should be noted that the input feature matrix X f is calculated as follows:

[0071] X f =[S (1) , S (2) ,..., S (n) , Voltage Data, Power Data]

[0072] In the formula: S (1) is the first-order sensitivity matrix, representing the direct effect of control variables on voltage; S (2) , S (n) are the second-order and higher-order sensitivity matrices, representing the indirect and non-linear effects of control variables on voltage; Voltage Data represents voltage data, including the voltage measurement values of each node in the power grid, reflecting the current voltage state of the power grid; Power Data: represents power data, including information such as active power and reactive power in the power grid.

[0073] Perform standardization processing on the input feature matrix to obtain the target input feature matrix, where the standardization processing is to ensure that data of different scales have the same influence;

[0074] Use the trained Transformer model to perform forward propagation calculation on the target input feature matrix, and output the voltage prediction value at the future moment.

[0075] It should be noted that the Transformer model mainly uses the self-attention mechanism to analyze the complex dependency relationships of each variable in the input feature matrix X. f The core calculation formula is as follows:

[0076]

[0077] In the formula: Q is the query matrix, representing the feature representation of the current node or variable; K is the key matrix, representing the features of the reference node or variable; K T represents the transpose of the key matrix K; V is the value matrix; d k is the dimension of the key vector; softmax is the normalization function that normalizes the attention weights to between 0 and 1, representing the dependency strength between variables.

[0078] Furthermore, it should be noted that the specific structure of the Transformer model adopted in this method includes:

[0079] 1. Input layer: The multi-level sensitivity matrix and target power grid data are used as inputs. Feature embedding is performed to map the data into a high-dimensional space.

[0080] 2. Use multi-head attention to calculate the dependency relationships from multiple different angles to improve the feature capture ability. The formula is as follows:

[0081] MultiHead(Q, K, V) = Concat(head1, head2,..., head h )W O

[0082] In the formula, represents the i-th attention head, representing feature extraction under different dimensions. are the linear transformation matrices of the query, key, and value respectively, used for feature transformation. Conat(·) is to concatenate the outputs of all attention heads together to integrate multi-dimensional feature information. W O is the output layer weight matrix, used to convert the concatenated result into the model output.

[0083] 3. Feed-Forward Network (FFN): A two-layer feed-forward network is connected after each Transformer layer. Specifically:

[0084] FFN(x) = ReLU(W1x + b1)W2 + b2

[0085] Where: x represents the input vector, W1 and W2 are weight matrices, b1 and b2 are bias terms, ReLU (Rectified Linear Unit) is the activation function, mainly used to introduce non-linear characteristics, and the final output forms the output of the FFN layer after being transformed by W2. It should be noted that the role of this layer is to perform further non-linear transformation on the features processed by the attention mechanism to enhance the model's expression ability.

[0086] 4. Normalization and residual connection: Layer Normalization is used to normalize the activation values. Residual Connection is used to accelerate gradient propagation and avoid gradient vanishing.

[0087] Furthermore, the training of the Transformer model includes

[0088] Obtaining historical power grid data and the corresponding multi-level sensitivity matrix data to construct a training set;

[0089] Adopting the method of supervised learning, using the mean squared error as the loss function, to train the Transformer model until the model converges.

[0090] It should be noted that the loss function is specifically:

[0091]

[0092] Where: N represents the number of samples; represents the voltage value of the i-th node predicted by the model; V i represents the voltage value of the i-th node measured actually; Loss is the loss function, representing the error between the model prediction value and the measured true value.

[0093] S3. According to the voltage prediction result, adopt the reinforcement learning optimization control strategy to obtain the optimized control strategy.

[0094] It should be noted that the reinforcement learning optimization process can be summarized into five core stages, including environment modeling, state perception, policy decision-making, learning optimization, and execution control. These stages interact with each other to jointly construct a complete intelligent optimization control system.

[0095] Furthermore, according to the voltage prediction result, adopting the reinforcement learning optimization control strategy to obtain the optimized control strategy includes

[0096] Defining the state space, action space, and reward function of the reinforcement learning environment.

[0097] It should be noted that in the environmental modeling stage, the state space, action space, and reward function of reinforcement learning are first defined so that the agent can effectively learn and optimize the power grid control strategy. The state space S mainly includes the node voltage, power load, and sensitivity matrix of the power grid, and these variables describe the current operating state of the power grid; the action space A allows the agent to adjust the active power, reactive power, and charge and discharge states of energy storage devices to achieve voltage regulation in an optimal manner; the reward function R is used to measure the effect of the optimization strategy, and its goal is to minimize the deviation between the node voltage and the target voltage V ref while taking into account the energy consumption cost C energy .

[0098] Based on the voltage prediction results, the Deep Deterministic Policy Gradient (DDPG) algorithm is used to optimize the control strategy, and the optimized control strategy is obtained.

[0099] It should be noted that in the state perception stage, the agent collects data such as voltage, power load, and energy storage device status in real time through the power grid monitoring system, and calculates the sensitivity matrix S (n) at the current moment to accurately describe the operating state of the power grid and generate the state vector S t as input data. On this basis, the DDPG algorithm is used in the policy decision-making stage to calculate the optimal control variable adjustment strategy. This algorithm calculates the optimal control action a t under the current state S t :

[0100]

[0101] where: μ(s t |θ μ ) is the decision value output by the policy network, and N t is the exploration noise, which ensures that the agent can try different control schemes during the learning process and avoid falling into local optima.

[0102] Meanwhile, the value network (Critic) evaluates the quality of this strategy and calculates the state-action value function Q(s t , a t ), that is, the expected return after taking the action a t in the current state S t .

[0103] Furthermore, in order to further optimize the learning process of the agent, in the learning optimization stage, the experience replay mechanism (Replay Buffer) is used to record the interaction data (s t , a t , r t , s t+1), and update the Critic network by minimizing the Temporal Difference (TD) error to enable it to more accurately estimate the impact of actions on the future:

[0104] y t = r t + γQ(s t+1 , μ(s t+1 |θ μ )|θ Q )

[0105] Where: y t represents the TD target value, r t represents the reward obtained at time t, γ represents the discount factor, Q(s t+1 , μ(s t+1 |θ μ )|θ Q ) represents the Q value obtained by evaluating the action generated by the policy network μ in the next state S t+1 through the value network. Among them, θ μ is the parameter of the policy network, and θ Q is the parameter of the value network.

[0106]

[0107] Where: Loss Q is used to measure the accuracy of the Critic network's estimation of Q(s t , a t ); N is the number of samples; y t represents the TD target value; (y t - Q(s t , a t |θ Q )) 2 represents the squared error between the TD target value of the t-th sample and the Q value estimation, which is used to calculate the loss function; Q(s t , a t |θ Q ) represents the Q value estimation of the t-th sample in state S t and action a t , which is calculated by the value network according to the parameter θ Q .

[0108] Furthermore, the gradient update of the policy network is based on the feedback of the value network to optimize the control strategy:

[0109]

[0110] Where: represents the gradient of the policy network; E s~DDenotes the expectation of the state s sampled from the data distribution D, i.e., the average value; Denotes the gradient of the Q-value with respect to the action a, i.e., the sensitivity of the Q-value to the change in the action, calculated by the value network according to the parameter θ Q ; Denotes the gradient of the action output by the policy network with respect to the policy network parameter θ μ , i.e., the sensitivity of the action to the change in the policy network parameter.

[0111] Finally, in the execution control phase, the agent dynamically adjusts the active power, reactive power, and the charge and discharge states of the energy storage device according to the currently optimized policy μ(s t |θ μ ), continuously optimizing the grid voltage.

[0112] Preferably, since this control strategy adopts a soft update mechanism, the parameter adjustment is smoother, avoiding violent oscillations in the control strategy and ensuring the stability of the grid operation. Further, this reinforcement learning optimization framework continuously operates in the grid dispatching system, and can complete the voltage optimization calculation within a millisecond-level response time, ensuring that the grid voltage is maintained within a reasonable range in real time.

[0113] Furthermore, the expression of the reward function is:

[0114]

[0115] where: V i is the voltage of the i-th node, V ref is the target voltage, λ is the penalty factor, and C energy is the energy consumption cost. By setting an appropriate value of λ, a balance can be found between grid stability and energy efficiency optimization.

[0116] S4. Dynamically adjust the charge and discharge states of the energy storage device according to the optimized control strategy and real-time data.

[0117] It should be noted that through real-time data monitoring, the charge and discharge states of the energy storage device will be adjusted according to the grid load changes and external disturbance conditions. The scheduling formula for the energy storage device is:

[0118] ΔP storage = f(V, P load , C energy )

[0119] where, ΔP storage is the charge and discharge power of the energy storage device, V is the grid voltage, P load is the grid load power, and C energyis the energy consumption cost. The basic idea of this scheduling formula is: by real-time monitoring the voltage and load changes of the power grid, dynamically adjusting the charge and discharge states of energy storage devices to maintain the stability of the power grid voltage, while minimizing energy consumption and improving the operation efficiency of the power grid.

[0120] Furthermore, to more intuitively understand this scheduling strategy, assume that in a typical power grid, the current load power P of the power grid load is 200 MW, the power grid voltage V is 220 kV, and the energy consumption cost C energy is 500 yuan / MWh. In this case, the energy storage device needs to charge or discharge to adjust the voltage and load balance. Assume that through calculation, the charge and discharge power of the energy storage device obtained by the scheduling strategy is ΔP storage = 50 MW. This means that the energy storage system needs to provide 50 MW of power during this time period to compensate for the fluctuations of the load power and the power grid voltage, ensuring the stable and efficient operation of the power grid.

[0121] In summary, the beneficial effects of the method for sensitivity quantization and adaptive optimization for voltage control in the present invention are as follows: by introducing a multi-level sensitivity matrix, it effectively solves the deficiencies of traditional methods in modeling the nonlinear characteristics of the power grid, improving the accuracy and reliability of the model. By using the Transformer model, it can capture the complex dependence relationships between various variables in the power grid, significantly improving the accuracy of voltage prediction and reducing the control error. Combining reinforcement learning for optimal control, especially the DDPG algorithm, it can adjust the power grid control strategy in real time, enhancing the response speed and adaptability of voltage regulation, while avoiding the problem that traditional heuristic algorithms are prone to falling into local optima, ensuring global optimal control. In addition, by dynamically adjusting the charge and discharge states of energy storage devices, the operation efficiency and stability of the power grid are further optimized.

[0122] Embodiment 2 is the second embodiment of the present invention. This embodiment provides a system for sensitivity quantization and adaptive optimization for voltage control, including a construction module for constructing a multi-level sensitivity matrix based on the obtained target power grid data;

[0123] a prediction module for performing voltage prediction using the trained Transformer model based on the target power grid data and the multi-level sensitivity matrix;

[0124] an optimization module for obtaining an optimized control strategy according to the voltage prediction result using a reinforcement learning optimal control strategy;

[0125] an adjustment module for dynamically adjusting the charge and discharge states of energy storage devices according to the optimized control strategy and real-time data.

[0126] Embodiment 3 is the third embodiment of the present invention. The difference from the previous two embodiments is:

[0127] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0128] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in combination with an instruction execution system, apparatus, or device.

[0129] More specific examples (non-exhaustive list) of computer-readable media include the following: electrical connection parts with one or more wirings (electronic devices), portable computer disk cartridges (magnetic devices), random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), optical fiber devices, and portable compact disc read-only memories (CDROMs). Additionally, a computer-readable medium can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or processing it in other suitable ways when necessary, and then storing it in a computer memory.

[0130] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0131] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A sensitivity quantification and adaptive optimization method for voltage control, characterized in that: including constructing a multi-level sensitivity matrix based on the obtained target power grid data; performing voltage prediction using the trained Transformer model based on the target power grid data and the multi-level sensitivity matrix; adopting a reinforcement learning optimization control strategy according to the voltage prediction result to obtain an optimized control strategy; dynamically adjusting the charge and discharge states of energy storage devices according to the optimized control strategy and real-time data.

2. The sensitivity quantization and adaptive optimization method for voltage control according to claim 1, characterized in that: Before constructing the multi-level sensitivity matrix based on the obtained target power grid data, it further includes obtaining power grid-related data, where the power grid-related data includes power grid topology data, operating state data, and distributed energy storage system data; preprocessing the obtained power grid-related data to obtain target power grid data, where the preprocessing includes data cleaning and standardization and normalization processing.

3. The sensitivity quantification and adaptive optimization method for voltage control according to claim 2, wherein: Constructing the multi-level sensitivity matrix based on the obtained target power grid data includes constructing a primary sensitivity matrix based on the relationship between voltage change and control variables; decomposing the primary sensitivity matrix by high-order derivative or numerical difference method to obtain a multi-level sensitivity matrix.

4. The sensitivity quantification and adaptive optimization method for voltage control according to claim 3, characterized in that: Performing voltage prediction using the trained Transformer model based on the target power grid data and the multi-level sensitivity matrix includes using the target power grid data and the multi-level sensitivity matrix as inputs to construct an input feature matrix; performing standardization processing on the input feature matrix to obtain a target input feature matrix; performing forward propagation calculation on the target input feature matrix using the trained Transformer model and outputting the voltage prediction value at a future moment.

5. The sensitivity quantization and adaptive optimization method for voltage control according to claim 4, characterized in that: Adopting a reinforcement learning optimization control strategy according to the voltage prediction result to obtain an optimized control strategy includes defining the state space, action space, and reward function of the reinforcement learning environment; optimizing the control strategy using the Deep Deterministic Policy Gradient (DDPG) algorithm based on the voltage prediction result to obtain an optimized control strategy.

6. The sensitivity quantization and adaptive optimization method for voltage control according to claim 5, characterized in that: The expression of the reward function is: Where: V i is the voltage of the i-th node, V ref is the target voltage, λ is the penalty factor, C energy is the energy consumption cost.

7. The sensitivity quantization and adaptive optimization method for voltage control according to claim 6, wherein: The training of the Transformer model includes obtaining historical power grid data and corresponding multi-level sensitivity matrix data to construct a training set; training the Transformer model in a supervised learning manner with the mean square error as the loss function until the model converges.

8. A sensitivity quantification and adaptive optimization system for voltage control, characterized in that, including: a construction module for constructing a multi-level sensitivity matrix based on the obtained target power grid data; a prediction module for performing voltage prediction using the trained Transformer model based on the target power grid data and the multi-level sensitivity matrix; an optimization module for adopting a reinforcement learning optimization control strategy according to the voltage prediction result to obtain an optimized control strategy; an adjustment module for dynamically adjusting the charge and discharge states of energy storage devices according to the optimized control strategy and real-time data.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.