Power system deep reinforcement learning voltage optimization method and device

Through the machine learning algorithm of deep reinforcement learning algorithm combined with gradient enhancement framework, the voltage optimization problem of power system under the influence of new energy power generation is solved, and the stable regulation and robustness of grid voltage are achieved.

CN120073752APending Publication Date: 2025-05-30NORTH CHINA BRANCH OF STATE GRID CORPORATION OF CHINA +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510224177.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When existing power systems deal with the impact of new energy power generation on the grid structure and operating stability, it is difficult to achieve global voltage optimization, and the automatic voltage control system is not robust enough.

Method used

A machine learning algorithm with a deep reinforcement learning algorithm combined with a gradient enhancement framework is used to train historical data and typical time period data to obtain an optimized model to optimize the power system voltage.

Benefits of technology

Voltage regulation in uncertain scenarios is realized, the power and stability of grid voltage regulation is improved, and the robustness and interpretability of artificial intelligence algorithms in power systems are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120073752A_ABST
    Figure CN120073752A_ABST
Patent Text Reader

Abstract

The invention provides a deep reinforcement learning voltage optimization method for a power system, and the method comprises the following steps: obtaining a historical data set of the power system based on the power system, and carrying out the training of a Light GBM algorithm, so as to obtain a test model; based on the test model, obtaining a typical time period output data set according to the typical time period input data set so as to obtain a node data set based on the time nodes; based on the power system, performing load flow calculation according to the node data set to obtain a node voltage set; and a DDQN algorithm is adopted, and an optimization model is obtained according to the node voltage set to optimize the power system. According to the invention, an artificial intelligence algorithm is adopted to progressively learn the optimization capacity, the optimal control strategy in the current learning stage is given in real time to optimize the voltage of the power system, the stability of the voltage of the power grid is effectively supported, and the regulation capacity of the voltage of the power grid is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of smart grids, and particularly to a method and device for voltage optimization of a power system based on deep reinforcement learning. Background Art

[0002] A large number of new energy generating units are continuously penetrating into the power system, and new energy power generation has an inevitable impact on the power system structure and operation stability. Voltage is not only one of the standards for grid quality, but also an important aspect for achieving high-quality power consumption safety. Load level, new energy installed capacity, thermal power starting capacity, received power, and switching of reactive power regulation equipment are all important factors affecting bus voltage, and the automatic voltage control system has become an important device for the power department to control voltage.

[0003] Since the late 1970s of the last century, reactive power voltage control has become a research hotspot in the field of power system operation and control. Among them, the interior point method can handle constrained optimization problems. Its characteristic is that the calculation time is not sensitive to the problem scale, but its analytical method depends highly on the accuracy of the grid structure, parameters, and operation measurement data. The complex iterative solution algorithm often has the disadvantage of poor robustness in the reactive power optimization process of the actual system. With the development of the power system, most substations are equipped with automatic voltage and reactive power control devices (VQC). VQC is based on local measurement information and adjusts the transformer tap and capacitor of the substation according to the multi-zone diagram principle (nine-zone, thirteen-zone, five-zone), so as to uniformly adjust the regional reactive power voltage control equipment. However, it often cannot set the zoning and adjustment criteria from the perspective of the whole network and is difficult to give a control strategy with global optimization characteristics for the regional power grid. At the same time, the rapid development of external measurement systems and communication technologies has brought about the accumulation of massive multi-source data, which are closely integrated with the new generation of artificial intelligence and provide a broad stage for the application of artificial intelligence in power system problems. The artificial intelligence algorithm uses the progressive learning and optimization ability to optimize the reactive power voltage control strategy of the regional power grid. It can give the best control strategy in the current learning stage in real time, ensure the robustness of the reactive power voltage control algorithm, and can implement coordinated control for multiple substations. Summary of the Invention

[0004] Aiming at the problems existing in the prior art, the present invention provides a method and device for voltage optimization of a power system based on deep reinforcement learning, which can give full play to the strong exploration ability of deep reinforcement learning in an uncertain scenario and achieve good voltage regulation effects.

[0005] To achieve the above object, the technical solution adopted by the present invention is as follows: A method for voltage optimization of a power system based on deep reinforcement learning includes the following steps:

[0006] Based on the power system, obtain a historical data set and a typical period input data set respectively;

[0007] Obtain a time node;

[0008] Adopt a machine learning algorithm based on the gradient boosting framework, and train a test model according to the historical data set;

[0009] Based on the test model, obtain a typical period output data set according to the typical period input data set;

[0010] Based on the time node, obtain a node data set according to the output data set, where the node data set is a set of typical period output data for each time node;

[0011] Based on the power system, perform a power flow calculation according to the node data set to obtain a node voltage set, where the node voltage set is a set of voltage data for each time node;

[0012] Adopt a deep reinforcement learning algorithm to obtain an optimization model according to the node voltage set;

[0013] Optimize the power system according to the optimization model.

[0014] In some embodiments, the machine learning algorithm based on the gradient boosting framework is the LightGBM algorithm.

[0015] In some embodiments, the historical data set includes a historical new energy output data set and a historical load data set.

[0016] In some embodiments, the step of performing a power flow calculation according to the node data set based on the power system to obtain a node voltage set is as follows:

[0017] Obtain a standard data set from the node data set, where the standard data set is a set of standard data of voltage data for each time node;

[0018] Obtain a first table file according to the standard data set;

[0019] Convert according to the first table file to obtain a case program;

[0020] Construct a case model according to the power system;

[0021] Run the case program according to the case model to obtain a result file;

[0022] Convert according to the result file to obtain a second table file;

[0023] Obtain the node voltage set according to the second table file.

[0024] In some of these embodiments, the conversion of the first tabular file to obtain the calculation example program and the conversion of the result file to obtain the second tabular file are respectively implemented using Python.

[0025] In some of these embodiments, the construction of the calculation example model based on the power system and the running of the calculation example program based on the calculation example model to obtain the result file are respectively implemented on the PSD-BPA software.

[0026] In some of these embodiments, the steps for constructing the calculation example model based on the power system are as follows:

[0027] Retrieve the original calculation example model in the PSD-BPA software, where the original calculation example model includes a number of initial PV nodes;

[0028] Based on the power system, set the installed capacity in the original calculation example model, add reactive power compensation devices and on-load tap-changers respectively, and change the corresponding initial PV nodes to PQ nodes to obtain the initial calculation example model;

[0029] Set the voltage change range of the calculation example nodes, set the tap ratio change range of the on-load tap-changer, and set the adjustable capacity of the reactive power compensation device in the initial calculation example model to obtain the calculation example model; the calculation example nodes include the PQ nodes and the unchanged initial PV nodes.

[0030] In some of these embodiments, the deep reinforcement learning algorithm is the DDQN algorithm.

[0031] A device for implementing the power system deep reinforcement learning voltage optimization method includes a data acquisition unit, a test unit, a calculation unit, and an optimization analysis unit;

[0032] The data acquisition unit:

[0033] For respectively acquiring a historical data set and a typical period input data set based on the power system and transmitting them to the test unit respectively;

[0034] For acquiring a time node and transmitting it to the calculation unit;

[0035] The test unit:

[0036] For training a test model using a machine learning algorithm based on the gradient boosting framework according to the historical data set;

[0037] For obtaining a typical period output data set based on the test model according to the typical period input data set and transmitting it to the calculation unit;

[0038] The calculation unit:

[0039] To obtain a node data set based on the time node and according to the output data set;

[0040] To perform power flow calculation based on the power system and according to the node data set to obtain a node voltage set and transmit it to the optimization analysis unit;

[0041] The optimization analysis unit:

[0042] To obtain an optimization model by using a deep reinforcement learning algorithm and according to the node voltage set;

[0043] To optimize the power system according to the optimization model.

[0044] In some embodiments, the calculation unit includes a data processing module and a power flow calculation module;

[0045] The data processing module:

[0046] To obtain a node data set based on the time node and according to the output data set;

[0047] To obtain a case program according to the node data set;

[0048] To transmit the node voltage set to the optimization analysis unit;

[0049] The power flow calculation module:

[0050] To construct a case model according to the power system;

[0051] To retrieve the case program based on the case model to obtain a result file;

[0052] To obtain the node voltage set according to the result file;

[0053] To transmit the node voltage set to the data processing module.

[0054] Compared with the prior art, the present invention has the following beneficial effects:

[0055] (1) The present invention adopts the progressive learning and optimization ability of the artificial intelligence algorithm, and gives the best control strategy in the current learning stage in real time to optimize the voltage of the power system, effectively supporting the stability of the grid voltage and improving the regulation ability of the grid voltage.

[0056] (2) The present invention uses a Python program to call the PSD-BPA software for batch power flow calculation, providing a means that is more suitable for the current actual operation of the power grid for power system related research.

[0057] (3) The present invention quantifies the number of device operation times of on-load tap-changing transformers and reactive power compensation capacitor banks during the voltage regulation process of the power system, provides a reference for voltage regulation means for dispatchers, improves the operation control efficiency of the power system, and enhances the interpretability of the application of artificial intelligence algorithms in the field of power systems. Description of the Drawings

[0058] Figure 1 It is a schematic flow chart of the power system deep reinforcement learning voltage optimization method of the present invention;

[0059] Figure 2 It is a schematic flow chart of the power flow calculation in the present invention;

[0060] Figure 3 It is a comparison chart of the load prediction value and the actual load value of the test model for a typical day in the numerical example of the present invention;

[0061] Figure 4 It is a distribution chart of the wind and light output and load for a typical day in the numerical example of the present invention;

[0062] Figure 5 It is a schematic diagram of the numerical example model improved based on the IEEE33 standard numerical example model in the numerical example of the present invention;

[0063] Figure 6 It is a comparison chart of the voltages at each node before optimization in the numerical example of the present invention;

[0064] Figure 7 It is a comparison chart of the voltages at each node after optimization in the numerical example of the present invention;

[0065] Figure 8 It is the state of change of the switching operation positions of the equipment in the numerical example of the present invention;

[0066] Figure 9 It is a schematic diagram of the device for implementing the power system deep reinforcement learning voltage optimization method of the present invention. Detailed Embodiments

[0067] Technologies such as big data and artificial intelligence can give full play to the value of various types of data, explore the complex relationships between various variables, and have broad application prospects in the fields of power system operation control, risk assessment, and auxiliary decision-making. Against this background, based on big data, applying modern advanced technologies such as artificial intelligence to study the influence mechanism and correlation effect between variables, and then studying the voltage regulation optimization strategy of regional power grids has important practical significance for supporting the voltage stability of the power grid and improving the voltage regulation ability of the power grid, and is conducive to enhancing the control ability of the dispatching department over the new power system and improving the operation control efficiency of the power system.

[0068] At present, artificial intelligence algorithms used in voltage regulation include reinforcement learning, deep reinforcement learning, machine learning, heuristic optimization algorithms, etc. Each algorithm has its own advantages and disadvantages. At the same time, most of the existing means of using artificial intelligence technology to study voltage regulation solve the power flow calculation through Matpower in Matlab or Pypower package in Python. In the actual power grid, the model of the power system is usually designed and operated in the format required by the PSD-BPA software. There is an urgent need for technical means that match its format. Compared with traditional methods, this voltage regulation model uses artificial intelligence algorithms and simultaneously batch calls the power system power flow calculation software PSD-BPA to achieve dynamic voltage optimization, which can effectively meet the voltage regulation requirements of regional power grids.

[0069] To clearly illustrate the technical characteristics of this solution, the following will combine the accompanying drawings and embodiments to detail the implementation manner of this application, so as to fully understand how this application uses technical means to solve technical problems and the implementation process of achieving corresponding technical effects and implement accordingly. The embodiments of this application and each feature in the embodiments can be combined with each other on the premise of not conflicting, and the formed technical solutions are all within the protection scope of this application.

[0070] See Figure 1 , the embodiment of the present invention provides a method for optimizing the voltage of a power system by deep reinforcement learning, including the following steps:

[0071] Based on the power system, obtain the historical data set and the typical period input data set respectively; the historical data set includes the historical new energy output data set and the historical load data set; obtain the time node; preferably, the typical period is a typical day, that is, 24 hours a day, and the time node is each whole point in a day, that is, the granularity is 1h;

[0072] Adopt a machine learning algorithm based on the gradient boosting framework, and train a test model according to the historical data set; preferably, the machine learning algorithm based on the gradient boosting framework is the LightGBM algorithm; the LightGBM algorithm is an efficient gradient boosting framework designed for large data sets and high-dimensional data sets;

[0073] The LightGBM algorithm uses a histogram-based decision tree construction method and introduces multiple optimizations in algorithm design, such as leaf node growth by depth (Leaf-wise Growth), GOSS (Gradient-based One-Side Sampling), etc., to improve the training speed and resource utilization rate. The main parameters include the maximum number of leaves, learning rate, number of iterations, regularization coefficient, etc., which need to be set accordingly. It can be evaluated by the mean absolute error (MAE), mean square error (MSE), root mean square error (RMSE) and R2 The fitting effect of judgment algorithms such as the coefficient of determination;

[0074] Based on the test model, obtain the typical period output data set according to the typical period input data set;

[0075] Based on time nodes, obtain the node data set according to the output data set. The node data set is a set of typical period output data for each time node;

[0076] According to the power system, perform power flow calculation on the node data set to obtain the node voltage set. The node voltage set is a set of voltage data for each time node; see Figure 2 , the process of batch power flow calculation is specifically implemented by Python calling PSD-BPA. The calling process of power flow calculation is automatically looped in Python. In addition, it can also be implemented by Matlab instead of Python. In some embodiments, the step of performing power flow calculation on the node data set according to the power system to obtain the node voltage set is:

[0077] Obtain the standard data set from the node data set. The standard data set is a set of standard data of voltage data for each time node;

[0078] Obtain the first table file according to the standard data set. Save the time node, data type (new energy output data and load data in the typical period) and typical period output data value in the Excel file in the standard format to obtain the first table file;

[0079] Through the code for converting Excel files and Dat files, use Python to convert the first table file to obtain an example program, which is an example program recognizable by the PSD-BPA software;

[0080] On the PSD-BPA software, construct an example model according to the power system; specifically:

[0081] Retrieve the original example model in the PSD-BPA software. The original example model includes several initial PV nodes. In the original example model, the node generator type of the initial PV nodes is thermal power generators, with no new energy penetration or extremely low new energy penetration rate, which does not match the power system and needs to be adjusted and improved;

[0082] Based on the power system, set the installed capacity in the original example model, add reactive power compensation equipment and on-load tap-changer respectively, and change the corresponding initial PV nodes to PQ nodes to obtain the initial example model; since the new energy generators in the example model to be constructed do not generate reactive power, the node type needs to be changed from PV nodes to PQ nodes accordingly;

[0083] In the initial case model, set the voltage change range of the case nodes, set the tap ratio change range of the on-load tap-changing transformer, and set the adjustable capacity of the reactive power compensation device to obtain the case model; the case nodes include PQ nodes and the unchanged initial PV nodes;

[0084] On the PSD-BPA software, run the case program according to the case model to obtain the result file; since the PSD-BPA software cannot automatically run the file and manual clicks are required, an automatic running code also needs to be set, which is prior art and will not be elaborated here; after running, a result file in the Pfo format is generated;

[0085] Use Python to convert according to the result file to obtain the second table file;

[0086] Obtain the node voltage set according to the second table file, and each case node corresponds to a node voltage;

[0087] Adopt the deep reinforcement learning algorithm to obtain the optimization model according to the node voltage set. Preferably, the deep reinforcement learning algorithm is the DDQN algorithm; the setting of the action space is mainly reflected in the tap ratio of the on-load tap-changing transformer and the switching capacity of the reactive power compensation device; automatically select the action with the largest return value to act. Compared with before optimization, the voltage is significantly improved. Specifically, the DDQN algorithm includes an objective function and a voltage barrier function; the objective function is:

[0088] F = min(V ave - V ref );

[0089] In the formula, F is the objective function, min represents the minimization function, V ave is the average value of the sum of the node voltages in the node voltage set, and V ref is the voltage reference value;

[0090] The voltage barrier function is:

[0091]

[0092] In the formula, f ν is the voltage barrier function, v i is the node voltage of the i-th case node, and N(v i , 0.1 2 ) is the density function that satisfies the expectation of v i , has a mean of 0.1 and satisfies the normal distribution, and b, c, d, and e are hyperparameters for adjusting the gradient respectively;

[0093] The steps to obtain the optimization model according to the node voltage set by using the DDQN algorithm are:

[0094] Step S1: Construct an initial optimization model, set the voltage reference value and initialize the hyperparameters;

[0095] Step S2: Let the initial optimization model be the current optimization model; Let the power system be the current power system;

[0096] Step S3: Obtain the node voltage set according to the current power system;

[0097] Step S4: Run the current optimization model according to the node voltage set to obtain the current optimization result;

[0098] Step S5: Update the current optimization model respectively according to the current optimization result;

[0099] If the current optimization result is not to terminate, update the current power system according to the current optimization model and repeat Steps S3 - S5;

[0100] If the current optimization result is to terminate, obtain the optimization model according to the current optimization model;

[0101] Optimize the power system according to the optimization model.

[0102] See Figure 9 , The embodiment of the present invention also provides a device for implementing the voltage optimization method of power system deep reinforcement learning, including a data acquisition unit, a test unit, a calculation unit and an optimization analysis unit;

[0103] Data acquisition unit:

[0104] For respectively acquiring the historical data set and the typical period input data set based on the power system and respectively transmitting them to the test unit;

[0105] For acquiring the time node and transmitting it to the calculation unit;

[0106] Test unit:

[0107] For training a test model by using a machine learning algorithm based on the gradient boosting framework according to the historical data set;

[0108] For obtaining the typical period output data set based on the test model according to the typical period input data set and transmitting it to the calculation unit;

[0109] Calculation unit:

[0110] For obtaining the node data set based on the time node according to the output data set;

[0111] For performing power flow calculation based on the power system according to the node data set to obtain the node voltage set and transmitting it to the optimization analysis unit;

[0112] Optimization analysis unit:

[0113] For obtaining an optimization model according to a node voltage set by using a deep reinforcement learning algorithm;

[0114] For optimizing a power system according to the optimization model.

[0115] In some embodiments, the calculation unit includes a data processing module and a power flow calculation module;

[0116] Data processing module:

[0117] For obtaining a node data set based on time nodes according to an output data set;

[0118] For obtaining a case program according to the node data set;

[0119] For transmitting the node voltage set to the optimization analysis unit;

[0120] Power flow calculation module:

[0121] For constructing a case model according to the power system;

[0122] For retrieving a case program based on the case model to obtain a result file;

[0123] For obtaining the node voltage set according to the result file;

[0124] For transmitting the node voltage set to the data processing module.

[0125] Case:

[0126] Select the real new energy output data and load data of a certain local power grid (the new energy equipment in this power grid includes wind power equipment and photovoltaic equipment) within one month as the historical data set, and use the LightGBM algorithm to train and obtain a test model according to the historical data set; the historical data set includes wind power output data, photovoltaic output data, and load data;

[0127] In the historical data set, the time node is 15 minutes, and there are 2880 pieces of data in total, of which 90% is the training set, a total of 2592 pieces, and the test set is 10%. The 2593-2880th pieces of data are the test set, a total of 288 pieces. The LightGBM algorithm parameters are set as follows: the maximum number of leaves: 31, the learning rate: 0.05, the number of trees (number of iterations): 200, the minimum amount of data per leaf node: 20, the percentage of features used in each iteration: 0.9, the percentage of samples used in each iteration: 0.8, the L1 regularization coefficient: 0.1, and the L2 regularization coefficient: 0.1. The result after running the LightGBM algorithm is as Figure 3 shown.

[0128] Based on the test model, the wind and solar power output data (wind power output data and photovoltaic power output data) of typical days (typical periods) are used to predict the load, and the fitting degree is very high. And the feature importance of wind power output is 2.6 - 3.3 times that of photovoltaic power output. The calculated results are as follows: Mean Absolute Error (MAE): 7.2259; Mean Squared Error (MSE): 426.0851; Root Mean Squared Error (RMSE): 20.6418; R 2 Coefficient of determination: 0.9996 (R 2 measures the fitting degree of the model to the data. The closer it is to 1, the better the fitting effect of the test model). The wind and solar power output data and load data of a certain typical day are distributed as Figure 4 shown.

[0129] Taking the 33 - node standard example model as the original example model, the topology diagram of this 33 - node standard example model is as Figure 5 shown.

[0130] Improve the 33 - node standard example model to obtain an example model: Connect photovoltaic generators with an installed capacity of 1 MW each at nodes 6, 12, 18, 22, and 25; connect a wind turbine generator with an installed capacity of 2 MW at node 32; connect an on - load tap changer (OLTC) between nodes 1 and 2. The tap - changing range has a total of 5 gears, and the adjustable voltage is 0.95 - 1.05 p.u., that is, set the range of the transformation ratio of the on - load tap changer; connect reactive power compensation equipment at nodes 21 and 24. Here, the reactive power compensation equipment is a shunt capacitor bank (SCB), with a total of 5 adjustable gears, and the adjustable capacity of each gear is 0.4 Mvar, that is, set the adjustable capacity of the reactive power compensation equipment.

[0131] Assign the data of a certain selected typical day to the example model, 24 times in total, and perform power flow calculations. The power flow calculation results converge. The node voltages of each example node are compared as Figure 6 shown. The abscissa is the time - node number, with 24 time nodes and a total of 24 voltage broken lines.

[0132] Introduce the DDQN algorithm for optimization training to obtain an optimized model. The values of the hyperparameters in the voltage barrier function of the optimized model are: b = 0.1; c = 0.05; d = 0.5; e = 0.1. That is:

[0133]

[0134] After optimizing the power system according to the optimized model, the voltages of each node after optimization are as Figure 7 shown, and the changes in the switching operation gears of the equipment after optimization are as Figure 8 shown. The statistics of the operation times of the on - load tap changer and reactive power compensation equipment are shown in Table 1.

[0135] Table 1 Statistics of Equipment Switching Times

[0136] Actuating device Switching times OLTC 3 SCB21 8 SCB24 7

[0137] The selection of operating equipment (on-load tap-changer and reactive power compensation equipment) is the action space of the DDQN algorithm when calculating the return value. The larger the return value, the more optimal the strategy is under the current objective function, which plays a very important auxiliary role in voltage optimization and helps dispatchers select the optimal voltage regulation scheme more efficiently and quickly. Moreover, under this optimization algorithm, all voltages are optimized between 0.95 p.u. and 1.05 p.u.; the average voltage deviation before optimization is 0.0289 p.u., and the average voltage deviation after optimization is 0.001 p.u. Compared with the two, the voltage has been significantly improved.

[0138] Finally, it should be noted that the above content is only used to illustrate the technical solution of the present invention, rather than a limitation on the protection scope of the present invention. Any simple modification or equivalent replacement of the technical solution of the present invention by those of ordinary skill in the art shall not depart from the essence and scope of the technical solution of the present invention.

Claims

1. A method for voltage optimization of a power system by deep reinforcement learning, characterized in that: The following steps are involved: Based on the power system, historical data sets and typical period input data sets are obtained respectively; Get the time node; Using a machine learning algorithm based on a gradient boosting framework, a test model is obtained by training the historical data set; Based on the test model, obtaining a typical period output data set according to the typical period input data set; Based on the time node, obtaining a node data set according to the output data set, wherein the node data set is a set of output data of a typical period of each time node; Based on the power system, a power flow calculation is performed according to the node data set to obtain a node voltage set, wherein the node voltage set is a set of voltage data of each time node; Using a deep reinforcement learning algorithm, an optimization model is obtained according to the node voltage set; The power system is optimized according to the optimization model.

2. The method for voltage optimization of a power system by deep reinforcement learning according to claim 1, characterized in that: The machine learning algorithm based on the gradient boosting framework is the LightGBM algorithm.

3. The method for voltage optimization of a power system by deep reinforcement learning according to claim 2, characterized in that: The historical data set includes a historical new energy output data set and a historical load data set.

4. The method for voltage optimization of a power system by deep reinforcement learning according to claim 3, characterized in that: Based on the power system, the steps of performing power flow calculation according to the node data set to obtain the node voltage set are: Obtaining a standard data set from the node data set, wherein the standard data set is a set of standard data of the voltage data of each time node; Obtaining a first table file according to the standard data set; Obtaining a calculation example program according to the first table file conversion; Constructing a calculation model according to the power system; Run the example program according to the example model to obtain a result file; Obtain a second table file according to the result file conversion; The node voltage set is obtained according to the second table file.

5. The method for voltage optimization of a power system by deep reinforcement learning according to claim 4, characterized in that: The example program is obtained by converting the first table file, and the second table file is obtained by converting the result file, respectively, using Python or Matlab.

6. The method for voltage optimization of a power system by deep reinforcement learning according to claim 4, characterized in that: Constructing a calculation example model according to the power system, and running the calculation example program according to the calculation example model to obtain a result file are respectively implemented on the PSD-BPA software.

7. The method for voltage optimization of a power system by deep reinforcement learning according to claim 6, characterized in that: The steps of constructing the example model according to the power system are as follows: Retrieving an original example model in the PSD-BPA software, wherein the original example model includes a plurality of initial PV nodes; Based on the power system, in the original calculation example model, the installed capacity is set, reactive compensation equipment and on-load tap-changing transformers are added respectively, and the corresponding initial PV nodes are changed to PQ nodes to obtain an initial calculation example model; The example model is obtained by setting the voltage variation range of the example nodes, the transformation ratio variation range of the on-load tap-changing transformer, and the adjustable capacity of the reactive compensation device in the initial example model; the example nodes include the PQ node and the unchanged initial PV node.

8. The method for voltage optimization of a power system by deep reinforcement learning according to claim 1, characterized in that: The deep reinforcement learning algorithm is the DDQN algorithm.

9. A device for implementing the power system deep reinforcement learning voltage optimization method according to any one of claims 1 to 8, characterized in that: It includes a data acquisition unit, a testing unit, a calculation unit and an optimization analysis unit; The data acquisition unit: Used to obtain historical data sets and typical period input data sets based on the power system and transmit them to the test unit respectively; Used to obtain the time node and transmit it to the computing unit; The test unit: Used to adopt a machine learning algorithm based on a gradient boosting framework to obtain a test model based on the historical data set training; for obtaining a typical period output data set according to the typical period input data set based on the test model and transmitting the data set to the calculation unit; The computing unit: Used to obtain a node data set based on the time node and according to the output data set; Used to perform power flow calculation based on the power system and the node data set to obtain a node voltage set and transmit it to the optimization analysis unit; The optimization analysis unit: Used to obtain an optimization model based on the node voltage set using a deep reinforcement learning algorithm; Used to optimize the power system according to the optimization model.

10. The device according to claim 9, characterized in that: The calculation unit includes a data processing module and a power flow calculation module; The data processing module: Used to obtain a node data set based on the time node and according to the output data set; Used to obtain a calculation example program according to the node data set; used for transmitting the node voltage set to the optimization analysis unit; The power flow calculation module: Used to construct a calculation model according to the power system; Used to call the example program based on the example model to obtain a result file; Used to obtain the node voltage set according to the result file; Used to transmit the node voltage set to the data processing module.