Power private network operation process dynamic reconstruction method based on digital twinborn model
By building a digital twin model of power private network and a hybrid strategy network, the problem of traditional power private network reconstruction methods relying on expert experience is solved, the intelligence and modernization of power private network is realized, and the efficiency and stability of power grid operation are improved.
Patent Information
- Application Number
- CN202510209977.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-07-11
AI Technical Summary
The traditional power private network operation process reconstruction method relies on expert experience and lacks system data-driven and automated decision-making support, which leads to low reconstruction efficiency and difficulty in dealing with complex and changing power grid conditions, affecting the accuracy and effectiveness of reconstruction.
The dynamic reconstruction method of the power private network operation process based on the digital twin model is adopted. By building the power private network digital twin model, 3D modeling, material attribute addition, behavioral constraints and control logic of device objects are realized, combined with data exchange and control command transmission, and dynamic reconstruction and optimization are used for hybrid policy networks and reinforcement learning.
Real-time simulation and fault prediction of the operation process of the power private network are realized, system performance and efficiency are improved, operating costs are reduced, and the intelligence and modernization of the power private network is improved.
Smart Images

Figure CN120300760A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to smart grid technology, and particularly to a method for dynamically reconstructing the operation process of a power private network based on a digital twin model. Background Art
[0002] A power private network is a dedicated network that provides communication services for the power system, mainly used for key functions such as power grid dispatching, control, information transmission, and relay protection. It has the characteristics of "high reliability, real-time performance, security, and stability". The power private network plays an important role in the modernization and intelligentization process of the power system. The operation process of the power private network mainly includes multiple links such as data acquisition, transmission, processing, control instruction issuance, operation monitoring, fault handling, maintenance management, operation optimization, and security protection. These links are interrelated and jointly ensure the stable, reliable, and efficient operation of the power grid. By reconstructing the power grid operation process, efficient operation of the power private network can be achieved, and real-time monitoring, remote control, and intelligent management of the power grid can be carried out, thereby improving the overall performance and operation efficiency of the power grid. However, the traditional method for reconstructing the operation process of the power private network overly relies on expert experience, lacks systematic data-driven and automated decision support; moreover, due to the low degree of automation and intelligence, the reconstruction efficiency is not high, and it is difficult to cope with complex and changing power grid conditions; affecting the accuracy and effectiveness of the reconstruction.
[0003] In view of the above problems, a method for dynamically reconstructing the operation process of a power private network based on a digital twin model is proposed. This method can accurately model the large-scale and complex power private network, and then use artificial intelligence methods to dynamically adjust the operation strategy, solve the problem of backward means in the traditional reconstruction of the operation process of the power private network, and thus promote the intelligentization and modernization of the power private network. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method for dynamically reconstructing the operation process of a power private network based on a digital twin model in view of the deficiencies in the prior art.
[0005] The technical solution adopted by the present invention to solve its technical problems is: A method for dynamically reconstructing the operation process of a power private network based on a digital twin model, comprising the following steps:
[0006] 1) Construct a digital twin model for the operation process of the power private network,
[0007] 1.1) For each device object in the power private network, establish a 3D model of various devices in the power private network as a basic unit model, and the device objects include: generators, transformers, switchgear, transmission lines, distribution equipment, and control and protection equipment;
[0008] 1.2) Add corresponding material attributes to the model according to the collected material data.
[0009] 1.3) Constrain the behaviors of devices, that is, constrain and stipulate the operating parameters of various devices during the operation of the dedicated power network in normal and abnormal conditions;
[0010] 1.4) Implement the control logic of each device object in the dedicated power network in the digital twin model;
[0011] 1.5) Assemble the basic unit models, assemble each basic unit model according to the layout and topological structure of the actual system to form a complete system model;
[0012] For the already established device models, perform position layout and connection in the order of generator - transformer - switchgear - transmission line - distribution equipment - control and protection equipment. The generator, as the starting point of the power system, is used to convert mechanical energy into electrical energy and is connected to the transformer through the generator outlet cabinet; the transformer is used to adjust the voltage up or down to meet the needs of different grid levels and is connected to the high - voltage transmission line or medium - voltage distribution line according to the grid plan; the switchgear is connected between the transformer and the transmission line, as well as between the transformer and the distribution line, and is used to control the on - off of the circuit; the distribution equipment is used to distribute electrical energy from the high - voltage or medium - voltage transmission line to the low - voltage distribution line for users or loads to use; the control and protection equipment is used to monitor and protect the safe operation of the power system.
[0013] After ensuring that the physical connections and logical links between the device models are correct, realize the data exchange and control command transmission between the device models. First, determine the data types and formats that each device model needs to exchange: the generator mainly considers the output voltage; the transformer considers the input and output voltages and currents; the switchgear considers the opening and closing state (open / closed); the transmission line considers the voltage and current; the distribution equipment considers the voltage and current; the control and protection equipment considers the state of the relay protection device (activated / deactivated). The control command transmission includes: adjusting the output voltage of the generator, adjusting the switch state, and adjusting the state of the control and protection equipment. For the control commands, build a human - machine interface to visually present various data during the operation of the dedicated power network, including voltage, current, power, switch state, etc., and the control data can be input through this interface.
[0014] 1.6) Verify the model to ensure the correctness and effectiveness of the model. The model verification, according to different requirements, checks whether the output of the model is consistent with the output of the physical object;
[0015] 1.7) Realize the two - way interaction between the digital twin model and the real - world device;
[0016] First, obtain and transmit data. By installing sensors and data acquisition devices on physical devices, the operating data of the devices is collected in real time, including the output voltage and temperature of generators, the input and output voltages of transformers, the voltage and current of transmission lines, the voltage and current of distribution lines, and the switch states of control and protection devices. The data is transmitted from the devices to the digital twin model through wired or wireless communication technologies. Then, the collected data is fused with the data in the digital twin model to update the model state in real time.
[0017] From the digital twin model to the display device, the commands of the control system are sent to the physical devices through communication technologies, and corresponding operations are performed, such as adjusting the output voltage of generators and controlling the closing of a certain switch. After the physical devices execute the control commands, the new state information (such as device operating status, performance parameters, etc.) is sent back to the digital twin model for model update and further control decision-making. Through these steps, it can be ensured that the sensors and data acquisition devices on the physical devices can accurately collect the operating data and transmit the data to the digital twin model, providing support for the analysis and optimization of the power system.
[0018] 2) Define the state space S of the digital twin model of the power private network; the state space should include all possible states that the agent can observe in the environment. In the digital twin model of the power private network, the state space mainly includes device status, grid status, and environmental factors.
[0019] 3) Define the action space of the digital twin model of the power private network; the action space defines all possible actions that the agent can take, and these actions are used to control grid devices or related parameters to achieve the optimal management of the grid. The action space of the power private network includes generator control, transformer control, transmission line control, distribution line control, load control, and protection circuit control.
[0020] 4) The neural architecture of the hybrid strategy network is composed of a multi-layer perceptron (MLP) and a classifier.
[0021] The construction of the MLP part includes an input layer, hidden layers, and an output layer. The input layer takes the grid state as the input, so the number of input layers is the number of grid states. There can be multiple hidden layers, and each layer contains multiple neurons. The number of hidden layers is taken as 3 layers. If more detailed grid scheduling or optimization problems need to be processed, the number of hidden layers is appropriately increased; the number of neurons in each layer is taken as 128, and the number of neurons is increased or decreased according to the convergence speed of the model and the complexity of the data. If the model does not converge for a long time, the number of neurons is decreased. If the model effect is not good or the data is more complex, the number of neurons is increased. The output layer is responsible for outputting the numerical values of continuous actions, including voltage, current, etc. Considering the rationality of the output values, the sampled action values need to be restricted after each sampling.
[0022] The classifier part includes an input layer, a hidden layer, and an output layer. The input layer takes the power grid state as input; the number of hidden layers is 3, and each layer has 128 neurons, which are used to learn the mapping relationship between the state and discrete actions; the output layer outputs the probability distribution of each discrete action, and the softmax function is used to ensure that the output is a probability distribution for selecting the action with the highest probability.
[0023] 5) Build a value network to estimate the expected return for a given state and action, thereby helping the agent evaluate the long-term value of different states and actions. The design of the value network should match the policy network, so a hybrid network is selected, which consists of an MLP and a classifier. The input of the value network is the current state, including the real-time state of power grid equipment, the power grid topology, and environmental factors; the MLP part directly outputs the expected return value, while the classifier part outputs the value estimate of each discrete action.
[0024] 6) Initialize the policy network and the value network. The policy network is used to select actions according to the current environmental state, with the current state s as its input and the probability distribution of action a as its output; the value network is used to evaluate the long-term value of state s. Among them, the training of the policy network includes defining a loss function, backpropagation, and an optimizer. The loss function needs to be defined separately for continuous actions and discrete actions. Mean Squared Error (MSE) is used for continuous actions, and the cross-entropy loss function is selected for discrete actions; then the backpropagation algorithm is used to optimize the network parameters and minimize the loss function; the Adam optimizer is selected to update the network parameters.
[0025] During the learning process, the agent obtains environmental information through the digital twin model of the power private network. The policy network will select the corresponding action a according to the current environmental state and the value evaluation of the value network, and then update the agent's state s, reward value reward, and determine whether to end the task at the next moment according to the current action a. At the same time, the three are stored in the data buffer for the agent's learning. When each round ends, the reward value is discounted and decayed, and the data in the buffer is extracted and input into the PPO algorithm for update, including calculating the policy loss, calculating the value loss, and updating the network parameters according to the loss. The network parameter weights are updated using the state s, action a, reward r, and the advantage function in each update iteration, and finally the deviation correction optimizes the objective function;
[0026] Among them, the reward function is designed as follows. The goal of the reward function design is to improve the energy efficiency ratio and enhance the system stability. Improving the energy efficiency ratio means increasing the energy conversion efficiency and reducing energy losses, which requires optimizing the load distribution and reducing energy waste in the system. Improving the system stability considers two cases: firstly, during the normal operation of the power grid, maintaining the voltage within the allowable range to avoid overvoltage or undervoltage; secondly, in the case of a circuit fault, being able to detect and locate the fault, isolate the fault area, and perform network reconstruction to finally quickly restore power supply to the area. Combining the above goals, the reward function is defined as:
[0027] R(s,a) = α·R efficiency (s,a) + β·R stability (s,a),
[0028] Among them, R efficiency (s,a) represents the reward regarding the energy efficiency ratio, and R stability (s,a) represents the reward regarding the stability, and α and β are the weighting coefficients respectively.
[0029]
[0030] Among them, V out represents the output voltage of the generator, and V in represents the input voltage of the transformer.
[0031] R stability (s,a) = R voltage_stability (s,a) + R fault_handling (s,a),
[0032] Among them, R voltage_stability (s,a) represents the reward related to voltage, and R fault_handling represents the reward for fault handling.
[0033]
[0034] Among them, r v is the reward for voltage stability, representing the reward value within the voltage range of the system.
[0035] R fault_handling (s,a) = r fh ,
[0036] Among them, r fh is the reward coefficient for fault handling, representing the performance of the system in fault handling.
[0037] 7) Deploy this model into the actual system, and the agent can select the optimal action according to the current state to achieve the dynamic reconstruction and optimization of the power private network.
[0038] The beneficial effects produced by the present invention are as follows:
[0039] 1. The present invention utilizes a digital twin model to real - time simulate the operation state of the power grid, predicts potential faults through agents, adjusts control parameters, improves the performance and efficiency of the system, and reduces operating costs.
[0040] 2. The present invention adopts a reinforcement learning method with policy optimization to dynamically reconstruct the operation process plan of the power special network in the current environment, enhancing the intelligence of the operation and maintenance process of the power special network. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:
[0042] Figure 1 is the method flow chart of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0044] As Figure 1 shown, a method for dynamically reconstructing the operation process of a power special network based on a digital twin model includes the following steps:
[0045] 1) Construct a digital twin model of the operation process of the power special network,
[0046] 1.1) For each device object in the power special network, establish 3D models of various devices in the power special network as basic unit models. The device objects include: generators, transformers, switchgear, transmission lines, distribution equipment, and control and protection equipment;
[0047] 1.2) According to the collected material data, add corresponding material attributes to the models;
[0048] 1.3) Constrain the device behavior, that is, constrain and specify the operation parameters of various devices in the power special network during normal and abnormal operations;
[0049] 1.4) Implement the control logic of each device object in the power special network in the digital twin model;
[0050] 1.5) Assemble the basic unit models, and assemble each basic unit model together according to the layout and topological structure of the actual system to form a complete system model;
[0051] For the established equipment model, perform the position layout and connection in the order of generator - transformer - switchgear - transmission line - distribution equipment - control and protection equipment. The generator, as the starting point of the power system, is used to convert mechanical energy into electrical energy and is connected to the transformer through the generator outgoing cabinet; the transformer is used to adjust the voltage increase or decrease to meet the needs of different grid levels and is connected to the high - voltage transmission line or medium - voltage distribution line according to the grid plan; the switchgear is connected between the transformer and the transmission line, as well as between the transformer and the distribution line, and is used to control the on - off of the circuit; the distribution equipment is used to distribute electrical energy from the high - voltage or medium - voltage transmission line to the low - voltage distribution line for users or loads; the control and protection equipment is used to monitor and protect the safe operation of the power system.
[0052] After ensuring the correct physical connection and logical link between the equipment models, realize the data exchange and control command transfer between the equipment models. First, determine the data types and formats that each equipment model needs to exchange: for the generator, mainly consider the output voltage; for the transformer, consider the input and output voltages and currents; for the switchgear, consider the opening and closing state (open / closed); for the transmission line, consider the voltage and current; for the distribution equipment, consider the voltage and current; for the control and protection equipment, consider the state of the relay protection device (activated / deactivated). The control command transfer includes: adjusting the output voltage of the generator, adjusting the switch state, and adjusting the state of the control and protection equipment. For the control commands, build a human - machine interface to visually present various data during the operation of the power private network, including voltage, current, power, switch state, etc., and the control data can be input through this interface.
[0053] 1.6) Verify the model to ensure its correctness and effectiveness. The model verification, according to different requirements, checks whether the output of the model is consistent with the output of the physical object.
[0054] 1.7) Realize the two - way interaction between the digital twin model and the real - world equipment.
[0055] First, acquire and transmit data. By installing sensors and data acquisition devices on the real - world equipment, collect the operation data of the equipment in real time, including the output voltage and temperature of the generator, the input and output voltages of the transformer, the voltage and current of the transmission line, the voltage and current of the distribution line, and the switch state of the control and protection equipment; transmit the data from the equipment to the digital twin model through wired or wireless communication technology; then fuse the collected data with the data in the digital twin model to update the model state in real time.
[0056] From the digital twin model to the display device, the commands of the control system are sent to the physical device through communication technology, and corresponding operations are performed, such as adjusting the output voltage of the generator and controlling the closing of a certain switch. After the physical device executes the control command, it sends the new state information (such as device operating status, performance parameters, etc.) back to the digital twin model for model update and further control decisions. Through these steps, it can be ensured that the sensors and data acquisition devices on the physical device can accurately collect operation data and transmit the data to the digital twin model to support the analysis and optimization of the power system.
[0057] Based on the established digital twin model of the power private network, dynamic reconstruction during the operation of the power private network is carried out;
[0058] 2) Define the state space S of the digital twin model of the power private network; the state space should include all possible states that the agent can observe in the environment. In the digital twin model of the power private network, the state space mainly includes device status, grid status, and environmental factors;
[0059] The device status includes generator status G, transformer status T, transmission line status L, distribution line status D, protection device status P, and control device status C;
[0060] The state space of the overall digital twin model of the power private network can be expressed as S = {G, T, L, D, P, C}. The generator status G = {V G , I G , P G , T G}, where V G is the output voltage of the generator, I G is the current of the generator, P G is the output power of the generator, and T G is the temperature of the generator; the transformer status T = {V Tin , V Tout , Load T}, where V Tin is the input voltage of the transformer, V Tout is the output voltage of the transformer, and Load T is the load of the transformer; the transmission line status L = {V L , I L , Z L , C L}, where V L is the voltage on the transmission line, I L is the current on the transmission line, Z L is the impedance of the transmission line, and C L is the capacitance of the transmission line; the distribution line status D = {V D , ID , Load D}, where V D is the voltage on the distribution line, and I D is the current on the distribution line, and Load D is the load condition of the distribution line; the protection device status P = {R1, R2,..., R n}, where R n represents the switch status of the nth relay; the control device status C = {S1, S2,..., S n}, where S n represents the switch status of the nth line.
[0061] The power grid status includes power grid parameters and power grid load conditions. The power grid parameters include the impedance and capacitance of each transmission line, and the power grid load includes the total load power, total load reactive power, and total load current of the power grid. The environmental factors mainly consider the environmental temperature Temperature and environmental humidity Humidity.
[0062] 3) Define the action space of the digital twin model of the power private network; the action space defines all possible actions that the agent can take, and these actions are used to control power grid equipment or related parameters to achieve the optimal management of the power grid. The action space of the power private network includes generator control, transformer control, transmission line control, distribution line control, load control, and protection circuit control.
[0063] The following is the detailed definition of the action space, using letters to represent each action. The action space is defined as
[0064] A = {A G , S G , A T , S T , A L , S L , A D , S D , A Load , A protection , S protection}, where
[0065] A G represents adjusting the output power of the generator, and S G represents switching the operating state of the generator (start / stop), and A T represents adjusting the transformation ratio of the transformer, and S T represents switching the operating state of the transformer (energized / withdrawn), and A L represents adjusting the transmission capacity of the transmission line, and S L represents switching the switch status of the transmission line (open / closed), and A DDenote the adjustment of the transmission capacity of the distribution line, S D Denote the switching state (open / closed) of the switch of the distribution line, A Load Denote the connection or disconnection of the load, A protection Denote the action parameters of the automatic reclosing device, S protection Denote the switching state (activated / closed) of the relay protection device.
[0066] 4) Construct a policy network, using a hybrid policy network; the neural architecture of the hybrid policy network consists of a multi-layer perceptron MLP and a classifier;
[0067] The construction of the MLP part includes an input layer, a hidden layer, and an output layer. The input layer takes the grid state as input, so the number of input layers is the number of grid states. The hidden layer can have multiple layers, and each layer contains multiple neurons. The number of hidden layers is taken as 3 layers. If more detailed power grid scheduling or optimization problems need to be processed, the number of hidden layers is appropriately increased; the number of neurons in each layer is taken as 128, and the number of neurons is increased or decreased according to the convergence speed of the model and the complexity of the data. If the model does not converge for a long time, the number of neurons is decreased. If the model performance is poor or the data is more complex, the number of neurons is increased. The output layer is responsible for outputting the numerical values of continuous actions, including voltage, current, etc. Considering the rationality of the output values, the sampled action values need to be restricted after each sampling.
[0068] The classifier part includes an input layer, a hidden layer, and an output layer. The input layer takes the grid state as input; the number of hidden layers is taken as 3, and each layer has 128 neurons, which are used to learn the mapping relationship between the state and the discrete action; the output layer outputs the probability distribution of each discrete action, and the softmax function is used to ensure that the output is a probability distribution, so as to select the action with the highest probability.
[0069] 5) Construct a value network, which is used to estimate the expected return of a given state and action, so as to help the agent evaluate the long-term value of different states and actions. The design of the value network should match the policy network, so a hybrid network is selected, which consists of an MLP and a classifier. The input of the value network is the current state, including the real-time state of grid equipment, the grid topology, and environmental factors; the MLP part directly outputs the expected return value, while the classifier part outputs the value estimation of each discrete action.
[0070] 6) Initialize the policy network and the value network. The policy network is used to select actions based on the current environmental state. Its input is the current state s, and its output is the probability distribution of action a. The value network is used to evaluate the long-term value of state s. Among them, the training of the policy network includes defining the loss function, backpropagation, and optimizer. The loss function needs to be defined separately for continuous actions and discrete actions. The mean squared error (MSE) is used for continuous actions, and the cross-entropy loss function is selected for discrete actions. Then, the backpropagation algorithm is used to optimize the network parameters and minimize the loss function. The Adam optimizer is selected to update the network parameters.
[0071] During the learning process, the agent obtains environmental information through the digital twin model of the power private network. The policy network will select the corresponding action a based on the current environmental state and the value evaluation of the value network. Then, according to the current action a, the agent updates the state s, the reward value reward, and determines whether to end the task at the next moment, and stores the three in the data buffer for the agent's learning. When each round ends, the reward value is discounted and decayed, and the data in the buffer is extracted and input into the PPO algorithm for update, including calculating the policy loss, calculating the value loss, and updating the network parameters according to the loss. In each update iteration, the network parameter weights are updated using the state s, action a, reward r, and the advantage function, and finally the deviation correction optimizes the objective function;
[0072] Among them, the reward function is designed as follows. The goal of the reward function design is to improve the energy efficiency ratio and improve the system stability. Improving the energy efficiency ratio means improving the energy conversion efficiency and reducing energy losses, which requires optimizing the load distribution and reducing energy waste in the system. Improving the system stability considers two cases: First, during the normal operation of the power grid, maintain the voltage within the allowable range to avoid overvoltage or undervoltage; Second, in the case of a circuit fault, be able to detect and locate the fault, isolate the fault area, perform network reconstruction, and finally quickly restore power supply to the area. Combining the above goals, the reward function is defined as:
[0073] R(s,a) = α·R efficiency (s,a) + β·R stability (s,a),
[0074] Among them, R efficiency (s,a) represents the reward regarding the energy efficiency ratio, and R stability (s,a) represents the reward regarding the stability, and α and β are the weighting coefficients respectively.
[0075]
[0076] Among them, V out represents the output voltage of the generator, and V in represents the input voltage of the transformer.
[0077] R stability (s,a) = R voltage_stability (s,a) + R fault_handling (s,a),
[0078] Among them, R voltage_stability (s,a) represents the reward related to voltage, and R fault_handling represents the reward for fault handling.
[0079]
[0080] Among them, r v is the reward for voltage stability, representing the reward value of the system within the voltage range.
[0081] R fault_handling (s,a) = r fh ,
[0082] Among them, r fh is the reward coefficient for fault handling, representing the performance of the system in fault handling.
[0083] 7) Conduct policy evaluation to verify the effect of the model in the actual power private network, deploy the verified model into the actual system, and the intelligent agent can select the optimal action according to the current state to achieve the dynamic reconstruction and optimization of the power private network.
[0084] After the policy network converges, it is necessary to verify the feasibility of the model, that is, evaluate its performance in the actual power private network environment. The higher the energy efficiency ratio and the higher the stability, the better the model.
[0085] It should be understood that those of ordinary skill in the art can make improvements or transformations according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.
Claims
1. A dynamic reconstruction method for the operation process of a power private network based on a digital twin model, characterized in that, It includes the following steps: 1) Construct a digital twin model for the operation process of the power private network based on the power private network; 2) Define the state space S of the digital twin model of the power private network; the state space should include all possible states that the agent can observe in the environment, and the state space mainly includes equipment status, power grid status, and environmental factors; 3) Define the action space of the digital twin model of the power private network; the action space defines all possible actions that the agent can take; the action space of the power private network includes generator control, transformer control, transmission line control, distribution line control, load control, and protection circuit control; 4) Construct a hybrid policy network; the policy network is used to select actions according to the current environmental state, and its input is the current state s, and the output is the probability distribution of the action a; The neural architecture of the hybrid policy network consists of a multi-layer perceptron MLP and a classifier; 5) Construct a value network, which is used to estimate the expected return of a given state and action, so as to help the agent evaluate the long-term value of different states and actions; 6) Initialize the policy network and the value network, and carry out the learning of the agent; During the learning process, the agent obtains environmental information through the digital twin model of the power private network. The policy network will select the corresponding action a according to the current environmental state and the value evaluation of the value network, and then update the state s, reward value reward, and determine whether to end the task of the agent at the next moment according to the current action a. At the same time, the three are stored in the data buffer for the learning of the agent; When each round ends, the reward value is discounted and decayed, and the data in the buffer is extracted and input into the policy optimization algorithm for update, including calculating the policy loss, calculating the value loss, and updating the network parameters according to the loss; In each update iteration, the network parameter weights are updated using the state s, action a, and reward r, and finally the optimization objective function is optimized through bias correction; 7) Deploy the optimized model into the actual system, and the agent selects the optimal action according to the current state to realize the dynamic reconstruction and optimization of the power private network.
2. The dynamic reconstruction method for the operation process of the power private network based on the digital twin model according to claim 1, characterized in that In step 1) above, to construct a digital twin model for the operation process of the power private network, the specific steps are as follows: 1.1) For each device object in the power private network, establish a 3D model of various devices in the power private network as the basic unit model, and the device objects include: generators, transformers, switchgear, transmission lines, distribution equipment, and control and protection equipment; 1.2) According to the collected material data, add the corresponding material attributes to the model; 1.3) Constrain the device behavior, that is, constrain and specify the operating parameters of various devices in the power private network during normal operation and abnormal conditions; 1.4) Implement the control logic of each device object in the power private network in the digital twin model; 1.5) Assemble the basic unit models, and assemble each basic unit model together according to the layout and topological structure of the actual system to form a complete system model; 1.6) Verify the model to ensure the correctness and effectiveness of the model. The model verification checks whether the output of the model is consistent with the output of the physical object according to different requirements; 1.7) Implement two-way interaction between the digital twin model and the real device.
3. The dynamic reconstruction method for the operation process of the power private network based on the digital twin model according to claim 1, characterized in that In step 1.5), specifically as follows: For the already established device models, perform position layout and connection in the order of generator - transformer - switchgear - transmission line - distribution equipment - control and protection equipment; the generator, as the starting point of the power system, is used to convert mechanical energy into electrical energy and is connected to the transformer through the generator outlet cabinet; the transformer is used to adjust the voltage increase or decrease to meet the needs of different grid levels and is connected to the high-voltage transmission line or medium-voltage distribution line according to the grid plan; The switchgear is connected between the transformer and the transmission line, and between the transformer and the distribution line, and is used to control the on / off of the circuit; The distribution equipment is used to distribute electrical energy from the high-voltage or medium-voltage transmission line to the low-voltage distribution line for users or loads; the control and protection equipment is used to monitor and protect the safe operation of the power system; After ensuring that the physical connections and logical links between the device models are correct, realize data exchange and control command transmission between the device models.
4. The dynamic reconstruction method for the operation process of the power private network based on the digital twin model according to claim 1, wherein In step 4), the MLP includes an input layer, a hidden layer, and an output layer; the input layer takes the grid state as input, and the number of input layers is the number of grid states; the number of hidden layers is 3, and the number of neurons in each layer is 128. The output layer is responsible for outputting the values of continuous actions, including voltage and current.
5. The dynamic reconstruction method for the operation process of the power private network based on the digital twin model according to claim 1, wherein In step 4), the classifier part includes an input layer, a hidden layer, and an output layer; the input layer takes the grid state as input; the number of hidden layers is 3, and each layer has 128 neurons, which are used to learn the mapping relationship between the state and the discrete action; the output layer outputs the probability distribution of each discrete action, and the softmax function is used to ensure that the output is a probability distribution for selecting the action with the highest probability.
6. The dynamic reconstruction method for the operation process of the power private network based on the digital twin model according to claim 1, wherein In step 5), the reward function is designed as follows: R(s,a) = α·R efficiency (s,a) + β·R stability (s,a), Among them, R efficiency (s,a) represents the reward regarding the energy efficiency ratio, and R stability (s,a) represents the reward regarding stability, where α and β are the weighting coefficients respectively; Among them, V out represents the output voltage of the generator, and V in represents the input voltage of the transformer; R stability (s,a) = R voltage_stability (s,a) + R fault_handling (s,a), Among them, R voltage_stability (s,a) represents the voltage-related reward, R fault_handling represents the reward for fault handling; Among them, r v is the reward for voltage stability, representing the reward value of the system within the voltage range; R fault_handling (s,a) = r fh , Among them, r fh is the reward coefficient for fault handling, indicating the system's performance in fault handling.
7. An electronic device, characterized in that it includes: one or more processors; and a storage device for storing one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the method according to any one of claims 1 to 6.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by the processor, implements the method according to any one of claims 1 to 6.