Power distribution system adaptive operation method considering differentiated demands based on joint learning
By building a recombinant soft switch R-SOP model and agent based on joint learning, the differentiated regulation needs of traditional distribution systems in the face of load changes and distributed energy access are solved, and the adaptive operation of the distribution system is realized, the flexibility and adaptability of the system is improved, power loss is reduced and voltage stability is maintained.
Patent Information
- Application Number
- CN202510766238.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-02
AI Technical Summary
When traditional distribution systems face changes in load characteristics, distributed energy access and grid topology evolution, it is difficult to meet differentiated regulation needs, especially the regulation needs of power loss, voltage stability and load balancing, and the capacity limitation of traditional soft switches affects the flexibility and economy of the system.
Using a joint learning method, a recombinant soft switch R-SOP model is constructed, and through a joint learning agent of the soft actor-critic network and a dual deep Q network, adaptive control of the recombinant soft switch is realized. Combined with loss reduction control, voltage optimization and demand response mode, an adaptive switching mechanism is designed to optimize feeder selection and port power scheduling.
It improves the flexibility and adaptability of the power distribution system, realizes rapid and accurate decision-making in a time-varying environment, reduces power loss, and maintains voltage stability, meeting differentiated regulation needs.
Smart Images

Figure CN120579854A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of power distribution system operation optimization, and specifically is a power distribution system adaptive operation method based on joint learning and considering differentiated needs. Background Art
[0002] Distribution systems must meet control requirements such as reactive power optimization, power loss minimization, voltage stability, and load balancing. These requirements are becoming increasingly complex with changing load characteristics, the high penetration of distributed energy resources, and evolving grid topologies. Simultaneously, the rapid development of power electronics technology has accelerated the evolution of traditional distribution networks into flexible ones. While the introduction of closed-loop structures improves system flexibility and controllability, it also introduces new operational characteristics such as bidirectional power flow, placing higher demands on control strategies. Furthermore, the large-scale integration of distributed energy resources further exacerbates system uncertainty, requiring control measures to be more flexible and adaptable. Traditional, single control modes are unable to meet the differentiated control requirements of power systems. In contrast, adaptive operation control strategies for power systems offer greater flexibility and cost-effectiveness. Different control modes can be adopted based on varying environmental conditions and control objectives, ultimately saving energy, reducing costs, and improving economic efficiency.
[0003] To improve system flexibility and stability, flexible interconnect devices (such as soft switches) are increasingly being applied to distribution systems. These devices can dynamically optimize power flow distribution, improve energy efficiency, and enhance operational adaptability. However, the power transmission and voltage / reactive power support capabilities of traditional soft switches are limited by the capacity of symmetrical converters. Due to the high cost of converters, distribution systems are constrained by the fixed capacity of their ports, hindering adaptability while balancing cost and flexibility. Improving system flexibility and cost-effectiveness while reducing power losses through the introduction of reconfigurable soft switches, combined with feeder selection and converters, remains a key technical challenge facing power networks. Summary of the Invention
[0004] In order to address the shortcomings of the above-mentioned existing technologies, the present invention proposes a distribution system adaptive operation method based on joint learning that takes into account differentiated needs, so as to achieve adaptive operation of the distribution system by adaptively adjusting the control mode, thereby improving the system flexibility and adaptability.
[0005] In order to achieve the above-mentioned object, the present invention adopts the following technical solutions:
[0006] The adaptive operation method of a power distribution system based on joint learning and considering differentiated needs of the present invention is characterized in that it includes the following steps:
[0007] Step 1: Establish a reconfigurable soft switch R-SOP model;
[0008] Step 2: Construct the state, action, and reward function of the joint learning agent based on the soft actor-critic network and the dual deep Q network at time t, and modify the reconfigurable soft switch R-SOP action output by the joint learning agent at time t;
[0009] Step 3: Based on step 2, the established joint learning agent is trained to obtain a trained joint learning agent;
[0010] Step 4: Following the process of step 3, the joint learning agent is trained under the control modes of loss reduction control P, voltage optimization U, and demand response D, respectively, to obtain joint learning agents under different control modes;
[0011] Step 5: Construct adaptive switching conditions among loss reduction control P, voltage optimization U, and demand response D. When the corresponding adaptive switching conditions are met, use the joint learning agent under the corresponding control mode to make decisions on the port power scheduling and feeder selection status of the reconfigurable soft switch R-SOP in a time-varying environment to achieve adaptive operation of the distribution system.
[0012] The adaptive operation method of the power distribution system considering differentiated needs based on joint learning of the present invention is also characterized in that, in step 1, a reconfigurable soft switch R-SOP model is established by equations (1) to (7):
[0013] (1)
[0014] (2)
[0015] (3)
[0016] (4)
[0017] (5)
[0018] (6)
[0019] (7)
[0020] In formula (1) to formula (7), 、 They represent the set of all converters in R-SOP and the set of nodes connected to the converters respectively; Indicates the inverter capacity; For the inverter The capacity ratio of is the total capacity of all converters; 、 and They are Always connected to the node Total losses, DC side power and AC side power of all converters; for All converters are connected to the node at the moment Reactive power on for All converters are connected to the node at the moment Apparent power on is the loss coefficient of R-SOP; The state of the feeder selection in R-SOP, indicating the node exist Is the moment consistent with the inverter Connectivity.
[0021] Furthermore, the step 2 includes the following steps:
[0022] Step 2.1: Construct a soft actor-critic network and a dual deep Q network, where the soft actor-critic network includes: Critic1 network, Critic2 network, target Critic1 network, target Critic2 network and Actor network;
[0023] The dual deep Q network includes: a main Q network and a target Q network with the same structure;
[0024] Step 2.2: Use formula (8) to construct the state of the joint learning agent at time t :
[0025] (8)
[0026] In formula (8), and Represents the nodes at time t Active load and reactive load consumed; Indicates that the photovoltaic power station injects into the node at time t Active power; Indicates the tap position of the on-load tap-changer at time t; represents the reactive power injected by the capacitor bank at time t;
[0027] Step 2.3: Use formula (9) to construct the action of the joint learning agent at time t :
[0028] (9)
[0029] In formula (9), are the sets of converters that transmit active power in R-SOP respectively; and are the active power transmission of converter m at time t in R-SOP and the reactive power support of converter n at time t in R-SOP respectively; is the feeder selection state of the port of converter n at time t;
[0030] Step 2.4: Use equations (10) and (11) to construct the reward function at time t :
[0031] (10)
[0032] (11)
[0033] In formula (10)-formula (11), is the control mode set; P, U and D represent loss reduction control, voltage optimization and demand response modes respectively; 、 and are the reward functions at time t for the control modes of loss reduction control P, voltage optimization U, and demand response D, respectively; Indicates any control mode at time t The flag, when any control mode When activated at time t, the corresponding flag is 1;
[0034] Step 2.5: Use equations (12)-(14) to Active power transmission in and reactive support Make corrections:
[0035] (12)
[0036] (13)
[0037] (14)
[0038] In formula (12)-formula (14), Indicates the converter for active power transmission capacity, represents the capacity of the last converter, Indicates the total number of converters in R-SOP;
[0039] Step 2.6: Use formula (15) to Feeder selection status in Make corrections:
[0040] (15)
[0041] In formula (15), Indicates feeder selection status The number of choices, Indicates the feeder selection status is regenerated.
[0042] Furthermore, the step 3 includes the following steps:
[0043] Step 3.1: Initialize the parameters of the Critic1 network in the soft actor-critic network , parameters of Critic2 network and the parameters of the Actor network ;
[0044] Initialize the parameters of the main Q network in the dual deep Q network ;
[0045] The parameters of the Critic1 network are Parameters assigned to the target Critic1 network , set the parameters of the Critic2 network Parameters assigned to the target Critic2 network , set the parameters of the main Q network Parameters assigned to the target Q network ;
[0046] Step 3.2: State of the moment Input to the joint learning agent and output Momentary action , by interacting with the environment, computing Reward value at the moment and State of the moment , thus obtaining a sample data And store it in the experience replay pool middle;
[0047] Step 3.3: Use Equation (16)-Equation (17) to calculate the loss function of Critic1 network respectively and Critic2 network in The total loss function at time , and thus use the gradient descent method to adjust the parameters of the Critic1 network , parameters of Critic2 network To update:
[0048] (16)
[0049] (17)
[0050] In formula (16)-formula (17), For Critic1 network and Critic2 network The total target Q value at the moment; For Critic1 network or Critic2 network The loss function at the moment; is the discount factor; and Critic1 network or Critic2 network respectively State of the moment and actions The Q value output under the target Critic1 network or the target Critic2 network is State of the moment and actions The Q value of the output; is the adjustment coefficient of temperature entropy; In the Actor network Constant action The probability distribution of From the experience replay pool The expectation of a sample data extracted from ;
[0051] Step 3.4: Backpropagate the Actor network and minimize the loss function of the Actor network using formula (18) , to update the parameters of the Actor network :
[0052] (18)
[0053] In formula (18), For Actor Network The loss function at the moment; For Actor network Momentary action The probability distribution of
[0054] Step 3.5: Backpropagate the soft actor-critic network and minimize the temperature entropy loss function using Equation (19) , to update the adjustment coefficient of temperature entropy :
[0055] (19)
[0056] In formula (19), for Momentary action The expectation of the probability distribution of ; is the target entropy;
[0057] Step 3.6: Use equations (20) and (21) to calculate the main Q network in the dual deep Q network. The loss function at the moment, so as to use the gradient descent method to adjust the parameters of the main Q network To update:
[0058] (20)
[0059] (twenty one)
[0060] In formula (20)-formula (21), is the target Q value of the main Q network in the dual deep Q network at time t; and They represent the state of the main Q network in the dual deep Q network at time t. and actions The Q value of the output and the state of the target Q network at time t+1 and actions The Q value of the output; For expectations;
[0061] Step 3.7: Use the soft update method shown in Equation (22)-Equation (23) to adjust the parameters of the target Critic1 network , parameters of the target Critic2 network and the parameters of the main Q network To update:
[0062] (twenty two)
[0063] (twenty three)
[0064] In formula (22)-formula (23), Indicates assignment, is the smoothing factor;
[0065] Step 3.8: Train the joint learning agent according to the process of steps 3.3 to 3.7 to obtain the trained joint learning agent.
[0066] Furthermore, the adaptive switching bar in step 5 is constructed according to the following steps:
[0067] Step 5.1: Use equations (24) and (25) to obtain the voltage regulation margin of node i at time t and R-SOP reactive flexibility :
[0068] (twenty four)
[0069] (25)
[0070] In formula (24)-formula (25), and Both are indicator functions. When the corresponding conditions are met, the value of the indicator function is 1, otherwise, the value of the indicator function is 0. represents the voltage value on node i at time t, and Respectively represent the upper and lower limit values of the voltage safety range, represents the active power transmitted by R-SOP on node i at time t, represents the reactive power provided or absorbed by the R-SOP on node i at time t, represents the R-SOP capacity of node i at time t;
[0071] Step 5.2: When and When , it means that the adaptive switching condition of loss reduction control P is reached, and the reward function at time t under loss reduction control P is constructed using equations (26) and (27) ;
[0072] (26)
[0073] (27)
[0074] In formula (26)-formula (27), Indicates that the voltage of node i exceeds the limit at time t; represents the set of all branches; Indicates the total number of nodes; represents the square of the current on the branch ij between nodes i and j at time t, Represents the resistance on branch ij;
[0075] when and When , it means that the adaptive switching condition of voltage optimization U is achieved, and the reward function at time t under voltage optimization U is constructed using equations (28) and (29) ;
[0076] (28)
[0077] (29)
[0078] In formula (28)-formula (29), Indicates the voltage deviation degree of node i at time t;
[0079] when and When , it means that the adaptive switching condition of demand response D is achieved, and the reward function at time t under demand response D is constructed using equations (28) and (29): .
[0080] An electronic device of the present invention includes a memory and a processor, and is characterized in that the memory is used to store a program that supports the processor to execute the adaptive operation method, and the processor is configured to execute the program stored in the memory.
[0081] The present invention provides a computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and the computer program executes the steps of the adaptive operation method when the computer program is executed by a processor.
[0082] Compared with the prior art, the present invention has the following beneficial effects:
[0083] 1. This invention establishes an adaptive operation framework that considers the differentiated needs of power distribution systems to address differentiated regulation requirements. Three control modes—loss reduction control, voltage optimization, and demand response—are designed for different operating scenarios. An adaptive switching mechanism is also designed to enable flexible switching between modes. This framework achieves differentiated operation of the power distribution system by adjusting operating strategies in real time, enhancing the system's flexibility and adaptability.
[0084] 2. This invention introduces a reconfigurable soft switch within an adaptive control framework, overcoming the fixed port capacity limitations of traditional multi-terminal soft switches. Furthermore, the reconfigurable nature of the reconfigurable soft switch allows for real-time optimization of converter port power scheduling and feeder selection, dynamically adapting to the differentiated control requirements of the distribution system.
[0085] 3. This paper constructs a joint learning agent based on a soft actor-critic and dual deep Q algorithm to simultaneously process both discrete variables (i.e., feeder selection) and continuous variables (i.e., power scheduling) in a reconfigurable soft switch. This joint learning agent supports the training of the reconfigurable soft switch, improving its accuracy and adaptability, and enabling rapid and accurate decision-making in time-varying environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0086] Figure 1 It is the topology diagram of the improved IEEE33-node power distribution system;
[0087] Figure 2 is a graph of the training of the joint learning agent;
[0088] Figure 3 It is the three-dimensional diagram of voltage under the adaptive operation strategy;
[0089] Figure 4 This is the result diagram of active transmission and reactive support of each feeder of R-SOP;
[0090] Figure 5 It is the allocation diagram of R-SOP port capacity;
[0091] Figure 6 It is the box plot of voltage distribution under different cases;
[0092] Figure 7 The voltage conditions of node 11 under different cases and the voltage curve of the specific time period 7:00-8:00 are shown;
[0093] Figure 8 This is a comparison chart of training results under different algorithms. DETAILED DESCRIPTION
[0094] In this embodiment, a method for adaptive operation of a power distribution system based on joint learning and considering differentiated needs is provided, and the specific steps are as follows:
[0095] Step 1: Establish a reconfigurable soft switch R-SOP model using equations (1) to (7):
[0096] (1)
[0097] (2)
[0098] (3)
[0099] (4)
[0100] (5)
[0101] (6)
[0102] (7)
[0103] In formula (1) to formula (7), 、 They represent the set of all converters in R-SOP and the set of nodes connected to the converters respectively; Indicates the inverter capacity; For the inverter The capacity ratio of is the total capacity of all converters; 、 and They are Always connected to the node Total losses, DC side power and AC side power of all converters; for All converters are connected to the node at the moment Reactive power on for All converters are connected to the node at the moment Apparent power on is the loss coefficient of R-SOP; The state of the feeder selection in R-SOP, indicating the node exist Is the moment consistent with the inverter Connectivity.
[0104] Step 2: Construct the state, action, and reward function of the joint learning agent based on the soft actor-critic network and the dual deep Q network at time t, and modify the reconfigurable soft switch R-SOP action output by the joint learning agent at time t;
[0105] Step 2.1: Construct a soft actor-critic network and a dual deep Q network, where the soft actor-critic network includes: Critic1 network, Critic2 network, target Critic1 network, target Critic2 network and Actor network;
[0106] The dual deep Q network includes: a main Q network and a target Q network with the same structure;
[0107] Step 2.2: Use formula (8) to construct the state of the joint learning agent at time t , which is composed of the injected active power and reactive power of each node in the distribution network, the active power injected by photovoltaics, and the actions of discrete devices:
[0108] (8)
[0109] In formula (8), and Represents the nodes at time t Active load and reactive load consumed; Indicates that the photovoltaic power station injects into the node at time t Active power; Indicates the tap position of the on-load tap-changer at time t; It represents the reactive power injected by the capacitor bank at time t.
[0110] Step 2.3: Use formula (9) to construct the action of the joint learning agent at time t , which consists of a soft actor-critic network controlling the active power transmission and reactive power support of R-SOP, and a dual deep Q network controlling the feeder selection state of R-SOP:
[0111] (9)
[0112] In formula (9), are the sets of converters that transmit active power in R-SOP respectively; and are the active power transmission of converter m at time t in R-SOP and the reactive power support of converter n at time t in R-SOP respectively; is the feeder selection state of the port of converter n at time t.
[0113] Step 2.4: Use equations (10) and (11) to construct the reward function at time t :
[0114] (10)
[0115] (11)
[0116] In formula (10)-formula (11), is the control mode set; P, U and D represent loss reduction control, voltage optimization and demand response modes respectively; 、 and are the reward functions at time t for the control modes of loss reduction control P, voltage optimization U, and demand response D, respectively; Indicates any control mode at time t The flag, when any control mode When activated at time t, the corresponding flag is 1.
[0117] Step 2.5: For the R-SOP port power output action, due to active transmission and reactive support Need to meet their respective capacity constraints. Using formula (12)-formula (14) for Active power transmission in and reactive support Make corrections:
[0118] (12)
[0119] (13)
[0120] (14)
[0121] In formula (12)-formula (14), Indicates the converter for active power transmission capacity, represents the capacity of the last converter, Indicates the total number of converters in R-SOP;
[0122] Step 2.6: The feeder selection state of R-SOP must ensure the rationality of the connection to avoid all ports being concentrated on the same feeder and affecting the power transmission capacity. Feeder selection status in Make corrections:
[0123] (15)
[0124] In formula (15), Indicates feeder selection status The number of choices, Indicates the regeneration of feeder selection state. The rationality of feeder selection can be judged by formula (15). If all ports are not concentrated on the same feeder, this action is adopted; otherwise, the feeder selection action is regenerated;
[0125] Step 3: Based on step 2, the established joint learning agent is trained to obtain a trained joint learning agent;
[0126] Step 3.1: Initialize the parameters of the Critic1 network in the soft actor-critic network , parameters of Critic2 network and the parameters of the Actor network ;
[0127] Initialize the parameters of the main Q network in the dual deep Q network ;
[0128] The parameters of the Critic1 network are Parameters assigned to the target Critic1 network , set the parameters of the Critic2 network Parameters assigned to the target Critic2 network , set the parameters of the main Q network Parameters assigned to the target Q network .
[0129] Step 3.2: State of the moment Input to the joint learning agent and output Momentary action , by interacting with the environment, computing Reward value at the moment and State of the moment , thus obtaining a sample data And store it in the experience replay pool middle.
[0130] Step 3.3: Use Equation (16)-Equation (17) to calculate the loss function of Critic1 network respectively and Critic2 network in The total loss function at time , and thus use the gradient descent method to adjust the parameters of the Critic1 network , parameters of Critic2 network To update:
[0131] (16)
[0132] (17)
[0133] In formula (16)-formula (17), For Critic1 network and Critic2 network The total target Q value at the moment; For Critic1 network or Critic2 network The loss function at the moment; is the discount factor; and Critic1 network or Critic2 network respectively State of the moment and actions The Q value output under the target Critic1 network or the target Critic2 network is State of the moment and actions The Q value of the output; is the adjustment coefficient of temperature entropy; In the Actor network Constant action The probability distribution of From the experience replay pool The expectation of a sample data extracted from .
[0134] Step 3.4: Backpropagate the Actor network and minimize the loss function of the Actor network using formula (18) , to update the parameters of the Actor network :
[0135] (18)
[0136] In formula (18), For Actor Network The loss function at the moment; For Actor network Momentary action The probability distribution of .
[0137] Step 3.5: Backpropagate the soft actor-critic network and minimize the temperature entropy loss function using Equation (19) , to update the adjustment coefficient of temperature entropy , the goal is to keep the entropy of the strategy in a suitable range and control the randomness of the optimal strategy:
[0138] (19)
[0139] In formula (19), for Momentary action The expectation of the probability distribution of ; is the target entropy.
[0140] Step 3.6: Use equations (20) and (21) to calculate the main Q network in the dual deep Q network. The loss function at the moment, so as to use the gradient descent method to adjust the parameters of the main Q network To update:
[0141] (20)
[0142] (twenty one)
[0143] In formula (20)-formula (21), is the target Q value of the main Q network in the dual deep Q network at time t; and They represent the state of the main Q network in the dual deep Q network at time t. and actions The Q value of the output and the state of the target Q network at time t+1 and actions The Q value of the output; For expectation.
[0144] Step 3.7: Use the soft update method shown in Equation (22)-Equation (23) to adjust the parameters of the target Critic1 network , parameters of the target Critic2 network and the parameters of the main Q network To update:
[0145] (twenty two)
[0146] (twenty three)
[0147] In formula (22)-formula (23), Indicates assignment, is the smoothing factor, which is used to indicate the update speed of the target network.
[0148] Step 3.8: Train the joint learning agent according to the process of steps 3.3 to 3.7 to obtain the trained joint learning agent.
[0149] Step 4: Following the process of step 3, the joint learning agent is trained under the control modes of loss reduction control P, voltage optimization U, and demand response D, respectively, to obtain the joint learning agent under different control modes.
[0150] Step 5: Construct adaptive switching conditions between loss reduction control (P), voltage optimization (U), and demand response (D). When the corresponding adaptive switching conditions are met, use the joint learning agent under the corresponding control mode to make decisions on the port power scheduling and feeder selection state of the reconfigurable soft switch (R-SOP) in a time-varying environment to achieve adaptive operation of the distribution system.
[0151] Step 5.1: Use equations (24) and (25) to calculate the voltage regulation margin of node i at time t and R-SOP reactive flexibility :
[0152] (twenty four)
[0153] (25)
[0154] In formula (24)-formula (25), and Both are indicator functions. When the corresponding conditions are met, the value of the indicator function is 1, otherwise, the value of the indicator function is 0. represents the voltage value on node i at time t, and Respectively represent the upper and lower limit values of the voltage safety range, represents the active power transmitted by R-SOP on node i at time t, represents the reactive power provided or absorbed by the R-SOP on node i at time t, represents the R-SOP capacity of node i at time t.
[0155] Step 5.2: When and When , it indicates that the adaptive switching condition of loss reduction control P is reached. The system selects this mode, giving priority to minimizing power loss, and uses equations (26)-(27) to construct the reward function at time t under loss reduction control P ;
[0156] (26)
[0157] (27)
[0158] In formula (26)-formula (27), Indicates that the voltage of node i exceeds the limit at time t; represents the set of all branches; Indicates the total number of nodes; represents the square of the current on the branch ij between nodes i and j at time t, Represents the resistance on branch ij.
[0159] when and When , it means that the adaptive switching condition of voltage optimization U is achieved. The system selects this mode to reduce the network voltage deviation and uses equations (28)-(29) to construct the reward function at time t under voltage optimization U ;
[0160] (28)
[0161] (29)
[0162] In formula (28)-formula (29), Indicates the voltage deviation degree of node i at time t;
[0163] when and When , it means that the adaptive switching condition of demand response D is achieved, and the reward function at time t under demand response D is constructed using equations (28) and (29): .
[0164] when and When , it means that the adaptive switching condition of demand response D is achieved, and the reward function at time t under voltage optimization U is constructed using equations (28) and (29): This mode dynamically adjusts loads by regulating interruptible loads, reducing non-critical loads, and incentivizing users to shift peak demand. This is consistent with the voltage optimization mode, which aims to ensure voltage security.
[0165] The above method takes into account the differentiated control needs of the distribution system, uses reconfigurable soft switches to improve system flexibility and adaptability, and realizes fast and accurate decision-making of reconfigurable soft switches in time-varying environments based on a joint deep reinforcement learning algorithm, thereby realizing adaptive operation of the distribution system.
[0166] In this embodiment, an electronic device includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.
[0167] In this embodiment, a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are executed.
[0168] To help those skilled in the art better understand the present invention, the example analysis includes the following components:
[0169] 1. Test system and parameter settings:
[0170] In order to verify its effectiveness, the present invention adopts Figure 1 The improved IEEE 33-node system with 6-port R-SOP is shown as a test system, and a case analysis is performed. Figure 1 In the test system shown, five PV panels were connected to the grid, installed at nodes 7, 10, 18, 24, and 27, with capacities of 300 kWp, 700 kWp, 300 kWp, 500 kWp, and 700 kWp, respectively. The capacity, operating parameters, and connection locations of the voltage control equipment are shown in Table 1. A set of 4-port R-SOPs with a capacity of 4 kVA was installed between nodes 12, 22, 25, and 29. The loss coefficient of each converter was 0.02. An 11-speed OLTC was installed between nodes 1 and 2, with each speed representing a 1% regulation. A 6-speed CB was located at node 33, with each switching unit capable of compensating 60 kV of reactive power. The operation times were limited to 6 and 8, respectively. The node voltage safety range was set between 0.95 pu and 1.05 pu. The dataset used to train the federated learning agent contained 115 days of measurement data, 110 of which were used for training and 5 for testing. The scheduling period is 24 hours and the sampling interval is 5 minutes. Table 2 lists the specific parameter settings of the federated learning agent.
[0171] Table 1 Voltage control equipment parameters
[0172]
[0173] Table 2 Hyperparameter settings of the joint learning agent
[0174]
[0175] 2. Model training process:
[0176] In the adaptive operation strategy, the constructed joint learning agent was trained 30,000 times. The training process is as follows Figure 2 As shown in the figure, the total time cost for offline training of the three modes on IEEE33 nodes is approximately 40 minutes. Offline training only needs to be performed once, and if the distribution system grid remains unchanged, no retraining is required. As can be seen from the figure, the total reward fluctuation of the three modes gradually increases. The cumulative reward curve shows that as the number of training steps increases, the reward value gradually stabilizes and approaches the maximum value, indicating that the model performance is becoming stable. Initial fluctuations are large because the model is exploring different strategies, but as training progresses, the reward fluctuations decrease and tend to converge, indicating that the strategy is gradually optimized and the decision-making ability is improved.
[0177] 3. Analysis of the effectiveness of adaptive operation strategy:
[0178] Figure 3 Figure 3 shows the voltage variation of 33 nodes in the IEEE distribution network over a day using the proposed R-SOP multi-mode control strategy. As can be seen in the figure, all nodes remained within the safe range of 0.95–1.05 pu throughout the day.
[0179] The R-SOP action results of the adaptive control strategy on the test set are as follows: Figure 4 As shown, Figure 4 Part (a) shows the reactive support of R-SOP. Figure 4 Part (b) shows the active power transmission of R-SOP. Green, red and yellow represent P-Mode, V-Mode and D-Mode respectively. Figure 4 The system initially operates in P-Mode, aiming to minimize network losses while keeping voltage deviations small. Around 12:00 PM, the PV penetration rate increased and voltage fluctuations increased, causing the voltage margin at some nodes to fall below the set value. Therefore, the adaptive switches switched to V-Mode to minimize voltage deviations.
[0180] After voltage mode optimization, at about 12:30, each port of R-SOP adjusted reactive power, voltage fluctuations decreased, the voltage margin of all nodes reached the set value, and the adaptive switch switched back to P-Mode. At 18:00 when power consumption peaked, the system switched to V-Mode again. When the load exceeded the upper limit from 18:10 to 18:30, the R-SOP reactive margin reached the upper limit, and the system adaptively switched to L-Mode for load reduction, and restored to V-Mode before 18:30. After 21:00, as power consumption decreased and voltage stabilized, the system switched back to P-Mode to reduce economic costs. Among them, Figure 5The data shows the temporal changes in the capacity of the feeders connected to the R-SOP during a specific time period (14:00-16:00). The R-SOP can change the topology between VSCs according to the changes in the load demand within the system and the PVs output power over time, adjust the feeder power transmission capacity in real time, and flexibly adjust the output power.
[0181] 4. Case Comparative Analysis:
[0182] In order to evaluate the effectiveness of the proposed method, this paper compares it with three other cases.
[0183] Case 1: Without R-SOP control, the initial state of the distribution network is obtained as a basic case.
[0184] Case 2: Based on the method proposed in this invention, the R-SOP is replaced with a 4-port symmetrical SOP, but the ports do not have the ability to reassemble.
[0185] Case 3: Based on the method proposed in this invention, only R-SOP control in a single loss reduction mode is performed to optimize the distribution network.
[0186] Case 4: Optimize the distribution network using the R-SOP adaptive control strategy method proposed in this invention.
[0187] Figure 6 The voltage distribution diagrams of different cases on the test set are shown. The results show that in Case 1, due to the large load and small DG power generation, there is a lack of R-SOP support, and the voltage of some nodes is lower than the lower limit (0.95 pu). In contrast, Case 2 uses multi-terminal SOP adjustment, and the overall voltage is improved, but some nodes still exceed the upper limit (1.05 pu) or the lower limit (0.95 pu). Case 3 introduces R-SOP single loss mode adjustment, and at some moments, over-limit phenomenon may still occur, which is mainly limited by voltage margin constraint and single mode regulation strategy. Compared with other methods, the method proposed in the present invention can adjust R-SOP active and reactive power in multiple modes, meet the operation requirements of different time periods, perform best, and the voltage curve is significantly improved. All node voltages are within a safe range throughout the day.
[0188] In addition, Table 3 shows the comparison of voltage and loss performance under different cases. It can be seen from the table that compared with case 1 without R-SOP regulation, the method proposed by the present invention adopts the adaptive operation control strategy containing R-SOP to regulate and control the maximum voltage deviation of the distribution network, which is reduced by 53.58%, and the power loss is reduced by 12.01%. The minimum voltage value in case 1 is 0.9061, which is much lower than the voltage lower limit. In contrast, the voltage curves of the method proposed by the present invention are all maintained within the range of safe voltage [0.95, 1.05], and the voltage average value is 1.0109, and the overall voltage situation is more concentrated, close to the per-unit value. Moreover, the maximum voltage deviation using the method proposed by the present invention is lower than that compared with case 2 and case 3.
[0189] Table 3 Comparison of voltage and loss performance under different cases
[0190]
[0191] When t=7:00 am, the voltage curves under different cases are as follows Figure 7 As shown in part (a) of the figure, it can be observed that when no control strategy is adopted, the voltage at nodes 7-17 falls below the lower limit. In contrast, in Case 2 and Case 3, the voltage remains within the safe range using multi-terminal SOP and single-mode R-SOP, respectively. The proposed data-driven multi-mode R-SOP control method significantly improves the voltage over-limit problem by providing reactive power support and flexible control of power output, making the node voltage more stable.
[0192] Figure 7 Part (b) shows the comparison of the voltage curves of node 11 under different cases. The results show that when there is no R-SOP support, the voltage of this node is lower than the lower limit (0.95pu) due to the heavy load and small DG power generation. Compared with the unregulated method, the method of the present invention significantly improves the voltage curve, and the overall voltage is raised and maintained within a safe range. Compared with the multi-terminal SOP in Case 2, R-SOP can flexibly switch feeder selection, enhancing the adaptability of the distribution network to load changes and renewable energy access. At the same time, the use of adaptive operation control strategy helps to balance power loss and voltage deviation while maintaining voltage safety.
[0193] To verify the advantages of the joint algorithm of soft actor-critic SAC and dual deep Q network DDQN, Figure 8The training performance of different algorithms for adaptive operational control of R-SOP was compared. While both the deep deterministic policy gradient (DDPG) and DDQN joint algorithms and the proposed SAC and DDQN joint algorithm achieved high cumulative rewards, the former exhibited significant oscillations in the early stages of training, reflecting unstable exploration. These oscillations gradually subsided as training progressed. In contrast, the proposed SAC and DDQN joint algorithm was more stable and achieved higher rewards in the later stages. Furthermore, SAC incorporates an entropy adjustment term to enhance exploration, effectively balancing exploration and exploitation, avoiding local optima, and improving policy diversity. Furthermore, SAC employs a dual Q-network to reduce Q-value overestimation, improve training stability, and avoid divergence. Using a hybrid soft actor-critic (HSAC) algorithm to handle both continuous (port power) and discrete (feeder selection) actions simplifies the model architecture, but its discrete action selection is less stable and efficient than DDQN, which can lead to increased training complexity. Furthermore, SAC without entropy adjustment performed worse than the SAC and DDQN joint algorithm in both reward and stability, demonstrating the importance of entropy adjustment.
[0194] In this specification, the schematic descriptions of the present invention do not necessarily refer to the same embodiment or example. Those skilled in the art may combine and combine the different embodiments or examples described in this specification. In addition, the contents of the embodiments in this specification are merely an enumeration of the implementation forms of the inventive concept. The scope of protection of the present invention should not be considered limited to the specific forms described in the implementation cases. The scope of protection of the present invention also includes equivalent technical means that can be conceived by those skilled in the art based on the inventive concept.
Claims
1. A method for adaptive operation of a power distribution system based on joint learning and considering differentiated needs, characterized in that: The following steps are involved: Step 1: Establish a reconfigurable soft switch R-SOP model; Step 2: Construct the state, action, and reward function of the joint learning agent based on the soft actor-critic network and the dual deep Q network at time t, and modify the reconfigurable soft switch R-SOP action output by the joint learning agent at time t; Step 3: Based on step 2, the established joint learning agent is trained to obtain a trained joint learning agent; Step 4: Following the process of step 3, the joint learning agent is trained under the control modes of loss reduction control P, voltage optimization U, and demand response D, respectively, to obtain joint learning agents under different control modes; Step 5: Construct adaptive switching conditions among loss reduction control P, voltage optimization U, and demand response D. When the corresponding adaptive switching conditions are met, use the joint learning agent under the corresponding control mode to make decisions on the port power scheduling and feeder selection status of the reconfigurable soft switch R-SOP in a time-varying environment to achieve adaptive operation of the distribution system.
2. The adaptive operation method of a power distribution system based on joint learning and considering differentiated needs according to claim 1, characterized in that: In step 1, the reconfigurable soft switch R-SOP model is established using equations (1) to (7): (1) (2) (3) (4) (5) (6) (7) In formula (1) to formula (7), 、 They represent the set of all converters in R-SOP and the set of nodes connected to the converters respectively; Indicates the inverter capacity; For the inverter The capacity ratio of is the total capacity of all converters; 、 and They are Always connected to the node Total losses, DC side power and AC side power of all converters; for All converters are connected to the node at the moment Reactive power on for All converters are connected to the node at the moment Apparent power on is the loss coefficient of R-SOP; The state of the feeder selection in R-SOP, indicating the node exist Is the moment consistent with the inverter Connectivity.
3. The adaptive operation method of a power distribution system based on joint learning and considering differentiated needs according to claim 2, characterized in that: The step 2 comprises the following steps: Step 2.1: Construct a soft actor-critic network and a dual deep Q network, where the soft actor-critic network includes: Critic1 network, Critic2 network, target Critic1 network, target Critic2 network and Actor network; The dual deep Q network includes: a main Q network and a target Q network with the same structure; Step 2.2: Use formula (8) to construct the state of the joint learning agent at time t : (8) In formula (8), and Represents the nodes at time t Active load and reactive load consumed; Indicates that the photovoltaic power station injects into the node at time t Active power; Indicates the tap position of the on-load tap-changer at time t; represents the reactive power injected by the capacitor bank at time t; Step 2.3: Use formula (9) to construct the action of the joint learning agent at time t : (9) In formula (9), are the sets of converters that transmit active power in R-SOP respectively; and are the active power transmission of converter m at time t in R-SOP and the reactive power support of converter n at time t in R-SOP respectively; is the feeder selection state of the converter n port at time t; Step 2.4: Use equations (10) and (11) to construct the reward function at time t : (10) (11) In formula (10)-formula (11), is the control mode set; P, U and D represent loss reduction control, voltage optimization and demand response modes respectively; 、 and are the reward functions at time t for the control modes of loss reduction control P, voltage optimization U, and demand response D, respectively; Indicates any control mode at time t The flag, when any control mode When activated at time t, the corresponding flag is 1; Step 2.5: Use equations (12)-(14) to Active power transmission in and reactive support Make corrections: (12) (13) (14) In formula (12)-formula (14), Indicates the converter for active power transmission capacity, represents the capacity of the last converter, Indicates the total number of converters in R-SOP; Step 2.6: Use formula (15) to Feeder selection status in Make corrections: (15) In formula (15), Indicates feeder selection status The number of choices, Indicates the feeder selection status is regenerated.
4. The adaptive operation method of a power distribution system based on joint learning and considering differentiated needs according to claim 1, characterized in that: The step 3 comprises the following steps: Step 3.1: Initialize the parameters of the Critic1 network in the soft actor-critic network , parameters of Critic2 network and the parameters of the Actor network ; Initialize the parameters of the main Q network in the dual deep Q network ; The parameters of the Critic1 network are Parameters assigned to the target Critic1 network , set the parameters of the Critic2 network Parameters assigned to the target Critic2 network , set the parameters of the main Q network Parameters assigned to the target Q network ; Step 3.2: State of the moment Input to the joint learning agent and output Momentary action , by interacting with the environment, computing Reward value at the moment and State of the moment , thus obtaining a sample data And store it in the experience replay pool middle; Step 3.3: Use Equation (16)-Equation (17) to calculate the loss function of Critic1 network respectively and Critic2 network in The total loss function at time , and thus use the gradient descent method to adjust the parameters of the Critic1 network , parameters of Critic2 network To update: (16) (17) In formula (16)-formula (17), For Critic1 network and Critic2 network The total target Q value at the moment; For Critic1 network or Critic2 network The loss function at the moment; is the discount factor; and Critic1 network or Critic2 network respectively State of the moment and actions The Q value output under the target Critic1 network or the target Critic2 network is State of the moment and actions The Q value of the output; is the adjustment coefficient of temperature entropy; In the Actor network Constant action The probability distribution of From the experience replay pool The expectation of a sample data extracted from ; Step 3.4: Backpropagate the Actor network and minimize the loss function of the Actor network using formula (18) , to update the parameters of the Actor network : (18) In formula (18), For Actor Network The loss function at the moment; For Actor network Momentary action The probability distribution of Step 3.5: Backpropagate the soft actor-critic network and minimize the temperature entropy loss function using Equation (19) , to update the adjustment coefficient of temperature entropy : (19) In formula (19), for Momentary action The expectation of the probability distribution of ; is the target entropy; Step 3.6: Use equations (20) and (21) to calculate the main Q network in the dual deep Q network. The loss function at the moment, so as to use the gradient descent method to adjust the parameters of the main Q network To update: (20) (21) In formula (20)-formula (21), is the target Q value of the main Q network in the dual deep Q network at time t; and They represent the state of the main Q network in the dual deep Q network at time t. and actions The Q value of the output and the state of the target Q network at time t+1 and actions The Q value of the output; For expectations; Step 3.7: Use the soft update method shown in Equation (22)-Equation (23) to adjust the parameters of the target Critic1 network , parameters of the target Critic2 network and the parameters of the main Q network To update: (22) (23) In formula (22)-formula (23), Indicates assignment, is the smoothing factor; Step 3.8: Train the joint learning agent according to the process of steps 3.3 to 3.7 to obtain the trained joint learning agent.
5. The adaptive operation method of a power distribution system based on joint learning and considering differentiated needs according to claim 1, characterized in that: The adaptive switching bar in step 5 is constructed according to the following steps: Step 5.1: Use equations (24) and (25) to obtain the voltage regulation margin of node i at time t and R-SOP reactive flexibility : (24) (25) In formula (24)-formula (25), and Both are indicator functions. When the corresponding conditions are met, the value of the indicator function is 1, otherwise, the value of the indicator function is 0. represents the voltage value on node i at time t, and Respectively represent the upper and lower limit values of the voltage safety range, represents the active power transmitted by R-SOP on node i at time t, represents the reactive power provided or absorbed by the R-SOP on node i at time t, represents the R-SOP capacity of node i at time t; Step 5.2: When and When , it means that the adaptive switching condition of loss reduction control P is reached, and the reward function at time t under loss reduction control P is constructed using equations (26) and (27) ; (26) (27) In formula (26)-formula (27), Indicates that the voltage of node i exceeds the limit at time t; represents the set of all branches; Indicates the total number of nodes; represents the square of the current on the branch ij between nodes i and j at time t, Represents the resistance on branch ij; when and When , it means that the adaptive switching condition of voltage optimization U is achieved, and the reward function at time t under voltage optimization U is constructed using equations (28) and (29) ; (28) (29) In formula (28)-formula (29), Indicates the voltage deviation degree of node i at time t; when and When , it means that the adaptive switching condition of demand response D is achieved, and the reward function at time t under demand response D is constructed using equations (28) and (29): .
6. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store a program that supports a processor to execute the adaptive operation method according to any one of claims 1 to 5, and the processor is configured to execute the program stored in the memory.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the adaptive operation method according to any one of claims 1 to 5 are executed.