Power control method and apparatus, computer device, and storage medium

By employing reinforcement learning algorithms in flexible interconnected distribution networks, the target active and reactive power can be autonomously learned and determined, thus solving the problem of low efficiency in centralized control and achieving rapid response and voltage distribution optimization.

CN114465237BActive Publication Date: 2026-01-23SHENZHEN POWER SUPPLY BUREAU
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111380947.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-20
Publication Date
2026-01-23
Estimated Expiration
2041-11-20

AI Technical Summary

Technical Problem

The existing centralized control methods for flexible interconnected distribution networks involve large amounts of computation and rely on information and communication networks, resulting in low computational efficiency and difficulty in responding quickly to sudden changes in system state.

Method used

By employing reinforcement learning algorithms, the target active power is determined based on the net active power of the converters and the active power of the flexible interconnection devices in the flexible interconnection distribution network system. Combined with the rated capacity of the converters and the reactive power ratio parameters, the power flow distribution and voltage fluctuations are adjusted in real time.

Benefits of technology

It reduces computation time, improves the system's adaptability to sudden changes in operating conditions, reduces network losses, and improves voltage distribution in the distribution network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114465237B_ABST
    Figure CN114465237B_ABST
Patent Text Reader

Abstract

The application relates to a power control method and device, a computer device and a storage medium. An enhanced learning algorithm is adopted to determine target active power of a flexible interconnection device according to net active power of a converter in a flexible interconnection power distribution network system and active power of the flexible interconnection device, the target active power is used for adjusting power flow distribution of the flexible interconnection power distribution network system, maximum reactive power of the converter is determined according to the target active power and rated capacity of the converter, target reactive power of the flexible interconnection device is determined according to the maximum reactive power of the converter and a preset reactive power proportion parameter, and the target reactive power is used for adjusting voltage fluctuation of the flexible interconnection power distribution network system. The application has less system operation information, greatly reduces the calculation processing time, improves the adaptability of the system to operation state mutation, can reduce network loss while improving voltage distribution of the power distribution network, and provides an efficient solution for real-time optimization operation of the flexible interconnection power distribution network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of power system technology, and in particular to a power control method, apparatus, computer equipment, and storage medium. Background Technology

[0002] With the rapid development of power electronic devices and control technologies, flexible interconnection devices (FIDs) based on fully controllable power electronics technology and oriented towards the distribution network level are attracting widespread attention. By flexibly interconnecting traditional distribution networks, FIDs can effectively improve the system's flexibility and controllability, achieving goals such as reducing network losses, improving voltage levels, and promoting the consumption of renewable energy. They are an important foundation for the future development of smart distribution networks.

[0003] Currently, the optimized operation of flexible interconnected distribution networks is mainly achieved through centralized control and local control. Among them, centralized control requires the measurement, collection, and processing of global system operation data, which is time-consuming, computationally intensive, and highly dependent on information and communication networks. Summary of the Invention

[0004] Therefore, it is necessary to provide a power control method, apparatus, computer equipment, and storage medium that can improve computing efficiency in response to the above-mentioned technical problems.

[0005] A power control method, the method comprising:

[0006] A reinforcement learning algorithm is used to determine the target active power of the flexible interconnection device based on the net active power of the converter and the active power of the flexible interconnection device in the flexible interconnection distribution network system; the target active power is used to adjust the power flow distribution of the flexible interconnection distribution network system.

[0007] The maximum reactive power of the converter is determined based on the target active power and the rated capacity of the converter.

[0008] The target reactive power of the flexible interconnection device is determined based on the maximum reactive power of the converter and the preset reactive power ratio parameter; the target reactive power is used to regulate the voltage fluctuation of the flexible interconnection distribution network system.

[0009] In one embodiment, the step of employing a reinforcement learning algorithm to determine the target active power of the flexible interconnection device based on the net active power of the converter and the active power of the flexible interconnection device in the flexible interconnection distribution network system includes:

[0010] The state set is determined based on the net active power of the converter, and the action set is determined based on the active power of the flexible interconnect device.

[0011] The target active power is obtained by using a reinforcement learning algorithm based on the state set and the action set.

[0012] In one embodiment, the step of employing a reinforcement learning algorithm to obtain the target active power based on the state set and the action set includes:

[0013] Substitute the state set and the action set into a preset Q-set function to solve for the maximum value;

[0014] The action corresponding to the maximum function value of the Q-set function in the action set is determined as the target active power.

[0015] In one embodiment, the Q-set function is a function that includes the correspondence between state variables, action variables, reward value variables, new environmental state variables after the action is performed, and probability variables. The reward value variables are determined by the lossy power of the flexible interconnected distribution network system, and the probability variables are variables corresponding to the probability of the state transitioning from the current state to the new environmental state.

[0016] In one embodiment, the method further includes:

[0017] Obtain the port voltage of the flexible interconnect device;

[0018] If the port voltage is outside the preset standard voltage range, the reactive power ratio parameter is determined based on the standard maximum voltage, standard minimum voltage of the standard voltage range, the system maximum voltage, system minimum voltage of the flexible interconnected distribution network system, and the port voltage.

[0019] In one embodiment, the method further includes:

[0020] If the reactive power ratio parameter is greater than zero, then the flexible interconnection device is controlled to inject reactive power into the flexible interconnection distribution network system according to the target reactive power.

[0021] If the reactive power ratio parameter is less than zero, the flexible interconnection device is controlled to absorb reactive power from the flexible interconnection distribution network system according to the target reactive power.

[0022] In one embodiment, the flexible interconnect device satisfies at least one of an active power transmission condition, a capacity constraint condition, and a port voltage constraint condition.

[0023] A power control device, the device comprising:

[0024] The first determining module is used to determine the target active power of the flexible interconnection device based on the net active power of the converter and the active power of the flexible interconnection device in the flexible interconnection distribution network system using a reinforcement learning algorithm; the target active power is used to adjust the power flow distribution of the flexible interconnection distribution network system.

[0025] The second determining module is used to determine the maximum reactive power of the converter based on the target active power and the rated capacity of the converter.

[0026] The third determining module is used to determine the target reactive power of the flexible interconnection device based on the maximum reactive power of the converter and the preset reactive power ratio parameter; the target reactive power is used to regulate the voltage fluctuation of the flexible interconnection distribution network system.

[0027] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program performing the following steps:

[0028] A reinforcement learning algorithm is used to determine the target active power of the flexible interconnection device based on the net active power of the converter and the active power of the flexible interconnection device in the flexible interconnection distribution network system; the target active power is used to adjust the power flow distribution of the flexible interconnection distribution network system.

[0029] The maximum reactive power of the converter is determined based on the target active power and the rated capacity of the converter.

[0030] The target reactive power of the flexible interconnection device is determined based on the maximum reactive power of the converter and the preset reactive power ratio parameter; the target reactive power is used to regulate the voltage fluctuation of the flexible interconnection distribution network system.

[0031] A computer-readable storage medium having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0032] A reinforcement learning algorithm is used to determine the target active power of the flexible interconnection device based on the net active power of the converter and the active power of the flexible interconnection device in the flexible interconnection distribution network system; the target active power is used to adjust the power flow distribution of the flexible interconnection distribution network system.

[0033] The maximum reactive power of the converter is determined based on the target active power and the rated capacity of the converter.

[0034] The target reactive power of the flexible interconnection device is determined based on the maximum reactive power of the converter and the preset reactive power ratio parameter; the target reactive power is used to regulate the voltage fluctuation of the flexible interconnection distribution network system.

[0035] The aforementioned power control method, device, computer equipment, and storage medium employ reinforcement learning algorithms to determine the target active power of the flexible interconnected device based on the net active power of the converter and the active power of the flexible interconnected device in the flexible interconnected distribution network system. This target active power is used to regulate the power flow distribution of the flexible interconnected distribution network system. Based on the target active power and the rated capacity of the converter, the maximum reactive power of the converter is determined. Based on the maximum reactive power of the converter and a preset reactive power ratio parameter, the target reactive power of the flexible interconnected device is determined. This target reactive power is used to regulate voltage fluctuations in the flexible interconnected distribution network system. Compared to existing control strategies, this application requires less system operation information, eliminates the need for global system coverage via a communication network, and continuously updates the control strategy through the autonomous learning of the learning agent. This significantly reduces computation time, improves the system's adaptability to sudden changes in operating conditions, and effectively reduces network losses while improving the voltage distribution of the distribution network, providing an efficient solution for real-time optimized operation of flexible interconnected distribution networks. Attached Figure Description

[0036] Figure 1.1 This is a diagram illustrating the application environment of the power control method in one embodiment;

[0037] Figure 1.2 This is a diagram illustrating the application environment of the power control method in another embodiment;

[0038] Figure 2 This is a flowchart illustrating a power control method in one embodiment;

[0039] Figure 3 This is a schematic diagram of reinforcement learning theory in one embodiment;

[0040] Figure 4 This is a flowchart illustrating the power control method in another embodiment;

[0041] Figure 5 This is a flowchart illustrating the power control method in another embodiment;

[0042] Figure 6 This is a flowchart illustrating the power control method in another embodiment;

[0043] Figure 7 This is a flowchart illustrating the power control method in another embodiment;

[0044] Figure 8 This is a structural block diagram of the power control device in one embodiment;

[0045] Figure 9 This is a structural block diagram of the power control device in one embodiment;

[0046] Figure 10 This is a structural block diagram of the power control device in one embodiment;

[0047] Figure 11 This is a structural block diagram of the power control device in one embodiment;

[0048] Figure 12 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0049] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0050] The power control method provided in this application can be applied to, for example... Figure 1.1 and 1.2 The application environment is shown. The flexible interconnected power distribution system 1 includes a converter 2 and a flexible interconnection device 3, with the flexible interconnection device 3 connected to points m and n in the system. Sensors 4 collect operating parameters of the converter 2 and the flexible interconnection device 3, and input these parameters to a server 5. Using a reinforcement learning algorithm, the target active power of the flexible interconnection device is determined based on the net active power of the converter and the active power of the flexible interconnection device in the flexible interconnected power distribution system. Based on the target active power and the rated capacity of the converter, the maximum reactive power of the converter is determined. Finally, based on the maximum reactive power of the converter and a preset reactive power ratio parameter, the target reactive power of the flexible interconnection device is determined, thereby regulating the power flow distribution and voltage fluctuations of the flexible interconnected power distribution system.

[0051] In one embodiment, such as Figure 3 As shown, a power control method is provided, which is applied to... Figure 1.2 Taking the server in the example, the following steps are included:

[0052] S201 employs a reinforcement learning algorithm to determine the target active power of the flexible interconnection device based on the net active power of the converter and the active power of the flexible interconnection device in the flexible interconnection distribution network system; the target active power is used to adjust the power flow distribution of the flexible interconnection distribution network system.

[0053] The diagram of reinforcement learning theory is shown below. Figure 3As shown, this involves two main entities: a learning agent or decision-making agent, and an environment. The learning agent performs an action *a* based on the currently perceived environmental state *s*, causing a change in the environmental state. The environment then provides feedback to the learning agent regarding the next state following the action and returns an immediate reward value *r*, used to evaluate the effectiveness of the action and form a new action policy. Subsequently, the learning agent performs another action based on the new action policy and the new environmental state. Through repeated interactions with the environment, the learning agent learns the optimal action policy for each environmental state, selecting actions based on maximizing the cumulative expected reward.

[0054] In this embodiment, voltage and current sensors can be used to collect voltage and current signals from the converter and flexible interconnect device, thereby obtaining the net active power of the converter and the active power of the flexible interconnect device. The net active power of the converter and the active power of the flexible interconnect device are then substituted into the reinforcement learning algorithm. In reinforcement learning, the environment perceives different net active power values ​​of the converter, representing different environmental states. For each net active power value, the learning agent finds the optimal active power value of the flexible interconnect device from the corresponding net active power value, which is then used as the target active power value.

[0055] S202, determine the maximum reactive power of the converter based on the target active power and the rated capacity of the converter.

[0056] In this embodiment, to address the frequent voltage fluctuations in the flexible interconnected distribution network system caused by renewable energy generation, the reactive power output of the flexible interconnected device is adjusted in real time based on the local information of each port, thereby achieving rapid compensation for the reactive power of the flexible interconnected distribution network system. In the reactive power control of the flexible interconnected device, the real-time power output is represented by the maximum reactive power that the converter can provide. Based on the above steps, the target active power of the flexible interconnected device can be determined. Furthermore, the maximum reactive power of the converter can be determined based on the target active power and the rated capacity of the converter. The maximum reactive power can be expressed by the following relationship (1):

[0057]

[0058] In the above formula (1): The maximum reactive power of the converter, S FID,k P is the rated capacity of the converter on this side. FID,k 'Target active power for flexible interconnected devices.'

[0059] S203, based on the maximum reactive power of the converter and the preset reactive power ratio parameters, determines the target reactive power of the flexible interconnection device; the target reactive power is used to regulate the voltage fluctuation of the flexible interconnection distribution network system.

[0060] In this embodiment, the maximum reactive power of the converter can be obtained according to the above steps. When the reactive power ratio parameter is known, the target reactive power of the flexible interconnection device can be determined by the following relationship (2), thereby adjusting the voltage fluctuation problem of the flexible interconnection distribution network system.

[0061]

[0062] In equation (2) above, Q FID,k The reactive power of the converter. This is the reactive power ratio parameter.

[0063] In this embodiment, when the voltage of the flexible interconnected distribution network system fluctuates, the target reactive power of the flexible interconnection device can be determined based on the maximum reactive power of the converter and a preset reactive power ratio parameter. For example, when the port voltage of the flexible interconnection device increases to a certain threshold, the flexible interconnection device can inject the corresponding target reactive power into the flexible interconnected distribution network system; or, when the port voltage of the flexible interconnection device decreases to a certain threshold, the flexible interconnection device can absorb the corresponding target reactive power from the flexible interconnected distribution network system, so that the voltage of the flexible interconnected distribution network system is within a specified range.

[0064] In the aforementioned power control method, a reinforcement learning algorithm is employed to determine the target active power of the flexible interconnected device based on the net active power of the converter and the active power of the flexible interconnection device in the flexible interconnected distribution network system. This target active power is used to regulate the power flow distribution of the flexible interconnected distribution network system. Based on the target active power and the rated capacity of the converter, the maximum reactive power of the converter is determined. Based on the maximum reactive power of the converter and a preset reactive power ratio parameter, the target reactive power of the flexible interconnection device is determined. This target reactive power is used to regulate voltage fluctuations in the flexible interconnected distribution network system. Compared to existing control strategies, this application requires less system operation information, eliminates the need for global system coverage via a communication network, and continuously updates the control strategy through the autonomous learning of the learning agent. This significantly reduces computation time, improves the system's adaptability to sudden changes in operating conditions, and effectively reduces network losses while improving the voltage distribution of the distribution network, providing an efficient solution for the real-time optimized operation of flexible interconnected distribution networks.

[0065] In the above Figure 2 The embodiments mainly introduce the power control method. The following focuses on how to determine the target active power, such as... Figure 4 As shown, it includes the following steps:

[0066] S301 determines the state set based on the net active power of the converter and the action set based on the active power of the flexible interconnect device.

[0067] In this embodiment, before the agent learns, the net active power needs to be divided into a series of continuous intervals within the feasible region, serving as the environmental state set; the action strategy set consists of the active power output of the flexible interconnect device. For ease of calculation, continuous integers of active power within its rated capacity are used for output. For the environmental state set consisting of M intervals and the action set consisting of N active power outputs, all the values ​​corresponding to the "state-action" can be used to establish an M*N matrix.

[0068] Furthermore, regarding the active power control of flexible interconnection devices, the net active power at the distribution network transformer or converter reflects the operating status of the flexible interconnection distribution network system. When the net active power is positive, it means that the load demand in the flexible interconnection distribution network system is greater than the output of distributed generation, and power needs to be purchased from the upstream grid. When the net active power is negative, it means that the output of distributed generation in the flexible interconnection distribution network system is greater than the load demand, and the flexible interconnection distribution network needs to feed power to the upstream grid.

[0069] S302 uses a reinforcement learning algorithm to obtain the target active power based on the state set and action set.

[0070] In this embodiment, the net active power of the converter and the active power of the flexible interconnect device are substituted into the reinforcement learning algorithm to obtain the target active power based on the calculation. For example, state s1 is 101W-120W, corresponding to action sets a11, a12, a13, a14, a15, and a16; state s2 is 121W-140W, corresponding to action sets a21, a22, a23, a24, a25, and a26; and state s3 is 141W-160W, corresponding to action sets a31, a32, a33, a34, a35, and a36. For state set s1, a11, a12, a13, a14, a15, a16, and state set s1 are substituted into the reinforcement learning algorithm to find the action that maximizes the expected benefit of reinforcement learning. If the result of the reinforcement learning algorithm corresponding to a13 is optimal, then a13 is the target active power in state s1. When the load in a flexible interconnected distribution network system increases, the net active power of the converter will also increase, and the state will change accordingly. If the state remains within the range of 101W-120W in state s1 after the change, it is not necessary to reacquire the target active power. If the load increase causes the state to change from s1 to s2 or s3, then it is necessary to reacquire the target active power under the new environmental state. The same logic applies when the load in the flexible interconnected distribution network system decreases.

[0071] In this embodiment, the net active power of the converter is used to determine the state set, and the active power of the flexible interconnection device is used to determine the action set. A reinforcement learning algorithm is used to obtain the target active power based on the state set and action set. This method uses a learning agent to learn autonomously and continuously update the control strategy to find the optimal strategy, thereby effectively improving the power flow distribution of the flexible interconnection distribution network system.

[0072] In the above Figure 4 The embodiment describes the process of determining the target active power. The following mainly explains the specific steps for obtaining the target active power using the Q-set function, such as... Figure 5 As shown, it includes the following steps:

[0073] S401, substitute the state set and action set into the preset Q-set function to solve for the maximum value.

[0074] The Q-set function is a function that includes the correspondence between state variables, action variables, reward value variables, new environmental state variables after the action is executed, and probability variables. The reward value variable is determined by the lossy power of the flexible interconnected distribution network system, and the probability variable is the variable corresponding to the probability of the state transitioning from the current state to the new environmental state. The state variable represents the net active power of the converter, the action variable represents the active power of the flexible interconnection device, the reward value variable represents the reward value obtained by the converter's net active power transitioning from the current net active power to the new net active power after executing the active power action of the flexible interconnection device, and the probability variable is the probability of the converter's net active power transitioning from the current net active power to the new net active power after executing the active power action of the flexible interconnection device.

[0075] In this embodiment, the Q-set function is based on the Discrete Time Markov Dispatch Process. The algorithm examines the Q-set functions corresponding to a series of "state-action" pairs. For the execution of action a in environment state s, the Q-set function can be represented by the following relation (3):

[0076] Q*(s,a)=r(s,s',a)+γ∑ s'∈S P ss' (s'|s,a)max a∈A Q*(s',a) (3);

[0077] In the formula: s' represents the new environmental state after the action is performed, and P ss'(s'|s,a) represents the probability that the state transitions from s to s' after performing action a, and r(s,s',a) is its corresponding immediate reward value; S and A represent the environmental state set and action set, respectively; γ∈[0,1] is the decay rate, γ=0 means that only the immediate reward is considered without considering the long-term reward, and γ=1 means that the long-term reward and the immediate reward are equally important. Based on the reward value, the Q-set function is updated, and the update results are shown in relation (4) and relation (5):

[0078] Q it+1 (s it ,a it )=Q it (s it ,a it )+α[r(s it ,s it+1 ,a it )+γmax a'∈A Q it (s it+1 ,a')-Q it (s it ,a it (4);

[0079]

[0080] In the above formula: it represents the number of iterations; α is the learning factor, the larger the value of α, the faster the convergence speed of the learning algorithm; a' represents the environment state determined by s. it Transition to s it 'All feasible action strategies at the time; In the it-th iteration, divide (s) it a it All "state-action" pairs other than )

[0081] In this embodiment, the return value variable is determined by the lossy power of the flexible interconnected distribution network system. The active power control of the flexible interconnection device aims to reduce network losses. The immediate return value reflects the degree of network loss reduction. Therefore, the immediate return value of the Q-set function can be expressed by the following relations (6) and (7):

[0082] r = (LP loss ) / L (6);

[0083]

[0084] In the above formula: P loss For the active power loss of the entire flexible interconnected distribution network system, N branch Total number of branch roads, I n Let R be the current in the nth branch. nLet L be its resistance value. Here, L is a given large value (much larger than the system network loss). It can be seen that the lower the network loss corresponding to the action strategy, the larger the immediate reward value.

[0085] S402, the action corresponding to the maximum function value of the Q set function in the action set is determined as the target active power.

[0086] In this embodiment, during the learning process of the agent interacting with the environment, the Q-set function will be continuously updated until the result stabilizes. Subsequent action selection will follow a greedy strategy, as shown in relation (8), selecting the action with the largest Q-value corresponding to each state, i.e.:

[0087] a*(s) = argmax a∈A Q it (s,a) (8);

[0088] In the above formula, a*(s) represents the action selected under the current environmental state s that maximizes the value of the Q-set function, i.e., the target active power.

[0089] In this embodiment, the state set and action set are substituted into a preset Q-set function to solve for the maximum value. The action in the action set that makes the Q-set function obtain the maximum function value is determined as the target active power. This method uses the Q-set function to determine the target active power, requires fewer parameters, greatly reduces the calculation processing time, and can quickly adjust the active power of the flexible interconnection device online according to the real-time operating status of the distribution network system.

[0090] The above Figure 5 The example primarily introduces the process of determining the target active power using the Q-set function. The following section focuses on how to calculate the reactive power ratio parameter, such as... Figure 6 As shown, it includes the following steps:

[0091] S501, obtain the port voltage of the flexible interconnect device.

[0092] Here, the port voltage is the voltage at the access point of the flexible interconnect device, as described above. Figure 1.1 As shown, points m and n are the access points of the flexible interconnection device in the distribution network system. The voltage at points m and n is collected and represents the port voltage of the flexible interconnection device.

[0093] In this embodiment, a voltage sensor can be used to acquire the port voltage of the flexible interconnect device. Optionally, the voltage sensor can be a resistive voltage divider, a capacitive voltage divider, an electromagnetic voltage transformer, a capacitive voltage transformer, a Hall voltage sensor, etc. Alternatively, a digital voltmeter can be used to acquire the port voltage of the flexible interconnect device.

[0094] S502 If the port voltage is outside the preset standard voltage range, the reactive power ratio parameter is determined based on the standard maximum voltage, standard minimum voltage of the standard voltage range, the system maximum voltage, system minimum voltage of the flexible interconnected distribution network system, and the port voltage.

[0095] The preset standard voltage range is within the maximum and minimum voltage range of the flexible interconnected distribution network system.

[0096] In this embodiment, the reactive power ratio parameter can be determined according to the following relationship (9).

[0097]

[0098] In equation (9) above, V FID,k V is the port voltage of the flexible interconnect device. min For the minimum system voltage of the flexible interconnected distribution network system, V max V is the maximum system voltage of the flexible interconnected distribution network system. min,d Standard minimum voltage, V max,d This is the standard maximum voltage.

[0099] For example, if the system minimum voltage is 20V, the system maximum voltage is 100V, the standard minimum voltage is 50V, the standard maximum voltage is 60V, and the obtained port voltage is 80V, it can be seen that the port voltage is between the standard maximum voltage and the system maximum voltage. Substituting the system minimum voltage of 20V, the system maximum voltage of 100V, the standard minimum voltage of 50V, the standard maximum voltage of 60V, and the port voltage of 80V into formula (9), the reactive power ratio parameter is -0.5. If the system minimum voltage is 20V, the system maximum voltage is 100V, the standard minimum voltage of 50V, the standard maximum voltage of 60V, and the obtained port voltage is 30V, it can be seen that the port voltage is between the standard minimum voltage and the system minimum voltage. Substituting the system minimum voltage of 20V, the system maximum voltage of 100V, the standard minimum voltage of 50V, the standard maximum voltage of 60V, and the port voltage of 30V into formula (9), the reactive power ratio parameter is 0.7.

[0100] In this embodiment, the reactive power ratio parameter is determined by acquiring the port voltage of the flexible interconnection device and further considering the standard maximum voltage, standard minimum voltage within the standard voltage range, the system maximum voltage, system minimum voltage of the flexible interconnection distribution network system, and the port voltage. This method calculates the reactive power ratio parameter based on various scenarios and takes into account variations in port voltage, resulting in more accurate calculations and laying the foundation for subsequent calculations of the target reactive power.

[0101] The above Figure 6The example primarily describes the process of calculating the reactive power ratio parameter. The following section focuses on how to adjust voltage fluctuations based on the target reactive power. For instance, adjusting voltage fluctuations based on the target reactive power can include the following two methods:

[0102] Firstly, if the reactive power ratio parameter is greater than zero, the flexible interconnection device is controlled to inject reactive power into the flexible interconnection distribution network system according to the target reactive power.

[0103] In this embodiment, if the reactive power ratio parameter obtained when the port voltage is 30V is 0.7, the reactive power ratio parameter is greater than zero. The corresponding target reactive power can be calculated using the above relationship (2) based on the maximum reactive power of the converter and the reactive power ratio parameter. The flexible interconnection device is controlled to inject reactive power into the flexible interconnection distribution network system according to the target reactive power, so that the port voltage of the flexible interconnection device is within the standard voltage range.

[0104] Secondly, if the reactive power ratio parameter is less than zero, the flexible interconnection device is controlled to absorb reactive power from the flexible interconnection distribution network system according to the target reactive power.

[0105] In this embodiment, if the reactive power ratio parameter obtained when the port voltage is 80V is -0.5, the reactive power ratio parameter is less than zero. The corresponding target reactive power can be calculated using the above relationship (2) based on the maximum reactive power of the converter and the reactive power ratio parameter. The flexible interconnection device is controlled to absorb reactive power from the flexible interconnection distribution network system according to the target reactive power, so that the port voltage of the flexible interconnection device is within the standard voltage range.

[0106] In this embodiment, the reactive power ratio parameter is determined based on the voltage of each port of the flexible interconnection device. The reactive power ratio parameter is positive or negative to control the absorption or injection of reactive power by the flexible interconnection device. This method uses local control to adjust the reactive power output of the flexible interconnection device in real time, so as to realize the rapid compensation of reactive power in the distribution network and thus alleviate the voltage fluctuation problem of the flexible interconnection distribution system.

[0107] In one embodiment, the flexible interconnect device satisfies at least one of an active power transmission condition, a capacity constraint condition, and a port voltage constraint condition.

[0108] In one embodiment, the flexible interconnect device satisfies at least one of an active power transmission condition, a capacity constraint condition, and a port voltage constraint condition. As described above... Figure 1.1 As shown, when a flexible interconnect device is used to connect node m and node n in a system, its operation must satisfy the following active power transmission constraints:

[0109]

[0110] In equation (10) above: PFID,k This represents the active power exchanged between the flexible interconnect device and node k. This refers to the internal losses of the converter on this side, where, It can be calculated from the transmission power and loss coefficient ρ of the flexible interconnect device, as shown in equation (11):

[0111]

[0112] In equation (11) above: Q FID,k This is for the reactive power exchange between the flexible interconnect device and node k.

[0113] Furthermore, the capacity constraint of the flexible interconnect device can be expressed by equation (3):

[0114]

[0115] In equation (12) above, S FID,k This is the rated capacity of the converter on this side.

[0116] In this embodiment, the voltage constraint conditions of each port of the flexible interconnect device can be expressed by equation (13):

[0117] V min ≤|V FID,k |≤V max (13);

[0118] In equation (13) above, V FID,k V is the voltage at node k. min V max These are the upper and lower limits of the allowable voltage fluctuation range in a flexible interconnected distribution network system.

[0119] Furthermore, such as Figure 7 As shown, the power control method further includes the following steps:

[0120] S601, the state set is determined by the net active power of the converter, and the action set is determined by the active power of the flexible interconnection device.

[0121] S602 uses a reinforcement learning algorithm to obtain the target active power based on the state set and action set;

[0122] S603, determine the maximum reactive power of the converter based on the target active power and the rated capacity of the converter;

[0123] S604, Obtain the port voltage of the flexible interconnect device;

[0124] S605 If the port voltage is outside the preset standard voltage range, the reactive power ratio parameter is determined based on the standard maximum voltage, standard minimum voltage of the standard voltage range, the system maximum voltage, system minimum voltage of the flexible interconnected distribution network system, and the port voltage.

[0125] S606, based on the maximum reactive power of the converter and the preset reactive power ratio parameters, determines the target reactive power of the flexible interconnection device; the target reactive power is used to regulate voltage fluctuations in the flexible interconnection distribution network system.

[0126] S607, if the reactive power ratio parameter is greater than zero, the flexible interconnection device is controlled to inject reactive power into the flexible interconnection distribution network system according to the target reactive power.

[0127] S608 If the reactive power ratio parameter is less than zero, the flexible interconnection device is controlled to absorb reactive power from the flexible interconnection distribution network system according to the target reactive power.

[0128] This application provides a power control method, comprising employing a reinforcement learning algorithm to determine the target active power of the flexible interconnected device based on the net active power of the converter and the active power of the flexible interconnection device in the flexible interconnected distribution network system. The target active power is used to regulate the power flow distribution of the flexible interconnected distribution network system. Based on the target active power and the rated capacity of the converter, the maximum reactive power of the converter is determined. Based on the maximum reactive power of the converter and a preset reactive power ratio parameter, the target reactive power of the flexible interconnection device is determined. The target reactive power is used to regulate voltage fluctuations in the flexible interconnected distribution network system. Compared with existing control strategies, this application requires less system operation information, eliminates the need for global system coverage by a communication network, and continuously updates the control strategy through the autonomous learning of the learning agent, significantly reducing computation time and improving the system's adaptability to sudden changes in operating state. It can effectively reduce network losses while improving the voltage distribution of the distribution network, providing an efficient solution for real-time optimized operation of flexible interconnected distribution networks.

[0129] It should be understood that, although Figure 2-7 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 2-7 At least some of the steps in the process may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but may be executed at different times. The execution order of these steps or stages is not necessarily sequential, but may be executed in turn or alternately with other steps or at least some of the steps or stages in other steps.

[0130] In one embodiment, such as Figure 8 As shown, a power control device is provided, comprising: a first determining module 11, a second determining module 12, and a third determining module 13, wherein:

[0131] The first determining module 11 is used to determine the target active power of the flexible interconnection device based on the net active power of the converter and the active power of the flexible interconnection device in the flexible interconnection distribution network system using a reinforcement learning algorithm; the target active power is used to adjust the power flow distribution of the flexible interconnection distribution network system.

[0132] The second determining module 12 is used to determine the maximum reactive power of the converter based on the target active power and the rated capacity of the converter.

[0133] The third determining module 13 is used to determine the target reactive power of the flexible interconnection device based on the maximum reactive power of the converter and the preset reactive power ratio parameters; the target reactive power is used to regulate the voltage fluctuation of the flexible interconnection distribution network system.

[0134] In one embodiment, such as Figure 9 As shown, the first determining module 11 includes:

[0135] The first determining unit 111 is used to determine a state set based on the net active power of the converter and to determine an action set based on the active power of the flexible interconnect device.

[0136] The acquisition unit 112 is used to acquire the target active power based on the state set and action set using a reinforcement learning algorithm.

[0137] In one embodiment, the acquisition unit 112 is used to substitute the state set and the action set into a preset Q-set function to solve for the maximum value; and to determine the action corresponding to the maximum function value of the Q-set function in the action set as the target active power.

[0138] In one embodiment, the Q-set function is a function that includes the correspondence between state variables, action variables, reward value variables, new environmental state variables after the action is performed, and probability variables. The reward value variables are determined by the lossy power of the flexible interconnected distribution network system, and the probability variables are the variables corresponding to the probability of the state transitioning from the current state to the new environmental state.

[0139] In one embodiment, such as Figure 10 As shown, a power control device is provided, which further includes:

[0140] Acquisition module 14 is used to acquire the port voltage of the flexible interconnect device;

[0141] The fourth determining module 15 is used to determine the reactive power ratio parameter based on the standard maximum voltage, standard minimum voltage, system maximum voltage, system minimum voltage of the flexible interconnected distribution network system, and port voltage if the port voltage is outside the preset standard voltage range.

[0142] In one embodiment, such as Figure 11 As shown, a power control device is provided, which further includes:

[0143] The injection module 16 is used to control the flexible interconnection device to inject reactive power into the flexible interconnection distribution network system according to the target reactive power if the reactive power ratio parameter is greater than zero.

[0144] The absorption module 17 is used to control the flexible interconnection device to absorb reactive power from the flexible interconnection distribution network system according to the target reactive power if the reactive power ratio parameter is less than zero.

[0145] In one embodiment, the flexible interconnect device satisfies at least one of an active power transmission condition, a capacity constraint condition, and a port voltage constraint condition.

[0146] For specific limitations regarding the power control device, please refer to the limitations on the power control method above, which will not be repeated here. Each module in the aforementioned power control device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of the processor in the computer device, or stored in software in the memory of the computer device, so that the processor can call and execute the operations corresponding to each module.

[0147] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 12 As shown, the computer device includes a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a power control method. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the computer device casing, or an external keyboard, touchpad, or mouse.

[0148] Those skilled in the art will understand that Figure 12 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0149] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0150] A reinforcement learning algorithm is used to determine the target active power of the flexible interconnection device based on the net active power of the converter and the active power of the flexible interconnection device in the flexible interconnection distribution network system; the target active power is used to adjust the power flow distribution of the flexible interconnection distribution network system.

[0151] Determine the maximum reactive power of the converter based on the target active power and the rated capacity of the converter;

[0152] The target reactive power of the flexible interconnection device is determined based on the maximum reactive power of the converter and the preset reactive power ratio parameters; the target reactive power is used to regulate the voltage fluctuation of the flexible interconnection distribution network system.

[0153] In one embodiment, the processor, when executing a computer program, also performs the following steps:

[0154] The state set is determined by the net active power of the converter, and the action set is determined by the active power of the flexible interconnection device.

[0155] A reinforcement learning algorithm is used to obtain the target active power based on the state set and action set.

[0156] In one embodiment, the processor, when executing a computer program, also performs the following steps:

[0157] Substitute the state set and action set into the preset Q-set function to solve for the maximum value;

[0158] The action corresponding to the maximum function value of the Q-set function in the action set is determined as the target active power.

[0159] In one embodiment, the Q-set function is a function that includes the correspondence between state variables, action variables, reward value variables, new environmental state variables after the action is performed, and probability variables. The reward value variables are determined by the lossy power of the flexible interconnected distribution network system, and the probability variables are the variables corresponding to the probability of the state transitioning from the current state to the new environmental state.

[0160] In one embodiment, the processor, when executing a computer program, also performs the following steps:

[0161] Obtain the port voltage of the flexible interconnect device;

[0162] If the port voltage is outside the preset standard voltage range, the reactive power ratio parameter is determined based on the standard maximum voltage, standard minimum voltage of the standard voltage range, the system maximum voltage, system minimum voltage of the flexible interconnected distribution network system, and the port voltage.

[0163] In one embodiment, the processor, when executing a computer program, also performs the following steps:

[0164] If the reactive power ratio parameter is greater than zero, the flexible interconnection device is controlled to inject reactive power into the flexible interconnection distribution network system according to the target reactive power.

[0165] If the reactive power ratio parameter is less than zero, the flexible interconnection device is controlled to absorb reactive power from the flexible interconnection distribution network system according to the target reactive power.

[0166] In one embodiment, the flexible interconnect device satisfies at least one of an active power transmission condition, a capacity constraint condition, and a port voltage constraint condition.

[0167] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0168] A reinforcement learning algorithm is used to determine the target active power of the flexible interconnection device based on the net active power of the converter and the active power of the flexible interconnection device in the flexible interconnection distribution network system; the target active power is used to adjust the power flow distribution of the flexible interconnection distribution network system.

[0169] Determine the maximum reactive power of the converter based on the target active power and the rated capacity of the converter;

[0170] The target reactive power of the flexible interconnection device is determined based on the maximum reactive power of the converter and the preset reactive power ratio parameters; the target reactive power is used to regulate the voltage fluctuation of the flexible interconnection distribution network system.

[0171] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:

[0172] The state set is determined by the net active power of the converter, and the action set is determined by the active power of the flexible interconnection device.

[0173] A reinforcement learning algorithm is used to obtain the target active power based on the state set and action set.

[0174] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:

[0175] The state set is determined by the net active power of the converter, and the action set is determined by the active power of the flexible interconnection device.

[0176] A reinforcement learning algorithm is used to obtain the target active power based on the state set and action set.

[0177] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:

[0178] Substitute the state set and action set into the preset Q-set function to solve for the maximum value;

[0179] The action corresponding to the maximum function value of the Q-set function in the action set is determined as the target active power.

[0180] In one embodiment, the Q-set function is a function that includes the correspondence between state variables, action variables, reward value variables, new environmental state variables after the action is performed, and probability variables. The reward value variables are determined by the lossy power of the flexible interconnected distribution network system, and the probability variables are the variables corresponding to the probability of the state transitioning from the current state to the new environmental state.

[0181] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:

[0182] Obtain the port voltage of the flexible interconnect device;

[0183] If the port voltage is outside the preset standard voltage range, the reactive power ratio parameter is determined based on the standard maximum voltage, standard minimum voltage of the standard voltage range, the system maximum voltage, system minimum voltage of the flexible interconnected distribution network system, and the port voltage.

[0184] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:

[0185] If the reactive power ratio parameter is greater than zero, the flexible interconnection device is controlled to inject reactive power into the flexible interconnection distribution network system according to the target reactive power.

[0186] If the reactive power ratio parameter is less than zero, the flexible interconnection device is controlled to absorb reactive power from the flexible interconnection distribution network system according to the target reactive power.

[0187] In one embodiment, the flexible interconnect device satisfies at least one of an active power transmission condition, a capacity constraint condition, and a port voltage constraint condition.

[0188] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:

[0189] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0190] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0191] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A power control method, characterized in that, The method includes: A state set is determined using the net active power of the converter in the flexible interconnected distribution network system as the state, and an action set is determined using the active power of the flexible interconnection device as the action. The state set and the action set are substituted into a preset Q-set function to solve for the maximum value. The action corresponding to the maximum function value of the Q-set function in the action set is determined as the target active power. The target active power is used to adjust the power flow distribution of the flexible interconnected distribution network system. The Q-set function is a function that includes the correspondence between state variables, action variables, reward value variables, new environmental state variables after the action is executed, and probability variables. The reward value variable is determined by the lossy power of the flexible interconnected distribution network system, and the probability variable is the variable corresponding to the probability of the state transitioning from the current state to the new environmental state. The maximum reactive power of the converter is determined based on the target active power and the rated capacity of the converter. The target reactive power of the flexible interconnection device is determined based on the maximum reactive power of the converter and the preset reactive power ratio parameter; the target reactive power is used to regulate the voltage fluctuation of the flexible interconnection distribution network system. The method further includes: Obtain the port voltage of the flexible interconnect device; If the port voltage is outside the preset standard voltage range, the reactive power ratio parameter is determined based on the standard maximum voltage, standard minimum voltage of the standard voltage range, the system maximum voltage, system minimum voltage of the flexible interconnected distribution network system, and the port voltage.

2. The method according to claim 1, characterized in that, The method further includes: If the reactive power ratio parameter is greater than zero, then the flexible interconnection device is controlled to inject reactive power into the flexible interconnection distribution network system according to the target reactive power. If the reactive power ratio parameter is less than zero, the flexible interconnection device is controlled to absorb reactive power from the flexible interconnection distribution network system according to the target reactive power.

3. The method according to claim 1 or 2, characterized in that, The flexible interconnect device satisfies at least one of the following conditions: active power transmission condition, capacity constraint condition, and port voltage constraint condition.

4. A power control device, characterized in that, The device includes: The first determining module is used to determine a state set based on the net active power of the converter in the flexible interconnected distribution network system, and to determine an action set based on the active power of the flexible interconnected device. The state set and the action set are substituted into a preset Q-set function to solve for the maximum value. The action corresponding to the maximum function value of the Q-set function in the action set is determined as the target active power. The target active power is used to adjust the power flow distribution of the flexible interconnected distribution network system. The Q-set function is a function that includes the correspondence between state variables, action variables, reward value variables, new environmental state variables after the action is executed, and probability variables. The reward value variables are determined by the lossy power of the flexible interconnected distribution network system, and the probability variables are variables corresponding to the probability of a state transitioning from the current state to a new environmental state. The second determining module is used to determine the maximum reactive power of the converter based on the target active power and the rated capacity of the converter. The third determining module is used to determine the target reactive power of the flexible interconnection device based on the maximum reactive power of the converter and the preset reactive power ratio parameter; the target reactive power is used to regulate the voltage fluctuation of the flexible interconnection distribution network system. The device further includes: The acquisition module is used to acquire the port voltage of the flexible interconnect device; The fourth determining module is used to determine the reactive power ratio parameter based on the standard maximum voltage, standard minimum voltage of the standard voltage range, the system maximum voltage, system minimum voltage of the flexible interconnected distribution network system, and the port voltage if the port voltage is outside the preset standard voltage range.

5. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 3.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Optimal scheduling method of flexible interconnection power distribution network, storage medium and processor

    CN111092429A

  • Direct-current interconnection-based power grid power flow regulation and control method

    CN113206503A