Multi-target electrical automation cooperative control method fusing reinforcement learning

Through the multi-objective electrical automation collaborative control method that integrates reinforcement learning, the problem of difficult to achieve equipment-level and system-level collaborative optimization control in microgrid systems is solved, and the efficient and stable operation of the system is achieved.

CN120103766AActive Publication Date: 2025-06-06SHANXI UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510293911.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-06
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

The prior art is difficult to achieve equipment-level and system-level collaborative optimization control in microgrid systems, and it is impossible to respond to various changes and uncertainties in the system in a timely and precise manner.

Method used

The multi-objective electrical automation collaborative control method is adopted to collect the electrical parameters and operating states of the distribution cabinet in real time, build the self and adjacent state vectors, determine the action vectors, and adjust the actions based on the self and the synergistic target reward function to achieve collaborative optimization between devices.

Benefits of technology

It realizes effective coordination between distribution cabinets in the power system, improves the operating stability and reliability of the system, reduces line losses, improves load balancing, and ensures efficient operation of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120103766A_ABST
    Figure CN120103766A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of electrical automation control, in particular to a reinforcement learning-fused multi-target electrical automation cooperative control method, which comprises the following steps of: acquiring parameters of each power distribution cabinet in a region to construct a self state vector and an adjacent state vector; constructing a preset target strategy according to the line loss and the output current and voltage of each power distribution cabinet, and determining an action vector based on the own state vector, the adjacent state vector and the preset target strategy; according to the action vector and a double-layer reward function, determining an influence direction on each power distribution cabinet after the action vector is executed, wherein the double-layer reward function comprises a self target reward function and a collaborative target reward function; and determining an action adjustment mode according to the influence direction. According to the invention, targets of the power distribution cabinets and system cooperation targets are fully considered, and effective cooperation among the power distribution cabinets in the power system is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of electrical automation control technology, and in particular to a multi-objective electrical automation collaborative control method integrating reinforcement learning. Background Art

[0002] With the rapid development of electrical automation technology, the scale and complexity of electrical automation systems are increasing. In actual operation, electrical automation systems need to meet multiple goals at the same time, such as efficient energy utilization, stable power supply quality, low equipment loss, and good economic benefits. Traditional control methods are usually designed for a single goal, and it is difficult to effectively balance and achieve multiple conflicting goals in a complex and changing operating environment. In addition, traditional control methods often rely on accurate system models and have poor adaptability to system parameter changes, uncertain factors, and real-time dynamic adjustments.

[0003] China Patent Publication No.: CN111200285B, discloses a hybrid coordinated control method for microgrids based on reinforcement learning and multi-agent theory, including: designing a transition voltage layer control strategy based on voltage stratification and a dual energy storage role control strategy: when the energy storage unit works in voltage stabilization mode, the two energy storages work separately; when energy storage assistance is required to continuously absorb power or supplement power, the two energy storage working modes are converted to coordinated charging / discharging. Constructing the action space and state space based on Q Learning: Designing a reinforcement learning control framework based on multi-agents: including designing the basic update rules of state-action pairs and selecting the corresponding value function; Designing the basic action selection mechanism and reward value strategy: including designing the selection strategy adopted by the system in the initial state and the reward value in each state; Designing the reinforcement learning algorithm flow: Based on the above strategies, design a suitable algorithm flow to implement the control strategy.

[0004] It can be seen that the technical solution has the following problems: although there is a dual energy storage role control strategy, in the entire microgrid system, the dynamic adjustment mechanism for other equipment and the overall system may be relatively concentrated in aspects such as the energy storage working mode conversion, which is not comprehensive and flexible enough, and cannot respond to various changes and uncertainties in the system in a timely and accurate manner. Summary of the invention

[0005] To this end, the present invention provides a multi-objective electrical automation collaborative control method integrating reinforcement learning to overcome the problem in the prior art that the electrical system cannot achieve device-level and system-level collaborative optimization control.

[0006] To achieve the above object, the present invention provides a multi-objective electrical automation collaborative control method integrating reinforcement learning, comprising: Step S1, real-time collection of electrical parameters and operating status of each distribution cabinet in the area to construct its own state vector, and construct neighboring state vectors according to the power transmission direction and line impedance between adjacent distribution cabinets, where the types of distribution cabinets include power access cabinets, load distribution cabinets and energy storage control cabinets; Step S2, constructing a preset target strategy according to the output current, output voltage, and line loss of each distribution cabinet, and determining the action vector of each distribution cabinet based on its own state vector, neighboring state vectors, and the preset target strategy; Step S3, determining the impact direction of the action vector on each distribution cabinet after execution based on the action vector, the own target reward function and the collaborative target reward function, wherein the impact direction includes an individual impact direction and a total impact direction; Among them, the self-target reward function is determined according to the type of distribution cabinet, including: the self-reward function of the power access cabinet includes a voltage deviation penalty term and a reactive output constraint term; the self-reward function of the load distribution cabinet includes a load balancing term and a tripping number penalty term; the self-reward function of the energy storage control cabinet includes an SOC health term and a charge and discharge efficiency term; The collaborative objective reward function is determined according to the total line loss, load balancing rate and communication efficiency between the distribution cabinets; Step S4, determining the action adjustment method according to the impact direction.

[0007] Specifically, in step S2, a harmonic target strategy is determined according to the output voltage and the output current, a line loss strategy is determined according to the line loss, and a preset target strategy is determined according to the harmonic target strategy and the line loss strategy.

[0008] Specifically, in step S2, constructing a harmonic target strategy includes: Determine the current harmonic content and the voltage harmonic content based on the output current and output voltage of each distribution cabinet; Determining a current influence coefficient based on the current harmonic content and the current harmonic threshold; Determining a voltage influence coefficient based on the voltage harmonic content and the voltage harmonic threshold; The harmonic target strategy is determined based on the current influence coefficient and the voltage influence coefficient.

[0009] Specifically, in step S2, the action vector corresponding to the harmonic target strategy includes: The power access cabinet injects reverse harmonics to offset the load harmonics; Disconnect the branch switch of the load distribution cabinet with high harmonic load or adjust the three-phase load distribution; The energy storage control cabinet offsets harmonics by charging / discharging energy storage.

[0010] Specifically, in step S2, constructing a line loss strategy includes: Determine the actual resistance of the line between the distribution cabinets based on the initial resistance of the line, the temperature influence coefficient and the aging influence coefficient; The total line loss is determined according to the actual line resistance and the line current.

[0011] Specifically, in step S2, the temperature influence coefficient is determined based on the historical ambient temperature and the historical line resistance, and the aging influence coefficient is determined based on the line usage time and the historical line resistance.

[0012] Specifically, in step S2, the action vector corresponding to the line loss strategy includes: Adjust the reactive power output of the power access cabinet to reduce the line current; Switch the load distribution cabinet switch combination to evenly distribute the load; Adjust the charging / discharging power of the energy storage control cabinet to optimize the power flow distribution.

[0013] Specifically, step S4 includes: Collect the output current of each load distribution cabinet in the area in real time, determine the average load of the load distribution cabinet based on the output current, and determine the load balancing rate based on the average load; Collect the data transmission volume and data reception volume of each power distribution cabinet in the area in real time, and determine the communication efficiency based on the data transmission volume and the data reception volume; The collaborative objective reward function is determined by weighting according to the load balancing rate and the communication efficiency.

[0014] Specifically, in step S5, the action adjustment method is determined, wherein: If the total impact direction is a positive reward, determining the action adjustment amplitude according to the individual impact direction; If the total impact direction is a negative reward, then iterate steps S2 to S5 until the total impact direction is a positive reward or the number of iterations is equal to the preset maximum number of iterations.

[0015] Specifically, in step S5, determining the action adjustment range includes: If the total impact direction is positive reward, and the individual impact direction corresponding to the power distribution cabinet is positive incentive, then the action adjustment amplitude of the action vector of the corresponding power distribution cabinet is increased; If the total impact direction is a positive reward and the individual impact direction corresponding to the distribution cabinet is a negative reward, the action adjustment amplitude of the action vector of the corresponding distribution cabinet is reduced.

[0016] Compared with the prior art, the beneficial effect of the present invention is that the present invention integrates multiple goals such as voltage stability, harmonic control, load balancing, etc. through a reinforcement learning framework to achieve effective coordination between distribution cabinets in the power system, including using its own state vector and neighboring state vectors to jointly represent the state of the device itself and reflect the network topology relationship, so as to achieve strong coupling of the power system. Constructing preset target strategies and determining action vectors based on multiple factors is more in line with actual needs. At the same time, a hierarchical reward mechanism is set up to distinguish between individual target rewards and collaborative target rewards, effectively evaluate the impact of action vectors on the system, and optimize device-level and system-level control strategies, thereby achieving multi-target collaborative control, improving the stability and reliability of electrical automation system operation, reducing line losses, improving load balancing, and ensuring efficient operation of the system.

[0017] Furthermore, the present invention takes harmonic factors into consideration and can analyze the current harmonics and voltage harmonics of the distribution cabinets in the area in a targeted manner, determine the current influence coefficient and the voltage influence coefficient according to the analysis results, and then provide accurate control strategies for harmonics for multi-objective electrical automation collaborative control, facilitate timely discovery and processing of harmonic problems, reduce the risk of failures caused by harmonics, achieve collaborative work between distribution cabinets, improve the operating efficiency and life of equipment, and ensure the stable operation of the electrical system.

[0018] Furthermore, the present invention considers factors such as temperature and aging, and determines the actual resistance of each line between the distribution cabinets based on the line current, temperature influence coefficient and aging influence coefficient, so as to determine the total line loss and build a line loss strategy to achieve targeted control of line loss. Under the premise of ensuring the safe operation of the line, the line loss is effectively reduced, providing a solid guarantee for the realization of the preset target strategy, improving the operating efficiency of the power system, and improving the reliability and stability of the entire power system.

[0019] Furthermore, the collaborative reward objective function of the present invention quantifies and weights the load balancing rate and communication efficiency, providing a clear and quantitative evaluation standard for the overall operating status of the power system, which is convenient for managers to intuitively understand the system performance and facilitate the subsequent collaborative objective reward function value to quickly determine whether the system is operating in the optimal state, providing a clear direction for the optimization and adjustment of the system, improving the operating efficiency of the power system, and enhancing the reliability and stability of the entire power system.

[0020] Furthermore, the present invention determines the corresponding adjustment method based on the total impact direction and the individual impact direction, including determining the action adjustment range under the total impact direction as a positive incentive, which can promote the overall performance of the system while taking into account the operating status of the individual distribution cabinet. Alternatively, repeated iterations are performed under the total impact direction as a negative reward, so as to prompt the system to find the optimal strategy in a complex and changeable operating environment, improve its own robustness and adaptability, achieve continuous performance improvement, improve the accuracy and effectiveness of control, realize multi-objective collaborative control, and ensure efficient operation of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 An algorithm diagram of a multi-objective electrical automation collaborative control method integrating reinforcement learning according to an embodiment of the present invention; Figure 2 A step diagram of a multi-objective electrical automation collaborative control method integrating reinforcement learning according to an embodiment of the present invention; Figure 3 A diagram of steps for constructing a harmonic target strategy for an embodiment of the present invention; Figure 4 A decision diagram for determining an action adjustment method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0022] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0023] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the protection scope of the present invention.

[0024] It should be noted that, in the description of the present invention, terms such as "up", "down", "left", "right", "inside" and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the drawings. This is merely for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present invention.

[0025] See also Figure 1 , Figure 2 As shown, Figure 1 Schematic diagram of the algorithm of the multi-objective electrical automation collaborative control method integrating reinforcement learning according to an embodiment of the present invention. Figure 2 This is a step diagram of a multi-objective electrical automation collaborative control method integrating reinforcement learning in an embodiment of the present invention.

[0026] Specifically, the present invention provides a multi-objective electrical automation collaborative control method integrating reinforcement learning, comprising: Step S1, real-time collection of electrical parameters and operating status of each distribution cabinet in the area to construct its own state vector, and construct neighboring state vectors according to the power transmission direction and line impedance between adjacent distribution cabinets, where the types of distribution cabinets include power access cabinets, load distribution cabinets and energy storage control cabinets; Step S2, constructing a preset target strategy according to the output current, output voltage, and line loss of each distribution cabinet, and determining the action vector of each distribution cabinet based on its own state vector, neighboring state vectors, and the preset target strategy; Step S3, determining the impact direction of the action vector on each distribution cabinet after execution based on the action vector, the own target reward function and the collaborative target reward function, wherein the impact direction includes an individual impact direction and a total impact direction; Among them, the self-target reward function is determined according to the type of distribution cabinet, including: the self-reward function of the power access cabinet includes a voltage deviation penalty term and a reactive output constraint term; the self-reward function of the load distribution cabinet includes a load balancing term and a tripping number penalty term; the self-reward function of the energy storage control cabinet includes an SOC health term and a charge and discharge efficiency term; The collaborative objective reward function is determined according to the total line loss, load balancing rate and communication efficiency between the distribution cabinets; Step S4, determining the action adjustment method according to the impact direction.

[0027] It can be understood that each distribution cabinet in the area is an intelligent agent in reinforcement learning. Step S1 is state perception, which monitors the equipment status of the distribution cabinet itself through its own state vector and neighboring state vector, and also reflects the network topology relationship in the distribution system. The neighboring state vector can accurately describe the dynamics of the distribution network based on the power transmission direction and line impedance. Steps S2 to S3 are strategy decisions. The preset target strategy is constructed based on the output current, voltage and line loss of the distribution cabinet, and then the action vector of each distribution cabinet is determined in combination with its own and neighboring state vectors. The action vector determines what operation the distribution cabinet performs. Based on the comparison result, it is determined whether to trigger the distribution cabinet to execute the corresponding action vector, so as to judge whether the current action meets the system requirements. Step S4 is reward evaluation, setting its own target reward and collaborative target reward, and determining the impact direction of the action vector on each distribution cabinet after execution (individual impact direction and total impact direction) through double-layer rewards. Step S5 is strategy optimization, which determines the action adjustment method based on the impact direction, thereby optimizing the control strategy of the distribution cabinet.

[0028] It is understandable that the preset target strategy is a preliminary decision-making plan based on rules and experience, which provides initial action guidance for the distribution cabinet. The self-target reward function and the collaborative target reward function are used to evaluate the results of the behavior, mainly by giving different reward values ​​to feedback whether the system state is developing in the expected direction. If the reward is low, it means that the preset target strategy may need to be adjusted.

[0029] In a specific embodiment, the acquisition frequency of the electrical parameters, operating status, power transmission direction, line impedance and other parameters of each distribution cabinet ranges from 20 times / min to 30 times / min. Preferably, the acquisition frequency is 25 times / min. The acquisition frequency can also be determined according to the rate of change of the electrical parameters, and the acquisition frequency is positively correlated with the rate of change of the electrical parameters. The self-reward function of the power access cabinet is the weighted sum of the voltage deviation penalty item and the reactive output constraint item of the power access cabinet; the self-reward function of the load distribution cabinet is the weighted sum of the load balancing item and the tripping number penalty item; the self-reward function of the energy storage control cabinet is the weighted sum of the SOC health item and the charging and discharging efficiency item. In implementation, the acquisition frequency value range and preferred value can be determined according to actual conditions, and the self-reward functions corresponding to the power access cabinet, load distribution cabinet and energy storage control cabinet can also be determined according to actual conditions. No specific limitation is made here and no further description is given.

[0030] The present invention integrates multiple goals such as voltage stability, harmonic control, and load balancing through a reinforcement learning framework to achieve effective coordination between distribution cabinets in the power system, including using its own state vector and neighboring state vectors to jointly represent the state of the device itself and reflect the network topology relationship, so as to achieve strong coupling of the power system. Constructing preset target strategies and determining action vectors based on multiple factors is more in line with actual needs. At the same time, a hierarchical reward mechanism is set up to distinguish between individual target rewards and collaborative target rewards, effectively evaluate the impact of action vectors on the system, and optimize device-level and system-level control strategies, thereby achieving multi-target collaborative control, improving the operational stability and reliability of the electrical automation system, reducing line losses, improving load balancing, and ensuring efficient operation of the system.

[0031] Specifically, in step S2, a harmonic target strategy is determined according to the output voltage and the output current, a line loss strategy is determined according to the line loss, and a preset target strategy is determined according to the harmonic target strategy and the line loss strategy.

[0032] It is understandable that using harmonic targets and line losses as preset target strategies can guide the intelligent agent (i.e., the control strategy of each distribution cabinet) to learn and adjust in the direction of reducing harmonics and line losses. Harmonics can cause problems such as equipment overheating, insulation aging, and malfunction. They can also increase line losses and voltage fluctuations and flickers, thereby affecting the stable operation of the entire electrical automation system. The harmonic target strategy helps to control the harmonic content within a reasonable range and reduce the adverse effects of harmonics on each distribution cabinet and other electrical equipment in the system. The line loss strategy can optimize the operating status and power distribution of each distribution cabinet, reduce energy losses in the line, and improve the transmission and utilization efficiency of electric energy.

[0033] See also Figure 3 As shown, it is a step diagram of constructing a harmonic target strategy in an embodiment of the present invention. Specifically, in step S2, constructing a harmonic target strategy includes: Step S211, determining the current harmonic content and the voltage harmonic content based on the output current and output voltage of each distribution cabinet; Step S212, determining a current influence coefficient based on the current harmonic content and the current harmonic threshold; Step S213, determining a voltage influence coefficient based on the voltage harmonic content and the voltage harmonic threshold; Step S214: determining a harmonic target strategy based on the current influence coefficient and the voltage influence coefficient.

[0034] It is understandable that in electrical systems, current harmonics and voltage harmonics will affect equipment operation and power quality. Determining the harmonic content from the output current and voltage can fully reflect the harmonic situation of the distribution cabinet. Comparing the harmonic content with the corresponding threshold to determine the influence coefficient, and then building a harmonic target strategy based on the influence coefficient can quantify the impact of harmonics on the system, provide a quantitative basis for subsequent control decisions, and make control more accurate.

[0035] In a specific embodiment, the output voltage and output current of the power distribution cabinet are subjected to a fast Fourier transform to convert the time domain signal into the frequency domain. In the frequency domain, the current harmonic content and the voltage harmonic content are determined respectively according to the amplitude and phase information of each frequency component. The value range of the current harmonic threshold is 5% to 8%, and preferably, the value of the current harmonic threshold is 5.5%. The value range of the voltage harmonic threshold is 3% to 5%, and preferably, the value of the voltage harmonic threshold is 4%. Current influence coefficient = |current harmonic threshold - current harmonic content | / current harmonic threshold; when the current harmonic content is closer to the current harmonic threshold, the current influence coefficient is smaller; the more the current harmonic threshold is exceeded, the greater the current influence coefficient. Voltage influence coefficient = |voltage harmonic threshold - voltage harmonic content | / voltage harmonic threshold. In implementation, the current influence coefficient and the voltage influence coefficient can also be determined according to the weighted calculation method, the electrical system simulation model, etc. The method of determining the current influence coefficient and the voltage influence coefficient can be adjusted according to the actual situation, which is not specifically limited here and will not be repeated.

[0036] In a specific embodiment, the harmonic target strategy is determined based on the current influence coefficient and the voltage influence coefficient, and the harmonic target strategy is characterized based on the harmonic comprehensive index. Specifically, the harmonic comprehensive index = current weight × current influence coefficient + voltage weight × voltage influence coefficient. The sum of the current weight and the voltage weight is 1, and the voltage weight is greater than the current weight. Preferably, the voltage weight takes a value of 0.7, and the current weight takes a value of 0.3. In implementation, the value range and preferred value of the voltage weight and the current weight can be determined according to actual conditions, and are not specifically limited here and will not be repeated.

[0037] Specifically, in step S2, the action vector corresponding to the harmonic target strategy includes: The power access cabinet injects reverse harmonics to offset the load harmonics; Disconnect the branch switch of the load distribution cabinet with high harmonic load or adjust the three-phase load distribution; The energy storage control cabinet offsets harmonics by charging / discharging energy storage.

[0038] It is understandable that nonlinear loads will generate a large number of harmonics, which will cause current and voltage waveform distortion and affect power quality. According to the principle of harmonic cancellation, the power access cabinet injects harmonics of equal magnitude and opposite direction to the load harmonics, which can cancel each other out, thereby reducing the total harmonic content in the power grid and making the current and voltage waveforms closer to sine waves.

[0039] It is understandable that high harmonic load is one of the main sources of harmonics in the power system. Disconnecting the corresponding branch switch can directly eliminate this part of the harmonic source and reduce the harmonic content. Unbalanced three-phase load will also cause harmonics. Unbalanced load will make the current size and phase of each phase different, thus causing current waveform distortion. By adjusting the three-phase load distribution, the three-phase load can be balanced and the harmonic generation caused by load imbalance can be reduced.

[0040] It is understandable that the energy storage control cabinet can adjust the power output and input during the charging / discharging process. When there are harmonics in the power grid, the energy storage control cabinet can control the energy storage device to perform charging / discharging operations according to the harmonic conditions, generate power components opposite to the harmonics, and thus offset some of the harmonics.

[0041] During implementation, active power filters or multi-functional inverters are required to achieve dynamic compensation capabilities.

[0042] The present invention takes harmonic factors into consideration and can analyze the current harmonics and voltage harmonics of the distribution cabinets in the area in a targeted manner, determine the current influence coefficient and the voltage influence coefficient according to the analysis results, and then provide accurate control strategies for harmonics for multi-objective electrical automation collaborative control, facilitate timely discovery and processing of harmonic problems, reduce the risk of failures caused by harmonics, achieve collaborative work between distribution cabinets, improve the operating efficiency and life of equipment, and ensure the stable operation of the electrical system.

[0043] Specifically, in step S2, the temperature influence coefficient is determined based on the historical ambient temperature and the historical line resistance, and the aging influence coefficient is determined based on the line usage time and the historical line resistance.

[0044] It is understandable that after long-term use, the electrical system lines will gradually age, the insulation performance will deteriorate, and leakage will easily occur. In addition, due to long-term oxidation, the resistivity of the lines will increase, resulting in increased resistance and increased line loss. In general, the resistance of the line will increase with the increase of temperature. The temperature is positively correlated with the line resistance, and the line loss is proportional to the resistance. Therefore, when determining the line loss strategy, it is necessary to consider the line usage time and ambient temperature.

[0045] In a specific embodiment, based on the physical properties of the metal conductor, the ambient temperature and the line resistance are linearly related, and the temperature influence coefficient is determined based on a linear regression model. The linear regression model is: Line resistance reference value = historical line resistance × [1 + temperature influence coefficient × (historical ambient temperature - temperature reference value)], wherein the line resistance reference value is the resistance reference value corresponding to the temperature reference value, and the temperature reference value is 25°C. Assuming that the line resistance increases exponentially over time (oxidation, fatigue factors), an exponential decay model is constructed: Historical line resistance = line resistance reference value × (1 + aging influence coefficient × [line usage time] ^ 2), where the aging influence coefficient is characterized as the growth rate of line resistance over time.

[0046] Specifically, in step S2, constructing a line loss strategy includes: Determine the actual resistance of the line between the distribution cabinets based on the initial resistance of the line, the temperature influence coefficient and the aging influence coefficient; The total line loss is determined according to the actual line resistance and the line current.

[0047] It can be understood that the line loss strategy is part of the preset target strategy. Based on the current state vector and neighboring state vector of the distribution cabinet, the line loss strategy can determine the action vector corresponding to the distribution cabinet, laying the foundation for subsequent data analysis. Line loss is mainly the heat generated by the resistance of the wire when the current passes through the wire, which consumes the electrical energy.

[0048] In a specific embodiment, the line resistance of the kth line = the initial line resistance of the kth line × [1 + temperature influence coefficient × temperature difference)] × (1 + aging influence coefficient × line usage time), and the temperature difference is the difference between the ambient temperature and the temperature reference value. The total line loss = ∑_(k=1)^n▒〖〖kth line current〗^2×〗the actual line resistance of the kth line.

[0049] Specifically, in step S2, the action vector corresponding to the line loss strategy includes: Adjust the reactive power output of the power access cabinet to reduce the line current; Switch the load distribution cabinet switch combination to evenly distribute the load; Adjust the charging / discharging power of the energy storage control cabinet to optimize the power flow distribution.

[0050] It is understandable that in the power system, the current is composed of active current and reactive current. The unreasonable distribution of reactive power will lead to an increase in the current in the line, thereby increasing the line loss. Therefore, by adjusting the reactive output of the power access cabinet, the reactive power distribution in the system can be changed, so that the power factor can be improved and the line current can be reduced.

[0051] It is understandable that the loads of the lines and distribution cabinets in the power system may be unbalanced, with some lines being overloaded and other lines being lightly loaded. Therefore, by switching the switch combination of the load distribution cabinet, the load distribution can be changed to make the load evenly distributed, thereby achieving current balance while avoiding excessive losses.

[0052] It is understandable that energy storage equipment plays a role in regulating power in the power system. In different periods and operating conditions, the power flow distribution may be unreasonable, resulting in excessive power transmission in some lines and large losses. By adjusting the charge / discharge power of the energy storage control cabinet, electric energy can be flexibly stored and released, the direction and size of the power flow in the system can be changed, the circuitous and unreasonable power transmission can be reduced, and the power loss in the line can be reduced.

[0053] The present invention considers factors such as temperature and aging, and determines the actual resistance of each line between the distribution cabinets based on the line current, temperature influence coefficient and aging influence coefficient, so as to determine the total line loss and build a line loss strategy to achieve targeted control of line loss. Under the premise of ensuring the safe operation of the line, the line loss is effectively reduced, providing a solid guarantee for the realization of the preset target strategy, improving the operating efficiency of the power system, and improving the reliability and stability of the entire power system.

[0054] Specifically, step S4 includes: Collect the output current of each load distribution cabinet in the area in real time, determine the average load of the load distribution cabinet based on the output current, and determine the load balancing rate based on the average load; Collect the data transmission volume and data reception volume of each power distribution cabinet in the area in real time, and determine the communication efficiency based on the data transmission volume and the data reception volume; The load balancing rate and communication efficiency are weighted to determine the collaborative objective reward function.

[0055] It is understandable that if the load distribution cabinet cannot achieve load balancing, it may cause insufficient power supply to some server cabinets, while other cabinets have excess power, affecting the overall performance and stability of the data center. Therefore, when setting the collaborative target reward function, priority is given to the overall load of the load distribution cabinet. Secondly, for the power distribution system, data exchange and communication between distribution cabinets are required to achieve coordinated control and status monitoring functions. High communication efficiency means that data can be transmitted quickly between distribution cabinets to ensure response speed. By comprehensively considering the load balancing rate and communication efficiency, the overall optimization of the regional distribution system in load distribution and communication coordination can be achieved, thereby improving the overall performance of the system.

[0056] In a specific embodiment, the average load is the output current mean of the load distribution cabinet in the area, and the load balancing rate is calculated by the standard deviation method. The load balancing rate = 1-output current standard deviation / average load, and the value range of the load balancing rate is 0-1. The closer to 1, the more balanced the load. The communication efficiency = data reception / data transmission, and the communication efficiency value range is 0-1. The higher the value, the higher the communication efficiency. The collaborative target reward function is the weighted sum of the load balancing rate and the communication efficiency, and the sum of the corresponding weights of the load balancing rate and the communication efficiency is 1. Preferably, the corresponding weights of the load balancing rate and the communication efficiency are both 0.5. The larger the value of the collaborative target reward function, the better the overall performance of the system, and the value of the collaborative target reward function is 0-1. In implementation, the corresponding weights of the load balancing rate and the communication efficiency can be determined according to actual conditions, and no specific limitation is made here. As long as the sum of the corresponding weights of the load balancing rate and the communication efficiency is 1, and the corresponding weight of the load balancing rate is less than the corresponding weight of the communication efficiency, it will not be repeated here.

[0057] The collaborative reward objective function of the present invention quantifies and weights the load balancing rate and communication efficiency, providing a clear and quantitative evaluation standard for the overall operating status of the power system, which is convenient for managers to intuitively understand the system performance and facilitate the value of the subsequent collaborative objective reward function to quickly determine whether the system is operating in the optimal state, providing a clear direction for the optimization and adjustment of the system, improving the operating efficiency of the power system, and enhancing the reliability and stability of the entire power system.

[0058] See also Figure 4 As shown, Figure 4 The decision diagram for determining the action adjustment method in the embodiment of the present invention is as follows. Specifically, in step S5, the action adjustment method is determined, wherein: If the total impact direction is a positive reward, determining the action adjustment amplitude according to the individual impact direction; If the total impact direction is a negative reward, then iterate steps S2 to S5 until the total impact direction is a positive reward or the number of iterations is equal to the preset maximum number of iterations.

[0059] It is understandable that when the total impact direction is a positive reward, it means that the action vector of the distribution cabinet has a positive impact on the overall system, but the individual impact direction may change with the execution of the distribution cabinet action vector. Therefore, after determining that the total impact direction is a positive reward, the action adjustment amplitude is determined according to the individual impact direction. When the total impact direction is a negative reward, it means that the current action vector and the decision based on the preset target strategy have not made the system develop in the expected direction, and the operating state of the system may become worse or fail to achieve the expected effect. At this time, by repeating the iterative steps S2 to S5, it is possible to re-evaluate the system state, adjust the preset target strategy, determine a new action vector, and find a strategy that can enable the system to obtain positive rewards.

[0060] It is understandable that the preset maximum number of iterations is set to prevent the algorithm from falling into an infinite loop. In actual applications, there may be some complex situations that make it difficult for the system to find a positive reward strategy in a short period of time. If the number of iterations is not limited, it may consume a lot of computing resources and time, and even cause the system to crash.

[0061] In a specific embodiment, the preset maximum number of iterations has a value range of 5000 to 10000 times, and preferably, the preset maximum number of iterations has a value of 6000. In implementation, the value range and preferred value of the preset maximum number of iterations can be determined according to actual conditions, and are not specifically limited here, nor are they described in detail.

[0062] In a specific embodiment, if the action vector taken by the distribution cabinet can reduce the harmonic content or line loss, and is close to the preset target strategy, the total impact direction is a positive reward, and the subsequent behavior of the distribution cabinet is determined according to the subsequent double-layer reward strategy. If the action vector taken by the distribution cabinet cannot reduce the harmonic content or line loss, and cannot be close to the preset target strategy, the total impact direction is a negative reward.

[0063] Specifically, in step S5, determining the action adjustment range includes: If the total impact direction is positive reward, and the individual impact direction corresponding to the power distribution cabinet is positive incentive, then the action adjustment amplitude of the action vector of the corresponding power distribution cabinet is increased; If the total impact direction is a positive reward and the individual impact direction corresponding to the distribution cabinet is a negative reward, the action adjustment amplitude of the action vector of the corresponding distribution cabinet is reduced.

[0064] It is understandable that when the total impact direction is positive reward, and the individual impact direction corresponding to the distribution cabinet is positive incentive, it means that the action taken by the distribution cabinet not only has a positive impact on the overall system, but also has a positive effect on improving its own state. For example, a load distribution cabinet adjusts the load distribution, so that the overall load balancing rate is improved (the total impact direction is positive reward), and at the same time, its own current, voltage and other parameters are more stable and reasonable (the individual impact direction is positive incentive), so the action adjustment amplitude of the action vector corresponding to the distribution cabinet is increased.

[0065] If the total impact direction is a positive reward, but the individual impact direction corresponding to the distribution cabinet is a negative reward, it means that although the action taken by the distribution cabinet has a positive impact on the overall system, it may have a negative effect on itself. For example, the energy storage control cabinet performs charging and discharging operations to optimize the system flow distribution, which reduces the overall system loss (the total impact direction is a positive reward), but its own battery life may be affected to a certain extent due to frequent charging and discharging (the individual impact direction is a negative reward). Reducing the action adjustment amplitude of the action vector corresponding to the distribution cabinet can ensure overall optimization while reducing the adverse effects on individual distribution cabinets.

[0066] In a specific embodiment, the adjustment percentage of the action adjustment amplitude of the action vector corresponding to the power distribution cabinet ranges from 4% to 8%, and preferably, the adjustment percentage corresponding to the action adjustment amplitude is 5%. In implementation, the value range and preferred value of the adjustment percentage can be determined according to actual conditions, and are not specifically limited here and will not be repeated.

[0067] The present invention determines the corresponding adjustment method based on the total impact direction and the individual impact direction, including determining the action adjustment range under the total impact direction as a positive incentive, which can promote the overall performance of the system while taking into account the operating status of the individual distribution cabinet. Alternatively, repeated iterations can be performed under the total impact direction as a negative reward, so as to prompt the system to find the optimal strategy in a complex and changeable operating environment, improve its own robustness and adaptability, achieve continuous performance improvement, improve the accuracy and effectiveness of control, realize multi-objective collaborative control, and ensure efficient operation of the system.

[0068] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.

Claims

1. A multi-objective electrical automation collaborative control method integrating reinforcement learning, characterized in that: include: Step S1, real-time collection of electrical parameters and operating status of each distribution cabinet in the area to construct its own state vector, and construct neighboring state vectors according to the power transmission direction and line impedance between adjacent distribution cabinets, where the types of distribution cabinets include power access cabinets, load distribution cabinets and energy storage control cabinets; Step S2, constructing a preset target strategy according to the output current, output voltage, and line loss of each distribution cabinet, and determining the action vector of each distribution cabinet based on its own state vector, neighboring state vectors, and the preset target strategy; Step S3, determining the impact direction of the action vector on each distribution cabinet after execution based on the action vector, the own target reward function and the collaborative target reward function, wherein the impact direction includes an individual impact direction and a total impact direction; Among them, the self-target reward function is determined according to the type of distribution cabinet, including: the self-reward function of the power access cabinet includes a voltage deviation penalty term and a reactive output constraint term; the self-reward function of the load distribution cabinet includes a load balancing term and a tripping number penalty term; the self-reward function of the energy storage control cabinet includes an SOC health term and a charge and discharge efficiency term; The collaborative objective reward function is determined according to the total line loss, load balancing rate and communication efficiency between the distribution cabinets; Step S4, determining the action adjustment method according to the impact direction.

2. The multi-objective electrical automation collaborative control method integrating reinforcement learning according to claim 1 is characterized in that: In step S2, a harmonic target strategy is determined according to the output voltage and the output current, a line loss strategy is determined according to the line loss, and a preset target strategy is determined according to the harmonic target strategy and the line loss strategy.

3. The multi-objective electrical automation collaborative control method integrating reinforcement learning according to claim 2 is characterized in that: In step S2, constructing a harmonic target strategy includes: Determine the current harmonic content and the voltage harmonic content based on the output current and output voltage of each distribution cabinet; Determining a current influence coefficient based on the current harmonic content and the current harmonic threshold; Determining a voltage influence coefficient based on the voltage harmonic content and the voltage harmonic threshold; The harmonic target strategy is determined based on the current influence coefficient and the voltage influence coefficient.

4. The multi-objective electrical automation collaborative control method integrating reinforcement learning according to claim 3 is characterized in that: In step S2, the action vector corresponding to the harmonic target strategy includes: The power access cabinet injects reverse harmonics to offset the load harmonics; Disconnect the branch switch of the load distribution cabinet with high harmonic load or adjust the three-phase load distribution; The energy storage control cabinet offsets harmonics by charging / discharging energy storage.

5. The multi-objective electrical automation collaborative control method integrating reinforcement learning according to claim 2 is characterized in that: In step S2, constructing a line loss strategy includes: Determine the actual resistance of the line between the distribution cabinets based on the initial resistance of the line, the temperature influence coefficient and the aging influence coefficient; The total line loss is determined according to the actual line resistance and the line current.

6. The multi-objective electrical automation collaborative control method integrating reinforcement learning according to claim 5 is characterized in that: In step S2, the temperature influence coefficient is determined based on the historical ambient temperature and the historical line resistance, and the aging influence coefficient is determined based on the line usage time and the historical line resistance.

7. The multi-objective electrical automation collaborative control method integrating reinforcement learning according to claim 5 is characterized in that: In step S2, the action vector corresponding to the line loss strategy includes: Adjust the reactive power output of the power access cabinet to reduce the line current; Switch the load distribution cabinet switch combination to evenly distribute the load; Adjust the charging / discharging power of the energy storage control cabinet to optimize the power flow distribution.

8. The multi-objective electrical automation collaborative control method integrating reinforcement learning according to claim 1 is characterized in that: The step S4 comprises: Collect the output current of each load distribution cabinet in the area in real time, determine the average load of the load distribution cabinet based on the output current, and determine the load balancing rate based on the average load; Collect the data transmission volume and data reception volume of each power distribution cabinet in the area in real time, and determine the communication efficiency based on the data transmission volume and the data reception volume; The collaborative objective reward function is determined by weighting according to the load balancing rate and the communication efficiency.

9. The multi-objective electrical automation collaborative control method integrating reinforcement learning according to claim 1 is characterized in that: In step S5, the action adjustment method is determined, wherein: If the total impact direction is a positive reward, determining the action adjustment amplitude according to the individual impact direction; If the total impact direction is a negative reward, then iterate steps S2 to S5 until the total impact direction is a positive reward or the number of iterations is equal to the preset maximum number of iterations.

10. The multi-objective electrical automation collaborative control method integrating reinforcement learning according to claim 9 is characterized in that: In step S5, determining the action adjustment range includes: If the total impact direction is positive reward, and the individual impact direction corresponding to the power distribution cabinet is positive incentive, then the action adjustment amplitude of the action vector of the corresponding power distribution cabinet is increased; If the total impact direction is a positive reward and the individual impact direction corresponding to the distribution cabinet is a negative reward, the action adjustment amplitude of the action vector of the corresponding distribution cabinet is reduced.

Citation Information

Patent Citations

  • A Hybrid Coordinated Control Method for Microgrids Based on Reinforcement Learning and Multi-Agent Theory

    CN111200285B

  • ANFIS-based electrical circuit aging degree online identification system and method

    CN113255203A

  • Double-layer cooperative control method for electricity-heat-gas comprehensive energy system

    CN115102158A

  • Multi-energy cooperative control method, device and equipment of power grid system and storage medium

    CN118040788A

  • Multi-source data fused wind turbine generator prediction optimization control method and system

    CN118523425A