Power quality compensation method and device of power grid, storage medium and electronic equipment

By collecting power grid data to determine power quality parameter values, formulating target compensation strategies, and using various power quality devices in a coordinated manner, the problem of unsatisfactory power quality compensation effects in the power grid has been solved, and intelligent power quality optimization of the power grid has been realized.

CN120978744APending Publication Date: 2025-11-18STATE GRID BEIJING ELECTRIC POWER CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511230192.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

In existing technologies, power quality compensation equipment for power grids lacks intelligent decision-making capabilities, making it difficult to cope with complex and ever-changing power grid environments and load demands, resulting in unsatisfactory power quality compensation effects.

Method used

By collecting initial operating status data of the power grid, power quality parameter values ​​are determined, target compensation strategies are formulated, and multiple power quality devices work together to perform power quality compensation until the deviation is less than a preset threshold.

Benefits of technology

It enables intelligent decision-making based on the real-time status of the power grid, optimizes the power quality compensation effect, and improves the stability and efficiency of power grid operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120978744A_ABST
    Figure CN120978744A_ABST
Patent Text Reader

Abstract

The invention discloses a power quality compensation method and device of a power grid, a storage medium and electronic equipment. The method comprises the following steps: collecting initial operation state data of a target power grid; determining an initial power quality parameter value of the target power grid based on the initial operation state data; determining a target compensation strategy of the target power grid based on the initial power quality parameter value; performing power quality compensation on the target power grid by adopting the compensation equipment according to the target compensation strategy to obtain a first power quality parameter value of the target power grid; based on the first electric energy quality parameter value, determining an electric energy quality deviation value of the target power grid, the electric energy quality deviation value being used for quantifying a difference between the first electric energy quality parameter value and an expected value; and when the electric energy quality deviation value is less than or equal to a preset threshold value, stopping performing electric energy quality compensation on the target power grid. The technical problem that the power quality compensation effect of the power grid is not ideal in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of power systems, and more specifically, to a method, apparatus, storage medium, and electronic equipment for power quality compensation in a power grid. Background Technology

[0002] As the two-way interaction between electric vehicles and the power grid deepens, the large-scale, unregulated integration of charging and discharging stations brings a series of power quality problems to the grid, such as voltage deviation, harmonic distortion, and three-phase imbalance. The power grid integrates power quality compensation devices such as active power filters (APFs) and static var compensators (SVCs), each with different compensation capabilities for specific power quality issues. Therefore, there is an urgent need for a multi-source collaborative adaptive power quality compensation method that utilizes the compensation capabilities of various power quality devices and the regulation capabilities of charging stations to solve power quality problems, promote vehicle-grid integration and interaction, and improve the stability of grid operation.

[0003] Current technologies primarily rely on rule-based control systems to compensate for power quality in the power grid. However, they lack the ability to make intelligent decisions based on the real-time operating status of the grid, making it difficult to cope with complex and ever-changing grid environments and load demands. Therefore, these technologies suffer from unsatisfactory power quality compensation effects.

[0004] There is currently no effective solution to the above problems. Summary of the Invention

[0005] This application provides a power quality compensation method, apparatus, storage medium, and electronic device for power grids, to at least solve the technical problem of unsatisfactory power quality compensation effect in related technologies.

[0006] According to one aspect of the embodiments of this application, a power quality compensation method for a power grid is provided, comprising: collecting initial operating state data of a target power grid; determining initial power quality parameter values ​​of the target power grid based on the initial operating state data, wherein the initial power quality parameter values ​​include initial voltage deviation, initial harmonic distortion rate, and initial three-phase imbalance; determining a target compensation strategy for the target power grid based on the initial power quality parameter values; using compensation equipment and following the target compensation strategy to perform power quality compensation on the target power grid, obtaining a first power quality parameter value of the target power grid; determining a power quality deviation of the target power grid based on the first power quality parameter value, wherein the power quality deviation is used to quantify the difference between the first power quality parameter value and an expected value; and stopping power quality compensation on the target power grid when the power quality deviation is less than or equal to a preset threshold.

[0007] According to another aspect of the embodiments of this application, a power quality compensation device for a power grid is provided, comprising: a data acquisition module for acquiring initial operating state data of a target power grid; a first determination module for determining initial power quality parameter values ​​of the target power grid based on the initial operating state data, wherein the initial power quality parameter values ​​include initial voltage deviation, initial harmonic distortion rate, and initial three-phase imbalance; a second determination module for determining a target compensation strategy for the target power grid based on the initial power quality parameter values; a compensation module for performing power quality compensation on the target power grid using compensation equipment according to the target compensation strategy to obtain a first power quality parameter value of the target power grid; a third determination module for determining a power quality deviation of the target power grid based on the first power quality parameter value, wherein the power quality deviation is used to quantify the difference between the first power quality parameter value and the expected value; and a judgment module for stopping power quality compensation on the target power grid when the power quality deviation is less than or equal to a preset threshold.

[0008] According to another aspect of the embodiments of this application, a non-volatile storage medium is provided, which stores a plurality of instructions adapted for a power quality compensation method for a power grid, any one of which is loaded by a processor.

[0009] According to another aspect of the embodiments of this application, an electronic device is provided, including: one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement any one of the following power quality compensation methods for a power grid.

[0010] According to another aspect of the embodiments of this application, a computer program product is provided, which, when executed on a data processing device, is a program adapted to perform the steps of a power quality compensation method for a power grid.

[0011] In this embodiment, initial operating status data of the target power grid is collected; based on the initial operating status data, initial power quality parameter values ​​of the target power grid are determined, wherein the initial power quality parameter values ​​include initial voltage deviation, initial harmonic distortion rate, and initial three-phase imbalance; based on the initial power quality parameter values, a target compensation strategy for the target power grid is determined; using compensation equipment, power quality compensation is performed on the target power grid according to the target compensation strategy to obtain the first power quality parameter value of the target power grid; based on the first power quality parameter value, the power quality deviation of the target power grid is determined, wherein the power quality deviation is used to quantify the difference between the first power quality parameter value and the expected value; if the power quality deviation is less than or equal to a preset threshold, power quality compensation for the target power grid is stopped. The goal is to determine the power quality compensation strategy for the power grid by collecting power grid operation status data, and to compensate the power grid according to the above compensation strategy until the power quality compensation condition is met, that is, the power quality deviation of the power grid is less than or equal to a preset threshold, at which point the power quality compensation of the power grid is stopped. This achieves the technical effect of optimizing the power quality compensation effect of the power grid, improving the satisfaction of the power quality compensation result, and thus solving the technical problem of unsatisfactory power quality compensation effect in related technologies. Attached Figure Description

[0012] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0013] Figure 1 This is a flowchart of a power quality compensation method for a power grid according to an embodiment of this application;

[0014] Figure 2 This is a schematic diagram of an optional power quality compensation device for a power grid according to an embodiment of this application. Detailed Implementation

[0015] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0016] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0017] According to an embodiment of this application, a method embodiment for power quality compensation of a power grid is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0018] Figure 1 This is a flowchart of a power quality compensation method for a power grid according to an embodiment of this application, such as... Figure 1 As shown, the method includes the following steps:

[0019] Step S102: Collect initial operating status data of the target power grid;

[0020] It is understandable that data acquisition devices, such as smart meters, voltage sensors, current sensors, and harmonic analyzers, are used to collect initial operating status data of the target power grid, including voltage values, current values, harmonic spectra, and three-phase balance. Accurate data acquisition can improve the quality of the initial operating status data of the target power grid, laying the foundation for the precise determination of subsequent initial power quality parameters and target compensation strategies.

[0021] Step S104: Based on the initial operating state data, determine the initial power quality parameter values ​​of the target power grid, wherein the initial power quality parameter values ​​include the initial voltage deviation, the initial harmonic distortion rate, and the initial three-phase imbalance.

[0022] It is understandable that, based on the initial operating status data of the target power grid, the initial power quality parameter values ​​are calculated, including the initial voltage deviation, initial harmonic distortion rate, and initial three-phase imbalance. Determining these initial power quality parameter values ​​provides crucial data support for power quality compensation, ensuring the accuracy and efficiency of the compensation results, and improving the operational stability and power quality compensation effectiveness of the target power grid.

[0023] Optionally, the voltage deviation mentioned above is used to measure the degree of difference between the actual voltage and the rated voltage in the target power grid. In the target power grid, the actual voltage should be as stable as possible near the rated value to ensure the normal operation of power equipment and the safety of electricity use. Excessive voltage deviation may affect the performance of power equipment, leading to overheating, damage, or unstable operation, and may also affect the stability and efficiency of the target power grid. Harmonic distortion rate is used to assess the degree of influence of harmonic currents or voltages on the fundamental frequency in the target power grid. Nonlinear loads in the target power grid (such as frequency converters, rectifiers, switching power supplies, etc.) generate harmonics, which can interfere with the normal operation of the target power grid, causing voltage fluctuations, current distortion, and even equipment damage. Monitoring the harmonic distortion rate is an important means of identifying and controlling harmonic sources. By reducing the harmonic distortion rate, power quality can be improved, power equipment can be protected, and the safety and stability of the target power grid can be ensured. Three-phase unbalance is used to measure the degree of difference between the voltages or currents of each phase in a three-phase power system. The three-phase power system in the target power grid should be balanced, that is, the voltages and currents of the three phases are equal in magnitude and differ in phase by 120 degrees. However, in actual operation, three-phase systems may become unbalanced due to factors such as uneven load distribution and equipment failure. Controlling the three-phase imbalance is crucial for maintaining the operating efficiency of the target power grid, reducing losses, and protecting equipment. Excessive imbalance can lead to problems such as equipment overheating, vibration, and reduced efficiency, negatively impacting the safety and stability of the target power grid.

[0024] Optionally, during power quality compensation for the target power grid, voltage deviation, harmonic distortion rate, and three-phase imbalance are key indicators for measuring power quality. By monitoring these indicators in real time, the operating status of the target power grid can be assessed, leading to the formulation and optimization of compensation strategies. For example, when a voltage deviation exceeds the allowable range, compensation equipment (such as a SVC) will provide or absorb reactive power to maintain voltage stability; when the harmonic distortion rate is too high, the APF will generate compensation current to eliminate harmonic interference; and when the three-phase imbalance increases, the output of the compensation equipment needs to be adjusted to balance the current or voltage of each phase. By comprehensively considering these three indicators, the power quality compensation system can be made more comprehensive and effective, ensuring high-quality and high-stability operation of the target power grid.

[0025] Step S106: Based on the initial power quality parameter values, determine the target compensation strategy for the target power grid;

[0026] It is understandable that, based on the initial power quality parameters of the target power grid, a target compensation strategy is determined to compensate the target power grid. This target compensation strategy may include the compensation sequence of the compensation equipment and the compensation actions used to quantify the amount of power quality compensation provided by the compensation equipment to the target power grid. By formulating a target compensation strategy based on the initial power quality parameters, the power quality problems of the target power grid can be addressed in a targeted manner, improving the satisfaction of the compensation effect, while enhancing the stability and efficiency of the power grid.

[0027] Optionally, a multi-objective optimization model of the power quality of the target power grid connected to the charging station can be established, and power quality compensation of the target power grid can be performed by optimizing this model. The multi-objective optimization model of the target power grid includes an objective function and constraints. The multi-objective optimization of the target power grid's power quality aims to maintain the harmonic distortion rate, voltage deviation, and three-phase imbalance of the target power grid within normal ranges by rationally configuring compensation equipment and optimizing the coordinated control strategy. The objective function F is defined as follows: t It can be constructed in the following way:

[0028] min F t =λ1△V t +λ2THD t +λ3ε t

[0029] Among them, F t Let represent the objective function at time t. The power quality deviation of the target power grid at time t can be characterized by the value of this objective function. △V t THD represents the voltage deviation of the target power grid at time t. t ε represents the harmonic distortion rate of the target power grid at time t. t Let λ represent the three-phase unbalance of the target power grid at time t, and let λ1, λ2, and λ3 represent the weights of voltage deviation, harmonic distortion rate, and three-phase unbalance, respectively.

[0030] Optionally, constraints may include power quality constraints, equipment capacity constraints, and power flow constraints.

[0031] Optionally, power quality constraints may include voltage deviation constraints, harmonic distortion rate constraints, and three-phase unbalance constraints. Voltage deviation constraints can be constructed as follows:

[0032] 0≤△V t =|V t -V t,ref |≤△V t,max

[0033] Among them, V t V represents the actual voltage of the target power grid at time t. t,refΔV represents the reference voltage value of the target power grid at time t. t,max This represents the upper limit of the voltage deviation of the target power grid at time t.

[0034] Harmonic distortion rate constraints can be constructed as follows:

[0035]

[0036] Among them, I 1,t THD represents the effective value of the fundamental current of the target power grid at time t. t,max I represents the upper limit of the harmonic distortion rate of the target power grid at time t. t,h Let represent the effective value of the h-th harmonic current of the target power grid at time t, and H represent the total number of H harmonic currents.

[0037] Three-phase unbalance constraints can be constructed as follows:

[0038]

[0039]

[0040] Among them, V A,t V B,t V C,t Let V represent the voltages of phases A, B, and C of the target power grid at time t. avg,t ε represents the average three-phase voltage of the target power grid at time t. t,max This represents the upper limit of the three-phase imbalance of the target power grid at time t.

[0041] Optionally, the equipment capacity constraints include the maximum compensation current constraint of the active filter, the reactive power output range constraint of the static var compensator, and the power constraint of the charging station.

[0042] The maximum compensation current constraint for the active power filter (APF) can be constructed as follows:

[0043]

[0044] in, This represents the actual compensation current value of the APF at time t. This represents the maximum allowable compensation current of the APF at time t, which is determined by the rated capacity, heat dissipation capacity, or operating strategy of the APF device, and is used to prevent overload.

[0045] The reactive power output range constraint of a static var compensator (SVC) can be constructed in the following way:

[0046]

[0047] Among them, Q t,SVCThis represents the actual reactive power output value of the SVC at time t. A positive value represents inductive reactive power compensation, and a negative value represents capacitive reactive power compensation. Let represent the minimum and maximum output reactive power of SVC at time t, respectively.

[0048] The power constraint of the charging station CS can be constructed as follows:

[0049]

[0050] Among them, P t,EV Q t,EV These represent the active power and reactive power of the EV (Electric Vehicle) during charging and discharging at time t (positive values ​​indicate charging, and negative values ​​indicate discharging); These represent the maximum active power and maximum reactive power of charging at time t, respectively. These represent the maximum active power and maximum reactive power of discharge at time t, respectively.

[0051] Alternatively, power flow constraints can be constructed as follows:

[0052]

[0053] Among them, Q i,t P i,t U represents the active power injected and the reactive power injected at node i at time t, respectively; i,t U j,t G represents the voltage values ​​at node i and node j at time t, respectively. ij θ represents the line conductance between node i and node j. ij,t B represents the voltage phase angle difference between node i and node j at time t, and n represents the total number of nodes. ij This represents the line susceptance between node i and node j.

[0054] In one optional embodiment, when there are multiple compensation devices, a target compensation strategy for the target power grid is determined based on the initial power quality parameter values, including: determining the compensation order of the multiple compensation devices based on the initial power quality parameter values; determining the target compensation actions corresponding to the multiple compensation devices based on the initial power quality parameter values ​​and the compensation order, wherein the target compensation actions are used to quantify the amount of power quality compensation provided by the corresponding compensation devices to the target power grid; and determining the target compensation strategy based on the compensation order and the target compensation actions corresponding to the multiple compensation devices.

[0055] It is understood that if multiple power quality compensation devices exist in the target power grid, such as charging stations (CS), active power filters, and static var compensators (SVCs), the target compensation strategy for the target power grid is determined as follows: First, based on the initial power quality parameters of the target power grid, the compensation sequence of the multiple compensation devices is determined. Second, based on the initial power quality parameters and the compensation sequence of the multiple compensation devices, the target compensation actions corresponding to each of the multiple compensation devices are determined. The compensation sequence of the multiple power quality compensation devices and the target compensation actions corresponding to each of the multiple compensation devices together constitute the target compensation strategy for the target power grid. The target compensation strategy determined through the above steps, with the collaborative operation of multiple compensation devices, can achieve efficient and accurate improvement of power quality, enhance the stability of the target power grid operation, optimize the resource allocation of the target power grid, and reduce the impact on power quality for users.

[0056] In one optional embodiment, determining the compensation order of multiple compensation devices based on initial power quality parameter values ​​includes: acquiring operating equipment data corresponding to each of the multiple compensation devices; determining the power quality problem type of the target power grid based on the initial power quality parameter values; determining the operating status of each of the multiple compensation devices based on the operating equipment data corresponding to each of the multiple compensation devices; and determining the compensation order based on the power quality problem type and the operating status of each of the multiple compensation devices.

[0057] To determine the compensation sequence of multiple compensation devices, the following steps are taken: First, obtain the operating data of each device. Second, based on the initial power quality parameters of the target grid, determine the type of power quality problem. For example, if the initial harmonic distortion rate is larger than the initial voltage deviation and initial three-phase imbalance, the harmonic problem in the target grid is more severe. Next, based on the operating data of each compensation device, determine its operating status, such as normal operation, maintenance, or inefficient operation. Finally, based on the power quality problem type and the operating status of each device, determine the compensation sequence. For example, if the main power quality problem in the target grid is severe harmonics, and the active power filter is operating normally, then the active power filter should be prioritized for power quality compensation. Determining the compensation sequence based on initial power quality parameters not only allows for efficient and rational allocation and use of target grid resources but also addresses power quality problems specifically, improving the satisfaction of compensation effectiveness and enhancing the operational stability and reliability of the target grid.

[0058] Optionally, the operating data of the aforementioned compensation equipment may include, but is not limited to, the current operating mode, remaining capacity, historical compensation effect, and response speed of the compensation equipment. For example, for an active filter, it is necessary to know its remaining compensation current capacity and historical harmonic suppression effect; for a static var compensator, it is necessary to know its current reactive power output level and remaining adjustment range.

[0059] In one optional embodiment, based on the initial power quality parameter values ​​and the compensation sequence, the target compensation actions corresponding to multiple compensation devices are determined, including: determining the first compensation action of the first-level compensation device among the multiple compensation devices based on the compensation sequence and the initial power quality parameter values; updating the initial power quality parameter values ​​based on the first compensation action to obtain the second power quality parameter values ​​of the target power grid; determining the second compensation action of the second-level compensation device among the multiple compensation devices based on the compensation sequence and the second power quality parameter values; and determining the target compensation actions corresponding to the multiple compensation devices by using the method of determining the second compensation action.

[0060] It is understandable that, based on the compensation sequence of multiple compensation devices and the initial power quality parameters of the target power grid, the first compensation action of the first-level compensation device among the multiple compensation devices is determined, that is, the amount of power quality compensation provided to the target power grid by the compensation device ranked first in the compensation sequence is quantified. After the first-level compensation device compensates for the power quality of the target power grid using the first compensation action, the power quality of the target power grid will change. In order to improve the accuracy of the determination results of the compensation actions of subsequent compensation devices, the initial power quality parameters of the target power grid are updated based on the first compensation action of the first-level compensation device, resulting in the second power quality parameter value of the target power grid. At this time, based on the compensation sequence and the second power quality parameter value, the second compensation action of the second-level compensation device among the multiple compensation devices is determined. Following the above method, the target compensation actions corresponding to multiple compensation devices are determined sequentially according to the compensation sequence. By gradually determining and optimizing the target compensation actions of the compensation devices, not only can the power quality problem be solved efficiently, but the optimal allocation of compensation device resources can also be promoted, thereby improving the power quality compensation effect of the target power grid.

[0061] Optionally, the power quality optimization problem of the target power grid can be transformed into a Markov decision process to determine the target compensation strategy for the target power grid. To solve the power quality problem of the target power grid connected to the charging station, the aforementioned multi-objective power quality optimization problem of the target power grid can be constructed as a partially observable Markov decision process containing (S, A, R, P, γ) and solved, where S represents the operating state of the target power grid, characterized by the operating state data of the target power grid; A represents the target compensation action of the compensation equipment; R represents the expected cumulative reward, calculated according to the reward function of the agent; P represents the state transition probability matrix; and γ represents the discount factor. In the operating state S of the target power grid... t Which compensation action a should be selected? t It is determined by the policy network π. Three Markov Decision Process (MDP) models are designed, using APF, SVG, and CS as agents respectively. Among them, S... t This represents the operating state of the target power grid at time t, characterized by the operating state data of the target power grid at time t. t This represents the target compensation action of the compensation device at time t. The agent observes the operating state of the target power grid and continuously interacts with it using the compensation actions determined by the policy network π to form a large number of empirical trajectories, thereby selecting the optimal compensation action and optimizing the power quality.

[0062] Optionally, an APF agent is constructed, including constructing a state space, an action space, a state transition probability matrix, and a reward function.

[0063] The state space of an APF agent can be constructed in the following way:

[0064]

[0065] Among them, s t,APF Let t represent the state space of the APF. P represents the h-th harmonic current after the charging pile operates at time t. t,load I represents the load power at time t. t,APF This represents the APF injection current at time t.

[0066] The state space of an APF agent can be constructed in the following way:

[0067]

[0068] Among them, a t,APF Let t represent the state space of the APF. This represents the actual compensation current value of the APF at time t. The actual compensation current value of the output is adjusted accordingly. To compensate for the current harmonics of the target power grid.

[0069] The state transition probability matrix of an APF agent can be constructed as follows:

[0070]

[0071] in, This indicates that at time t, the running state s t and action a t Transition to running state s t+1 The probability, s t ′ represents the updated operating state of the target power grid at time t after the operation of the previous level compensation equipment.

[0072] The reward function for an APF agent can be constructed as follows:

[0073] r t =-(ω1×△V) t +ω2×ε t +ω3×THD t )

[0074] Where, r t ω1, ω2, and ω3 represent the instantaneous reward at time t, respectively, and represent the reward weights for voltage deviation, harmonic distortion rate, and three-phase imbalance.

[0075] Optionally, an SVC agent is constructed, including constructing a state space, an action space, a state transition probability matrix, and a reward function.

[0076] The state space of an SVC agent can be constructed in the following way:

[0077] s t,SVG =[V t CS-a Q t,SVC ,P t,load ]

[0078] Among them, s t,SVG V represents the state space of SVC at time t. t CS-a Q represents the node voltage after the charging station agent's action at time t. t,SVC This represents the output power of the SVC at time t.

[0079] The action space of an SVC agent can be constructed in the following way:

[0080]

[0081] Among them, a t,SVC Denotes the action space of SVC at time t. This represents the actual reactive power compensation of the SVC at time t, used for voltage support and imbalance adjustment.

[0082] The state transition probability matrix of an SVC agent can be constructed as follows:

[0083]

[0084] Where s′ represents the updated operating state of the target power grid after the operation of the previous level compensation equipment.

[0085] The reward function for an SVC agent can be constructed as follows:

[0086] r t =-(ω1×△V) t +ω2×ε t +ω3×THD t )

[0087] Optionally, a charging station intelligent agent is constructed, including constructing a state space, an action space, a state transition probability matrix, and a reward function.

[0088] The state space of the charging station agent can be constructed in the following way:

[0089] s t,CS =[V t ,I t,h C t,rem ,P t,load ]

[0090] Among them, s t,CS Let C represent the state space of CS at time t. t,rem This indicates the remaining adjustable capacity (reactive or active) of the charging station at time t.

[0091] The action space of the charging station's intelligent agent can be constructed in the following way:

[0092] a t,CS =[Q t,CS ,P t,CS ,I t,CS ]

[0093] Among them, a t,CS Let Q represent the action space of CS at time t. t,CS P represents the amount of reactive power compensation called at time t. t,CS I represents the amount of active power peak reduction reserved at time t. t,CS This represents the harmonic current of the compensation output at time t.

[0094] The state transition probability matrix of the charging station agent can be constructed as follows:

[0095]

[0096] The reward function for the charging station agent can be constructed as follows:

[0097] r t =-(ω1×△V) t +ω2×ε t +ω3×THD t )

[0098] Optionally, the target compensation actions corresponding to multiple compensation devices can be determined and the initial power quality parameter values ​​updated in the following manner. If the compensation sequence is as follows: charging station as the first-level compensation device, static var compensator as the second-level compensation device, and active power filter as the third-level compensation device, the coordinated power quality compensation process includes initializing residuals (i.e., initial power quality parameter values), priority charging station (CS) output, static var compensator (SVC) secondary compensation, active power filter (APF) final-level compensation, and action issuance.

[0099] Optionally, the residuals are initialized first. Residual g t,0 The following methods can be used to determine this:

[0100] g t,0 =g t

[0101] Among them, g t,0 [ΔV] represents the power quality residual vector of the target power grid at time t after level 0 compensation. t THD t ,ε t ] T g t This represents the initial power quality parameter value of the target power grid in vector form at time t.

[0102] Optionally, the charging station (CS) should prioritize power output. The target compensation action a for the charging station is determined. t,CS a t,CS The following methods can be used to determine this:

[0103]

[0104] Among them, a t,CS π represents the target compensation action of the charging station at time t. CS The policy network Actor, representing the intelligent agent of the charging station, will define the state space s. t,CS and residual g t,0 Input is given to the policy network Actor, and output is the target compensation action a. t,CS . a CSThis represents the lower limit of the compensation action of the CS agent (i.e., the charging station agent). This represents the upper limit of the compensation actions of the CS agent, which is determined by the device capacity constraint.

[0105] Update residual g t,0 The updated residual g is obtained. t,1 (The second power quality parameter value of the target power grid in vector form at time t). g t,1 The following methods can be used to determine this:

[0106] g t,1 =g t,0 -J CS a t,CS

[0107]

[0108] Among them, J i This represents the sensitivity matrix of agent i to the residual, with dimension 1. d i Let J be the vector length of the action space of agent i, represented in vector form. CS This represents the sensitivity matrix of the CS agent to the residuals. This represents the d-th action in the action space of agent i, expressed in vector form. i A compensation action.

[0109] Optionally, the next step is secondary compensation using a static var compensator (SVC). The target compensation action a of the SVC is determined. t,SVC a t,SVC The following methods can be used to determine this:

[0110]

[0111] Among them, a t,SVC π represents the target compensation action of SVC at time t. SVC The policy network Actor, representing the SVC agent, defines the state space s. t,SVC and residual g t,1 Input is given to the policy network Actor, and output is the target compensation action a. t,SVC a SVC This represents the lower bound of the compensation action of the SVC agent. This indicates the upper limit of the compensation actions of the SVC agent, which is determined by the device capacity constraint.

[0112] Update residual g t,1 The updated residual g is obtained. t,2 g t,2 The following methods can be used to determine this:

[0113] g t,2 =g t,1 -J SVC a t,SVC

[0114] Among them, J SVC This represents the sensitivity matrix of the SVC agent to the residual.

[0115] Optionally, this is followed by final-stage compensation of the active power filter (APF). The target compensation action a for the active power filter is determined. t,APF a t,APF The following methods can be used to determine this:

[0116]

[0117] Among them, a t,APF π represents the target compensation action of the APF at time t. APF The policy network Actor, representing the APF agent, will have a state space s t,APF and residual g t,2 Input is given to the policy network Actor, and output is the target compensation action a. t,APF a APF This represents the lower bound of the compensation action of the APF agent. This indicates the upper limit of compensation actions for the APF agent, which is determined by the device capacity constraint.

[0118] Optionally, the final action is issued, sending {a} all at once. t,APF ,a t,SVC ,a t,CS The information is then distributed to the corresponding compensation equipment.

[0119] Optionally, the above {a t,APF ,a t,SVC ,a t,CS The model can be trained using the MADDPG (Multi-Agent Deep Deterministic Policy Gradient) algorithm. The agent's objective is determined. Each agent's objective is to maximize its expected cumulative reward. Among them, R t r is the expected cumulative reward of agent i at time t. i t γ is the instantaneous reward for agent i at time t, and γ is the discount factor.

[0120] Optionally, a Critic network (i.e., a target reward evaluation network, used to evaluate the compensation actions determined by the Actor network, and thus calculate the joint Q-value of all agent actions) is determined. The Critic network Q-value for each agent i is... i=(s,a1,…,a N Estimate the global running state and the joint Q-value of all agent actions. The parameters θ of the Critic network for agent i can be updated using the Bellman equation. i θ i Update using the following method:

[0121]

[0122] y = r i +γQ′ i (s',a′1,…,a' N ;θ i )

[0123] Among them, Q i This indicates that under running state s, N agents are performing compensation actions (a1, ..., a...). N The joint Q-value under ( ); y represents Q i Expected value; a N Represents the compensation action of agent N; θ i The Critic network parameters of agent i are represented by r. i Q represents the immediate reward of agent i; i 'Indicates running state s', N agents are performing compensation actions (a′1, ..., a′) N The joint Q value under ().

[0124] Optionally, an Actor network (i.e., a policy network used to determine the compensation actions of the compensation equipment based on the operating state of the target power grid) is defined. The parameters of the Actor network for each agent i are... It is updated by maximizing the expectation of its Critic function. Update using the following method:

[0125]

[0126] in, π represents the parameters of the policy network Actor of agent i. i Denotes the policy network Actor, o of agent i. i a represents the local observation of agent i. i This represents the compensation action of agent i. The parameters are adjusted using the policy gradient method. Update.

[0127] Alternatively, the Actor network and Critic network can be trained in the following manner. First, initialize the Actor network π. i and Critic network Q iand their target Actor network π' i and the target Critic network Q' i Secondly, agent i interacts with the target power grid. At each time t, agent i, according to its policy network π, i Select target compensation action a i The target power grid returns to its next state (i.e., according to the target compensation action a). i Updated target grid operating status (used to update power quality parameter values ​​of the target grid) and immediate reward r i Then, the running state, target compensation action, immediate reward, and next running state are stored in the experience replay pool. Next, a batch is sampled from the experience replay pool, and the Critic network parameters and Actor network parameters for each agent are updated using the aforementioned formula. Finally, the target network is softly updated, with its parameters θ updated slowly. θ is updated as follows:

[0128]

[0129] in, Let θ represent the updated parameters, τ represent the parameters before the update, and τ represent the soft update coefficients of the target reward evaluation network parameters. Repeat the interaction and update process until training is complete.

[0130] Step S108: Using compensation equipment, power quality compensation is performed on the target power grid according to the target compensation strategy to obtain the first power quality parameter value of the target power grid.

[0131] It is understandable that, after determining the target compensation strategy for the target power grid, compensation equipment is used to perform power quality compensation on the target power grid according to the aforementioned target compensation strategy, and the initial power quality parameter values ​​of the target power grid are updated to obtain the first power quality parameter values ​​of the target power grid. After the compensation equipment performs compensation according to the target compensation strategy, it can significantly reduce the voltage deviation, harmonic distortion rate, and three-phase imbalance of the target power grid, improve power quality, reduce power equipment losses and faults, and ensure the normal operation of the power system.

[0132] Optionally, each compensation device performs corresponding power quality compensation based on the received control signal. For example, the charging station adjusts its charging and discharging power according to its target compensation action to balance the load of the target power grid; the APF generates and injects compensation current to filter out harmonics in the target power grid according to the target compensation action; and the SVC outputs reactive power to stabilize the voltage deviation of the target power grid or balance the three phases according to the target compensation action.

[0133] Step S110: Based on the first power quality parameter value, determine the power quality deviation of the target power grid, wherein the power quality deviation is used to quantify the difference between the first power quality parameter value and the expected value.

[0134] It is understandable that, based on the first power quality parameter value of the target power grid, a power quality deviation is determined, which quantifies the difference between the first power quality parameter value and the expected value. By calculating the power quality deviation, a quantitative assessment of the compensation effect on the target power grid can be provided, which is of great significance for continuously optimizing compensation strategies, improving power grid power quality, enhancing system stability, and increasing user satisfaction.

[0135] In one optional embodiment, determining the power quality deviation of the target power grid based on the first power quality parameter value includes: determining the updated voltage deviation, updated harmonic distortion rate, and updated three-phase imbalance of the target power grid based on the first power quality parameter value; and determining the power quality deviation by weighted summation based on the updated voltage deviation, updated harmonic distortion rate, updated three-phase imbalance, the weight of the updated voltage deviation, the weight of the updated harmonic distortion rate, and the weight of the updated three-phase imbalance.

[0136] It is understandable that, based on the first power quality parameter value, the updated voltage deviation, updated harmonic distortion rate, and updated three-phase imbalance of the target power grid are determined. The power quality deviation of the target power grid is then determined by weighted calculation based on the updated voltage deviation, updated harmonic distortion rate, and updated three-phase imbalance, along with their respective weights. Determining the power quality deviation based on the first power quality parameter value not only allows for a comprehensive and accurate evaluation of the effectiveness of power quality compensation strategies but also guides the continuous optimization of these strategies, ultimately improving the power quality of the target power grid and enhancing the efficiency and stability of grid operation.

[0137] Alternatively, the above objective function F can be used. t The function value represents the power quality deviation. After each power quality compensation of the target grid is completed, the objective function F is used with updated operating status data. t The power quality deviation of the target power grid is calculated, and by comparing it with a preset threshold, it is determined whether power quality compensation for the target power grid needs to continue.

[0138] Optionally, methods such as expert scoring and the analytic hierarchy process (AHP) can be used to determine the subjective weights of power quality parameters such as voltage deviation, harmonic distortion rate, and three-phase imbalance. Expert scoring involves inviting experts in the field to rate the importance of each power quality parameter based on their expertise and experience, thereby determining the subjective weight of the parameter. This method directly utilizes the experience and expertise of industry experts, but it is highly subjective, with potential discrepancies in scores among experts and the influence of personal biases. The AHP uses pairwise comparisons to determine the relative importance of power quality parameters. This method provides a systematic approach to handling subjective judgments and can quantify the relative importance of power quality parameters. However, the calculation process is more complex, and the reasonableness of the expert scores significantly impacts the results.

[0139] Optionally, objective evaluation methods such as Criteria Importance Through Intercriteria Correlation (CRITIC) and entropy weighting can be used to determine the objective weights of power quality parameters such as voltage deviation, harmonic distortion rate, and three-phase imbalance. Criteria Importance Through Intercriteria Correlation (CRITIC) comprehensively considers the correlation between power quality parameters and the variability of their values ​​to determine their objective weights. It is suitable for handling decision-making problems involving multiple power quality parameters, where these parameters may be interrelated, and their importance varies depending on the dispersion of their values. Entropy weighting is a weight determination method based on the principle of information entropy. It measures the uncertainty or information content of power quality parameter values; the lower the information entropy, the higher the certainty of the information provided by the power quality parameter, and its weight should be greater.

[0140] Optionally, methods such as fuzzy comprehensive evaluation, weighted index method, and weighted average method can be used to determine the weights for updating voltage deviation, harmonic distortion rate, and three-phase imbalance based on the subjective and objective weights of voltage deviation, harmonic distortion rate, and three-phase imbalance. The fuzzy comprehensive evaluation method uses fuzzy set theory and fuzzy mathematics to comprehensively evaluate the weights, taking into account the uncertainty of subjective and objective weights, thereby determining the weights for updating voltage deviation, harmonic distortion rate, and three-phase imbalance. The weighted index method uses an exponential function to strengthen or weaken subjective and objective weights, then performs a weighted merging to obtain the weights for updating voltage deviation, harmonic distortion rate, and three-phase imbalance. The weighted average method sums the subjective and objective weights according to a certain weighted ratio to obtain the weights for updating voltage deviation, harmonic distortion rate, and three-phase imbalance. The choice of which method to combine subjective and objective weights depends on factors such as the specific application scenario, the quality and complexity of the data, and the knowledge level and availability of experts. In determining the weights for updating voltage deviation, updating harmonic distortion rate, and updating three-phase imbalance, if expert resources are abundant and have sufficient industry experience, the analytic hierarchy process (AHP) can be given priority. However, for situations with limited data or where the uncertainty of weight information needs to be processed, fuzzy comprehensive evaluation or weighted index methods can be prioritized.

[0141] Step S112: If the power quality deviation is less than or equal to a preset threshold, stop power quality compensation for the target power grid.

[0142] It is understandable that after using a target compensation strategy to compensate the power quality of the target power grid, the power quality deviation is compared with a preset threshold. If the power quality deviation is less than or equal to the preset threshold, it indicates that the power quality of the target power grid has met the requirements, and power quality compensation for the target power grid is stopped. By setting and comparing the power quality deviation with a preset threshold, precise control of the compensation process can be achieved, unnecessary compensation can be avoided, thereby optimizing the compensation effect of the target power grid's power quality and improving satisfaction with the compensation results.

[0143] In an optional embodiment, when the power quality deviation exceeds a preset threshold, the method further includes: determining an update compensation strategy for the target power grid based on a first power quality parameter value; using a compensation device to perform power quality compensation on the target power grid according to the update compensation strategy to obtain a third power quality parameter value for the target power grid; updating the power quality deviation based on the third power quality parameter value to obtain an updated power quality deviation; and stopping power quality compensation on the target power grid when the updated power quality deviation is less than or equal to the preset threshold.

[0144] It is understandable that if the power quality deviation exceeds a preset threshold, it indicates that the power quality of the target grid does not meet the requirements, and power quality compensation for the target grid still needs to continue. First, based on the first power quality parameter value, the target compensation strategy is updated to obtain the updated compensation strategy for the target grid. Second, according to the updated compensation strategy, compensation equipment is used to perform power quality compensation on the target grid, and the first power quality parameter value is updated to obtain the third power quality parameter value for the target grid. Next, based on the third power quality parameter value, the power quality deviation is updated to obtain the updated power quality deviation. The updated power quality deviation is compared with the preset threshold. If the updated power quality deviation is less than or equal to the preset threshold, power quality compensation for the target grid is stopped; if the updated power quality deviation is greater than the preset threshold, power quality compensation for the target grid continues in the above manner until the updated power quality deviation is less than or equal to the preset threshold, at which point power quality compensation for the target grid is stopped. Through the above process, the compensation strategy can be flexibly adjusted according to the real-time changes in the power quality of the target power grid, enhancing its adaptability and response speed to complex situations in the target power grid. At the same time, iteratively updating the compensation strategy until the power quality deviation reaches the standard can ensure the accuracy and effectiveness of power quality compensation for the target power grid, avoid power quality parameters from exceeding the allowable range, and improve the safety and stability of the target power grid operation.

[0145] Through the above steps S102 to S112, the goal of collecting power grid operating status data, determining the power quality compensation strategy of the power grid, and compensating the power grid according to the above compensation strategy until the power quality compensation condition is met, that is, the power quality deviation of the power grid is less than or equal to a preset threshold, and then stopping the power quality compensation of the power grid, can be achieved. This achieves the technical effect of optimizing the power quality compensation effect of the power grid, improving the satisfaction of the power quality compensation result of the power grid, and thus solving the technical problem of unsatisfactory power quality compensation effect of the power grid in related technologies.

[0146] Based on the above embodiments and optional embodiments, this application proposes an implementation method for an optional power quality compensation method for a power grid. This implementation method can be understood as a multi-source coordinated charging and discharging power quality adaptive compensation method. The steps of this method include:

[0147] Step S1: Establish a multi-objective optimization model for the power quality of the target power grid to which the charging station is connected.

[0148] Step S11: Construct the objective function.

[0149] Multi-objective optimization of power quality in the target power grid involves rationally configuring compensation equipment and optimizing the coordinated control strategy to maintain the harmonic distortion rate, voltage deviation, and three-phase imbalance of the target power grid within normal ranges. An objective function is defined based on this objective. The objective function F is... t Construct it in the following way:

[0150] min F t =λ1△V t +λ2THD t +λ3ε t

[0151] Among them, F t Let represent the objective function at time t. The power quality deviation of the target power grid at time t can be characterized by the value of this objective function. △V t THD represents the voltage deviation of the target power grid at time t. t ε represents the harmonic distortion rate of the target power grid at time t. t Let λ represent the three-phase unbalance of the target power grid at time t, and let λ1, λ2, and λ3 represent the weights of voltage deviation, harmonic distortion rate, and three-phase unbalance, respectively.

[0152] Step S12: Construct constraints.

[0153] Consider power quality constraints, equipment capacity constraints, and power flow constraints.

[0154] Power quality constraints include voltage deviation constraints, harmonic distortion rate constraints, and three-phase unbalance constraints. Voltage deviation constraints are constructed as follows:

[0155] 0≤△V t =|V t -V t,ref |≤△V t,max

[0156] Among them, V t V represents the actual voltage of the target power grid at time t. t,ref ΔV represents the reference voltage value of the target power grid at time t. t,max This represents the upper limit of the voltage deviation of the target power grid at time t.

[0157] Harmonic distortion rate constraints are constructed as follows:

[0158]

[0159] Among them, I 1,t THD represents the effective value of the fundamental current of the target power grid at time t. t,max I represents the upper limit of the harmonic distortion rate of the target power grid at time t. t,hLet represent the effective value of the h-th harmonic current of the target power grid at time t, and H represent the total number of H harmonic currents.

[0160] The three-phase unbalance constraint is constructed as follows:

[0161]

[0162] Among them, V A,t V B,t V C,t Let V represent the voltages of phases A, B, and C of the target power grid at time t. avg,t ε represents the average three-phase voltage of the target power grid at time t. t,max This represents the upper limit of the three-phase imbalance of the target power grid at time t.

[0163] Equipment capacity constraints include the maximum compensation current constraint of the active filter, the reactive power output range constraint of the static var compensator, and the power constraint of the charging station.

[0164] The maximum compensation current constraint of the active power filter (APF) is constructed as follows:

[0165]

[0166] in, This represents the actual compensation current value of the APF at time t. This represents the maximum allowable compensation current of the APF at time t, which is determined by the rated capacity, heat dissipation capacity, or operating strategy of the APF device, and is used to prevent overload.

[0167] The reactive power output range constraint of the Static Var Compensator (SVC) is constructed in the following way:

[0168]

[0169] Among them, Q t,SVC This represents the actual reactive power output value of the SVC at time t. A positive value represents inductive reactive power compensation, and a negative value represents capacitive reactive power compensation. Let represent the minimum and maximum output reactive power of SVC at time t, respectively.

[0170] The power constraint of the charging station CS is constructed as follows:

[0171]

[0172] Among them, P t,EV Q t,EV These represent the active power and reactive power of the EV (Electric Vehicle) during charging and discharging at time t (positive values ​​indicate charging, and negative values ​​indicate discharging); These represent the maximum active power and maximum reactive power of charging at time t, respectively. These represent the maximum active power and maximum reactive power of discharge at time t, respectively.

[0173] Current flow constraints are constructed in the following way:

[0174]

[0175]

[0176] Among them, Q i,t P i,t U represents the active power injected and the reactive power injected at node i at time t, respectively; i,t U j,t G represents the voltage values ​​at node i and node j at time t, respectively. ij θ represents the line conductance between node i and node j. ij,t B represents the voltage phase angle difference between node i and node j at time t, and n represents the total number of nodes. ij This represents the line susceptance between node i and node j.

[0177] Step S2 transforms the power quality optimization problem of the target power grid into a Markov decision process.

[0178] To address the power quality issue of the target power grid connected to the charging station, the multi-objective optimization problem of the target power grid's power quality can be constructed as a partially observable Markov decision process containing (S, A, R, P, γ) and solved. Here, S represents the operating state of the target power grid, characterized by its operating state data; A represents the target compensation action of the compensation equipment; R represents the expected cumulative reward, calculated based on the agent's reward function; P represents the state transition probability matrix; and γ represents the discount factor. Under the operating state S of the target power grid... t Which compensation action a should be selected? t It is determined by the policy network π. Three Markov Decision Process (MDP) models are designed, using APF, SVG, and CS as agents respectively. Among them, S... t This represents the operating state of the target power grid at time t, characterized by the operating state data of the target power grid at time t. t This represents the target compensation action of the compensation device at time t. The agent observes the operating state of the target power grid and continuously interacts with it using the compensation actions determined by the policy network π to form a large number of empirical trajectories, thereby selecting the optimal compensation action and optimizing the power quality.

[0179] Step S21: Construct the APF agent, including constructing the state space, action space, state transition probability matrix, and reward function.

[0180] The state space of an APF agent is constructed as follows:

[0181]

[0182] Among them, s t,APF Let t represent the state space of the APF. P represents the h-th harmonic current after the charging pile operates at time t. t,load I represents the load power at time t. t,APF This represents the APF injection current at time t.

[0183] The state space of an APF agent is constructed as follows:

[0184]

[0185] Among them, a t,APF Let t represent the state space of the APF. This represents the actual compensation current value of the APF at time t. The actual compensation current value of the output is adjusted accordingly. To compensate for the current harmonics of the target power grid.

[0186] The state transition probability matrix of the APF agent is constructed as follows:

[0187]

[0188] in, This indicates that at time t, the running state s t and action a t Transition to running state s t+1 The probability, s t ′ represents the updated operating state of the target power grid at time t after the operation of the previous level compensation equipment.

[0189] The reward function for an APF agent is constructed as follows:

[0190] r t =-(ω1×△V) t +ω2×ε t +ω3×THD t )

[0191] Where, r t ω1, ω2, and ω3 represent the instantaneous reward at time t, respectively, and represent the reward weights for voltage deviation, harmonic distortion rate, and three-phase imbalance.

[0192] Step S22: Construct the SVC agent, including constructing the state space, action space, state transition probability matrix, and reward function.

[0193] The state space of an SVC agent is constructed as follows:

[0194] s t,SVG =[V t CS-a Q t,SVC ,P t,load ]

[0195] Among them, s t,SVG V represents the state space of SVC at time t. t CS-a Q represents the node voltage after the charging station agent's action at time t. t,SVC This represents the output power of the SVC at time t.

[0196] The action space of an SVC agent is constructed in the following way:

[0197]

[0198] Among them, a t,SVC Denotes the action space of SVC at time t. This represents the actual reactive power compensation of the SVC at time t, used for voltage support and imbalance adjustment.

[0199] The state transition probability matrix of the SVC agent is constructed as follows:

[0200]

[0201] Where s′ represents the updated operating state of the target power grid after the operation of the previous level compensation equipment.

[0202] The reward function for the SVC agent is constructed as follows:

[0203] r t =-(ω1×△V) t +ω2×ε t +ω3×THD t )

[0204] Step S23: Construct the charging station intelligent agent, including constructing the state space, action space, state transition probability matrix, and reward function.

[0205] The state space of the charging station's intelligent agent is constructed as follows:

[0206] s t,CS =[V t ,I t,h C t,rem,P t,load ]

[0207] Among them, s t,CS Let C represent the state space of CS at time t. t,rem This indicates the remaining adjustable capacity (reactive or active) of the charging station at time t.

[0208] The action space of the charging station's intelligent agent is constructed in the following way:

[0209] a t,CS =[Q t,CS ,P t,CS ,I t,CS ]

[0210] Among them, a t,CS Let Q represent the action space of CS at time t. t,CS P represents the amount of reactive power compensation called at time t. t,CS I represents the amount of active power peak reduction reserved at time t. t,CS This represents the harmonic current of the compensation output at time t.

[0211] The state transition probability matrix of the charging station agent is constructed as follows:

[0212]

[0213] The reward function for the charging station agent is constructed as follows:

[0214] r t =-(ω1×△V) t +ω2×ε t +ω3×THD t )

[0215] Step S24, the coordinated power quality compensation process, includes initializing residuals (i.e., initial power quality parameter values), priority charging station (CS) output, static var compensator (SVC) secondary compensation, active power filter (APF) final-stage compensation, and action issuance (i.e., the compensation sequence is: charging station as the first-stage compensation device, static var compensator as the second-stage compensation device, and active power filter as the third-stage compensation device).

[0216] Step S241, initialize the residual. Residual g t,0 The following method is used to determine:

[0217] g t,0 =g t

[0218] Among them, g t,0 [ΔV] represents the power quality residual vector of the target power grid at time t after level 0 compensation. t THD t,ε t ] T g t This represents the initial power quality parameter value of the target power grid in vector form at time t.

[0219] Step S242, the charging station (CS) prioritizes power output. Determine the target compensation action a for the charging station. t,CS a t,CS The following method is used to determine:

[0220]

[0221] Among them, a t,CS π represents the target compensation action of the charging station at time t. CS The policy network Actor, representing the intelligent agent of the charging station, will define the state space s. t,CS and residual g t,0 Input is given to the policy network Actor, and output is the target compensation action a. t,CS . a CS This represents the lower limit of the compensation action of the CS agent (i.e., the charging station agent). This represents the upper limit of the compensation actions of the CS agent, which is determined by the device capacity constraint.

[0222] Update residual g t,0 The updated residual g is obtained. t,1 (The second power quality parameter value of the target power grid in vector form at time t). g t,1 The following method is used to determine:

[0223] g t,1 =g t,0 -J CS a t,CS

[0224]

[0225] Among them, J i This represents the sensitivity matrix of agent i to the residual, with dimension 1. d i Let J be the vector length of the action space of agent i, represented in vector form. CS This represents the sensitivity matrix of the CS agent to the residuals. This represents the d-th action in the action space of agent i, expressed in vector form. i A compensation action.

[0226] Step S243, Static Var Compensator (SVC) Secondary Compensation. Determine the target compensation action a of the Static Var Compensator. t,SVC a t,SVC The following method is used to determine:

[0227]

[0228] Among them, a t,SVC π represents the target compensation action of SVC at time t. SVC The policy network Actor, representing the SVC agent, defines the state space s. t,SVC and residual g t,1 Input is given to the policy network Actor, and output is the target compensation action a. t,SVC a SVC This represents the lower bound of the compensation action of the SVC agent. This indicates the upper limit of the compensation actions of the SVC agent, which is determined by the device capacity constraint.

[0229] Update residual g t,1 The updated residual g is obtained. t,2 g t,2 The following method is used to determine:

[0230] g t,2 =g t,1 -J SVC a t,SVC

[0231] Among them, J SVC This represents the sensitivity matrix of the SVC agent to the residual.

[0232] Step S244, active power filter (APF) final stage compensation. Determine the target compensation action a for the active power filter. t,APF a t,APF The following method is used to determine:

[0233]

[0234] Among them, a t,APF π represents the target compensation action of the APF at time t. APF The policy network Actor, representing the APF agent, will have a state space s t,APF and residual g t,2 Input is given to the policy network Actor, and output is the target compensation action a. t,APF a APF This represents the lower bound of the compensation action of the APF agent. This indicates the upper limit of compensation actions for the APF agent, which is determined by the device capacity constraint.

[0235] Step S245, Action issuance: Issue {a} in one go. t,APF ,a t,SVC ,a t,CS The information is then distributed to the corresponding compensation equipment.

[0236] Step S3: Use the MADDPG algorithm (Multi-Agent Deep Deterministic Policy Gradient) to solve and train the model.

[0237] Step S31: Determine the agent's objective. Each agent's objective is to maximize its expected cumulative reward. Among them, R t r is the expected cumulative reward of agent i at time t. i t γ is the instantaneous reward for agent i at time t, and γ is the discount factor.

[0238] Step S32: Determine the Critic network (i.e., the target reward evaluation network, used to evaluate the compensation actions determined by the Actor network, and thus calculate the joint Q-value of all agent actions). The Critic network Q-value for each agent i is determined. i =(s,a1,...,a N Estimate the global running state and the joint Q-value of all agent actions. The parameters θ of the Critic network for agent i can be updated using the Bellman equation. i θ i Update using the following method:

[0239]

[0240] y = r i +γQ′ i (s',a′1,…,a' N ;θ i )

[0241] Among them, Q i This indicates that under running state s, N agents are performing compensation actions (a1, ..., a...). N The joint Q-value under ( ); y represents Q i Expected value; a N Represents the compensation action of agent N; θ i The Critic network parameters of agent i are represented by r. i Q represents the immediate reward of agent i; i 'Indicates running state s', N agents are performing compensation actions (a′1, ..., a′) N The joint Q value under ( ).

[0242] Step S33: Determine the Actor network (i.e., the policy network, used to determine the compensation actions of the compensation equipment based on the operating state of the target power grid). Parameters of the Actor network for each agent i. It is updated by maximizing the expectation of its Critic function. Update using the following method:

[0243]

[0244] in, π represents the parameters of the policy network Actor of agent i. i Denotes the policy network Actor, o of agent i. i a represents the local observation of agent i. i This represents the compensation action of agent i. The parameters are adjusted using the policy gradient method. Update.

[0245] Step S34, Algorithm Flow.

[0246] Step S341, Initialization. Initialize the Actor network π. i and Critic network Q i and their target Actor network π' i and the target Critic network Q' i .

[0247] Step S342, Interaction. Agent i interacts with the target power grid. At each time t, agent i, according to its policy network π, i Select target compensation action a i The target power grid returns to its next state (i.e., according to the target compensation action a). i Updated target grid operating status (used to update power quality parameter values ​​of the target grid) and immediate reward r i .

[0248] Step S343, Store Experience. Store the running status, target compensation action, immediate reward, and next running status into the experience replay pool.

[0249] Step S344, Sampling and Update. Sample a batch from the experience replay pool and update the Critic network parameters and Actor network parameters for each agent using the aforementioned formula.

[0250] Step S345, soft update the target network. Update the parameters θ of the target network in a slow manner. θ is updated as follows:

[0251]

[0252] in, Let θ represent the updated parameters, θ represent the parameters before the update, and τ represent the soft update coefficient of the target reward evaluation network parameters.

[0253] Step S346, repeat. Repeat the interaction and update process until training is complete.

[0254] The above-mentioned optional implementation methods achieve at least the following effects: through the coordinated work of multiple compensation devices, power quality can be improved efficiently and accurately, the stability of the target power grid operation can be enhanced, the resource allocation of the target power grid can be optimized, and the impact on power quality of users can be reduced; by gradually determining and optimizing the target compensation actions of the compensation devices, not only can power quality problems be solved efficiently, but also the optimal allocation of compensation device resources can be promoted, thereby improving the power quality compensation effect of the target power grid.

[0255] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0256] This embodiment also provides a power quality compensation device for a power grid, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the terms "module" and "device" can refer to a combination of software and / or hardware that performs a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0257] According to an embodiment of this application, an apparatus embodiment for implementing a power quality compensation method for a power grid is also provided. Figure 2 This is a schematic diagram of a power quality compensation device for a power grid according to an embodiment of this application, as shown below. Figure 2 As shown, the power quality compensation device for the power grid includes a data acquisition module 202, a first determination module 204, a second determination module 206, a compensation module 208, a third determination module 210, and a judgment module 212. The device will be described below.

[0258] The data acquisition module 202 is used to acquire the initial operating status data of the target power grid;

[0259] The first determining module 204 is connected to the data acquisition module 202 and is used to determine the initial power quality parameter values ​​of the target power grid based on the initial operating state data. The initial power quality parameter values ​​include the initial voltage deviation, the initial harmonic distortion rate, and the initial three-phase imbalance.

[0260] The second determining module 206, connected to the first determining module 204, is used to determine the target compensation strategy of the target power grid based on the initial power quality parameter values.

[0261] The compensation module 208 is connected to the second determining module 206 and is used to perform power quality compensation on the target power grid according to the target compensation strategy using compensation equipment to obtain the first power quality parameter value of the target power grid.

[0262] The third determining module 210, connected to the compensation module 208, is used to determine the power quality deviation of the target power grid based on the first power quality parameter value, wherein the power quality deviation is used to quantify the difference between the first power quality parameter value and the expected value.

[0263] The judgment module 212, connected to the third determination module 210, is used to stop power quality compensation for the target power grid when the power quality deviation is less than or equal to a preset threshold.

[0264] The power quality compensation device for a power grid provided in this application embodiment, by setting up a data acquisition module 202, a first determination module 204, a second determination module 206, a compensation module 208, a third determination module 210, and a judgment module 212, achieves the purpose of determining the power quality compensation strategy of the power grid by collecting the operating status data of the power grid, and performing power compensation on the power grid according to the above compensation strategy until the power quality compensation condition is met, that is, the power quality deviation of the power grid is less than or equal to a preset threshold, and then the power quality compensation on the power grid is stopped. This achieves the technical effect of optimizing the power quality compensation effect of the power grid, improving the satisfaction of the power quality compensation result, and thus solving the technical problem of unsatisfactory power quality compensation effect of the power grid in related technologies.

[0265] It should be noted that the above modules can be implemented by software or hardware. For example, for the latter, it can be implemented in the following ways: the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.

[0266] It should be noted that the data acquisition module 202, the first determining module 204, the second determining module 206, the compensation module 208, the third determining module 210, and the judgment module 212 mentioned above correspond to steps S102 to S112 in the embodiments. The instances and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the content disclosed in the above embodiments. It should be noted that the above modules, as part of the device, can run in a computer terminal.

[0267] It should be noted that the optional or preferred implementation methods of this embodiment can be found in the relevant descriptions in the embodiments, and will not be repeated here.

[0268] The power quality compensation device of the power grid mentioned above may also include a processor and a memory. The data acquisition module 202, the first determination module 204, the second determination module 206, the compensation module 208, the third determination module 210, the judgment module 212, etc. are all stored in the memory as program units. The processor executes the above program units stored in the memory to realize the corresponding functions.

[0269] The processor contains a core that retrieves the corresponding program unit from memory. One or more cores may be configured. Memory may include non-persistent memory in computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory includes at least one memory chip.

[0270] This application provides a non-volatile storage medium storing a program that, when executed by a processor, implements a power quality compensation method for the power grid.

[0271] This application provides an electronic device including a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, it performs the following steps: collecting initial operating state data of a target power grid; determining initial power quality parameter values ​​of the target power grid based on the initial operating state data, wherein the initial power quality parameter values ​​include initial voltage deviation, initial harmonic distortion rate, and initial three-phase imbalance; determining a target compensation strategy for the target power grid based on the initial power quality parameter values; using compensation equipment and following the target compensation strategy, performing power quality compensation on the target power grid to obtain a first power quality parameter value for the target power grid; determining a power quality deviation of the target power grid based on the first power quality parameter value, wherein the power quality deviation is used to quantify the difference between the first power quality parameter value and the expected value; and stopping power quality compensation on the target power grid when the power quality deviation is less than or equal to a preset threshold. The device described herein can be a server, PC, etc.

[0272] This application also provides a computer program product, which, when executed on a data processing device, is suitable for executing an initialization program comprising the following steps: collecting initial operating state data of a target power grid; determining initial power quality parameter values ​​of the target power grid based on the initial operating state data, wherein the initial power quality parameter values ​​include initial voltage deviation, initial harmonic distortion rate, and initial three-phase imbalance; determining a target compensation strategy for the target power grid based on the initial power quality parameter values; using compensation equipment and following the target compensation strategy, performing power quality compensation on the target power grid to obtain a first power quality parameter value of the target power grid; determining a power quality deviation of the target power grid based on the first power quality parameter value, wherein the power quality deviation is used to quantify the difference between the first power quality parameter value and the expected value; and stopping power quality compensation on the target power grid when the power quality deviation is less than or equal to a preset threshold.

[0273] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0274] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0275] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0276] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0277] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0278] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0279] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0280] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0281] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0282] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A power quality compensation method for an electrical grid, characterized by, The method comprises the following steps: collecting initial operation state data of a target power grid; determining initial power quality parameter values of the target power grid based on the initial operation state data, wherein the initial power quality parameter values include initial voltage deviation, initial harmonic distortion rate and initial three-phase imbalance degree; determining a target compensation strategy of the target power grid based on the initial power quality parameter values; compensating power quality of the target power grid according to the target compensation strategy by using a compensation device to obtain first power quality parameter values of the target power grid; determining a power quality deviation of the target power grid based on the first power quality parameter values, wherein the power quality deviation is used to quantify the gap between the first power quality parameter values and expected values; stopping the compensation of power quality of the target power grid when the power quality deviation is less than or equal to a preset threshold.

2. The method of claim 1, wherein, When the compensation device is multiple, the step of determining the target compensation strategy of the target power grid based on the initial power quality parameter values comprises: determining a compensation sequence of the multiple compensation devices based on the initial power quality parameter values; determining target compensation actions corresponding to the multiple compensation devices respectively based on the initial power quality parameter values and the compensation sequence, wherein the target compensation actions are used to quantify the power quality compensation amount of the corresponding compensation device to the target power grid; determining the target compensation strategy based on the compensation sequence and the target compensation actions corresponding to the multiple compensation devices respectively.

3. The method of claim 2, wherein, The step of determining the compensation sequence of the multiple compensation devices based on the initial power quality parameter values comprises: obtaining operation device data corresponding to the multiple compensation devices respectively; determining a power quality problem type of the target power grid based on the initial power quality parameter values; determining operation states of the multiple compensation devices respectively based on the operation device data corresponding to the multiple compensation devices respectively; determining the compensation sequence based on the power quality problem type and the operation states of the multiple compensation devices respectively.

4. The method of claim 2, wherein, The step of determining the target compensation actions corresponding to the multiple compensation devices respectively based on the initial power quality parameter values and the compensation sequence comprises: determining a first compensation action of a first-level compensation device in the multiple compensation devices based on the compensation sequence and the initial power quality parameter values; updating the initial power quality parameter values based on the first compensation action to obtain second power quality parameter values of the target power grid; determining a second compensation action of a second-level compensation device in the multiple compensation devices based on the compensation sequence and the second power quality parameter values; determining the target compensation actions corresponding to the multiple compensation devices respectively in the same way as determining the second compensation action. The step of determining the power quality deviation of the target power grid based on the first power quality parameter values comprises:

5. The method of claim 1, wherein, determining updated voltage deviation, updated harmonic distortion rate and updated three-phase imbalance degree of the target power grid based on the first power quality parameter values; ​ Based on the updated voltage deviation, the updated harmonic distortion rate, the updated three-phase imbalance, the weight of the updated voltage deviation, the weight of the updated harmonic distortion rate, and the weight of the updated three-phase imbalance, a weighted sum is used to determine the power quality deviation.

6. The method according to any one of claims 1 to 5, characterized in that, In the case where the power quality deviation is greater than the preset threshold, the method further comprises: Based on the first power quality parameter value, an updated compensation strategy for the target power grid is determined; Using the compensation device, the target power grid is compensated for power quality according to the updated compensation strategy, to obtain a third power quality parameter value of the target power grid; Based on the third power quality parameter value, the power quality deviation is updated to obtain an updated power quality deviation; In the case where the updated power quality deviation is less than or equal to a preset threshold, the compensation for the target power grid is stopped.

7. A power quality compensation device for an electrical grid, characterized by Comprise: A data acquisition module for acquiring initial operating state data of a target power grid; A first determination module for determining an initial power quality parameter value of the target power grid based on the initial operating state data, wherein the initial power quality parameter value includes an initial voltage deviation, an initial harmonic distortion rate, and an initial three-phase imbalance; A second determination module for determining a target compensation strategy for the target power grid based on the initial power quality parameter value; A compensation module for using a compensation device to compensate for the power quality of the target power grid according to the target compensation strategy, to obtain a first power quality parameter value of the target power grid; A third determination module for determining a power quality deviation of the target power grid based on the first power quality parameter value, wherein the power quality deviation is used to quantify the difference between the first power quality parameter value and an expected value; A judgment module for stopping the compensation for the power quality of the target power grid in the case where the power quality deviation is less than or equal to a preset threshold.

8. A non-volatile storage medium, comprising: The non-volatile storage medium stores a plurality of instructions, which are suitable for being loaded and executed by the processor to implement the power quality compensation method of the power grid according to any one of claims 1 to 6.

9. An electronic device, comprising: Comprise: One or more processors and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the power quality compensation method of the power grid according to any one of claims 1 to 6.

10. A computer program product comprising computer instructions, characterized in that, The computer instructions are executed by the processor to implement the power quality compensation method of the power grid according to any one of claims 1 to 6.