Method for adjusting compensation parameters of active filter based on Q learning and related equipment

By adjusting the compensation parameters of the active power filter using the Q-learning algorithm, the problem of the difficulty in automatically adjusting the compensation parameters of traditional active power filters in complex power grid environments is solved. This achieves autonomous learning and optimal compensation, thereby improving the efficiency of power grid harmonic compensation.

CN121769946APending Publication Date: 2026-03-31SHENZHEN QIDIAN NEW ENERGY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Traditional active power filters struggle to automatically adjust harmonic compensation parameters in complex power grid environments, requiring frequent manual modifications to achieve the desired effect, resulting in low efficiency.

Method used

The Q-learning algorithm is used to adjust the compensation parameters of the active filter. By acquiring the load current and the compensation output current, the harmonic characteristics are analyzed, the action is defined and the Q value is calculated, the Q table is updated, and the optimal control parameters are selected to achieve optimal harmonic compensation.

Benefits of technology

It realizes the autonomous learning and optimization capabilities of active filters, automatically adjusts to the optimal operating point, improves the efficiency and effect of harmonic compensation, and adapts to complex working conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121769946A_ABST
    Figure CN121769946A_ABST
Patent Text Reader

Abstract

The invention provides a method for adjusting compensation parameters of an active filter based on Q learning and related equipment. The method comprises the following steps: acquiring load current in a power grid and compensation output current of the active filter; obtaining harmonic characteristics based on load current analysis, and defining the harmonic characteristics as a current environment state; defining the adjustment amount of the to-be-adjusted control parameter in the active filter as an action; in the current environment state, executing an action to adjust the to-be-adjusted control parameter, and calculating to obtain a Q value according to a difference value between the load current and the compensation output current; using a Q-learning algorithm to update and store a Q table of the corresponding relationship between the current environment state, the action and the Q value according to the Q value; according to the updated Q table, the control parameter corresponding to the action enabling the Q value to be maximum is selected in the current environment state to be used for controlling a loop, and therefore optimal compensation of harmonic waves is achieved. According to the method, the active filter has autonomous learning and optimizing capabilities, the optimal compensation strategy can be automatically found, and the compensation efficiency and effect are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power grid technology, and in particular to a method and related equipment for adjusting the compensation parameters of an active filter based on Q-learning. Background Technology

[0002] With the development of technology, human electricity consumption is increasing day by day, and the power grid faces many challenges. For example, many industrial loads generate a large amount of harmonics. Harmonics are like an "invisible killer," and their harm is systemic and cumulative. They not only directly damage equipment and increase operating costs, but also pose a serious threat to the core interests of enterprises by causing production interruptions and product quality problems.

[0003] Active power filters are widely used for harmonic compensation, but most devices rely on manual adjustment of harmonic current control parameters, such as PI current loops and voltage loops, some of which cannot be changed at all. However, traditional active power filter control struggles to adapt to various complex situations, such as in the environment of an induction furnace. In such environments, the high harmonic content and voltage dips in the mains power grid pose a significant challenge to the machine's performance. As the mains harmonic voltage and load harmonic current change, the harmonic compensation effect of the control loop formed by the original control parameters becomes less ideal. In various complex environments, repeated manual and frequent modifications of control parameters are necessary to achieve the desired compensation effect.

[0004] Therefore, existing technologies still need to be improved and developed. Summary of the Invention

[0005] This invention provides a method and related equipment for adjusting the compensation parameters of an active filter based on Q-learning. The main objective of this invention is to solve the technical problems mentioned in the background section of the prior art.

[0006] The first aspect of this invention provides a method for adjusting the compensation parameters of an active filter based on Q-learning, comprising: Obtain the load current in the power grid and the compensation output current of the active filter; Harmonic characteristics are obtained based on the load current analysis, and these harmonic characteristics are defined as the current environmental state. The adjustment amount of the control parameter to be adjusted in the active filter is defined as the action; Under the current environmental conditions, the action is performed to adjust the control parameter to be adjusted, and the Q value used to evaluate the merits of the compensation strategy is calculated based on the difference between the load current and the compensation output current. Using the Q-learning algorithm, the Q-table, which stores the correspondence between the current environment state, the action, and the Q-value, is updated according to the Q-value; Based on the updated Q table, under the current environmental conditions, the control parameter corresponding to the action that maximizes the Q value is selected for the control loop of the active filter to achieve optimal harmonic compensation.

[0007] In an optional embodiment of the first aspect of the present invention, obtaining harmonic features based on the load current analysis includes: extracting the amplitude and phase information of each harmonic in the load current using a fast Fourier transform algorithm, as the harmonic features.

[0008] In an optional embodiment of the first aspect of the invention, the Q value is calculated using the following formula: Q = 1 / abs(I _load -I _inv ), where I _load I represents the instantaneous or effective value of the load current. _inv This refers to the instantaneous or effective value of the compensated output current.

[0009] In an optional embodiment of the first aspect of the present invention, the control parameters to be adjusted include the resonant gain parameter and the resonant bandwidth parameter for a specific subharmonic in the quasi-PR controller.

[0010] In an optional embodiment of the first aspect of the present invention, the control parameter to be adjusted further includes the proportional coefficient and integral coefficient of the PI control loop.

[0011] In an optional embodiment of the first aspect of the invention, the control parameters to be adjusted further include the amplitude and phase of a reference harmonic current for generating the compensation current.

[0012] In an optional embodiment of the first aspect of the invention, the Q-learning algorithm updates the Q-table using the following update formula: Q(s,a)=Q(s,a)+α×[R+γ×max(Q(s',a'))-Q(s,a) ], where s is the current environment state, a is the current action, α is the learning rate, R is the reward value calculated based on the current Q value, γ is the discount factor, s' is the next environment state, and a' is all possible actions in the next environment state.

[0013] A second aspect of the present invention provides an apparatus for adjusting the compensation parameters of an active filter based on Q-learning, the apparatus comprising: The current acquisition module is used to acquire the load current in the power grid and the compensation output current of the active filter. A state parameter definition block is used to obtain harmonic characteristics based on the load current analysis and define the harmonic characteristics as the current environmental state; The action parameter definition module is used to define the adjustment amount of the control parameter to be adjusted in the active filter as an action; The Q-value calculation module is used to perform the action to adjust the control parameter to be adjusted under the current environmental state, and calculate the Q-value for evaluating the merits of the compensation strategy based on the difference between the load current and the compensation output current. The Q-table update module is used to update the Q-table storing the current environment state, the action and the correspondence between the Q-values ​​according to the Q-values ​​using the Q-learning algorithm. The harmonic compensation module is used to select the control parameters corresponding to the action that maximizes the Q value under the current environmental conditions, based on the updated Q table, for use in the control loop of the active filter, so as to achieve optimal harmonic compensation.

[0014] A third aspect of the present invention provides a device for adjusting the compensation parameters of an active filter based on Q-learning. The device for adjusting the compensation parameters of an active filter based on Q-learning includes: a memory and at least one processor, wherein the memory stores instructions, and the memory and the at least one processor are interconnected via a line. The at least one processor invokes the instructions in the memory to cause the device for adjusting the active filter compensation parameters based on Q-learning to perform the method for adjusting the active filter compensation parameters based on Q-learning as described in any one of the first aspects of the invention.

[0015] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method for adjusting the compensation parameters of an active filter based on Q-learning as described in any one of the first aspects of the present invention.

[0016] Beneficial Effects: This invention provides a method and related equipment for adjusting the compensation parameters of an active power filter based on Q-learning. The method includes acquiring the load current in the power grid and the compensation output current of the active power filter; analyzing the load current to obtain harmonic characteristics and defining these characteristics as the current environmental state; defining the adjustment amount of the control parameter to be adjusted in the active power filter as an action; executing the action to adjust the control parameter under the current environmental state, and calculating the Q value based on the difference between the load current and the compensation output current; using a Q-learning algorithm to update the Q table storing the correspondence between the current environmental state, actions, and Q values; and selecting the control parameter corresponding to the action that maximizes the Q value under the current environmental state for use in the control loop to achieve optimal harmonic compensation. This invention enables the active power filter to have autonomous learning and optimization capabilities, automatically finding the optimal compensation strategy and improving compensation efficiency and effectiveness. Attached Figure Description

[0017] Figure 1 This is a schematic diagram of an embodiment of a method for adjusting the compensation parameters of an active filter based on Q-learning according to the present invention; Figure 2 A proportionality coefficient Kp and parameter exemplified by the present invention A diagram of Bird; Figure 3 This is a schematic diagram of the process of load current passing through a quasi-PR controller according to the present invention; Figure 4 This is a schematic diagram illustrating the changes in environmental state and Q value according to the present invention; Figure 5 This is a schematic diagram of an embodiment of the device for adjusting the compensation parameters of an active filter based on Q-learning according to the present invention; Figure 6 This is a schematic diagram of an embodiment of a device for adjusting the compensation parameters of an active filter based on Q-learning according to the present invention. Detailed Implementation

[0018] The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0019] For ease of understanding, the specific process of the embodiments of the present invention is described below, see [link to documentation]. Figure 1 The first aspect of this invention provides a method for adjusting the compensation parameters of an active power filter based on Q-learning. This method can be operated within a power grid system with an active power filter. The power grid system includes a current transformer for acquiring grid current, a controller containing a digital signal processor or an FPGA programmable logic chip, and a power inverter unit for performing harmonic compensation. The method for adjusting the compensation parameters of the active power filter based on Q-learning includes: S100: Obtain the load current in the power grid and the compensation output current of the active power filter (APF). In this invention, after the system starts, the controller samples the load current in real time at high frequency through an external current transformer (CT) and samples the compensation output current of the APF itself through an internal current transformer (CT).

[0020] S200. Based on the load current analysis, harmonic characteristics are obtained and defined as the current environmental state. In this step, the DSP in the controller performs a Fast Fourier Transform (FFT) algorithm on the acquired load current signal to accurately analyze the amplitude A of each harmonic (such as the 5th, 7th, 11th, 13th, and other key harmonics). _n and phase φ _n In an optional embodiment of the first aspect of the present invention, obtaining harmonic features based on the load current analysis includes: extracting the amplitude and phase information of each harmonic in the load current using a fast Fourier transform algorithm as the harmonic features. The obtained harmonic feature information together constitutes a quantitative description of the current power grid environment. In the present invention, the combination of these information is defined as a multidimensional environmental state vector s (i.e., the current environmental state), for example, s={A5, φ5, A7, φ7, ...}. The degree of discretization of the environmental state vector s can be set according to the actual accuracy requirements.

[0021] S300, The adjustment amount of the control parameter to be adjusted in the active filter is defined as an action. In this invention, the control parameter to be adjusted includes the resonant gain parameter and resonant bandwidth parameter for a specific harmonic in the quasi-PR controller, the proportional coefficient and integral coefficient of the PI control loop, and the amplitude and phase of the reference harmonic current used to generate the compensation current.

[0022] Specifically, in this invention, "action" refers to making small, discrete adjustments to one or more key control parameters within the APF (Active Power Filter). These parameters collectively determine the compensation performance of the APF. In an optional embodiment of this invention, the action space may include, but is not limited to, the adjustment of the following parameters: Quasi-PR controller parameter adjustment: For each nth harmonic that needs compensation, the quasi-PR controller has a resonant gain parameter K. r_n and bandwidth parameter ω c_n The action can be to K r_n or ω c_n To increase or decrease by one unit step.

[0023] PI controller parameter adjustment: For PI controllers with DC-side voltage loop or current inner loop, the action can be adjusted by adjusting the proportional coefficient K. p Or the integral coefficient K i To increase or decrease by one unit step.

[0024] Reference current signal fine-tuning: The compensation reference current synthesized from the harmonic information extracted by FFT (Fast Fourier Transform) can be fine-tuned in terms of its total amplitude or phase.

[0025] S400. Under the current environmental conditions, the action is performed to adjust the control parameter to be adjusted, and a Q value for evaluating the merits of the compensation strategy is calculated based on the difference between the load current and the compensated output current. In this invention, the Q value is calculated using the following formula: Q = 1 / abs(I _load -I _inv ), where I _load I represents the instantaneous or effective value of the load current. _inv This refers to the instantaneous or effective value of the compensated output current.

[0026] Specifically, under the current environmental conditions, the control algorithm will select an action from the action set (initially randomly or based on a greedy strategy), for example, selecting the action "to reduce the K of the 5th harmonic". r_5 After increasing the value by 0.1, the controller applies the adjusted parameters to the control loop. After a very short settling time (such as one or several power frequency cycles), the system collects I again. _load and I _inv And calculate the compensated error e=I _load -I _inv Next, the system calculates the Q value, which is used to evaluate the effectiveness of the action just performed under the current environmental conditions. The Q value is preferably defined as the reciprocal of the integer absolute value of the residual current, i.e., Q = 1 / abs(I _load -I _inv The formula intuitively demonstrates that the smaller the compensation error, the larger the Q value obtained.

[0027] S500. Using the Q-learning algorithm, update the Q-table storing the correspondence between the current environment state, the action, and the Q-value according to the Q-value. In this invention, the Q-learning algorithm updates the Q-table using the following update formula: Q(s,a)=Q(s,a)+α×[R+γ×max(Q(s',a'))-Q(s,a)], where s is the current environment state, a is the current action, α is the learning rate, R is the reward value calculated based on the current Q-value, γ is the discount factor, s' is the next environment state, and a' is all possible actions for the next environment state.

[0028] Specifically, in this invention, the system maintains a Q-table in memory. Each row of the table represents a discretized environment state, each column represents an executable action, and the value Q in the table is the long-term expected reward of performing the action in the given environment state. After calculating the Q-value of the action in step S300, the system uses the core update formula of Q-learning to update the corresponding value in the Q-table. The update formula is Q(s,a)=Q(s,a)+α×[R+γ×max(Q(s',a'))-Q(s,a)], where s is the current environment state, a is the current action, α is the learning rate, R is the reward value calculated based on the current Q-value, γ is the discount factor (e.g., 0.9), s' is the next environment state, and a' is all possible actions in the next environment state. By repeatedly going through steps S100 to S500, the APF (Active Filter) continuously tries new actions and uses known optimal actions during operation. The data in the Q table will gradually converge due to continuous iterative updates, ultimately reflecting the optimal parameter adjustment strategy under various harmonic states.

[0029] S600. Based on the updated Q-table, under the current environmental state, the control parameters corresponding to the action that maximizes the Q-value are selected for the control loop of the active filter to achieve optimal harmonic compensation. In this invention, after a period of online learning, the Q-table gradually matures and enters the APF (Active Filter for Power) utilization stage. When the system detects a new environmental state, it no longer tries randomly but directly queries the Q-table. In the row corresponding to the environmental state, it finds the action with the maximum Q-value, and the system then executes the parameter adjustment corresponding to that action. Since this action is the optimal action learned from historical experience, the APF can quickly and automatically adjust to the optimal operating point to achieve optimal compensation for the current harmonics. The entire process requires no manual intervention.

[0030] To better understand the technical solution of this invention, in conjunction with Figures 2-4 The quasi-PR controller transfer function is as follows: As an example, the effect of a quasi-PR controller is mainly affected by the proportional coefficient K. p and parameters ,parameter Influence, The parameters are proportional to the peak gain of the controller, while the parameters It not only affects the controller's gain, but also the bandwidth of the controller's cutoff frequency. With the increase of [value], both the controller gain and bandwidth will increase (fundamental frequency gain is [value]). constant).

[0031] Increase This will simultaneously amplify harmonics at other frequencies, polluting the controller output and increasing... This will cause drastic changes in the output, affecting the PI loop. After the load current passes through the quasi-PR controller, only the corresponding harmonics remain.

[0032] Increase This will simultaneously amplify harmonics at other frequencies, polluting the controller output. For example, increasing the quasi-PR of the second harmonic... The increased second harmonic gain, due to the increased gain bandwidth, will also increase the third harmonic current gain. The other harmonics will experience similar increases. This will increase the gain of the harmonic, but at the same time it will cause drastic changes in the amplitude of the output result, affecting the PI loop.

[0033] Q-learning can learn by analyzing each harmonic... , The change in current is expressed as the reciprocal of the absolute value of the difference between the output current and the load current, where Q = 1 / abs(I). _load -I _inv The smaller the error value, the larger the Q value, and the better the compensation strategy. Each Nth harmonic corresponds to the environmental state s. n (ω c_n , K r_n ), ω c_n With K r_n Change to action a n Then calculate the Q of each harmonic. n Values, and based on each harmonic Q n The optimal strategy for determining the value size (when each Q) n When the mean is at its highest, adjust the parameters.

[0034] Q-learning can utilize the K-axis of the PI ring. p k i The change is expressed as the reciprocal of the absolute value of the difference between the output current and the load current, where Q = 1 / abs(I). _load -I _inv Each value corresponds to the environment state s(K). p K i ), K p K i The change is to action a, then the total Q value is calculated, and the optimal strategy is determined based on the size of the Q value, and the parameters are adjusted accordingly.

[0035] In summary, this invention, by introducing the Q-learning algorithm, transforms the parameter adjustment process of an APF (Active Power Filter) into a Markov decision process involving interactive learning between an agent and the environment, endowing the APF with unprecedented adaptive adjustment and autonomous optimization capabilities. This method can effectively cope with complex operating conditions, significantly improve harmonic compensation effects and system robustness, and has extremely high industrial application value.

[0036] See Figure 5 The second aspect of the present invention provides an apparatus for adjusting the compensation parameters of an active filter based on Q-learning, the apparatus comprising: The current acquisition module 10 is used to acquire the load current in the power grid and the compensation output current of the active filter. The state parameter definition module 20 is used to obtain harmonic characteristics based on the load current analysis and define the harmonic characteristics as the current environmental state; Action parameter definition module 30 is used to define the adjustment amount of the control parameter to be adjusted in the active filter as an action; Q-value calculation module 40 is used to perform the action to adjust the control parameter to be adjusted under the current environmental state, and calculate the Q-value for evaluating the merits of the compensation strategy based on the difference between the load current and the compensation output current. Q-table update module 50 is used to update the Q-table storing the current environment state, the correspondence between the action and the Q-value using the Q-learning algorithm; The harmonic compensation module 60 is used to select, based on the updated Q table and under the current environmental condition, the control parameters corresponding to the action that maximizes the Q value for use in the control loop of the active filter, so as to achieve optimal harmonic compensation.

[0037] In an optional embodiment of the second aspect of the present invention, the state parameter definition module includes: a harmonic feature calculation unit, used to extract the amplitude and phase information of each harmonic in the load current through a fast Fourier transform algorithm, as the harmonic feature.

[0038] In an optional embodiment of the second aspect of the invention, the Q value is calculated using the following formula: Q = 1 / abs(I _load -I _inv ), where I _load I represents the instantaneous or effective value of the load current. _inv This refers to the instantaneous or effective value of the compensated output current.

[0039] In an optional embodiment of the second aspect of the invention, the control parameters to be adjusted include the resonant gain parameter and the resonant bandwidth parameter for a specific subharmonic in the quasi-PR controller.

[0040] In an optional embodiment of the second aspect of the present invention, the control parameter to be adjusted further includes the proportional coefficient and integral coefficient of the PI control loop.

[0041] In an optional embodiment of the second aspect of the invention, the control parameters to be adjusted further include the amplitude and phase of a reference harmonic current for generating the compensation current.

[0042] In an optional embodiment of the second aspect of the invention, the Q-learning algorithm updates the Q-table using the following update formula: Q(s,a)=Q(s,a)+α×[R+γ×max(Q(s',a'))-Q(s,a) ], where s is the current environment state, a is the current action, α is the learning rate, R is the reward value calculated based on the current Q value, γ is the discount factor, s' is the next environment state, and a' is all possible actions in the next environment state.

[0043] Figure 6 This is a schematic diagram of a device for adjusting active filter compensation parameters based on Q-learning, provided in an embodiment of the present invention. This device can vary significantly depending on its configuration or performance, and may include one or more processors 70 (central processing units, CPUs) (e.g., one or more processors) and memory 80, and one or more storage media 90 (e.g., one or more mass storage devices) for storing application programs or data. The memory and storage media can be temporary or persistent storage. The program stored in the storage media may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the device for adjusting active filter compensation parameters based on Q-learning. Furthermore, the processor may be configured to communicate with the storage media and execute the series of instruction operations in the storage media on the device for adjusting active filter compensation parameters based on Q-learning.

[0044] The device for adjusting the compensation parameters of an active filter based on Q-learning, as described in this invention, may further include one or more power supplies 100, one or more wired or wireless network interfaces 110, one or more input / output interfaces 120, and / or one or more operating systems, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will understand that... Figure 6The illustrated device structure for adjusting active filter compensation parameters based on Q-learning does not constitute a limitation on devices for adjusting active filter compensation parameters based on Q-learning. It may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.

[0045] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when the instructions are executed on a computer, cause the computer to perform the steps of the method for adjusting the compensation parameters of an active filter based on Q-learning.

[0046] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system or system / unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0047] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0048] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for adjusting compensation parameters of an active filter based on Q-learning, characterized in that, The method comprises: obtaining a load current in a power grid and a compensation output current of an active filter; analyzing a harmonic feature based on the load current, and defining the harmonic feature as a current environment state; defining an adjustment amount of a control parameter to be adjusted in the active filter as an action; under the current environment state, performing the action to adjust the control parameter to be adjusted, and calculating a Q value for evaluating the compensation strategy based on a difference between the load current and the compensation output current; updating a Q table storing a correspondence between the current environment state, the action and the Q value according to the Q-learning algorithm and the Q value; selecting a control parameter corresponding to the action with the maximum Q value under the current environment state S according to the updated Q table, and using the control parameter in a control loop of the active filter to achieve optimal compensation of harmonics.

2. The method for adjusting compensation parameters of active filter based on Q-learning according to claim 1, characterized in that, The harmonic feature is obtained by analyzing the load current, and the harmonic feature is extracted by a fast Fourier transform algorithm to obtain amplitude and phase information of each harmonic in the load current.

3. The method for adjusting compensation parameters of active filter based on Q-learning according to claim 1, characterized in that, The Q value is calculated by the following formula: Q = 1 / abs(I _load -I _inv ), wherein I _load is the instantaneous value or effective value of the load current, and I _inv is the instantaneous value or effective value of the compensation output current.

4. The method for adjusting compensation parameters of active filter based on Q-learning of claim 1, wherein, The control parameter to be adjusted includes a resonance gain parameter and a resonance bandwidth parameter of a specific harmonic in a quasi-PR controller.

5. The method for adjusting compensation parameters of active filter based on Q-learning according to claim 4, characterized in that, The control parameter to be adjusted also includes a proportional coefficient and an integral coefficient of a PI control loop.

6. The method for adjusting compensation parameters of active filter based on Q-learning according to claim 5, characterized in that, The control parameter to be adjusted also includes an amplitude and a phase of a reference harmonic current used to generate a compensation current.

7. The method for adjusting compensation parameters of active filter based on Q-learning of claim 1, wherein, The Q-learning algorithm updates the Q table according to the following update formula: Q(s,a)=Q(s,a)+α×[R+γ×max(Q(s’,a’))-Q(s,a) ], wherein s is a current environment state, a is a current action, α is a learning rate, R is a reward value calculated based on a current Q value, γ is a discount factor, s' is a next environment state, and a' is all possible actions of the next environment state.

8. A device for adjusting the compensation parameters of an active filter based on Q-learning, characterized in that, The device for adjusting compensation parameters of an active filter based on Q-learning comprises: a current acquisition module for obtaining a load current in a power grid and a compensation output current of an active filter; a state parameter definition block for analyzing a harmonic feature based on the load current, and defining the harmonic feature as a current environment state; an action parameter definition module for defining an adjustment amount of a control parameter to be adjusted in the active filter as an action; a Q value calculation module for, under the current environment state, performing the action to adjust the control parameter to be adjusted, and calculating a Q value for evaluating the compensation strategy based on a difference between the load current and the compensation output current; a Q table updating module for updating a Q table storing a correspondence between the current environment state, the action and the Q value according to the Q-learning algorithm and the Q value; a harmonic compensation module for selecting a control parameter corresponding to the action with the maximum Q value under the current environment state S according to the updated Q table, and using the control parameter in a control loop of the active filter to achieve optimal compensation of harmonics.

9. A device for adjusting the compensation parameters of an active filter based on Q-learning, characterized in that, The device for adjusting compensation parameters of an active filter based on Q-learning comprises a memory and at least one processor, the memory having instructions stored therein, the memory and the at least one processor being interconnected by a line; The at least one processor invokes the instructions in the memory to cause the device for adjusting compensation parameters of an active filter based on Q-learning to perform the method for adjusting compensation parameters of an active filter based on Q-learning according to any one of claims 1-7.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by a processor, implements the method for adjusting compensation parameters of an active filter based on Q-learning according to any one of claims 1-7.