Electric energy quality regulation and control method, device, equipment, medium and program product

By combining a multi-agent deep reinforcement learning model with a data processing module, the problem of weak collaboration in decentralized governance methods in low-voltage distribution networks is solved, enabling the optimization and real-time control of power quality across the entire network and improving the overall power quality effect.

CN120933978APending Publication Date: 2025-11-11BEIJING SMARTCHIP MICROELECTRONICS TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511116608.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing decentralized governance methods have weak collaboration capabilities and limited global performance in low-voltage distribution networks, making it difficult to guarantee optimal power quality across the entire network and resulting in poor practical application effects.

Method used

A multi-agent deep reinforcement learning model is adopted. Through a centralized training-distributed execution framework, Markov decision processes are used for power quality regulation. Combined with Bayesian inference and adversarial network generation layers to process data, multi-agent collaborative operation is achieved.

Benefits of technology

It effectively improves the power quality of low-voltage distribution networks, and has the advantages of no communication required, strong real-time performance and no reliance on accurate power flow models, thereby enhancing the overall network's power quality optimization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120933978A_ABST
    Figure CN120933978A_ABST
Patent Text Reader

Abstract

The invention discloses an electric energy quality regulation and control method, device and equipment, a medium and a program product. The method comprises the following steps: obtaining power distribution network observation data of a regulation and control node; inputting the observation data of the power distribution network into the multi-agent deep reinforcement learning model to obtain a power quality regulation action; wherein the multi-agent deep reinforcement learning model makes a decision based on a Markov decision process; the multi-agent deep reinforcement learning model is obtained through training by the regulation and control center according to empirical values determined by the agents based on local target sample data; and executing an electric energy quality regulation and control action on the power distribution network where the regulation and control node is located. A power quality regulation and control problem is converted into a Markov decision problem, a centralized training-decentralized execution framework is adopted, a control strategy is continuously optimized in continuous interaction between local observation data of each agent and a power distribution network, and accurate value estimation is provided according to the global evaluation capability of a regulation and control center. And multi-agent cooperative operation is realized while local observation is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power distribution Internet of Things (IoT) technology, and in particular to a power quality control method, device, equipment, medium, and program product. Background Technology

[0002] With the large-scale integration of distributed photovoltaic power and new loads, power quality problems in low-voltage distribution networks are becoming increasingly prominent, mainly manifested as voltage deviation, harmonic distortion, and three-phase imbalance. These problems not only affect users' electricity experience but may also cause equipment to disconnect from the grid, resulting in economic losses.

[0003] Traditional centralized or distributed control methods rely on accurate power flow models and strong communication capabilities, making them ill-suited to the characteristics of low-voltage distribution networks, such as weak communication conditions, limited equipment computing power, and complex and diverse power quality issues. Existing decentralized governance methods suffer from weak collaboration capabilities, limited global performance, and complete reliance on local information, making it difficult to guarantee optimal power quality across the entire network and resulting in poor practical application effects in low-voltage distribution networks. Summary of the Invention

[0004] This invention provides a power quality control method, device, equipment, medium, and program product to solve the problems of weak collaboration capabilities, limited global performance, complete reliance on local information, difficulty in ensuring optimal power quality across the entire network, and poor practical application effect in low-voltage distribution networks.

[0005] In a first aspect, embodiments of the present invention provide a power quality control method applied to a power quality control system, the power quality control system comprising a control center and multiple control nodes deploying intelligent agents; each of the intelligent agents is respectively deployed with a trained multi-agent deep reinforcement learning model; the method is executed by the control nodes, and the method includes:

[0006] Obtain the distribution network observation data of the control node;

[0007] The power distribution network observation data is input into the multi-agent deep reinforcement learning model to obtain power quality control actions; wherein, the multi-agent deep reinforcement learning model makes decisions based on Markov decision processes; the multi-agent deep reinforcement learning model is trained by the control center based on the experience values ​​determined by each agent based on local target sample data;

[0008] The power quality control action is performed on the distribution network where the control node is located.

[0009] Secondly, embodiments of the present invention provide a power quality control device applied to a power quality control system, the power quality control system including a control center and multiple control nodes deploying intelligent agents; each of the intelligent agents is respectively deployed with a trained multi-agent deep reinforcement learning model; the device is deployed in the control node, the device comprising:

[0010] The observation data acquisition module is used to acquire the distribution network observation data of the control node;

[0011] The control action determination module is used to input the distribution network observation data into the multi-agent deep reinforcement learning model to obtain power quality control actions; wherein, the multi-agent deep reinforcement learning model makes decisions based on Markov decision processes; the multi-agent deep reinforcement learning model is trained by the control center based on the experience values ​​determined by each agent based on local target sample data;

[0012] The power quality control module is used to perform the power quality control action on the distribution network where the control node is located.

[0013] Thirdly, embodiments of the present invention provide an electronic device, the electronic device comprising:

[0014] At least one processor;

[0015] and a memory communicatively connected to the at least one processor;

[0016] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the power quality control method according to any embodiment of the present invention.

[0017] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing computer instructions, which are used to cause a processor to execute the power quality control method described in any embodiment of the present invention.

[0018] The technical solution of this invention acquires distribution network observation data from control nodes; inputs the distribution network observation data into a multi-agent deep reinforcement learning model to obtain power quality control actions; wherein, the multi-agent deep reinforcement learning model makes decisions based on Markov decision processes; the multi-agent deep reinforcement learning model is trained by the control center based on the experience values ​​determined by each agent based on local target sample data; and executes power quality control actions on the distribution network where the control nodes are located. By transforming the power quality control problem into a Markov decision problem, a centralized training-distributed execution framework is adopted. It continuously optimizes the control strategy through the interaction between each agent's local observation data and the distribution network, and provides accurate value estimates based on the control center's global evaluation capabilities. While ensuring local observation, it achieves multi-agent collaborative operation, solving the problems of weak collaboration, limited global performance, complete reliance on local information, difficulty in guaranteeing optimal power quality across the entire network, and poor practical application effects in low-voltage distribution networks in existing decentralized governance methods. It effectively improves the power quality of low-voltage distribution networks and has the advantages of no communication required, strong real-time performance, and independence from accurate power flow models.

[0019] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart of a power quality control method provided in Embodiment 1 of the present invention;

[0022] Figure 2 This is a flowchart of a power quality control method provided in Embodiment 2 of the present invention;

[0023] Figure 3 This is a schematic diagram of the structure of a power quality control device provided in Embodiment 3 of the present invention;

[0024] Figure 4 A schematic diagram of the structure of an electronic device for implementing the power quality control method of this invention. Detailed Implementation

[0025] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0026] It should be noted that the terms "comprising" and "having" and any variations thereof in the specification, claims and accompanying drawings of this invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such process, method, product or device.

[0027] Example 1

[0028] Figure 1 This is a flowchart of a power quality control method provided in Embodiment 1 of the present invention. This embodiment is applicable to the control of power quality in distribution networks. The method can be executed by a power quality control device, which can be implemented in hardware and / or software. The power quality control method is applied to a power quality control system, which includes a control center and multiple control nodes (e.g., k control nodes, k≥1) deploying intelligent agents. Each intelligent agent is equipped with a pre-trained multi-agent deep reinforcement learning model. The method is executed by the control nodes. The power quality control device can be configured at the control nodes in the power quality control system, and the control nodes can exist in the form of electronic devices.

[0029] like Figure 1 As shown, the method includes:

[0030] S110. Obtain distribution network observation data from the control nodes.

[0031] In this context, the control node can be considered a key component playing a crucial role in the power quality control of the distribution network. Distribution network observation data can be considered as the distribution network data observed by the control node in the local distribution network it controls, which may include fundamental voltage, harmonic voltage, current data, frequency data, and power data.

[0032] Specifically, control nodes can measure distribution network observation data by connecting to measuring devices in the distribution network, or they can monitor equipment in the distribution network and collect distribution network observation data through Supervisory Control and Data Acquisition (SCADA).

[0033] S120. Input the distribution network observation data into the multi-agent deep reinforcement learning model to obtain power quality control actions; wherein, the multi-agent deep reinforcement learning model makes decisions based on the Markov decision process; the multi-agent deep reinforcement learning model is trained by the control center based on the experience values ​​determined by each agent based on local target sample data.

[0034] Among them, the Multi-Agent Deep Reinforcement Learning (MADRL) model combines the representational capabilities of deep learning with the distributed decision-making advantages of multi-agent systems. It enables agents to learn to make decisions in complex environments and continuously optimize these decisions through interaction with the environment. In deep reinforcement learning, the control center trains the MADRL model based on the empirical values ​​determined by each agent using local target sample data. The agents then choose appropriate actions based on the state of the environment to achieve the expected goals and continuously adjust and improve their actions by observing environmental feedback. Local target sample data can be considered as sample data composed of distribution network data collected by the control nodes within the controlled distribution network.

[0035] In this embodiment, the multi-agent deep reinforcement learning model adopts the centralized training distributed execution (CTDE) paradigm, modeling the distribution network regulation problem as a partially observable Markov decision process (POMDP). The Markov decision process can be represented by a six-tuple (S, [O... i,t ] k [A] t ] k [R] t ] k , P, [F i ] k ) represents the state space, [O i,t ] k Let [A] represent the observation space of the i-th agent among k agents at time t. t ] k Let R be the joint action space of k agents at time t. t ] k Let P represent the reward function of k agents at time t, and let P be the state transition function of the multi-agent fusion environment. i ]k Let be the action function of the i-th agent among k agents at time t.

[0036] In power quality control in low-voltage distribution networks, the specific meanings of each element in the Markov decision process are as follows: (1) State space (S): used to represent the state matrix of the distribution network, which includes the three-phase active and reactive power matrix, the three-phase fundamental and harmonic voltage matrix, and the three-phase fundamental and harmonic voltage phase angle matrix of each control node at time t. (2) Joint observation space ([O t ] K ): Composed of the local observation matrices of each agent, the local observation matrix includes the observation values ​​of the agents in each control node at time t [o i,t ] K (e.g., three-phase active power and three-phase reactive power, three-phase fundamental voltage and three-phase harmonic voltage (including amplitude and phase angle)) and the operating state of each control node at time t-1 [a i,t-1 ] K (3) Joint action space ([A) t ] K Since the actions of each agent are independent of each other, a joint action space is used to represent the joint actions of all agents in order to describe the dynamic interaction process between each agent and the environment. (4) Reward function ([R t ] K ): To ensure that voltage deviation, harmonic distortion and three-phase imbalance are effectively suppressed, strong penalties are applied outside the power quality limits and weak penalties are applied within the power quality limits, so that each agent can provide sufficient operational safety margin for the low-voltage distribution network while ensuring that each power quality index meets the constraint requirements. In addition, since the power quality regulation problem is carried out in a fully cooperative manner, each agent shares the same reward so that single-phase agents can effectively perceive the changes in the power quality index of different phases. (5) State transition function (P): Represents the probability of performing the current action in the current state to transition to the state at the next moment. (6) Action function ([F i ] k ): For an agent located at a single control node, it is used to determine the action taken by the agent under the current observation.

[0037] Specifically, the agent in each control node will input the locally observed distribution network observation data into the multi-agent deep reinforcement learning model deployed within the agent. After passing through the action network and value network in the multi-agent deep reinforcement learning model, the agent will output power quality control actions.

[0038] S130. Perform power quality control actions on the distribution network where the control node is located.

[0039] Specifically, each control node executes power quality control actions output by the multi-agent deep reinforcement learning model on the distribution network it is located in, thereby realizing power quality control and management of the distribution network.

[0040] The technical solution of this invention involves acquiring distribution network observation data from control nodes; inputting the distribution network observation data into a multi-agent deep reinforcement learning model to obtain power quality control actions; wherein, the multi-agent deep reinforcement learning model makes decisions based on Markov decision processes; the multi-agent deep reinforcement learning model is trained by the control center based on the empirical values ​​determined by each agent based on local target sample data; and power quality control actions are executed on the distribution network where the control nodes are located. By transforming the power quality control problem into a Markov decision problem, a centralized training-distributed execution framework is adopted. The control strategy is continuously optimized through the interaction between each agent's local observation data and the distribution network, and accurate value estimates are provided based on the control center's global evaluation capabilities. This achieves multi-agent collaborative operation while ensuring local observation, effectively improving the power quality of low-voltage distribution networks. It also has advantages such as no communication required, high real-time performance, and independence from accurate power flow models.

[0041] Example 2

[0042] Figure 2 This is a flowchart of a power quality control method according to Embodiment 2 of the present invention. Based on the above embodiments, this embodiment further refines the training steps of the multi-agent deep reinforcement learning model. Specifically, the multi-agent deep reinforcement learning model includes an action network and a value network; the training steps of the multi-agent deep reinforcement learning model include: each agent in the control node inputs local target sample data into the action network to obtain experience values; the experience values ​​are stored in an experience replay buffer; the control center randomly selects a batch of experience values ​​from the experience replay buffer, calculates action gradients based on the selected experience values, and iteratively adjusts the weight parameters of the action network based on the action gradients; the loss function value is calculated based on the selected experience values, and the weight parameters of the value network are iteratively adjusted based on the loss function value.

[0043] like Figure 2 As shown, the method includes:

[0044] S210. Train a multi-agent deep reinforcement learning model and deploy it on each control node; wherein, the multi-agent deep reinforcement learning model makes decisions based on a Markov decision process; the multi-agent deep reinforcement learning model is trained by the control center based on the experience values ​​determined by each agent based on local target sample data; the multi-agent deep reinforcement learning model includes an action network and a value network.

[0045] The value network comprises multiple fully connected layers, ReLU activation layers, and one fully connected layer. The feature vector input to the value network can include: the observation value at the current time step [o...]. i,t ] K and action value [a i,t ] K And the action value at the next moment [a] i,t+1 ] K The structure is composed of an output feature vector that represents the action value.

[0046] In this embodiment, the control center trains the network parameters of the action network and the value network based on the experience values ​​determined by each agent based on local target sample data, thus obtaining a multi-agent deep reinforcement learning model. The multi-agent deep reinforcement learning model improves its strategies through interactive training of multiple agents. The multi-agent deep reinforcement learning model includes an action network and a value network. The action network generates corresponding actions based on the observations of each local node. i,t ] K The goal is to maximize the value of the action function. The value function is accurately estimated by the control center based on the joint observations and joint actions of all agents.

[0047] In an optional embodiment, the training steps of the multi-agent deep reinforcement learning model include:

[0048] A. The agent in each of the control nodes inputs the sample data of its node into the action network to obtain an experience value; and stores the experience value in the experience playback buffer.

[0049] The Experience Replay Buffer is a key component in deep reinforcement learning, used to store and reuse historical experiences of the agent's interactions with the environment to improve training efficiency and stability.

[0050] Specifically, for each control node, the agent acquires historical data from the local power distribution network as sample data, inputs the sample data into the action network, obtains the empirical value output by the action network, and stores the empirical value in the experience replay buffer. It can be understood that the experience replay buffer stores the empirical values ​​obtained by each agent at multiple different times.

[0051] Optionally, the step of inputting local target sample data into the action network to obtain experience value includes: inputting the local target sample data into the action network to obtain the current action with the highest action value selected by the action network based on the current state in the local target sample data; determining the current reward and target state obtained by executing the current action; and determining the experience value composed of the current state, current action, current reward, and target state.

[0052] The action network comprises multiple fully connected layers, a ReLU activation layer, and a boundary constraint layer. The feature vectors input to each feature layer of the action network can include: the maximum power, capacity information, fundamental frequency information, and harmonic information of each control node at the current moment. It can also incorporate the actions from the previous moment, giving the action network temporal characteristics and optimizing the temporal decision-making process.

[0053] Among them, the experience value can include: the current state (s) t ), current action (a t ), current reward (r) t ) and target state (s t+1 A quadruple. The current state can be considered as the state of the distribution network observed by the control node at the current moment. The current action can be considered as the action selected from the action space in the current state. The current reward can be considered as the immediate reward obtained by performing the selected action in the current state, indicating the quality of the action. The optimization objective of the power quality control system is to optimize power quality. The current reward can be determined based on power quality assessment indicators, such as a single power quality assessment data point or a weighted value of multiple power quality assessment data points. The target state can be considered as the state transitioning to the next moment by performing the selected current action in the current state.

[0054] Specifically, for each agent, the agent's local target sample data is input into the action network. The action network determines executable actions based on the current state in the local target sample data, and determines the action value of each executable action. The executable action with the highest action value is selected as the current action. The network also determines the current reward obtained by the agent in the current state from performing the current action, and the target state to which the current state transitions. Finally, an experience value is determined, consisting of the current state, current action, current reward, and target state.

[0055] In this embodiment, each control node outputs experience values ​​based on local target sample data and local action network and stores them in the experience replay buffer, providing experience for the control center to train the action network and value network.

[0056] B. The control center randomly selects a batch of experience values ​​from the experience playback buffer, calculates the action gradient based on the selected experience values, and iteratively adjusts the weight parameters of the action network based on the action gradient; it calculates the loss function value based on the selected experience values, and iteratively adjusts the weight parameters of the value network based on the loss function value.

[0057] The action gradient can be considered as the gradient of the action network.

[0058] In this embodiment, the control center randomly selects a batch of experience values ​​from the experience replay buffer to train the action network and value network in the multi-agent deep reinforcement learning model. Specifically, the action gradient of the action network is calculated based on the selected experience values, and the weight parameters of the action network are adjusted based on the action gradient to update the action network. Furthermore, the loss function value of the value network is calculated based on the selected experience values, and the weight parameters of the value network are adjusted based on the loss function value.

[0059] Optionally, the step of calculating the action gradient based on the extracted empirical values ​​and iteratively adjusting the weight parameters of the action network based on the action gradient includes: calculating an objective function based on the extracted empirical values, wherein the objective function is the expected cumulative reward; calculating the action gradient of the objective function; and iteratively updating the weight parameters of the action network based on the action gradient using the gradient ascent method.

[0060] Among them, the expected cumulative reward can be considered as the expectation of the cumulative reward.

[0061] For example, the objective function is defined as the expected cumulative reward under the action, such as...

[0062]

[0063] Where τ=(s0, a0, r0, s1, ......s t a t r t s t+1 ), where γ∈[0,1) is the discount factor. The goal is to maximize J(θ) by adjusting the policy parameter θ.

[0064] According to the policy gradient theorem, the gradient of the objective function is:

[0065]

[0066] Due to p(s0) and state transition p(s) t+1 |s t a t (Irrelevant to θ, only the action term π is retained) θ (a t |s t ):

[0067]

[0068] Replace R(τ) with the cumulative reward Q starting from time t. π (s t |a t ):

[0069]

[0070] Update the weight parameters θ of the action network along the gradient ascent direction: α is the learning rate, which controls the update step size. If the action has high value, gradient update increases the probability of that action; if the action has low value, gradient update suppresses the probability of that action.

[0071] Optionally, calculating the loss function value based on the extracted experience value includes: calculating the current action value based on the current state and current action in the extracted experience value; calculating the action value of all actions corresponding to the target state in the extracted experience value; calculating the target action value based on the maximum value among all action values ​​and the reward in the extracted experience value; and calculating the loss function value between the target action value and the current action value.

[0072] The current action value can be considered as the value of the current action. The target action value can be considered as the value of the target action.

[0073] Specifically, the value of the current action is determined as follows: Here are the weight parameters of the current value network. The current action value is used to evaluate the reward for performing the current action in the current state. The target action value is determined as follows: Here are the parameters of the target value network, and a′∈A represents the target state s. t+1 The corresponding action, A is the target state s t+1 The corresponding action set is generated by the action network. The target action value Q is... target and the current action value Q current The loss function value can be represented by the mean squared error, for example:

[0074]

[0075] In this embodiment, by separating the current value network and the target value network and iteratively updating based on the loss function value, the deep reinforcement learning model can stably approximate the optimal action value function.

[0076] In another optional embodiment, the multi-agent deep reinforcement learning model further includes: a data processing module; the data processing module includes: a Bayesian inference layer and an adversarial network generation layer; the data processing module processes the input local initial sample data to obtain high-confidence local target sample data; the local initial sample data includes missing values ​​and noise;

[0077] The Bayesian inference layer uses a Bayesian probability matrix factorization model to estimate the posterior distribution of missing values ​​in the local initial sample data, and samples imputation values ​​from the local initial sample data to fill the missing values ​​according to the posterior distribution; noise separation is performed on the imputed local initial sample data to obtain local denoised sample data; the adversarial network generation layer uses the posterior distribution as a condition to generate high-confidence local target sample data based on the local denoised sample data.

[0078] Among them, the initial local sample data can be considered as unprocessed local sample data. Local sample data can be considered as sample data based on local distribution network data at the control node. Local denoised sample data can be considered as denoised local sample data. Local target sample data can be considered as processed local sample data.

[0079] Specifically, in the Bayesian inference layer, assuming the local initial sample data matrix is... The Bayesian Probabilistic Matrix Factorization (BPMF) model is X≈UV. T For the latent factor matrix, and It follows the Wishart prior. Based on Markov Monte Carlo, it approximates the posterior through Gibbs sampling, generating multiple sets of filler values ​​from the posterior distribution. Imputing missing values ​​yields the local initial sample data X after imputation. fill The padded local initial sample data can be represented as: X fill =Y + N, where N is noise. Noise can be filtered using wavelet algorithms to obtain the local denoised sample data Y. Local target sample data conforming to the posterior distribution is generated based on the adversarial network generation layer. The adversarial network generation layer includes a generator and a discriminator. The generator can be represented as G(z|Y, σ), where the noise Z ~ N(0, 1), and σ is the uncertainty, which can be represented by the standard deviation of the posterior distribution. The discriminator can be represented as D(X, σ), and the discriminator weights can be set as follows: Finally, the output of the adversarial network generation layer was used as the local target sample data with high confidence.

[0080] This embodiment takes into account the potential for missing data or noise in actual power distribution network applications. Based on the multi-agent deep reinforcement learning model, a data processing module consisting of a Bayesian inference layer and an adversarial network generation layer is added to preprocess the data, thereby improving data quality and thus improving the accuracy of determining power quality control actions.

[0081] S220. Obtain distribution network observation data from the control nodes.

[0082] S230. Input the distribution network observation data into the multi-agent deep reinforcement learning model to obtain power quality regulation actions.

[0083] In this embodiment, the multi-agent deep reinforcement learning model includes a data processing module and an action network. The agent inputs local distribution network observation data into the data processing module, where it is processed sequentially by a Bayesian inference layer and an adversarial network generation layer to obtain high-confidence distribution network observation data. This high-confidence data is then input into the action network to obtain power quality control actions.

[0084] S240. Perform power quality control actions on the distribution network where the control node is located.

[0085] The technical solution of this invention involves training a multi-agent deep reinforcement learning model and deploying it at each control node. The multi-agent deep reinforcement learning model makes decisions based on a Markov decision process. The model is trained by the control center using empirical values ​​determined by each agent based on local target sample data. The model includes an action network and a value network. Distribution network observation data is acquired at the control nodes. This data is then input into the model to obtain power quality control actions. These actions are then executed on the distribution network where the control nodes are located. By transforming the power quality control problem into a Markov decision problem, and employing a centralized training-distributed execution framework, the control strategy of each agent is continuously optimized through its local observation data and the ongoing interaction with the distribution network. Accurate value estimates are provided based on the control center's global evaluation capabilities. This approach ensures coordinated operation of multiple agents while maintaining local observations, effectively improving the power quality of low-voltage distribution networks. It also offers advantages such as no communication required, high real-time performance, and independence from precise power flow models. Furthermore, a data processing module consisting of a Bayesian inference layer and an adversarial network generation layer was added to the multi-agent deep reinforcement learning model, which deeply integrates reinforcement learning agents with power quality governance and constructs an integrated "perception-decision-execution" control system, breaking through the limitations of existing methods in terms of real-time performance, coordination and adaptability.

[0086] Example 3

[0087] Figure 3 This is a schematic diagram of a power quality control device according to Embodiment 3 of the present invention. The device is applied to a control node in a power quality control system, which includes a control center and multiple control nodes deploying intelligent agents; each intelligent agent is equipped with a pre-trained multi-agent deep reinforcement learning model. Figure 3 As shown, the device includes: an observation data acquisition module 310, a control action determination module 320, and a power quality control module 330; wherein:

[0088] The observation data acquisition module 310 is used to acquire the distribution network observation data of the control node;

[0089] The control action determination module 320 is used to input the distribution network observation data into the multi-agent deep reinforcement learning model to obtain power quality control actions; wherein, the multi-agent deep reinforcement learning model makes decisions based on Markov decision processes; the multi-agent deep reinforcement learning model is trained by the control center based on the experience values ​​determined by each agent based on local target sample data;

[0090] The power quality control module 330 is used to perform the power quality control action on the distribution network where the control node is located.

[0091] Optionally, the multi-agent deep reinforcement learning model includes an action network and a value network; the training steps of the multi-agent deep reinforcement learning model include:

[0092] The agent in each of the control nodes inputs local target sample data into the action network to obtain experience values; the experience values ​​are then stored in the experience replay buffer.

[0093] The control center randomly selects a batch of experience values ​​from the experience replay buffer, calculates the action gradient based on the selected experience values, and iteratively adjusts the weight parameters of the action network based on the action gradient; it also calculates the loss function value based on the selected experience values ​​and iteratively adjusts the weight parameters of the value network based on the loss function value.

[0094] Optionally, the step of inputting local target sample data into the action network to obtain empirical values ​​includes:

[0095] The local target sample data is input into the action network to obtain the current action with the highest action value selected by the action network based on the current state in the local target sample data.

[0096] Determine the current reward and target state obtained by performing the current action;

[0097] Determine the experience value comprised of the current state, the current action, the current reward, and the target state.

[0098] Optionally, the step of calculating the action gradient based on the extracted empirical values ​​and iteratively adjusting the weight parameters of the action network based on the action gradient includes:

[0099] The objective function is calculated based on the extracted experience values, and the objective function is the expected cumulative reward.

[0100] Calculate the action gradient of the objective function;

[0101] The weight parameters of the action network are iteratively updated based on the action gradient using the gradient ascent method.

[0102] Optionally, calculating the loss function value based on the extracted empirical values ​​includes:

[0103] Calculate the value of the current action based on the current state and current action in the extracted experience values;

[0104] Calculate the action value of all actions corresponding to the target state in the extracted empirical values;

[0105] The target action value is calculated based on the maximum value among all action values ​​and the current reward from the extracted experience value.

[0106] Calculate the loss function value between the target action value and the current action value.

[0107] Optionally, the multi-agent deep reinforcement learning model further includes: a data processing module; the data processing module includes: a Bayesian inference layer and an adversarial network generation layer; the data processing module processes the input local initial sample data to obtain high-confidence local target sample data; the local initial sample data includes missing values ​​and noise;

[0108] The Bayesian inference layer uses a Bayesian probability matrix factorization model to estimate the posterior distribution of missing values ​​in the local initial sample data, and samples imputation values ​​from the local initial sample data to fill the missing values ​​according to the posterior distribution; noise separation is performed on the imputed local initial sample data to obtain local denoised sample data.

[0109] The adversarial network generation layer generates high-confidence local target sample data based on the local denoised sample data, conditioned on the posterior distribution.

[0110] The power quality control device provided in the embodiments of the present invention can execute the power quality control method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method.

[0111] Example 4

[0112] Figure 4 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0113] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0114] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0115] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as power quality control methods.

[0116] In some embodiments, the power quality control method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded into and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the power quality control method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the power quality control method by any other suitable means (e.g., by means of firmware).

[0117] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0118] In some embodiments, the power quality control method may be implemented as a computer program, which is implicitly included in a computer program product. When executed by a processor, the computer program implements the power quality control method of the present invention. The computer program product can be understood as a software product that primarily implements its solution through a computer program. The computer program used to implement the method of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on a machine, partially on a machine, partially on a remote machine as a standalone software package, or entirely on a remote machine or server.

[0119] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0120] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0121] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0122] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0123] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0124] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A power quality control method, characterized in that, The system is applied to a power quality control system, which includes a control center and multiple control nodes with deployed intelligent agents. Each of the aforementioned agents is equipped with a pre-trained multi-agent deep reinforcement learning model; The method is executed by the control node, and the method includes: Obtain the distribution network observation data of the control node; The power distribution network observation data is input into the multi-agent deep reinforcement learning model to obtain power quality control actions; wherein, the multi-agent deep reinforcement learning model makes decisions based on Markov decision processes; the multi-agent deep reinforcement learning model is trained by the control center based on the experience values ​​determined by each agent based on local target sample data; The power quality control action is performed on the distribution network where the control node is located.

2. The method according to claim 1, characterized in that, The multi-agent deep reinforcement learning model includes an action network and a value network; The training steps of the multi-agent deep reinforcement learning model include: The agent in each of the control nodes inputs local target sample data into the action network to obtain experience values; the experience values ​​are then stored in the experience replay buffer. The control center randomly selects a batch of experience values ​​from the experience replay buffer, calculates the action gradient based on the selected experience values, and iteratively adjusts the weight parameters of the action network based on the action gradient; it also calculates the loss function value based on the selected experience values ​​and iteratively adjusts the weight parameters of the value network based on the loss function value.

3. The method according to claim 2, characterized in that, The step of inputting local target sample data into the action network to obtain empirical values ​​includes: The local target sample data is input into the action network to obtain the current action with the highest action value selected by the action network based on the current state in the local target sample data. Determine the current reward and target state obtained by performing the current action; Determine the experience value comprised of the current state, the current action, the current reward, and the target state.

4. The method according to claim 3, characterized in that, The step of calculating the action gradient based on the extracted empirical values ​​and iteratively adjusting the weight parameters of the action network based on the action gradient includes: The objective function is calculated based on the extracted experience values, and the objective function is the expected cumulative reward. Calculate the action gradient of the objective function; The weight parameters of the action network are iteratively updated based on the action gradient using the gradient ascent method.

5. The method according to claim 3, characterized in that, The step of calculating the loss function value based on the extracted empirical values ​​includes: Calculate the value of the current action based on the current state and current action in the extracted experience values; Calculate the action value of all actions corresponding to the target state in the extracted empirical values; The target action value is calculated based on the maximum value among all action values ​​and the current reward from the extracted experience value. Calculate the loss function value between the target action value and the current action value.

6. The method according to claim 2, characterized in that, The multi-agent deep reinforcement learning model further includes a data processing module; the data processing module includes a Bayesian inference layer and an adversarial network generation layer; the data processing module processes the input local initial sample data to obtain high-confidence local target sample data; the local initial sample data includes missing values ​​and noise; The Bayesian inference layer uses a Bayesian probability matrix factorization model to estimate the posterior distribution of missing values ​​in the local initial sample data, and samples imputation values ​​from the local initial sample data to fill the missing values ​​according to the posterior distribution; noise separation is performed on the imputed local initial sample data to obtain local denoised sample data. The adversarial network generation layer generates high-confidence local target sample data based on the local denoised sample data, using the posterior distribution as a condition.

7. A power quality control device, characterized in that, A control node applied in a power quality control system, wherein the power quality control system includes a control center and multiple control nodes with deployed intelligent agents; Each of the aforementioned agents is equipped with a pre-trained multi-agent deep reinforcement learning model; the device includes: The observation data acquisition module is used to acquire the distribution network observation data of the control node; The control action determination module is used to input the distribution network observation data into the multi-agent deep reinforcement learning model to obtain power quality control actions; wherein, the multi-agent deep reinforcement learning model makes decisions based on Markov decision processes; the multi-agent deep reinforcement learning model is trained by the control center based on the experience values ​​determined by each agent based on local target sample data; The power quality control module is used to perform the power quality control action on the distribution network where the control node is located.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the power quality control method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that are used to cause a processor to execute the power quality control method according to any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the power quality control method according to any one of claims 1-6.