Method and device for multi-agent power grid voltage optimization control based on large language model, electronic equipment and readable storage medium

By employing a multi-agent power grid voltage optimization control method based on a large language model, and utilizing prompt word engineering to generate diverse training data and combining it with a distributed decision-making process, the voltage optimization problem of the power grid in complex scenarios is solved, achieving safe and stable power grid voltage and minimizing network losses.

CN122159234APending Publication Date: 2026-06-05ZHIBO ENERGY TECHNOLOGY (JIANGSU) CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHIBO ENERGY TECHNOLOGY (JIANGSU) CO LTD
Filing Date
2026-02-03
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing grid voltage optimization control methods have shortcomings in terms of data dependence, generalization ability, and adaptability to complex scenarios. In particular, they are difficult to achieve safe and stable voltage and minimize network losses in situations with a high proportion of distributed energy access and complex operating environments.

Method used

A multi-agent power grid voltage optimization control method based on a large language model is adopted. The large language model is guided to generate diverse voltage control scenarios through prompt word engineering. Combined with distributed partially observable Markov decision process and multi-agent TD3 algorithm, a multi-agent system is constructed to realize centralized training and distributed execution of voltage regulation strategy.

Benefits of technology

It improves the robustness and adaptability of grid voltage optimization control, effectively copes with node voltage fluctuations, load uncertainties and inter-regional coupling effects, achieves safe and stable grid voltage and minimizes network losses, and has good scalability and real-time control capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122159234A_ABST
    Figure CN122159234A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of power grid optimization control and artificial intelligence, and particularly relates to a multi-agent power grid voltage optimization control method and device based on a large language model, an electronic device and a readable storage medium, which comprises the following steps: using prompt word engineering technology, inputting power grid environment information, task target optimization model and device characteristics as context information into a large language model to generate time sequence running data sets covering different working conditions; constructing a multi-agent system based on a distributed partially observable Markov decision process mechanism and training the same by using a TD3 (double-delay deep deterministic policy gradient) algorithm; and deploying the multi-agent system in a power grid control system, and each regional agent executes voltage regulation actions according to real-time collected local state information. The application can realize safe stability and network loss minimization of power grid voltage without relying on accurate system mechanism models / equations, and has good scalability and engineering application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of power grid optimization control and artificial intelligence technology, specifically relating to a multi-agent power grid voltage optimization control method, device, electronic equipment and readable storage medium based on a large language model. Background Technology

[0002] Current power grid voltage regulation primarily relies on mechanistic model-based optimization and control methods, such as mathematical programming, model predictive control, and heuristic algorithms. These methods require accurate modeling of the power grid's physical parameters, topology, and operating characteristics. However, in actual operation, due to the large-scale integration of renewable energy and distributed power sources, as well as cross-regional coupling and complex operating environments, grid parameters are difficult to obtain precisely, and models often contain biases, leading to a gap between control performance and actual operational needs. Furthermore, with the increasing uncertainty and volatility of renewable energy output, traditional methods bear a heavy computational burden when dealing with dynamic and complex scenarios, lacking real-time performance and adaptability.

[0003] In recent years, deep reinforcement learning, as a data-driven control method, has been introduced into power grid voltage optimization and regulation. It overcomes the dependence on precise mechanistic models and can learn approximately optimal control strategies through interaction with the environment. It exhibits good adaptability and rapid solution capabilities in complex nonlinear and uncertain scenarios, and is well-suited to the characteristics of coordinated regulation of multiple regions, nodes, and types of resources in the power grid. However, its application also has shortcomings, especially in the lack of extreme operating condition samples due to the fact that the power grid typically operates within a stable range. This results in weak generalization ability of the agent, which may lead to erroneous decisions when encountering new operating conditions, thus affecting the safe and stable operation of the power grid.

[0004] Therefore, existing power grid voltage optimization control methods still have significant shortcomings in terms of data dependence, generalization ability, and adaptability to complex scenarios. There is an urgent need for an improved method that can enhance the robustness and reliability of the strategy under diverse operating scenarios. Summary of the Invention

[0005] To address the aforementioned issues, embodiments of this application provide a multi-agent power grid voltage optimization control method, apparatus, electronic device, and readable storage medium based on a large language model. This method utilizes a large language model for data augmentation during the training phase to generate diverse voltage control scenarios, thereby improving the generalization and robustness of the multi-agent deep reinforcement learning strategy, overcoming or at least partially overcoming the shortcomings of existing technologies.

[0006] Firstly, this application provides a multi-agent power grid voltage optimization control method based on a large language model, including: The prompt information input step utilizes prompt word engineering technology to input power grid environment information, task objective optimization model, and equipment characteristics as contextual information into the large language model; In the data generation step, the large language model generates a time-series running dataset covering different working conditions based on the input information, which serves as the training dataset for the multi-agent system. The multi-agent training steps involve constructing a multi-agent system based on a distributed partially observable Markov decision process mechanism and training it using the TD3 algorithm. The control implementation steps are optimized by deploying the multi-agent system in the power grid control system, and each regional agent performs voltage regulation actions based on the real-time collected local state information.

[0007] Secondly, this application also provides a multi-agent power grid voltage optimization control device based on a large language model, the device comprising: The prompt information input unit is used to input power grid environment information, task objective optimization model, and equipment characteristics as context information into the large language model using prompt word engineering technology; The data generation unit is used by the large language model to generate time-series operation datasets covering different working conditions based on the input information, which serve as training datasets for multiple agents. A multi-agent training unit is used to construct a multi-agent system based on a distributed partially observable Markov decision process mechanism and to train it using the TD3 algorithm. An optimized control implementation unit is used to deploy the multi-agent system in the power grid control system, whereby each regional agent performs voltage regulation actions based on real-time collected local state information.

[0008] Thirdly, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described multi-agent power grid voltage optimization control method based on a large language model.

[0009] Fourthly, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described multi-agent power grid voltage optimization control method based on a large language model.

[0010] The above-mentioned at least one technical application used in the embodiments of this application can achieve the following beneficial effects: This application guides a large language model through prompt word engineering to generate diverse operational data that conforms to power physics constraints and meteorological characteristics, thereby achieving high-quality data augmentation. The agents in each power grid control area learn regional collaborative voltage control strategies based on the augmented data, and complete real-time voltage optimization through a combination of centralized training and distributed execution.

[0011] This application combines a distributed partially observable Markov decision process with the multi-agent TD3 algorithm. During the centralized training phase, global rewards are used to achieve cross-regional collaboration, while during the execution phase, real-time voltage regulation can be achieved by relying only on local observations. The large language model can generate diverse and realistic operating scenarios based on seasons, climates, and load patterns, thereby significantly improving the robustness of the control strategy and its adaptability to unseen operating conditions. The multi-agent TD3 training process combines centralized training and distributed execution, which ensures global collaborative performance while reducing communication and computational overhead, thus meeting the real-time control requirements of the power grid.

[0012] In summary, this application effectively addresses the control difficulties arising from fluctuations in node voltage due to various power sources, load uncertainties, and inter-regional coupling in power grids with a high proportion of distributed energy resources. It overcomes the shortcomings of traditional centralized control systems that rely on precise models in terms of real-time performance and adaptability. This application achieves secure and stable grid voltage and minimizes network losses without relying on precise system mechanism models / equations, demonstrating good scalability and promising engineering applications. Attached Figure Description

[0013] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A schematic diagram of the basic structure of a power grid voltage control framework according to an embodiment of this application is shown; Figure 2 A flowchart illustrating a multi-agent power grid voltage optimization control method based on a large language model according to an embodiment of this application is shown. Figure 3 A schematic diagram of a process for generating a training dataset using prompt word engineering guided by an embodiment of this application is shown. Figure 4 A schematic diagram of a multi-agent power grid voltage optimization control device based on a large language model according to an embodiment of this application is shown. Figure 5 A schematic diagram of the resulting electronic device according to an embodiment of this application is shown. Detailed Implementation

[0014] To make the objectives, technical claims, and advantages of this application clearer, the technical application of this application will be clearly and completely described below with reference to specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0015] The purpose of this application is to overcome the shortcomings of existing AI-based power grid voltage optimization control methods, particularly their strong dependence on historical data during training, limited available samples, and poor adaptability under unfamiliar operating conditions. To address these issues, this application proposes a multi-agent deep reinforcement learning-based power grid voltage optimization control method assisted by a large language model. This method utilizes the generative capabilities of the large language model to expand and enhance the training dataset, thereby improving the adaptability and reliability of the deep reinforcement learning algorithm under diverse operating scenarios, ultimately achieving stable voltage operation and optimized network losses.

[0016] Figure 1This diagram illustrates the basic structure of a power grid voltage optimization control framework according to an embodiment of this application. This framework aims to achieve efficient and stable regulation of the power grid voltage. The entire power grid voltage control framework mainly consists of a data acquisition layer, an analysis and decision-making layer, and an execution layer. The data acquisition layer uses various sensors to collect real-time data such as voltage and current from the power grid. The analysis and decision-making layer receives the collected data, analyzes and processes it, and generates control strategies based on preset rules and algorithms. The execution layer receives the control strategies and implements voltage control by adjusting transformer taps, switching capacitor banks, etc. The data acquisition layer is located at the bottom of the framework and is the data source for the entire framework. It collects real-time operating data of the power grid, including parameters such as voltage, current, and power, through various sensors such as smart meters, current transformers, and voltage transformers, providing basic data for subsequent analysis and decision-making. The analysis and decision-making layer is located in the middle of the framework and receives data from the data acquisition layer. This layer uses big data analytics, machine learning, and artificial intelligence technologies to process and analyze the collected data, extracting valuable information to formulate reasonable voltage control strategies and decisions. Large language models can be integrated into the analysis and decision-making layer, leveraging their powerful language understanding, generation, and reasoning capabilities to help process and analyze complex power grid data, assisting in the formulation of more precise control strategies. For example, large language models can perform semantic analysis on large amounts of power grid operation data, extracting key information to support decision-making. The execution layer can be located at the upper layer of the framework, operating and controlling equipment in the power grid based on the control strategies generated by the analysis and decision-making layer, such as adjusting transformer taps and switching capacitor banks, to achieve optimized control of power grid voltage. Multiple agents can be distributed between the analysis and decision-making layer and the execution layer. In the analysis and decision-making layer, agents can autonomously analyze tasks, plan strategies, and dynamically adjust execution schemes, integrating multiple models, APIs, and external tools to complete complex tasks, improving the efficiency and accuracy of decision-making; in the execution layer, agents can automatically execute corresponding control operations based on decision results, and through continuous performance optimization, better adapt to the dynamic changes of the power grid.

[0017] Figure 2 This paper illustrates a flowchart of a multi-agent power grid voltage optimization control method based on a large language model according to an embodiment of this application. Figure 1 As can be seen, this embodiment includes steps S110 to S140: Prompt information input step S110: Using prompt word engineering technology, power grid environment information, task objective optimization model, and equipment characteristics are input into the large language model as context information.

[0018] A systematic prompt word engineering framework is constructed, incorporating power grid environmental parameters, task requirements, and the characteristics of adjustable generator sets and adjustable equipment as contextual information into the large language model, thus clarifying their roles and objectives in training data generation.

[0019] Specifically, in some embodiments of this application, the power grid environment parameters include, but are not limited to: power grid topology, number of nodes, and regional division method; the task objective is to minimize network losses while meeting voltage safety constraints; the equipment characteristics include, but are not limited to: the adjustment capability range of adjustable generator sets (conventional units, renewable energy units), energy storage devices, flexible AC transmission systems (FACTS), and DC transmission systems, etc.

[0020] The objective of voltage optimization control is to keep voltage fluctuations at all nodes within acceptable limits during the scheduling cycle while minimizing network losses. The optimization model achieves this objective by adjusting the active and reactive power of adjustable generation resources. This model can control both conventional and renewable energy units and can coordinate with energy storage devices, flexible DC transmission, and FACTS (Flexible Interchange Controllers) and other flexible regulation equipment. Its objective function is constructed as follows: ; Among them, the t Network loss during scheduling period Represented as: ; Among them, T Indicates the total number of time periods in the scheduling cycle; This represents the total number of nodes in the power grid. , They are nodes and nodes The voltage; For nodes and nodes voltage phase angle difference, , They are nodes and nodes The electrical conductance and susceptance between them.

[0021] In some embodiments of this application, the constraints of the above-mentioned task objective optimization model include: power flow constraints to ensure power balance; node voltage constraints to ensure that the voltage amplitude does not exceed the limit; adjustable power supply output constraints; and energy storage state constraints.

[0022] Among them, power flow constraints are used to ensure node power balance and satisfy the relationship between active and reactive power flow, and their expression is as follows: ; in, , For nodes The active and reactive contributions of adjustable resources , For nodes The active and reactive power requirements of the load.

[0023] Node voltage constraints, used to ensure that voltage amplitude does not exceed limits, are expressed as follows: ,| | , ; in, This is the lower limit of the node voltage. This represents the upper limit of the node voltage. This is the reference voltage. This is the maximum allowable deviation from the reference voltage.

[0024] The expression for the adjustable power supply output constraint is as follows: , , ; in, , To contribute the most effective and ineffective effort, This is for the minimum reactive power output.

[0025] Energy storage state constraints, namely energy storage state updates and capacity limitations, are expressed as follows: ; ; in, For nodes m Energy status of energy storage devices The energy storage charging and discharging power (>0 indicates charging, <0 indicates discharging). For charging efficiency, For discharge efficiency, This is the upper limit of energy storage capacity. This represents the lower limit of energy storage capacity.

[0026] Through the above constraints and objective function, the model can optimize the power generation and energy storage scheduling of the entire network while ensuring the safety of node voltage, thereby minimizing network losses and taking into account the operational limitations of various power sources.

[0027] To ensure the above generation process meets specific physical constraints and scenario requirements, it needs to be precisely guided by systematic prompt word engineering. Prompt word engineering, through the design of precise prompts, guides the model to generate high-quality results that meet the requirements. The core of prompt word engineering lies in conveying the task context and generation requirements to the large model through carefully designed prompt words, ensuring that the generated content effectively serves the agent's training objectives. For large language models, prompt word engineering guides them to generate training data that conforms to the characteristics of power grid operation, thereby enhancing the generalization of training data, mitigating performance fluctuations caused by insufficient data, and optimizing the model's resource utilization efficiency.

[0028] Please refer to Figure 3 , Figure 3 This illustration shows a flowchart of a prompt word engineering process for generating a training dataset for a large language model according to an embodiment of this application. Figure 3 As can be seen, the prompt word engineering can generate prompt words based on the initial dataset and user needs, input the prompt words into the large model, and the large model generates a new dataset based on the prompt words.

[0029] The prompt word engineering, through three modules—task description, generation logic constraints, and output quality control—guides the large language model to generate diverse training data that conforms to the physical laws and operational characteristics of the power system. Specifically, the task description module clarifies the type, capacity, and spatial distribution characteristics of power grid units, and specifies the format and temporal granularity of the output power data to guide the large language model in generating training data that matches the operational characteristics of the power system. The generation logic and constraint module embeds power balance constraints, power supply capacity boundary conditions, and seasonal meteorological influences into the prompts to ensure that the generated data meets the physical operational constraints of the power grid. The output quality control module sets data accuracy levels and consistency verification rules to improve the fidelity and applicability of the generated data, thereby enhancing the state coverage of the experience replay pool and the generalization ability of voltage optimization strategies during agent training.

[0030] As can be seen, the task description module can clearly express any background and data generation requirements in the prompts, including the type, capacity and distribution characteristics of the power grid units, as well as the data format and time granularity of the required output power.

[0031] The generation logic and constraints can prompt the model to describe the characteristics and constraints of the target data, guiding the large model to generate reasonable data based on the overall adjustable power sources and operating characteristics of the power grid. The power balance, capacity boundaries, and seasonal meteorological conditions that must be met should be listed to ensure that the generated data conforms to the power system operating constraints.

[0032] The output quality control module can prompt users to specify the accuracy requirements of the generated data, ensuring that the output meets the training requirements for data accuracy.

[0033] Through the above process, the agent can learn from the diverse and high-fidelity operational data generated by the large language model during the training phase, expand the state coverage of the experience replay pool, alleviate overfitting, and improve the generalization ability of the final voltage optimization control strategy.

[0034] Data generation step S120: The large language model generates a time-series running dataset covering different working conditions based on the input information, which serves as the training dataset for the multi-agent system.

[0035] By leveraging a large language model guided by prompt word engineering, customized datasets can be generated based on input information, producing high-fidelity time-series operational data covering different operating conditions. This data can be used to train reinforcement learning agents, construct diverse operational scenarios, and achieve data augmentation.

[0036] Considering that the actual operating state of the power grid is affected by multiple factors such as climate, season, load structure, and equipment maintenance, the number of effective samples in historical operating data that truly contain issues such as node voltage exceeding limits and excessive fluctuations is very limited. These critical operating conditions occur infrequently in actual power grids, but they have a decisive impact on the robustness of voltage control strategies. Traditional training methods that rely on fixed operating modes or simple random disturbances often fail to generate sufficient voltage problem data, making it difficult to cover the diverse scenarios that may occur in the power grid under complex and changing conditions. This leads to overfitting during agent training, insufficient generalization ability, and poor adaptability to unseen operating conditions.

[0037] To overcome this problem, this application introduces a large language model during the experience replay pool construction stage of agent training. The large language model possesses long-term reasoning and data generation capabilities, and under the guidance of cue word engineering, it can generate diverse operational data covering different seasons, load levels, and abnormal operating conditions, combined with the physical constraints of the power system, unit characteristics, and meteorological patterns. In particular, this application can automatically expand and enhance scarce scenario samples such as voltage limit exceedances through the large language model, enabling the training dataset to maintain realistic statistical regularities while having a wider range of operating conditions. This effectively improves the generalization and robustness of deep reinforcement learning strategies, providing more reliable support for grid voltage optimization control.

[0038] First, a large language model is constructed. This model can be implemented based on any existing generative large language network technology, and it can be viewed as a conditional probability model for time series data. The model is as follows: ; Each of them This corresponds to the output of adjustable generator units within a specific time range. The large language model learns the statistical and semantic characteristics of historical data to predict the state at the next time step.

[0039] The goal of large language training is to maximize the likelihood function with respect to real data, as shown below: .

[0040] The mechanism of the aforementioned large language model ensures that the generated data not only conforms to historical statistical patterns, but also allows for reasonable inferences about system behavior under unobserved conditions.

[0041] Then, the large language model is trained and optimized, learning a wealth of textual and time-series knowledge about conventional generating units, various adjustable power sources, overall grid operation, and meteorological changes. This knowledge is embedded in the data generation process through specially designed prompts, ensuring that the generated unit output data conforms to the physical and operational laws of the power system, thus improving coverage and generalization ability. Compared to traditional methods that rely on fixed distributions, this generation mechanism can dynamically simulate diverse scenarios, avoiding excessive reliance on manually set parameters.

[0042] Through carefully designed prompts, the large language model generates operational data covering multiple scenarios, outputting 24-hour or finer-grained time-series data, encompassing capacity constraints and power fluctuation characteristics, and constructing various typical operational scenarios. Compared to traditional Monte Carlo methods or manually set distributions, this generation mechanism is more flexible, automatically adjusting data distribution according to the overall power grid's operational characteristics, reducing reliance on expert parameter settings. Leveraging its rich knowledge base, the large language model can provide more diverse data with less human intervention.

[0043] Finally, the trained large language model is put into use. When generating data, the large language model must meet the numerical constraints explicitly given in the prompts, such as capacity limits and voltage ranges. It also infers implicit conditions such as meteorological patterns, seasonal regularities, and changes in the overall network operation status based on the latent knowledge gained during training, making the generated data closer to real historical records and possessing stronger generalization capabilities. Simultaneously, the data generated by the large language model should maintain the basic statistical characteristics of various generating units, such as meeting average output and fluctuation range requirements, while incorporating real output characteristics under the constraints of the prompts, thereby achieving a higher order of distribution consistency.

[0044] As can be seen from the above, the large language model in this application can generate diverse time-series data containing different seasons, load levels, and abnormal operating conditions by introducing conditional constraints through prompt word engineering when generating time-series operational datasets, based on the physical constraints of the power system, the operating characteristics of generating units, and meteorological patterns. The generation process satisfies explicit numerical constraints such as the upper limit of generator unit capacity and the safe voltage range, and combines the historical statistical features and implicit operating patterns learned by the model to ensure that the generated data maintains the consistency of basic statistical attributes such as average output and power fluctuation characteristics, while enhancing the sample coverage density of key events such as node voltage exceeding limits and large power fluctuations in scarce scenarios, thereby constructing a training dataset that combines realism, diversity, and high generalization ability.

[0045] Multi-agent training step S130: Construct a multi-agent system based on a distributed partially observable Markov decision process mechanism and train it using the TD3 algorithm.

[0046] Considering the complexity of actual power grid physical parameters and the uncertainty of load and various (new energy) power source outputs, this application adopts a multi-agent deep reinforcement learning algorithm. This method is entirely data-driven, requiring no precise system model or equations, and possesses excellent online learning and optimization control capabilities. The power grid is divided into several control regions, each configured with one agent. Each agent observes only the local state within its region, and the agents cooperate to achieve the control objective. Therefore, the voltage optimization control task is modeled as a distributed partially observable Markov decision process, consisting of three elements: state space, actions, and reward function.

[0047] In some embodiments of this application, time is defined t ,area r (Intelligent agent) r The observable local state includes the voltage magnitudes of each node within that region. Adjustable resources have active power Adjustable resource reactive power Active power of load Reactive power of load and energy storage level The state space can then be represented as: ; in, For the region r The set of nodes it contains.

[0048] In some embodiments of this application, the intelligent agent r At any moment k The actions include adjustable reactive and active control quantities within the region; therefore, the action space can be represented as: ; in, These represent the active and reactive power outputs of adjustable power generation or other controllable resources, respectively. , This refers to the active and reactive power regulation of the energy storage device.

[0049] In some embodiments of this application, the reward function is a globally shared reward. At time t, the globally shared reward is used to guide the agent to collaboratively reduce network loss and penalize voltage exceedances. The globally shared reward can be expressed as:

[0050] in, , This is the penalty coefficient for network loss and voltage exceeding limits.

[0051] By using a customized global shared reward function, node voltage over-limit control and network loss optimization are considered in a unified manner, ensuring voltage stability and operating efficiency under the fluctuation of multiple power sources.

[0052] In this application, the TD3 (Dual Delay Deep Deterministic Policy Gradient) algorithm is used to train each agent.

[0053] In this application, the training of the agents employs a scheme of centralized training and distributed execution. During the offline centralized training phase, agents corresponding to each region collaboratively learn voltage regulation strategies in an offline environment using global information. In the subsequent online execution phase, the agents do not need to interact with other regions; they only need to utilize local observation information and the strategies obtained from offline training to make decisions in real time. This model effectively reduces training computation while ensuring collaborative performance and mitigates the curse of dimensionality caused by the increased number of agents and information dimensionality.

[0054] To address the continuous action space and enhance stability, the multi-agent TD3 algorithm is used for centralized training. In some embodiments of this application, the core steps are as follows: Initialize the network parameters and experience replay pool for all agents, and set the step size for each training epoch to be [value missing]. T Each step corresponds to a specific time t in a day. At each time t, all agents interact with the power grid of their respective responsible areas in the simulation environment and output actions based on the policy, while adding random Gaussian noise. The environment performs flow calculations based on the actions of all agents, updates rewards and states, and generates an experience tuple consisting of state, action, reward, and the next state. , , , The data is stored in the replay pool. Each agent performs batch sampling from the experience replay pool to update the network parameters. A soft update method is used to update the target network parameters.

[0055] Soft update is a parameter update strategy that improves training stability by slowly and smoothly adjusting the target network parameters to avoid drastic fluctuations in target values.

[0056] The core principle of soft update lies in: target network parameters θ' The update is achieved by weighting the main network parameter θ with its own old parameters: θ' = τθ + (1-τ)θ' ,in τ This is used to update the coefficients (e.g., a typical value of 0.001). Each time the main network is updated, the target network is only fine-tuned to ensure it slowly follows the changes in the main network, rather than directly copying the parameters.

[0057] Soft updates reduce training oscillations by suppressing sudden changes in target values, making them particularly suitable for environments with continuously changing action spaces and dynamically changing sample distributions.

[0058] Optimized control implementation step S140: The multi-agent system is deployed in the power grid control system, and each regional agent performs voltage regulation actions based on the real-time collected local state information.

[0059] After training, each agent outputs control actions in real time based solely on its local observation state, achieving distributed online voltage regulation without the need for real-time communication. During the runtime phase, each agent performs distributed online execution based on offline-learned strategies, achieving voltage regulation and loss optimization. This method effectively addresses the problems of insufficient training samples and limited generalization ability while ensuring collaborative performance.

[0060] This application guides a large language model through prompt word engineering to generate diverse operational data that conforms to power physics constraints and meteorological characteristics, thereby achieving high-quality data augmentation. The agents in each power grid control area learn regional collaborative voltage control strategies based on the augmented data, and complete real-time voltage optimization through a combination of centralized training and distributed execution.

[0061] This application combines a distributed partially observable Markov decision process with the multi-agent TD3 algorithm. During the centralized training phase, global rewards are used to achieve cross-regional collaboration, while during the execution phase, real-time voltage regulation can be achieved by relying only on local observations. The large language model can generate diverse and realistic operating scenarios based on seasons, climates, and load patterns, thereby significantly improving the robustness of the control strategy and its adaptability to unseen operating conditions. The multi-agent TD3 training process combines centralized training and distributed execution, which ensures global collaborative performance while reducing communication and computational overhead, thus meeting the real-time control requirements of the power grid.

[0062] In summary, this application effectively addresses the control difficulties arising from fluctuations in node voltage due to various power sources, load uncertainties, and inter-regional coupling in power grids with a high proportion of distributed energy resources. It overcomes the shortcomings of traditional centralized control systems that rely on precise models in terms of real-time performance and adaptability. This application achieves secure and stable grid voltage and minimizes network losses without relying on precise system mechanism models / equations, demonstrating good scalability and promising engineering applications.

[0063] Figure 4 A schematic diagram of a multi-agent power grid voltage optimization control device based on a large language model according to an embodiment of this application is shown. Figure 4 It can be seen that the multi-agent power grid voltage optimization control device 400 based on a large language model includes: The prompt information input unit 410 is used to input power grid environment information, task target optimization model and equipment characteristics as context information into the large language model using prompt word engineering technology; The data generation unit 420 is used by the large language model to generate a time-series running dataset covering different working conditions based on the input information, which serves as a training dataset for the multi-agent system. The multi-agent training unit 430 is used to construct a multi-agent system based on a distributed partially observable Markov decision process mechanism and is trained using the TD3 algorithm. The optimized control implementation unit 440 is used to deploy the multi-agent system in the power grid control system, and each regional agent performs voltage regulation actions based on the real-time collected local state information.

[0064] In some embodiments of this application, in the above-described apparatus, the objective function of the task objective optimization model is: ; Among them, the t Network loss during scheduling period Represented as: ; in, T Indicates the total number of time periods in the scheduling cycle; This represents the total number of nodes in the power grid. , They are nodes and nodes The voltage; For nodes and nodes voltage phase angle difference, , They are nodes and nodes The electrical conductance and susceptance between them.

[0065] In some embodiments of this application, the constraints of the task objective optimization model in the above-described apparatus include: power flow constraints, node voltage constraints, adjustable power output constraints, and energy storage state constraints. The expression for the power flow constraint is: ; in, , For nodes Adjustable resources contribute both active and reactive energy. , For nodes The active and reactive power demands of the load; The expression for the node voltage constraint is: ,| | , ; in, This is the lower limit of the node voltage. This is the upper limit of the node voltage. For reference voltage, The maximum allowable deviation from the reference voltage; The expression for the adjustable power supply output constraint is: , , ; in, , To contribute the most effective and ineffective effort, Minimize reactive power output; The expression for the energy state constraint of energy storage is: ; ; in, For nodes m Energy status of energy storage devices For energy storage charging and discharging power, For charging efficiency, For discharge efficiency, This is the upper limit of energy storage capacity. This represents the lower limit of energy storage capacity.

[0066] In some embodiments of this application, the prompt word engineering in the above-described apparatus includes a task description module, a generation logic constraint module, and an output quality control module; Specifically, by clearly defining the type, capacity, and spatial distribution characteristics of power grid units in the task description module, and specifying the format and time granularity of the output power data, the large language model is guided to generate training data that conforms to the operating characteristics of the power system. By embedding power balance constraints, power capacity boundary conditions, and seasonal meteorological factors in the prompts through the generation logic and constraint module, the generated data is ensured to meet the physical operation constraints of the power grid. By setting data accuracy levels and consistency verification rules through the output quality control module, the fidelity and applicability of the generated data are improved, thereby enhancing the state coverage of the experience replay pool and the generalization ability of the voltage optimization strategy during the training of the agent.

[0067] In some embodiments of this application, in the above-described apparatus, the large language model is a conditional probability model for time series data. The model is as follows: ; Each of them The output of an adjustable generator set corresponding to a specific time period and region; The goal of the large language training is to maximize the likelihood function with respect to the real data, as follows: .

[0068] In some embodiments of this application, in the above-described apparatus, the distributed partially observable Markov decision mechanism consists of three elements: a state space, actions, and a reward function, wherein the reward function is defined at time t. t ,area r (Intelligent agent) r The observable local state includes the voltage magnitudes of each node within that region. Adjustable resources have active power Adjustable resource reactive power Active power of load Reactive power of load and energy storage level The state space is then represented as: ; in, For the region r The set of nodes contained therein; intelligent agent r At any moment k The actions include adjustable reactive and active control quantities within the region; therefore, the action space is represented as: ; in, These represent the active and reactive power outputs of adjustable power generation or other controllable resources, respectively. , This refers to the active and reactive power regulation of the energy storage device.

[0069] The reward function is a globally shared reward, which can be represented as:

[0070] in, , This is the penalty coefficient for network loss and voltage exceeding limits.

[0071] In some embodiments of this application, in the above-described apparatus, the multi-agent system is trained according to the following method: Initialize the network parameters and experience replay pool of all agents, set the step size of each training round to T, and each step corresponds to a specific time t in a day; At each time t, all agents interact with their respective regional power grids in the simulation environment and output actions based on the policy, while adding random Gaussian noise. ; The environment performs flow calculations based on the actions of all agents, updates rewards and states, and stores the experience tuple consisting of state, action, reward, and the next state into the replay pool. Each agent performs batch sampling from the experience replay pool to update the network parameters; The target network parameters are updated using a soft update method.

[0072] It should be noted that the aforementioned multi-agent power grid voltage optimization control device based on a large language model can implement the aforementioned multi-agent power grid voltage optimization control method based on a large language model. Details will not be elaborated here.

[0073] Figure 5 This invention illustrates a schematic diagram of the structure of an electronic device according to an embodiment of the present application. Figure 5 As shown, the electronic device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used for communication with external devices via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a multi-agent power grid voltage optimization control method based on a large language model.

[0074] In one embodiment, the electronic device provided in this application includes a memory and a processor. The memory stores a database and a computer program that can run on the processor. When the processor executes the computer program, it implements the steps of the aforementioned multi-agent power grid voltage optimization control method based on a large language model.

[0075] The above is as stated in this application. Figure 4 The method for multi-agent power grid voltage optimization control based on a large language model, as disclosed in the illustrated embodiments, can be applied to a processor or implemented by a processor. During implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The steps of the method disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.

[0076] In one embodiment, a computer-readable storage medium is also provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the aforementioned multi-agent power grid voltage optimization control method based on a large language model.

[0077] It should be noted that the functions or steps that the above-mentioned electronic devices or computer-readable storage media can achieve can be referred to the relevant descriptions in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0078] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0079] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above.

[0080] The above-described embodiments are only used to illustrate the technical application of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical applications described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical applications to deviate from the spirit and scope of the technical applications of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A multi-agent power grid voltage optimization control method based on a large language model, characterized in that, include: The prompt information input step utilizes prompt word engineering technology to input power grid environment information, task objective optimization model, and equipment characteristics as contextual information into the large language model; In the data generation step, the large language model generates a time-series running dataset covering different working conditions based on the input information, which serves as the training dataset for the multi-agent system. The multi-agent training steps involve constructing a multi-agent system based on a distributed partially observable Markov decision process mechanism and training it using the TD3 algorithm. The control implementation steps are optimized by deploying the multi-agent system in the power grid control system, and each regional agent performs voltage regulation actions based on the real-time collected local state information.

2. The method according to claim 1, characterized in that, The objective function of the task objective optimization model is: ; Among them, the t Network loss during scheduling period Represented as: ; in, T Indicates the total number of time periods in the scheduling cycle; This represents the total number of nodes in the power grid. , They are nodes and nodes The voltage; For nodes and nodes voltage phase angle difference, , They are nodes and nodes The electrical conductance and susceptance between them.

3. The method according to claim 1 or 2, characterized in that, The constraints of the task objective optimization model include: power flow constraints, node voltage constraints, adjustable power output constraints, and energy storage state constraints. The expression for the power flow constraint is: ; in, , For nodes The active and reactive contributions of adjustable resources , For nodes The active and reactive power demands of the load; The expression for the node voltage constraint is: ,| | , ; in, This is the lower limit of the node voltage. This is the upper limit of the node voltage. For reference voltage, The maximum allowable deviation from the reference voltage; The expression for the adjustable power supply output constraint is: , , ; in, , To contribute the most effective and ineffective effort, Minimize reactive power output; The expression for the energy state constraint of energy storage is: ; ; in, For nodes m Energy status of energy storage devices For energy storage charging and discharging power, For charging efficiency, For discharge efficiency, This is the upper limit of energy storage capacity. This represents the lower limit of energy storage capacity.

4. The method according to claim 1, characterized in that, The prompting project includes a task description module, a generation logic constraint module, and an output quality control module; Specifically, by clearly defining the type, capacity, and spatial distribution characteristics of power grid units in the task description module, and specifying the format and time granularity of the output power data, the large language model is guided to generate training data that conforms to the operating characteristics of the power system. By embedding power balance constraints, power capacity boundary conditions, and seasonal meteorological factors in the prompts through the generation logic and constraint module, the generated data is ensured to meet the physical operation constraints of the power grid. By setting data accuracy levels and consistency verification rules through the output quality control module, the fidelity and applicability of the generated data are improved, thereby enhancing the state coverage of the experience replay pool and the generalization ability of the voltage optimization strategy during the training of the agent.

5. The method according to claim 1, characterized in that, The large language model is a conditional probability model for time series data. The model is as follows: ; Each of them The output of an adjustable generator set corresponding to a specific time period and region; The goal of the large language training is to maximize the likelihood function with respect to the real data, as follows: 。 6. The method according to claim 1, characterized in that, The distributed partially observable Markov decision mechanism consists of three elements: a state space, actions, and a reward function, where is defined at time . t ,area r The observable local state includes the voltage magnitude of each node in the region. Adjustable resources have active power Adjustable resource reactive power Active power of load Reactive power of load and energy storage level The state space is then represented as: , ; in, For the region r The set of nodes contained therein; intelligent agent r At any moment k The actions include adjustable reactive and active control quantities within the region; therefore, the action space is represented as: ; in, These represent the active and reactive power outputs of adjustable power generation or other controllable resources, respectively. , The active and reactive power regulation power of the energy storage device; The reward function is a globally shared reward, which can be represented as: - ; in, , This is the penalty coefficient for network loss and voltage exceeding limits.

7. The method according to claim 1, characterized in that, The multi-agent system is trained according to the following method: Initialize the network parameters and experience replay pool for all agents, and set the step size for each training epoch to be [value missing]. T Each step corresponds to a specific time of day. t ; At every moment t In the simulation environment, all agents interact with the power grid of their respective areas and output actions based on policies, while adding random Gaussian noise. ; The environment performs flow calculations based on the actions of all agents, updates rewards and states, and stores the experience tuple consisting of state, action, reward, and the next state into the replay pool. Each agent performs batch sampling from the experience replay pool to update the network parameters; The target network parameters are updated using a soft update method.

8. A multi-agent power grid voltage optimization control device based on a large language model, characterized in that, The device includes: The prompt information input unit is used to input power grid environment information, task objective optimization model, and equipment characteristics as context information into the large language model using prompt word engineering technology; The data generation unit is used by the large language model to generate time-series operation datasets covering different working conditions based on the input information, which serve as training datasets for multiple agents. A multi-agent training unit is used to construct a multi-agent system based on a distributed partially observable Markov decision process mechanism and to train it using the TD3 algorithm. An optimized control implementation unit is used to deploy the multi-agent system in the power grid control system, whereby each regional agent performs voltage regulation actions based on real-time collected local state information.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the multi-agent power grid voltage optimization control method based on a large language model as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the multi-agent power grid voltage optimization control method based on a large language model as described in any one of claims 1 to 7.