Grid Voltage Control Task Offloading and Computing Resource Allocation Method, Device and Medium

By setting up a partition model in the power grid and dynamically adjusting the partitioning strategy and computing resource allocation using deep reinforcement learning algorithms, the problem of increasing computing load concentration and processing delay in the traditional centralized voltage control mode is solved, efficient processing of grid voltage control tasks and optimized allocation of computing resources is achieved, and the adaptability and robustness of the power grid are improved.

CN119883654BActive Publication Date: 2025-06-24HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510364619.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-24
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

In a large-scale distributed grid architecture, the traditional centralized voltage control model faces problems such as centralized computing load, increased processing delay, and limited system scalability. How to achieve efficient processing of voltage control tasks and optimal allocation of computing resources under the constraints of limited computing resources has become a key problem affecting the efficiency of grid voltage regulation.

Method used

The following steps are performed by computer equipment: set the grid voltage node partition model according to the number of nodes and load conditions, and calculate the calculation delay of each partition voltage control task; based on the deep reinforcement learning algorithm, determine the objective functions and constraints for minimizing delay and energy consumption, dynamically adjust the partition strategy and computing resource allocation, and realize the offload of grid voltage control tasks and the optimal allocation of computing resources.

Benefits of technology

It realizes that while meeting the delay constraints, it reduces overall computing energy consumption, improves the adaptability, robustness and computing efficiency of smart grid voltage control, and helps the power grid control system to develop in a more intelligent, efficient and low-carbon direction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119883654B_ABST
    Figure CN119883654B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device and medium for grid voltage control task offloading and computing resource allocation, including: setting a grid voltage node partition model according to the number of nodes and the load condition, aiming at the number of nodes and computing complexity of different partitions, and obtaining the computing delay of the voltage control task for each partition; obtaining the global task computing delay based on the computing delays of each partition; after obtaining the global computing delay, obtaining the energy consumption during the task computing process; after obtaining the energy consumption during the task computing process, performing modeling to minimize the delay and energy consumption; finally, determining the objective function and constraint conditions based on the minimization of the delay and energy consumption, and using a deep reinforcement learning algorithm to solve, so as to obtain the grid voltage control task offloading and computing resource allocation strategy; the present invention optimizes the grid partition strategy, reasonably allocates computing resources, thereby minimizing the computing delay and computing energy consumption cost of the voltage control task, and enhancing the self-adaptability and robustness of the grid regulation system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power system resource scheduling and optimization, and specifically relates to a method, device, and medium for offloading grid voltage control tasks and allocating computing resources. Background Art

[0002] In the context of the global energy system's deep evolution towards intelligence and digitalization, the automation and intelligence levels of the power grid are continuously improving, and the construction of smart grids has become a key direction for the development of modern power systems. By integrating advanced power regulation technologies and information and communication technologies, smart grids can accurately perceive, efficiently calculate, and optimize the control of the power grid's operating state, thereby enhancing the security, stability, and economy of the power grid operation. Among them, the voltage control system based on the edge computing architecture relies on distributed computing capabilities to effectively integrate various computing resources, achieve rapid response and precise execution of voltage regulation tasks, and provide technical support for the efficient operation of the power grid.

[0003] However, with the continuous expansion of the power grid scale and the increasing proportion of distributed energy access, the topological structure of the power grid becomes more complex, and the traditional centralized voltage control mode faces many challenges such as concentrated computing loads, increased processing delays, and limited system scalability. Especially in a multi-node distributed architecture, how to efficiently process voltage control tasks and achieve optimal computing resource allocation under the constraint of limited computing resources has become a key problem affecting the voltage regulation efficiency of the power grid.

[0004] Currently, grid voltage control tasks mainly rely on centralized scheduling strategies, that is, data from all nodes need to be transmitted to the central server for calculation and decision-making. However, this method is prone to increased calculation delays and system response lags in the face of dynamic load changes, network congestion, or communication interruptions, thereby affecting the real-time performance and stability of voltage control. In addition, due to the characteristics of high computational complexity and strict timeliness requirements for voltage control tasks, how to optimize the zoning strategy of grid nodes and reasonably allocate computing resources to minimize calculation delay and calculation energy consumption cost is the core issue in the current research on smart grid calculation optimization.

[0005] In view of the above problems, it is urgent to explore a method for offloading voltage control tasks and optimizing the allocation of computing resources suitable for large-scale distributed power grid architectures, which can dynamically adjust the zoning strategy to reasonably allocate computing resources among different task units, thereby meeting the delay constraint while reducing the overall computing energy consumption. The breakthrough of this technology will significantly improve the adaptability, robustness, and computing efficiency of smart grid voltage control, contribute to the development of the power grid regulation system towards a more intelligent, efficient, and low-carbon direction, and provide important support for building a new energy Internet.

[0006] Moreover, the above content is only used to assist in understanding the technical solution of this application, and does not represent an admission that the above content is the closest prior art. Summary of the Invention

[0007] The main object of the present invention is to provide a method for offloading grid voltage control tasks and allocating computing resources to solve the technical problems proposed in the background art.

[0008] The present invention adopts the following technical solutions to solve the above technical problems:

[0009] A method for offloading grid voltage control tasks and allocating computing resources, which is executed by a computer device through the following steps:

[0010] S1. According to the number of nodes and the load condition, set up a grid voltage node partition model, aiming at the number of nodes and the calculation complexity of different partitions, and calculate the calculation delay of the voltage control task for each partition;

[0011] S2. Based on the calculation delays of each partition obtained in S1, calculate the global task calculation delay;

[0012] S3. After obtaining the global calculation delay, calculate the energy consumption during the task calculation process;

[0013] S4. After calculating the energy consumption during the task calculation process, perform modeling to minimize the delay and energy consumption;

[0014] S5. Finally, based on the minimization of the delay and energy consumption, determine the objective function and the constraint conditions, and use the deep reinforcement learning algorithm to solve to obtain the grid voltage control task offloading and computing resource allocation strategy.

[0015] Further, step S1 specifically includes,

[0016] According to the number of nodes and the load condition, divide the nodes in the physical area into several subsets, and each subset is regarded as an independent task unit to independently calculate the corresponding voltage control task;

[0017] Suppose there are nodes in the power grid, and the macro base station needs to partition them. The number of nodes in each partition is , and there are partitions, then the partitioning strategy is expressed as:

[0018] , (1)

[0019] Among them, is the partition set, and is the partition index.

[0020] Furthermore, step S1 further includes

[0021] allocating computing resources according to the number of nodes and computational complexity of different partitions to minimize the task processing delay and cost;

[0022] Let the computing power of the base station be , and allocate computing resources for each partition to satisfy . The allocation strategy affects the task execution delay and computing energy consumption. The computing delay of each partition is expressed as:

[0023] (2)

[0024] where is the computational complexity of the voltage control task for partition .

[0025] Furthermore, step S2 specifically includes

[0026] taking the maximum delay in each partition as the global task computing delay , which is expressed as:

[0027] (3).

[0028] Furthermore, step S3 specifically includes

[0029] During the task calculation, the server provides computing resources for the partition tasks. The energy consumption associated with this link, the computing energy consumption of partition is :

[0030] (4)

[0031] where is the energy consumption factor.

[0032] Furthermore, step S4 specifically includes

[0033] The problem of minimizing the calculation delay and energy consumption is expressed as follows:

[0034] (5)

[0035] where and are weight factors used to represent the importance of delay and energy consumption; is the delay constraint of the global voltage control task; represents the computing resource constraint, and the sum of the computing resources given by the server to each partition does not exceed its own computing resources; It represents the node number constraint, where the sum of the node numbers in each partition is consistent with the total number of nodes; It represents the time delay constraint, and the completion time delay of the global voltage control task shall not exceed the maximum time delay tolerance of the global voltage control; It represents the weight constraint, and the sum of the weight factors does not exceed 1.

[0036] Furthermore, the specific steps of step S5 include

[0037] 1) Initialize the power grid system parameters, including: power grid nodes, the upper limit of the base station computing resources, and the task time delay threshold; Define and initialize the deep reinforcement learning model. In the deep reinforcement learning framework, the state space of each partition is:

[0038] (6)

[0039] where is the total number of nodes, is the node partition situation, is the total amount of computing resources of the macro base station, is the historical decision feedback information;

[0040] The action space is:

[0041] (7)

[0042] where represents adjusting the number of partition nodes, represents adjusting the computing resource allocation;

[0043] For partition the reward function is:

[0044] (8)

[0045] where is the weight parameter, representing the influence weight of the computing time delay and the computing energy consumption;

[0046] The policy network for making decisions on the computing resource allocation strategy is ; Initialize the global model and the buffer , where the global model is used to store the aggregated global parameters, and the buffer is used to temporarily store the local model parameters collected from each partition;

[0047] 2) After initializing the power grid system parameters, use proximal policy optimization for learning, specifically optimize the power grid regulation strategy based on the policy-evaluation architecture. The policy network is responsible for generating the optimal action, that is, the power grid partition scheme and the computing resource allocation ratio:

[0048] ​(9)

[0049] Proximal Policy Optimization objective:

[0050] (10)

[0051] Wherein, is the policy change ratio, evaluates the superiority of the current action, is the limit range of policy change;

[0052] The evaluation network is responsible for evaluating the current state value and guiding the learning of the policy network:

[0053] (11)

[0054] The evaluation network is optimized using mean squared error:

[0055] (12)

[0056] Wherein, ;

[0057] 3) After learning using proximal policy optimization, training and optimization are carried out, specifically including initializing the parameters of the policy network and the evaluation network, setting the initial power grid partition scheme and the computing resource allocation scheme, taking actions based on the current state to update the power grid partition and resource allocation; calculating the reward and storing the experience data, namely the state, action, reward, and next state; updating the evaluation network to optimize the state value estimation; calculating the advantage estimation and updating the policy network; repeating the above steps until the algorithm converges or reaches the upper limit of the training rounds.

[0058] On the other hand, the present invention also discloses a computer-readable storage medium, characterized in that it stores a computer program, and when the computer program is executed by a processor, the processor is caused to execute the steps of the above method.

[0059] On yet another aspect, the present invention also discloses a computer device, including a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of the above method.

[0060] As can be seen from the above technical solutions, the present invention provides a method for power grid voltage control task offloading and computing resource allocation, which aims to optimize the power grid partition strategy and reasonably allocate computing resources, so as to minimize the computing delay and computing energy consumption cost of the voltage control task and improve the self-adaptability and robustness of the power grid regulation system.

[0061] Specifically, the advantages of the present invention are as follows:

[0062] 1. Real-time partition adjustment: Dynamically divide the power grid partitions through deep reinforcement learning, and adjust the partition strategy according to real-time data such as the output fluctuations of renewable energy and load changes, so as to solve the problem of rigidity in traditional static partitioning.

[0063] 2. Elastic resource allocation: Dynamically allocate computing resources based on task complexity and node load, achieve collaborative optimization of "partition-resource", and adapt to load mutations at the second level (such as fast charging of electric vehicles and industrial impact loads).

[0064] 3. Multi-objective optimization: Minimize both computing latency and energy consumption simultaneously to avoid risks of resource waste or overload.

[0065] 4. Closed-loop control strategy: Form a closed loop from data acquisition, task offloading to control feedback, and enhance the anti-interference ability of the power grid under extreme fluctuations (such as sudden drops in wind and light, and load mutations).

[0066] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Of course, any product implementing the present invention does not necessarily need to achieve all the advantages described above simultaneously. Brief Description of the Drawings

[0067] The accompanying drawings forming a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0068] Figure 1 The system architecture of the embodiments of the present invention;

[0069] Figure 2 The flowchart of the system of the present invention. Detailed Description of the Embodiments

[0070] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0071] In the embodiments, refer in detail to Figures 1 to 2 .

[0072] Such as Figure 1As shown in the figure, the grid voltage control task offloading and computing resource allocation method proposed in the embodiment of the present invention is based on a grid voltage control system. The grid voltage control system consists of a base station and multiple distributed grid nodes, and adopts a closed-loop control strategy of zoning-resource optimization-task offloading-control feedback to achieve efficient voltage regulation. The system includes the following key components:

[0073] 1. Zoning management module: According to the load conditions of grid nodes and the computing capabilities of servers, the entire power grid is divided into multiple zones. Each zone is regarded as an independent task unit, and the base station coordinates computing resources for optimization calculations. The optimization goal of the zoning strategy is to minimize the computing delay and computing energy consumption cost. The core optimization variables include the number of nodes in each zone and the load balance between zones.

[0074] 2. Computing resource allocation module: The base station is responsible for allocating computing resources to each zone to ensure that computing tasks are completed under the time delay constraint and optimal energy consumption. The computing resource allocation strategy takes into account the number of zoning nodes, task complexity, and computing capabilities of each zone to dynamically adjust the resource allocation scheme and avoid waste or overload of computing resources.

[0075] 3. Task offloading and execution module: Grid nodes transmit the collected voltage data to the base station. The base station generates corresponding task calculation models according to the zoning situation, and uses a deep reinforcement learning strategy to optimize the number of zoning nodes and computing resource scheduling.

[0076] 4. Voltage control instruction generation and feedback module: After the base station completes the task calculation, it sends the optimized voltage control instructions to the control devices in each zone, and the controller executes the voltage regulation to achieve dynamic optimization of the grid voltage.

[0077] As Figure 2 shown in the figure, the grid voltage control task offloading and computing resource allocation method proposed in the embodiment of the present invention specifically includes the following steps:

[0078] Step 1: Node zoning model

[0079] According to the number of nodes and load conditions, the nodes in the physical area are divided into several subsets, and each subset is regarded as an independent task unit for independent calculation of the corresponding voltage control task. Assume that there are nodes in the power grid, and the macro base station needs to zone them. The number of nodes in each zone is , and there are zones. Then the zoning strategy can be expressed as:

[0080] , (1)

[0081] where is a partition set, is a partition index.

[0082] Step 2: Calculate the delay model

[0083] For the number of nodes and computational complexity in different partitions, reasonably allocate computing resources to minimize the task processing delay and cost. Assume the computing power of the base station is , allocate computing resources for each partition , (satisfying ). The allocation strategy affects the task execution delay and computing energy consumption. The computing delay of each partition can be expressed as:

[0084] (2)

[0085] where is the computational complexity of the voltage control task for partition .

[0086] Step 3: Global delay model

[0087] Take the maximum delay value in each partition as the global task computing delay , which can be expressed as:

[0088] (3)

[0089] Step 4: Calculate the energy consumption model

[0090] During the task calculation process, the server provides computing resources for the partition tasks. The energy consumption associated with this link, the computing energy consumption of partition is :

[0091] (4)

[0092] where is the energy consumption factor.

[0093] Step 5: Determine the objective function and constraints

[0094] The present invention jointly optimizes the number of node partitions and the computing resource allocation of the server for each partition. The problem of minimizing the computing delay and energy consumption of the strategy of the present invention is expressed as follows:

[0095] (5)

[0096] where and are weight factors used to represent the importance of delay and energy consumption; is the delay constraint of the global voltage control task. It represents the computing resource constraint, where the sum of the computing resources given by the server to each partition does not exceed its own computing resources; It represents the node number constraint, where the sum of the node numbers in each partition is consistent with the total number of nodes; It represents the time delay constraint, where the completion time delay of the global voltage control task shall not exceed the maximum time delay tolerance of the global voltage control; It represents the weight constraint, where the sum of the weight factors does not exceed 1.

[0097] Step 6: Solve using the deep reinforcement learning algorithm

[0098] Under the established constraint framework, the present invention uses the deep reinforcement learning algorithm to solve the objective function, aiming to discover the optimal allocation scheme of the partition node numbers and computing resources. The specific steps are as follows:

[0099] The system initializes and defines the power grid parameters and model architecture, providing the basic rules for state space, action space, and reward calculation for decision optimization. Initialize the power grid system parameters, including: power grid nodes, the upper limit of the base station computing resources, and the task time delay threshold. Define and initialize the deep reinforcement learning model. In the deep reinforcement learning framework, each partition The state space is:

[0100] (6)

[0101] Among them, is the total number of nodes, is the node partition situation, is the total amount of computing resources of the macro base station, is the historical decision feedback information.

[0102] The action space is:

[0103] (7)

[0104] Among them, represents adjusting the partition node number, represents adjusting the computing resource allocation.

[0105] Partition The reward function is:

[0106] (8)

[0107] Among them, is the weight parameter, representing the influence weight of computing time delay and computing energy consumption. The objective of the present invention is to minimize the time delay and energy consumption, so it takes a negative value.

[0108] The policy network for making decisions on the computing resource allocation strategy is . Initialize the global model and buffer , the global model is used to store the aggregated global parameters, and the buffer is used to temporarily store the local model parameters collected from each partition.

[0109] The actions generated by the decision optimization policy network are used in the actual power grid environment, and the generated interaction data is fed back to the training process to form a "decision - execution - feedback" closed loop. Specifically, proximal policy optimization (PPO) is used for learning, and the power grid regulation strategy is optimized based on the policy - evaluation architecture. The policy network is responsible for generating the optimal actions, that is, the power grid partition scheme and the calculation resource allocation ratio:

[0110] (9)

[0111] Proximal policy optimization objective:

[0112] (10)

[0113] Among them, is the policy change ratio, evaluates the superiority of the current action, is the limit range of policy change.

[0114] The evaluation network is responsible for evaluating the current state value to guide the learning of the policy network:

[0115] (11)

[0116] The evaluation network uses mean squared error for optimization:

[0117] (12)

[0118] Among them, .

[0119] During training, the policy network and the evaluation network are alternately updated - the policy network optimizes the action generation strategy, and the evaluation network optimizes the state value estimation. The two are tightly coupled through the advantage function to jointly drive the convergence of the global optimal solution for multiple objectives (delay, energy consumption). By periodically synchronizing the global model and continuously collecting real - time data, the system can adapt to the power grid load fluctuations and achieve dynamic adjustment of the partition strategy and resource allocation.

[0120] Specifically, initialize the parameters of the policy network and the evaluation network, and set the initial power grid partition scheme and calculation resource allocation scheme. 2. Based on the current state take the action , and update the power grid partition and resource allocation. 3. Calculate the reward and store the experience data (state, action, reward, next state). 4. Update the evaluation network to optimize the state value estimation. 5. Calculate the advantage estimation And update the policy network. 6. Repeat steps 2 - 5 until the algorithm converges or reaches the upper limit of the number of training rounds.

[0121] The above system initialization parameters are directly used as the input and constraint conditions for decision optimization; the actions generated by decision optimization are used for environmental interaction to generate training data; the training process uses the interaction data to update the policy and evaluation network to form a closed-loop feedback; finally, the converged policy feeds back to the power grid system to achieve the dynamic regulation goal.

[0122] Generally speaking, the advantages of the embodiments of the present invention are as follows:

[0123] Real-time partition adjustment: Dynamically divide the power grid partitions through deep reinforcement learning, and adjust the partition strategy according to real-time data such as the output fluctuation of renewable energy and load changes, to solve the problem of rigidity of traditional static partitions. 2. Elastic resource allocation: Dynamically allocate computing resources based on task complexity and node load to achieve collaborative optimization of "partition - resource", and adapt to load mutations at the second level (such as fast charging of electric vehicles and industrial impact loads). 3. Multi-objective optimization: Minimize both computing latency and energy consumption simultaneously to avoid risks of resource waste or overload. 4. Closed-loop control strategy: Form a closed-loop from data collection, task offloading to control feedback to enhance the anti-interference ability of the power grid under extreme fluctuations (such as sudden drops in wind and light, load mutations). Based on the above theory, taking an application scenario as an example, the entire practical process is described as follows: The power generation of renewable energy such as wind and light has strong volatility, resulting in frequent voltage fluctuations in the power grid. Traditional static partition strategies are difficult to adapt to dynamic load changes and are prone to cause local voltage over-limit or resource allocation imbalance. The peak-to-valley difference of the daytime load in the commercial area reaches 3:1. There is a communication delay in traditional centralized control, and the idle rate of computing resources is high during light load at night. The impact load of the steel plant causes voltage flicker, the computing power of the local controller is insufficient, and there are data security risks in outsourcing cloud computing. Therefore, the present invention is applicable to actual scenarios such as urban high-density distribution networks and industrial park microgrids.

[0124] On the other hand, the present invention also discloses a computer-readable storage medium storing a computer program, which when executed by a processor, causes the processor to execute the steps of the above method.

[0125] On yet another aspect, the present invention also discloses a computer device including a memory and a processor, where the memory stores a computer program, and when the computer program is executed by the processor, it causes the processor to execute the steps of the above method.

[0126] In another embodiment provided by the present application, there is also provided a computer program product containing instructions, which when running on a computer, causes the computer to execute any one of the power grid voltage control task offloading and computing resource allocation methods in the above embodiments.

[0127] It is understandable that the system provided by the embodiments of the present invention corresponds to the method provided by the embodiments of the present invention. For the explanations, examples, and beneficial effects of related content, reference can be made to the corresponding parts in the above method.

[0128] An embodiment of the present application further provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus.

[0129] The memory is used to store a computer program.

[0130] When the processor is used to execute the program stored on the memory, it implements the above-mentioned power grid voltage control task offloading and computing resource allocation method.

[0131] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc.

[0132] The communication interface is used for communication between the above electronic device and other devices.

[0133] The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0134] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0135] It should also be noted that the electronic device further includes a terminal device, which can also be referred to as a terminal, user equipment (UE), mobile station (MS), mobile terminal (MT), etc. The terminal device can be a mobile phone, smart TV, wearable device, tablet computer (Pad), computer with wireless transceiver function, virtual reality (VR) terminal device, augmented reality (AR) terminal device, wireless terminal in industrial control, wireless terminal in self-driving, wireless terminal in remote medical surgery, wireless terminal in smart grid, wireless terminal in transportation safety, wireless terminal in smart city, wireless terminal in smart home, and so on. The embodiments of the present application do not limit the specific technologies and specific device forms adopted by the terminal device.

[0136] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0137] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

[0138] In addition, it should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of the present invention, the directional indications are only used to explain the relative positional relationship and movement conditions between components in a specific posture. If the specific posture changes, the directional indications will also change accordingly.

[0139] In addition, if there are descriptions such as "first" and "second" involved in the embodiments of the present invention, the descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the meaning of "and / or" appearing throughout the text includes three parallel solutions. Taking "A and / or B" as an example, it includes solution A, or solution B, or the solution where A and B are satisfied simultaneously. In addition, in the embodiments of the present invention, "a plurality of" means two or more. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.

Claims

1. A method for unloading power grid voltage control tasks and allocating computing resources, characterized in that: S1. According to the number of nodes and load conditions, the grid voltage node partition model is set, and the number of nodes and computational complexity of different partitions are considered, and the computational delay of the voltage control task of each partition is calculated; S2, based on S1, obtains the computation delay of each partition and then calculates the global task computation delay; S3. After the global computing delay is obtained, the energy consumption during the task computing process is obtained; S4. After finding the energy consumption in the task calculation process, modeling is performed to minimize the delay and energy consumption; S5. Finally, the objective function and constraints are determined based on the minimization of delay and energy consumption, and the deep reinforcement learning algorithm is used to solve the grid voltage control task unloading and computing resource allocation strategy; Step S1 specifically includes: According to the number of nodes and load conditions, the nodes in the physical area are divided into several subsets, each of which is regarded as an independent task unit to independently calculate the corresponding voltage control task; Assume that there is The macro base station needs to partition the nodes. The number of nodes in ,have partitions, the partition strategy is expressed as: , (1) in, is a set of partitions, is the partition index; The specific steps of step S5 include: 1) Initialize the power grid system parameters, including: power grid nodes, base station computing resource upper limit and task delay threshold; define and initialize the deep reinforcement learning model. In the deep reinforcement learning framework, each partition The state space is: (6) in, is the total number of nodes, For node partitioning, Calculate the total amount of resources for the macro base station, Provide feedback for historical decision making; The action space is: (7) in, Indicates adjusting the number of partition nodes. Indicates adjustment of computing resource allocation; Partition The reward function is: (8) in, is a weight parameter, representing the influence weight of computing delay and computing energy consumption; The policy network for decision-making computing resource allocation strategy is ; Initialize the global model and buffer ,The global model is used to store the aggregated global parameters, and the buffer is used to temporarily store the local model parameters collected from each partition; 2) After initializing the parameters of the power grid system, proximal policy optimization is used for learning. Specifically, the power grid control strategy is optimized based on the policy-evaluation architecture. The policy network is responsible for generating the optimal action, that is, generating the power grid partitioning scheme and the computing resource allocation ratio. 3) After learning using proximal policy optimization, training and optimization are performed until the algorithm converges or the upper limit of the number of training rounds is reached.

2. The method for unloading grid voltage control tasks and allocating computing resources according to claim 1, characterized in that: Step S1 also includes, Allocate computing resources according to the number of nodes and computing complexity of different partitions to minimize task processing latency and cost; Assume the computing power of the base station is , allocate computing resources to each partition ,satisfy ,The allocation strategy affects the task execution delay and computing energy consumption. The computing delay of each partition is expressed as: (2) in, For partition The computational complexity of the voltage control task.

3. The method for unloading grid voltage control tasks and allocating computing resources according to claim 2, characterized in that: Step S2 specifically includes: Take the maximum latency in each partition as the global task calculation latency , expressed as: (3)。 4. The method for unloading grid voltage control tasks and allocating computing resources according to claim 3, characterized in that: Step S3 specifically includes: During the task calculation process, the server provides computing resources for the partitioned tasks. The energy consumption associated with this process is The computational energy consumption is : (4) in, is the energy consumption factor.

5. The method for unloading grid voltage control tasks and allocating computing resources according to claim 4, characterized in that: Step S4 specifically includes: The problem of minimizing computational delay and energy consumption is stated as follows: (5) in, and is a weight factor, used to indicate the importance of latency and energy consumption; is the delay constraint of the global voltage control task; Indicates computing resource constraints. The sum of computing resources that the server assigns to each partition must not exceed its own computing resources. Indicates the node number constraint. The sum of the number of nodes in each partition is consistent with the summary number of nodes. represents the delay constraint. The completion delay of the global voltage control task must not exceed the maximum delay tolerance of the global voltage control; Represents a weight constraint, the sum of weight factors does not exceed 1.

6. The method for unloading grid voltage control tasks and allocating computing resources according to claim 5, characterized in that: After initializing the power grid system parameters in step S5, proximal strategy optimization is used for learning. Specifically, the power grid control strategy is optimized based on the strategy-evaluation architecture. The strategy network is responsible for generating the optimal action, that is, the power grid partitioning scheme and the computing resource allocation ratio are specifically: (9) Proximal strategy optimization goals: (10) in, is the strategy change ratio, Evaluate the superiority of the current action, To limit the scope of policy changes; The evaluation network is responsible for evaluating the value of the current state and guiding the strategy network to learn: (11) The evaluation network uses mean square error optimization: (12) in, ; After learning with proximal strategy optimization, training and optimization are performed, including initializing the strategy network and evaluating network parameters, setting the initial grid partitioning scheme and computing resource allocation scheme, and Take Action , update grid partitioning and resource allocation; calculate rewards And store experience data, i.e. state, action, reward, next state; update the evaluation network, optimize the state value estimate; calculate the advantage estimate And update the policy network; repeat the above steps until the algorithm converges or reaches the upper limit of the number of training rounds.

7. A computer-readable storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 6.

8. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-target joint optimization task unloading strategy based on deep reinforcement learning in Internet of Vehicles

    CN116321298A

  • Core particle-oriented deep large model fault-tolerant deployment optimization method and system

    CN117632148A