Edge computing power grid time delay optimization method based on reinforcement learning

By employing the DDPG algorithm based on deep reinforcement learning in smart grids to optimize task offloading and power allocation, the problems of latency and power constraints in collaborative computing scenarios are solved, thereby improving system performance and achieving adaptive optimization.

CN121924499APending Publication Date: 2026-04-24ANSHAN POWER SUPPLY COMPANY OF STATE GRID LIAONING ELECTRIC POWER COMPANY +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ANSHAN POWER SUPPLY COMPANY OF STATE GRID LIAONING ELECTRIC POWER COMPANY
Filing Date
2025-12-11
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing technologies struggle to adapt to changes in the network environment in smart grid collaborative computing scenarios, causing task offloading and power allocation optimization to fall into local optima, failing to meet latency and power constraints, and impacting system performance.

Method used

We adopt an edge computing method based on deep reinforcement learning. By constructing a deep deterministic policy gradient (DDPG) algorithm framework and combining a data transmission model and a latency model, we optimize the task offloading ratio and transmission power, and establish a comprehensive reward function to meet the constraints of latency, power and task allocation ratio.

Benefits of technology

It enables adaptive optimization of task offloading and power allocation in a dynamic power grid environment, reduces total system latency, takes power consumption into account, and improves the overall service quality and operating efficiency of the edge computing power grid.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121924499A_ABST
    Figure CN121924499A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent power grid communication, in particular to an edge computing power grid time delay optimization method based on reinforcement learning, which comprises the following steps: data acquisition and preprocessing: multi-target time delay optimization modeling: establishing a data transmission model, a time delay model and constraint conditions; the joint optimization of the task unloading proportion and the transmitting power is converted into a quantifiable optimization problem; multi-objective optimization solution based on deep reinforcement learning: adopting a deep deterministic strategy gradient (DDPG) algorithm framework, and learning an optimal task allocation and power allocation strategy through interaction with the environment; and outputting and executing an optimization decision. The method has the advantages that the task unloading proportion and the transmitting power of each link can be jointly optimized under a unified model, and compared with a scheme which only aims at a single index or adopts a simple weighted summation mode, the method is beneficial to reducing the total time delay of the system on the premise of meeting constraint conditions, gives consideration to power consumption, and improves the reliability of the system. Therefore, the comprehensive service quality in the edge computing power grid scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart grid communication technology, and in particular to an edge computing method for optimizing grid latency based on reinforcement learning. Background Technology

[0002] With the continuous evolution of communication technologies, a large number of new applications with high latency sensitivity and stringent reliability requirements have emerged in the power industry, such as real-time monitoring, fault location, and protection control in smart grids. These applications typically require data acquisition, transmission, and processing to be completed within milliseconds or less to ensure the safe and stable operation of the power system. Existing solutions mostly rely on remote cloud computing centers to complete complex computing tasks. While cloud computing centers possess powerful computing capabilities, their long transmission paths and high risk of network congestion easily introduce significant latency, making it difficult to meet the low latency and high reliability requirements of these new smart grid applications.

[0003] To address these challenges, Mobile Edge Computing (MEC) technology has been introduced into power communication systems. By moving computing and storage resources from remote cloud centers to edge nodes closer to power equipment, MEC enables some data processing to be performed locally or locally. This effectively shortens the data transmission path, reduces the impact of network congestion, and significantly reduces end-to-end latency and improves system response speed. The ubiquitous sensing, adaptive scheduling, and intelligent collaboration capabilities of MEC provide crucial technical support for building an efficient and flexible intelligent power grid system.

[0004] However, introducing edge computing into smart grids also brings new challenges. In complex grid operation environments, multiple power devices often need to collaborate to complete specific tasks, such as collaborative fault diagnosis between adjacent monitoring terminals and joint criterion calculation between protection devices. In this case, tasks can be processed not only by local devices but also through direct-to-device (D2D) communication or offloaded to MEC servers via base stations, forming a multi-path, multi-node collaborative computing model. In these scenarios, the task offloading ratio is coupled with the transmission power of each link. How to achieve optimal global resource allocation while satisfying latency and power constraints becomes a key issue in edge computing for smart grids. Existing research often employs heuristic algorithms and rule-driven scheduling strategies for resource allocation. While these methods are relatively simple to implement, they heavily rely on expert experience, are prone to getting trapped in local optima, and lack the ability to continuously adapt and evolve in dynamic environments, making it difficult to adapt to rapid changes in grid conditions and business needs.

[0005] Deep reinforcement learning, as an important branch of machine learning, learns optimal policies through continuous interaction between an agent and its environment, employing a trial-and-error approach. It possesses end-to-end feature extraction and multi-objective collaborative optimization capabilities. Its core idea is to continuously adjust policy parameters based on given state information and immediate reward feedback to maximize long-term cumulative rewards. Compared to traditional heuristic methods, deep reinforcement learning can search for policies in continuous state and action spaces, making it suitable for handling complex decision-making problems where task allocation ratios and power allocations are continuously adjustable. Furthermore, deep reinforcement learning can update policies online under constantly changing network environments and workloads, maintaining high decision-making accuracy and robustness even in power grid edge computing systems with random channels, random task arrivals, and time-varying resource constraints. Therefore, applying deep reinforcement learning to latency optimization problems in smart grid edge computing scenarios has significant technical potential and application value.

[0006] In summary, existing technologies generally suffer from insufficient real-time performance, susceptibility to local optima, and weak adaptability when addressing latency optimization issues in smart grid collaborative computing. There is an urgent need for a novel technical solution that can adapt to changes in the network environment, achieve joint optimization of task offloading and power allocation, and improve overall service quality while meeting latency and power constraints. Summary of the Invention

[0007] The purpose of this invention is to provide a reinforcement learning-based edge computing power grid latency optimization method to solve the problem of jointly optimizing the task offloading ratio and the transmission power of multiple communication links in the scenario of collaborative computing of power equipment. By constructing a deep reinforcement learning decision model at the edge, and considering latency constraints, power constraints and task allocation ratio constraints, the method adaptively adjusts task allocation and power allocation, so as to significantly reduce the total system latency while taking into account power consumption, thereby improving the overall performance and operating efficiency of the smart grid edge computing system.

[0008] To achieve the above objectives, the present invention provides the following technical solution: A reinforcement learning-based edge computing method for optimizing power grid latency includes: S1. Data Acquisition and Preprocessing: Real-time acquisition of power grid communication-related data; cleaning, normalization, and feature extraction of the power grid communication-related data to form a standardized state vector; S2. Multi-objective time delay optimization modeling: Establish data transmission model, latency model, and constraints to transform the joint optimization of task offloading ratio and transmit power into a quantifiable optimization problem; S3. Multi-objective optimization solution based on deep reinforcement learning: We adopt the Deep Deterministic Policy Gradient (DDPG) algorithm framework, which learns the optimal task allocation and power allocation strategy through interaction with the environment. S4. Optimize decision output and execution.

[0009] In S1, the power grid communication-related data includes channel gain, equipment computing power, and task characteristics; Channel gain includes: channel state information between two power devices. Channel state information between power equipment A and the base station Channel state information between power equipment B and the base station ; The equipment's computing capabilities include: the local computing frequency of the two power devices. The computing frequency of mobile edge computing servers, also known as MEC servers. ; Task characteristics include: the total amount of task data generated by power equipment. The computational complexity of tasks generated by power equipment .

[0010] In S2, the data transmission model includes the D2D communication transmission rate and the device-to-base station transmission rate, and is modeled based on Shannon's formula. The latency model includes D2D transmission latency, base station transmission latency, edge server processing latency, and local device processing latency. The total latency is the sum of the maximum value of the latency of each path and the processing latency. The constraints include latency constraints, power constraints, and task allocation ratio constraints; Establishing an optimization problem involves integrating the objective and constraints into a complete mathematical optimization problem.

[0011] The D2D communication transmission rate is expressed as: (1); In formula (1): Indicates bandwidth, unit is ; Noise power spectral density, in units of ; The signal transmission power for D2D communication between two electrical devices, in units of ; This refers to the channel status information between two electrical devices; This refers to the D2D communication transmission rate, measured in units of... ; The transmission rate from the device to the base station is expressed as: (2); (3); In formulas (2) and (3): The transmission power of power device A sending the portion of the task data allocated to the MEC server for processing to the MEC server, in units of ; This refers to the channel state information between power equipment A and the base station; The transmission power (in units of ) for power equipment B to send the portion of the task data allocated to the MEC server for processing to the MEC server. ; This refers to the channel state information between power equipment B and the base station; The uplink transmission rate from power device A to the base station is expressed in units of 1. ; The uplink transmission rate from power equipment B to the base station is expressed in units of 1. .

[0012] D2D transmission delay, expressed as: (4); In formula (4): The total amount of task data generated by power equipment, in units of ; This represents the proportion of task data transmitted from the power equipment to the peer equipment, with a value range of [value range missing]. ; This refers to the D2D transmission delay, measured in units of... ; Base station transmission delay, expressed as: (5); In formula (5): The transmission delay from power equipment A to the base station is expressed in units of 1. ; The transmission delay from power equipment B to the base station is expressed in units of 1. ; This represents the proportion of task data transmitted from power equipment to the MEC server, with a value range of [value missing]. ; The edge server handles latency, expressed as: (6); In formula (6): The computing frequency of the MEC server, in units of ; Latency processing for edge servers, in units of ; The computational complexity of tasks generated by power equipment, in units of ; The local device processing latency is expressed as: (7); In formula (7): The calculated frequency for two electrical devices, in units of ; Local device processing latency, in units of ; Total delay, expressed as: (8); In formula (8): Total delay, in units of .

[0013] Delay constraints are constraints on the completion of tasks by electrical equipment within a specified time, i.e., there is a maximum delay limit, expressed as: (9); In formula (9), The maximum allowable delay for the task, in units of ; Power constraints mean that the transmit power of each electrical device cannot exceed the maximum power allowed by its hardware. The expression is: (10); The task allocation ratio constraint sets the total data ratio of each power device to 1. It is divided into two parts: one part is transmitted to the peer for comparison and analysis, and the other part is transmitted to the MEC server for comparison and analysis. The expression is: (11); The objective and constraints are integrated into a complete mathematical optimization problem. The objective function aims to minimize the total data transmission and processing delay between the power equipment and the MEC server, and its expression is: (12); The constraints are time delay constraint, power constraint, and task allocation ratio constraint, expressed as: (13).

[0014] In S3, the DDPG algorithm framework includes state space design, continuous action space design, and reward function design. 1) State-space design: State-space design includes three dimensions: communication channel state, device computing state, and task characteristics, expressed as: (14); In formula (14), This represents the task allocation decision and power allocation made at the previous moment; 2) Continuous motion space design: The continuous action space includes the task allocation ratio and the transmit power of each link, which are continuous variables, and the expression is: (15); In formula (15), the action The decisions are made by the intelligent agent, which is a decision-making module / control policy learning module built based on the DDPG algorithm and deployed on an edge server or control center. 3) Reward function design: The reward function includes a basic reward term and a constraint penalty term. The basic reward is negatively correlated with the total latency, and the penalty term is used to address constraint violations. Considering power efficiency, a negative reward term for power consumption is introduced to explicitly penalize high-power strategies during reinforcement learning. This encourages the agent to choose scheduling schemes with lower power consumption while meeting latency constraints, thus demonstrating the technical effect of this invention in improving energy efficiency and reducing power consumption while ensuring service latency performance. The expression is as follows: (16); In formula (16): ; ; These are weighting coefficients used to balance the importance of time delay and power, satisfying... ; Total power consumption, in units of ; The constraint penalty term is expressed as follows: (17); In formula (17): ; The penalty coefficient is much larger than the typical value of the base reward; The value is 1 if the indicator function condition is true, and 0 otherwise. The value of the constraint penalty item; 4) The final reward is the sum of the basic reward and the constraint penalty, expressed as: (18); In formula (18): This is the final total reward value; Basic reward value; The value of the constraint penalty item.

[0015] In S3, the training process of the Deep Deterministic Policy Gradient (DDPG) algorithm is as follows: An experience replay mechanism is used to store and sample training data; Exploration was conducted using Ornstein-Uhlenbeck process noise; Update the target network parameters using a soft update mechanism; During training, the Critic network is updated using mean squared error loss, and the Actor network is updated using policy gradient.

[0016] In S4, optimizing decision output and execution includes: The action vectors output by the Actor network are parsed into specific task allocation ratios and power control instructions. The optimization effect is evaluated through key performance indicators, including latency optimization rate, power efficiency improvement, and constraint satisfaction rate.

[0017] Compared with the prior art, the beneficial effects of the present invention are: 1. By constructing a comprehensive reward function that simultaneously considers indicators such as total system latency and transmit power, and explicitly integrating latency constraints, power constraints, and task allocation ratio constraints into the reinforcement learning framework, this invention can jointly optimize the task offloading ratio and transmit power of each link under a unified model. Compared with schemes that only target a single indicator or use a simple weighted summation method, this invention is beneficial to reduce total system latency and take power consumption into account while meeting the constraints, thereby improving the overall service quality in edge computing power grid scenarios. 2. This invention employs a deep reinforcement learning agent, which learns task offloading and power allocation strategies by repeatedly interacting with a near-real power grid communication and computing environment. It does not rely on a completely accurate analytical mathematical model or complex channel statistical priors. When the operating environment, such as channel conditions and service load, changes, the agent can adaptively adjust its decisions based on the new observation state. This makes the proposed method more adaptable and robust than schemes based on fixed rules or static optimization in dynamic power grid scenarios. 3. This invention constructs an optimization solver in a continuous action space based on the DDPG algorithm, uses a deep neural network to represent the state of a high-dimensional system, and directly performs policy search in the continuous action space through deterministic policy gradient. Compared with methods that discretize continuous control quantities before optimization or rely on partial heuristic search, this invention can reduce the quantization error caused by action discretization, reduce the risk of getting trapped in local suboptimal solutions, and more easily obtain task offloading and power allocation schemes that meet the overall performance requirements of the system. 4. At the algorithm level, stability mechanisms such as experience replay, soft update of the target network, and exploration noise are introduced, and preprocessing steps such as data standardization and outlier handling are combined to ensure stable convergence of the training process. In addition, this method does not rely on an accurate analytical model, making it easier to deploy and apply in complex real-world power grid environments, and has good engineering application value and promotion prospects. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the overall system flow for a reinforcement learning-based edge computing smart grid latency optimization method.

[0019] Figure 2 This is a schematic diagram of a collaborative computing scenario for smart grids.

[0020] Figure 3 This is a schematic diagram of the reinforcement learning training process.

[0021] Figure 4 This is the network structure diagram of the DDPG algorithm.

[0022] Figure 5 This is a comparison chart of latency performance in the embodiments. Detailed Implementation

[0023] The present invention will now be described in detail with reference to the accompanying drawings, but it should be noted that the implementation of the present invention is not limited to the following embodiments.

[0024] The following embodiments are implemented based on the technical solution of the present invention, providing detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments. Unless otherwise specified, the methods used in the following embodiments are conventional methods.

[0025] Example 1 A reinforcement learning-based edge computing method for optimizing power grid latency, see [link to relevant documentation]. Figure 1 ,include: Step S1, Data Acquisition and Preprocessing: Real-time acquisition of power grid communication-related data specifically includes periodically collecting channel status information, such as channel gain, through sensing modules or network management systems in the power communication network; collecting the computing resource status of each power device, such as the currently available computing frequency (CPU cycles); and collecting feature data of the computing tasks to be processed, such as the task data volume (number of bits) and computational complexity (number of CPU cycles required per bit). The collected raw data is preprocessed, including data cleaning (removing outliers), normalization (scaling data of different dimensions to a similar range to facilitate neural network model training), and feature extraction, forming a standardized state vector that can be used by the reinforcement learning model.

[0026] Channel gain is a key parameter for measuring the quality of a communication link, and it mainly includes the channel state information between two power devices. Channel state information between power equipment A and the base station Channel state information between power equipment B and the base station .

[0027] Equipment computing power refers to the number of instructions that equipment can process per unit of time, mainly including the local computing frequency of the two power devices. And the computing frequency of mobile edge computing servers, i.e., MEC servers. .

[0028] Task characteristics describe the nature of the computational task to be processed, mainly including the total amount of task data generated by power equipment. And the computational complexity of processing this task. .

[0029] Step S2: Multi-objective delay optimization modeling. Establish an accurate mathematical model that reflects the actual operating conditions of the system, transforming delay constraints, power constraints, and task allocation constraints into quantifiable optimization problems. This model is the foundation for subsequent reinforcement learning algorithm design, and its accuracy directly affects the final optimization effect. Therefore, this model is established in the context of smart grid collaborative computing, as shown in Figure 2. Specifically, it includes: establishing a data transmission model, a delay model, and constraints, transforming the joint optimization of task offloading ratio and transmission power into a quantifiable optimization problem. S21. Establish a data transmission model to calculate the data transmission rate on each communication link; In this information exchange system, power device A and power device B are configured to use specific D2D technology for information transmission, avoiding the roundabout transmission of data via base stations and thus reducing latency. Device A can directly send some data to device B for processing, and vice versa. The instantaneous transmission rate of this D2D communication process is determined by Shannon's formula, expressed as: (1); In formula (1): Indicates bandwidth, unit is ; Noise power spectral density, in units of ; The signal transmission power for D2D communication between two electrical devices, in units of ; This refers to the channel status information between two electrical devices; This refers to the D2D communication transmission rate, measured in units of... ; 2) Both power devices also need to upload another portion of data to the MEC server for processing; the upload process passes through the base station, and their uplink transmission rates are calculated separately. This includes the transmission rate from power device A to the base station and the transmission rate from power device B to the base station, expressed as: (2); (3); In formulas (2) and (3): The transmission power of power device A sending the portion of the task data allocated to the MEC server for processing to the MEC server, in units of ; This refers to the channel state information between power equipment A and the base station; The transmission power (in units of ) for power equipment B to send the portion of the task data allocated to the MEC server for processing to the MEC server. ; This refers to the channel state information between power equipment B and the base station; The uplink transmission rate from power device A to the base station is expressed in units of 1. ; The uplink transmission rate from power equipment B to the base station is expressed in units of 1. .

[0030] S22. Establish a latency model to calculate the total time taken from the start of transmission to the completion of processing of the task; The latency model includes D2D transmission latency, base station transmission latency, edge server processing latency, and local device processing latency. The total latency is the sum of the maximum value of the latency of each path and the processing latency. 1) The time delay for data transmission between two electrical devices via D2D technology is expressed as: (4); In formula (4): The total amount of task data generated by power equipment, in units of ; This represents the proportion of task data transmitted from the power equipment to the peer equipment, with a value range of [value range missing]. ; This refers to the D2D transmission delay, measured in units of... ; 2) Base station transmission delay includes the transmission delay from power equipment A to the base station and the transmission delay from power equipment B to the base station, expressed as: (5); In formula (5): The transmission delay from power equipment A to the base station is expressed in units of 1. ; The transmission delay from power equipment B to the base station is expressed in units of 1. ; This represents the proportion of task data transmitted from power equipment to the MEC server, with a value range of [value missing]. ; 3) Edge server processing latency, expressed as: (6); In formula (6): The computing frequency of the MEC server, in units of ; Latency processing for edge servers, in units of ; The computational complexity of tasks generated by power equipment, in units of ; 4) Local device processing latency, expressed as: (7); In formula (7): The calculated frequency for two electrical devices, in units of ; Local device processing latency, in units of ; 5) The total latency is the sum of the data transmission and processing latency of the two power devices A and B, and the data transmission and processing latency of the MEC server. For the MEC server, it needs to compare the data from the two power devices. Since the task data of both power devices must be transmitted completely before the next step of calculation and analysis can proceed, the total transmission time from the power devices to the base station is taken as the maximum of the transmission times of the two devices. The expression is: (8); In formula (8): Total delay, in units of .

[0031] S23. Constraints include delay constraints, power constraints, and task allocation ratio constraints; 1) Time delay constraint means that the task of the power equipment needs to be completed within a specified time, that is, there is a maximum time delay limit, expressed as: (9); In formula (9), The maximum allowable delay for the task, in units of ; 2) Power constraints mean that the transmission power of each electrical device cannot exceed the maximum power allowed by its hardware. In this information exchange system, the total transmission power of each power device is divided into two parts: one part is used to send data to the base station, and the other part is used to send data to the peer via D2D communication technology. The power of both parts must not exceed the maximum transmission power, as expressed by: (10); 3) The task allocation ratio constraint sets the total data ratio of each power device to 1. It is divided into two parts: one part is transmitted to the peer for comparison and analysis, and the other part is transmitted to the MEC server for comparison and analysis. The expression is: (11); S24. Establishing an optimization problem involves integrating the objective and constraints into a complete mathematical optimization problem, where: The objective and constraints are integrated into a complete mathematical optimization problem. The objective function aims to minimize the total data transmission and processing delay between the power equipment and the MEC server, and its expression is: (12); The constraints are time delay constraint, power constraint, and task allocation ratio constraint, expressed as: (13).

[0032] Step S3, the core process of solving the above complex optimization problem based on deep reinforcement learning, see... Figure 3 Specifically, it includes: Step S3.1: Construct a reinforcement learning environment; The reinforcement learning environment is established based on the collaborative computing scenario of smart grids, and the state space, action space and state transition mechanism are fully defined.

[0033] S3.1.1: Design State Space The state space comprises three dimensions: communication channel state, device computing state, and task characteristics, expressed as: (14); In formula (14), This represents the task allocation decision and power allocation made at the previous moment; S3.1.2: Designing the motion space: action This is the decision made by the agent. Since the task allocation ratio and power value are continuous variables, the action space is continuous, and the expression is: (15); In formula (15), the action The decisions made by the agent are based on the decision-making module / control policy learning module built on the DDPG algorithm and deployed on the edge server or control center. These actions need to meet certain constraints; at the output of the neural network, appropriate activation functions (such as the Sigmoid function for scaling and normalized power) can be used to ensure that the action values ​​fall within the legal range. The constraints are as follows: .

[0034] S3.1.3: Define the state transition function State transitions are described by Markov decision processes, i.e., the next state. From the current state and current action Decisions are made and are subject to environmental randomness. The effects (such as channel fading, arrival of new tasks) can be expressed as: ; in: This is the state transition function; The environmental randomness vector includes, but is not limited to: random fast and slow fading variations in channel gain, uncertainty in the arrival process of tasks, and random factors such as load fluctuations of available computing resources on the device. This represents the system state vector at the current moment. This is the system state vector for the next time step; This is the action vector at the current moment.

[0035] Step S3.2 Construct the agent model; The DDPG (Deep Deterministic Policy Gradient) algorithm is used as the core architecture of the intelligent agent. (See...) Figure 4 It contains four neural networks: the current Actor network, the target Actor network, the current Critic network, and the target Critic network, specifically including: S3.2.1: Design the Actor network structure; Actor Network It is a deterministic policy network with input state Output action This invention employs a deep neural network structure, expressed as follows: ; in: A deterministic policy function implemented for the Actor network; This is the parameter vector (weights and biases) of the Actor network. Network forward propagation, expressed as: ; in: These are the parameters of the Actor network; It is the ReLU activation function; For the Actor network Layer weight matrix; For the Actor network The bias vector of the layer.

[0036] S3.2.2: Design the Critic network structure; Critic Network It is an action-value function network, with the state as the input. and actions Output a scalar Q value to evaluate the state. Next action The expression for good or bad is: ; in: The action value function output by the Critic network; These are the parameters for the Critic network.

[0037] Network forward propagation, expressed as: ; in: These are the parameters for the Critic network. This involves concatenating the state and action vectors. For the Critic network Layer weight matrix; For the Critic network The bias vector of the layer.

[0038] S3.2.3: Design the target network and soft update mechanism; To ensure stable training, DDPG introduces a target network; the target networks for Actor and Critic (with parameters...) and The structure is the same as the current network; parameter updates use a soft update method instead of direct copying, and the expression is: ; in: This is the soft update coefficient, usually set to 0.001, which allows the target network parameters to change slowly, greatly improving the stability of the learning process. This is the parameter vector (weights and biases) of the current network for the Actor. The parameter vector of the Actor target network; This is the parameter vector of the current Critic network; This is the parameter vector of the Critic target network.

[0039] Step S3.3: Design the reward function; reward function This is crucial for guiding the agent's learning; it requires accurately and efficiently transforming the optimization objective into signals that the agent can understand. The design process must comprehensively consider latency performance, power efficiency, and constraint satisfaction, specifically including: S3.3.1: Reward Function Design: Define the basic reward function: The base reward is directly negatively correlated with the optimization objective—total latency; the lower the latency, the higher the reward. Simultaneously, considering power efficiency, a negative reward term for power consumption is introduced, expressed as: (16); In formula (16): ; ; These are weighting coefficients used to balance the importance of time delay and power, satisfying... ; Total power consumption, in units of ; S3.3.2: Design of constraint penalty terms; To ensure that the policy learned by the agent satisfies the constraints, a large negative reward (penalty) is given when the constraints are violated, expressed as: (17); In formula (17): ; The penalty coefficient is much larger than the typical value of the base reward; The value is 1 if the indicator function condition is true, and 0 otherwise. The value of the constraint penalty item; S3.3.3: Determine the final reward function; The final reward is the sum of the base reward and the constraint penalty, expressed as: (18); In formula (18): This is the final total reward value; Basic reward value; The value of the constraint penalty item.

[0040] Step S3.4: Training the agent. This is a process of iteratively optimizing network parameters through a large amount of data, specifically including: S3.4.1: Establishment of the Experience Replay Mechanism The agent will use the experience tuples from each interaction step. Stored in a fixed-size playback buffer In training, a mini-batch of experience is randomly sampled from the buffer for network updates. This approach breaks down the correlation between data points, improving sample efficiency and learning stability. To further improve efficiency, Prioritized Experience Replay can be used, which adjusts the sampling probability based on the importance of the experience (such as the absolute value of the Temporal Difference Error, TD-Error). Therefore: (19); in For TD error, This is the priority coefficient.

[0041] S3.4.2: Critic Network Update The objective of the Critic network is to minimize the temporal difference error. value The formula is calculated using the target network as follows: (20); in It is a discount factor that measures the importance of future rewards. Then, by minimizing the current... Values ​​and Objectives value The mean squared error between the two is used to update the Critic network: (twenty one); Gradient descent update: (twenty two); is the learning rate for Critic.

[0042] S3.4.3: Actor Network Update The update objective of the Actor network is to maximize the expected cumulative reward, i.e. Value. Using the policy gradient theorem, the gradient provided by the Critic network is used to update the Actor network: (twenty three); Parameter update: (twenty four); The learning rate of the Actor is typically less than .

[0043] S3.4.4: Exploration Strategy; To allow for thorough exploration of the environment in the early stages of learning, exploratory noise needs to be added to the deterministic actions output by the Actor network; this invention employs time-dependent Ornstein-Uhlenbeck (OU) process noise. It can generate inertial exploration that corresponds to physical processes (such as channel changes), expressed as: (25); The OU process satisfies: (26); in: It is a Wiener process, and the noise amplitude can be gradually reduced as training progresses; The regression velocity coefficient for OU noise; This represents the long-term mean of the OU noise. OU noise intensity coefficient; This refers to noise in the OU process.

[0044] S3.4.5: Training algorithm flow. The complete DDPG algorithm for smart grid latency optimization is as follows: 1) Initialization: Randomly initialize the Critic network and Actor Network ; Initialize the target network ; Initialize the experience replay buffer ; 2) For each episode: Initialize the environment and obtain the initial state. For each time step: : a. Select Action b. Execution of actions Observation and reward and new status c. Storage experience arrive d. Randomly sample mini-batch experience e. Update the Critic network: f. Update the Actor network: g. Update the target network: 3) End training: When the policy converges or the maximum number of training rounds is reached. S3.4.6: Convergence analysis; The DDPG algorithm can converge to a local optimum under certain conditions (such as a sufficiently small learning rate and a sufficiently strong function approximator).

[0045] Define the value function as follows: ; in: In strategy Next state State value function; For instant rewards; Discount factor; The strategy improvement theorem guarantees that the expression is: ; in: For the first Substitution strategy; For the first Substitution strategy; When the policy gradient When the algorithm converges to a local optimum, it will eventually converge to a local optimum.

[0046] S3.4.7: Hyperparameter settings; Learning rate: ; Discount factor: ; Soft update coefficient: ; Batch size: ; Buffer capacity: ; Exploration parameters: ; A smart grid delay optimization system based on reinforcement learning, according to the present invention, includes: An edge computing power grid delay optimization system based on reinforcement learning includes: a data acquisition and preprocessing module, a delay optimization modeling module, an optimization solution module, and a decision execution module; 1) The data acquisition and preprocessing module is used to acquire and process channel status, device status, and task data in real time; 2) The latency optimization modeling module is used to construct data transmission models, latency models, and constraints; 3) The optimization solution module is used to train the agent based on the DDPG algorithm and output the optimization policy; the optimization solution module includes: state space construction unit, action space construction unit, reward function design unit, and network training unit; State space building blocks are used to define system states; Action space building blocks are used to define continuous control actions; The reward function design unit is used to construct multi-objective reward mechanisms; The network training unit is used to implement the training and convergence of the DDPG algorithm; 4) The decision execution module is used to distribute optimization strategies to power equipment and edge servers for execution.

[0047] Example 2 In this embodiment, the edge computing power grid latency optimization method based on reinforcement learning is the same as in Embodiment 1, but with the addition of the overall system flow of the edge computing smart grid latency optimization method based on reinforcement learning.

[0048] The overall system flow of the edge computing-based smart grid latency optimization method based on reinforcement learning starts with data acquisition and preprocessing, covering key stages such as system modeling, reinforcement learning framework construction, optimization decision generation and execution. Subsequently, intelligent decision-making is achieved through deep reinforcement learning algorithms, with particular emphasis on the advantages of the DDPG algorithm in continuous action spaces. In the technical implementation section, a verification mechanism, such as model comparison and convergence analysis, is constructed to ensure the reliability of the optimization results. Finally, in the application verification phase, specific embodiments demonstrate the effectiveness and superiority of the method in smart grid monitoring scenarios. The method includes the following steps: Step 1: System Model Construction In the edge computing environment of smart grids, latency optimization is a complex multi-objective optimization problem. To accurately describe this problem, this invention establishes a systematic model that includes a data transmission model, a latency model, and constraints, comprehensively characterizing the latency formation mechanism in smart grid collaborative computing scenarios.

[0049] System model construction includes the following three steps: I. Data Transmission Model: This model is used to calculate the data transmission rate on each communication link and is the basis for latency calculation.

[0050] (1) D2D communication transmission rate; The D2D communication transmission rate between device A and device B is determined by Shannon's formula, which is expressed as: in; Bandwidth, unit: ; Noise power spectral density, in units of ; The signal transmission power for D2D communication, in units of ; Channel gain between devices; This refers to the D2D communication transmission rate, measured in units of... .

[0051] (2) Transmission rate from device to base station; The transmission rate from device A to the base station is expressed as: The transmission rate from device B to the base station is expressed as: ; in: and These represent the transmit power from the device to the base station, in units of... ; and For the corresponding channel gain; and These represent the uplink transmission rates from the device to the base station, in units of... .

[0052] II. Delay Model: This model is used to calculate the total time a task takes from the start of transmission to completion of processing.

[0053] (1) D2D transmission delay; The delay in D2D communication between devices is expressed as: ; in: The amount of data for the task, in units of ; This represents the proportion of task data transmitted from power equipment to the peer equipment; it is a dimensionless quantity with a value range of [value range missing]. ; This refers to the D2D communication transmission rate, measured in units of... ; This refers to the D2D transmission delay, measured in units of... .

[0054] (2) Base station transmission delay; Transmission latency from device to base station: ; The transmission delay from power equipment A to the base station is expressed in units of 1. ; The transmission delay from power equipment B to the base station is expressed in units of 1. ; The proportion of tasks allocated to MEC servers, with a value range of [value missing]. ; The uplink transmission rate from power device A to the base station is expressed in units of 1. ; The uplink transmission rate from power equipment B to the base station is expressed in units of 1. .

[0055] (3) Handling latency Edge server processing latency: ; Latency processing for edge servers, in units of ; The computing frequency of the MEC server, in units of .

[0056] Local device processing latency: ; Local device processing latency, in units of ; The calculated frequency for two electrical devices, in units of (4) Total delay; ; Total delay, in units of .

[0057] III. Constraints: The actual physical constraints that the optimization problem must satisfy.

[0058] (1) Delay constraint: The total delay cannot exceed the maximum allowable delay, expressed as: ; in: The maximum allowable delay for the task, in units of ; (2) Power constraint: The transmission power cannot exceed the maximum power of the equipment, expressed as: ; in: The transmission power of each electrical device must not exceed the maximum power allowed by its hardware, in units of... ; (3) Task allocation constraints: The task allocation ratio must meet the following requirements: .

[0059] in: The proportion of tasks allocated to MEC servers.

[0060] Step 2: Building a Reinforcement Learning Framework Reinforcement learning frameworks are a core solution to latency optimization problems, learning optimal policies through the interaction between agents and their environment. This framework includes state space design, action space design, reward function construction, and the implementation of the DDPG algorithm.

[0061] I. State-space design: The state space comprises three dimensions: communication channel state, device computing state, and task characteristics. ; in: Channel gain; To calculate the frequency; and As a task characteristic; and This is a historical decision value.

[0062] II. Motion space design: The action space is continuous and includes task allocation and power allocation decisions, expressed as: ; Activation functions such as Sigmoid are used to ensure that action values ​​fall within the legal range.

[0063] III. Reward Function Design: The reward function guides the agent's learning, taking into account both latency performance and constraint satisfaction.

[0064] (1) The basic reward function is expressed as: ; in: and For the weighting coefficients, satisfying ; Basic reward value; The maximum acceptable latency for the task, in units of ; Total power consumption, in units of ; The transmission power of each electrical device must not exceed the maximum power allowed by its hardware, in units of... .

[0065] (2) Constraint penalty term, the expression is: ; in: This is the penalty coefficient; For indicator functions; The value of the constraint penalty item.

[0066] (3) The total reward function is expressed as: .

[0067] IV. DDPG Algorithm Implementation: The DDPG algorithm includes an Actor network, a Critic network, a target network, and an experience replay mechanism.

[0068] (1) The Actor network structure is expressed as: ; in: A deterministic policy function implemented for the Actor network; These are the parameters of the Actor network.

[0069] (2) The Critic network structure is expressed as: ; in: The action value function output by the Critic network; These are the parameters for the Critic network.

[0070] (3) Target network soft update, the expression is: .

[0071] in: This is the parameter vector (weights and biases) of the current network for the Actor. The parameter vector of the Actor target network; This is the parameter vector of the current Critic network; For the parameter vector of the Critic target network; This is the soft update coefficient.

[0072] Step 3: Optimize Decision Making and Quantification Optimization decision-making and quantification are the core steps to achieving latency optimization. The optimal decision is generated through a well-trained reinforcement learning model, and the optimization effect is quantitatively evaluated.

[0073] I. Decision Generation: Input the current state into the Actor network to obtain the optimal action decision: ; in: The optimal action instruction; For policy functions; It is a state vector; These are the policy network parameters.

[0074] Post-process the output actions to ensure that the constraints are met.

[0075] II. Decision Implementation: Decompose the motion vector into specific control commands: task allocation ratio and power distribution scheme , , The command is then distributed to the relevant power equipment and edge servers for execution.

[0076] III. Quantification of Results: Evaluate the optimization effect using key performance indicators: (1) Delay optimization rate: ; in: For latency optimization rate; The average total delay of the baseline scheme is given in units of 1. ; To optimize the average total latency of the solution, the unit is... .

[0077] (2) Improved power efficiency: ; in: For power efficiency improvement rate; The average total power of the baseline scheme is expressed in units of... ; To optimize the average total power of the scheme, the unit is... .

[0078] (3) Constraint satisfaction rate: .

[0079] in: The constraint satisfaction rate; This represents the total number of tasks. The number of tasks to satisfy the constraints.

[0080] Step 4: Algorithm Flow and Model Validation Algorithm flow and model verification are key steps to ensure the reliability of latency optimization methods. Rigorous training processes and verification mechanisms are used to guarantee model performance.

[0081] I. Data Preprocessing: Standardize the raw data to eliminate dimensional differences: ; in: This is the original data; The mean; Standard deviation; This is the data after standardization.

[0082] A sliding window is used to smooth and suppress noise, and outliers are identified and corrected based on the IQR method.

[0083] II. Model Training Process: S1. Initialize the Critic network and Actor Network The parameters.

[0084] S2. Initialize target network parameters: .

[0085] S3. Initialize the experience playback buffer .

[0086] S4.Forepisode=1toM: 1. Initialize the environment and obtain the initial state. .

[0087] 2. Fort = 1 to T: Select actions based on the current strategy and noise exploration. .

[0088] 2. Perform the action Environmental feedback rewards and new status .

[0089] 3. Experience Store in buffer .

[0090] 4. From The experience of randomly sampling a small batch.

[0091] 5. Update the Critic and Actor networks.

[0092] 6. Target network soft update.

[0093] 3. EndFor. S5.EndFor. III. Convergence Analysis: The convergence of the algorithm is evaluated using the reward curve and performance metrics: . IV. Performance Verification: (1) Benchmark Comparison: Compare latency performance with traditional optimization algorithms; (2) Ablation experiment: to verify the contribution of each module; (3) Robustness test: Test the model's adaptability under different network conditions. Example 3 In this embodiment, the edge computing power grid latency optimization method based on reinforcement learning is the same as in Embodiment 1, but with the addition of a specific application of the edge computing smart grid latency optimization method based on reinforcement learning on the basis of Embodiment 1 and / or Embodiment 2.

[0094] In a smart grid monitoring scenario, latency optimization is performed for two power devices working collaboratively.

[0095] The system parameters are set as follows: bandwidth: ; Noise power spectral density: ; Maximum transmit power: ; Task data volume: ; Computational complexity: ; Local computing frequency: ; Edge server computing frequency: ; Delay constraints: .

[0096] After adopting a latency optimization method based on deep reinforcement learning: First, the system collects channel status information in real time. The device calculates its status and performs preprocessing. Then the modeling module establishes a delay optimization model, including transmission rate calculation, delay calculation, and constraints; Next, the optimization and solution module uses the DDPG algorithm to train the agent to learn the optimal task allocation ratio. and power distribution scheme .

[0097] During training, the Actor network uses a three-layer fully connected neural network (128-64-32 nodes), and the Critic network uses a similar structure. The learning rate is set to 0.001, and the discount factor is... Soft update parameters .

[0098] After 500 rounds of training, the algorithm stably converges to the optimal strategy.

[0099] In actual operation, the agent can dynamically adjust task migration decisions based on real-time network conditions: When channel conditions are good, the proportion of D2D direct communication should be appropriately increased; When the channel quality is poor, more computing tasks are adaptively offloaded to edge servers for processing.

[0100] Experimental results show that, compared with traditional optimization methods, this method achieves better results with a lower data upload volume. Within the specified range, the system's average total latency decreased by approximately 28%, power consumption decreased by approximately 18%, and the latency constraint satisfaction rate remained above 97%.

[0101] Simulation results yielded a comparison chart of the time delay performance of this method, see below. Figure 5 It clearly demonstrates that, under different data upload volumes, this method can effectively maintain a low latency level, exhibiting good adaptability and system stability.

[0102] This invention constructs a comprehensive reward function that simultaneously considers indicators such as total system latency and transmit power, and explicitly integrates latency constraints, power constraints, and task allocation ratio constraints into a reinforcement learning framework. This invention enables joint optimization of task offloading ratios and transmit power of each link under a unified model. Compared to schemes that only target a single indicator or use simple weighted summation, this invention helps reduce total system latency while considering power consumption, thereby improving the overall service quality in edge computing power grid scenarios. This invention employs a deep reinforcement learning agent that learns task offloading and power allocation strategies through repeated interactions with a near-realistic power grid communication and computing environment. It does not rely on completely accurate analytical mathematical models or complex channel statistical priors. When the operating environment, such as channel conditions and service load, changes, the agent can adaptively adjust its decisions based on the new observation state, making the proposed method more effective than traditional methods in dynamic power grid scenarios. This invention demonstrates better adaptability and robustness for schemes with fixed rules or static optimization. Based on the DDPG algorithm, it constructs an optimization solver in a continuous action space, utilizes deep neural networks to represent the state of high-dimensional systems, and directly performs policy search within the continuous action space through deterministic policy gradients. Compared to methods that discretize continuous control quantities before optimization or rely on partial heuristic search, this invention reduces quantization errors caused by action discretization, lowers the risk of getting trapped in local suboptimal solutions, and more easily obtains task offloading and power allocation schemes that meet the overall system performance requirements. At the algorithm level, it introduces stability mechanisms such as experience replay, soft update of the target network, and exploration noise, combined with preprocessing steps such as data standardization and outlier handling, ensuring stable convergence during training. Furthermore, this method does not rely on precise analytical models, making it easier to deploy and apply in complex real-world power grid environments, and possesses good engineering application value and promising prospects for widespread adoption.

Claims

1. A method for optimizing power grid latency using edge computing based on reinforcement learning, characterized in that, include: S1. Data Acquisition and Preprocessing: Real-time acquisition of power grid communication-related data; cleaning, normalization, and feature extraction of the power grid communication-related data to form a standardized state vector; S2. Multi-objective time delay optimization modeling: Establish data transmission model, latency model, and constraints to transform the joint optimization of task offloading ratio and transmit power into a quantifiable optimization problem; S3. Multi-objective optimization solution based on deep reinforcement learning: We adopt the Deep Deterministic Policy Gradient (DDPG) algorithm framework, which learns the optimal task allocation and power allocation strategy through interaction with the environment. S4. Optimize decision output and execution.

2. The edge computing power grid delay optimization method based on reinforcement learning according to claim 1, characterized in that, In S1, the power grid communication-related data includes channel gain, equipment computing power, and task characteristics; Channel gain includes: channel state information between two power devices. Channel state information between power equipment A and the base station Channel state information between power equipment B and the base station ; The equipment's computing capabilities include: the local computing frequency of the two power devices. The computing frequency of mobile edge computing servers, also known as MEC servers. ; Task characteristics include: the total amount of task data generated by power equipment. The computational complexity of tasks generated by power equipment .

3. The edge computing power grid delay optimization method based on reinforcement learning according to claim 1, characterized in that, In S2, the data transmission model includes the D2D communication transmission rate and the device-to-base station transmission rate, and is modeled based on Shannon's formula. The latency model includes D2D transmission latency, base station transmission latency, edge server processing latency, and local device processing latency. The total latency is the sum of the maximum value of the latency of each path and the processing latency. The constraints include latency constraints, power constraints, and task allocation ratio constraints. The optimization problem described above integrates the objective and constraints into a complete mathematical optimization problem.

4. The edge computing power grid delay optimization method based on reinforcement learning according to claim 3, characterized in that, The D2D communication transmission rate is expressed as: (1); In formula (1): Indicates bandwidth, unit is ; Noise power spectral density, in units of ; The signal transmission power for D2D communication between two electrical devices, in units of ; This refers to the channel status information between two electrical devices; This refers to the D2D communication transmission rate, measured in units of... ; The transmission rate from the device to the base station is expressed as: (2); (3); In formulas (2) and (3): The transmission power of power device A sending data from the task data assigned to the MEC server for processing to the MEC server, in units of ; This refers to the channel state information between power equipment A and the base station; The transmission power (in units of ) for power equipment B to send the portion of the task data allocated to the MEC server for processing to the MEC server. ; This refers to the channel state information between power equipment B and the base station; The uplink transmission rate from power device A to the base station is expressed in units of 1. ; The uplink transmission rate from power equipment B to the base station is expressed in units of 1. .

5. The edge computing power grid delay optimization method based on reinforcement learning according to claim 3, characterized in that, The D2D transmission delay is expressed as follows: (4); In formula (4): The total amount of task data generated by power equipment, in units of ; This represents the proportion of task data transmitted from the power equipment to the peer equipment, with a value range of [value range missing]. ; This refers to the D2D transmission delay, measured in units of... ; The base station transmission delay is expressed as follows: (5); In formula (5): The transmission delay from power equipment A to the base station is expressed in units of 1. ; The transmission delay from power equipment B to the base station is expressed in units of 1. ; This represents the proportion of task data transmitted from power equipment to the MEC server, with a value range of [value missing]. ; The edge server processing latency is expressed as follows: (6); In formula (6): The computing frequency of the MEC server, in units of ; Latency processing for edge servers, in units of ; The computational complexity of tasks generated by power equipment, in units of ; The local device processing latency is expressed as: (7); In formula (7): The calculated frequency for two electrical devices, in units of ; Local device processing latency, in units of ; The total delay is expressed as: (8); In formula (8): Total delay, in units of .

6. The edge computing power grid delay optimization method based on reinforcement learning according to claim 3, characterized in that, The aforementioned time delay constraint means that the task of the power equipment needs to be completed within a specified time, i.e., there is a maximum time delay limit, expressed as: (9); In formula (9), The maximum allowable delay for the task, in units of ; The power constraint means that the transmission power of each power device cannot exceed the maximum power allowed by its hardware. The expression is: (10); The task allocation ratio constraint sets the total data ratio of each power device to 1. It is divided into two parts: one part is transmitted to the peer for comparison and analysis, and the other part is transmitted to the MEC server for comparison and analysis. The expression is: (11); The aforementioned integration of objectives and constraints into a complete mathematical optimization problem, where the objective function aims to minimize the total data transmission and processing delays between power equipment and the MEC server, is expressed as: (12); The constraints are time delay constraint, power constraint, and task allocation ratio constraint, expressed as: (13)。 7. The edge computing power grid delay optimization method based on reinforcement learning according to claim 1, characterized in that, In S3, the DDPG algorithm framework includes state space design, continuous action space design, and reward function design. 1) State-space design: State-space design includes three dimensions: communication channel state, device computing state, and task characteristics, expressed as: (14); In formula (14), This represents the task allocation decision and power allocation made at the previous moment; 2) Continuous motion space design: The continuous action space includes the task allocation ratio and the transmit power of each link, which are continuous variables, and the expression is: (15); In formula (15), the action The decisions are made by the intelligent agent, which is a decision-making module / control policy learning module built based on the DDPG algorithm and deployed on an edge server or control center. 3) Reward function design: The reward function includes a basic reward term and a constraint penalty term. The basic reward is negatively correlated with the total delay, and the penalty term is used to address constraint violations. Considering power efficiency, a negative reward term for power consumption is introduced, expressed as: (16); In formula (16): ; ; These are weighting coefficients used to balance the importance of time delay and power, satisfying... ; Total power consumption, in units of ; The constraint penalty term is expressed as follows: (17); In formula (17): ; The penalty coefficient is much larger than the typical value of the base reward; The value is 1 if the indicator function condition is true, and 0 otherwise. The value of the constraint penalty item; 4) The final reward is the sum of the basic reward and the constraint penalty, expressed as: (18); In formula (18): This is the final total reward value; Basic reward value; The value of the constraint penalty item.

8. The edge computing power grid delay optimization method based on reinforcement learning according to claim 1, characterized in that, In S3, the training process of the Deep Deterministic Policy Gradient (DDPG) algorithm is as follows: An experience replay mechanism is used to store and sample training data; Exploration was conducted using Ornstein-Uhlenbeck process noise; Update the target network parameters using a soft update mechanism; During training, the Critic network is updated using mean squared error loss, and the Actor network is updated using policy gradient.

9. The edge computing power grid delay optimization method based on reinforcement learning according to claim 1, characterized in that, In S4, the optimization decision output and execution include: The action vectors output by the Actor network are parsed into specific task allocation ratios and power control instructions. The optimization effect is evaluated through key performance indicators, including latency optimization rate, power efficiency improvement, and constraint satisfaction rate.