A Virtual Power Plant Terminal Resource Scheduling Method and System Based on Mobile Edge Cloud Computing
By constructing a resource scheduling method based on mobile edge cloud and deep reinforcement learning in a virtual power plant, optimizing hotspot selection and task offloading, the problem of low resource allocation efficiency in traditional edge computing in virtual power plants is solved, and the benefits of virtual power plants are maximized.
Patent Information
- Application Number
- CN202311601125.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-28
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2043-11-28
AI Technical Summary
Traditional edge computing struggles to flexibly provide services based on different task characteristics in virtual power plants, especially for latency-sensitive and data-intensive tasks. Furthermore, the low efficiency of network resource allocation in mobile edge clouds impacts the performance of virtual power plants.
By using a virtual power plant terminal resource scheduling method based on mobile edge cloud computing, target service hotspots and task offloading strategies are determined, a benefit optimization problem is constructed, and deep reinforcement learning is used to find the optimal solution, thereby optimizing hotspot selection and task offloading to maximize the benefits of the virtual power plant.
This approach enables the rational allocation of network resources based on task priority, thereby improving the overall efficiency of the virtual power plant. It also optimizes hotspot selection and task offloading strategies, enhancing the performance of the virtual power plant.
Smart Images

Figure CN117634796B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of edge computing technology, and in particular to a method and system for scheduling virtual power plant terminal resources based on mobile edge cloud computing. Background Technology
[0002] The construction of a virtual power plant is a process of digitizing, intelligentizing, and internetizing the traditional power grid. Through advanced information and communication technologies and software systems, it achieves the aggregation and coordinated optimization of distributed energy sources such as distributed power sources, energy storage systems, controllable loads, and electric vehicles, thus functioning as a power coordination and management system that participates in the electricity market and grid operation as a special type of power plant.
[0003] With the rapid development of information technology and the Internet of Vehicles, the scenarios in which electric vehicles and other mobile energy sources participate in the regulation of virtual power plants have increased significantly. These applications mostly generate latency-sensitive and data-intensive tasks, which increases the burden on traditional edge computing. Furthermore, the characteristics of tasks within the network vary, and traditional static edge clouds, due to their limited coverage and lack of flexibility, find it difficult to flexibly provide services to virtual power plant terminals based on different task characteristics.
[0004] To address this challenge, mobile edge cloud has been introduced into the network. By equipping edge servers on mobile devices such as drones, vehicles, and robots, it can overcome coverage bottlenecks and flexibly utilize communication and computing resources to provide mobile computing services to terminals. However, in related research on mobile edge computing, how should mobile edge cloud rationally allocate network resources to maximize the performance of virtual power plants? Summary of the Invention
[0005] This invention provides a method and system for scheduling virtual power plant terminal resources based on mobile edge cloud computing, which solves the technical problem of how mobile edge cloud should reasonably allocate network resources in order to maximize the performance of virtual power plants.
[0006] In view of this, the first aspect of the present invention provides a method for scheduling virtual power plant terminal resources based on mobile edge cloud computing, comprising the following steps:
[0007] In response to task computation requests from terminals in various hotspot areas of the virtual power plant, the terminal of the target service hotspot and the corresponding task offloading strategy of the terminal of the target service hotspot are determined according to task priority.
[0008] The execution latency of the computing tasks of each terminal in the target service hotspot is determined based on the task offloading strategy of the terminal in the target service hotspot.
[0009] The total benefit of the virtual power plant is determined by the benefit function of the task whose priority is determined based on the execution delay of each task;
[0010] With the goal of maximizing the total benefits of the virtual power plant, and with the hotspot selection strategy and task offloading strategy of the virtual power plant as decision variables, we construct the benefit optimization problem of the virtual power plant.
[0011] The optimization problem of the virtual power plant's efficiency is solved by finding the optimal solution, and the corresponding hotspot selection strategy and task offloading strategy for the virtual power plant are determined based on the optimal solution.
[0012] Preferably, the step of determining the execution latency of the computing task of each terminal in the target service hotspot according to the task offloading strategy of the terminal in the target service hotspot specifically includes:
[0013] The unloading latency of the UAV is determined by the task space size and data transmission rate of the computing task of the terminal of the target service hotspot.
[0014] The computational latency of the drone when the computational tasks of the terminal at the target service hotspot are offloaded to the drone for computation is determined by the task space size of the computational tasks of the terminal at the target service hotspot, the CPU computation cycle, and the computational capability allocated by the drone to the tasks offloaded to the drone.
[0015] The local computation latency of a terminal choosing to offload its tasks to local computation is determined by the task space size, CPU computation cycle and local computing capability of the terminal's computing tasks at the target service hotspot.
[0016] The execution latency of each computing task is determined based on the unloading latency of the drone at the target service hotspot terminal, the drone's computing latency, and the local computing latency.
[0017] Preferably, the step of determining the total benefit of the virtual power plant by determining the benefit function of the task prioritizing each task based on the execution delay of each task specifically includes:
[0018] Based on the execution latency of each task, the benefit functions corresponding to high-priority, medium-priority, and low-priority tasks are determined as follows:
[0019]
[0020]
[0021]
[0022] In the formula, τ represents the benefit function values for high-priority, medium-priority, and low-priority tasks, respectively. m,l Let I(x) be the tolerance latency threshold for terminal m within hotspot l, and let x be the indicator function. If x is true, then I(x) = 1; if x is false, then I(x) = 0. H As a negative penalty factor, v m,l For task priority, v m,l When v = 1, it is a high-priority task. m,l =0 indicates a low-priority task, v m,l When ∈(0,1), it is a medium priority task, Γ L For positive rewards, α is the exponential decay factor;
[0023] The total benefit of the virtual power plant is determined based on the benefit functions corresponding to high-priority, medium-priority, and low-priority tasks:
[0024]
[0025] In the formula, U sys M represents the total benefit value, L represents the number of terminals, and L represents the number of hotspots.
[0026] Preferably, the benefit optimization problem includes an objective function and constraints, wherein the objective function and the constraints are respectively:
[0027]
[0028] In the formula, P represents the objective function, st represents the constraint, C1 represents the task unloading constraint, C2 represents the hotspot selection constraint, and C3 represents the constraint that the UAV can only select a service hotspot.
[0029] Preferably, the method further includes:
[0030] The optimization problem of the virtual power plant's efficiency is solved by using deep reinforcement learning. Based on the optimal solution, the corresponding hotspot selection strategy and task offloading strategy of the virtual power plant are determined.
[0031] Preferably, the step of using deep reinforcement learning to find the optimal solution to the virtual power plant's efficiency problem, and determining the corresponding hotspot selection strategy and task offloading strategy for the virtual power plant based on the optimal solution, specifically includes:
[0032] The initial experience pool capacity is set to E;
[0033] Randomly initialize the weight parameters θ of the evaluation network, and initialize the weight parameters θ' = θ of the target network;
[0034] The efficiency optimization problem is constructed as a Markov process. According to the Markov process definition, the state space S(t) and action space A(t) are defined, where the state space S(t) is defined as:
[0035] S(t)={s(t)|s(t)={p l (t),b m,l (t)}}
[0036] In the formula, s(t) represents the state at time t, and p l (t) represents the coordinates of the hotspot selected by the drone at time t, p l (t)={x l (t),y l (t)},x l (t),y l 9t) represent the x and y coordinates of the hotspot area selected by the drone at time t, respectively, and b m,l (t) represents the task offloading strategy of terminal m within hotspot l at time t;
[0037] The action space A(t) is defined as:
[0038]
[0039] In the formula, a(t) represents the action at time t. This indicates the action of selecting a hotspot. This indicates the action of unloading the task;
[0040] Initialize the state space, obtain all actions through the evaluation network based on the current input state, select an action to execute using a greedy algorithm, and record the state and reward value of the action at the next moment after execution.
[0041] The state at the next moment is used as the current state. Return to the initial state space. Based on the current input state, obtain all actions through the evaluation network. Use a greedy algorithm to select the steps of an action. Store the current state, the state at the next moment, the selected action, and the reward value for executing the action into the experience pool E until the experience pool E is full. Then execute the next step.
[0042] Randomly sample a small batch of the current state, the next state, the selected action, and the executed action from the experience pool E, and iteratively update the weight parameters θ of the evaluation network by calculating the loss function;
[0043] After iterative updates to the preset maximum number of iterations, the iteration stops, and the corresponding action space is output to determine the hotspot selection strategy and task unloading strategy for the corresponding virtual power plant.
[0044] Secondly, the present invention also provides a virtual power plant terminal resource scheduling system based on mobile edge cloud computing, comprising:
[0045] The task response module is used to respond to the task calculation requests of terminals in each hotspot area of the virtual power plant, and to determine the terminal of the target service hotspot and the task unloading strategy corresponding to the terminal of the target service hotspot according to the task priority.
[0046] The latency calculation module is used to determine the execution latency of the computing tasks of the terminal in the target service hotspot under each task priority according to the task offloading strategy of the terminal in the target service hotspot.
[0047] The benefit calculation module is used to determine the benefit function of each task based on the execution delay of each task and the priority of each task, so as to determine the total benefit of the virtual power plant.
[0048] The benefit optimization module is used to construct the benefit optimization problem of the virtual power plant with the goal of maximizing the total benefit of the virtual power plant and the hot spot selection strategy and task unloading strategy of the virtual power plant as decision variables.
[0049] The problem-solving module is used to find the optimal solution to the efficiency optimization problem of the virtual power plant, and determine the corresponding hot spot selection strategy and task unloading strategy of the virtual power plant based on the optimal solution.
[0050] Thirdly, the present invention also provides an electronic device, the electronic device including a memory and a processor;
[0051] The memory is used to store programs;
[0052] The processor executes the program to implement the above-described method.
[0053] Fourthly, the present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0054] Fifthly, the present invention also provides a computer program product comprising at least one computer-readable storage medium having computer-executable program code instructions stored therein, the computer-executable program code comprising:
[0055] The first program code section is configured to respond to task calculation requests from terminals in each hotspot area of the virtual power plant, and determine the terminal of the target service hotspot and the task offloading strategy corresponding to the terminal of the target service hotspot according to the task priority.
[0056] The second program code section is configured to determine the execution latency of the computing tasks of each terminal in the target service hotspot based on the task offloading strategy of the terminal in the target service hotspot.
[0057] The third part of the program code is configured to determine the total benefit of the virtual power plant by determining the benefit function of the task priority based on the execution delay of each task.
[0058] The fourth part of the program code is configured to construct the efficiency optimization problem of the virtual power plant with the goal of maximizing the total benefits of the virtual power plant and the hot spot selection strategy and task unloading strategy of the virtual power plant as decision variables.
[0059] The fifth program code section is configured to find the optimal solution to the efficiency optimization problem of the virtual power plant, and determine the corresponding hotspot selection strategy and task unloading strategy of the virtual power plant based on the optimal solution.
[0060] As can be seen from the above technical solutions, the present invention has the following advantages:
[0061] This invention determines the terminals of the target service hotspot and their corresponding task offloading strategies based on task priorities. It then determines the execution latency of the computing tasks for each terminal in the target service hotspot based on the task offloading strategies. Finally, it determines the benefit function of each task priority based on the execution latency of each computing task to determine the total benefit of the virtual power plant. With maximizing the total benefit of the virtual power plant as the objective condition, and using the hotspot selection strategy and task offloading strategy of the virtual power plant as decision variables, it constructs and solves the benefit optimization problem of the virtual power plant to determine the optimal hotspot selection strategy and task offloading strategy. This allows for network resource allocation of tasks with different priorities through mobile edge cloud, thereby maximizing the performance of the virtual power plant. Attached Figure Description
[0062] Figure 1 A flowchart illustrating a virtual power plant terminal resource scheduling method based on mobile edge cloud computing, provided as an embodiment of the present invention;
[0063] Figure 2 This is a schematic diagram of the structure of a drone-assisted edge computing-based virtual power plant provided in an embodiment of the present invention;
[0064] Figure 3 This is a schematic diagram illustrating the simulation results of the virtual power plant's efficiency under varying training rounds, as provided in an embodiment of the present invention.
[0065] Figure 4 A schematic diagram illustrating the simulation results of the virtual power plant benefits under varying drone computing power, provided by an embodiment of the present invention.
[0066] Figure 5A schematic diagram illustrating the simulation results of the virtual power plant benefits under varying wireless bandwidth, provided by an embodiment of the present invention.
[0067] Figure 6 This is a schematic diagram of the structure of a virtual power plant terminal resource scheduling system based on mobile edge cloud computing, provided as an embodiment of the present invention. Detailed Implementation
[0068] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0069] For easier understanding, please refer to Figure 1 The present invention provides a method for scheduling virtual power plant terminal resources based on mobile edge cloud computing, comprising the following steps:
[0070] S1. In response to the task calculation requests from terminals in each hotspot area of the virtual power plant, determine the terminal of the target service hotspot and the corresponding task offloading strategy based on the task priority.
[0071] It should be noted that within each time slot, the drone establishes a connection with terminals in each hotspot area of the virtual power plant via a wireless channel. Terminals in each hotspot area will issue task computing requests, and the tasks have different priorities. The drone collects the task request information from the terminals. The drone is used for mobile edge cloud computing.
[0072] In one example, priorities include high, medium, and low. Generally, the higher the priority of a task issued by a terminal in a hotspot, the more likely it is to be selected to serve the corresponding terminal in the target service hotspot. The task offloading strategy refers to where the task of the terminal in the target service hotspot is offloaded for computation. Task offloading strategies include two types: offloading to a drone for computation via a wireless channel and offloading to the terminal's local machine for computation.
[0073] S2. Determine the execution latency of the computing tasks of each terminal in the target service hotspot based on the task offloading strategy of the terminal in the target service hotspot.
[0074] It should be noted that since the task offloading strategy of the terminal at the target service hotspot includes two offloading strategies, if the computation is offloaded to the drone via the wireless channel, it involves the terminal's wireless bandwidth, wireless channel parameters, and data transmission rate; if the computation is offloaded to the terminal's local machine, it involves local computing capabilities, etc. Therefore, it is necessary to determine the execution latency of each task based on the task offloading strategy of the terminal at the target service hotspot.
[0075] S3. Determine the total benefit of the virtual power plant by determining the benefit function of the task priority based on the execution delay of each task.
[0076] Specifically, the benefit expressions for tasks with different priorities are different. High-priority tasks generally occur in real-time control operations, such as frequency modulation control. They typically have strict execution delays; if a high-priority task is not completed within the specified time, it will fail and its benefits will be significantly reduced. Low-priority tasks generally occur in invitation-based request responses with more lenient delay requirements. Their computational delay requirements are relatively tolerant. When the execution time of a low-priority task exceeds the given delay, the computational result is still usable, but the task's benefits will decrease as the execution delay increases. Medium-priority tasks exist in some communication scenarios, such as security in many public places. Medium-priority tasks fall between high-priority and low-priority tasks, but their urgency is not as high as that of high-priority tasks.
[0077] It should be noted that the benefits of the virtual power plant are related to task priority. When a high-priority task is completed within the specified time delay, its benefits increase as the completion delay decreases. However, if a task is not completed within the specified time, its benefits will be significantly reduced. When a low-priority task is completed within the specified time, it will receive a fixed positive benefit value. If a task is not completed within the specified time, the calculation result is still usable, but the task benefit will decrease as the execution delay increases.
[0078] S4. With the goal of maximizing the total benefits of the virtual power plant, and with the hotspot selection strategy and task offloading strategy of the virtual power plant as decision variables, construct the benefit optimization problem of the virtual power plant.
[0079] It should be noted that the research found that when the computing power and terminal wireless bandwidth allocated by the drone to the tasks offloaded to the drone are large, their impact on the utility of the virtual power plant can be ignored. Therefore, the effectiveness of the virtual power plant depends on the hotspot selection and task offloading strategy.
[0080] S5. Solve the optimization problem of the virtual power plant's efficiency, and determine the corresponding hot spot selection strategy and task unloading strategy for the virtual power plant based on the optimal solution.
[0081] It should be noted that this invention determines the terminals of the target service hotspot and the corresponding task offloading strategy based on task priority. It then determines the execution latency of each task based on the task offloading strategy of the target service hotspot terminals, and determines the benefit function of each task priority based on the execution latency of the computational tasks of each target service hotspot terminal. This determines the total benefit of the virtual power plant. With maximizing the total benefit of the virtual power plant as the objective condition, and using the hotspot selection strategy and task offloading strategy of the virtual power plant as decision variables, it constructs a benefit optimization problem for the virtual power plant and performs problem optimization to determine the optimal hotspot selection strategy and task offloading strategy. This allows for network resource allocation of tasks with different priorities through mobile edge cloud, thereby maximizing the performance of the virtual power plant.
[0082] In one specific embodiment, step S2 specifically includes:
[0083] 201. Determine the unloading latency of the UAV when the computing tasks of the terminal at the target service hotspot are offloaded to the UAV for computing based on the task space size and data transmission rate of the computing tasks of the terminal at the target service hotspot.
[0084] Specifically, the formula for calculating the unloading delay of a drone is:
[0085]
[0086] In the formula, l represents the hotspot index, and m represents the terminal index. Indicates the unloading delay, s m,l Indicates the task space size, r m,l Represents the data transmission rate, where the data transmission rate r m,l The calculation is as follows:
[0087]
[0088] In the formula, B represents the terminal's wireless bandwidth, and P... m,l h represents the transmission power of terminal m within hotspot l. m,l This represents the wireless channel parameters σ of the terminal m within the hotspot l connecting to the drone. 2 Let V be the variance of the Gaussian white noise.
[0089] 202. Determine the computation latency of the UAV when the computation tasks of the terminal at the target service hotspot are offloaded to the UAV for computation by considering the task space size of the computation tasks of the terminal at the target service hotspot, the CPU computation cycle, and the computational capability allocated by the UAV to the tasks offloaded to the UAV.
[0090] Specifically, the formula for calculating the computational latency of a drone is:
[0091]
[0092] In the formula, The computational latency of the drone is represented by ζ, where ζ represents the CPU computation cycle, and f is the CPU computation time. uav This indicates the computing power allocated by the drone to the tasks unloaded onto it.
[0093] 203. Determine the local computation latency when the terminal chooses to offload its tasks to local computation based on the task space size, CPU computation cycle, and local computing capability of the terminal's computation task at the target service hotspot.
[0094] The local computation latency is:
[0095]
[0096] In the formula, f represents the local computation latency. m,l This represents the local computing power of terminal m within hotspot l;
[0097] It should be noted that when the task is unloaded to the local machine for computation, there is no need to consider the task unloading latency.
[0098] 204. Determine the execution latency of each computing task based on the unloading latency of the drone at the target service hotspot terminal, the computing latency of the drone, and the local computing latency.
[0099] Specifically, the formula for calculating the execution latency of each computational task is as follows:
[0100]
[0101] In the formula, T m,l c represents the execution latency of terminal m within hotspot l. l This indicates the hotspot selection strategy, c l =1 indicates that hotspot l is selected, c l =0 indicates that hotspot l was not selected, b m,l Indicates the task unloading policy, b m,l =1 indicates that task m within hotspot l is offloaded to the drone for computation, b m,l =0 indicates that task m within hotspot l is unloaded to the local machine for computation.
[0102] In one specific embodiment, step S3 specifically includes:
[0103] 301. Based on the execution latency of each task, determine the benefit functions corresponding to high-priority, medium-priority, and low-priority tasks, respectively:
[0104]
[0105]
[0106]
[0107] In the formula, τ represents the benefit function values for high-priority, medium-priority, and low-priority tasks, respectively. m,l Let I(x) be the tolerance latency threshold for terminal m within hotspot l, and let x be the indicator function. If x is true, then I(x) = 1; if x is false, then I(x) = 0. H As a negative penalty factor, v m,l For task priority, v m,l When v = 1, it is a high-priority task. m,l =0 indicates a low-priority task, v m,l When ∈(0,1), it is a medium priority task, Γ L For positive rewards, α is the exponential decay factor;
[0108] It should be noted that after obtaining the benefit functions corresponding to high-priority and low-priority tasks, the benefit function of medium-priority tasks can be obtained through linear interpolation.
[0109] 302. Based on the benefit functions corresponding to high-priority, medium-priority, and low-priority tasks, the total benefit of the virtual power plant is determined as follows:
[0110]
[0111] In the formula, U sys M represents the total benefit value, L represents the number of terminals, and L represents the number of hotspots.
[0112] It should be noted that, based on the expression for the total benefit of the virtual power plant, many factors influence its effectiveness, such as wireless bandwidth, the computing power of the drone, and hotspot selection and task offloading strategies. The following analysis examines the impact of wireless bandwidth and drone computing power, and then optimizes the virtual power plant performance by improving hotspot selection and task offloading strategies.
[0113] Expanding on the formula for calculating execution latency, the execution latency T of each task can be calculated. m,l Represented as:
[0114]
[0115] From the above formula, we can see that the task execution time varies with the computing power f of the task assigned to the drone. uav It decreases as the terminal's wireless bandwidth B increases. When fuav When the value is large enough, the task execution latency T m,l It can be approximated as:
[0116]
[0117] In the formula,
[0118] Based on the above formula, we can conclude that when f uav When the size is large enough, the computational latency of the task on the drone is negligible, and the offloading latency dominates the task processing latency, resulting in f uav The impact on the profitability of virtual power plants becomes minimal.
[0119] Similarly, when the terminal's wireless bandwidth B is sufficiently large, the task execution latency T m,l It can be approximated as:
[0120]
[0121] Based on the above formula, we can also conclude that when the terminal wireless bandwidth B is large enough, the task offloading delay is very small. At this time, the task processing time mainly depends on the computation delay. Therefore, the impact of the terminal wireless bandwidth B on the virtual power plant's efficiency is negligible.
[0122] In summary, when the computing power and terminal wireless bandwidth allocated to the tasks offloaded to the drone by the drone are large, their impact on the utility of the virtual power plant can be ignored. Therefore, the effectiveness of the virtual power plant depends on the hotspot selection and task offloading strategy.
[0123] In one specific embodiment, the benefit optimization problem includes an objective function and constraints, wherein the objective function and constraints are as follows:
[0124]
[0125] In the formula, P represents the objective function, st represents the constraint, C1 represents the task unloading constraint, C2 represents the hotspot selection constraint, and C3 represents the constraint that the UAV can only select a service hotspot.
[0126] In one feasible approach, deep reinforcement learning is used to find the optimal solution to the efficiency optimization problem of a virtual power plant, and the corresponding hotspot selection strategy and task offloading strategy of the virtual power plant are determined based on the optimal solution.
[0127] Specifically, the steps of using deep reinforcement learning to find the optimal solution to the virtual power plant's efficiency optimization problem, and determining the corresponding hotspot selection strategy and task offloading strategy based on the optimal solution, include:
[0128] 501. Initialize the experience pool capacity to E;
[0129] 502. Randomly initialize the weight parameters θ of the evaluation network, and initialize the weight parameters θ' = θ of the target network;
[0130] 503. Construct the efficiency optimization problem as a Markov process. According to the Markov process definition, the state space S(t) and action space A(t) are defined, where the state space S(t) is defined as:
[0131] S(t)={s(t)|s(t)={p l (t),b m,l (t)}}
[0132] In the formula, s(t) represents the state at time t, and p l (t) represents the coordinates of the hotspot selected by the drone at time t, p l (t)={x l (t),y l (t)},x l (t),y l (t) represents the x and y coordinates of the hotspot area selected by the drone at time t, respectively. m,l (t) represents the task offloading strategy of terminal m within hotspot l at time t;
[0133] The action space A(t) is defined as:
[0134]
[0135] In the formula, a(t) represents the action at time t. This indicates the action of selecting a hotspot. This indicates the action of unloading the task;
[0136] 504. Initialize the state space, obtain all actions through the evaluation network based on the current input state, select an action to execute using a greedy algorithm, and determine the state and reward value of the action at the next moment after execution.
[0137] The process of selecting an action using a greedy algorithm is represented as follows:
[0138]
[0139] In the formula, a(t) represents the selected action, ε represents the probability, ε∈[0,1], and Q(S(t),a;θ) represents the Q evaluation network;
[0140] It should be noted that after executing the selected action a(t), the current state S(t) will transition to the next state S(t+1), and the reward value for executing action a(t) will be obtained:
[0141]
[0142] In the formula, δ1, δ2, and δ3 are all positive constants, U sys (t) and U sys (t+1) represents the total benefits of the virtual power plant at time t and time t+1, respectively.
[0143] 505. Take the state of the next moment as the current state, return to the initial state space, obtain all actions through the evaluation network based on the current input state, select the steps of an action using a greedy algorithm, and store the current state, the state of the next moment, the selected action and the reward value of the action into the experience pool E until the experience pool E is full, and then proceed to the next step.
[0144] 506. Randomly sample a small batch of the current state, the next state, the selected action, and the executed action from the experience pool E, and iteratively update the weight parameters θ of the evaluation network by calculating the loss function.
[0145] The loss function L(t) is expressed as:
[0146] L(t)=(Z(t)-Q(S(t),a(t);θ)) 2
[0147] In the formula, Z(t) represents a function of the target network. η is the discount parameter;
[0148] 507. After iterative updates to the preset maximum number of iterations, the iteration stops, and the corresponding action space is output to determine the hotspot selection strategy and task unloading strategy of the corresponding virtual power plant.
[0149] The following examples, based on the virtual power plant terminal resource scheduling method based on mobile edge cloud computing provided in this embodiment, illustrate some computational examples to verify the effectiveness of the hotspot selection strategy and task offloading strategy for the virtual power plant determined by this invention. In the following examples, a virtual power plant based on edge computing assisted by a drone is applied, such as... Figure 2 As shown, Figure 2 The diagram illustrates the structure of a virtual power plant based on edge computing.
[0150] Example 1
[0151] In a Python simulation environment, a computer was used to simulate the change in the virtual power plant benefits of the proposed method with the number of training rounds of the deep reinforcement learning algorithm. The simulation results are as follows: Figure 3 As shown. The parameters in the simulation experiment are set as follows: P m,l =2W and σ 2 =1×10 -2 W, ζ = 5 cycles / bit, f m,l =1×10 7 cycle / s, f uav =6×10 8 The network has a cycle / s, B = 30MHz, L = 6, and M = 5. Furthermore, the task sizes in the network follow a uniform distribution s. m,l ~U(5l,5(l+1)), the priority follows a uniform distribution v m,l ~U(l / L, (l+1) / L).
[0152] By comparing the proposed method with three other methods—random hotspot selection, random task offloading, and both random hotspot selection and task offloading—it is found that the proposed method for scheduling virtual power plant terminal resources based on mobile edge cloud computing can learn effective hotspot selection and task offloading strategies. Furthermore, the virtual power plant efficiency is superior to the other three methods, thus verifying the effectiveness of the proposed method.
[0153] Example 2
[0154] In a Python simulation environment, a computer was used to simulate the virtual power plant benefits of the proposed method as a function of the UAV's computing power. The simulation results are as follows: Figure 4 As shown. The parameters in the simulation experiment are set as follows: P m,l =2W and σ 2 =1×10 -2 W, ζ = 5 cycles / bit, f m,l =1×10 7 The network has a cycle / s, B = 30MHz, L = 6, and M = 5. Furthermore, the task sizes in the network follow a uniform distribution s. m,l ~U(5l,5(l+1)), the priority follows a uniform distribution v m,l ~U(l / L, (l+1) / L). Through Figure 4 We found that the benefits of the virtual power plant change with the computing power of the drone, which is consistent with the conclusions analyzed in this invention. That is, the benefits of the virtual power plant depend on the hotspot selection and task offloading strategies. Furthermore, by comparing it with three methods, namely random hotspot selection, random task offloading, and random hotspot selection and task offloading, the terminal resource scheduling method of the virtual power plant based on mobile edge cloud computing proposed in this invention is superior to the other three methods, thus verifying the effectiveness of this method.
[0155] Example 3
[0156] In a Python simulation environment, a computer was used to simulate the virtual power plant benefits of the proposed method as a function of wireless bandwidth. The simulation results are as follows: Figure 5 As shown. The parameters in the simulation experiment are set as follows: P m,l =2W and σ 2 =1×10 - 2 W, ζ = 5 cycles / bit, f m,l =1×10 7 cycle / s, f uav =6×10 8 The network has a cycle / s, L=6, and M=5. Furthermore, the task sizes in the network follow a uniform distribution s. m,l ~U(5l,5(l+1)), the priority follows a uniform distribution v m,l ~U(l / L, (l+1) / L). Through Figure 5 We found that the benefits of the virtual power plant vary with wireless bandwidth in line with the conclusions analyzed in this invention, namely, the benefits of the virtual power plant depend on hotspot selection and task offloading strategies. Furthermore, by comparing it with three methods—random hotspot selection, random task offloading, and both random hotspot selection and task offloading—the terminal resource scheduling method for the virtual power plant based on mobile edge cloud computing proposed in this invention is superior to the other three methods, thus verifying the effectiveness of this method.
[0157] The above is a detailed description of an embodiment of a virtual power plant terminal resource scheduling method based on mobile edge cloud computing provided by the present invention. The following is a detailed description of an embodiment of a virtual power plant terminal resource scheduling system based on mobile edge cloud computing provided by the present invention.
[0158] For easier understanding, please refer to Figure 6 The present invention also provides a virtual power plant terminal resource scheduling system based on mobile edge cloud computing, comprising:
[0159] The task response module 100 is used to respond to the task calculation requests of terminals in each hotspot area of the virtual power plant, and to determine the terminal of the target service hotspot and the task unloading strategy corresponding to the terminal of the target service hotspot according to the task priority.
[0160] The latency calculation module 200 is used to determine the execution latency of the computing tasks of the terminals of the target service hotspot under each task priority according to the task offloading strategy of the terminals of the target service hotspot.
[0161] The benefit calculation module 300 is used to determine the benefit function of each task based on the execution delay of each task and the priority of each task to determine the total benefit of the virtual power plant.
[0162] The benefit optimization module 400 is used to construct the benefit optimization problem of the virtual power plant with the goal of maximizing the total benefit of the virtual power plant and the hot spot selection strategy and task unloading strategy of the virtual power plant as decision variables.
[0163] The problem-solving module 500 is used to find the optimal solution to the efficiency optimization problem of the virtual power plant, and to determine the corresponding hot spot selection strategy and task unloading strategy of the virtual power plant based on the optimal solution.
[0164] The present invention also provides an electronic device, which includes a memory and a processor;
[0165] Memory is used to store programs;
[0166] The processor executes the program to implement the above method.
[0167] The present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0168] The present invention also provides a computer program product, comprising at least one computer-readable storage medium having computer-executable program code instructions stored therein, the computer-executable program code including:
[0169] The first program code section is configured to respond to task calculation requests from terminals in each hotspot area of the virtual power plant, and determine the terminal of the target service hotspot and the corresponding task offloading strategy based on task priority.
[0170] The second part of the program code is configured to determine the execution latency of the computing tasks of each terminal in the target service hotspot based on the task offloading strategy of the terminal in the target service hotspot.
[0171] The third part of the program code is configured to determine the total benefit of the virtual power plant by determining the benefit function of the task priority based on the execution delay of each task.
[0172] The fourth part of the program code is configured to construct the efficiency optimization problem of the virtual power plant with the goal of maximizing the total benefits of the virtual power plant and the hot spot selection strategy and task unloading strategy of the virtual power plant as decision variables.
[0173] The fifth part of the program code is configured to find the optimal solution to the efficiency optimization problem of the virtual power plant, and determine the corresponding hot spot selection strategy and task unloading strategy of the virtual power plant based on the optimal solution.
[0174] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, electronic devices, computer-readable storage media, and computer program products described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0175] In the embodiments provided by this invention, it should be understood that the disclosed systems, electronic devices, computer-readable storage media, computer program products, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.
[0176] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0177] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0178] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for executing all or part of the steps of the methods of the various embodiments of this invention through a computer device (which may be a personal computer, server, or network device, etc.). The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing program code.
[0179] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for scheduling virtual power plant terminal resources based on mobile edge cloud computing, characterized in that, Includes the following steps: In response to task computation requests from terminals in various hotspot areas of the virtual power plant, the terminal of the target service hotspot and the corresponding task offloading strategy of the terminal of the target service hotspot are determined according to task priority. The execution latency of the computing tasks of each terminal in the target service hotspot is determined based on the task offloading strategy of the terminal in the target service hotspot. The total benefit of the virtual power plant is determined by the benefit function of tasks whose priorities are determined based on the execution latency of each task, including: Based on the execution latency of each task, the benefit functions corresponding to high-priority, medium-priority, and low-priority tasks are determined as follows: ; ; ; In the formula, , , These represent the benefit function values for high-priority, medium-priority, and low-priority tasks, respectively. Hot topic The tolerance latency threshold for internal terminal m, Indicates hot topics The execution latency of terminal m within the terminal. (x) is an indicator function, where x is the indicator parameter. If x is true, then (x) = 1, if x is false, then (x) = 0, As a negative penalty factor, As a task priority, When the value is 1, it indicates a high-priority task. A value of 0 indicates a low-priority task. Tasks with a priority of ∈ (0,1) are considered medium priority tasks. For positive rewards, α is the exponential decay factor; The total benefit of the virtual power plant is determined based on the benefit functions corresponding to high-priority, medium-priority, and low-priority tasks: ; In the formula, M represents the total benefit value, L represents the number of terminals, and L represents the number of hotspots. ; With maximizing the total benefit of the virtual power plant as the objective condition, and using the hotspot selection strategy and task offloading strategy of the virtual power plant as decision variables, a benefit optimization problem for the virtual power plant is constructed. The benefit optimization problem includes an objective function and constraints, wherein the objective function and the constraints are as follows: ; ; ; ; In the formula, P represents the objective function. Indicates constraints. This indicates the task unloading constraints. This indicates the constraints for hotspot selection. This indicates that the only constraint for drones to select service hotspots is... This indicates the hot topic selection strategy. Indicates hot topics Being chosen, Indicates hot topics Not selected Indicates the task unloading policy. Indicates hot topics Task m within the system is offloaded to the drone for computation. Indicates hot topics Task m within the system is unloaded and computed locally. The optimization problem of the virtual power plant's efficiency is solved by finding the optimal solution, and the corresponding hotspot selection strategy and task offloading strategy for the virtual power plant are determined based on the optimal solution.
2. The virtual power plant terminal resource scheduling method based on mobile edge cloud computing according to claim 1, characterized in that, The step of determining the execution latency of the computing task of each terminal in the target service hotspot according to the task offloading strategy of the terminal in the target service hotspot specifically includes: The unloading latency of the UAV is determined by the task space size and data transmission rate of the computing task of the terminal of the target service hotspot. The computational latency of the drone when the computational tasks of the terminal at the target service hotspot are offloaded to the drone for computation is determined by the task space size of the computational tasks of the terminal at the target service hotspot, the CPU computation cycle, and the computational capability allocated by the drone to the tasks offloaded to the drone. The local computation latency of a terminal choosing to offload its tasks to local computation is determined by the task space size, CPU computation cycle and local computing capability of the terminal's computing tasks at the target service hotspot. The execution latency of each computing task is determined based on the unloading latency of the drone at the target service hotspot terminal, the drone's computing latency, and the local computing latency.
3. The virtual power plant terminal resource scheduling method based on mobile edge cloud computing according to claim 1, characterized in that, Also includes: The optimization problem of the virtual power plant's efficiency is solved by using deep reinforcement learning. Based on the optimal solution, the corresponding hotspot selection strategy and task offloading strategy of the virtual power plant are determined.
4. The virtual power plant terminal resource scheduling method based on mobile edge cloud computing according to claim 3, characterized in that, The steps of using deep reinforcement learning to find the optimal solution to the virtual power plant's efficiency problem, and determining the corresponding hotspot selection strategy and task offloading strategy based on the optimal solution, specifically include: The initial experience pool capacity is set to E; Randomly initialize the weight parameters of the evaluation network Initialize the weight parameters of the target network. ; The efficiency optimization problem is constructed as a Markov process. According to the Markov process definition, the state space S(t) and action space A(t) are defined, where the state space S(t) is defined as: ; In the formula, Indicates the state at time t. This indicates the coordinates of the hotspot the drone selected for service at time t. , These represent the x and y coordinates of the hotspot area selected by the drone at time t. This indicates the task offloading policy of terminal m within hotspot l at time t; The action space A(t) is defined as: ; In the formula, Indicates the action at time t. This indicates the action of selecting a hotspot. , This indicates the action of unloading the task; Initialize the state space, obtain all actions through the evaluation network based on the current input state, select an action to execute using a greedy algorithm, and record the state and reward value of the action at the next moment after execution. The state at the next moment is used as the current state. Return to the initial state space. Based on the current input state, obtain all actions through the evaluation network. Use a greedy algorithm to select the steps of an action. Store the current state, the state at the next moment, the selected action, and the reward value for executing the action into the experience pool E until the experience pool E is full. Then execute the next step. Randomly sample a small batch of the current state, the next state, the selected action, and the executed action from the experience pool E, and iteratively update the weight parameters of the evaluation network by calculating the loss function. ; After iterative updates to the preset maximum number of iterations, the iteration stops, and the corresponding action space is output to determine the hotspot selection strategy and task unloading strategy for the corresponding virtual power plant.
5. A virtual power plant terminal resource scheduling system based on mobile edge cloud computing, characterized in that, include: The task response module is used to respond to the task calculation requests of terminals in each hotspot area of the virtual power plant, and to determine the terminal of the target service hotspot and the task unloading strategy corresponding to the terminal of the target service hotspot according to the task priority. The latency calculation module is used to determine the execution latency of the computing tasks of the terminal in the target service hotspot under each task priority according to the task offloading strategy of the terminal in the target service hotspot. The benefit calculation module is used to determine the benefit function of each task based on the execution delay of each task and the priority of each task, so as to determine the total benefit of the virtual power plant. The total benefit of the virtual power plant is determined by the benefit function of tasks whose priorities are determined based on the execution latency of each task, including: Based on the execution latency of each task, the benefit functions corresponding to high-priority, medium-priority, and low-priority tasks are determined as follows: ; ; ; In the formula, , , These represent the benefit function values for high-priority, medium-priority, and low-priority tasks, respectively. Hot topic The tolerance latency threshold for internal terminal m, Indicates hot topics The execution latency of terminal m within the terminal. (x) is an indicator function, where x is the indicator parameter. If x is true, then (x) = 1, if x is false, then (x) = 0, As a negative penalty factor, As a task priority, When the value is 1, it indicates a high-priority task. A value of 0 indicates a low-priority task. Tasks with a priority of ∈ (0,1) are considered medium priority tasks. For positive rewards, α is the exponential decay factor; The total benefit of the virtual power plant is determined based on the benefit functions corresponding to high-priority, medium-priority, and low-priority tasks: ; In the formula, M represents the total benefit value, L represents the number of terminals, and L represents the number of hotspots. ; The benefit optimization module is used to construct a benefit optimization problem for a virtual power plant, with the goal of maximizing the total benefit of the virtual power plant and the hotspot selection strategy and task offloading strategy of the virtual power plant as decision variables. The benefit optimization problem includes an objective function and constraints, wherein the objective function and the constraints are as follows: ; ; ; ; In the formula, P represents the objective function. Indicates constraints. This indicates the task unloading constraints. This indicates the constraints for hotspot selection. This indicates that the only constraint for drones to select service hotspots is... This indicates the hot topic selection strategy. Indicates hot topics Being chosen, Indicates hot topics Not selected Indicates the task unloading policy. Indicates hot topics Task m within the system is offloaded to the drone for computation. Indicates hot topics Task m within the system is unloaded and computed locally. The problem-solving module is used to find the optimal solution to the efficiency optimization problem of the virtual power plant, and determine the corresponding hot spot selection strategy and task unloading strategy of the virtual power plant based on the optimal solution.
6. An electronic device, characterized in that, The electronic device includes a memory and a processor; The memory is used to store programs; The processor executes the program to implement the method of any one of claims 1 to 4.
7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 4.
8. A computer program product comprising at least one computer-readable storage medium, characterized in that, The computer-readable storage medium has computer-executable program code instructions stored therein, the computer-executable program code including: The first program code section is configured to respond to task calculation requests from terminals in each hotspot area of the virtual power plant, and determine the terminal of the target service hotspot and the task offloading strategy corresponding to the terminal of the target service hotspot according to the task priority. The second program code section is configured to determine the execution latency of the computing tasks of each terminal in the target service hotspot based on the task offloading strategy of the terminal in the target service hotspot. The third part of the program code is configured to determine the total benefit of the virtual power plant by determining the benefit function of the task priority based on the execution delay of each task. The total benefit of the virtual power plant is determined by the benefit function of tasks whose priorities are determined based on the execution latency of each task, including: Based on the execution latency of each task, the benefit functions corresponding to high-priority, medium-priority, and low-priority tasks are determined as follows: ; ; ; In the formula, , , These represent the benefit function values for high-priority, medium-priority, and low-priority tasks, respectively. Hot topic The tolerance latency threshold for internal terminal m, Indicates hot topics The execution latency of terminal m within the terminal. (x) is an indicator function, where x is the indicator parameter. If x is true, then (x) = 1, if x is false, then (x) = 0, As a negative penalty factor, As a task priority, When the value is 1, it indicates a high-priority task. A value of 0 indicates a low-priority task. Tasks with a priority of ∈ (0,1) are considered medium priority tasks. For positive rewards, α is the exponential decay factor; The total benefit of the virtual power plant is determined based on the benefit functions corresponding to high-priority, medium-priority, and low-priority tasks: ; In the formula, M represents the total benefit value, L represents the number of terminals, and L represents the number of hotspots. ; The fourth part of the program code is configured to construct a virtual power plant benefit optimization problem with the goal of maximizing the total benefit of the virtual power plant and the hotspot selection strategy and task offloading strategy of the virtual power plant as decision variables. The benefit optimization problem includes an objective function and constraints, wherein the objective function and the constraints are as follows: ; ; ; ; In the formula, P represents the objective function. Indicates constraints. This indicates the task unloading constraints. This indicates the constraints for hotspot selection. This indicates that the only constraint for drones to select service hotspots is... This indicates the hot topic selection strategy. Indicates hot topics Being chosen, Indicates hot topics Not selected Indicates the task unloading policy. Indicates hot topics Task m within the system is offloaded to the drone for computation. Indicates hot topics Task m within the system is unloaded and computed locally. The fifth program code section is configured to find the optimal solution to the efficiency optimization problem of the virtual power plant, and determine the corresponding hotspot selection strategy and task unloading strategy of the virtual power plant based on the optimal solution.
Citation Information
Patent Citations
Mobile edge computing task unloading method and device based on transfer learning
CN113504987A
End-side cloud collaborative scheduling method and system based on DDQN (Double Data Quality Network) in edge environment of Internet of Vehicles
CN115243217A