A method, apparatus, equipment, and medium for power task offloading based on cloud-edge collaboration.

By using a cloud-edge collaborative reinforcement learning method, the environmental state of the MEC wireless network is determined and an objective function is constructed. This solves the problems of computational offloading latency and energy consumption in the power IoT in existing technologies, and realizes efficient offloading of power tasks and resource optimization.

CN116390160BActive Publication Date: 2026-05-26INFORMATION & TELECOMM COMPANY SICHUAN ELECTRIC POWER
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INFORMATION & TELECOMM COMPANY SICHUAN ELECTRIC POWER
Filing Date
2023-01-17
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

The existing computing offloading mechanism of the power Internet of Things fails to effectively consider the dynamic nature of the environment, resulting in increased latency and energy consumption for terminal devices when selecting MEC servers, which cannot meet the power service requirements of low latency and high bandwidth.

Method used

A cloud-edge collaborative reinforcement learning approach is adopted. By determining the environmental state of the MEC wireless network, an objective function is constructed to minimize the power task offloading delay. The solution is obtained using a pre-set reinforcement learning offloading model, and a suitable MEC server is selected to offload the computational tasks.

Benefits of technology

It effectively reduced the processing delay of power tasks, improved system performance, adapted to changing network environments, and optimized the use of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116390160B_ABST
    Figure CN116390160B_ABST
Patent Text Reader

Abstract

This invention discloses a method, apparatus, device, and medium for power task offloading based on cloud-edge collaboration. The power task offloading method provided by this invention considers a multi-user multi-access edge computing (MEC) wireless network, where multiple mobile devices can associate via wireless channels and offload tasks to an MEC server connected to a base station (BS) for execution. The decision of whether to execute the computing task locally at the user equipment or offload it to the MEC server should adapt to time-varying network dynamics, taking into account the dynamic nature of the environment, including computing resources, bandwidth resources, and channel conditions, selecting appropriate edge server computing tasks, and minimizing computational cost in terms of total latency by combining the DQN algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of communication technology, and specifically relates to a method, apparatus, equipment and medium for power task offloading based on cloud-edge collaboration. Background Technology

[0002] The Power Internet of Things (PIoT) is the application of the Internet of Things (IoT) in smart grids. It can play a role in all aspects of the smart grid, including power generation, transmission, distribution, and consumption. By integrating communication and power infrastructure resources and utilizing advanced information and communication technologies, PIoT allows information exchange between interconnected devices and further improves the unit efficiency of the power system. However, facing challenges such as differentiated business needs and diverse application scenarios, the construction of the Power IoT requires the introduction of key 5G technologies such as edge computing and low-latency technology to achieve flexible and differentiated communication capabilities. 5G networks, due to their low latency, low cost, and high security in carrying control and data acquisition services, can meet the diverse critical communication needs of the power grid.

[0003] The power industry encompasses generation, transmission, transformation, distribution, and consumption, with numerous business types across each stage, including electricity consumption information collection, precise load control, mobile inspection, and distribution network electrification. For the most latency-critical services within the power Internet of Things (IoT), deploying edge computing platforms can reduce latency in wireless-to-core transmission and signaling processing. Multi-access edge computing (MEC) places data center-level computing, storage, and network resources closer to users and end devices, at the network edge. MEC servers are designed to serve consumer and enterprise applications requiring low latency and high bandwidth. Devices can offload computing tasks to nearby MEC servers to reduce processing latency and conserve battery power.

[0004] Existing technologies utilize collaborative compute offloading mechanisms between centralized cloud servers and distributed MEC server resources. These mechanisms consider collaboration between end users, MEC servers, and the centralized cloud, and employ a partial compute offloading architecture where user computational tasks are partially split and offloaded to MEC servers or the centralized cloud, while the non-offloadable portions are executed locally. However, most existing compute offloading mechanisms are based on one-time optimizations and cannot characterize long-term compute offloading performance. They also fail to consider environmental dynamics, such as bandwidth resources, time-varying channel conditions, energy availability at mobile devices, and the computing power of different MEC servers. If the selected MEC server experiences heavy workloads and degraded channel conditions, the terminal device may take longer to offload data and receive results. Summary of the Invention

[0005] The purpose of this invention is to provide a method, apparatus, device and medium for power task offloading based on cloud-edge collaboration, so as to reduce the processing delay of power tasks.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] Firstly, a method for offloading power tasks based on cloud-edge collaboration includes the following steps:

[0008] Determine the environmental status of the MEC wireless network within the target area;

[0009] Based on the aforementioned environmental conditions, an objective function is constructed with the goal of minimizing the total delay time for power task unloading.

[0010] The objective function is solved using a pre-set reinforcement learning unloading model to obtain the action policy corresponding to the minimum objective function;

[0011] Execute the aforementioned action strategy to complete the power task offloading.

[0012] Furthermore, in the step of determining the environmental status of the MEC wireless network within the target area, the environmental status specifically includes: channel condition status, actual available energy of user equipment, and computing resources that can be allocated to the MEC server.

[0013] Furthermore, in the step of constructing the objective function with the goal of minimizing the total delay time of power task unloading, the objective function is as follows:

[0014]

[0015] Where, α k Let α represent the task offloading decision for user device k. k =1 indicates that the task is offloaded to the MEC server, α k =0 indicates that the task is computed locally; Indicates task T on user device k k Total local execution time; This represents the total unloading and computation delay for user device k task unloading.

[0016] Furthermore, the total unloading and computational delay of the user equipment k task unloading Calculated using the following formula:

[0017]

[0018] in, This represents the task T that user equipment k offloads on MEC m. k Total execution time; This represents the total execution time of tasks offloaded from user device k to MEC m and n; This represents the task T that user device k uninstalls in the cloud. k Total execution time; τ k Indicates the deadline for task calculation.

[0019] Furthermore, the task T that user equipment k unloads on MEC m k Total execution time as follows:

[0020]

[0021] The total execution time for tasks offloaded from user device k to MEC m and n is:

[0022]

[0023] Task T, which is uninstalled by user device k in the cloud k The total execution time is:

[0024]

[0025] In the above formula, This represents the transmission delay of offloading the task from user equipment k to MEC m; Indicates task T on MEC m k The computational delay; This indicates the forwarding delay between MEC m and MEC n; Indicates task T on MEC n k Execution delay; This indicates the forwarding latency between MEC m and the cloud server; It's an execution delay on the cloud server.

[0026] Furthermore, the task T on the user equipment k k Total local execution time Represented as:

[0027]

[0028] in, It is task T k Average waiting time; k Indicates task T k Execution delay on user device k.

[0029] Furthermore, in the step of solving the objective function using a pre-set reinforcement learning unloading model, the training method of the reinforcement learning unloading model specifically includes the following steps:

[0030] The channel condition status of the wireless network in the target area, the actual available energy of user equipment, and the computing power resources that the MEC server can allocate are obtained in order to construct the state vector of the environment.

[0031] Based on the constructed state vector of the environment, corresponding actions are generated according to the reinforcement learning strategy;

[0032] A control strategy is formulated based on the state vector of the environmental state and the actions; wherein, the control strategy includes fixed task unloading and server selection strategies;

[0033] The control strategy is executed, and the resulting reward is calculated based on the feedback from the environment after execution.

[0034] The updated Q-function is calculated based on the environmental state, the action, and the obtained reward.

[0035] Based on the updated Q-function, the loss function is set and the gradient is calculated. The model with the smallest loss function is used as the final reinforcement learning unloading model.

[0036] Secondly, a power task offloading device based on cloud-edge collaboration includes:

[0037] The environment determination module is used to determine the environmental status of the MEC wireless network within the target area;

[0038] The function construction module is used to construct an objective function based on the environmental state, with the goal of minimizing the total delay time of power task unloading;

[0039] The solution module is used to solve the objective function using a pre-set reinforcement learning unloading model to obtain the action policy corresponding to the minimum objective function;

[0040] The execution module is used to execute the action strategy and complete the power task unloading.

[0041] Thirdly, an electronic device includes a processor and a memory, the processor being configured to execute a computer program stored in the memory to implement the aforementioned cloud-edge collaborative power task offloading method.

[0042] Fourthly, a computer-readable storage medium stores at least one instruction that, when executed by a processor, implements the aforementioned cloud-edge collaborative power task offloading method.

[0043] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0044] The cloud-edge collaborative power task offloading method provided by this invention considers a multi-user multi-access edge computing (MEC) wireless network, in which multiple mobile devices can be associated through wireless channels and offload tasks to an MEC server connected to a base station (BS) for execution. The decision of whether to execute computational tasks locally at the user equipment or offload them to the MEC server should adapt to time-varying network dynamics, taking into account the dynamic nature of the environment, including computing resources, bandwidth resources, and channel conditions. Appropriate edge server computational tasks are selected, and the DQN algorithm is combined to minimize computational cost in terms of total latency. Attached Figure Description

[0045] The accompanying drawings, which form part of this application, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:

[0046] Figure 1 This is a flowchart of a power task offloading method based on cloud-edge collaboration according to an embodiment of the present invention;

[0047] Figure 2 This is a schematic diagram of strategy generation for the reinforcement learning unloading model in an embodiment of the present invention;

[0048] Figure 3 This is a schematic diagram of the system model in an embodiment of the present invention;

[0049] Figure 4 This is a simulation comparison diagram in an embodiment of the present invention;

[0050] Figure 5 This is a structural block diagram of a power task offloading device based on cloud-edge collaboration according to an embodiment of the present invention;

[0051] Figure 6 This is a structural block diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0052] The present invention will now be described in detail with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other.

[0053] The following detailed description is exemplary and intended to provide further detailed explanation of the invention. Unless otherwise specified, all technical terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. The terminology used in this invention is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention.

[0054] Example 1

[0055] like Figures 1-4As shown, a power task offloading method based on cloud-edge collaboration includes the following steps:

[0056] S1. Determine the environmental status of the MEC wireless network within the target area.

[0057] Specifically, the MEC wireless network within the target area involved in this scheme is a wireless network system consisting of M base stations (BSs), each base station connecting to an MEC server. Considering a group of K users, where each user device k∈K can connect to any MEC server within range, the system model is as follows: Figure 3 As shown. User equipment has limited computing resources, so it will offload data to the MEC server for processing. In this scheme, at each time t, each user equipment k∈K has a computational task that needs to be computed, such as electricity consumption information collection, precise load control, mobile inspection, and distribution network electrification.

[0058] More specifically, a task can be computed locally on a user device or offloaded to an MEC server or a cloud server. For each user device k, a task T is defined. k ={s(d k ),τ k ,z k}, k∈K, where s(d k ) is the task data size of user equipment k, expressed in bits as the computational input, τ k It is the task calculation deadline, z k This refers to the computational workload or intensity, measured in CPU cycles required per bit. Furthermore, all MEC servers in this scheme belong to the same network operator, allowing computational data to be partitioned among the MEC servers for collaborative execution. When a shardable task is offloaded to an MEC server, a decision is made based on whether the task is executed on a single MEC server or a portion of the task is forwarded to other MEC servers or a remote cloud, taking into account the workload and computational resource status of each MEC server.

[0059] More specifically, a user equipment (UE) can associate with any MEC server based on wireless channel conditions to offload data to that MEC server. However, the associated MEC server may have limited resources (wireless or computing resources) available to the user. If the MEC server lacks sufficient computing resources for an UE's offloading task, it can partially process the task and forward the remainder to other MEC servers or cloud servers for further processing. In summary, there are four possible decisions: 1) execute the task locally; 2) execute the task on the MEC server; 3) forward the task to other MEC servers; 4) offload the task to a cloud server.

[0060] Executing tasks locally offers the advantage of reduced latency; however, this decision is limited by the availability of local resources. On the other hand, offloading tasks to cloud servers is frequently used in practice due to their computing power, but this leads to high latency. Offloading tasks to MEC servers is considered the most efficient decision because it can utilize available computing resources at the network edge to reduce latency. The switching time in this invention, i.e., the connection time between the user and the MEC server, is negligible compared to the transmission and execution time of the computing task. System time is divided into continuous time frames of equal length t in proportions of several seconds.

[0061] MEC server resources are virtualized and shared by multiple users. Resource demands that cannot be met by one MEC server can be met by any other MEC server. The resource allocation function for each user device k in MEC m uses... It means that, among them Used to represent the computing resources allocated within the MEC server. Used to indicate the proportion of bandwidth resources allocated by MEC m to user equipment k.

[0062] Specifically, the environmental conditions involved in this step include: channel conditions, the actual available energy of user equipment, and the computing resources that the MEC server can allocate.

[0063] S2. Based on the environmental state, construct an objective function with the goal of minimizing the total delay time of power task unloading.

[0064] The server selection variable is defined as follows: in This indicates that user equipment k selects MEC m, otherwise Given that the radio channel between the user equipment (UE) and its associated MEC is a time-varying channel, the channel power gain of the radio link between UE k and MEC m is modeled as a random variable. The value range can be divided and quantized into L discrete levels. Each level corresponds to a state of the Markov chain, thus forming an L-element state space G = {G0, G1, ..., G...}. L-1}; The channel state at time t can be represented as:

[0065] When user device k first associates with an MEC server, user device k will comprehensively consider the available computing resources of all possible MEC servers. Bandwidth resources and channel condition state Select the MEC server with the best overall performance. The overall performance is determined using a utility function. The formula is as follows:

[0066]

[0067] Among them, β1, β2, and β3 are three scaling factors.

[0068] The uplink transmission rate of user equipment k associated with MEC m can be given by the following formula:

[0069]

[0070] Among them, P k It is the transmission power of user equipment k. It is the power of the Gaussian noise at point k in the user equipment. This interference originates from other user equipment k' that shares the same frequency band as user equipment k. The transmit power of the user equipment is set to a fixed value; therefore, for simplicity, the interference term... It can be absorbed into the noise term. middle.

[0071] The instantaneous data rate of user equipment k associated with MEC m is given by the following formula:

[0072]

[0073] Where K is the set of all user devices, Choose variables for the server, W m For the total bandwidth of MEC m, each user equipment k associated with MEC m is allocated a proportion of bandwidth W.

[0074] In densely populated scenarios, due to spectrum resource constraints, all base stations share the same bandwidth W, i.e., general spectrum sharing. Furthermore, the selection of a server for offloading is only accepted when there are sufficient spectrum resources to meet its needs.

[0075] The task offloading decision for user equipment k is represented as α. k , where α k =1 indicates that the task is offloaded to the MEC server, α k =0 indicates that the task is computed locally.

[0076] Based on the defined instantaneous data rate, the transmission delay of the task offloading from user equipment k to MEC m is expressed as:

[0077]

[0078] in, K represents the transmission delay of offloading a task from user equipment k to MEC m.m It is a group of user devices that provide MEC m services.

[0079] When MEC m does not have enough resources to meet the needs of user equipment, MEC m forwards the request to another MEC n with sufficient resources via the link.

[0080] The proportion of user equipment forwarding tasks is expressed as follows: Therefore, user equipment requests are served by different MEC servers at different latency costs. The forwarding latency between MEC m and MEC n is... To represent, as shown below:

[0081]

[0082] in, It is the transmission rate of the link between MEC m and MEC n; This indicates the forwarding delay between MEC m and MEC n; This indicates the proportion of tasks forwarded by user devices.

[0083] When no MEC server with sufficient computing resources can complete the remaining tasks of user device k, MEC m will forward the request to a remote cloud server via a wired backhaul link. The forwarding latency between MEC m and the cloud server is defined as... in Given by the following formula:

[0084]

[0085] in, It is the transmission rate between MEC m and the remote cloud; This indicates the forwarding latency between MEC m and the cloud server.

[0086] Assume that each user device k∈K has a task T k It requires the use of device k's local computing resources. Task T k Local computation execution time l k Therefore, task T k The calculation requires CPU energy η k The CPU computing power consumption at user equipment k is expressed as:

[0087]

[0088] Where v is a constant parameter related to the CPU hardware architecture. For local computing resources.

[0089] Besides CPU power consumption, task T k The calculation also requires execution time. k Therefore, task T k The execution latency of user equipment k is given by the following formula:

[0090]

[0091] The actual available energy f of user equipment k in each time slot t k (t) can be modeled as a random variable and divided into discrete levels, denoted by F = {F0, F1, ..., F...} N-1} represents the number of available energy states.

[0092] Task T on user device k k Total local execution time It is given by the following formula:

[0093]

[0094] in, It is task T k The average waiting time until it is executed locally by device k; k Indicates task T k Execution delay on user device k.

[0095] The computing resources of MEC m allocated to user device k are represented as follows: This can be measured in CPU cycles. In a network model, multiple user devices can access the same MEC server at a given time and share the computing resources on the MEC server. Therefore, computing power... It can be modeled as a random variable and divided into discrete levels, denoted by H = {H0, H1, ..., H...} J-1} represents the number of available computing power states.

[0096] In the model, if computing resources available for local processing are limited or scarce, the user device will offload the task to a MEC server. If sufficient resources are available, the terminal device will detect all MEC servers within its service range, comprehensively consider their computing and operational resources, and select the optimal MEC server to offload the task. k Uninstall it on the server and execute it.

[0097] The available computing resources (i.e., CPU cycle frequency) of MEC m are represented by Δ. m The total calculation allocation must satisfy:

[0098]

[0099] MEC m on task T k computation delay It is given by the following formula:

[0100]

[0101] Therefore, the task T that user equipment k offloads on MEC m k The total execution time is as follows:

[0102]

[0103] However, if (That is, MEC m does not have enough computing resources to meet the computation deadline), MEC m will postpone task T k The remaining data is forwarded to other MECs n with sufficient resources to meet the demand. It is a task T on MEC n k The execution delay can be calculated using the following formula.

[0104]

[0105] Therefore, the total execution time of the tasks offloaded from user equipment k to MEC m and n becomes:

[0106]

[0107] When no single MEC has sufficient computing resources to meet the task computation deadline, i.e. MEC m forwards the task to the cloud server. Therefore, the task T is unloaded by user device k in the cloud. k The total execution time becomes:

[0108]

[0109] In the above formula, This represents the transmission delay of offloading the task from user equipment k to MEC m; Indicates task T on MEC m k The computational delay; This indicates the forwarding delay between MEC m and MEC n; Indicates task T on MEC n k Execution delay; This indicates the forwarding latency between MEC m and the cloud server; It's an execution delay on the cloud server.

[0110] in It can be calculated from (13). Furthermore, the total unloading and computational delay for user equipment k are as follows:

[0111]

[0112] in, This represents the task T that user equipment k offloads on MEC m. k Total execution time; This represents the total execution time of tasks offloaded from user device k to MEC m and n; This represents the task T that user device k uninstalls in the cloud. k Total execution time; τ k Indicates the deadline for task calculation.

[0113] This solution treats server selection, resource allocation, computational task offloading, and switching as a deep reinforcement learning process. An agent is responsible for collecting state information from each user device and MEC server, then aggregating all this information into a constructed system state. Once the optimal strategies for server selection, resource allocation, and task offloading are determined, the agent sends a notification to the user. The computational cost in terms of total latency is considered as the total computational task T. k The total time required (including offloading latency). To minimize computational latency costs, the total latency D(x,α,ω,P) is used for tasks offloading to the MEC server or remote cloud; therefore, the objective function constructed in this application is as follows:

[0114]

[0115] Where, α k Let α represent the task offloading decision for user device k. k =1 indicates that the task is offloaded to the MEC server, α k =0 indicates that the task is computed locally; Indicates task T on user device k k Total local execution time; This represents the total unloading and computation delay for user device k task unloading.

[0116] The following procedures were established for joint server selection, coordinated unloading, and failover:

[0117]

[0118] Constraint (18a) guarantees that the total computing resource allocation for each MEC server does not exceed the maximum computing capacity of that MEC server. Constraint (18b) guarantees that each user device can associate with at most one MEC at a time. Constraint (18c) guarantees that the sum of the spectrum allocations for all user devices must be less than or equal to the total available spectrum for each MEC m. If a user device cannot associate with any MEC at time t, i.e. It will compute the task locally, namely αk =0.

[0119] Decomposing the problem, once the server selection x and the unloading decision α are given, problem (18) can be reformulated as a convex problem, as follows:

[0120]

[0121] Where, φ k It is a weight, relative to the data z of each user device. k s(d k Proportional to It is a scaling factor.

[0122] S3. Solve the objective function using a pre-set reinforcement learning unloading model to obtain the action policy corresponding to the minimum objective function.

[0123] It's important to note that in reinforcement learning, the agent learns optimal policies by interacting with its environment. These optimal policies are achieved by recursively correcting the agent's errors after each interaction. The time-varying channel conditions between the mobile device and the MEC server, the available energy of the mobile device, the computational workload, and the computing power of different MEC servers are all dynamically changing. The goal is then to determine whether the mobile user should offload its computational tasks to the MEC server or execute them locally, and based on the current state, which MEC servers and their resources are best suited to serve the user's computational task requests. Formally, each action taken by the agent is defined by probability and associated with the agent's policy. When the agent interacts with the environment based on a certain action, it is rewarded and its state is changed. Therefore, the goal of each agent is to implement an optimal policy that maximizes the total reward.

[0124] Within the framework, the agent learns its optimal control policy, thereby minimizing the computational cost in terms of total latency, as defined in (17).

[0125] As an optional implementation of the present invention, the training method of the reinforcement learning unloading model specifically includes the following steps:

[0126] (1) Obtain the channel condition status of the wireless network in the target area, the actual available energy of the user equipment, and the computing power resources that the MEC server can allocate, so as to construct the state vector of the environment status.

[0127] Specifically, the network state of user equipment k in time slot t can be determined by random variables. status random variable f k State F k (t), and random variables status Therefore, the state vector can be described as...

[0128] (2) Based on the state vector of the constructed environment state, generate the corresponding action according to the reinforcement learning strategy;

[0129] Specifically, at the beginning of time slot t, the agent strategically determines the action of user equipment k according to a fixed control policy.

[0130] (3) Formulate control strategies based on the state vector of the environment and actions; among which, the control strategies include fixed task unloading and server selection strategies;

[0131] Specifically, joint control operations Observe the network state χ at the beginning of each decision time slot t k (t) is then generated according to Φ, where (Φ (α) ,Φ (x) These are fixed task uninstallation and server selection strategies, respectively.

[0132] (4) Implement control strategies and calculate the resulting rewards based on feedback from the environment after implementation;

[0133] Specifically, after taking action, the AI ​​will receive a reward or penalty from the environment, treating a negative reward as an estimated total computational delay cost. If a task fails to meet its computation deadline, it is considered unsuccessful, and a penalty value p will be added to the reward. This penalty value will be updated based on the task's deadline value.

[0134] (5) Calculate the updated Q function based on the environmental state, actions, and rewards obtained;

[0135] Specifically, in the Q-learning method, the Q function is updated at the beginning of each time step. This is based on the network state χ... k Given the network state χ(t), action y(t), received computational cost D(χ(t), y(t)), and the resulting network state χ(t+1) at the next time step t+1, the agent updates its Q-function as follows:

[0136]

[0137] Where γ is the discount factor, δ t ∈[0,1] is a learning rate that varies over time.

[0138] More specifically, to mitigate the costs associated with Q-learning, a Deep Q-Network (DQN) is employed to estimate the Q-function online. In DQN, Q(χ,y) is approximated by Q(χ,y,θ), i.e., Q(χ,y)≈Q(χ,y,θ), where θ is the weight set of the DQN.

[0139] (6) Based on the updated Q function, set the loss function and calculate the gradient. The model with the smallest loss function is used as the final reinforcement learning unloading model.

[0140]

[0141] Specifically, the transition is a tuple consisting of four parts: the current state χ(t), the current action y(t), the current reward D(χ(t), y(t)), and the next state χ(t+1), denoted by ζ. (t) =(χ(t),y(t),D(χ(t),y(t)),χ(t+1)) means that at the end of each duration t, it is stored in an experience replay pool of finite size U.

[0142] More specifically, let ψ (t) ={ζ (t-U+1) ,…,ζ (t)} represents the memory pool. The agent randomly samples experience once at each time point t, i.e., from ψ... (t) The initial S transition mini-batch DQN is trained in the direction that minimizes the loss function, and its gradient is calculated.

[0143] DQN is a set of parameters The target DQN is used to evaluate action values, specifically... The function is computed based on this single objective DQN, and Q(χ,y; θ) (t+1) The target DQN is computed based on the DQN trained on it. The target DQN is a copy of an earlier DQN; in the next iteration, the weight set of the target DQN is... The weight set of the DQN updated to time t, i.e., θ (t) .

[0144] S4. Execute the action strategy to complete the power task unloading.

[0145] To verify the performance of the server selection and collaborative offloading method based on edge computing proposed in this invention, three simulation schemes were designed for comparison: server computing: all user device tasks are offloaded to MEC servers or cloud servers; local computing: all user device tasks are executed locally on the user devices; random computing: all user device tasks are randomly executed locally on the user devices or offloaded to MEC servers and cloud servers. The specific simulation process is as follows:

[0146] First, initialize the system model. Consider a wireless network consisting of 5 base stations (BS), each of which is connected to an MEC server. Consider a group of 10 user devices, where each user device can connect to any MEC server within range. At each time t, each user device has a computational task that needs to be computed.

[0147] When selecting edge servers, the available computing power, bandwidth, and channel conditions of the edge servers are comprehensively considered. User equipment determines whether its own available energy and computing power can meet the task deadline; if not, the task is offloaded to an edge server for execution. The edge server decides, based on its own computing power, whether to collaborate with other edge servers or with cloud servers to minimize execution time. This system model is as follows: Figure 3 As shown.

[0148] To verify the performance of the server selection collaborative offloading method based on edge computing proposed in this invention, it was simulated and compared with server-based computing, local computing, and random computing under the same environment.

[0149] The system simulation environment parameters are set as follows:

[0150] 1) A wireless network consisting of 5 base stations (BS) and 10 user equipments. Each user equipment can connect to any MEC server within range, and each user equipment has a computing task that needs to be computed.

[0151] 2) The task data size of the user equipment is randomly generated in the range of 1MB-10MB, the task calculation deadline is randomly generated in the range of 0.1s-1s, and the task workload is evenly distributed in the range of 452.5 cycles / bit to 737.5 cycles / bit.

[0152] 3) For each user equipment, computing resources are randomly generated in the range of 0.5 GHz to 1.0 GHz, and the local computing power of each user equipment per cycle is set to equal 10. -8 Joules per cycle, the available energy of the user equipment is evenly distributed in the range of 1 joule / second to 10 joules / second;

[0153] 4) The computing resources of the MEC server are assumed to be in the range of 2.3GHz to 2.8GHz, with 24 cores / 48 hyperthreads;

[0154] 5) The transmission power of all terminal devices is 1W;

[0155] 6) All MECs share the same bandwidth, randomly generated from 25-32MHz;

[0156] 7) The transmission rate between edge servers is randomly generated within the range of 20-25 Mbps;

[0157] 8) The transmission rate between the cloud server and the edge server is randomly generated within the range of 50-120Mbps.

[0158] Simulation results are as follows Figure 4 As shown, from Figure 4 As can be seen, the delay used by the method of the present invention after 2000 episodes is less than that of the other three comparative methods.

[0159] The simulation results show that by employing a deep Q-learning method to learn the task scheduling strategy, this edge computing-based server selection and cooperative offloading method can be used for joint server selection and cooperative offloading in multi-access edge wireless networks. Numerical results demonstrate that the proposed DRL-based algorithm outperforms other computational methods in terms of latency, thus improving system performance.

[0160] Example 2

[0161] like Figure 5 As shown, a power task offloading device based on cloud-edge collaboration includes:

[0162] The environment determination module is used to determine the environmental status of the MEC wireless network within the target area;

[0163] The function building module is used to construct an objective function based on the environment state, with the goal of minimizing the total delay time of power task unloading;

[0164] The solution module is used to solve the objective function using a pre-set reinforcement learning offloading model to obtain the action policy corresponding to the minimum objective function;

[0165] The execution module is used to execute action strategies and complete the unloading of power tasks.

[0166] Example 3

[0167] like Figure 6As shown, the present invention also provides an electronic device 100 for implementing a cloud-edge collaborative power task offloading method; the electronic device 100 includes a memory 101, at least one processor 102, a computer program 103 stored in the memory 101 and executable on the at least one processor 102, and at least one communication bus 104. The memory 101 can be used to store the computer program 103, and the processor 102 implements the steps of the cloud-edge collaborative power task offloading method of Embodiment 1 by running or executing the computer program stored in the memory 101 and calling data stored in the memory 101.

[0168] The memory 101 may primarily include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created based on the use of the electronic device 100 (such as audio data), etc. In addition, the memory 101 may include non-volatile memory, such as hard disk, RAM, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other non-volatile solid-state storage device.

[0169] At least one processor 102 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 102 may be a microprocessor or any conventional processor. Processor 102 is the control center of electronic device 100, connecting various parts of electronic device 100 via various interfaces and lines.

[0170] The memory 101 in the electronic device 100 stores multiple instructions to implement a power task offloading method based on cloud-edge collaboration, and the processor 102 can execute multiple instructions to achieve the following:

[0171] Determine the environmental status of the MEC wireless network within the target area;

[0172] Based on the environmental conditions, an objective function is constructed with the goal of minimizing the total delay time of power task unloading.

[0173] The objective function is solved using a pre-set reinforcement learning offloading model to obtain the action policy corresponding to the minimum objective function;

[0174] Execute the action strategy to complete the power task unloading.

[0175] Example 4

[0176] If the modules / units integrated in the electronic device 100 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, and read-only memory (ROM).

[0177] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0178] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0179] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0180] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0181] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0182] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A method for offloading power tasks based on cloud-edge collaboration, characterized in that, Includes the following steps: Determine the environmental status of the MEC wireless network within the target area. The MEC wireless network consists of M base stations, each connected to an MEC server, serving K user devices. The environmental conditions include: the channel condition status of the time-varying wireless channel between the user equipment and the MEC server. Actual available energy of user equipment and the computing resources that can be allocated to the MEC server. Channel condition states, actual available energy, and allocable computing resources are all quantified as discrete-level random variables. Based on the environmental conditions, an objective function is constructed with the goal of minimizing the total delay time of power task unloading. The objective function is: ; in, Indicates the total delay. Indicates user equipment k Task unloading decision, This indicates that the task has been offloaded to the MEC server. This indicates that the task is computed locally; Indicates user equipment k On the task Total local execution time, , It is a task The average waiting time until it is received by the device. k Execution continues locally. Indicates task In user equipment k Execution delay; Indicates user equipment k The total unloading and computation delay for task unloading is: in, Indicates user equipment k In MEC m Upload / Unload Tasks Total execution time ; For user equipment k Uninstall to MEC m and n The total execution time of the task. ; For user equipment k Tasks unloaded in the cloud Total execution time ; Indicates the deadline for task calculation; Indicates task The CPU energy required for computation; This represents the transmission delay of offloading the task from user equipment k to MEC m; Indicates a task on MEC m The computational delay; This represents the forwarding delay between MEC m and MEC n; This represents a task on MEC n. Execution delay; This indicates the forwarding latency between MEC m and the cloud server; This indicates execution latency on the cloud server; The objective function is solved using a pre-defined reinforcement learning offloading model to obtain the action policy corresponding to the minimum objective function; the deep reinforcement learning model is a deep Q-network (DQN), and the reinforcement learning offloading model is based on the state vector. Actions are generated according to the strategy. Actions include unloading decisions Server selection The reward is a negative value of the total delay D; in, Channel state, The actual available energy state of user equipment. The status of allocable computing resources for the MEC server. For task unloading decisions, Select variables for the server; the reinforcement learning offloading model is a model based on a deep Q-network (DQN); Execute the action strategy to complete the power task offloading; If a task does not meet its deadline, it is considered unsuccessful and a penalty value p will be added to the reward. The penalty value will be updated based on the task's deadline value.

2. A power task offloading device based on cloud-edge collaboration, used to implement the power task offloading method based on cloud-edge collaboration as described in claim 1, characterized in that, include: The environment determination module is used to determine the environmental status of the MEC wireless network within the target area; The function construction module is used to construct an objective function based on the environmental state, with the goal of minimizing the total delay time of power task unloading; The solution module is used to solve the objective function using a pre-set reinforcement learning unloading model to obtain the action policy corresponding to the minimum objective function; The execution module is used to execute the action strategy and complete the power task unloading.

3. An electronic device, characterized in that, It includes a processor and a memory, the processor being used to execute a computer program stored in the memory to implement the power task offloading method based on cloud-edge collaboration as described in claim 1.

4. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction, which, when executed by a processor, implements the power task offloading method based on cloud-edge collaboration as described in claim 1.