Task unloading optimization method based on adaptive attention and reinforcement learning
By adopting adaptive attention and reinforcement learning methods in the cloud-edge collaborative computing environment, dynamically generate task unloading strategies, solving the problems of insufficient real-time and poor dynamic adaptability in the existing technology, and achieving efficient task unloading and resource utilization.
Patent Information
- Application Number
- CN202510337787.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-24
AI Technical Summary
The existing task offloading methods have problems such as insufficient real-time, poor dynamic adaptability and low resource utilization efficiency in complex and dynamic network environments.
The task offload optimization method based on adaptive attention and reinforcement learning is adopted. By collecting feature data of the task and environment, combining adaptive attention mechanism and reinforcement learning technology, the optimal offload strategy is dynamically generated, and the computing tasks are reasonably allocated to local, edge nodes or cloud processing.
It improves the dynamic adaptability of the task unloading strategy, realizes real-time adjustments based on task characteristics and environment changes, reduces the computational complexity of unloading decisions, and improves the overall performance and stability of the system.
Smart Images

Figure CN120196442A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cloud-edge collaborative edge computing, and particularly to a task offloading optimization method based on adaptive attention and reinforcement learning. Background Art
[0002] In a cloud-edge fusion computing environment, the task offloading technology refers to dynamically allocating the computing tasks of user devices to appropriate computing nodes (such as local, edge nodes or the cloud) for processing, so as to optimize the system performance. This technology is particularly suitable for devices with limited resources, such as smart phones, Internet of Things devices, etc., which can significantly reduce the computing burden of the devices, reduce energy consumption, and improve the overall task processing efficiency. With the development of edge computing and cloud computing, task offloading has been widely applied in fields such as intelligent transportation, Internet of Things, industrial automation, etc., and has played an important role especially in scenarios with real-time computing and low latency requirements. However, in the face of complex and dynamic network environments, existing task offloading methods still have some challenges.
[0003] Currently, existing task offloading methods mainly include offloading decisions based on static rules and offloading strategies based on optimization algorithms. Traditional static rule methods determine offloading targets through predefined decision rules (such as task complexity, network bandwidth, device load, etc.). These methods are simple to implement and have low computational overhead, and are suitable for stable environments. However, static rule methods cannot effectively adapt to the dynamic changes of network and task loads, and are prone to performance degradation in complex scenarios. To solve this problem, some methods introduce optimization algorithms (such as genetic algorithms, particle swarm optimization, etc.) for multi-objective optimization. These methods can improve the flexibility of task offloading to a certain extent, but have high computational overhead and are difficult to respond to environmental changes in real time, which limits their popularization in practical applications.
[0004] With the development of deep learning and reinforcement learning technologies, some emerging methods have begun to apply these technologies to the optimization of task offloading. Through technologies such as deep reinforcement learning (DRL), the model can dynamically adjust offloading decisions according to real-time environmental data, improving the system's adaptive ability. However, existing deep reinforcement learning methods still face problems such as high computational complexity, long training time, and poor real-time performance in complex multi-node cloud-edge collaborative environments, especially in dealing with large-scale and rapidly changing task scenarios, and their effects are limited. The following are the related researches and applications of two existing technologies:
[0005] CN114489977A, "A Task Offloading Method and Device for a Mobile Edge Computing System", proposes a task offloading method based on deep reinforcement learning. The core technology lies in optimizing the task offloading process through the Deep Q-Network (DQN) algorithm, aiming to simultaneously minimize the energy consumption and latency of mobile devices and improve the privacy protection level during the task offloading process. This method can perform joint optimization among multiple objectives, and better balance the latency, energy consumption, and privacy protection of tasks. However, this method still has certain limitations. First, although deep reinforcement learning has strong adaptive capabilities, when the DQN algorithm is used to handle task offloading problems in large-scale and complex environments, the training process is relatively complex and the computational overhead is large, making it difficult to meet real-time requirements. Second, the DQN algorithm lacks precise weighted processing of multi-dimensional features, and may not fully consider the priorities of task features in different scenarios, thus affecting the accuracy of the task offloading strategy. Therefore, in practical applications, especially in scenarios that require efficient real-time response, there is still room for optimization of this method.
[0006] CN114564304A, "A Task Offloading Method for Edge Computing", proposes a method to optimize the task offloading process by constructing a network model and a computing model. This method relies on the latency and energy consumption calculation models, and combines the particle swarm algorithm to optimize the task offloading location to achieve the optimal allocation of tasks, thereby optimizing energy consumption, latency, and satisfaction. By constructing a scenario satisfaction model and a penalty function, this method can find the best offloading solution in multi-objective optimization, thus improving the efficiency of mobile edge computing. However, the deficiency of this method is that although the particle swarm algorithm can find an approximate optimal solution, when dealing with complex and dynamically changing network environments, the computational complexity of the algorithm is relatively high, which may lead to insufficient real-time performance, especially in the case of large-scale tasks and multi-node cooperation. In addition, based on the pre-set computing model and static optimization strategy, there is a lack of adaptive adjustment to task characteristics and environmental changes, which may lead to a decline in performance in actual dynamic environments. Therefore, this method still needs to be further improved in terms of real-time performance and adaptability, especially in high-dynamic and variable task offloading scenarios. Summary of the Invention
[0007] To solve the problems of real-time performance, dynamic adaptability, and resource utilization efficiency of task offloading in the cloud-edge collaborative computing environment, the present invention provides a task offloading optimization method based on adaptive attention and reinforcement learning. This method collects the characteristic data of tasks and the environment, combines the adaptive attention mechanism and reinforcement learning technology, dynamically generates the optimal offloading strategy, and reasonably allocates computing tasks to be processed locally, at edge nodes, or in the cloud. This method can improve the dynamic adaptability of the task offloading strategy, enabling it to adjust offloading decisions in real time according to task characteristics and environmental changes. Enhance the feature processing mechanism, extract the key features of tasks and the environment through the adaptive attention module, and improve the decision-making efficiency and accuracy. Reduce the computational complexity of offloading decisions and achieve lightweight real-time deployment on edge devices. Introduce a dynamic optimization mechanism to enable the offloading scheme to adapt to dynamic environmental changes through real-time learning and updating, and enhance the overall performance and stability of the network system.
[0008] The technical solution of the present invention to solve the above technical problems is to design a task offloading optimization method based on adaptive attention and reinforcement learning, characterized in that the method includes the following steps:
[0009] Step 1: Data collection and preprocessing
[0010] Collect the characteristic data of a certain number of tasks, where the characteristic data includes task characteristics and environmental characteristics; among them, task characteristics include computing requirements, data volume, and latency requirements; environmental characteristics include network bandwidth, edge node load, and cloud response time; perform preprocessing of cleaning and standardization on the collected characteristic data; use the task type ID as a label, and combine the computing requirements, data volume, latency requirements, bandwidth, node load, and cloud response time of a task after completing preprocessing into a feature vector to obtain a training dataset;
[0011] Step 2: Weighted processing of feature vectors
[0012] Step 2.1: Feature weight calculation
[0013] First, based on the adaptive attention mechanism, dynamically calculate the weights of task and environmental characteristics; assume that the feature vector of a task is:
[0014] x = [x1, x2,..., x i ,..., x n
[0015] where x i represents the normalized feature value of the i-th feature, and n represents the number of features of a task; the adaptive attention mechanism calculates the importance weights of each feature through a trainable attention network, and the specific calculation formula is as follows:
[0016]
[0017] where α i is the weight of the i-th feature, W i is the trainable parameter of the attention network, and exp represents the exponential function; at the same time, through normalization, the sum of all weights is made equal to 1;
[0018] Step 2.2: Multiply each eigenvalue by its corresponding normalized weight to obtain the weighted feature vector of this task;
[0019] Step Three: Task offloading model training
[0020] Step 3.1: Define the action set, that is, the possible offloading choices; the action set includes three task execution actions:
[0021] Action A1: Local execution, the task is processed locally on the device;
[0022] Action A2: Edge offloading, the task is offloaded to the edge node for execution;
[0023] Action A3: Cloud offloading, the task is uploaded to the cloud for processing;
[0024] Step 3.2: Task offloading model training
[0025] The task offloading model adopts the Actor-Critic framework, including two neural network models, namely the Actor module and the Critic module; the Actor module is a policy network, according to the input task state feature s, within the action set in Step 3.1, generates an offloading policy, and outputs the action probability distribution π(a|s; θ actor ), that is, the probability distribution of selecting action a under the given task state feature s, θ actor is the policy parameter of the Actor module and is a trainable parameter;
[0026] The task state feature s refers to the weighted feature vector of the computing requirements, data volume, latency requirements, network bandwidth, edge node load, and cloud response time of a task; the Critic module is a value function network, adopting the evaluation point based on the TD error. The Critic module calculates the TD error through a neural network and updates its network parameters by gradient, and at the same time, the Actor module updates its network parameters according to this TD error by gradient;
[0027] The specific process of task offloading model training is as follows:
[0028] Step 3.2.1: Set the maximum number of iterations T, action set A, decay factor γ, exploration rate ε, learning rate; initialize the network parameters of the Actor module as the network parameters of the Critic module as and all trainable parameters of the attention network;
[0029] Step 3.2.2: Process the feature vector of a task in the training dataset of Step 1 through the initialized attention network in Step 2, and obtain the weighted feature vector of this task according to the method in Step 2; then use the weighted feature vector of this task as the starting state feature x of this task 0 , input it into the initialized Actor module. The Actor module outputs the probability distribution of actions through forward propagation; adopt the ε-greedy strategy, and select the action A with the highest probability with a probability of 1 - ε * , and its probability value is denoted as π(A * |x 0 );
[0030] Step 3.2.3: Execute action A * , the network bandwidth and the edge node load value are updated. Update the feature data of this task according to the current network bandwidth and edge node load value, then standardize the updated feature data of this task as in Step 1, and then process it through the initialized attention network in Step 2. According to the method in Step 2, perform weighted processing of the feature vector to obtain the once-updated state feature x of this task 1 ; and obtain the actual delay D * of executing action A real , energy consumption E real , and whether the evaluation is completed C real , and calculate the reward value R:
[0031] R = -w1·D real - w2·E real + w3·C real
[0032] where D real is the actual completion time of the task, in milliseconds; E real is the total energy consumption of task execution, in joules; C real is whether the evaluation is completed, taking the value of 1 when the task is completed on time and 0 when it times out; w1, w2, and w3 are hyperparameters and are set values;
[0033] Step 3.2.4: Use the starting state feature x of this task 0 and the once-updated state feature x in Step 3.2.3 1 as the inputs of the Critic module respectively, and obtain the state values V(x 0 ), V(x 1 ) respectively, and then calculate the TD error δ:
[0034] δ = R + γV(x 1 ) - V(x 0 )
[0035] where R is the reward value in step 3.2.3, and γ is the attenuation factor;
[0036] Step 3.2.5: According to the TD error δ in step 3.2.4, update all the trainable parameters of the attention network in step two through the backpropagation algorithm;
[0037] Step 3.2.6: Take the square of the TD error δ of this task as the update gradient of the network parameters of the Critic module to obtain the network parameters after one update
[0038]
[0039] where α c is the learning rate of the Critic module;
[0040] Meanwhile, perform one update on the network parameters of the Actor module:
[0041]
[0042] where α a is the learning rate of the Actor module;
[0043] Complete the iterative training of one task of the task offloading model;
[0044] Step 3.2.7: Take the network parameter values of the task offloading model and the trainable parameter values of the attention network when one task's iterative training is completed as the initial parameter values for the iterative training of the next task, and continuously repeat the process of steps 3.2.2 - 3.2.7 until the weighted feature vectors of the last task in the training dataset of step two are trained, completing one round of training;
[0045] Step 3.2.8: Take the network parameter values of the task offloading model and the trainable parameter values of the attention network when the previous round of iterative training is completed as the initial parameter values for the next round of iterative training, and continuously repeat the process of step 3.2.7. When the maximum number of training rounds is reached or the average reward change in 10 consecutive rounds of training is less than 0.01 and the task delay < 200ms, the energy consumption < 50J, and the task completion rate > 95%, the training of the attention network and the task offloading model in step two is completed;
[0046] Step Four: Execute task offloading
[0047] Obtain the computing requirements, data volume, and latency requirements of the task to be offloaded; collect the current environmental characteristics, including network bandwidth, edge node load, and cloud response time, to obtain the characteristic data of the task to be offloaded; then standardize the characteristic data of the task to be offloaded according to Step 1, and then process it through the trained attention network in Step 2. According to the method in Step 2, perform weighted processing of the feature vectors to obtain the initial state characteristics of the task to be offloaded; then input the initial state characteristics into the Actor module of the trained task offloading model in Step 3. The Actor module outputs the probability distribution of actions, and selects the action with the highest probability value to execute task offloading, that is, completes task offloading;
[0048] Step 5: Dynamic Optimization of Task Offloading
[0049] After the task to be offloaded is completed, obtain the current values of the network bandwidth and edge node load, update the characteristic data of this task to be offloaded, and then standardize the updated characteristic data of this task to be offloaded as in Step 1, and then process it through the trained attention network in Step 2. According to the method in Step 2, perform weighted processing of the feature vectors to obtain the once-updated state characteristics of this task to be offloaded; and obtain the actual latency, energy consumption, and whether the evaluation is completed after the task to be offloaded is completed, and calculate the reward value R; then according to Step 3.2.4, calculate the TD error δ after the task to be offloaded is completed, and then according to Steps 3.2.5 and 3.2.6, perform an update on the trainable parameters of the attention network and the network parameters of the task offloading model in Step 2 to complete a dynamic update of the attention network and the task offloading model; then input the initial state characteristics of the next task to be offloaded into the Actor module of the task offloading model that has completed a dynamic update, and sequentially execute Step 4 and Step 5 to enable the attention network and the task offloading model to achieve dynamic updates; each time a task offloading is executed, the attention network and the task offloading model are dynamically updated once, that is, dynamic optimization of task offloading is achieved.
[0050] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0051] 1. Improve the adaptability of task offloading decisions. By introducing an adaptive attention mechanism, the present invention can dynamically calculate feature weights according to the characteristics of tasks and environments, thereby adjusting offloading decisions in real time. This mechanism can accurately identify the key features of tasks in different scenarios, improving the accuracy and adaptability of task offloading decisions, especially performing more excellently in complex and changing network environments.
[0052] 2. Optimize the multi-objective task offloading strategy. Combining the Actor-Critic framework of reinforcement learning, the present invention can flexibly adjust the offloading strategy according to task characteristics and real-time environmental states in multi-objective optimization (such as latency, energy consumption, and resource utilization). Through the dynamic learning ability of reinforcement learning, the balance between latency and energy consumption can be achieved, while improving resource utilization, and a better task offloading strategy can be realized in complex environments.
[0053] 3. Reduce computational overhead and improve real-time performance. The present invention adopts a scheme combining reinforcement learning and an adaptive attention mechanism. Utilizing the step-by-step optimization characteristics of reinforcement learning, the need for large-scale calculations is avoided, the computational overhead is significantly reduced, and the real-time performance is remarkably improved, enabling it to handle high-dynamic task offloading scenarios.
[0054] 4. Enhance the robustness during the task offloading process. Through the combination of reinforcement learning and an adaptive attention mechanism, the present invention can not only provide reasonable offloading decisions in the initial stage but also continuously adapt to environmental changes through dynamic learning and optimization, enhancing the robustness in dynamic and complex environments, thereby ensuring efficient operation under different task characteristics and network states. Brief Description of the Drawings
[0055] Figure 1 It is a flowchart of the steps of a task offloading optimization method based on adaptive attention and reinforcement learning according to the present invention. Detailed Embodiment
[0056] The technical solutions of the present invention will be described in detail below with reference to the drawings and embodiments.
[0057] The present invention provides a task offloading optimization method based on adaptive attention and reinforcement learning. Referring to Figure 1 , the method includes the following steps:
[0058] Step 1: Data collection and preprocessing
[0059] Collect the feature data of a certain number of tasks (i.e., offloading tasks), where the feature data includes task features and environmental features. Among them, task features include computational requirements, data volume, and latency requirements. Computational requirements use GFLOPS (floating-point operations per second) to describe the computational complexity of the task. For example, a video processing task requires 100 GFLOPS. The data volume represents the size of the data to be processed or transmitted by the task, with the unit of MB. For example, the input data size of a task is 200 MB. The latency requirement is the response time requirement of the task, with the unit of millisecond (ms). For example, an image recognition task needs to be completed within 300 ms.
[0060] Environmental characteristics describe the current system status, including network bandwidth, edge node load, and cloud response time. Network bandwidth reflects the available bandwidth between the edge node and the cloud, in Mbps, for example, the current network bandwidth is 50Mbps. Edge node load represents the real-time computing resource occupancy rate of the edge node, expressed as a percentage, for example, the current load of the edge node is 70%. Cloud response time refers to the total delay after the task is offloaded to the cloud, including transmission time and processing time, usually in milliseconds (ms), for example, 500ms. The characteristic data of some tasks finally collected are shown in Table 1.
[0061] Table 1 Characteristic data
[0062]
[0063] The collected feature data is cleaned and standardized to ensure data availability and consistency. First, invalid data is removed and outliers are corrected. For example, if the bandwidth value is negative, the task is directly removed; if the load data is missing, it is supplemented by interpolation based on historical data. Then feature standardization is performed. The specific calculation process is shown in the following formula:
[0064]
[0065] Where x is the value of a feature of a task, x max 、x min are the maximum and minimum values of the feature in all tasks respectively. The same is true for other features. The standardized data after preprocessing is shown in Table 2.
[0066] Table 2 Normalization results of feature data
[0067]
[0068] Finally, the task type ID is used as a label to combine the preprocessing computing requirements, data volume, latency requirements, bandwidth, node load, and cloud response time of a task into a feature vector to obtain the training data set. For example, for task T001, its feature vector is:
[0069] x T001 =[0.5,0.667,0.25,0.5,0.75,0.6]
[0070] Step 2: Weighted processing of feature vectors
[0071] In task offloading optimization, the importance of task and environmental characteristics for decision-making varies depending on the scenario. For example, latency-sensitive tasks may be more concerned about network bandwidth, while compute-intensive tasks are more concerned about the load of edge nodes. To accurately extract key features and adapt to dynamic environments, this step assigns weights to features through an adaptive attention mechanism to generate an optimized weighted feature vector. This process includes feature weight calculation and weighted feature vector generation, ensuring that the model input can accurately reflect task requirements and environmental states, thereby enhancing the decision-making ability of reinforcement learning.
[0072] Step 2.1: Feature weight calculation; First, based on the adaptive attention mechanism, the weights of task and environmental characteristics are dynamically calculated. Assume that the feature vector of a task is:
[0073] x = [x1, x2,..., x i ,..., x n
[0074] where x i represents the normalized feature value of the i-th feature, and n represents the number of features of a task. In this embodiment, n = 6. The adaptive attention mechanism calculates the importance weights of each feature through a trainable attention network. The specific calculation formula is as follows:
[0075]
[0076] where α i is the weight of the i-th feature, W i is the trainable parameter of the attention network, exp represents the exponential function, ensuring the positive nature of the weights, and at the same time making the sum of all weights equal to 1 through normalization. Through this mechanism, important features (such as task latency requirements or network bandwidth) will obtain higher weights, while the weights of secondary features are relatively low.
[0077] Step 2.2: Multiply each feature value x i by its corresponding normalized weight α' i to obtain the weighted feature vector x' of this task. The calculation formula of the weighted feature vector is:
[0078] x′ = [α′1·x1, α′2·x2,…, α′ i ·x i ,…, α′ n ·x n
[0079] The goal of this step is to amplify the features that have a greater impact on decision-making and weaken the secondary features, thereby optimizing the input data of the model. Taking an actual task as an example, assume that the task feature vector is:
[0080] x = [0.5, 0.667, 0.25, 0.5, 0.75, 0.6]
[0081] The weights calculated by the adaptive attention mechanism are as follows:
[0082] α = [0.2, 0.4, 0.1, 0.2, 0.05, 0.05]
[0083] Then the weighted feature vector is:
[0084] x' = [0.1, 0.267, 0.025, 0.1, 0.0375, 0.03]
[0085] The weighted feature vector can more accurately reflect the key characteristics of the task and the environment, improving the adaptability of the model to complex scenarios. For example, in latency-sensitive scenarios, the importance of network bandwidth (with a relatively high corresponding feature weight) for decision-making is highlighted, while the impact of secondary features (such as cloud response time) is weakened.
[0086] Step 3: Task offloading model training
[0087] Step 3.1: Define the action set, that is, the possible offloading options. The action set includes three task execution actions:
[0088] Action A1: Local execution, where the task is processed locally on the device. Generally applicable to scenarios with strict latency requirements and low local resource load. Specifically, when the latency requirement of the task is less than 200 ms and the CPU load of the device is below 40%, local execution is selected. When exceeding this threshold, the possibility of local execution decreases.
[0089] Action A2: Edge offloading, where the task is offloaded to the edge node for execution. Generally applicable to scenarios with high task computing requirements and sufficient network bandwidth. When the task computing requirement should be greater than 50 GFLOPS and the available bandwidth of the edge node is greater than 30 Mbps, edge offloading is selected. If the bandwidth is below this value or the computing requirement is low, the possibility of edge offloading will decrease.
[0090] Action A3: Cloud offloading, where the task is uploaded to the cloud for processing. Generally applicable to scenarios with large amounts of data and idle cloud resources. When the data volume is greater than 500 MB and the cloud load is below 50%, cloud offloading is selected. If the cloud load is too high or the network bandwidth is unstable, the possibility of cloud offloading decreases.
[0091] Step 3.2: Task offloading model training
[0092] The task offloading model adopts the Actor-Critic framework, which includes two neural network models, namely the Actor module and the Critic module. The Actor module is a policy network that generates an offloading policy within the action set in step 3.1 based on the input task state feature s, and outputs the action probability distribution π(a|s; θ actor ) of the current task, that is, the probability distribution of selecting action a given the task state feature s, where θ actor is the policy parameter of the Actor module and is a trainable parameter.
[0093] For example, the task selects edge offloading with a probability of 70%, local execution with a probability of 20%, and cloud offloading with a probability of 10%.
[0094] The task state feature s refers to a weighted feature vector of a task's computing requirements, data volume, latency requirements, network bandwidth, edge node load, and cloud response time. The action a refers to one of the three possible task offloading methods: edge offloading, local execution, or cloud offloading.
[0095] The Critic module is a value function network that uses an evaluation point based on the TD error. The Critic module calculates the TD error through a neural network and updates its network parameters by gradient, while the Actor module updates its network parameters according to this TD error by gradient.
[0096] The specific process of training the task offloading model is as follows:
[0097] Step 3.2.1: Set the maximum number of iterations T, the action set A (i.e., A1, A2, A3), the decay factor γ, the exploration rate ε, and the learning rate α c ; Initialize the network parameters of the Actor module as the network parameters of the Critic module as and all trainable parameters of the attention network.
[0098] Step 3.2.2: Process the feature vector of a task in the training dataset of step 1 through the initialized attention network in step 2, and obtain the weighted feature vector of this task according to the method in step 2; then use the weighted feature vector of this task as the starting state feature x 0 of this task, input it into the initialized Actor module, and the Actor module outputs the probability distribution of actions through forward propagation. For example:
[0099] π(α|x 0 ) = [0.2, 0.7, 0.1]
[0100] It represents that the probability of local execution is 20%, the probability of edge offloading is 70%, and the probability of cloud offloading is 10%.
[0101] Adopt the ε-greedy strategy, and select the action A with the highest probability with a probability of 1-ε * , and its probability value is denoted as π(A * |x 0 ).
[0102] Step 3.2.3: Execute action A * , the network bandwidth and the edge node load value are updated. According to the current network bandwidth and edge node load value, the characteristic data of this task (computing requirements, data volume, latency requirements, network bandwidth, edge node load, and cloud response time) is updated, and then the updated characteristic data of this task is standardized as in Step 1, and then processed by the initialized attention network in Step 2. According to the method in Step 2, weighted processing of the feature vectors is performed to obtain the once-updated state feature x of this task 1 ; and obtain the actual latency D * of executing action A real , energy consumption E real , and whether the evaluation is completed C real , and calculate the reward value R:
[0103] R = -w1·D real - w2·E real + w3·C real
[0104] where D real is the actual completion time (ms) of the task, and the smaller the value, the higher the reward. E real is the total energy consumption (J) of task execution, and the smaller the value, the higher the reward. C real is whether the evaluation is completed. When the task is completed on time, the value is 1, and when it times out, the value is 0. A higher reward is given for on-time completion. w1, w2, and w3 are hyperparameters and are adjusted according to the specific scenario. For example, in a latency-sensitive scenario (e.g., for tasks with a latency requirement less than 200 ms, a higher weight will be set for the latency part, and (w1 = 0.6, w2 = 0.3, w3 = 0.1) will be set); in an energy-consumption-sensitive scenario (e.g., when the energy consumption during task execution exceeds a certain set threshold, the weight of energy consumption will increase, thus encouraging the system to choose a low-energy offloading method, then (w1 = 0.3, w2 = 0.6, w3 = 0.1) will be set).
[0105] Step 3.2.4: Use the initial state feature x 0 of this task and the once-updated state feature x 1 in Step 3.2.3 as the inputs of the Critic module respectively, and obtain the state value V(x0 )、V(x 1 ), then calculate the TD error (Temporal Difference Error) δ:
[0106] δ = R + γV(x 1 ) - V(x 0 )
[0107] where R is the reward value in step 3.2.3, and γ is the attenuation factor used to measure the impact of future rewards;
[0108] Step 3.2.5: According to the TD error δ in step 3.2.4, update all the trainable parameters of the attention network in step two once through the backpropagation algorithm;
[0109] Step 3.2.6: Take the square of the TD error δ of this task as the update gradient of the network parameters of the Critic module to obtain the network parameters after one update
[0110]
[0111] where α c is the learning rate of the Critic module;
[0112] Meanwhile, perform one update on the network parameters of the Actor module:
[0113]
[0114] where α a is the learning rate of the Actor module.
[0115] Complete the iterative training of one task of the task offloading model.
[0116] Step 3.2.7: Take the network parameter values of the task offloading model and the trainable parameter values of the attention network when the iterative training of one task is completed as the initial parameter values for the iterative training of the next task, and continuously repeat the process of steps 3.2.2 - 3.2.7 until the weighted feature vectors of the last task in the training dataset in step two are trained, thus completing one round of training.
[0117] Step 3.2.8: Use the network parameter values of the task offloading model and the trainable parameter values of the attention network at the end of the previous round of iterative training as the initial parameter values for the next round of iterative training. Continuously repeat the process of Step 3.2.7. When the maximum number of training rounds is reached or the average reward (the mean value of R after all tasks are executed for offloading) of 10 consecutive rounds of training changes by less than 0.01 and the task latency < 200 ms, the energy consumption < 50 J, and the task completion rate > 95%, the training of the attention network and the task offloading model in Step 2 is completed.
[0118] Step Four: Execute task offloading
[0119] Obtain the computing requirements (GFLOPS), data volume (MB), and latency requirements (ms) of the task to be offloaded. Collect the current environmental characteristics, including network bandwidth (Mbps), edge node load (%), and cloud response time (ms), to obtain the feature data of the task to be offloaded. Then, standardize the feature data of the task to be offloaded according to Step 1, and then process it through the trained attention network in Step 2. According to the method in Step 2, perform weighted processing of the feature vectors to obtain the starting state features of the task to be offloaded. Then, input the starting state features into the Actor module of the trained task offloading model in Step 3. The Actor module outputs the probability distribution of the actions, and select the action with the highest probability value to execute task offloading, that is, complete task offloading.
[0120] Step Five: Dynamic optimization of task offloading
[0121] After the task to be unloaded is completed, the current values of the network bandwidth and edge node load value are obtained, and the characteristic data of the task to be unloaded (computing requirements, data volume, delay requirements, network bandwidth, edge node load and cloud response time) are updated. Then, the updated characteristic data of the task to be unloaded is normalized as in step 1, and then processed by the trained attention network in step 2. According to the method in step 2, the characteristic vector is weighted to obtain the updated state characteristics of the task to be unloaded; and the actual delay, energy consumption, and completion evaluation after the task to be unloaded are obtained, and the reward value R is calculated. Then, according to step 3.2.4, the TD error δ after the task to be unloaded is completed is calculated, and then according to steps 3.2.5 and 3.2.6, the trainable parameters of the attention network in step 2 and the network parameters of the task unloading model are updated once, completing a dynamic update of the attention network and the task unloading model. Then, the starting state features of the next task to be offloaded are input into the Actor module of the task offloading model that has completed a dynamic update, and steps 4 and 5 are executed in sequence to enable the attention network and the task offloading model to be dynamically updated. Each time a task offloading is performed, the attention network and the task offloading model are dynamically updated, thereby achieving dynamic optimization of task offloading.
[0122] Any matters not described in the present invention are applicable to the prior art.
Claims
1. A task offloading optimization method based on adaptive attention and reinforcement learning, characterized in that: The method comprises the following steps: Step 1: Data collection and preprocessing Collect feature data of a certain number of tasks, wherein the feature data includes task features and environment features; wherein task features include computing requirements, data volume, and delay requirements; and environment features include network bandwidth, edge node load, and cloud response time; perform preprocessing of the collected feature data by cleaning and standardization; use the task type ID as a label, and combine the computing requirements, data volume, delay requirements, bandwidth, node load, and cloud response time of completing the preprocessing of a task into a feature vector to obtain a training data set; Step 2: Weighted processing of feature vectors Step 2.1: Feature weight calculation First, based on the adaptive attention mechanism, the weights of task and environment features are dynamically calculated; assuming that the feature vector of a task is: x=[x1,x2,...,x i ,...,x n ] Among them, x i represents the normalized feature value of the i-th feature, and n represents the number of features of a task. The adaptive attention mechanism calculates the importance weight of each feature through a trainable attention network. The specific calculation formula is as follows: where α i is the weight of the i-th feature, W i is the trainable parameter of the attention network, exp represents the exponential function; at the same time, the sum of all weights is normalized to 1; Step 2.2: Multiply each eigenvalue by its corresponding normalized weight to obtain the weighted eigenvector of the task; Step 3: Task offloading model training Step 3.1: Define an action set, i.e., possible offloading options; the action set includes three task execution actions: Action A1: Local execution, the task is processed locally on the device; Action A2: Edge offloading, the task is offloaded to the edge node for execution; Action A3: Cloud unloading, the task is uploaded to the cloud for processing; Step 3.2: Task offloading model training The task offloading model adopts the Actor-Critic framework, which includes two neural network models, namely the Actor module and the Critic module. The Actor module is a policy network, which generates an offloading strategy based on the input task state feature s in the action set in step 3.1 and outputs the action probability distribution π(a|s; θ actor ), that is, the probability distribution of selecting action a given task state feature s, θ actor It is the strategy parameter of the Actor module, which is a trainable parameter; The task state feature s refers to the weighted feature vector of the computing demand, data volume, delay requirement, network bandwidth, edge node load and cloud response time of a task; the Critic module is a value function network, which uses an evaluation point based on TD error. The Critic module calculates the TD error through a neural network and updates its network parameters by gradient, while the Actor module updates its network parameters by gradient according to the TD error; The specific process of task offloading model training is as follows: Step 3.2.1: Set the maximum number of iterations T, action set A, decay factor γ, exploration rate ε, and learning rate; initialize the network parameters of the Actor module to The network parameters of the Critic module are and all trainable parameters of the attention network; Step 3.2.2: Process the feature vector of a task in the training data set of step 1 through the attention network initialized in step 2, and obtain the weighted feature vector of the task according to the method in step 2; then use the weighted feature vector of the task as the starting state feature x of the task 0 , input it into the initialized Actor module, the Actor module outputs the probability distribution of the action through forward propagation; the ε-greedy strategy is adopted to select the action A with the highest probability with a probability of 1-ε * , whose probability value is denoted as π(A * |x 0 ); Step 3.2.3: Execute Action A * , the network bandwidth and edge node load values are updated, and the feature data of the task is updated according to the current network bandwidth and edge node load values. Then, the updated feature data of the task is normalized as in step 1, and then processed by the initialized attention network in step 2. According to the method in step 2, the feature vector is weighted to obtain the updated state feature x of the task. 1 ; and get the execution action A * The actual delay D real , Energy consumption E real 、Whether the evaluation is completed C real , calculate the reward value R: R=-w1·D real -w2·E real +w3·C real Among them, D real is the actual completion time of the task, in ms; E real is the total energy consumption of task execution, in J; C real is the evaluation of whether the task is completed. When the task is completed on time, the value is 1, and when it is overtime, the value is 0; w1, w2, and w3 are hyperparameters, which are set values; Step 3.2.4: Set the starting state feature x of the task 0 and the updated state feature x in step 3.2.3 1 As the input of the Critic module, we can get the state value V(x 0 )、V(x 1 ), and then calculate the TD error δ: δ=R+γV(x 1 )-V(x 0 ) Where R is the reward value in step 3.2.3, and γ is the decay factor; Step 3.2.5: According to the TD error δ in step 3.2.4, update all trainable parameters of the attention network in step 2 through the back propagation algorithm; Step 3.2.6: The square of the TD error δ of the task is used as the update gradient of the network parameters of the Critic module Get the updated network parameters Among them, α c is the learning rate of the Critic module; At the same time, update the network parameters of the Actor module: Among them, α a is the learning rate of the Actor module; Complete iterative training of a task for the task offloading model; Step 3.2.7: Use the network parameter values of the task offloading model and the trainable parameter values of the attention network when the iterative training of one task is completed as the initial parameter values for the iterative training of the next task, and repeat the process of steps 3.2.2 to 3.2.7 until the weighted feature vector of the last task in the training data set of step 2 is trained, completing one round of training; Step 3.2.8: Use the network parameter values of the task offloading model and the trainable parameter values of the attention network at the completion of the previous round of iterative training as the initial parameter values for the next round of iterative training, and repeat the process of step 3.2.
7. When the maximum number of training rounds is reached or the average reward change for 10 consecutive rounds of training is less than 0.01 and the task delay is less than 200ms, the energy consumption is less than 50J, and the task completion rate is greater than 95%, the training of the attention network and task offloading model in step 2 is completed; Step 4: Execute task uninstallation Obtain the computing requirements, data volume, and latency requirements of the task to be offloaded; collect current environmental characteristics, including network bandwidth, edge node load, and cloud response time, to obtain feature data of the task to be offloaded; then standardize the feature data of the task to be offloaded according to step one, and then process it through the trained attention network in step two, and perform weighted processing of the feature vector according to the method in step two to obtain the initial state characteristics of the task to be offloaded; then input the initial state characteristics into the Actor module of the task offloading model trained in step three, and the Actor module outputs the probability distribution of the action, and selects the action with the highest probability value to perform task offloading, that is, complete the task offloading; Step 5: Dynamic optimization of task offloading After the task to be unloaded is completed, the current values of the network bandwidth and the edge node load value are obtained, and the feature data of the task to be unloaded is updated. Then, the updated feature data of the task to be unloaded is normalized as in step one, and then processed by the trained attention network in step two. According to the method in step two, the feature vector is weighted to obtain an updated state feature of the task to be unloaded; and the actual delay, energy consumption, and completion evaluation after the task to be unloaded are obtained, and the reward value R is calculated; then according to step 3.2.4, the TD error δ after the task to be unloaded is completed is calculated, and then according to steps 3.2.5 and 3.2.6, the trainable parameters of the attention network in step two and the network parameters of the task unloading model are updated once, and a dynamic update of the attention network and the task unloading model is completed; then the starting state features of the next task to be unloaded are input into the Actor module of the task unloading model that has completed a dynamic update, and steps four and five are executed in sequence to realize dynamic update of the attention network and the task unloading model; each time a task unloading is performed, the attention network and the task unloading model are dynamically updated once, that is, dynamic optimization of task unloading is realized.
2. The task offloading optimization method based on adaptive attention and reinforcement learning according to claim 1, characterized in that: The computing requirement uses GFLOPS to describe the computational complexity of the task; the data volume represents the size of the data that needs to be processed or transmitted by the task, in MB; the delay requirement is the response time requirement of the task, in milliseconds.
3. The task offloading optimization method based on adaptive attention and reinforcement learning according to claim 1, characterized in that: The unit of network bandwidth is Mbps; the edge node load refers to the real-time computing resource occupancy rate of the edge node, expressed in percentage; the cloud response time refers to the total delay after the task is offloaded to the cloud, including transmission time and processing time, expressed in milliseconds (ms).
4. The task offloading optimization method based on adaptive attention and reinforcement learning according to claim 1, characterized in that: In step one, cleaning means removing invalid data and correcting outliers.
5. The task offloading optimization method based on adaptive attention and reinforcement learning according to claim 1, characterized in that: In step 1, the specific calculation process of feature standardization is shown in the following formula: Where x is the value of a feature of a task, x max 、x min are the maximum and minimum values of this feature in all tasks respectively.
6. The task offloading optimization method based on adaptive attention and reinforcement learning according to claim 1, characterized in that: In step 2, n=6.
7. The task offloading optimization method based on adaptive attention and reinforcement learning according to claim 1, characterized in that: In step 3, when the delay requirement is less than 200ms, set w1=0.6, w2=0.3, and w3=0.
1.
8. The task offloading optimization method based on adaptive attention and reinforcement learning according to claim 1, characterized in that: In step 3, when the energy consumption during task execution exceeds a certain set threshold, w1=0.3, w2=0.6, and w3=0.1 are set.
Citation Information
Patent Citations
Task unloading method and device for mobile edge computing system
CN114489977A