Task unloading method and system using deep reinforcement learning and attention mechanism

CN120029694AActive Publication Date: 2025-05-23QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Patent Information

Application Number
CN202510511736.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-05-23
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

In an edge computing environment, how to flexibly schedule heterogeneous resources, maintain the flexibility and adaptability of task offloading strategies, and find a balance between performance and energy consumption to achieve economic and environmental sustainability of task offloading.

Method used

The task unloading method with deep reinforcement learning and attention mechanism is adopted, and the task unloading strategy based on directed acyclic graph (DAG) and the asymptotic rectangular window attention mechanism are formed to optimize task unloading in the edge computing environment.

Benefits of technology

It improves the comprehensive performance of task offloading, enhances data security and privacy protection, improves system resource utilization, achieves relatively efficient and energy-saving offloading, and has good generalization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029694A_ABST
    Figure CN120029694A_ABST
Patent Text Reader

Abstract

The invention relates to a task unloading method and system using deep reinforcement learning and an attention mechanism, and belongs to the technical field of task unloading. Comprising the steps of data acquisition and preprocessing; the data comprises edge server information, user information and task information; and inputting the preprocessed data into the trained task unloading model, and realizing task unloading based on a DPAQN algorithm. The DPAQN algorithm has obvious advantages in the aspect of optimizing the comprehensive performance of task unloading, and the DPAQN algorithm is about 20.71%-30.39% superior to an existing algorithm on average.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a task offloading method and system utilizing deep reinforcement learning and attention mechanism, and belongs to the technical field of task offloading. Background Art

[0002] In recent years, edge computing has rapidly emerged as an emerging computing paradigm. Its core advantage is to migrate data processing and computing tasks to the edge of the network, that is, close to the terminal device. This layout significantly reduces the latency and energy consumption during task computing, which is crucial for applications that require instant response and helps to efficiently process resource-intensive and time-sensitive tasks. At the same time, edge computing can also reduce the burden on data centers, reduce network congestion, and enhance data security and privacy protection by processing data locally. This architecture not only speeds up system response, but also provides a solid foundation for building smarter and more efficient Internet of Things (IoT) applications. Therefore, edge computing shows great potential in task offloading and is expected to become an important development direction of future computing technology.

[0003] With the rapid development of edge computing technology, computing offloading is becoming a key technology for edge computing. Offloading the computing tasks of IoT terminals to the edge environment can effectively solve the deficiencies in computing performance and energy efficiency. However, while offloading tasks to the edge environment brings convenience, it also brings a series of challenges:

[0004] 1) How to flexibly schedule the heterogeneous resources of edge servers to adapt to different tasks, including differences in computing, processing speed, and offloading requirements, to ensure that the task matches the best server.

[0005] 2) How to keep the task offloading strategy flexible and adaptable in the face of network fluctuations and changes in user demand to maintain system stability and performance.

[0006] 3) How to find a balance between improving performance and reducing energy consumption to achieve economic and environmental sustainability of task offloading.

[0007] In response to the above challenges, researchers have proposed a variety of solutions, which are mainly divided into traditional methods and artificial intelligence (AI) methods.

[0008] Traditional methods include mathematical optimization, Lyapunov optimization, and heuristic algorithms, which have their own advantages in ensuring optimality and stability, but also have problems of scalability, complexity, and suboptimal solutions. Peng et al. proposed an intelligent computing offloading and resource allocation method (EECCT) based on end-edge-cloud collaboration to optimize computing offloading and resource allocation in the Industrial Internet of Things (IIoT). Nguyen et al. studied a fast and efficient algorithm (FEA) to maximize the total computing rate of wireless powered mobile edge computing (WPMEC) in IoT networks by jointly optimizing the transmission power, backscattering coefficient, time division ratio, and computing mode decision of the OFDMA system. Wu et al. proposed a nonlinear exponential inertia weighted particle swarm optimization algorithm (PSO) for solving the edge-cloud collaborative multi-task computing offloading model. The algorithm can make up for the defect of premature convergence of the standard PSO by dynamically adjusting the inertia weight and effectively avoid falling into the local optimal solution. Zhang et al. proposed a new task scheduling scheme (GA-EC) based on genetic algorithm (GA) for the efficient scheduling of resource-intensive tasks in edge computing environments. The advantage of this solution is that it can achieve multi-objective optimization and improve the robustness and adaptability of the system through dynamic scheduling strategies.

[0009] Traditional methods are superior in fast deployment and low-complexity scenarios due to their simplicity, but they are not adapted to the dynamics of multi-user and resource-constrained environments. In contrast, AI-based methods provide real-time decision support in resource-constrained and complex network environments through good adaptability and predictive capabilities, effectively process high-dimensional data and learn feature representations, and optimize task offloading. However, AI methods also face challenges in edge computing, mainly including the limited computing and storage resources of edge devices that limit model complexity, as well as the high requirements for algorithm robustness and real-time performance due to network instability and frequent device changes.

[0010] Artificial intelligence-based methods, such as Q-learning, DRL, DQN, SVM, and ANN, have developed rapidly since 2019. They can dynamically evaluate network conditions and user behaviors to achieve intelligent task allocation and resource optimization. Li et al. proposed a Q-learning-based and DRL-based model, respectively, taking the sum of all delays and energy consumption as the total optimization target, effectively reducing the cost. Xiao et al. designed a DRL-based mobile resource offloading scheduling for edge cloud environments to reduce delays and energy consumption, in which the actor-network is designed to select scheduling strategies and the critic-network updates the actor-network parameters to improve computing performance. Wang et al. proposed an algorithm based on meta-reinforcement learning (MRL) to solve the task offloading problem, which can quickly adapt to new environments with small samples and effectively reduce task delays. Liu et al. proposed a multi-workflow scheduling method based on DQN, which can cope with the situation where edge service performance fluctuates over time. Sheng et al. proposed a task scheduling method for IoT edge computing based on deep reinforcement learning, which was implemented by modeling the problem as an MDP and applying a policy-based deep reinforcement learning algorithm. The method performed well on multiple evaluation indicators. Heidari et al. combined double Q-learning and deep post-decision state (PDS) learning to propose a new learning method called D2QDPO3 to deal with the problem of dynamic computing offloading in IoT edge computing.

[0011] Deep learning automatically extracts features, reinforcement learning learns optimal strategies through interaction, and SVM and neural networks process high-dimensional data to optimize task offloading decisions. However, the complexity of models that can be deployed by AI-based methods is often limited by the limited computing and storage resources of edge devices. Summary of the invention

[0012] In view of the shortcomings of the prior art, the present invention proposes a task offloading method using deep reinforcement learning and attention mechanism; The present invention also proposes a task offloading system using deep reinforcement learning and attention mechanism.

[0013] In order to solve the task offloading problem more practically, in the present invention: 1) Considering the heterogeneity of edge servers and user diversity, a task offloading strategy based on directed acyclic graph (DAG) is proposed to optimize the task offloading problem in edge computing environment.

[0014] 2) This paper first applies the asymptotic rectangular window attention mechanism to edge computing task offloading, which can improve decision quality. Combining it with deep reinforcement learning to form the DPAQN algorithm is expected to make new contributions to optimizing the task offloading problem in edge computing environments.

[0015] 3) The DPAQN algorithm proposed in this paper comprehensively considers the three indicators of total task completion time, energy consumption and throughput, and is committed to studying a more efficient and energy-saving offloading algorithm. By comparing with five algorithms, it is proved that DPAQN improves the comprehensive performance of task offloading.

[0016] The technical solution of the present invention is: Task offloading methods using deep reinforcement learning and attention mechanisms, including: Data acquisition and preprocessing; data includes edge server information, user information and task information; The preprocessed data is input into the trained task offloading model, and task offloading is realized based on the DPAQN algorithm.

[0017] Further preferably, the pretreatment comprises: Noise addition; Dynamic environment simulation: By continuously generating new tasks, a dynamic edge computing environment is simulated to ensure that the task flow is continuous; Parameter settings.

[0018] According to the preferred embodiment of the present invention, the objective function in the task offloading model is It is expressed as: ; in, ; Represents the maximum completion time of all completed tasks. Represents the number of completed tasks; Represents the total completion time of the task; ; Represents the maximum completion time of all completed tasks. represents the total energy consumption of the task; w1,w2,w3 are weights, It refers to the total throughput.

[0019] Further preferably, the values ​​of w1, w2, and w3 are 0.25, 0.5, and 0.25 respectively.

[0020] Preferably, according to the present invention, the total task completion time TT includes the transmission time and the running time of all tasks; it is specifically defined as: ; Where n is the number of users, l is the number of subtasks for each user, represents the data volume of the i-th task of user device k, represents the bandwidth allocated to the i-th task of user equipment k, is the transmission time of the i-th task of user device k, represents the time required for the i-th task of user device k to execute locally, represents the ratio of local and remote execution rates, is the running time of the i-th task of user device k; The total energy consumption of a task includes the communication energy consumption and operation energy consumption of all tasks; it is specifically defined as: ; in, is the transmission power of the ith task of user equipment k, represents the communication time of the i-th task of user device k; represents the communication energy consumption of the i-th task of user device k, that is, ; represents the running energy consumption of the i-th task of user device k.

[0021] According to the preferred embodiment of the present invention, in the edge computing environment, task offloading refers to the process of migrating computing tasks from user devices to edge servers or edge nodes for execution. The task offloading decision is regarded as an MDP, represented by a five-tuple To express; Among them, s is the state space, which is used to describe the state of the system environment; according to the current state, a task is selected from all tasks waiting to be assigned, and after the task is unloaded to the edge server, the next state is obtained; a is the action space, including all actions selected in the current state; r is the immediate reward, which is used to evaluate the feedback of the action; p is the state transition probability, which describes the probability of transitioning from one state to another; It is a discount factor that measures the importance of immediate rewards and long-term rewards. It belongs to the interval [0,1], where 0 means that only immediate rewards are considered, and 1 means that long-term rewards and immediate rewards are considered equally important.

[0022] Further preferably, the calculation formula of r is as follows: ; in, is the transmission time of the ith task of user device k, is the running time of the i-th task of user device k, is the communication energy consumption of the ith task of user device k, ) is the running energy consumption of the i-th task of user device k.

[0023] Preferably, according to the present invention, the task offloading model includes an action network and a target network; The action network and the target network have the same structure; In the action network, it is used to evaluate the Q value of each action in the current state and select the optimal action. Its parameters are continuously updated through training; In the target network, it is used to provide a stable training target. The parameters of the target network are copied from the action network Q network every C steps.

[0024] Preferably, according to the present invention, the action network includes several neural network layers, specifically including: an input layer, a first hidden layer, a second hidden layer, an asymptotic rectangular window attention mechanism, and an output layer; The input layer of the action network receives the feature vector, and after the linear transformation and relu activation function processing of the first hidden layer, the output is passed to the second hidden layer; after the second hidden layer, the asymptotic rectangular window attention mechanism is introduced, which divides the output of the second hidden layer into multiple windows and performs weighted summation on the features in each window; specifically: 1) First, the features in each rectangular window, that is, the features processed by the linear transformation of the first hidden layer and the relu activation function, are transformed into feature representations through linear transformation; then, nonlinear processing is performed using the tanh activation function; then, the dot product similarity between the feature representation and the learnable vector is calculated and normalized by the Softmax function to obtain the attention weight in each window; finally, the weighted features of all windows are concatenated to form the output of the asymptotic rectangular window attention mechanism; 2) The output of the asymptotic rectangular window attention mechanism passes through the first hidden layer, the second hidden layer and the asymptotic rectangular window attention mechanism again to further extract and abstract features; The predicted Q value of each action is output through the last linear transformation plus the relu activation function. The predicted Q value represents the expected return of taking each possible action in the current state.

[0025] Preferably, according to the present invention, task offloading is implemented based on the DPAQN algorithm, including: First, initialize the edge computing environment, DQN parameters, and state s; Using the current state s as input, the action network evaluates the predicted Q-values ​​of all possible actions, i.e. ; and select action a with the highest Q value to execute; Execute action a to interact with the environment, observe the reward r, and obtain the next state s' and the iteration end flag done; Then, the current state s, action a, reward r, next state s' and iteration end flag done are stored in the experience replay buffer for subsequent learning; a small batch of data is randomly sampled from the experience replay for training; Take the next state s' as the input of the target network Q' and calculate the future ; Combine the reward r and the iteration end flag done to calculate , as shown below; ; in, It is the parameter of the target network, and the value of done is 0 or 1; The action network receives the state s and action a in the mini-batch sampling, and the Q value of the action network is combined with the action a in the mini-batch sampling to obtain ; Then use the Mean Squared Error (MSE) loss function to calculate and Loss , as shown below; ; m represents the number of samples in the mini-batch sampling; Update the parameters of the action network through back propagation according to the obtained Loss, and copy the parameters of the action network Q to the target network Q' every C steps; Determine whether the termination condition is met. If so, the process ends; if not, the steps after returning to the initial state continue.

[0026] A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of a task offloading method using deep reinforcement learning and an attention mechanism are implemented.

[0027] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a task offloading method using deep reinforcement learning and an attention mechanism.

[0028] Task offloading system using deep reinforcement learning and attention mechanism, including: The data acquisition and preprocessing module is configured to: data acquisition and preprocessing; The task offloading module is configured to input the preprocessed data into the trained task offloading model and realize task offloading based on the DPAQN algorithm.

[0029] The beneficial effects of the present invention are: 1. Enhanced data security and privacy protection: By performing task offloading on the edge server, the present invention can process data close to the data source, reducing the amount of data transmission in the network, thereby enhancing data security and privacy protection.

[0030] 2. Comprehensive optimization of task offloading performance: The DPAQN algorithm proposed in this paper comprehensively considers the three key indicators of total task completion time, energy consumption and throughput, and significantly improves the quality of task offloading decisions by introducing the asymptotic rectangular window attention mechanism in DQN. Experimental results show that the DPAQN algorithm has obvious advantages in optimizing the comprehensive performance of task offloading, and is on average about 20.71% to 30.39% better than existing algorithms (such as DQN, Double DQN, Dueling DQN, Prioritized Replay and PPO).

[0031] 3. Improve system resource utilization: This invention takes into account the heterogeneity of edge servers and user diversity, and through a task offloading strategy based on a directed acyclic graph (DAG), rationally distributes tasks to different edge servers for execution, thereby achieving optimal resource allocation. This not only improves the resource utilization of edge servers, but also reduces the burden on data centers and reduces network congestion.

[0032] 4. Improve the generalization ability of the algorithm: By conducting comparative experiments under different random seeds and multiple user number configurations, the good generalization of the DPAQN algorithm was verified. The experimental results show that DPAQN can maintain stable performance in different scenarios, which can alleviate the common problems of weak robustness and poor generalization of artificial intelligence algorithms to a certain extent, and is suitable for a variety of practical application scenarios.

[0033] 5. Assist the intelligent development of industrial Internet of Things: This invention has significant application value in the industrial Internet of Things environment. Intelligent manufacturing systems need to process a large amount of sensor and equipment data, and have high requirements for energy consumption, task completion time and throughput. The DPAQN algorithm aims to optimize the comprehensive performance of task offloading and dynamically adjust task allocation, which is expected to provide support for efficient and intelligent production processes. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 A schematic diagram of the overall architecture of task offloading in edge computing environment; Figure 2 This is a flowchart of task offloading based on DQN; Figure 3 It is a schematic diagram of the action network structure; Figure 4 Schematic diagram of TEH of five different algorithms under different numbers of users Figure 1 ; Figure 5 Schematic diagram of TEH of five different algorithms under different numbers of users Figure 2 . DETAILED DESCRIPTION

[0035] The present invention will be further defined below in conjunction with the accompanying drawings and embodiments, but is not limited thereto.

[0036] Example 1 Task offloading methods using deep reinforcement learning and attention mechanisms, including: Data acquisition and preprocessing; data includes edge server information, user information and task information; Edge server information, including the following: Basic information: ID, number of users, basic rate ratio of local and remote execution, user list object, total bandwidth, system status refresh frequency, uninstall frequency, bandwidth allocation method; Real-time information: list of tasks being executed, list of tasks being transmitted, remote waiting list, user priority list, occupied bandwidth, system running time, current system CPU usage, last system action, last system score; Log information: uninstalled task records, total transmission time, total transmission energy consumption, total execution time, total number of uninstalled completed tasks, number of failed tasks, and total rewards.

[0037] User information, including the following: Basic information: ID, user task list, task transmission data volume, task transmission time distribution, task local execution time, task CPU occupancy rate, initial task number, task transmission and local execution energy consumption, task deadline, user priority; Log information: throughput, energy consumption, total communication time.

[0038] Task information, including the following: Basic information: user ID, task ID, number of task instructions, size of transmitted data, upload time, download time, local processing time, CPU resource requirements, transmission energy consumption, local execution energy consumption, task type, and task status; Real-time information: remaining remote execution time of the task, remaining transmission time of the task, bandwidth allocated to the task, task start time, task completion time, task start unloading time, task execution time, and task transmission time.

[0039] The preprocessed data is input into the trained task offloading model, and task offloading is realized based on the DPAQN algorithm.

[0040] Example 2 The task offloading method using deep reinforcement learning and attention mechanism described in Example 1 is different in that: Pre-processing, including: Noise addition: To simulate the randomness and uncertainty in a real environment, random noise (ξ) is added to the computation time of the task, ranging from 1 to 5. This noise simulates the computation time fluctuations that may exist in a real environment.

[0041] Dynamic environment simulation: By continuously generating new tasks (when a task is completed, a new task is immediately generated), a dynamic edge computing environment is simulated to ensure that the task flow is continuous; Parameter setting. According to the requirements of the simulation environment, some key parameters are set, such as: network bandwidth (50Mbps), number of user devices (ranging from 20 to 90), minimum computing time unit of the task (25 milliseconds) and minimum data volume unit (6 milliseconds).

[0042] Objective function in the task offloading model It is expressed as: ; in, ; Represents the maximum completion time of all completed tasks. Represents the number of completed tasks; Represents the total completion time of the task; ; Represents the maximum completion time of all completed tasks. represents the total energy consumption of the task; w1,w2,w3 are weights, It refers to the total throughput.

[0043] is a comprehensive optimization objective used to measure the performance of the task offloading strategy in the edge computing environment. The total completion time of the task ( )、Total energy consumption( ) and total throughput ( ) into a single objective function. Among them, w1, w2, and w3 are weight coefficients, which respectively represent the importance of different indicators.

[0044] The values ​​of w1, w2, and w3 are 0.25, 0.5, and 0.25 respectively. The purpose is to optimize the energy consumption of task offloading, while optimizing the total completion time and throughput of the task offloading, and strive to build an energy-saving and efficient offloading model. Among them, the smaller the index value TEH, the better the overall performance.

[0045] In order to transform the multi-objective optimization problem into a single-objective optimization problem, the weighted sum method (WSM) is introduced, that is, the values ​​of the three objective functions are normalized and a set of weights are used to sum them. , , to form a new overall optimization objective. , , The range is , and the sum of the three is 1, which reflects the importance of different goals.

[0046] According to the introduction of the above overall model, the DAG-based tasks to be processed by users are offloaded to the edge server, with the goal of comprehensively optimizing the three indicators of task completion time, energy consumption and throughput. The optimization problem is described below.

[0047] The total task completion time TT includes the transmission time and running time of all tasks; it is specifically defined as: ; Where n is the number of users, l is the number of subtasks for each user, represents the data volume of the i-th task of user device k, represents the bandwidth allocated to the i-th task of user device k, is the transmission time of the i-th task of user device k, represents the time required for the i-th task of user device k to execute locally, Represents the ratio of local and remote execution rates, is the running time of the i-th task of user device k; The total energy consumption of a task includes the communication energy consumption and operation energy consumption of all tasks; it is specifically defined as: ; in, is the transmission power of the ith task of user equipment k, which depends on the transmission distance and the power control strategy of the device. represents the communication time of the i-th task of user device k; represents the communication energy consumption of the i-th task of user device k, that is, ; Represents the running energy consumption of the i-th task of user device k. It is pre-set according to user settings and task characteristics.

[0048] Throughput refers to the amount of data successfully transmitted in a certain period of time. For each user, the sum of the data volume of all its tasks and the total time of the simulation run are recorded. At the end of the simulation, the total data volume of each user is calculated and divided by the simulation time to obtain the average throughput. The present invention defines the total throughput as .

[0049] This paper is committed to giving full play to the advantages of artificial intelligence algorithms, while taking into account the limited resources of edge devices and the instability of the network, and designs a DPAQN algorithm that combines deep reinforcement learning and asymptotic rectangular window attention mechanism. The algorithm comprehensively considers the actual situation of task offloading in the edge computing environment, and considers the three optimization goals of latency, energy consumption and throughput, striving to make the model as stable and efficient as possible. The final experiment proves its effectiveness.

[0050] In an edge computing environment, task offloading is a key optimization problem, which involves allocating computing tasks between terminal devices and edge servers. The design goal of the present invention is how to perform energy-efficient task offloading in an edge computing environment while taking into account the heterogeneity of edge servers and user diversity. The heterogeneity of edge servers is reflected in the differences in computing resources, processing speed, and task offloading requirements of edge servers. Since different edge servers usually have different processing capabilities, storage capacities, and network bandwidths, they show diversity in task execution and resource allocation. In addition, the number of tasks that need to be offloaded by different user devices, the time required for offloading, and the energy consumption are also different. However, most of the current research on task offloading methods assumes that edge servers are homogeneous devices, and does not take into account the differences in user devices in actual application scenarios.

[0051] Figure 1 The overall architecture of task offloading in edge computing environment is shown. It includes different subtasks to be processed, different user devices, and heterogeneous edge servers. Among them, the relationship between subtasks can be represented by a directed acyclic graph DAG. The vertices of DAG represent subtasks, and the directed edges between vertices represent the dependencies between subtasks. In DAG, subtasks without dependencies are executed in parallel to improve efficiency. Subtasks with dependencies need to be executed serially. During task execution, subtasks can be dynamically scheduled according to the topological structure of DAG to adapt to resource changes and task execution. Therefore, before the user device sends a task request, the task is analyzed in detail by the analyzer, including evaluating resource requirements and resolving dependencies. Then, according to the predetermined offloading strategy, these subtasks are assigned to heterogeneous edge servers, which receive and process tasks according to their own characteristics and current status. Finally, the execution results will be returned to the user device through the feedback mechanism to ensure the accuracy and timeliness of data transmission.

[0052] In edge computing environments, task offloading refers to the process of migrating computing tasks from user devices to edge servers or edge nodes for execution. Markov decision process is a mathematical framework used to model decision-making problems in uncertain environments. In edge computing, task offloading decisions are viewed as an MDP, represented by a five-tuple To express;

[0053] Among them, s is the state space, which is used to describe the state of the system environment; according to the current state, a task is selected from all tasks waiting to be assigned, and after the task is unloaded to the edge server, the next state is obtained; a is the action space, including all actions selected in the current state; after selecting the subtask to be executed, the present invention needs to select an edge server for it to execute the task. Therefore, the index of each edge server is an action.

[0054] r is an immediate reward, which is used to evaluate the feedback of the action. Because the optimization goal of the present invention comprehensively considers the total completion time, energy consumption and throughput of the task, the immediate reward of the present invention also considers these three factors.

[0055] p is the state transition probability, which describes the probability of transitioning from one state to another. In task offloading, it involves the probability of task transition between different offloading decisions.

[0056] It is a discount factor that measures the importance of immediate rewards and long-term rewards. It belongs to the interval [0,1], where 0 means that only immediate rewards are considered, and 1 means that long-term rewards and immediate rewards are considered equally important.

[0057] Because the optimization goal of the present invention comprehensively considers the total completion time, energy consumption and throughput of the task, the instant reward of the present invention also considers these three factors. The calculation formula of r is as follows: ; in, is the transmission time of the ith task of user device k, is the running time of the i-th task of user device k, is the communication energy consumption of the ith task of user device k, ) is the running energy consumption of the i-th task of user device k.

[0058] The task offloading model includes an action network and a target network; The action network and the target network have the same structure; In the action network, it is used to evaluate the Q value of each action in the current state and select the optimal action. Its parameters are continuously updated through training; In the target network, it is used to provide a stable training target. The parameters of the target network are copied from the action network Q network every C steps.

[0059] The target network Q' is a copy of the action network, with the same structure as the action network, but with a lower frequency of parameter updates. The role of the target network is to provide a stable training target to avoid frequent updates of the action network that lead to unstable training.

[0060] like Figure 3 As shown in the figure, the action network includes several neural network layers, including: input layer, first hidden layer, second hidden layer, asymptotic rectangular window attention mechanism, output layer; the attention mechanism is integrated into it to enhance the model's ability to capture key information.

[0061] The input layer of the action network receives the feature vector, which is a collection of attributes describing the task, the state of the edge server, and the characteristics of the user device. After the linear transformation and relu activation function processing of the first hidden layer, the output is passed to the second hidden layer; after the second hidden layer, the asymptotic rectangular window attention mechanism is introduced, which divides the output of the second hidden layer into multiple windows and performs weighted summation on the features in each window to highlight important features. Specifically:

[0062] 1) First, the features in each rectangular window, that is, the features processed by the linear transformation of the first hidden layer and the relu activation function, are transformed into feature representations through linear transformation; then, the tanh activation function is used for nonlinear processing; the expressive power of the features is enhanced. Next, the attention weights in each window are obtained by calculating the dot product similarity between the feature representation and the learnable vector and normalizing it through the Softmax function; these weights are used to weight the features in the window, thereby highlighting the key features in the window. A learnable vector refers to a vector parameter that can be automatically learned and adjusted during model training. In the asymptotic rectangular window attention mechanism, the role of this learnable vector is to perform dot product similarity calculations with the feature representation in each rectangular window, thereby obtaining the attention weights in each window. Specifically, this learnable vector can be regarded as a query vector, which represents the key feature direction that the model focuses on in the current task offloading decision process. By learning this learnable vector, the model can automatically identify which features are more important in the current decision scenario and give them higher attention weights. Finally, the weighted features of all windows are concatenated to form the output of the asymptotic rectangular window attention mechanism; this process not only effectively captures local features, but also enables the model to better adapt to the extraction of multi-scale features by gradually processing windows at different positions.

[0063] 2) The output of the asymptotic rectangular window attention mechanism passes through the first hidden layer, the second hidden layer and the asymptotic rectangular window attention mechanism again to further extract and abstract features; providing richer information representation for subsequent decisions while further enhancing the network's ability to capture local features.

[0064] The predicted Q value of each action is output through the last linear transformation plus the relu activation function. The predicted Q value represents the expected return of taking each possible action in the current state. The entire network structure is trained by the gradient descent method to minimize the difference between the predicted Q value and the target Q value, so as to make better decisions. Among them, the target Q value is calculated by the Bellman equation, which is the target value in the DQN training process and is used to guide the update of network parameters.

[0065] Through this structure, the task offloading model can focus on different parts of the input features through the asymptotic rectangular window attention mechanism while maintaining the good adaptability of the DQN algorithm to dynamic environments, thereby better learning and predicting.

[0066] In order to optimize the above TEH index and realize more energy-saving and efficient task offloading in edge computing environment, the present invention proposes a DPAQN algorithm that combines the asymptotic rectangular window attention mechanism and DQN. The basic framework of the DPAQN algorithm is that the present invention models the above problem as a Markov decision process (MDP), forms an algorithm for task offloading based on DQN, and then integrates the asymptotic rectangular window attention mechanism on the basis of the DQN algorithm.

[0067] Task offloading is achieved based on the DPAQN algorithm, including: First, initialize the edge computing environment and DQN parameters and state s (such as Figure 2 (center ①); Edge computing environment refers to a distributed computing architecture in which computing tasks can be performed on edge devices (such as edge servers) close to the data source. Specifically, it includes: heterogeneous edge servers with limited resources, user devices with diverse requirements, tasks with different attributes, and fluctuating network environments;

[0068] DQN parameters, including: number of neural network layers, number of neurons in each layer, weights and biases in the neural network, learning rate, discount factor, experience replay buffer size, target network update frequency, mini-batch sampling size, exploration rate; Status s, specifically including: task characteristics, edge server status, network status, and user device status; Using the current state s as input, the action network evaluates the predicted Q-values ​​of all possible actions, i.e. ; and select the action a with the highest Q value to execute (such as Figure 2(as shown in ②); Execute action a to interact with the environment, observe the reward r, and obtain the next state s' and the iteration end flag done; including: first, select the optimal action a based on the current state s and the evaluation results of the action network. This action usually means offloading the task to a certain edge server for execution. Then, the system executes action a, the task starts to be transmitted, and the reward r is calculated based on the completion time, energy consumption and throughput of the task. As the task is executed, the system state changes, including the update of the task queue, the load change of the edge server, the change of the network state, and the resource occupancy of the user device. These changes constitute the next state s'. Finally, the iteration end flag done is set according to whether the iteration termination condition is met. If the task is completed or the termination condition is met, done = 1; otherwise done = 0. This process enables the DQN algorithm to dynamically adjust the task offloading strategy through interaction with the environment to optimize the TEH comprehensive index value of the task.

[0069] Then, the current state s, action a, reward r, the next state s' and the iteration end flag done are stored in the experience replay buffer (such as Figure 2 (as shown in ③) for subsequent learning; randomly sample a small batch of data from the experience replay for training; The next state s' is used as the input of the target network Q' (such as Figure 2 As shown in ④), calculate the future ;include: The Q value of the optimal action a' predicted by the target network at the next state s'. This value is calculated by the target network Q' and represents the maximum future reward that can be obtained at state s'. is the maximum value among these Q values, namely: ;

[0070] Combine the reward r and the iteration end flag done (such as Figure 2 Calculate as shown in ⑤ , as shown below; ; in, It is the parameter of the target network, and the value of done is 0 or 1; The target network calculates the Q value. The target Q value combines the immediate reward r and the discounted value of future rewards to calculate the loss function.

[0071] The action network receives the state s and action a in the mini-batch sample (such as Figure 2 As shown in ⑥, the Q value of the action network (such as Figure 2 ⑦ in the figure) combined with action a in the mini-batch sampling (as shown in Figure 2As shown in ⑧ in the figure, we can get ; Then use the Mean Squared Error (MSE) loss function to calculate and Loss (like Figure 2 As shown in ⑨ in the figure), as shown below; ; m represents the number of samples in the mini-batch sampling; According to the obtained Loss, the parameters of the action network are updated by back propagation, and the parameters of the action network Q are copied to the target network Q' every C steps (such as Figure 2 to stabilize the training process.

[0072] Determine whether the termination condition is met. If so, the process ends; if not, the steps after returning to the initial state continue.

[0073] In this experiment, the present invention constructs a dynamic simulation environment of edge computing to study the offloading decision of user tasks on heterogeneous edge servers. The data set used is based on the real measurement data of offloading in image recognition calculations. At the same time, the impact of changes in the system environment such as the network in the real environment on the execution and travel time of the task is considered, and the simulation data is generated by adding noise. The environment initializes a certain number of users, and each user is assigned tasks with different attributes, including the number of instructions of the task, the size of the transmitted data, the expected running time, energy consumption, and CPU resource requirements. Considering that in some common scenarios such as industrial automation or enterprises, an edge server is usually responsible for a small number of users, so the present invention tests the number of users for 20, 30, 40, 50, 60, 70, 80, and 90. The specific parameters are shown in Table 1.

[0074] Table 1 Experimental configuration parameters;

[0075] The proposed DPAQN algorithm is averaged using the average TEH of the application. It is compared with DQN, DoubleDQN, DuelingDQN, Prioritized Replay and PPO algorithms. The performance improvement is calculated using the average value, and the average performance improvement can be defined as:

[0076] ; in, and They are the average TEH obtained by other algorithms and the average TEH obtained by the DPAQN algorithm, respectively.

[0077] This paper compares the DPAQN algorithm with five algorithms: DQN, Double Q-learing, DuelingDQN, Prioritized RePlay, and Proximal Policy Optimization (PPO), and tests the TEH index values ​​of four edge servers with different configurations scheduling different numbers of user tasks on the same data set. , the smaller the TEH value, the better the performance. The TEH values ​​of each algorithm in these eight cases on the four different edge servers are as follows Figure 4 shown.

[0078] The simulation results show that, in general, the TEH index value of DPAQN on four edge servers is about 24.51096269% better than DQN on average, about 23.87575690 better than Double Q-learing on average, about 23.52633988 better than Dueling DQN on average, about 30.39486988 better than Prioritized RePlay on average, and about 28.82079803% better than PPO on average. This shows that the DPAQN algorithm proposed in the present invention has certain advantages in optimizing the comprehensive performance of task offloading problems in edge computing environments. Specifically, since the network status and resource availability are dynamically changing in the edge computing environment, and the data in the edge computing environment is sensitive, the TEH index values ​​of different algorithms with different numbers of users vary greatly. However, in these eight cases, except for the case where the TEH index value of DPAQN is higher than that of the PPO algorithm when the number of users is 20, the TEH index value of the DPAQN algorithm is always better than the other five algorithms. Therefore, the DPAQN algorithm is expected to achieve more efficient and energy-saving offloading and optimize the load balancing of edge servers to a certain extent.

[0079] At the same time, in order to verify the generalization of the results of the present invention, the present invention also conducted comparative experiments under different random seeds. The results are as follows: Figure 5 As shown. In this case, DPAQN is generally better than DQN by about 28.93454619% on average, better than Double Q-learing by about 20.71041776% on average, better than Dueling DQN by about 29.19596098% on average, better than Prioritized RePlay by about 28.71283886% on average, and better than PPO by about 29.46298988% on average. This verifies the good generalization of DPAQN to a certain extent. This shows that the DPAQN algorithm proposed in the present invention can alleviate the problems of weak robustness and poor generalization of artificial intelligence algorithms to a certain extent.

[0080] In the industrial Internet of Things environment, smart manufacturing systems need to process a large amount of data generated by sensors and devices to achieve real-time monitoring, fault prediction and quality control of the production process. These tasks have high requirements on energy consumption and task completion time, and also require sufficient throughput to ensure production efficiency. The DPAQN algorithm can offload some computationally intensive tasks to edge servers by optimizing task offloading strategies, which is expected to reduce equipment energy consumption and increase task completion speed to a certain extent. For example, in a smart factory, sensors can send data to edge servers for analysis, and the DPAQN algorithm will dynamically adjust task offloading decisions based on the current network status and resource conditions. This optimization method is expected to reduce equipment energy consumption while improving the overall system throughput, providing certain support for the efficiency and intelligence of the production process.

[0081] The present invention takes into account the heterogeneity of edge servers and the diversity of users, and proposes the DPAQN algorithm for the problem of task offloading in edge computing environments. It retains the general advantages of artificial intelligence algorithms, while alleviating the instability that is common in artificial intelligence algorithms to a certain extent. The present invention comprehensively considers the three indicators of task completion time, energy consumption and throughput to experiment with the model. By comparing with the five algorithms of DQN, Double DQN, Dueling DQN, Prioritized Replay and PPO, it is proved that the comprehensive performance of the algorithm is good, and it is 20.71%-30.39% better than the other five algorithms on average.

[0082] Example 3 A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the task offloading method using deep reinforcement learning and attention mechanism described in Example 1 or 2 are implemented.

[0083] Example 4 A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of the task offloading method using deep reinforcement learning and attention mechanism described in Example 1 or 2 are implemented.

[0084] Example 5 Task offloading system using deep reinforcement learning and attention mechanism, including: The data acquisition and preprocessing module is configured to: data acquisition and preprocessing; The task offloading module is configured to input the preprocessed data into the trained task offloading model and realize task offloading based on the DPAQN algorithm.

Claims

1. A task offloading method using deep reinforcement learning and attention mechanism, characterized in that: include: Data acquisition and preprocessing; data includes edge server information, user information and task information; The preprocessed data is input into the trained task offloading model, and task offloading is realized based on the DPAQN algorithm.

2. The task offloading method using deep reinforcement learning and attention mechanism according to claim 1, characterized in that: Preprocessing, including: noise addition; dynamic environment simulation: by continuously generating new tasks, simulate a dynamic edge computing environment to ensure that the task flow is continuous; parameter setting.

3. The task offloading method using deep reinforcement learning and attention mechanism according to claim 1, characterized in that: Objective function in the task offloading model It is expressed as: ; in, ; Represents the maximum completion time of all completed tasks. Represents the number of completed tasks; Represents the total completion time of the task; ; Represents the maximum completion time of all completed tasks. represents the total energy consumption of the task; w1,w2,w3 are weights, It refers to the total throughput; The values ​​of w1, w2, and w3 are 0.25, 0.5, and 0.25 respectively.

4. The task offloading method using deep reinforcement learning and attention mechanism according to claim 1, characterized in that: The total task completion time TT includes the transmission time and running time of all tasks; it is specifically defined as: ; Where n is the number of users, l is the number of subtasks for each user, represents the data volume of the i-th task of user device k, represents the bandwidth allocated to the i-th task of user device k, is the transmission time of the i-th task of user device k, represents the time required for the i-th task of user device k to execute locally, represents the ratio of local and remote execution rates, is the running time of the i-th task of user device k; The total energy consumption of a task includes the communication energy consumption and operation energy consumption of all tasks; it is specifically defined as: ; in, is the transmission power of the ith task of user equipment k, represents the communication time of the i-th task of user device k; represents the communication energy consumption of the i-th task of user device k, that is, ; represents the running energy consumption of the i-th task of user device k.

5. The task offloading method using deep reinforcement learning and attention mechanism according to claim 1, characterized in that: In the edge computing environment, task offloading refers to the process of migrating computing tasks from user devices to edge servers or edge nodes for execution. The task offloading decision is regarded as an MDP, represented by a five-tuple To express; Among them, s is the state space, which is used to describe the state of the system environment; according to the current state, a task is selected from all tasks waiting to be assigned, and after the task is unloaded to the edge server, the next state is obtained; a is the action space, including all actions selected in the current state; r is the immediate reward, which is used to evaluate the feedback of the action; p is the state transition probability, which describes the probability of transitioning from one state to another; It is a discount factor that measures the importance of immediate rewards and long-term rewards. It belongs to the interval [0,1], where 0 means that only immediate rewards are considered, and 1 means that long-term rewards and immediate rewards are considered equally important.

6. The task offloading method using deep reinforcement learning and attention mechanism according to claim 5, characterized in that: The calculation formula for r is as follows: ; in, is the transmission time of the ith task of user device k, is the running time of the i-th task of user device k, is the communication energy consumption of the ith task of user device k, ) is the running energy consumption of the i-th task of user device k.

7. The task offloading method using deep reinforcement learning and attention mechanism according to claim 1, characterized in that: The task offloading model includes an action network and a target network; The action network and the target network have the same structure; In the action network, it is used to evaluate the Q value of each action in the current state and select the optimal action. Its parameters are continuously updated through training; In the target network, it is used to provide a stable training target. The parameters of the target network are copied from the action network Q network every C steps.

8. The task offloading method using deep reinforcement learning and attention mechanism according to claim 1, characterized in that: The action network includes several neural network layers, including: input layer, first hidden layer, second hidden layer, asymptotic rectangular window attention mechanism, and output layer; The input layer of the action network receives the feature vector, and after the linear transformation and relu activation function processing of the first hidden layer, the output is passed to the second hidden layer; after the second hidden layer, the asymptotic rectangular window attention mechanism is introduced, which divides the output of the second hidden layer into multiple windows and performs weighted summation on the features in each window; specifically: 1) First, the features in each rectangular window, that is, the features processed by the linear transformation of the first hidden layer and the relu activation function, are transformed into feature representations through linear transformation; then, nonlinear processing is performed using the tanh activation function; then, the dot product similarity between the feature representation and the learnable vector is calculated and normalized by the Softmax function to obtain the attention weight in each window; finally, the weighted features of all windows are concatenated to form the output of the asymptotic rectangular window attention mechanism; 2) The output of the asymptotic rectangular window attention mechanism passes through the first hidden layer, the second hidden layer and the asymptotic rectangular window attention mechanism again to further extract and abstract features; The predicted Q value of each action is output through the last linear transformation plus the relu activation function. The predicted Q value represents the expected return of taking each possible action in the current state.

9. The task offloading method using deep reinforcement learning and attention mechanism according to any one of claims 1 to 8, characterized in that: Task offloading is achieved based on the DPAQN algorithm, including: First, initialize the edge computing environment, DQN parameters, and state s; Using the current state s as input, the action network evaluates the predicted Q-values ​​of all possible actions, i.e. ; and select action a with the highest Q value to execute; Execute action a to interact with the environment, observe the reward r, and obtain the next state s' and the iteration end flag done; Then, the current state s, action a, reward r, next state s' and iteration end flag done are stored in the experience replay buffer for subsequent learning; a small batch of data is randomly sampled from the experience replay for training; Take the next state s' as the input of the target network Q' and calculate the future ; Combine the reward r and the iteration end flag done to calculate , as shown below; ; in, It is the parameter of the target network, and the value of done is 0 or 1; The action network receives the state s and action a in the mini-batch sampling, and the Q value of the action network is combined with the action a in the mini-batch sampling to obtain ; Then use the Mean Squared Error loss function to calculate and Loss , as shown below; ; m represents the number of samples in the mini-batch sampling; Update the parameters of the action network through back propagation according to the obtained Loss, and copy the parameters of the action network Q to the target network Q' every C steps; Determine whether the termination condition is met. If so, the process ends; if not, the steps after returning to the initial state continue.

10. Task offloading system using deep reinforcement learning and attention mechanism, characterized by: include: The data acquisition and preprocessing module is configured to: data acquisition and preprocessing; The task offloading module,is configured as; The preprocessed data is input into the trained task offloading model, and task offloading is realized based on the DPAQN algorithm.

Citation Information

Patent Citations

  • Resource joint allocation method based on deep reinforcement learning in Internet of Vehicles

    CN112995950A

  • Unmanned aerial vehicle edge calculation unloading method based on multi-target deep reinforcement learning

    CN115827108A

  • Calculation unloading method and system for task with dependency relationship in edge calculation

    CN116755882A

  • Reinforcement learning workflow task unloading method and device based on attention mechanism

    CN118349293A

  • Calculation unloading optimization strategy based on multi-agent deep reinforcement learning

    CN119322681A

Cited By

  • Task scheduling optimization method and system of edge system based on reinforcement learning

    CN121092295A

  • A method and system for task scheduling optimization of an edge system based on reinforcement learning

    CN121092295B