A Task Offloading Decision Method for Multiple Intelligent Devices Based on Deep Reinforcement Learning

By applying the multi-intelligent device task offload decision method based on deep reinforcement learning in a dynamic MEC environment, the problems of multi-user task computing offloading and resource allocation optimization are solved, and efficient data transmission and user experience improvement are achieved.

CN115065678BActive Publication Date: 2025-06-17SOUTHEAST UNIV

Patent Information

Application Number
CN202210362289.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-07
Publication Date
2025-06-17
Estimated Expiration
2042-04-07

AI Technical Summary

Technical Problem

In dynamic MEC environments, multi-user task computing offloading and resource allocation optimization are still a serious problem. Traditional methods such as game theory and optimization theory encounter obstacles in complex scenarios, and deep reinforcement learning (DRL) has problems of environmental instability and feedback overhead in dynamic environments of multi-intelligent devices.

Method used

A multi-intelligent device task offload decision method based on deep reinforcement learning is proposed. By obtaining the configuration information of smart terminal devices and environments, a multi-intelligent terminal device offload model aimed at the maximum data transmission rate is established. Mobile users are trained to find the best offload decision, so that the system can achieve the optimal channel and power distribution scheme.

Benefits of technology

The optimization of multi-user task computing in a dynamic MEC environment is achieved, which improves data transmission rate and user experience, reduces network congestion and energy consumption, and solves the limitations of traditional methods in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115065678B_ABST
    Figure CN115065678B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-intelligent device task offloading decision method based on deep reinforcement learning, which is used to solve the problems of hybrid task offloading and resource allocation in a cloud-edge-end fusion network with multiple intelligent terminal devices in the Internet of Things environment. This method uses the data transmission rate as the performance evaluation index. First, obtain the configuration information of each intelligent terminal device and the environment in this environment, then establish a multi-intelligent terminal device offloading model with the maximum data transmission rate as the goal, and finally solve the optimal task offloading scheme based on the deep reinforcement learning method, and then perform task offloading according to the overall optimal offloading scheme. The present invention can effectively solve the problems of multi-intelligent device hybrid task offloading and resource allocation in a multi-access edge computing (MEC) network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a multi-intelligent device task offloading decision method based on deep reinforcement learning, belonging to the technical fields of Internet of Things and artificial intelligence. Background Art

[0002] With the popularization of general intelligent terminals and ubiquitous network technologies, the traditional centralized cloud computing model has some deficiencies in meeting the different needs of the rapidly growing new delay-sensitive, computationally intensive network applications and services. In recent years, widely distributed and geographically distributed mobile devices and Internet of Things devices have replaced traditional large cloud data centers and created more and more data; therefore, the cloud computing model with centralized computing and storage as its basic characteristics is difficult to meet the needs of various technologies and application scenarios. Compared with network expansion and optimization, data offloading is an effective strategy to reduce network congestion and can significantly and effectively solve the overload problem. Therefore, multi-access edge computing (MEC), as a new computing paradigm, has been proposed to alleviate the network congestion caused by directly transferring computing tasks from resource-limited mobile terminals to cloud computing centers and is considered a promising technology for 5G heterogeneous networks. In addition, MEC reduces latency for mobile terminal users by providing edge storage and computing resources, provides faster responses, and thus improves the quality of the user experience (QoE).

[0003] Due to the advantages of MEC, there is an increasing interest in MEC in both the academic and industrial communities. However, in an actual MEC system, since offloading highly depends on the efficiency of wireless data transmission and its influencing factors are multi-dimensional, random, and time-varying, the MEC system must manage wireless resources and computing resources simultaneously. Therefore, traditional methods such as game theory and optimization theory encounter obstacles and limitations in solving the computational offloading optimization problem in complex scenarios. Fortunately, deep reinforcement learning (DRL) can effectively solve the above problems in MEC. It is an emerging and effective method to maximize long-term rewards and obtain optimal decision-making strategies. In previous MEC research, especially in large MEC networks, the resource allocation and offloading computation problems are usually formulated as mixed integer programming (MIP) problems, which are solved by introducing dynamic programming or branch and bound algorithms. Unfortunately, the computational complexity of these two methods is particularly high. Although heuristic local search and convex relaxation can reduce the computational complexity, both require multiple iterations to achieve the expected local optimum. Therefore, compared with these traditional algorithms, DRL has been widely recognized as a new method that enables mobile devices and mobile edge servers to make adaptive and effective decisions on resource allocation problems. However, most of the existing work focuses on static MEC single-user computing optimization. In addition, in reinforcement learning algorithms, each intelligent terminal is constantly learning and improving its own strategy. Then, from the perspective of each intelligent terminal, the environment is unstable, which is not conducive to convergence. Traditional deep reinforcement learning applications such as Deep Q-networks (DQN) have the following disadvantages in a multi-intelligent device dynamic environment: 1) Due to the unstable environment, it is impossible to directly use the experience replay techniques DQN and DDQN. 2) Due to the inevitable feedback overhead caused by the interaction of many intelligent terminals, the use of MDP will suffer from the curse of dimensionality. 3) Due to the instability of the environment, it is impossible to adapt to the dynamic and unstable environment only by changing the policy of the intelligent terminal itself. Therefore, achieving multi-user computing optimization in a dynamic MEC environment remains a serious problem. Summary of the Invention

[0004] Object of the Invention: Aiming at the problems and deficiencies in the prior art, the present invention proposes a multi-intelligent device task offloading decision method based on deep reinforcement learning, which solves the problems of multi-user task computational offloading and resource allocation optimization in a dynamic MEC environment.

[0005] Technical Solution: A multi-intelligent device task offloading decision method based on deep reinforcement learning according to the present invention includes the following steps:

[0006] Step 1: Obtain the configuration information of each intelligent terminal device and the environment in this environment. According to the situation of each intelligent terminal in the actual environment, collect their configuration information for subsequent modeling.

[0007] Step 2: Establish a multi-intelligent terminal device offloading model with the maximum data transmission rate as the goal. When the task is small enough, the user can choose local computing; when the user is close to the base station, the task can be offloaded to the base station for computing; when the user is far from the base station, this model can transfer the tasks of the mobile terminal to neighboring devices closer to the base station. Adjacent devices can make a secondary decision based on the amount of task data: if the task volume is small, the device can use its own idle computing resources for computing, and if the task volume is large, the device will offload the task to the nearby base station server network for computing.

[0008] Step 3: Solve the optimal task offloading scheme based on the deep reinforcement learning method. In this system model, the entire edge network is modeled as a Markov decision process (MDP), where the state represents the channel gain, the action represents the power allocation, and the overheads (such as data transmission rate and energy consumption) are designed as a reward function. Through the training of the deep neural network, the mobile user can find the best offloading decision, enabling the system to achieve the optimal channel and power allocation scheme. Step 3 is divided into two parts: centralized training and distributed application. During training, each intelligent terminal can obtain the system reward through learning, and then intelligent terminal t adjusts its behavior by updating the DQN to adapt to the optimal policy π * . In the application, each intelligent terminal makes a partial observation of the dynamic MEC environment and then selects the action corresponding to the small-scale channel fading according to its own training experience.

[0009] Step 4: Perform task offloading according to the overall optimal offloading scheme. In the real MEC network, mobile users have strong dynamics. As the user's position changes, the intensity of the communication signal also changes accordingly. For specific situations, it is necessary to return to Step 3 for mobile devices with corresponding configuration information changes to redesign the offloading scheme suitable for the current situation.

[0010] Compared with the prior art, the present invention has the following advantages:

[0011] 1) Aiming at the fact that the static MEC system cannot reflect the mobility and environmental adaptability of mobile users in the real network, a dynamic MEC system model that combines offloading computing and power resource allocation is proposed. This model fully considers the impact of path loss on the task data transmission rate.

[0012] 2) The present invention adopts energy harvesting technology to obtain radio frequency energy from external interference, improve the energy of mobile devices, and at the same time improve the computing efficiency of the devices.

[0013] 3) The research on multi-intelligent terminal applications based on discrete decisions is a very complex issue. Current multi-intelligent terminal algorithms have various limitations in value-based application scenarios, such as the inability to achieve effective convergence. The present invention first applies the ideas of centralized training and distributed applications to the value-based multi-intelligent terminal MEC system. A two-stage training model is adopted to share training data by centrally training multi-intelligent terminals, ensuring that the algorithm can be applied to MEC network scenarios of different scales and maximizing the benefits of multi-intelligent terminal collaborative training. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 It is a structural example diagram of a task offloading decision method for multi-intelligent devices based on deep reinforcement learning.

[0015] Figure 2 It is an example diagram of hybrid offloading of mobile terminal devices in the actual MEC network of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0016] To deepen the understanding and recognition of the present invention, the present invention will be further clarified below in conjunction with specific embodiments.

[0017] Embodiment 1: Refer to Figure 1 、 Figure 2 , a task offloading decision method for multi-intelligent devices based on deep reinforcement learning, includes the following steps:

[0018] Step 1: Obtain the configuration information of each intelligent terminal device and the environment in this environment. In this instance, it is first necessary to obtain the specific configuration information in the cloud-edge-end fusion environment in the Internet of Things environment, such as the number N of intelligent terminal devices, and the set of intelligent terminals is {x1, x2,..., x N ,}, the current state S t at a given time step t, and the current state S t includes all channel state information. Each mobile intelligent terminal x i obtains the observation value O(S t , i), and then executes to generate the joint action A i , and the reward of the intelligent terminal x i within the time step t can be defined as: ω1, ω2, and ω3 are rate weights that balance local computing and M2B and U2U objectives. In addition, all intelligent terminals obtain the reward r t and transfer to the next state S t+1 with a probability of p(S t+1 , r t | S t , A), and the next observation O(S t+1, i) Obtained from all intelligent terminals. The operations of each intelligent terminal are unknown to each other. Therefore, each intelligent terminal needs to observe to understand the entire environment. Each intelligent terminal x i 's observation includes: channel gain g i , interference channel gain from other mobile terminals interference channel gain from the base station Within each time step t, the above channel gains can be accurately calculated. Therefore, the observation value O(S t , i) is defined as: In the present invention, we consider that all channels in the propagation scenario are dominated by line-of-sight (LOS) links. The proposed channel model is applicable to wireless network systems with a frequency range of 2 - 6 GHz and a radio frequency bandwidth of up to 100 MHz.

[0019] Step 2: Establish a multi-intelligent terminal device offloading model with the goal of maximizing the data transmission rate. In a real MEC network, there are a large number of idle mobile devices. Mobile devices have corresponding computing performance and storage capacity, and can perform local computing for tasks with relatively small computing loads. To maximize the utilization of these idle mobile device resources, when the communication distance between a mobile user and the base station is too long, the mobile user can transfer tasks to the edge server through idle neighbor devices (similar to a cognitive radio network). In addition, in the case of a relatively small amount of computation, idle proximal mobile devices can also choose to directly complete the task offloading calculation of remote devices, thereby reducing the delay and energy consumption of task calculation.

[0020] The data transmission rates in the three task calculation modes are respectively:

[0021] (4) The data transmission rate of mobile user i in the local computing mode:

[0022]

[0023] Among them, f i represents the CPU frequency, and U i represents the computing power required to process 1 bit of data in time step t.

[0024] (5) The data transmission rate of mobile user i in the user-to-base station (M2B) mode:

[0025]

[0026] Among them, W represents the bandwidth, represents the transmission power between user i and the base station, P loss represents the loss power, represents the channel gain between user i and the base station, σ2 represents the noise power, represents the interference between the MEC server and mobile device i.

[0027] (6) Data transmission rate of mobile user i in the machine-to-machine (M2M) mode:

[0028]

[0029] Therefore, in the MEC network, the purpose of offloading users' tasks to the MEC server is to improve the data transmission rate in mobile computing, that is, to maximize the sum of data transmission rates γ i :

[0030]

[0031] where x i ∈ {0, 1}, representing the indicator of local computing, and y i ∈ {0, 1} represents the indicator of M2B offloading computing and M2M offloading computing, represents the data transmission rate of mobile user i in the local computing mode, represents the data transmission rate of mobile user i in the machine-to-base station (M2B) mode, represents the data transmission rate of mobile user i in the machine-to-machine (M2M) mode.

[0032] The resource allocation problem is modeled as follows:

[0033] (1) Channel resource allocation. For all users i ∈ I and sub-channels k ∈ K, there is represents the sub-channel between the user and the server and the sub-channel between users. Assume that a mobile device can only access one sub-channel (which can be regarded as a resource block), that is

[0034]

[0035] (2) Power resource allocation. For all users i ∈ I, there is respectively represent the transmission power between the user and the server and the transmission power between users.

[0036] Step 3: Solve the optimal task offloading scheme based on the deep reinforcement learning method.

[0037] In this step, the present invention models the dynamic multi-intelligent terminal system model as a discrete-time stochastic control process model, i.e., a Markov decision process (MDP) model. The MDP can be described as <S, A, r, P>, where S is the set of states, A is the set of actions, r is the reward received after the execution of an action, and P is the transition probability. Let π: S → A denote the policy. The goal of the optimal task offloading scheme is to obtain the optimal policy π: S → A by maximizing the cumulative reward function, i.e., the maximum data transmission rate obtained by all intelligent terminals in the dynamic MEC system. The cumulative reward function for T steps is where γ is the discount factor, a t = π * (s t ).

[0038] This step is divided into two sub-steps:

[0039] Sub-step 3-1: Centralized learning.

[0040] Let V π : S1×S2×…×S n → R represent the value function, and its calculation formula is as follows:

[0041] V π (S) = E π [r t (S t , A t ) + γV π (S)|S0 = S]

[0042] where r t (S t , A t ) represents the reward of the intelligent terminal when the state is S t and the action is A t at time step t, and γ is the discount factor.

[0043] The calculation formula of the optimal value function is as follows:

[0044]

[0045] The optimal value function can be transformed to solve the optimal Q function as follows,

[0046]

[0047] This process is solved iteratively and simplified to find the optimal Q function Q*(S, A) for all (S, A).

[0048] In traditional DRL algorithms, to improve the training efficiency in high-dimensional state space and action space environments, DNNs are usually introduced to replace Q-tables for deriving the approximate formula Q*(S,A); the result is called DQN. In the application of reinforcement learning, there is a difficult problem to solve; that is, the training network data collection speed is too slow. Therefore, DQNs use the experience replay mechanism to solve this problem. However, the performance of DQNs is not ideal because of the overestimation problem of Q-values. In fact, whether the algorithm is Q-learning algorithm or traditional DQN, the target Q-value is calculated by the greedy policy.

[0049] Although using argmax can quickly make the Q-value approach the optimal value, this usage is likely to lead to overestimation and introduce a large deviation to the final result. To solve this problem, DDQN decouples two steps: selecting the target Q-value action and calculating the target Q-value to eliminate the overestimation problem.

[0050] Similar to traditional DQNs, DDQNs have two Q-network structures: the target Q-network and the evaluation Q-network. However, DDQNs will not find the maximum value maxQ(S t ,A t ) in the target network, but directly search for the action corresponding to the maximum value maxQ(S t ,A t ) in the evaluation network, and the calculation is as follows:

[0051]

[0052] where w2 is the weighted parameter of the current network, S i+1 represents the state at time step i + 1, and a i+1 represents the action at time step i + 1.

[0053] Next, use the selected to solve the expected value yi in the target network:

[0054]

[0055] where R i represents the T-step cumulative reward function, w1 is the weighted parameter of the evaluation network, w2 is the weighted parameter of the current network, S i+1 represents the state at time step i + 1, a i+1 represents the action at time step i + 1, and γ is the discount factor. =

[0056] The calculation formula of the loss function of the Q-network is as follows:

[0057]

[0058] where R iDenote the cumulative reward function at T steps, w1 is the weighted parameter of the evaluation network, w2 is the weighted parameter of the current network, S j+1 Denote the state at time step j + 1, a j+1 Denote the action at time step j + 1, S j Denote the state at time step j, a j Denote the action at time step j.

[0059] To update the parameter w2 in the loss function, the commonly used update algorithm is Stochastic Gradient Descent (SGD). However, due to the frequent updates of SGD, the cost function will be severely affected. In addition, SGD has a large amount of noise; although the training speed is fast, the accuracy is reduced, so SGD does not tend to global optimization in each iteration. Therefore, we consider using the adaptive learning rate method RMSprop to solve the global optimization problem and significantly reduce the learning rate. The update strategy of the RMSprop algorithm is expressed as:

[0060]

[0061] where Denote the gradient of the loss function L at time slot t, w t Denote the parameters of the deep neural network in time slot t, Denote the expected value of the square of the previous gradient, β represents the momentum factor, α represents the global learning rate, and μ is a very small constant to avoid division by zero.

[0062] Although the DDQN algorithm for a single intelligent terminal has good convergence performance, there are still some limitations, especially when this algorithm is directly applied to a multi-intelligent terminal environment. For example, when there are two intelligent terminals A and B in the MEC environment, for intelligent terminal A, intelligent terminal B is also part of its environment. Therefore, when intelligent terminal B performs an operation that changes its environment, the environment of intelligent terminal A will also change. Therefore, for multi-intelligent terminals that collaborate to make decisions, when one of the intelligent terminals changes its own environment, the decisions made by the remaining intelligent terminals are no longer applicable. For DRL, the application scenarios of multi-intelligent terminal collaboration usually produce a large variance because the rewards obtained by each intelligent terminal depend on the joint actions of multi-intelligent terminals. However, the variability of the rewards of intelligent terminals themselves leads to an increase in the variance of their gradients.

[0063] In addition, in the application of multi-intelligent device collaboration, each intelligent terminal should consider the states of other intelligent terminals; doing so involves the communication problem between intelligent terminals. Each intelligent terminal needs to know all the real-time information of other intelligent terminals, but this requirement is difficult to meet in practical applications.

[0064] To solve the above problems, we made the following improvements:

[0065] a) We adopt the method of centralized learning to train the DNN. In a distributed application, each intelligent terminal only needs to know local information to make a decision, but all intelligent terminals share all learning experiences, thus improving their next decision level.

[0066] b) We improved the experience replay mechanism of DRL. To adapt to the dynamic MEC environment, we reset the format of the experience replay pool sample information to: (S1,…,S t+1 ,A1,…,A t ,R1,…,R t ), where S t =(O1,…,O n ) represents the observations of n intelligent terminals at time slot t, A t =(a1,…,a n ) represents the actions of n intelligent terminals, and R t =(r1,…,r n ) represents the rewards obtained by n intelligent terminals at time slot t.

[0067] c) Policy integration optimization. For the policy of each intelligent terminal, use the policy experiences of all intelligent terminals for global optimization to improve the stability and robustness of the algorithm.

[0068] Sub-step 3-2: Multi-user distributed application.

[0069] After the learning process is completed, it enters the multi-user distributed application stage. In the implementation stage, each mobile user estimates the channel state value and obtains local observations in time slot t. Next, intelligent terminal i will select an action with argmax Q according to its trained Q network After that, all users start to make computing decisions (offloading computing or local computing). In addition, the power and channel gain required for task computing are determined by the actions they select.

[0070] In addition, the proposed multi-user intensive training is executed offline for different channel conditions in different network topologies. Therefore, only when the network environment changes greatly (resulting in a great change in channel conditions), all training experiences need to be updated again.

[0071] The intelligent terminal that realizes the maximum benefit must also contribute to the system. In other words, in the MEC system, each user should help other users achieve the maximum transmission efficiency of the system while achieving its own maximum transmission rate. However, it is difficult to achieve this goal with the traditional overall reward method.

[0072] Therefore, we considered a two-step training model: 1) In the first step, let a single intelligent terminal achieve the best performance in a dynamic environment; 2) The second step is to initialize the single intelligent terminal that implements the optimal policy in the first step as multiple intelligent terminals in a dynamic environment. Different from the traditional multi-intelligent device algorithm with sequential selection strategy, the intelligent terminals in our algorithm make decisions according to their goals in the second step. To accurately observe the contribution of each intelligent terminal to the system, we introduced a credit assignment mechanism and a trust function Ω m (s,a n ) to evaluate each target θ m and action A n . Let π(a m ) = π(a n |θ m ,o n ), where o n is an observation of intelligent terminal n, and are the goals of all intelligent terminals.

[0073] In a multi-intelligent device system, a n ∈ A n ; s ∈ S; for intelligent terminal n, the trust function Ω m (s,a n ) is calculated as follows:

[0074]

[0075] where s represents the state, a n represents the action, γ is the discount factor, and R represents the reward.

[0076] According to the iterative calculation rules mentioned above, the trust function Ω m (s,a n ) satisfies the Bellman equation:

[0077]

[0078] where s represents the state, a n represents the action, θ m represents the goal, and γ is the discount factor.

[0079] Then V m (s) is calculated as follows,

[0080]

[0081] where s represents the state, a represents the action, o n (s) is an observation of intelligent terminal n in state s, and θ m represents the goal.

[0082] Ω m (s, a n ) is used to evaluate each target and action, so that

[0083] In a multi-intelligent device system, let θ parameterize π and make it any natural number; then, the gradient of the total objective is:

[0084]

[0085] where s represents the state, a represents the action, o n is an observation of intelligent terminal n, θ m represents the target, π represents the policy, Ω m (s, a) represents the trust function.

[0086] The trust function Ω m (s, a) and the value function V m (s) are converted into each other in the following way:

[0087]

[0088] where R represents the reward, s represents the state, a represents the action, θ m represents the target, γ is the discount factor, P represents the state transition probability, and π represents the policy.

[0089] The method proposed by the present invention first uses centralized training to train a single intelligent terminal, and finally prepares for improving the efficiency of distributed multi-intelligent terminal decision-making.

[0090] Step 4: Perform task offloading according to the overall optimal offloading scheme. The most important feature of deep reinforcement learning in solving optimization problems is the flexible reward design pattern. When the reward setting model is reasonable, high-efficiency learning performance can be obtained. Our goal is to maximize the data transmission rate of the MEC system and increase the success rate of user task delivery. Each user helps other users achieve the maximum transmission efficiency of the system while achieving their own maximum transmission rate. According to the optimal task offloading scheme obtained in step 3, each intelligent device selects the optimal offloading scheme, thereby reducing the delay and energy consumption of task calculation.

[0091] Those of ordinary skill in the art will realize that the embodiments described herein are for helping the reader understand the principles of the present invention. It should be understood that the embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent forms of modification of the present invention fall within the scope defined by the appended claims of this application.

Claims

1. A task offloading decision method for multiple intelligent devices based on deep reinforcement learning, characterized in that, It includes the following steps: Step 1: Obtain the configuration information of each smart terminal and the environment in the environment; Step 2: Establish a multi-smart terminal offloading model with the data transmission rate as the performance evaluation index; Step 3: Solve the optimal task offloading scheme based on the deep reinforcement learning method; Step 4: Perform task offloading according to the overall optimal offloading scheme; Among them, Step 1: Obtain the configuration information of each smart terminal and the environment in the environment, specifically as follows: (1) In a multi-intelligent device scenario, each mobile user is an intelligent terminal, exploring the unknown environment simultaneously. At the current state S at a given time step t, t , each mobile user i obtains the observation O(S t , i), and then executes to generate the joint action A i . In addition, all intelligent terminals obtain the reward r t and transfer to the next state S t+1 with a probability of p(S t+1 , r t | S t , A). The next observation O(S t+1 , i) is obtained by all intelligent terminals; (2) In the real MEC scenario, the current state S t includes all the channel state information. In addition, the operations of each intelligent terminal are unknown to each other. Therefore, each intelligent terminal needs to observe to understand the entire environment. The observations of each intelligent terminal i include: the channel gain g i , the interference channel gain from other mobile users , the interference channel gain from the base station At each time step t, the above channel gains can be accurately calculated. The observation value O(S t , i) is defined as: The reward of the intelligent terminal i within the time step t can be defined as: ω1, ω2, and ω3 are the rate weights for balancing local computing and the Mobile user to Base station (M2B) and user to user (U2U) objectives, respectively; where i = 1, …… N, t = 0, …… T; Step 2: Establish a multi-smart terminal offloading model with the maximum data transmission rate as the goal, The data transmission rates under the three task calculation modes are respectively: (1) The data transmission rate of mobile user i in the local calculation mode: Among them, f i represents the CPU frequency, and U i represents the computing power required to process 1 bit of data in the time step t. (2) The data transmission rate of mobile user i in the user-to-base station M2B mode: Where W represents the bandwidth, represents the transmission power between the mobile user i and the base station, P loss represents the loss power, represents the channel gain between the mobile user i and the base station, σ 2 represents the noise power, represents the interference between the MEC server and the mobile device i, (3) The data transmission rate of mobile user i in the user-to-user U2U mode: Therefore, in the MEC network, the purpose of offloading the user's tasks to the MEC server is to improve the data transmission rate in mobile computing, that is, to maximize the sum of data transmission rates γ i : where x i ∈ {0, 1}, is an indicator representing local computation, and y i ∈ {0, 1} is an indicator representing M2B offloading computation and U2U offloading computation, represents the data transmission rate of mobile user i in the local computation mode, represents the data transmission rate of mobile user i in the user-to-base station M2B mode, represents the data transmission rate of mobile user i in the user-to-user U2U mode.

2. The task offloading decision method for multiple intelligent devices based on deep reinforcement learning according to claim 1, characterized in that, Step 3: Solve the optimal task offloading scheme based on the deep reinforcement learning method, specifically as follows, (1) Model the dynamic multi-smart terminal system model as a discrete-time stochastic control process model, that is, a Markov decision process MDP model. The MDP is described as <S, A, r, P>, where S is the set of states, A is the set of actions, r is the reward received after the action is executed, P is the transition probability, and let π: S → A represent the policy; (2) The goal of the optimization problem is to obtain the optimal policy by maximizing the cumulative reward function, and the cumulative function is where γ is the discount factor, a t = π * (S t ).

3. The task offloading decision method for multiple intelligent devices based on deep reinforcement learning according to claim 1, characterized in that, Step 4: Perform task offloading according to the overall optimal offloading scheme, that is, the task offloading scheme that can maximize the data transmission rate of the task by reasonably allocating limited computing resources. According to the optimal task offloading scheme obtained in Step 3, make a decision on the multi-smart device task offloading. First, according to the configuration information in Step 2, apply Step 3 to construct the model. After obtaining the optimal offloading scheme in Step 4, implement the smart terminal task offloading.

Citation Information

Patent Citations

  • Mobile edge computing task unloading method and device based on transfer learning

    CN113504987A

Cited By

  • Task unloading and cache updating method based on federal reinforcement learning in Internet of Vehicles

    CN116437481A

  • Task offloading and cache update method based on federated reinforcement learning in internet of vehicles

    CN116437481B

  • Task scheduling and resource allocation method based on federal reinforcement learning in Internet of Vehicles

    CN116709378A