A Task Offloading and Resource Allocation Method Based on Mobile Edge Computing

The DSRA algorithm addresses the challenge of dynamic task offloading and resource allocation in decentralized MEC systems by using multi-agent DRL with LSTM networks, optimizing task processing delay and resource allocation in MEC systems.

CN116137724BActive Publication Date: 2025-07-15CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310138344.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-20
Publication Date
2025-07-15
Estimated Expiration
2043-02-20

AI Technical Summary

Technical Problem

In mobile edge computing environments, it is difficult for the existing technology to achieve dynamic and flexible task offloading and resource allocation, especially in decentralized network environments, where task processing delays, resource allocation is uneven, and the single agent deep reinforcement learning algorithm is not effective in dynamic environments.

Method used

Multi-agent deep reinforcement learning algorithm (DSRA) is adopted and combined with LSTM network to construct part of the observable Markov decision-making process. Through the base station as the agent, we learn task offloading and resource allocation strategies, considering the time dependence of user service requests and service cache relationships, and optimizing task processing delay.

Benefits of technology

It achieves lower task processing delay and higher cache hit rate, and makes resource allocation more reasonable, adapts to dynamic network environments and meets user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116137724B_ABST
    Figure CN116137724B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of wireless communication technologies, and particularly relates to a task offloading and resource allocation method based on mobile edge computing; the method includes: constructing a mobile edge computing system model; constructing a service caching model and a service assignment model based on the mobile edge computing system model; establishing task offloading and resource allocation constraint conditions based on the service caching model and the service assignment model; constructing a task offloading and resource allocation joint optimization problem with the goal of minimizing the task processing delay according to the task offloading and resource allocation constraint conditions; using the DSRA algorithm to solve the task offloading and resource allocation joint optimization problem to obtain a task offloading and resource allocation strategy; the present invention can achieve low latency and high cache hit rate, and realize on-demand allocation of resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of wireless communication, and particularly relates to a task offloading and resource allocation method based on mobile edge computing. Background Technique

[0002] With the rapid development of the Internet of Things and the explosive growth of intelligent mobile devices (MDs), new applications featuring big data and intelligence have emerged continuously (such as online games, virtual reality (VR), augmented reality (AR), remote medical treatment, etc.), and these application services usually have the characteristics of being computationally intensive and latency-sensitive. However, limited by the volume, computing power, storage capacity, and battery power of mobile devices, MDs usually have problems such as insufficient computing power, large latency, and low battery life when processing high-energy-consuming and high-complexity computing tasks. Mobile Edge Computing (MEC) has been proposed as an advanced computing method to achieve the vision of ultra-large capacity, ultra-low latency, ultra-high bandwidth, and low-energy-consuming data processing at the network edge. MEC sinks resources such as computing power and storage in the cloud center to the network edge and drives users to offload computing tasks to the network edge to enjoy a high-performance computing service experience.

[0003] Deep Reinforcement Learning (DRL) combines the perception ability of deep learning and the decision-making ability of reinforcement learning and can effectively handle various decision-making problems in the MEC system. For example, a resource management method for computing deep reinforcement learning in a vehicle multi-access edge computing in the prior art studied the joint allocation problem of spectrum, computing, and storage resources in the MEC vehicle network, and used DDPG and hierarchical learning to achieve rapid resource allocation, meeting the quality of service requirements of vehicle applications. A dynamic computing offloading and resource allocation method based on deep reinforcement learning in a cache-aided mobile edge computing system studied the problems of dynamic caching, computing offloading, and resource allocation in the cache-aided MEC system and proposed an intelligent dynamic scheduling strategy based on DRL. However, the above methods all adopt the deep reinforcement learning algorithm of a single agent. The deep reinforcement learning algorithm of a single agent requires the environment to be stable, while the real network environment is often dynamically changing and the environment is unstable, which is not conducive to convergence. At the same time, it will also make techniques such as experience replay unable to be directly used.

[0004] Therefore, in the future edge network where the network structure is becoming increasingly dense and heterogeneous and resource deployment is decentralized, it is of great significance to design and implement a more dynamic and flexible distributed computing offloading and resource allocation strategy. At the same time, considering the impact of characteristics such as the partial observability of the network environment and the time dependence of service requests on network service orchestration and computing network resource allocation, the task offloading and multi-dimensional resource allocation problems in the decentralized MEC scenario have important research value. Summary of the Invention

[0005] Aiming at the deficiencies of the existing technology, the present invention proposes a task offloading and resource allocation method based on mobile edge computing, and the method includes:

[0006] S1: Construct a mobile edge computing system model;

[0007] S2: Construct a service caching model and a service assignment model based on the mobile edge computing system model;

[0008] S3: Based on the service caching model and the service assignment model, establish task offloading and resource allocation constraint conditions;

[0009] S4: According to the task offloading and resource allocation constraint conditions, construct a task offloading and resource allocation joint optimization problem with the goal of minimizing task processing delay;

[0010] S5: Use the DSRA algorithm to solve the task offloading and resource allocation joint optimization problem to obtain a task offloading and resource allocation strategy.

[0011] Preferably, step S1 specifically includes: constructing a mobile edge computing system model, including M base stations BS, and the base station set is expressed as Each base station is equipped with an MEC server; for the base station it has N m user equipment MD, and the user set is expressed as The system runs in discrete time slots, and the time set T = {0, 1, 2,...} is defined; for a user under the base station BS m the computing-intensive task generated in time slot t (t ∈ T) is defined as where, represents the data volume size of the task, represents the maximum tolerable delay of the task, represents the number of CPU cycles required to process a unit bit of the task, represents the service type required to process the task; all tasks generated by users under the base station BS are expressed as m

[0012] Preferably, the construction of the service caching model in step S2 specifically includes: defining the service type set as Let a k,m (t) ∈ {0, 1} represent the caching indication function of service k in the BS at time slot t. m a k,m (t) = 1 indicates that service k is cached in the BS, m otherwise the BS m will not cache service k; the service caching policy set of the base station BS m at time slot t is represented as a m (t) = {a 1,m (t), …, a k,m (t), …, a K,m (t)}.

[0013] Preferably, the construction of the service assignment model in step S2 specifically includes: for any user There are four task processing methods, and different task processing methods have different processing delays; the four task processing methods are: local computing, offloading to the associated BS m for processing, forwarding the offloaded task to other BSs for processing through the associated base station, and offloading to the cloud center for processing.

[0014] Furthermore, the task processing delay of the user is expressed as:

[0015]

[0016] Among them, represents the task processing delay of user m under the base station BS at time slot t, represents the task processing delay when the user performs local computing, represents the transmission delay when the task is offloaded to the associated base station, represents the delay of the associated base station in processing the task, T tr,m (t) represents the delay of the task being forwarded by the associated base station, represents the delay of other base stations in processing the task, T m,c (t) represents the transmission delay of the task being forwarded to the cloud center through the associated base station, represents the local task processing policy, represents the policy of offloading the task to the associated base station for processing, represents the policy of offloading the task to other base stations for processing, represents the policy of offloading the task to the cloud center for processing.

[0017] Preferably, the joint optimization problem of task offloading and resource allocation is expressed as:

[0018]

[0019] Among them, T represents the system running time, M represents the number of base stations, represents the user m under the base station BS task processing delay at time slot t, a(t) represents the base station service caching policy, b(t) represents the task offloading policy, α(t) represents the spectrum resource allocation policy, β(t) represents the base station computing power resource allocation policy, N m represents the number of user equipment under the m-th base station, represents the base station BS m under the user maximum tolerable task delay, represents the user local task processing policy, represents the user policy for offloading tasks to associated base stations for processing, represents the user policy for offloading tasks to other base stations for processing, represents the user policy for offloading tasks to the cloud center for processing, a k,m (t) represents the caching indication function of the m-th base station BS m for service k at time slot t, K represents the number of service types, l k represents the storage space size occupied by service k for processing tasks, R m represents the storage space size of the m-th MEC server, represents BS m assigned to at time slot t represents BS m assigned to CPU frequency allocation coefficient at time slot t.

[0020] Preferably, the process of using the DSRA algorithm to solve the joint optimization problem of task offloading and resource allocation includes: abstracting the joint optimization problem of task offloading and resource allocation into a partially observable Markov decision process, with the base station acting as the agent, and constructing the corresponding observation space, action space, and reward function; each agent has an actor network and a critic network embedded with an LSTM network; the actor network generates corresponding actions based on the current local observation state of a single agent and updates the reward function according to the actions, entering the next state; the critic network estimates the strategies of other agents based on the global observation state and actions; generates experience information according to the current state, next state, actions, and reward values; samples multiple pieces of experience information to train the actor network and the critic network, updates the network parameters, and obtains the trained actor network and critic network; obtains the task offloading and resource allocation strategy according to the training results of the actor network.

[0021] Further, the reward function is expressed as:

[0022]

[0023] where r m (t) represents the reward value of the base station BS at time slot t, T represents the system running time, M represents the number of base stations, N m represents the number of user equipments under the m-th base station, m represents the task processing delay of the user under the base station BS at time slot t, m Y (t) represents the reward when the task processing delay meets the delay constraint, and U m (t) represents the reward when the cache does not exceed the storage capacity limit of the edge server. m

[0024] The beneficial effects of the present invention are as follows: The present invention aims at the service orchestration and computing and network resource allocation problems in the decentralized MEC scenario, with the goal of minimizing the task processing delay, and proposes a task offloading and resource allocation method based on mobile edge computing; considering the time dependence of user service requests and the coupling relationship between service requests and service caches, an LSTM network is introduced to extract historical state information about service requests, enabling users to make better decisions by learning this historical information. Through simulation experiments, this method can achieve lower latency and higher cache hit rates, realizing on-demand allocation of resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is a flowchart of the task offloading and resource allocation method based on mobile edge computing in the present invention;

[0026] Figure 2 Schematic diagram of the mobile edge computing system model in the present invention;

[0027] Figure 3 Block diagram of the DSRA algorithm in the present invention;

[0028] Figure 4 Graph showing the change process of the average delay of the DSRA algorithm and the comparison algorithm in the present invention with the number of training iterations;

[0029] Figure 5 Graph showing the change process of the average cache hit rate of the DSRA algorithm and the comparison algorithm in the present invention with the number of training iterations. Detailed implementation manner

[0030] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0031] The present invention proposes a task offloading and resource allocation method based on mobile edge computing, as Figure 1 shown, the method includes the following contents:

[0032] S1: Construct a mobile edge computing system model.

[0033] As Figure 2 shown, the present invention considers a typical MEC system, which includes M base stations (BaseStation, BS), and the base station set is defined as Each BS is configured with an MEC server with certain computing and storage resources; under the mth base station there are N m user equipment MDs, and the user set under the mth base station is represented as The system operates in discrete time slots, and the time set is defined as For the ith user device under BS m the computation-intensive task generated in time slot t is defined as wherein, represents the data volume size of the task, with the unit of bit; represents the maximum tolerable delay of the task, represents the number of CPU cycles required to process a unit-bit task; represents the service type required to process the task. Then the tasks generated by all users under BS m are represented as

[0034] S2: Build a service caching model and a service assignment model based on the mobile edge computing system model.

[0035] Specifically, building the service caching model includes:

[0036] In the present invention, a service refers to a specific program or data required to run various types of tasks (such as games, virtual / augmented reality). At any time slot, only the MEC server that caches the corresponding service can provide computing services for the offloading tasks of the MD. Assume that there are a total of K different types of services in the network, and define the service type set as Let a k,m (t) ∈ {0, 1} represent the caching indication function of the BS m for service k at time slot t. a k,m (t) = 1 means that service k is cached in the BS m , otherwise the BS m will not cache service k; the service caching policy set of the base station BS m at time slot t is represented as a m (t) = {a 1,m (t), …, a k,m (t), …, a K,m (t)}.

[0037] Specifically, building the service assignment model includes:

[0038] If the BS m caches the service type required to process the task, then the task can be processed by the BS . Otherwise, the task can only be processed locally on the device or offloaded to other servers. For any m there are four task processing methods, and different task processing methods have different processing delays; the four task processing methods are: 1) local computing; 2) offloading to the associated BS for processing; 3) forwarding the offloaded task to other BSs for processing through the associated base station; 4) offloading to the cloud center for processing. Let m represent the task offloading policy of at time slot t. Among them, represents the local task processing policy of , indicating that the task can be processed locally. Similarly, represents the policy of offloading the task to the associated base station for processing, represents the policy of offloading the task to the neighboring base station for processing, represents the policy of offloading the task to the cloud center for processing, Represents the policy for task offloading to the cloud center for processing; at time slot t, base station BS m The task offloading policies of all users under BS

[0039] 1) The task is computed locally

[0040] When the task is processed locally, that is Let Represent The local CPU frequency of, then the processing time of the task locally can be expressed as Represents the data volume size of the task, in bits, Represents the number of CPU cycles required to process a unit-bit task.

[0041] 2) The task is offloaded to the associated base station for processing

[0042] If The associated base station BS of m Caches service k, then The task of can be offloaded to BS through the wireless link m For processing, that is According to Shannon's formula, from To BS m The uplink transmission rate of is Among them, B m Is the bandwidth of BS m Of, Is BS m At time slot t, the spectrum resource allocation coefficient assigned to Satisfies Is BS m Allocated to The bandwidth of, then BS m The spectrum resource allocation policy can be expressed as Represent The transmission power of, Represent And BS m The channel gain between, σ 2 (t) represents the additive white Gaussian noise power at time slot t. Then the transmission delay of the task is

[0043] BS m The time for BS to process the task is Among them, f m Represents the CPU frequency of BS m Of, Is BS m At time slot t, the CPU frequency allocation coefficient assigned to Satisfies Indicates BS m Allocated to CPU frequency, then BS m The computing power resource allocation strategy can be expressed as The processing result of the task is usually much smaller than the uploaded data, and the present invention ignores the time delay of result transmission back.

[0044] As can be seen from the above analysis, The task of is offloaded to the associated base station BS m The time delay for processing is

[0045] 3) The task is migrated to the nearby base station for processing

[0046] If The associated base station BS of m Does not cache service k, but the nearby base station BS n (n ∈ {1, 2,..., M} and n ≠ m) Caches service k, then The task of can be forwarded by the associated base station BS m To be forwarded to other nearby base stations BS n For processing, that is At time slot t, the transmission rate of the task from the associated base station to the nearby base station is Among them, ω m Is the bandwidth when the base station m forwards the task, P m Is the forwarding power of the base station m, G m,n Is the channel gain between the base station m and the base station n, then the time for the task to be forwarded by the associated base station is:

[0047] As can be seen from the above analysis, BS n The time for processing the task is Therefore, the task is forwarded to BS n The computing offloading time delay for processing is

[0048] 4) The task is offloaded to the cloud center for processing

[0049] If The associated base station BS of m Does not cache the relevant service for processing this task, then this task can also be forwarded by the associated base station BS m To the cloud center for processing, that is The cloud center has rich computing resources and storage resources, and the present invention ignores the task processing time and result transmission back time of the cloud center.

[0050] The task of passes through the associated base station BSm The computing offloading time forwarded to the cloud center is where r m,c (t) is the transmission rate of the BS m forwarding the task to the cloud center. The delay for the task to be offloaded to the cloud center for processing is

[0051] To sum up, at time slot t, the task processing delay of the user is expressed as:

[0052]

[0053] where represents the task processing delay of the user m under the BS at time slot t, represents the task processing delay of the user m under the BS when performing local computing, represents the transmission delay of the user m under the BS offloading the task to the associated BS, represents the delay for the associated BS to process the task, T tr,m (t) represents the delay for the task to be forwarded by the associated BS, represents the delay for other BSs to process the task, T m,c (t) represents the transmission delay of the task of the user m under the BS being forwarded to the cloud center through the associated BS.

[0054] S3: Based on the service cache model and service assignment model, establish the task offloading and resource allocation constraints.

[0055] The storage space of the MEC server is limited, and the storage space occupied by the cached services cannot exceed the storage capacity of the MEC server. Define the size of the storage space of the m-th MEC server MECm as Rm, then there is where l k represents the size of the storage space occupied by the service for processing this task.

[0056] At time slot t, it satisfies

[0057] The processing delay of the task cannot exceed the maximum tolerable delay:

[0058] The total sum of the allocated spectrum resources should not be greater than the base station bandwidth:

[0059] The total allocated computing resources should not be greater than the computing resources of the base station:

[0060] S4: According to the task offloading and resource allocation constraints, a joint optimization problem of task offloading and resource allocation is constructed with the goal of minimizing the task processing delay.

[0061] Limited by the resources of the server (such as computing, spectrum, and storage space), at the same time, task offloading and resource allocation are coupled. In view of this, with the goal of minimizing the long-term processing delay of tasks, the present invention establishes a joint optimization problem of service caching and computing network resource allocation, expressed as:

[0062]

[0063] Among them, T represents the system running time, M represents the number of base stations, represents the task processing delay of user at time slot t, a(t) = {a1(t), …, a M (t)} represents the service caching strategy of the base station, b(t) = {b1(t), …, b M (t)} represents the task offloading strategy, α(t) = {α1(t), …, α M (t)} represents the spectrum resource allocation strategy, β(t) = {β1(t), …, β M (t)} represents the computing power resource allocation strategy of the base station, N m represents the number of user equipment under the m-th base station, represents the user m under the base station BS at time slot t, represents the maximum tolerable delay of the task of user m under the base station BS at time slot t, represents the local task processing strategy of user whose task is offloaded to the associated base station for processing, represents the strategy for user whose task is offloaded to other base stations for processing, represents the strategy for user whose task is offloaded to the cloud center for processing, a k,m (t) represents the caching indication function of the m-th base station BS m for service k at time slot t, K represents the number of service types, l k represents the storage space size occupied by service k for processing the task, R m represents the storage space size of the m-th MEC server, represents BS m allocated to Spectrum resource allocation coefficient denotes BS m assigned to at time slot t CPU frequency allocation coefficient

[0064] S5: Use the DSRA algorithm to solve the joint optimization problem of task offloading and resource allocation, and obtain the task offloading and resource allocation strategy

[0065] In the edge network environment, characteristics such as the decentralization of computing and network resource deployment, the high dynamics of the network environment, and the increasing density of the network structure make the centralized management method unable to well cope with the highly dynamic decentralized MEC environment. It is necessary to design a more dynamic and flexible distributed computing offloading and resource allocation strategy. As a distributed DRL algorithm, multi-agent deep reinforcement learning can be well applied to problem solving in the decentralized MEC environment. In view of this, the present invention designs a distributed intelligent service orchestration and computing and network resource allocation algorithm (Distributed Service Arrangement and Resource Allocation Algorithm, DSRA), where the base station acts as an agent to learn the task offloading strategy, service caching strategy, and computing and network resource allocation strategy. At the same time, considering the time dependence of user service requests and the coupling relationship between service requests and service caching, the LSTM network is used to extract historical state information about service requests. By learning this historical information, the agent can better understand the future environmental state and thus make better decisions. As Figure 3 shown, it specifically includes the following contents

[0066] Abstract the joint optimization problem of task offloading and resource allocation into a partially observable Markov decision process (POMDP), with the base station acting as an agent, and construct the corresponding observation space, action space, and reward function; define the tuple to describe the above Markov game process, where represents the global state space, and the environment at time slot t is the global state is the set of the agent's observation space is the set of the global action space is the reward set. At time slot t, agent m takes the strategy according to the local observation to select the corresponding action and thus obtain the corresponding reward

[0067] 1) Environmental status

[0068] At time slot t, the agent can receive detailed task information of mobile devices within its coverage area, including the amount of task data, the maximum tolerable delay, the number of CPU cycles required to process a single-bit task, and the required service type. The environmental status can be defined as s(t) = {d1, d2, …, d M , P1, P2, …, P M , f1, f2, …, f M , B1, B2, …, B M , G1, G2, …, G M}, where represents the tasks generated by all users under BS m , f m represents the CPU frequency of BS m . is the set of transmission powers of all users under BS m . is the set of channel gains between all users under BS m and BS m . At time slot t, the environmental status observed by agent m is defined as follows:

[0069]

[0070] 2) Action space

[0071] Agent m selects a corresponding action from the action space according to the observed environmental status o m (t) and the current policy π m . At time slot t, the action of agent m is defined as follows:

[0072]

[0073] a 1,m (t), a 2,m (t), …, a K,m (t)}

[0074] Relax the binary variables a k,m (t), and to real-valued variables and a′ k,m (t) > 0.5 indicates that BS m caches service k, otherwise BS m will not cache service k. For and The task will select the offloading mode corresponding to the maximum value for computing offloading. According to the definition of the action space and the value range of each element in a m (t), it can be known that the action space is a continuous set.

[0075] 3) Reward function

[0076] The reward function measures the effect brought by an agent taking a certain action in a given state. During the training process, when the agent takes a certain action at time slot t - 1, the corresponding reward will be returned to the agent at time slot t. Based on the obtained reward, the agent will update its policy to obtain the optimal result. Since the reward causes each agent to reach its optimal policy, and the policy directly determines the computing and network resource allocation policy, computing offloading policy, and service caching policy of the corresponding MEC server, the reward function should be designed according to the original optimization problem. The reward function constructed in the present invention includes three parts: the first part is the reward for task processing time, the second part is the reward for the task processing delay satisfying the delay constraint, that is The third part is the reward for the cache not exceeding the storage capacity limit of the edge server, that is The optimization goal is to minimize the long-term processing delay of the task and maximize the long-term return. Therefore, the cumulative reward of agent m should be:

[0077]

[0078] where H(·) is the Heaviside step function; λ1 and λ2 respectively represent the first and second weight coefficients, Y m (t) represents the reward for the task processing delay satisfying the delay constraint, and U m (t) represents the reward for the cache not exceeding the storage capacity limit of the edge server.

[0079] Each base station has an actor network and a critic network embedded with an LSTM network. Both the actor network and the critic network include a current network and a target network. The framework of the DSRA algorithm consists of an environment and M agents, i.e., base stations. Each agent has a centralized training phase and a decentralized execution phase. During training, centralized learning is used to train the critic network and the actor network. When training the critic network, the state information of other agents needs to be used. During decentralized execution, the actor network only needs to know local information. That is, each agent will utilize the global state and actions during the training process to estimate the strategies of other agents, and adjust its local strategy according to the estimated strategies of other agents to achieve global optimality. The Multi-agent Deep Deterministic Policy Gradient (MADDPG) algorithm can handle the case where the environment is fully observable well. However, the real environment state is often partially observable. To cope with the partial observability of the environment and the time-dependence of service requests, the present invention incorporates the Long Short-Term Memory network (LSTM) into the actor network and the critic network. LSTM is a recurrent neural network that can extract historical state information about service requests. By learning this historical information, agents can better understand future states and make better decisions.

[0080] The actor network generates corresponding actions based on the current local observation state of a single agent. Specifically, the actor network obtains the current task offloading and resource allocation strategy according to the local observation state, and can generate corresponding actions from the action space according to the task offloading and resource allocation strategy. The agent enters the next state.

[0081] Update the reward function according to the action; generate experience information according to the current state, the next state, the action, and the reward value; sample multiple pieces of experience information to train the actor network and the critic network, and update the network parameters to obtain a trained actor network. Specifically, during the training process, let and respectively represent the historical information about service requests of the actor network and the critic network before and after taking the action, and use the experiences from the experience replay memory D to iteratively update the DSRA algorithm. The experience replay memory D of agent m contains a set of experience tuples, where o m (t) represents the observation state of agent m at time slot t, a m (t) represents the action taken by agent m at time slot t based on the current observation o m (t), r m (t) represents the reward obtained by agent m at time slot t when taking the action a mThe reward obtained after (t), o' m (t + 1) represents the state of agent m at time slot t + 1, represents the historical information of the actor network regarding the service request at time slot t, represents the historical information of the critic network regarding the service request at time slot t, represents the historical information of the actor network regarding the service request at time slot t + 1, represents the historical information of the critic network regarding the service request at time slot t + 1.

[0082] In the decentralized execution phase, at time slot t, the actor network of each agent selects an action according to the local observation state o m (t), the current historical state information and its own policy to select an action

[0083] In the centralized training phase, each critic network can obtain the observations o m (t) and actions a m (t) of other agents, then the Q - function of agent m can be expressed as

[0084] The Q - function evaluates the actions of the actor network from a global perspective and guides the actor network to select better actions. During training, the critic network updates the network parameters by minimizing the loss function, and the loss function is defined as follows:

[0085]

[0086] where γ is the discount factor. At the same time, the actor network updates the network parameters θ based on the centralized Q - function calculated by the critic network and its own observation information, and outputs the action a. The actor network parameters θ are updated by maximizing the policy gradient, that is:

[0087]

[0088]

[0089]

[0090]

[0091] The parameters of the target network are updated by soft - update, that is:

[0092] After the actor network is trained, the task offloading, service caching, and resource allocation strategies within the time period T can be obtained based on the actions taken by the actor network. Task offloading according to the task offloading and resource allocation strategies can minimize the total processing delay of tasks while meeting various constraints.

[0093] Evaluate the present invention:

[0094] Compare the present invention with the multi-agent deep deterministic policy gradient algorithm MADDPG (Multi-agent Deep Deterministic Policy Gradient), the single-agent deep deterministic policy gradient algorithm SADDPG (Single agent Deep Deterministic Policy Gradient), and the single-agent deep deterministic policy gradient algorithm TADPG based on LSTM. As Figure 4 shown, it can be seen that as the number of training episodes increases, the average processing delay of tasks continuously decreases and gradually stabilizes, eventually reaching convergence. The DSRA algorithm has the smallest delay, indicating that the DSRA algorithm can make better offloading and computing network resource allocation decisions, thus obtaining a smaller delay, realizing the on-demand allocation of resources, and proving the effectiveness of the algorithm. From Figure 5 it can be seen that as the number of episodes increases, the cache hit rate curve shows an upward trend and finally reaches convergence, and the DSRA has the largest cache hit rate, proving the effectiveness of the algorithm.

[0095] The above-mentioned embodiments further elaborate on the purpose, technical solution, and advantages of the present invention. It should be understood that the above-mentioned embodiments are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A task offloading and resource allocation method based on mobile edge computing, characterized in that, Including: S1: Construct a mobile edge computing system model; S2: Construct a service caching model and a service assignment model based on the mobile edge computing system model; S3: Establish task offloading and resource allocation constraint conditions based on the service caching model and the service assignment model; S4: According to the task offloading and resource allocation constraint conditions, construct a joint optimization problem of task offloading and resource allocation with the goal of minimizing the task processing delay; the joint optimization problem of task offloading and resource allocation is expressed as: Among them, T represents the system running time, and M represents the number of base stations. represents the task processing delay of user m under base station BS at time slot t, a(t) represents the base station service caching policy, b(t) represents the task offloading policy, α(t) represents the spectrum resource allocation policy, β(t) represents the base station computing power resource allocation policy, and N m represents the number of user equipment under the m-th base station. represents the maximum tolerable delay of the task of user m under base station BS at time slot t. represents the local task processing policy of user . represents the policy of offloading the task of user to the associated base station for processing. represents the policy of offloading the task of user to other base stations for processing. represents the policy of offloading the task of user to the cloud center for processing. a k,m (t) represents the caching indication function of the m-th base station BS m for service k at time slot t, K represents the number of service types, and l k represents the storage space size occupied by service k for processing the task, and R m represents the storage space size of the m-th MEC server. represents the spectrum resource allocation coefficient allocated by BS m to at time slot t. represents the CPU frequency allocation coefficient allocated by BS m to at time slot t. S5: Use the DSRA algorithm to solve the joint optimization problem of task offloading and resource allocation to obtain a task offloading and resource allocation strategy; the process of using the DSRA algorithm to solve the joint optimization problem of task offloading and resource allocation includes: abstracting the joint optimization problem of task offloading and resource allocation into a partially observable Markov decision process, with the base station acting as an agent, and constructing a corresponding observation space, action space, and reward function; each agent has an actor network and a critic network embedded with an LSTM network; the actor network generates corresponding actions according to the current local observation state of a single agent and updates the reward function according to the actions, entering the next state; the critic network estimates the strategies of other agents according to the global observation state and actions; generate experience information according to the current state, next state, actions, and reward values; sample multiple pieces of experience information to train the actor network and the critic network, update the network parameters, and obtain the trained actor network and critic network; obtain the task offloading and resource allocation strategy according to the training results of the actor network; the reward function is expressed as: where r m 9t) represents the reward value of the base station BS at time slot t m , Y m (t) represents the reward when the task processing delay meets the delay constraint, and U m (t) represents the reward when the cache does not exceed the storage capacity limit of the edge server.

2. The task offloading and resource allocation method based on mobile edge computing according to claim 1, characterized in that, Step S1 specifically includes: constructing a mobile edge computing system model, which includes M base stations BS, and the base station set is represented as Each base station is equipped with an MEC server; for the base station There are N m user devices MD under it, and the user set is represented as The system runs in discrete time slots, and the time set T = {0, 1, 2,...} is defined; for the base station BS m A user under The computationally intensive task generated in time slot t (t ∈ T) is defined as where represents the data volume size of the task, represents the maximum tolerable delay of the task, represents the number of CPU cycles required to process a unit-bit task, represents the service type required to process the task; all tasks generated by users under the base station BS m are represented as 3. The task offloading and resource allocation method based on mobile edge computing according to claim 1, characterized in that The construction of the service cache model in step S2 specifically includes: defining the service type set as Let a k,m (t) ∈ {0, 1} represent the cache indication function of service k in the BS at time slot t. m a k,m (t) = 1 indicates that service k is cached in the BS. m Otherwise, the BS m will not cache service k; the service cache policy set of the base station BS m at time slot t is represented as a m (t) = {a 1,m (t), …, a k,m (t), …, a K,m (t)}.

4. A task offloading and resource allocation method based on mobile edge computing according to claim 1, characterized in that The specific construction of the service assignment model in step S2 includes: for any user There are four task processing methods, and different task processing methods have different processing delays; the four task processing methods are: local computing, offloading to the associated BS m for processing, forwarding the offloaded task to other BSs through the associated base station for processing, and offloading to the cloud center for processing.

5. A task offloading and resource allocation method based on mobile edge computing according to claim 4, characterized in that, The task processing delay of the user is expressed as: Among them, represents the task processing delay of user m under the base station BS at time slot t, and represents the task processing delay when the user performs local computing, represents the transmission delay for offloading the task to the associated base station, represents the delay for the associated base station to process the task, and T tr,m (t) represents the delay for the associated base station to forward the task, represents the delay for other base stations to process the task, and T m,c (t) represents the transmission delay for the task to be forwarded to the cloud center through the associated base station, represents the local task processing strategy, represents the strategy for offloading the task to the associated base station for processing, represents the strategy for offloading the task to other base stations for processing, represents the strategy for offloading the task to the cloud center for processing.