A method for modeling an optimization model of a mobile edge computing system and an optimization method thereof

By constructing a mobile edge computing system and optimizing power allocation and computation offloading using the D3T3 framework, the uncertainty of computing resource allocation in multi-edge server and multi-mobile device environments is solved, thereby reducing task costs and improving system stability.

CN119233317BActive Publication Date: 2025-12-19GUANGZHOU HUASU INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411217273.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-02
Publication Date
2025-12-19
Estimated Expiration
2044-09-02

AI Technical Summary

Technical Problem

Existing technologies cannot effectively solve the problems of computation offloading and resource allocation in mobile edge computing systems with multiple edge servers and multiple mobile devices, resulting in uncertainty and instability in computation resource allocation, which affects the spatiotemporal heterogeneous requirements for task completion.

Method used

A mobile edge computing system is constructed, consisting of N edge servers and M mobile devices that interact using time-division multiplexing. An optimization problem model is built by combining channel gain, signal-to-noise ratio, data transmission rate, total latency, and energy consumption. Power allocation, computation offloading, and resource allocation are optimized through the D3T3 framework, and decision optimization is performed using D3QN and TD3 algorithms.

Benefits of technology

In multi-edge server and multi-mobile device environments, it achieves reduced computational accuracy and task costs, and improves system stability and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119233317B_ABST
    Figure CN119233317B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of mobile edge computing, and discloses a mobile edge computing system, an optimization model modeling method thereof and an optimization method thereof, which comprises the following specific steps: a mobile edge computing system composed of N edge servers and M mobile devices is constructed, each mobile device alternately interacts with the environment in a time division multiplexing mode; the mobile devices are randomly distributed in a closed area and randomly move in the area; and an optimization problem model is constructed by comprehensively considering the channel gain between the edge servers and the mobile devices, the signal-to-noise ratio of the edge servers and the mobile devices, the data transmission rate of the edge servers and the mobile devices, the total delay and total energy consumption of edge computing, and the task cost of edge computing. The application solves the problem that the prior art is not applicable to the environment with multiple edge servers and multiple mobile devices, and has the characteristics of accurate calculation and effective reduction of the task cost.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of mobile edge computing, and more particularly to a mobile edge computing system, an optimization model modeling method thereof, and an optimization method thereof. BACKGROUND

[0002] With the rapid development of emerging technologies such as the Internet of Things, artificial intelligence, and augmented reality, people's demand for real-time, low latency, and high bandwidth is increasing. However, the traditional cloud computing model has certain limitations in meeting these demands, mainly in terms of large data transmission latency, network congestion, privacy security, and the like. Therefore, mobile edge computing (MEC) has emerged as a new computing model. Its core idea is to push computing resources and services to the network edge closer to users, thereby achieving lower network latency, higher bandwidth utilization, and better user experience. Specifically, the traditional cloud computing model deploys data processing and services in remote data centers, resulting in the problem of large data transmission latency, especially in application scenarios such as intelligent transportation and virtual reality that require real-time and low latency. The emergence of MEC provides new ideas and possibilities for solving this problem. By deploying applications and services on edge devices such as base stations and edge servers, MEC can provide data processing and services closer to the edge of the user network, greatly reducing the distance and time of data transmission, and achieving more immediate and efficient data processing and service response. Thus, the quality of service (QoS) is improved. However, in a MEC system with multiple mobile devices and multiple edge servers, the frequent movement of users can cause high-speed changes in system state and uneven task completion requirements in time and space. In fact, the location of the edge server is fixed in most cases, i.e., the computing resources are determined in time and space. This may cause uncertainty and instability in edge server computing offloading and computing resource allocation when facing spatio-temporal heterogeneous task requirements. For example, there are a large number of users around a certain edge server who need to complete tasks at a certain time, while the total amount of computing resources of the edge server is relatively fixed, so it is impossible to completely meet the needs of all users. At this time, the problems of computing offloading decision and computing resource allocation arise. Power allocation is also one of the factors affecting task latency and needs to be handled together with computing offloading and computing resource allocation. In addition, tasks need to be completed within a maximum tolerable latency, which further increases the complexity of the problem.

[0003] It is a challenge to design an effective solution to jointly handle the power allocation and computation offloading and computation resource allocation problems in MEC systems. In recent years, with the continuous development and innovation of machine learning (ML), learning-based algorithms provide new possibilities for solving the problems of MEC systems, because ML can more effectively optimize difficult problems. In existing learning frameworks, DRL as a new branch of ML is the most promising choice to solve this problem, because DRL can extract hidden information from a large number of complex states and use valuable information therein. Since the joint optimization problem involves discrete actions and continuous actions, one approach is to use different DRL frameworks to handle different types of actions separately and then combine them together. A common DRL framework based on discrete actions is deep Q network (DQN), which can effectively optimize computation offloading decisions. Another DRL framework based on continuous actions is deep deterministic policy gradient (DDPG), which can optimize computation resource allocation decisions. The combination of the two can effectively solve the joint optimization problem.

[0004] The prior art has a mobile edge computing task offloading method for cloud-edge-end collaboration in the metaverse. The method comprises: calculating total time delay consumption and total energy consumption consumed in all mobile edge computing task offloading processes in a cloud-edge-end collaborative system according to time delay consumption and energy consumption of user terminal devices, edge servers and cloud servers; establishing a target function for mobile edge computing offloading in the metaverse with the goal of minimizing the total time delay consumption and total energy consumption; and using an improved cloud-edge-end collaborative network computing offloading algorithm to solve the target function to obtain an optimal task offloading strategy of the user terminal device.

[0005] However, the prior art only considers a single-edge server and multiple mobile device environment, and has problems that are not applicable to a multi-edge server and multi-mobile device environment, so how to invent a mobile edge computing method that adapts to a multi-edge server and multi-mobile device environment is a technical problem that needs to be solved in the technical field. SUMMARY

[0006] To solve the problem that the prior art is not applicable to a multi-edge server and multi-mobile device environment, the present application provides a mobile edge computing system, an optimization model modeling method thereof and an optimization method thereof, which have the characteristics of accurate calculation and effective reduction of task cost.

[0007] To achieve the above-mentioned purposes of the present application, the technical solutions adopted are as follows:

[0008] A mobile edge computing system is composed of N edge servers and M mobile devices, each of which interacts with the environment in turn using time division multiplexing; the mobile devices are randomly distributed in a closed area and move randomly within the area; the edge servers ={1, 2,..., N} are uniformly distributed in a fixed area, each edge server is equipped with a computing unit ; the mobile devices ={1, 2,..., M} are randomly positioned and move randomly in a fixed area, each mobile device is equipped with a computing unit ; and further comprising a transmitting antenna with a maximum transmitting power .

[0009] A modeling method of an optimization model of a mobile edge computing system, comprising the following specific steps:

[0010] Constructing the mobile edge computing system;

[0011] Integrating the channel gain between the edge servers and the mobile devices, the signal-to-noise ratio of the edge servers and the mobile devices, the data transmission rate of the edge servers and the mobile devices, the total delay and total energy consumption of edge computing, and the task cost of edge computing to construct an optimization problem model.

[0012] Preferably, each mobile device interacts with the environment in turn using time division multiplexing, specifically: the system time is uniformly divided into T discrete time slots, denoted as T={1, 2,..., T}, and each time slot has a length of ζ seconds; in time slot t, the system uses time division multiplexing to let each mobile device interact with the system environment in turn, and each mobile device in each time slot has a task arriving, the mobile device can perform either local computing task or task offloading to edge server for computing, if the task is offloaded to edge server for computing, the task will occupy the computing resources allocated by ES to MD until the end of the time slot; each mobile node moves randomly within the time slot, and the moving speed and direction follow a mobile model.

[0013] Further, the mobile model is specifically:

[0014]

[0015] wherein, is the maximum moving speed, L is the side length of the closed square area, , represents the velocity vector of the mth mobile device, represents the velocity of the mobile device, represents the moving direction of the mobile device, and respectively represent the coordinates of the mth mobile device.

[0016] Further, the channel gain between the edge server and the mobile device, the signal-to-noise ratio of the edge server and the mobile device, the data transmission rate of the edge server and the mobile device, the total delay and total energy consumption of edge computing, and the task cost of edge computing are integrated to construct an optimization problem model, and the specific steps are as follows:

[0017] The computing task generated by the mobile device m at time slot t is represented as , representing the task data size, representing the number of cycles required to process one bit of data, representing the maximum allowable delay, representing the value of the task;

[0018] The channel gain between the edge server n and the mobile device m at time slot t is represented as:

[0019]

[0020] wherein represents the antenna gain, represents the speed of light, f represents the carrier frequency, and d represents the distance between the edge server n and the mobile device m;

[0021] The transmit power of the mobile device m at time slot t is defined as , and the signal-to-noise ratio of the mobile device m to the edge server n is represented as:

[0022]

[0023] wherein is the additional noise power;

[0024] The data transmission rate of the mobile device m to the edge server n is represented as:

[0025]

[0026] wherein is the bandwidth of the wireless channel;

[0027] At the beginning of each time slot t, each mobile device will be connected to its nearest edge server; the computing offloading strategy of the connected edge server and the mobile device m at time slot t can be represented as and ; when , the mobile device m selects to compute the task locally, when , the mobile device m selects to offload the task to the connected edge server n; when , and , the mobile device m chooses to offload the task to ES n, the connected edge server n' helps the MD to forward the task;

[0028] When , the local computation latency is denoted as:

[0029]

[0030] The energy consumption of the local computation is denoted as:

[0031]

[0032]

[0033] where is the power of the mobile device m, ξ is the energy consumption coefficient;

[0034] When , and , the mobile device m chooses to connect the edge server n' and offload the task to the edge server n to compute the task; at this time, the transmission latency is denoted as:

[0035]

[0036] where I is the indicator function, is the data transmission rate between edge servers;

[0037] The energy consumption of the offloading is denoted as:

[0038]

[0039] The computation latency of the task on the edge server n is denoted as:

[0040]

[0041] where is the computation resource of the edge server n allocated to the MD, the capability consumption is denoted as:

[0042]

[0043]

[0044] where is the power of the ES n;

[0045] The total latency and total energy consumption of the edge computation at the time slot t are denoted as:

[0046]

[0047]

[0048] The task cost is defined as the weighted sum of the task delay and the energy consumption, and the task cost of the mobile device m at the time slot t is represented as:

[0049]

[0050]

[0051] wherein and are the weights of the task delay cost and the energy consumption cost, respectively, and satisfy the condition:

[0052]

[0053] Based on the above, the optimization problem model P is constructed:

[0054]

[0055] wherein and represent the minimum allocable computing resource and the maximum allocable computing resource of the ES n, P m MD represents the power of the mth mobile device, represents the task delay of the mobile device m at the time slot t:

[0056] .

[0057] The optimization method of the power allocation of the mobile edge computing system comprises the following specific steps:

[0058] An optimization problem model is constructed by the modeling method;

[0059] The task cost related to the power is defined as ; when , and t are determined, the is represented as:

[0060]

[0061]

[0062]

[0063] wherein is the offloading transmission delay, is the data transmission rate between edge servers, and are the weights of task delay cost and energy cost respectively, is the energy consumption of local computing, n and n' represent different edge servers respectively;

[0064] According to the optimization problem model, the power optimization sub-problem is constructed:

[0065]

[0066] where, P MD is the power of mobile device, is the edge server delay;

[0067] When , is convex;

[0068] Take the derivative of :

[0069]

[0070] When is minimum, , , we can get:

[0071]

[0072]

[0073]

[0074]

[0075]

[0076]

[0077]

[0078] Introducing the Lambert W function, we get

[0079]

[0080]

[0081]

[0082] According to the constraint condition of the power optimization sub-problem, the optimal power is obtained:

[0083]

[0084] wherein is:

[0085] .

[0086] The optimization method of the mobile edge computing system's computing offloading and computing resource allocation includes the following specific steps:

[0087] An optimization problem model is constructed by the modeling method;

[0088] According to the optimization problem model, a computing offloading and computing resource allocation sub-problem P2 is constructed:

[0089]

[0090] A Markov decision process composed of four tuples is constructed , wherein is the current state space, a is the current action space, r is the reward, is the next state space; In the Markov decision process, at the beginning of each time slot, the current state of the system environment is observed and interacted with through actions, then the system environment enters the next state and accepts the reward, and so on;

[0091] Solving the computing offloading and computing resource allocation sub-problem, the optimal allocation scheme of power allocation, computing offloading and computing resource allocation is obtained.

[0092] Preferably, the Markov decision Specifically:

[0093] In the state space, from time slot t, each edge server independently observes the task attributes sent by the connected mobile devices and the remaining resources of all ESs, , wherein is a vector; denotes the distance between the MD and the connected ES, denotes the remaining computing resources;

[0094] The action space includes the optimization variables of P2, specifically including the task offloading and computing resource allocation decisions of all MDs, ;

[0095] After taking the action, the agent immediately obtains the reward, and the reward function adopted is specifically:

[0096] .

[0097] Further, when solving the sub-problems of computing offloading and computing resource allocation to obtain the optimal allocation scheme of power allocation, computing offloading and computing resource allocation, a D3T3 framework including duel double deep Q network D3QN and double delayed deep deterministic policy gradient TD3 is adopted, and the specific steps are as follows:

[0098] Initialize the environment experience replay buffer of the Q network D3QN and the environment experience replay buffer of the double delayed deep deterministic policy gradient TD3 ;

[0099] For each mobile device:

[0100] Generate a task;

[0101] The edge server connected by the mobile device observes the current system state s;

[0102] The agent on the edge server obtains a computing offloading policy action b by inputting s into the state-action value network Q of the D3QN according to -greedy;

[0103] If b is equal to 0, the system environment directly obtains a reward r according to b, and updates the system state to obtain s′ ;

[0104] If b is not equal to 0, s and b are spliced to obtain s, a computing resource action f is obtained according to A( s) + N(0, σ1), N is noise; the system environment obtains a reward r according to b and f; f changes the remaining computing resource of the edge server, and the system environment state changes according to b and f to obtain s′ ; the agent on the edge server obtains s′ using Q( b′ ); s′ and b′ are spliced to obtain s′; the four-tuple ( s , f, r, s ′ ) is stored into ; the four-tuple ( s, b, r, s′ ) is stored into ; and the experience ( s, b, r, s′ ), ( s , f, r, s ′ ) is obtained;

[0105] Wherein:

[0106]

[0107]

[0108]

[0109]

[0110]

[0111]

[0112]

[0113]

[0114]

[0115] When the experience replay buffer stores a set number of experiences, the D3T3 framework is trained, and the optimal allocation scheme of power allocation and computation offloading and computation resource allocation is obtained by solving the calculation offloading and computation resource allocation sub-problems through the trained D3T3 framework.

[0116] Further, the D3T3 framework is trained, and the specific steps are as follows:

[0117] Initialize 8 networks, including 4 online networks and 4 target networks; the online networks include the state-action value network Q of D3QN, the actor network A of TD3, and two independent critic networks C1 and C2; the target networks include four networks Q', A', C1', and C2' which have the same structure as the online networks;

[0118] For each time slot, loop:

[0119] Obtain experiences (s, a, r, s') from the experience replay buffer; s, b, r, s′ , ( s , f, r, s ′ );

[0120] Store the experiences in the respective buffer pools;

[0121] If the time slot t is in a multiple relationship with 1, then: tf Randomly take out n experiences from the buffer pool

[0122] K ( K )= s, b, r, s′ ( ) from the buffer pool; .Sample K

[0123] ​​According to the formula y = r + γ 1 Q′ ( s′, arg max b Q ( s′,b )) Calculate the target value y;

[0124] Update the parameters of Q based on the target value y;

[0125] If time slot t is a multiple of 2, then copy the parameters of Q to Q′;

[0126] If time slot t and tf If the numbers are multiples of each other, then:

[0127] K experiences are randomly selected from the buffer pool. K ( s ,f,r, s ′ ) = .Sample ( K) ;

[0128] According to the formula y = r + γ 2min i=1,2 C i ′ ( s ′,A′ ( s ′ )+ N (0 ,σ 2) Calculate the target value y;

[0129] Update based on the target value y C 1, C 2 parameters;

[0130] If time slot t is a multiple of 2, then update using a soft update method. A , A′ , C 1 ′ and C 2 ′ Parameters;

[0131] The soft update process is as follows:

[0132]

[0133]

[0134] wherein theta is a neural network parameter, is a soft update coefficient.

[0135] The beneficial effects of the present application are as follows:

[0136] The present application constructs a mobile edge computing system composed of N edge servers and M mobile devices, and comprehensively constructs an optimization problem model of channel gain between edge servers and mobile devices, signal-to-noise ratio of edge servers and mobile devices, data transmission rate of edge servers and mobile devices, total delay and total energy consumption of edge computing, and task cost of edge computing. Thus, the present application provides a model that can perform calculation in a multi-edge server and multi-mobile device environment, solves the problem that the prior art is not applicable to a multi-edge server and multi-mobile device environment, and has accurate calculation and can effectively reduce the task cost. BRIEF DESCRIPTION OF DRAWINGS

[0137] Figure 1 is a flowchart of the modeling method of the optimization model of the mobile edge computing system.

[0138] Figure 2 is a schematic diagram of a mobile edge computing system.

[0139] Figure 3 is a schematic diagram of the interaction of the D3T3 framework and the environment.

[0140] Figure 4 is a learning schematic diagram of the D3QN network.

[0141] Figure 5 is a learning schematic diagram of the TD3 network.

[0142] Figure 6 is a performance comparison schematic diagram of D3T3 under the optimized power and the maximum power.

[0143] Figure 7 is a comparison diagram of the reward of the comparative experiment in Example 4 under different methods.

[0144] Figure 8 is a reward comparison diagram of the comparative experiment in Example 4 under different numbers of mobile devices.

[0145] Figure 9 is an average TCR comparison diagram of the comparative experiment in Example 4 under different methods.

[0146] Figure 10 is an average RCR comparison diagram of the comparative experiment in Example 4 under different methods. DETAILED DESCRIPTION

[0147] The present application will be described in detail below in conjunction with the drawings and specific embodiments.

[0148] Example 1

[0149] like Figure 2 As shown, the mobile edge computing system consists of N edge servers and M mobile devices. Each mobile device interacts with the environment in turn using time-division multiplexing. The mobile devices are randomly distributed in a closed area and move randomly within the area. In this system, the edge servers... ={1, 2, ..., N} are evenly distributed within a fixed area, and each edge server is equipped with a computing unit. ;mobile device ={1, 2, ..., M} are randomly positioned and moved within a fixed area, and each mobile device is equipped with a computing unit. It also includes a maximum transmission power of The transmitting antenna.

[0150] Example 2

[0151] like Figure 1 As shown, the modeling method for the optimization model of a mobile edge computing system includes the following specific steps:

[0152] Construct a mobile edge computing system consisting of N edge servers and M mobile devices. Each mobile device takes turns interacting with the environment using time-division multiplexing. The mobile devices are randomly distributed in a closed area and move randomly within the area.

[0153] An optimization problem model is constructed by considering the channel gain between edge servers and mobile devices, the signal-to-noise ratio between edge servers and mobile devices, the data transmission rate between edge servers and mobile devices, the total latency and total energy consumption of edge computing, and the task cost of edge computing.

[0154] In one specific embodiment, each mobile device interacts with the environment in turn using time-division multiplexing. Specifically, the system time is uniformly divided into T discrete time slots, denoted as T={1, 2, ..., T}, with each time slot having a length of ζ seconds. In time slot t, the system uses time-division multiplexing to allow each mobile device to interact with the system environment in turn. Assume that each mobile device in each time slot has a task arriving. The mobile device can perform either local computation or offload the task to an edge server for computation. If the task is offloaded to the edge server for computation, the task will occupy the computing resources allocated to the MD by ES until the time slot ends. Each mobile node moves randomly within the time slot, and the movement speed and direction follow the movement model.

[0155] In one specific embodiment, the moving model is as follows:

[0156]

[0157] wherein, is the maximum moving speed, L is the side length of the closed square region, , denotes the speed vector of the mth mobile device, denotes the speed of the mobile device, denotes the moving direction of the mobile device, and denote the coordinates of the mth mobile device, respectively.

[0158] In one specific embodiment, the channel gain between the edge server and the mobile device, the signal-to-noise ratio of the edge server and the mobile device, the data transmission rate of the edge server and the mobile device, the total delay and total energy consumption of edge computing, and the task cost of edge computing are integrated to construct an optimization problem model, and the specific steps are as follows:

[0159] The computing task generated by the mobile device m at time slot t is denoted as , denotes the task data size, denotes the number of cycles required to process one bit of data, denotes the maximum allowable delay, denotes the value of the task;

[0160] The channel gain between the edge server n and the mobile device m at time slot t is denoted as:

[0161]

[0162] wherein denotes the antenna gain, denotes the speed of light, f denotes the carrier frequency, and d denotes the distance between the edge server n and the mobile device m;

[0163] The transmit power of the mobile device m at time slot t is defined as , and the signal-to-noise ratio of the mobile device m to the edge server n is denoted as:

[0164]

[0165] wherein is the additional noise power;

[0166] The data transmission rate of the mobile device m to the edge server n is denoted as:

[0167]

[0168] wherein is the bandwidth of the wireless channel;

[0169] At the beginning of each time slot t, let each mobile device connect to its nearest edge server; the computing offloading strategy of the connected edge server and mobile device m at time slot t can be represented as and ; at , the mobile device m chooses to compute the task locally, when , the mobile device m chooses to offload the task to the connected edge server n; when , and , the mobile device m chooses to offload the task to the ES n, and the connected edge server n' will help the MD forward the task;

[0170] When , the local computing delay is represented as:

[0171]

[0172] The energy consumption of local computing is represented as:

[0173]

[0174]

[0175] where is the power of the mobile device m, ξ is the energy consumption coefficient;

[0176] When , and , the mobile device m chooses to connect the edge server n' and offload the task to the edge server n to compute the task; at this time, the transmission delay is represented as:

[0177]

[0178] where I is the indicator function, is the data transmission rate between edge servers;

[0179] The energy consumption of offloading is represented as:

[0180]

[0181] The computing delay of the task on the edge server n is represented as:

[0182]

[0183] where is the computing resource allocated by the edge server n to the MD, and the capability consumption is represented as:

[0184]

[0185]

[0186] wherein is the power of ES n;

[0187] The total latency and total energy consumption of edge computing at time slot t is represented as:

[0188]

[0189]

[0190] The task cost is defined as the weighted sum of the task latency and energy consumption, and the task cost of mobile device m at time slot t is represented as:

[0191]

[0192]

[0193] wherein and are the weights of the task latency cost and energy consumption cost respectively, satisfying the condition:

[0194]

[0195] Based on the above, the optimization problem model P is constructed:

[0196]

[0197] wherein and represent the minimum allocable computing resource and the maximum allocable computing resource of ES n[1], P m MD represents the power of the mth mobile device, represents the task latency of mobile device m at time slot t:

[0198] .

[0199] Embodiment 3

[0200] The optimization method for power allocation of the mobile edge computing system comprises the following specific steps:

[0201] An optimization problem model is constructed by the modeling method;

[0202] The task cost related to power is defined as ; when , and t is determined, is expressed as:

[0203]

[0204]

[0205]

[0206] wherein is the offloading transmission delay, is the data transmission rate between edge servers, and are the weights of the task delay cost and the energy consumption cost respectively, is the energy consumption of local computing, and n and n' represent different edge servers respectively;

[0207] According to the optimization problem model, a power optimization sub-problem is constructed:

[0208]

[0209] wherein, P MD is the power of the mobile device, is the edge server delay;

[0210] When , is convex;

[0211] Derivation is performed on :

[0212]

[0213] When is the minimum, , , it can be obtained that:

[0214]

[0215]

[0216]

[0217]

[0218]

[0219]

[0220]

[0221] Lambert W function is introduced, and the following equation is obtained

[0222]

[0223]

[0224]

[0225] According to the constraint condition of the power optimization sub-problem, the optimal power is obtained:

[0226]

[0227] wherein is:

[0228] .

[0229] Embodiment 4

[0230] The optimization method of the computing offloading and computing resource allocation of the mobile edge computing system comprises the following specific steps:

[0231] An optimization problem model is constructed by the modeling method;

[0232] According to the optimization problem model, a computing offloading and computing resource allocation sub-problem P2 is constructed:

[0233]

[0234] A Markov decision process composed of four tuples is constructed wherein is the current state space, a is the current action space, r is the reward, is the next state space; in the Markov decision process, at the beginning of each time slot, the current state of the system environment is observed and interacted with through actions, then the system environment enters the next state and accepts the reward, and so on;

[0235] The computing offloading and computing resource allocation sub-problem is solved to obtain the optimal allocation scheme of power allocation and computing offloading and computing resource allocation.

[0236] In one specific embodiment, the Markov decision process Specifically:

[0237] In the state space, starting from time slot t, each edge server independently observes the task attributes sent by the connected mobile devices and the remaining resources of all ESs, wherein is a vector; denotes the distance between the MD and the connected ES, denotes the remaining computing resources;

[0238] The action space includes the optimization variables of P2, specifically including the task offloading and computing resource allocation decisions of all MDs, ;

[0239] After taking an action, the agent immediately obtains a reward, and the reward function is specifically:

[0240] .

[0241] In one specific embodiment, when solving the computing offloading and computing resource allocation sub-problems to obtain the optimal allocation scheme of power allocation and computing offloading and computing resource allocation, as shown in Figure 3 , a D3T3 framework including duel double deep Q network D3QN and double delayed deep deterministic policy gradient TD3 is adopted, and the specific steps are as follows:

[0242] As shown in Figure 4 , initialize the environment experience replay buffer of the Q network D3QN and the environment experience replay buffer of the double delayed deep deterministic policy gradient TD3;

[0243] For each mobile device:

[0244] Generate a task;

[0245] The edge server connected by the mobile device observes the current system state s;

[0246] The agent on the edge server obtains a computing offloading policy action b by inputting s into the state-action value network Q of D3QN according to -greedy;

[0247] If b is equal to 0, the system environment directly obtains a reward r according to b, and updates the system state to obtain s′ ;

[0248] If b is not equal to 0, concatenate s and b to obtain s, obtain the allocated computing resource action f according to A( s) + N(0, σ1), N is noise; the system environment obtains a reward r according to b and f; f changes the remaining computing resources of the edge server, and the system environment state changes according to b and f, to obtain s′ ; the agent on the edge server obtains s′ using Q( b′ ); concatenate s′ and b′ to obtain s′; obtain the four-tuple ( s , f, r, s′ ) into the experience replay buffer ; and the quadruple ( s, b, r, s′ ) into the experience replay buffer ; obtaining the experience ( s, b, r, s′ ), ( s , f, r, s ′ );

[0249] wherein:

[0250]

[0251]

[0252]

[0253]

[0254]

[0255]

[0256]

[0257]

[0258]

[0259] When the experience replay buffer stores a set number of experiences, the D3T3 framework is trained, and the trained D3T3 framework is used to solve the computational offloading and computational resource allocation sub-problems to obtain the optimal allocation scheme of power allocation and computational offloading and computational resource allocation.

[0260] In this embodiment, D3QN is used for computational offloading and TD3 is used for computational resource allocation. D3QN takes the current state as input, processes the information, and iteratively helps MDs to formulate appropriate computational offloading strategies. Then, if the computational offloading strategy selects edge computing, we connect the current state and the computational offloading strategy as the input of TD3. TD3 synthesizes these information and provides appropriate computational resource allocation scheme. Otherwise, if the computational offloading strategy selects local computing, only D3QN is needed.

[0261] In this embodiment, D3T3 adopts an experience replay buffer mechanism, so at the beginning of DRL, D3T3 first needs to interact with the environment to obtain useful experiences. Only when there are enough experiences in the experience replay buffer can the training of the deep neural network be supported. The pseudo code of the interaction between MDs and ESs and the environment is shown in Algorithm 1. At the beginning of each interaction, D3QN will follow The `-greedy` option executes a computation offloading strategy. At this point, the agent determines whether the action is a local or edge computation strategy. If it's the latter, TD3 will perform a noisy computational resource allocation action. Finally, after completing the interaction, the agent receives a reward and stores the experience in an experience replay buffer.

[0262] In one specific embodiment, such as Figure 5 As shown, the specific steps for training the D3T3 framework are as follows:

[0263] Initialize 8 networks, including 4 online networks and 4 target networks; the online networks include the D3QN state-action-value network Q, the TD3 actor network A, and two independent critic networks C1 and C2; the target networks contain four networks Q′, A′, C1′, and C2′ with the same structure as the online networks.

[0264] For each time slot, loop:

[0265] Gain experience ( s, b, r, s′ ), ( s , f, r, s ′ );

[0266] Each person stores their experience in their own buffer pool;

[0267] If time slot t and tf If they are multiples of each other, then:

[0268] Randomly retrieved from the buffer pool K experience K ( s, b, r, s′ )= .Sample ( K );

[0269] According to the formula y = r + γ 1 Q′ ( s′, arg max b Q ( s′,b )) Calculate the target value y;

[0270] Update the parameters of Q based on the target value y;

[0271] If time slot t is a multiple of 2, then copy the parameters of Q to Q′;

[0272] If time slot t and tf If the numbers are multiples of each other, then:

[0273] Randomly take out K experiences from the buffer pool K ( s ,f,r, s ′ ) = .Sample ( K) ;

[0274] According to the formula y = r + γ 2min i=1,2 C i ′ ( s ′,A′ ( s ′ )+ N (0 ,σ 2))Calculate the target value y;

[0275] Update the parameters of C 1, C 2according to the target value y;

[0276] If the time slot t is a multiple of 2, update the parameters of A , A′ , C 1 ′ and C 2 ′ by soft updating;

[0277] The soft updating process is as follows:

[0278]

[0279]

[0280] Where θ is the neural network parameter, is the soft updating coefficient.

[0281] In this embodiment, the performance of D3T3 under the optimized power and the maximum power is compared as shown in Figure 6 .

[0282] In this embodiment, as shown in Figure 7 , Figure 8 , Figure 9 , Figure 10 In order to evaluate the performance of D3T3, the following baseline algorithm is also proposed for comparison experiment.

[0283] •D3DG: DRL based on D3QN-based computing offloading strategy and DDPG-based computing resource allocation strategy.

[0284] • DT3: DRL based on DQN-based computing offloading policy and TD3-based computing resource allocation policy.

[0285] • DDG: DRL based on DQN-based computing offloading policy and DDPG-based computing resource allocation policy.

[0286] • GAT3: DRL based on greedy algorithm-based computing offloading policy and DDPG-based computing resource allocation policy. The greedy algorithm will make the agent offload tasks to the nearest connected edge server.

[0287] • GADG: DRL based on greedy algorithm-based computing offloading policy and DDPG-based computing resource allocation policy.

[0288] The experimental settings and network structures of the above baseline algorithms are the same as D3T3.

[0289] In addition to the reward of DRL, the following indicators are used to further evaluate the performance of D3T3.

[0290] • Task completion rate (TCR): the percentage of completed tasks generated by all mobile devices.

[0291] • Remaining computing resources (RCR): the average number of remaining computing resources of edge servers after each episode ends.

[0292] As can be seen from the comparative experiments, the application has the advantages of effectively coordinating the task offloading and completion of mobile devices in MEC, being able to reduce the task delay and energy consumption of the overall MEC system, and obtaining more benefits.

[0293] Obviously, the above embodiments of the application are only examples for clearly illustrating the application, and are not intended to limit the embodiments of the application. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the application shall be included in the protection scope of the claims of the application.

Claims

1. A modeling method of an optimization model of a mobile edge computing system, characterized by comprising: The method comprises the following specific steps. A mobile edge computing system is constructed, which is composed of N edge servers and M mobile devices, and each mobile device is arranged to interact with the environment in a time division multiplexing manner. The mobile devices are randomly distributed in a closed area and randomly move in the area; in the system, the edge servers ={1, 2,..., N} are uniformly distributed in a fixed area, and each edge server is equipped with a computing unit ; the mobile devices ={1, 2,..., M} are randomly positioned and randomly move in the fixed area, and each mobile device is equipped with a computing unit ; and further include a transmitting antenna with a maximum transmitting power of . Each mobile device interacts with the environment in a time division multiplexing manner, specifically: the system time is uniformly divided into T discrete time slots, denoted as T={1, 2,..., T}, and each time slot has a length of ζ seconds; in time slot t, the system uses time division multiplexing to let each mobile device interact with the system environment in turn, and it is assumed that each mobile device has a task arriving in each time slot; the mobile device can execute a local computing task or offload the task to an edge server for computing; if the task is offloaded to the edge server for computing, the task will occupy the computing resources allocated to the MD by the ES until the end of the time slot; each mobile node moves randomly within the time slot, and the moving speed and direction follow a movement model; The movement model is specifically: wherein, is the maximum moving speed, L is the side length of the closed square region, , denotes the speed vector of the mth mobile device, denotes the speed of the mobile device, denotes the moving direction of the mobile device, and denote the coordinates of the mth mobile device, respectively; An optimization problem model is constructed by comprehensively considering the channel gain between the edge server and the mobile device, the signal-to-noise ratio of the edge server and the mobile device, the data transmission rate of the edge server and the mobile device, the total delay and total energy consumption of edge computing, and the task cost of edge computing. 2.The modeling method of an optimization model of a mobile edge computing system according to claim 1, characterized in that: An optimization problem model is constructed by comprehensively considering the channel gain between the edge server and the mobile device, the signal-to-noise ratio of the edge server and the mobile device, the data transmission rate of the edge server and the mobile device, the total delay and total energy consumption of edge computing, and the task cost of edge computing, and the task cost of edge computing. Let Tm(t) denote the computation task generated by mobile device m at time slot t , denotes the task data size, denotes the number of cycles required to process one bit of data, denotes the maximum allowable delay, denotes the value of the task; The channel gain between the edge server n and the mobile device m at time t is represented as: wherein denotes the antenna gain, denotes the speed of light, f denotes the carrier frequency, and d denotes the distance between the edge server n and the mobile device m; The transmit power of the mobile device m at time slot t is defined as The signal-to-noise ratio of the mobile device m to the edge server n is expressed as: wherein is the additional noise power; The data transmission rate of the mobile device m to the edge server n is represented as: wherein is the bandwidth of the wireless channel; At the beginning of each time slot t, let each mobile device connect to its nearest edge server; the computing offloading strategy of the connected edge server and mobile device m at time slot t can be represented as and ; at , the mobile device m selects to compute the task locally, when , the mobile device m selects to offload the task to the connected edge server n; when , and , the mobile device m selects to offload the task to the ES n, and the connected edge server n' will help the MD to forward the task; When the local computation delay is represented as: the local computation delay is represented as: The energy consumption of local computing is represented as: wherein is the power of the mobile device m, The energy consumption of offloading is represented as: is the energy consumption coefficient; When , and , the mobile device m selects to connect the edge server n' and offloads the task to the edge server n to compute the task; at this time, the transmission delay is represented as: where I is an indicator function, is the data transfer rate between edge servers; The computing delay of the task on the edge server n is represented as: The total delay and total energy consumption of edge computing at time t are represented as: wherein allocating computing resources to the MD for the edge server n, the capacity consumption is expressed as: wherein is the power of ES n; The task cost is defined as the weighted sum of the task delay and the energy consumption, and the task cost of the mobile device m at time t is represented as: Based on the above, the optimization problem model P is constructed: wherein and are the weights of the task latency cost and the energy cost, respectively, satisfying the condition: The method comprises the following specific steps. wherein and denote the minimum and maximum allocable computing resources of ES n, P m MD denotes the power of the mth mobile device, denotes the task latency of mobile device m at time slot t: 。 3. A method for optimization of power distribution of a mobile edge computing system, characterized in that: The optimization problem model is constructed by the modeling method in claim 2; According to the optimization problem model, a power optimization sub-problem is constructed: The power-related task cost is defined as ; when , and t are determined, the is expressed as: wherein is the offloading transmission delay, is the data transfer rate between edge servers, and are the weights of the task latency cost and the energy cost, respectively, is the energy consumption of local computation, and n and n' represent different edge servers, respectively. The Lambert W function is introduced to obtain wherein, P MD power for the mobile device, delay for the edge server; When time, convex; For derivative: When At the minimum, , , we have: According to the constraint condition of the power optimization sub-problem, the optimal power is obtained: The method comprises the following specific steps. wherein is: 。 4. A method for optimization of compute offloading and compute resource allocation for mobile edge computing systems, characterized in that: The optimization problem model is constructed by the modeling method in claim 2; According to the optimization problem model, a computing offloading and computing resource allocation sub-problem P2 is constructed: The computing offloading and computing resource allocation sub-problem is solved to obtain the optimal allocation scheme of power allocation, computing offloading and computing resource allocation. Constructing a Markov decision process consisting of a quadruple wherein is a current state space, a is a current action space, r is a reward, is a next state space; in a Markov decision process, at the beginning of each time slot, the current state of the system environment is observed and interacted with by an action, then the system environment enters a next state and receives a reward, and so on; After the action, the agent immediately obtains a reward, and the reward function used is specifically:

5. The method of optimization of computation offloading and computation resource allocation for mobile edge computing systems according to claim 4, characterized in that: The Markov decision Specifically: In the state space, from time slot t, each edge server independently observes the task attributes sent by the connected mobile devices and the remaining resources of all ESs, where is a vector; denotes the distance between the MD and the connected ES, denotes the remaining computing resources; The action space includes optimization variables of P2, specifically including task offloading and computing resource allocation decisions for all MDs, ; When the computing offloading and computing resource allocation sub-problem is solved to obtain the optimal allocation scheme of power allocation, computing offloading and computing resource allocation, the D3T3 framework including duel double deep Q network D3QN and double delay deep deterministic policy gradient TD3 is used, and the specific steps are as follows: 。 6. The method of optimization of computation offloading and computation resource allocation for mobile edge computing systems according to claim 5, characterized in that: For each mobile device: An environment experience replay buffer to initialize a Q-network, D3QN An environment experience replay buffer to initialize a Q-network, D3QN ; A task is generated; ​ The edge server connected by the mobile device observes the current system state s; The agent on the edge server obtains the action b according to , by inputting s into the state-action value network Q of the D3QN to obtain a calculation of the offloading strategy action b; If b is equal to 0, the system environment gets a reward r directly according to b, and updates the system state to get s′ ; If b is not equal to 0, then s and b are concatenated to obtain s, according to A( s) + N(0, σ1) to obtain the allocated computing resource action f, N is noise; the system environment obtains a reward r according to b and f; f changes the remaining computing resources of the edge server, and the system environment state changes according to b and f to obtain s′ ; the agent on the edge server obtains Q( s′ ) from b′ ; s' and b' are concatenated to obtain s'; the four-tuple ( s , f, r, s ′ ) is stored into ; the four-tuple ( s, b, r, s′ ) is stored into ; experience ( s, b, r, s′ ), ( s , f, r, s ′ ) is obtained; Wherein: When the experience replay buffer stores a set number of experiences, the D3T3 framework is trained, and the optimal allocation scheme of power allocation and computation offloading and computation resource allocation is obtained by solving the calculation offloading and computation resource allocation sub-problems through the trained D3T3 framework.

7. The method of optimization of computation offloading and computation resource allocation for mobile edge computing systems according to claim 6, characterized in that: The D3T3 framework is trained, and the specific steps are as follows: 8 networks are initialized, including 4 online networks and 4 target networks; the online network includes the state-action value network Q of D3QN, the actor network A of TD3, and two independent critic networks C1 and C2; the target network includes four networks Q', A', C1', and C2' which have the same structure as the online network; For each time slot, loop: Gaining experience s, b, r, s′ , ( s , f, r, s ′ ); Store the experience in the respective buffer pool; If the time slot t is in a multiple of tf 1 relationship with 1, then: randomly fetched from the buffer pool K a bar of experience K ( s, b, r, s′ )= .Sample ( K ); The target value y is calculated according to the formula y = r + γ 1 Q′ ( s′, arg max b Q ( s′,b )) Update the parameters of Q according to the target value y; If the time slot t is a multiple of 2, copy the parameters of Q to Q'; If the time slot t is in a multiple of 2 relationship with tf 2, then: Randomly take K experiences from the buffer pool K ( s , f, r, s ′ ) = .Sample ( K) ; According to the formula y = r + γ 2 min i=1,2 C i ′ ( s ′, A′ ( s ′ )+ N (0 ,σ 2)) calculate the target value y; update the target value y according to the parameter of the target value y C 1, C 2; If the time slot t is a multiple of 2, the parameters of the soft update are updated by the way of soft update A , A′ , C 1 ′ and C 2 ′ ​ The soft update process is as follows: where θ is the neural network parameter, is a soft update coefficient.

Citation Information

Patent Citations

  • Mobile edge computing resource allocation method for ultra-dense network

    CN111800828A

  • Method, system and equipment for unloading partial calculation of single server in mobile edge environment

    CN113950066A