A mobile edge computing resource allocation method based on game theory
By employing a method based on Stackelberg game theory and deep reinforcement learning, the problems of mobile device computing resource limitations and dynamic user demand interaction were solved, achieving efficient allocation of mobile edge computing resources and maximizing profits, thus reaching the global optimal solution.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY
- Filing Date
- 2023-03-02
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies have failed to effectively address the limitations of mobile device computing resources and the dynamic interaction of user needs, resulting in the inability to dynamically adjust computing resource pricing strategies and the inability to achieve globally optimal computing results.
We adopt a mobile edge computing resource allocation method based on Stackelberg game theory. By establishing a system model and deep reinforcement learning algorithm, we construct a game model between MEC servers and mobile users. This model enables MEC servers to set prices as leaders and mobile users to adjust their offloading strategies as followers, thereby optimizing resource allocation.
It achieves efficient computational offloading for mobile users and maximizes server profits under limited resources and transmission loss environments, reducing computational costs and achieving Stackelberg equilibrium.
Smart Images

Figure CN116501484B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of mobile edge computing, and particularly relates to a mobile edge computing resource allocation method based on game theory. BACKGROUND
[0002] In the past decade, mobile devices, including mobile phones, wearable devices, tablets, etc., have become increasingly popular. According to the report of Ericsson, it is predicted that mobile devices will reach about 20 billion by 2023. The rapid growth and progress of mobile devices have driven the diversity and complexity of mobile and Internet of Things applications. Many applications on mobile devices are resource-consuming, computation-intensive and high-energy-consuming. Due to limited computing power and battery power, it becomes increasingly difficult and impractical for mobile devices to run these applications. As a key technology of 5G mobile networks, edge computing can transfer computing tasks from mobile devices to edge servers with cloud computing capabilities, thereby solving the limitations of the resource conditions of mobile devices themselves. During the interaction between mobile users and service providers, mobile users purchase services of service providers, and service providers sell services to mobile users, both of which consider their own utility problems, highlighting the game relationship between them. In order to make mobile users and service providers achieve their expected utility, it is necessary to find a balance between them. Most of the current researches do not consider the dynamic problem of real-time interaction between user demand and computing resources. The pricing strategy of computing resources cannot be dynamically adjusted according to user demand, and the computing result cannot approach the global optimal solution. SUMMARY
[0003] The application aims to provide a mobile edge computing resource allocation method based on game theory to realize the maximization of interests of service providers and mobile users.
[0004] To solve the above technical problems, the technical solution adopted by the application is
[0005] S1, a system model is established according to actual application, and the unloading time delay t of user k is determined k and energy consumption e k ;
[0006] S2, a Stackelberg game model is established
[0007] The computing unloading decision and the computing resource allocation problem are established as a Stackelberg game model, denoted as Δ=(S, N, Y, Z, {U(Y)},{U k (Y k ,Z k}, where S is the leader-MEC server, N is the follower-mobile user set; Y, Z are the strategy sets of the leader and follower, respectively, denoted as Y = {Y1, Y2, …, Y N}, Z = {Z1, Z2, …, Z N} ;
[0008] {U(Y)},{U k (Y k ,Z k )} are the utility functions of the leader and follower, respectively;
[0009] The MEC server as the leader of the game, through the pricing of computing resources, sells to mobile users to maximize its own interests; the utility function of the leader can be expressed as:
[0010]
[0011] Y k is the price of the server that user K offloads tasks, mobile user k as the follower of the game, by adjusting its own offloading strategy according to the pricing strategy given by the leader to maximize its own interests; the utility function of the follower consists of two parts, one is the weighted sum of task offloading delay and energy consumption, and the weight i1i2 can be set to 0.5, the other part is the cost of purchasing MEC server computing resources, expressed as:
[0012] U k (Y k Z k ) = i1t k +i2e k +η k λ k ρkY k
[0013] In Stackelberg game, the interests of the leader and the follower are coupled, and the pricing strategy of the leader and the offloading strategy of the mobile user influence each other.
[0014] The total optimization problem can be expressed as
[0015]
[0016]
[0017] The application constructs the interaction process of the service provider and the user into a Stackelberg game model according to the different utilities pursued by the service provider and the user, i.e., maximizing the revenue of the service provider and minimizing the computing overhead of the mobile user, adopts a deep reinforcement learning algorithm to solve the model, and can reach the Stackelberg equilibrium within a limited number of games. The application can realize efficient computing offloading of the user and maximize the profit of the server, so as to minimize the computing cost under the condition of considering the limited resources of the server, data transmission loss and server computing cost environment. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0019] Figure 1 is the flow chart of the DDQN algorithm. DETAILED DESCRIPTION
[0020] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.
[0021] S1, establish a system model according to actual application, determine the offloading delay t of the user k k and energy consumption e k ;
[0022] In the whole edge computing system, the service provider is equipped with M MEC servers, which are used to provide computing and communication services for the application programs of N mobile users, wherein the set of MEC servers is represented as m={1, 2, …, M}, the set of mobile users is represented as u={1, 2, …, N}, and each mobile user k has a computing task W k =(λ k ,ρ k ,η k ), i.e., task state information, wherein λ k is in bit unit, represents the data amount required for processing the task, ρ k is in CPU cycle unit, represents the number of CPU cycles required for processing each unit of data of the task, and η kdenotes the proportion of data selected by mobile user k to be offloaded to the MEC server, and λ k η k is the amount of data reaching the MEC server.
[0023] The time delay of mobile user k for local computation is denoted as
[0024]
[0025] where f k denotes the computing capability of mobile user k, i.e., CPU processing capability. Meanwhile, the energy consumption of mobile user k for local computation is denoted as
[0026]
[0027] where w k denotes the energy consumed by mobile user for local computation of 1 bit data.
[0028] When mobile user offloads tasks to the MEC server, the time delay for transmission is
[0029]
[0030] where r k denotes the data transmission rate, and the calculation formula is
[0031]
[0032] where B denotes the wireless channel bandwidth between mobile user k and the MEC server; N denotes the channel noise power density; p k is the transmission power of mobile user k per unit time; g k denotes the channel gain, in dB, g k The calculation formula is:
[0033] g k = 127 + 25 x lgD
[0034] where D denotes the communication distance, in m, and if the communication distance does not change, the channel gain is a constant, r k is a constant value, in bit / s.
[0035] The energy consumption of mobile user k in the process of uploading computing data through the wireless channel is denoted as:
[0036]
[0037] The time delay of the MEC server for executing the computing task is denoted as
[0038]
[0039] where F k represents the computing resource allocated to mobile user k from MEC server, i.e. the number of CPU cycles per unit time.
[0040] In summary, the task offloading delay of mobile user k can be expressed as
[0041]
[0042] The energy consumption of mobile user k is expressed as
[0043]
[0044] S2, Stackelberg game model is established
[0045] The computing offloading decision and computing resource allocation problem is established as a Stackelberg game model, denoted as Δ = (S, N, Y, Z, {U(Y)}, {U k (Y k ,Z k )}), where S is the leader - MEC server, N is the follower - mobile user set; Y, Z are the strategy sets of the leader and the follower, denoted as Y = {Y1, Y2, …, Y N}, Z = {Z1, Z2, …, Z N};
[0046] {U(Y)}, {U k (Y k ,Z k )} are the utility functions of the leader and the follower, respectively.
[0047] The MEC server as the leader of the game, through the pricing of computing resources, sells to mobile users to maximize its own benefits. The utility function of the leader can be expressed as:
[0048]
[0049] Y k is the price of the server to which user K offloads tasks, mobile user k as the follower of the game, by adjusting its own offloading strategy according to the pricing strategy given by the leader to maximize its own benefits; the utility function of the follower consists of two parts, one is the weighted sum of task offloading delay and energy consumption, and the weight i1i2 can be set to 0.5, the other part is the cost of purchasing MEC server computing resources, expressed as:
[0050] U k (Y k Z k) = i1t k + i2e k + η k λ k ρ k Y k
[0051] In Stackelberg game, the interests of the leader and the follower are coupled, and the pricing strategy of the leader and the offloading strategy of the mobile user influence each other.
[0052] The total optimization problem can be expressed as
[0053]
[0054]
[0055] S3, taking deep reinforcement learning to optimize the Stackelberg game model
[0056] Specifically, each mobile user determines its offloading strategy by pricing the computing resource, and the MEC server solves the optimization of the pricing strategy according to the offloading strategy of the mobile user. Since the action spaces of the leader and the follower are different, the present application proposes a double-layer deep reinforcement learning algorithm, in which the deep deterministic policy gradient (DDPG) algorithm is used to optimize the leader problem, and the DDQN (double DQN) algorithm is used for the follower problem.
[0057] The DDQN algorithm specifically uses the action a to represent the offloading decision of the mobile user, and the state s to represent the computing overhead of the mobile user. Compared with the DQN algorithm, the DDQN algorithm mainly aims at the overestimation problem of the latter and changes the calculation method of the target value. The DDQN can effectively solve the overestimation problem and has better decision-making ability by constructing two action value functions, one for estimating the action and the other for estimating the value of the action. The specific algorithm process is as shown in Figure 1 .
[0058] In the case where the MEC server gives the pricing strategy, the mobile user can learn whether the action of itself is valuable to optimize the offloading strategy through the DDQN structure without learning the effect of each action of each state. The DDQN algorithm mainly includes the following steps:
[0059] (1) initialization.
[0060] (2) input the state S into the evaluation network to obtain the action, reward and next state S_ and store them.
[0061] (3) randomly extract the stored values to train the neural network.
[0062] (4) Add the parameters of the evaluation model to the target model.
[0063] DDPG algorithm, specifically, the action a is the pricing strategy of the MEC server, which is a continuous action space, the state s is the revenue of the MEC server, and the reward function is proportional to the revenue of the MEC server. DDPG outputs a determined action according to the policy function, which can reduce the sampling amount and improve the efficiency of the algorithm. The actor network and critic network in the DDPG algorithm both contain two networks with the same structure: the main network and the target network. The actor selects a determined action according to the policy function. In the network updating process, the actor network uses the policy gradient to update the parameters to determine the optimal action in a certain state. The critic network uses the loss function to update the parameters to evaluate the policy function generated by the actor.
[0064] The total code pseudo of the double-layer deep reinforcement learning algorithm of Stackelberg game is as follows:
[0065] Input: system parameters
[0066] Output: optimal offloading strategy and pricing strategy
[0067] 1. Initialize parameters, give a reasonable pricing strategy of the MEC server;
[0068] 2. Repeat:
[0069] 3. Mobile users use the DDQN algorithm to optimize the offloading strategy according to the pricing strategy of the MEC server
[0070] 4. The MEC server calculates the offloading strategy according to the feedback of the mobile users and uses the DDPG algorithm to optimize the pricing strategy
[0071] 5. Until: reach Stackelberg equilibrium.
[0072] The application researches the problem of joint MEC server computing resource allocation and computing offloading strategy in a heterogeneous network. The computing capacity of the MEC server is limited, and the MEC server needs to sell computing resources to mobile users based on a pricing strategy to obtain revenue. Mobile users need to decide the purchase quantity according to the pricing strategy of the MEC server to make offloading decisions. The interaction between the MEC server and the mobile user adopts a Stackelberg game model, and they are the leader and follower of the game respectively. Further, a double-layer reinforcement learning algorithm is proposed to optimize the pricing strategy of the MEC server and the computing offloading decision of the mobile user according to the different action spaces of the leader and the follower in the Stackelberg game, which can effectively improve the revenue of the MEC server and reduce the computing overhead of the system.
[0073] Each of the embodiments in the specification is described in a related manner, and the same and similar parts between the embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. Especially, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the related parts can be referred to the part of the method embodiment.
[0074] The above only describes the preferred embodiments of the application and is not used to limit the protection scope of the application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the application shall be included in the protection scope of the application.
Claims
1.A method for mobile edge computing resource allocation based on game theory, characterized in that, The method comprises the following steps: S1, establish a system model according to actual application, determine the user k's unloading time delay With energy consumption ; S2, establishing a Stackelberg game model The computing offloading decision and the computing resource allocation problem are established as a Stackelberg game model, denoted as Δ = (S, N, Y, Z, {U(Y)},{U k (Y k ,Z k )}), where S is the leader-MEC server, N is the follower-mobile user set; Y and Z are the strategy sets of the leader and the follower, denoted as Y = {Y 1, Y 2,….., Y N}, Z = {Z 1, Z 2,….., Z N}; {U(Y)},{U k (Y k ,Z k )} are the utility functions of the leader and the follower, respectively; The MEC server acts as a leader of the game, sells to the mobile user by pricing the computing resource, and obtains the maximum benefit; the utility function of the leader can be expressed as: ; Y k is the price of the server that user K offloads tasks, mobile user k as a follower of the game, by adjusting its own offloading strategy according to the pricing strategy given by the leader to obtain its own profit maximization; the utility function of the follower consists of two parts, one is the weighted sum of task offloading delay and energy consumption, and the weight can be set to 0.5, and the other part is the cost of purchasing MEC server computing resources, which is represented as: ; In the Stackelberg game, the benefits of the leader and the follower are coupled with each other, and the pricing strategy of the leader and the offloading strategy of the mobile user influence each other; The total optimization problem can be expressed as ; ; In the step S1, in the whole edge computing system, the service provider is equipped with M MEC servers for providing computing and communication services for the application programs of N mobile users, wherein the set of MEC servers is represented as m={1, 2, …, M}, the set of mobile users is represented as u={1, 2, …, N}, and each mobile user k has a computing task W k = (λ k , ρ k , η k ), i.e. task state information, wherein λ k is in bit unit, ρ k is in CPU cycle unit, η k represents the data proportion selected by the mobile user k to be unloaded to the MEC server, λ k η k is the data amount reaching the MEC server; The time delay of the mobile user k in local computing is expressed as ; where f k represents the computing capability of mobile user k, i.e., CPU processing capability; meanwhile, the energy consumption of mobile user k in the local computing part is represented as: ; wherein, represents the energy consumed by the mobile user to compute 1 bit of data locally; When the mobile user offloads the task to the MEC server, the time delay for transmission is ; where r k is expressed as a data transfer rate, the calculation formula is ; where B denotes the wireless channel bandwidth between the mobile user k and the MEC server; N denotes the channel noise power density; the transmit power of the mobile user k per unit time; denotes the channel gain, in dB, The calculation formula is: ; Wherein, D represents the communication distance, unit is m, if the communication distance does not change, the channel gain is constant, Is a constant value, unit is bit / s; The energy consumption of the mobile user k in uploading the computing data through the wireless channel is expressed as: ; The time delay of the MEC server in executing the computing task after the task arrives at the MEC server is expressed as ; wherein, denotes the number of CPU cycles per unit time allocated to mobile user k from the MEC server, i.e. According to the above, the task offloading time delay of the mobile user k is expressed as ; The energy consumption of the mobile user k is expressed as 。 2. The mobile edge computing resource allocation method of claim 1, wherein, The method further comprises the following step S3: taking deep reinforcement learning to optimize the Stackelberg game model, using a deep deterministic policy gradient algorithm to optimize the leader problem, and using a DDQN algorithm to optimize the follower problem.
Citation Information
Patent Citations
Joint Resource Allocation Based on Hierarchical Game in Mobile Edge Computing System
CN108990159A
Wireless body area network resource allocation and task unloading algorithm with maximum system revenue
CN111163519A