Calculation unloading and resource allocation joint optimization method and device

By building a model in an edge computing system and using reinforcement learning methods, combined with preset resource allocation methods, joint optimization of computing offloading and resource allocation is achieved, the impact of computing offloading and resource allocation strategies on performance in edge computing is solved, and system performance and user experience are improved.

CN120104304APending Publication Date: 2025-06-06BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510029363.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the field of edge computing, differences in computing offloading and resource allocation strategies will lead to changes in performance metrics such as latency and energy consumption, affecting the quality of user experience.

Method used

A joint optimization method for computing offloading and resource allocation is proposed. By building an edge computing system model, it is transformed into an online optimization problem. The computation offloading strategy is determined using reinforcement learning method, and based on this strategy, the optimal resource allocation strategy is determined using a preset resource allocation method.

Benefits of technology

By comprehensively considering factors such as resources, energy consumption and delay, joint optimization of computing offloading and resource allocation can be achieved, and the overall performance and service quality of the system are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104304A_ABST
    Figure CN120104304A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a calculation unloading and resource allocation joint optimization method and device, and the method comprises the steps: constructing an edge calculation system model, determining a long-term optimization problem of calculation unloading and resource allocation based on the model, and carrying out the calculation unloading and resource allocation joint optimization through employing a preset optimization method. A long-term optimization problem is converted into an online optimization problem for determining a calculation unloading strategy and a resource allocation strategy of each time slot, the calculation unloading strategy of each time slot is determined based on a reinforcement learning method for the online optimization problem, and an optimal resource allocation strategy is determined by adopting a preset resource allocation method based on the calculation unloading strategy. According to the method, factors influencing performance, such as resources, energy consumption and time delay, are comprehensively considered to carry out joint optimization of calculation unloading and resource allocation, and the overall performance and service quality of the system can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of edge computing technology, and in particular to a method and device for joint optimization of computing offloading and resource allocation. Background Art

[0002] Edge computing can meet the needs of low latency and high computing power for computationally intensive and delay-sensitive tasks in IoT terminal devices, and has been widely used in fields such as autonomous driving, smart healthcare, and intelligent industry. Computation offloading and resource allocation are two important issues in edge computing. Computation offloading means that the terminal device offloads part or all of the tasks to the edge server close to the terminal device based on the amount of computing required for the task and the current network environment; resource allocation means that the edge server receives the task sent by the terminal device, allocates resources to the task of the terminal device based on the task characteristics and its own resource conditions, and ensures that the task is completed efficiently and reliably. Different computation offloading decisions and resource allocation strategies will lead to changes in performance indicators such as latency and energy consumption, and provide different user experience quality (Quality of Experience, QoE), which is a key issue that needs to be solved in the field of edge computing. Summary of the invention

[0003] In view of this, an object of an embodiment of the present application is to propose a method and device for joint optimization of computation offloading and resource allocation.

[0004] Based on the above objectives, an embodiment of the present application provides a method for joint optimization of computing offloading and resource allocation, including:

[0005] Build an edge computing system model and determine the long-term optimization problem of computation offloading and resource allocation based on the model;

[0006] Using a preset optimization method, the long-term optimization problem is converted into an online optimization problem for determining a computation offloading strategy and a resource allocation strategy for each time slot;

[0007] For the online optimization problem, a computation offloading strategy for each time slot is determined based on a reinforcement learning method;

[0008] Based on the computation offloading strategy, a preset resource allocation method is adopted to determine the optimal resource allocation strategy.

[0009] Optionally, the online optimization problem includes a problem P4 for solving an optimal resource allocation strategy for an edge server and a problem P5 for solving an optimal computing offloading strategy; wherein the problem P4 is:

[0010]

[0011] Question P5 is:

[0012]

[0013] Where K is the number of terminal devices, D k (t) is the processing delay of the t-th time slot task between the terminal device and the edge server, f e (t) is the resource allocation strategy; γ t is the discount factor of the reward obtained in time slot t, r t is the reward after taking action in time slot t, T is the number of time slots, The expected value of the long-term cumulative reward;

[0014] Constraint C1 is used to ensure the legitimacy of the decision, C2 is used to ensure that the execution time of the task does not exceed the maximum delay, C3 is used to ensure that the computing resources allocated to each task on the edge server do not exceed the maximum computing capacity of the edge server, C4 is used to ensure that the computing resources allocated to each task on the terminal device do not exceed the maximum computing capacity of the terminal device, C5 is used to ensure the stability of the task queue used by the edge server to cache tasks, C6 is used to ensure the stability of the task waiting queue used by the terminal device to cache tasks, C7 is used to ensure that the total resources allocated to the edge server do not exceed its maximum resources, and C8 is used to ensure that the edge server does not exceed its energy consumption budget.

[0015] Optionally, a computation offloading strategy for each time slot is determined based on a reinforcement learning method, including:

[0016] Using a pre-trained computing offloading strategy model, inputting the current system state into the computing offloading strategy model, and the computing offloading strategy model outputs an optimal computing offloading strategy;

[0017] Among them, the state of the system includes the arrival rate of tasks on each terminal device, the wireless transmission bandwidth, the computing resources of the edge server, the computing resources of the terminal device, the length of the energy deficit queue of the edge server, and the length of the energy deficit queue of the terminal device; the length change rule of the energy deficit queue of the edge server is determined based on the energy consumption of the edge server processing tasks and the preset energy consumption budget, and the length change rule of the energy deficit queue of the terminal device is determined based on the energy consumption of the terminal device processing tasks and the preset energy consumption budget.

[0018] Optionally, determining an optimal resource allocation strategy based on the computing offloading strategy using a preset resource allocation method includes:

[0019] According to the state of the system and the computation offloading strategy, a preset resource allocation algorithm is used to solve problem P4 to obtain an optimal resource allocation strategy.

[0020] Optionally, according to the state of the system and the computing offloading strategy, a preset resource allocation algorithm is used to solve problem P4 to obtain an optimal resource allocation strategy, including:

[0021] Determining the computing resources required for the tasks offloaded by each terminal device according to the computing resource amount of each terminal device and the computing offloading decision;

[0022] Determine the computing resources initially allocated by the edge server to each task based on the computing resources required by the tasks offloaded by each terminal device and the task arrival rate;

[0023] Calculate the processing delay of each task from the corresponding terminal device to the edge server;

[0024] For each terminal device offload task:

[0025] Reducing and increasing the initially allocated computing resources by a preset resource adjustment amount respectively to obtain a first resource amount and a second resource amount;

[0026] Calculate a corresponding first processing delay and a second processing delay respectively according to the first resource amount and the second resource amount;

[0027] Determining an increase in the first processing delay according to the processing delay and the first processing delay;

[0028] Determining a reduction amount of the second processing delay according to the processing delay and the second processing delay;

[0029] Determine a minimum first processing delay increase according to the first processing delay increase corresponding to each task;

[0030] Determine a maximum second processing delay reduction amount according to the second processing delay reduction amount corresponding to each task;

[0031] If the maximum second processing delay reduction is greater than the minimum first processing delay increase, the initially allocated computing resources are adjusted for the task corresponding to the maximum second processing delay reduction and the task corresponding to the minimum first processing delay increase.

[0032] Optionally, the resource adjustment amount is a minimum computing resource block that can be adjusted by the edge server; and the adjusting of the initially allocated computing resources includes:

[0033] For the task corresponding to the maximum second processing delay reduction amount, reducing the initially allocated computing resources by the resource adjustment amount;

[0034] For the task corresponding to the minimum first processing delay increase, the initially allocated computing resources are increased by the resource adjustment amount.

[0035] Optionally, the method for calculating the processing delay of each task from the corresponding terminal device to the edge server is:

[0036]

[0037] in, Processing task delay for terminal device k, is the data transmission delay between the terminal device and the edge server, Processing task delay for edge servers, D k (t) is the processing delay of the t-th time slot task between the terminal device and the edge server.

[0038] Optionally, the calculation offloading strategy is an offloading ratio of the terminal device to offload tasks to the edge server.

[0039] Optionally, the building of the edge computing system model includes:

[0040] The edge computing system is modeled using the M / D / 1 queue model.

[0041] The embodiment of the present application also provides a device for joint optimization of computation offloading and resource allocation, including:

[0042] A model building module, which is used to build an edge computing system model and determine the long-term optimization problem of computing offloading and resource allocation based on the model;

[0043] A problem conversion module, used to convert the long-term optimization problem into an online optimization problem for determining a computation offloading strategy and a resource allocation strategy for each time slot by using a preset optimization method;

[0044] A computing offloading strategy solving module, used for determining the computing offloading strategy for each time slot based on a reinforcement learning method for the online optimization problem;

[0045] The resource allocation strategy solving module is used to determine the optimal resource allocation strategy based on the computing offloading strategy by using a preset resource allocation method.

[0046] From the above, it can be seen that the method and device for joint optimization of computational offloading and resource allocation provided by the embodiment of the present application, by constructing an edge computing system model, determines the long-term optimization problem of computational offloading and resource allocation based on the model, and uses a preset optimization method to convert the long-term optimization problem into an online optimization problem of determining the computational offloading strategy and resource allocation strategy for each time slot. For the online optimization problem, the computational offloading strategy for each time slot is determined based on the reinforcement learning method, and the optimal resource allocation strategy is determined based on the computational offloading strategy using a preset resource allocation method. The present application comprehensively considers factors that affect performance, such as resources, energy consumption, and latency, to perform joint optimization of computational offloading and resource allocation, which can improve the overall system performance and service quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0048] Figure 1 A schematic diagram of a method flow of an embodiment of the present application;

[0049] Figure 2 A schematic diagram of a joint optimization model according to an embodiment of the present application;

[0050] Figure 3 A schematic diagram of a reinforcement learning model according to an embodiment of the present application;

[0051] Figure 4 A pseudo code diagram of a training strategy network model according to an embodiment of the present application;

[0052] Figure 5 A pseudo code diagram of a heuristic resource allocation algorithm according to an embodiment of the present application;

[0053] Figure 6 A pseudo code diagram of the P2PHRA algorithm of an embodiment of the present application;

[0054] Figure 7 This is a block diagram of the device structure of an embodiment of the present application;

[0055] Figure 8 This is a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0056] In order to make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.

[0057] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should be the usual meanings understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the embodiments of the present application do not represent any order, quantity or importance, but are only used to distinguish different components. "Including" or "comprising" and similar words mean that the elements or objects appearing in front of the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connecting" or "connected" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to represent relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0058] As described in the background technology section, computation offloading and resource allocation are key issues that edge computing needs to solve. Computation offloading and resource allocation affect each other. Computation offloading decisions will directly affect resource allocation strategies, and resource allocation strategies will also affect the effect of computation offloading. If the terminal device offloads the task to the edge server, the edge server should allocate appropriate resources for this task to ensure the quality of the task, which will lead to changes in the resource allocation strategy; if the resource allocation is unreasonable and cannot meet the task requirements, the effect of computing offloading to the edge server will be reduced, resulting in problems such as extended task processing time, which will affect system performance and user experience. Therefore, in order to achieve the optimal overall performance of edge computing, it is necessary to jointly optimize computation offloading and resource allocation.

[0059] In view of this, an embodiment of the present application provides a method for joint optimization of computing offloading and resource allocation, constructs an edge computing system model, determines the long-term optimization problem of computing offloading and resource allocation based on the model, and converts the long-term optimization problem into an online optimization problem that can be solvable in each time slot. The computing offloading strategy for each time slot is determined based on the reinforcement learning method, and the optimal resource allocation strategy is determined by using a heuristic resource allocation method based on the computing offloading strategy, thereby achieving joint optimization of computing offloading and resource allocation and improving the overall performance and service quality of the edge computing system.

[0060] The technical solution of the present application is further described in detail below through specific embodiments.

[0061] like Figure 1 , 2 As shown, the embodiment of the present application provides a method for joint optimization of computing offloading and resource allocation, including:

[0062] S101: construct an edge computing system model, and determine the long-term optimization problem of computing offloading and resource allocation based on the model;

[0063] In this embodiment, the edge computing scenario includes multiple edge servers and multiple terminal devices. For an edge computing system consisting of an edge server and multiple terminal devices that can access the edge server, each terminal device builds a task waiting queue for caching tasks that have arrived but not processed, and the tasks in the task waiting queue can be partially or completely unloaded to the edge server for execution. The edge server builds a task queue for each terminal device and uses the task queue to cache the tasks unloaded to the edge server by the corresponding terminal device. The network status, resource status, etc. of the edge computing system are all dynamically changing. At the same time, in order to ensure the long-term operation of the system, the terminal devices and edge servers must meet long-term energy consumption restrictions, that is, during the operation of the system, the sum of the energy consumption of each time slot cannot exceed the overall energy consumption budget.

[0064] To describe the above edge computing system, the edge computing system is modeled by the M / D / 1 queue model. Define the terminal device set K is the number of terminal devices in the system, and the computing task of terminal device k is in, is the maximum execution delay of the task on terminal device k (in seconds), S k is the data volume of the task on terminal device k (in Mb), C k The amount of computing resources (including CPU cycles) required to complete the task for terminal device k, C k It can be split in any proportion. When a task reaches the terminal device, it is added to the task waiting queue. Through computational offloading, part of the task can be offloaded to the edge server, that is, the terminal device executes part of the task and the rest is offloaded to the edge server, which executes some of the tasks in the task queue; or, the terminal device offloads all the tasks to the edge server, which executes all the tasks in the task queue.

[0065] In some embodiments, the operation process of the edge computing system is divided into multiple time slots. In each time slot, it is assumed that the task arrival of the terminal device follows the Poisson distribution. At the beginning of each time slot, the computing offloading strategy of the terminal device and the resource allocation strategy of the edge server are determined. In the entire time slot, the computing offloading strategy and the resource allocation strategy remain unchanged, and the network status, channel bandwidth and other system information of the terminal device and the edge server remain unchanged. Then, in the tth time slot, the system state is defined as follows:

[0066] The task arrival rate of terminal device k is λ k (t), vector λ(t)={λ 1 (t),…,λ k (t)} represents the task arrival rate of each terminal device in the tth time slot; the computing resources available for tasks on terminal device k are vector represents the computing resources available for tasks in each terminal device in the tth time slot. The current computing offloading decision of terminal device k is x k (t)∈[0,1],x k (t) = 0 means that all tasks are executed by the terminal device, x k (t) = 1 means that all tasks are offloaded to the edge server. A value between 0 and 1 means that the terminal device performs part of the task and offloads part of the task to the edge server. The specific value represents the offloading ratio of the task. The vector x(t) = (x 1 (t),x 2 (t),…,x K (t)) represents the computation offloading decision of each terminal device in the tth time slot. The computing resources allocated by the edge server to the task offloaded by terminal device k in the tth time slot, vector represents the amount of computing resources allocated by the edge server to each terminal device to offload tasks in the tth time slot. The communication bandwidth between each terminal device and the edge server is W(t); the energy consumption budget of the terminal device and the edge server in a single time slot is δ e .

[0067] The processing delay of the t-th time slot task between the terminal device and the edge server is determined by the terminal device processing task delay, the data transmission delay between the terminal device and the edge server, and the edge server processing task delay. Among them, the terminal device processing task delay includes task queuing delay and processing delay, and the calculation method is:

[0068]

[0069] in, Processing task delays for terminal devices, Indicates the service rate of tasks in the task waiting queue. It represents the utilization rate of tasks in the task waiting queue, which is the ratio of the task arrival rate to the task service rate. k When (t) = 1, all tasks are offloaded to the edge server, then When 0≤x k When (t)<1, Indicates the processing delay of the task, Indicates the queuing delay in the queue.

[0070] When the edge server is required to perform part of the task, the terminal device needs to transmit the intermediate result data of the executed part of the task to the edge server, and the edge server will continue to perform the remaining tasks based on the intermediate result. Assuming that the data transmission rate in a time slot remains unchanged, taking wireless transmission as an example, the wireless transmission rate in the tth time slot is defined as W(t), and its value range is [W min (t),W max (t))], the calculation method of wireless data transmission delay is:

[0071]

[0072] in, is the data transmission delay, ξ(x k (t)) is the independent variable to calculate the unloading ratio x k (t) is a function of the ratio of tasks performed by the terminal device. The size of the intermediate result data generated may be different from x k (t)S k It is a nonlinear relationship, ξ(x k (t)) is used to characterize the size of the intermediate result data generated under different computation offloading decisions.

[0073] In order to keep the task waiting queue stable, it is necessary to ensure the utilization rate of tasks in the task waiting queue When it is less than 1, the task leaving process will also approach the rate of λ k (t) is a Poisson process. Since the system information remains unchanged, the arrival process of the offload task at the edge server also approaches the rate λ k (t), the calculation method of the edge server processing task delay is:

[0074]

[0075] in, Processing task delay for edge servers, when When , it means that the task is offloaded to the edge server and the service rate of the task on the edge server is greater than the service rate of the task on the terminal device. In this case, the task does not need to queue on the edge server and can be processed immediately, that is, Indicates the service rate of tasks in the task queue, Indicates the utilization rate of tasks in the task queue.

[0076] Therefore, the processing delay of the t-th time slot task in the terminal device and the edge server is expressed as:

[0077]

[0078] In some embodiments, the energy consumption of the terminal device and the edge server is related to the computing load. The energy consumption of the terminal device k processing the task in the tth time slot is defined as:

[0079]

[0080] Where τ is the time slot length, is the energy consumption per unit of computing resources of the terminal device. k When (t) = 1, all tasks are unloaded, and the energy consumption of terminal device k is 0; when 0≤x k When (t)<1, in order to ensure the stability of the service queue, all tasks arriving in time slot t will be processed. The number of tasks arriving is τλ(t), and the computational complexity of each task on the terminal device is [1-x k (t)]C k .

[0081] The energy consumption of the edge server processing the task offloaded by terminal device k is:

[0082]

[0083] in, is the energy consumption per unit of computing resources of the edge server.

[0084] The total energy consumption of the edge server in time slot t is the sum of the energy consumption of processing each task, expressed as:

[0085]

[0086] In this embodiment, in order to achieve joint optimization of computation offloading and resource allocation, a long-term optimization problem is constructed with the goal of minimizing the long-term average end-to-end (from terminal device to edge server) delay of all tasks under a given energy budget:

[0087]

[0088] Among them, constraint C1 ensures the legitimacy of the decision, C2 ensures that the execution time of the task does not exceed the maximum delay, C3 ensures that the computing resources allocated to each task on the edge server do not exceed the maximum computing capacity of the edge server, C4 ensures that the computing resources allocated to each task on the terminal device do not exceed the maximum computing capacity of the terminal device, C5 ensures the stability of the task queue, C6 ensures the stability of the task waiting queue, C7 ensures that the total resources allocated to the edge server do not exceed its maximum resources, C8 ensures that the edge server does not exceed its energy consumption budget, and C9 ensures that each terminal device does not exceed the energy consumption budget.

[0089] S102: using a preset optimization method, converting the long-term optimization problem into an online optimization problem for determining a computation offloading strategy and a resource allocation strategy for each time slot;

[0090] In this embodiment, considering that problem P1 shown in formula (8) is a long-term optimization problem, its constraints C8 and C9 both involve the limit when T tends to infinity. Solving it requires knowing the global information of time slots t=0 to t=T, and it is difficult to solve it directly. Therefore, the Lyapunov optimization method is used to construct the energy deficit queue of the terminal device and the edge server, and the long-term optimization problem is converted into an online optimization problem of the computation offloading strategy and resource allocation strategy for each time slot, guiding the computation offloading decision to follow the energy consumption constraint, while ensuring the minimization of the long-term end-to-end delay.

[0091] In some embodiments, changes in the energy deficit queue of the constructed edge server and the energy deficit queue of the terminal device follow the following rules:

[0092] q e (t+1)=max{q e (t)-δ e +E e (t),0} (9)

[0093]

[0094] Among them, q e (t) is the length of the energy deficit of the edge server at the end of time slot t, is the length of the energy deficit of the terminal device at the end of time slot t. Both are initialized to 0. They are used to characterize the deviation between the current energy consumption and the long-term energy consumption constraint.

[0095] Problem P1 can be solved by, in each time slot t, We can approximate this by solving problem P2. This allows us to achieve a balance between energy deficit queue stability and end-to-end delay minimization. Trade-off. Problem P2 is expressed as:

[0096]

[0097] in, represents Lyapunov drift, is the penalty function, and V is a non-negative weight parameter used to adjust the trade-off between minimizing the Lyapunov drift and minimizing the penalty function.

[0098] It can be seen that problem P2 no longer has the problem of infinite time slots involved in constraints C8 and C9, and problem P2 is an online optimization problem for each time slot.

[0099] S103: For the online optimization problem, determine the computation offloading strategy for each time slot based on the reinforcement learning method;

[0100] In this embodiment, the analysis is based on the online optimization problem P2, which consists of two mutually coupled sub-problems. e (t), problem P2 is equivalent to the computational offloading decision subproblem about the continuous variable x(t). The objective function of this subproblem is non-convex and is a typical NP-hard problem. When the variable x(t) is given, problem P2 is equivalent to the computational offloading decision subproblem about the discrete variable f e The resource allocation subproblem of (t) belongs to the integer linear programming problem and is also a typical NP-hard problem. Therefore, problem P2 is a mixed integer nonlinear programming problem, which is an NP-hard problem. It is impossible to find the optimal solution of problem P2 in polynomial time.

[0101] In order to find the optimal solution to problem P2, based on the reinforcement learning method, the optimal strategy is determined by maximizing the long-term cumulative reward of the time slot in the Markov decision process. For the joint optimization problem, a Markov decision process is constructed

[0102]

[0103] in, is the state space of the edge computing system, is the action space of the system. is the state transition probability, is the range of the reward function r.

[0104] Specifically, at the beginning of time slot t, the state of the edge computing system can be defined as: the arrival rate λ(t) of tasks on each terminal device, the wireless transmission bandwidth W(t), the computing resources F(t) of the edge server, and the e (t), computing resources F of the terminal device u (t), the length of the energy deficit queue of the edge server q e (t), the length q of the energy deficit queue of the terminal device u (t), expressed as:

[0105] s t ={λ(t),W(t),F e (t),F u (t),q e (t),q u (t)} (13)

[0106] After observing the environment state at time slot t, the computation offloading decision and resource allocation action are determined, which is expressed as:

[0107] a t ={x(t),f e (t)} (14)

[0108] Since the Markov decision process is to maximize the cumulative reward, problem P2 is to minimize the objective function. Therefore, the constructed reward function is opposite to the drift plus penalty function in problem P2. The reward function is expressed as:

[0109]

[0110] After constructing the Markov decision process, problem P2 is transformed into problem P3:

[0111]

[0112] Among them, γ t is the discount factor of the reward obtained in time slot t, r t is the reward obtained after taking action in time slot t, is the expected value of the long-term cumulative reward.

[0113] Constraints C2, C5, and C6 greatly increase the difficulty of training the neural network to learn the optimal strategy. If the reinforcement learning algorithm is used to solve problem P3 directly, the training time cost of the neural network is too high and even difficult to converge. Further analysis, according to formulas (5)-(7), for time slot t, given the computing resources F of the edge server e (t), resource allocation vector f e (t) does not affect the energy consumption of the edge server, because no matter how the computing resources are allocated, the sum of its elements is equal to F e (t), in other words, the energy consumption of the system is only related to the computation offloading decision and has nothing to do with the resource allocation strategy. Therefore, the resource allocation vector can be separated from the reinforcement learning strategy training, and the computation offloading decision x(t) of time slot t can be determined by the reinforcement learning method. Then, based on the determined computation offloading decision, the optimal resource allocation strategy of the edge server can be determined by solving problem P4. Problem P4 is expressed as:

[0114]

[0115] After removing the resource allocation vector from the reinforcement learning policy training, problem P3 is transformed into problem P5:

[0116]

[0117] Problem P5 only has constraint C1 left. We only need to determine the computational offloading decision for each time slot task based on the reinforcement learning method. As long as the legality of the strategy output action is guaranteed, constraint C1 can be met.

[0118] Through the above analysis, this embodiment transforms the long-term optimization problem into an online optimization problem P2 with a complex coupling relationship, and further decomposes the problem P2 into simplified online optimization problems P4 and P5. The optimal computational offloading strategy is obtained by solving the problem P5 respectively, and the optimal resource allocation strategy is obtained by solving the problem P4.

[0119] In some embodiments, to solve problem P5 to obtain the optimal computation offloading strategy, a pre-trained computation offloading strategy model is used to transform the current system state s t The computing offloading strategy model is input, and the computing offloading strategy model outputs the optimal computing offloading strategy.

[0120] In some embodiments, the policy network model is trained using a proximal policy optimization algorithm. In each iteration, the policy network model is optimized by limiting the difference between the new policy and the old policy, maintaining the stability of the policy and avoiding large fluctuations during the training process. Figure 3 , 4 As shown in Figure 1, the training process mainly includes:

[0121] In each training round, the environment is initialized, and the model parameters of the policy network model and the value network model are randomly initialized. The policy network model provides an initial computation offloading decision based on the initial parameters. t In the input strategy network model and value network model, the strategy network model is based on the system state s t Determine the computational offloading decision vector x(t) and the action probability ratio logπ of the task old (a t |s t ), the value network model estimates the value function v(t) of the corresponding state. After determining the calculation of the unloading decision vector x(t), the resource allocation vector f is determined using the preset heuristic resource allocation algorithm. e (t), get action a t =(x(t),f e (t)), and calculate the reward value r according to formula (15) t+1 , update the next state s t+1 ; Construct state transfer sample (s t ,a t ,s t+1 ,r t+1 ,logπ old (a t |s t )), and store the samples in the playback buffer for updating the network model.

[0122] When a certain number of state transition samples are stored in the replay buffer, the network model is updated. The specific process is as follows: read the state transition samples from the replay buffer, use the preset policy gradient formula to perform gradient ascent, and update the policy network model; use the preset temporal difference error function to perform gradient descent and update the value network model. After a predetermined number of updates, the replay buffer is cleared and the next sample collection is performed. After a certain number of rounds of training, the final policy network model is obtained, and the trained policy network model is used as the computational offloading policy model to determine the optimal computational offloading strategy.

[0123] S104: Determine an optimal resource allocation strategy using a preset resource allocation method based on the computation offloading strategy.

[0124] In this embodiment, after the optimal computing offloading strategy is determined by using the computing offloading strategy model, a heuristic resource allocation algorithm is used to solve problem P4 to obtain the optimal resource allocation strategy, that is, according to the current system state s t And calculate the unloading strategy x(t), and use the preset resource allocation algorithm to determine the optimal resource allocation strategy f e (t), thus obtaining the optimal action a t ={x(t),f e (t)}.

[0125] In some embodiments, according to the current system status s t And calculate the unloading strategy x(t), and use the preset resource allocation algorithm to determine the optimal resource allocation strategy f e (t), including:

[0126] Determine the computing resources required for the tasks offloaded from each terminal device based on the computing resources of each terminal device and the computing offloading decision;

[0127] Determine the computing resources initially allocated by the edge server to each task based on the computing resources required by the tasks offloaded by each terminal device and the task arrival rate;

[0128] Calculate the processing delay of each task from the corresponding terminal device to the edge server;

[0129] For each terminal device offload task:

[0130] Reducing and increasing the initially allocated computing resources by a preset resource adjustment amount respectively to obtain a first resource amount and a second resource amount;

[0131] Calculate the corresponding first processing delay and second processing delay respectively according to the first resource amount and the second resource amount;

[0132] Determining an increase in the first processing delay according to the processing delay and the first processing delay;

[0133] Determining a reduction amount of the second processing delay according to the processing delay and the second processing delay;

[0134] Determine a minimum first processing delay increase according to the first processing delay increase corresponding to each task;

[0135] Determine a maximum second processing delay reduction amount according to the second processing delay reduction amount corresponding to each task;

[0136] If the maximum second processing delay reduction is greater than the minimum first processing delay increase, the initially allocated computing resources are adjusted for the task corresponding to the maximum second processing delay reduction and the task corresponding to the minimum first processing delay increase.

[0137] In this embodiment, the resource allocation of the edge server to the tasks offloaded by the two terminal devices is adjusted by using a point-to-point heuristic resource allocation algorithm (P2PHRA). Figure 5 , 6 As shown, specifically:

[0138] According to the computing resource amount and computing offloading decision of each terminal device, the computing resources required for the task offloaded by each terminal device are calculated; wherein the computing resources required for the task offloaded by terminal device k are:

[0139]

[0140] According to the computing resources required for the tasks unloaded by each terminal device and the task arrival rate, the computing resources initially allocated by the edge server to each task are determined; wherein the computing resources required to be allocated by the edge server to the task unloaded by the terminal device k are:

[0141]

[0142] According to the system status and computing offloading strategy, the processing delay from each terminal device to the edge server is calculated according to formula (1), and the processing delay of each terminal device forms a delay vector

[0143] The computing resources initially allocated to any task By reducing the amount of resource adjustments Get the first resource amount f k ′ , by increasing the amount of resource adjustment Get the second resource amount f k ′′ ;in, It is the smallest computing resource block that can be adjusted by the edge server.

[0144] According to the first resource amount, the corresponding first processing delay D is calculated according to formula (1): ′ k , that is, after reducing resource allocation, the edge server processing task delay shown in formula (3) changes, and the processing delay increases. and the first processing delay D after reducing resource allocation ′ k , calculate the increase in the first processing delay A vector recording the first processing delay increase corresponding to the tasks unloaded by all terminal devices The smallest first processing delay increase is selected from the vector, and the task unloaded by the terminal device corresponding to the smallest first processing delay increase is recorded.

[0145] According to the second resource amount, the corresponding second processing delay D is calculated according to formula (1): ″ k ″ , that is, the processing delay caused by increasing resource allocation, according to the processing delay and the second processing delay D after adding resource allocation ″ k ″ , calculate the second processing delay reduction A vector recording the second processing delay reduction corresponding to the tasks unloaded by all terminal devices The maximum increase in the second processing delay is selected from the vector, and the task unloaded by the terminal device corresponding to the maximum decrease in the second processing delay is recorded.

[0146] Compare the minimum increase in the first processing delay with the maximum decrease in the second processing delay. If the maximum decrease in the second processing delay is greater than the minimum increase in the first processing delay, it indicates that the total delay of the two tasks will be reduced, and the total delay can be reduced by adjusting the resource allocation of the two tasks. That is, for the task unloaded by the terminal device corresponding to the minimum increase in the first processing delay and the task unloaded by the terminal device corresponding to the maximum decrease in the second processing delay, by adjusting the computing resources initially allocated to the edge server, the resource adjustment amount is reduced for the task corresponding to the minimum increase in the first processing delay, and the resource adjustment amount is increased for the task corresponding to the maximum increase in the second processing delay. This can reduce the total delay of the tasks unloaded by the two terminal devices, and is an optimization solution for resource allocation.

[0147] The heuristic resource allocation algorithm provided in this embodiment can ensure that the constraints C2-C8 are satisfied. The reason is that: given x(t), the initialization method of the algorithm can ensure that the initial resource allocation is legal; if f k v Less than It will cause the execution time of the task to exceed the maximum tolerable delay, and the corresponding task will not be selected as a choice of reducing computing resources, and will only accept adjustments to increase computing resources. If I is the maximum number of iterations of the algorithm, the time complexity of the algorithm is O(I×K).

[0148] The method for joint optimization of computing offloading and resource allocation provided in the embodiment of the present application determines the long-term optimization problem of computing offloading and resource allocation based on the constructed edge computing system model, and transforms the long-term optimization problem into an online optimization problem of determining the computing offloading strategy and resource allocation strategy for each time slot through optimization. For the online optimization problem of each time slot, according to the state of the system, the optimal computing offloading strategy is determined by using a computing offloading strategy model based on reinforcement learning. According to the state of the system and the computing offloading strategy, the optimal resource allocation strategy is determined by using a heuristic resource allocation method, thereby determining the optimal computing offloading strategy and resource allocation strategy of the system, comprehensively considering factors affecting performance such as resources, energy consumption, and latency, and being able to improve the overall performance and service quality of the system.

[0149] It should be noted that the method of the embodiment of the present application can be performed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only perform one or more steps in the method of the embodiment of the present application, and the multiple devices will interact with each other to complete the described method.

[0150] It should be noted that the above is a description of a specific embodiment of the present specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0151] like Figure 7 As shown, a device for joint optimization of computing offloading and resource allocation includes:

[0152] A model building module, which is used to build an edge computing system model and determine the long-term optimization problem of computing offloading and resource allocation based on the model;

[0153] A problem conversion module, used to convert the long-term optimization problem into an online optimization problem for determining a computation offloading strategy and a resource allocation strategy for each time slot by using a preset optimization method;

[0154] The computation offloading strategy solving module is used to determine the computation offloading strategy for each time slot based on the reinforcement learning method for the online optimization problem;

[0155] The resource allocation strategy solving module is used to determine the optimal resource allocation strategy based on the computing offloading strategy and using a preset resource allocation method.

[0156] For the convenience of description, the above devices are described in terms of functions divided into various modules. Of course, when implementing the embodiments of the present application, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0157] The apparatus of the above-mentioned embodiment is used to implement the corresponding method in the above-mentioned embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.

[0158] Figure 8 A more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment is shown, and the device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are connected to each other through the bus 1050 in the device.

[0159] The processor 1010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0160] The memory 1020 may be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.

[0161] The input / output interface 1030 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0162] The communication interface 1040 is used to connect a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).

[0163] The bus 1050 includes a path that transmits information between the various components of the device (eg, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0164] It should be noted that, although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040 and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiments of the present specification, and does not necessarily include all the components shown in the figure.

[0165] The electronic device of the above embodiment is used to implement the corresponding method in the above embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.

[0166] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0167] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples. Based on the concept of the present disclosure, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.

[0168] In addition, to simplify the description and discussion, and in order not to make the embodiments of the present application difficult to understand, the known power supply / ground connection with the integrated circuit (IC) chip and other components may or may not be shown in the provided drawings. In addition, the device can be shown in the form of a block diagram to avoid making the embodiments of the present application difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform to be implemented in the embodiments of the present application (that is, these details should be fully within the scope of understanding of those skilled in the art). In the case of elaborating specific details (e.g., circuits) to describe exemplary embodiments of the present disclosure, it is obvious to those skilled in the art that the embodiments of the present application can be implemented without these specific details or when these specific details are changed. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0169] Although the present disclosure has been described in conjunction with specific embodiments of the present disclosure, many replacements, modifications and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.

[0170] The embodiments of the present application are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application should be included in the scope of protection of the present disclosure.

Claims

1. A method for joint optimization of computation offloading and resource allocation, characterized in that: include: Build an edge computing system model and determine the long-term optimization problem of computation offloading and resource allocation based on the model; Using a preset optimization method, the long-term optimization problem is converted into an online optimization problem for determining a computation offloading strategy and a resource allocation strategy for each time slot; For the online optimization problem, a computation offloading strategy for each time slot is determined based on a reinforcement learning method; Based on the computation offloading strategy, a preset resource allocation method is adopted to determine the optimal resource allocation strategy.

2. The method according to claim 1, characterized in that The online optimization problem includes problem P4 for solving the optimal resource allocation strategy of the edge server and problem P5 for solving the optimal computing offloading strategy; wherein problem P4 is: Question P5 is: Where K is the number of terminal devices, D k (t) is the processing delay of the t-th time slot task between the terminal device and the edge server, f e (t) is the resource allocation strategy; γ t is the discount factor of the reward obtained in time slot t, r t is the reward after taking action in time slot t, T is the number of time slots, The expected value of the long-term cumulative reward; Constraint C1 is used to ensure the legitimacy of the decision, C2 is used to ensure that the execution time of the task does not exceed the maximum delay, C3 is used to ensure that the computing resources allocated to each task on the edge server do not exceed the maximum computing capacity of the edge server, C4 is used to ensure that the computing resources allocated to each task on the terminal device do not exceed the maximum computing capacity of the terminal device, C5 is used to ensure the stability of the task queue used by the edge server to cache tasks, C6 is used to ensure the stability of the task waiting queue used by the terminal device to cache tasks, C7 is used to ensure that the total resources allocated to the edge server do not exceed its maximum resources, and C8 is used to ensure that the edge server does not exceed its energy consumption budget.

3. The method according to claim 2, characterized in that The computation offloading strategy for each time slot is determined based on the reinforcement learning method, including: Using a pre-trained computing offloading strategy model, inputting the current system state into the computing offloading strategy model, and the computing offloading strategy model outputs an optimal computing offloading strategy; Among them, the state of the system includes the arrival rate of tasks on each terminal device, the wireless transmission bandwidth, the computing resources of the edge server, the computing resources of the terminal device, the length of the energy deficit queue of the edge server, and the length of the energy deficit queue of the terminal device; the length change rule of the energy deficit queue of the edge server is determined based on the energy consumption of the edge server processing tasks and the preset energy consumption budget, and the length change rule of the energy deficit queue of the terminal device is determined based on the energy consumption of the terminal device processing tasks and the preset energy consumption budget.

4. The method according to claim 3, characterized in that Based on the computing offloading strategy, a preset resource allocation method is used to determine an optimal resource allocation strategy, including: According to the state of the system and the computation offloading strategy, a preset resource allocation algorithm is used to solve problem P4 to obtain an optimal resource allocation strategy.

5. The method according to claim 4, characterized in that According to the state of the system and the computing offloading strategy, a preset resource allocation algorithm is used to solve problem P4 to obtain an optimal resource allocation strategy, including: Determining the computing resources required for the tasks offloaded by each terminal device according to the computing resource amount of each terminal device and the computing offloading decision; Determine the computing resources initially allocated by the edge server to each task based on the computing resources required by the tasks offloaded by each terminal device and the task arrival rate; Calculate the processing delay of each task from the corresponding terminal device to the edge server; For each terminal device offload task: Reducing and increasing the initially allocated computing resources by a preset resource adjustment amount respectively to obtain a first resource amount and a second resource amount; Calculate a corresponding first processing delay and a second processing delay respectively according to the first resource amount and the second resource amount; Determining an increase in the first processing delay according to the processing delay and the first processing delay; Determining a reduction amount of the second processing delay according to the processing delay and the second processing delay; Determine a minimum first processing delay increase according to the first processing delay increase corresponding to each task; Determine a maximum second processing delay reduction amount according to the second processing delay reduction amount corresponding to each task; If the maximum second processing delay reduction is greater than the minimum first processing delay increase, the initially allocated computing resources are adjusted for the task corresponding to the maximum second processing delay reduction and the task corresponding to the minimum first processing delay increase.

6. The method according to claim 5, characterized in that The resource adjustment amount is the minimum computing resource block that can be adjusted by the edge server; the adjusting of the initially allocated computing resources includes: For the task corresponding to the maximum second processing delay reduction amount, reducing the initially allocated computing resources by the resource adjustment amount; For the task corresponding to the minimum first processing delay increase, the initially allocated computing resources are increased by the resource adjustment amount.

7. The method according to claim 5, characterized in that The method for calculating the processing delay of each task from the corresponding terminal device to the edge server is as follows: in, Processing task delay for terminal device k, is the data transmission delay between the terminal device and the edge server, Processing task delay for edge servers, D k (t) is the processing delay of the t-th time slot task between the terminal device and the edge server.

8. The method according to any one of claims 1 to 5, characterized in that: The calculation offloading strategy is the offloading ratio of the terminal device to offload tasks to the edge server.

9. The method according to claim 1, characterized in that: The edge computing system model is constructed as follows: The edge computing system is modeled using the M / D / 1 queue model.

10. A device for joint optimization of computation offloading and resource allocation, characterized in that: include: A model building module, which is used to build an edge computing system model and determine the long-term optimization problem of computing offloading and resource allocation based on the model; A problem conversion module, used to convert the long-term optimization problem into an online optimization problem for determining a computation offloading strategy and a resource allocation strategy for each time slot by using a preset optimization method; A computing offloading strategy solving module, used for determining the computing offloading strategy for each time slot based on a reinforcement learning method for the online optimization problem; The resource allocation strategy solving module is used to determine the optimal resource allocation strategy based on the computing offloading strategy by using a preset resource allocation method.

Citation Information

Cited By

  • Computing task allocation method based on end-cloud fusion, computer equipment and medium

    CN120723488A