Partial task offloading and resource allocation method based on LSTM-DDPG

By adopting the mobile edge computing method based on LSTM-DDPG in the Internet of Vehicles, partial offloading of tasks and optimized allocation of resources are achieved, and the delay and energy consumption problems of computing-intensive applications in the Internet of Vehicles are solved, which significantly improves system performance.

CN115243220BActive Publication Date: 2025-05-16ZHONGYUN ZHIWANG DATA IND (CHANGZHOU) CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210861273.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-22
Publication Date
2025-05-16
Estimated Expiration
2042-07-22

AI Technical Summary

Technical Problem

In the Internet of Vehicles, the prior art is difficult to effectively solve the delay and energy consumption problems of computing-intensive and delay-sensitive applications on in-vehicle terminals with limited resources.

Method used

Using the mobile edge computing method based on LSTM-DDPG, by deploying LSTM neural network and DDPG algorithm in the Internet of Vehicles MEC network, partial offloading of tasks and optimized allocation of resources, weighing the delay and energy consumption of task processing.

Benefits of technology

In the MEC-enabled Internet of Vehicles scenario, the trade-off performance of task processing delay and energy consumption is significantly improved, and the average task delay and energy consumption are significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115243220B_ABST
    Figure CN115243220B_ABST
Patent Text Reader

Abstract

The present invention relates to a partial task offloading and resource allocation method based on LSTM-DDPG, comprising: creating a vehicle network MEC network model for partial task offloading and resource allocation; converting the partial task offloading and resource allocation problems into a reinforcement learning model; and introducing the LSTM neural network into the actor network and critic network of the DDPG algorithm. Compared with the existing DRL-based algorithm, the LSTM-DDPG-based mobile edge computing algorithm of the present invention solves the task offloading and resource allocation problems, can effectively achieve a trade-off between latency and energy consumption, and has significant performance improvements in terms of average task latency and average task energy consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical field:

[0001] The present invention relates to the field of deep reinforcement learning and vehicle networking technology, and in particular to a method for partial task offloading and resource optimization allocation of mobile edge computing based on LSTM-DDPG. Background technology:

[0002] With the development of intelligent transportation systems in recent years, the Internet of Vehicles has attracted widespread attention due to its role in many applications such as autonomous driving and intelligent traffic management. These computationally intensive and delay-sensitive applications require a large amount of computing resources for real-time processing, which poses a huge challenge to the resource-limited vehicle terminals.

[0003] In order to meet the needs of computing-intensive and latency-sensitive applications in the Internet of Vehicles, applying mobile edge computing (MEC) to the Internet of Vehicles scenario is considered to be a very promising solution. Mobile edge computing sinks the computing resources of the remote cloud to the network edge close to the vehicle user, allowing the vehicle to offload tasks to nearby MEC servers for storage and processing, which can effectively reduce the latency and energy consumption of vehicle applications.

[0004] Task offloading is one of the key issues in MEC. According to different processing methods of computing tasks, three computing task offloading decisions can be divided into local computing, full task offloading and partial task offloading. Among them, the partial task offloading strategy is to split the task and process it on the local and surrounding MEC servers at the same time. The computing task offloading performance depends largely on the wireless transmission of data offloading from the local vehicle to the MEC server. Therefore, an effective task offloading scheduling and resource allocation scheme is needed to improve system performance.

[0005] The Internet of Vehicles (MEC) can effectively solve the problem of limited computing resources on vehicle terminals, but there are still problems with multi-dimensional resource allocation caused by vehicle mobility and resource imbalance. Considering that the resource allocation decision in the Internet of Vehicles (MEC) scenario can be regarded as a Markov decision process (MDP), the reinforcement learning (RL) method is used to solve the task offloading and resource allocation problems in the Internet of Vehicles (MEC) network. Deep reinforcement learning (DRL) uses neural networks to fit strategies, which can effectively solve the problem of high-dimensional states.

[0006] DDPG is a deep reinforcement learning method that can generate continuous actions, realize continuous resource allocation, and improve the practicality of the system. Considering that the environmental state in the network model has correlation in the time dimension, the LSTM neural network that can effectively obtain the correlation of time series is incorporated into the network structure of DDPG, and the LSTM unit is used to obtain the correlation of historical information in the time dimension in the vehicle network MEC scenario, so as to perform state prediction and improve the performance of some task offloading decisions and resource allocation. Summary of the invention:

[0007] The technical problem to be solved by the present invention is to overcome the shortcomings of the prior art and provide a partial task offloading and resource allocation method based on LSTM-DDPG to achieve a trade-off between task processing delay and energy consumption in a MEC-enabled Internet of Vehicles system.

[0008] The present invention is achieved through the following technical solutions:

[0009] A partial task offloading and resource allocation method based on LSTM-DDPG, comprising the following steps:

[0010] First, consider a vehicle network MEC network model that performs partial task offloading and resource allocation. The network model contains a base station connected to all RSUs in the scene, which can collect status information of vehicles and RSUs. The RSUs (roadside units) deployed at intersections are equipped with MEC servers, which can provide computing power for task vehicles within the coverage area. The set of task vehicles is denoted as C H ={1,2,...,H}. The system operation process is divided into a series of frames with a time length of ξ, and the frame number is l∈N + At the beginning of each frame, the base station will perceive the network environment status from a global perspective and select appropriate partial task offloading and resource allocation strategies.

[0011] The task vehicle can divide the task into three parts, which are processed locally, offloaded to surrounding idle vehicles through V2V communication, and offloaded to the MEC server through V2I communication for processing. Vehicles may be blocked by buildings when performing V2V communication, so the influence of line-of-sight (LOS) communication and non-line-of-sight (NLOS) communication is considered in the network model. Considering the strict requirements of vehicle networking applications on latency and reliability, partial offloading and resource allocation of tasks are subject to the latency and reliability constraints of URLLC.

[0012] The network model in this article considers two transmission modes, namely V2V transmission and V2I transmission:

[0013] 1. V2I transmission:

[0014] In the lth frame, the signal-to-interference-noise ratio of the uplink when the task vehicle h (1≤h≤H) performs V2I task offloading It can be expressed as

[0015]

[0016] in, represents the transmission power of the task vehicle h when performing V2I communication in the lth frame, represents the V2I channel gain of the task vehicle h at the lth frame, σ 2 Represents the noise power. represents the interference caused by other mission vehicles in different RSUs to the V2I communication of mission vehicle h in the lth frame. When mission vehicle x interferes with the V2I communication of mission vehicle h, is 1, otherwise it is 0. Considering the influence of the finite code length (FBL) mechanism, the transmission rate (bit / s) when the task vehicle h performs V2I task offloading in the lth frame is

[0017]

[0018] Among them, V k represents the channel dispersion function, Q(·) represents the Gaussian Q function, ε represents the probability of data transmission decoding error, and n 0 Indicates the code length for uplink transmission.

[0019] 2. V2V transmission:

[0020] At the lth frame, the signal-to-interference-noise ratio of the uplink of the task vehicle h during V2V task offloading It can be expressed as

[0021]

[0022] in, Indicates the transmission power during V2V communication, represents the channel gain of the V2V channel, Indicates the interference caused by other mission vehicles to the V2V communication of this vehicle. represents the interference caused by other mission vehicles in different RSUs to the V2V communication of mission vehicle h. When mission vehicle y interferes with the V2V communication of mission vehicle h, is 1, otherwise it is 0. In the lth frame, the transmission rate (bit / s) of the task vehicle h when performing V2V task offloading is

[0023]

[0024] In the lth frame, the calculation of task data in the task vehicle h can be divided into three parts: local calculation, MEC server calculation and target idle vehicle calculation.

[0025] 1. Local computing latency is

[0026]

[0027] in, represents the amount of task data assigned to the task vehicle h for calculation in the first frame, N represents the number of CPU cycles required to calculate each bit of data, Represents the computing power of the mission vehicle h at the lth frame.

[0028] The energy consumption of local computing is

[0029]

[0030] Wherein, k′ represents the chip architecture coefficient, which depends on the chip architecture of the CPU.

[0031] 2. The delay of V2I transmission and calculation (MEC server calculation) is

[0032]

[0033] in, represents the amount of task data that is allocated to the MEC server for calculation by the task vehicle h in the first frame, Indicates the CPU computing resources allocated by the MEC server to the task data unloaded by the task vehicle h in the lth frame.

[0034] The energy consumption of V2I transmission and calculation is

[0035]

[0036] 3. The delay between V2V transmission and calculation (target idle vehicle calculation) is

[0037]

[0038] in, represents the amount of task data that is calculated by task vehicle h in the target idle vehicle at the lth frame, It represents the computing power of the target idle vehicle selected when the task vehicle h performs V2V task unloading in the lth frame.

[0039] The energy consumption of V2V transmission and calculation is

[0040]

[0041] Ignoring the delay and energy consumption in the process of transmitting the task results back to the task vehicle, the overall delay can be expressed as

[0042]

[0043] in, represents the part of the task processing delay that exceeds the task vehicle h in the lth frame. When the task processing delay exceeds one frame, the task will be discarded. The total energy consumption of task processing in task vehicle h in the lth frame can be expressed as

[0044]

[0045] Considering latency and energy consumption as optimization targets for task offloading decisions and resource allocation, in order to minimize the system cost, the optimization problem in this paper can be expressed as

[0046]

[0047] Among them, β represents the weight coefficient of energy consumption.

[0048] Second, the partial task offloading and resource allocation problems in the network model are transformed into a reinforcement learning model. Considering the time correlation of the environmental state in the scene, a partial task offloading decision and resource allocation method based on LSTM-DDPG is proposed.

[0049] Firstly, the partial task offloading and resource allocation problems in this paper are transformed into a reinforcement learning model, which mainly includes the design of state space, action space and reward function.

[0050] (1) State space: Before the agent makes an action decision, it must first obtain the current state input from the environment. in:

[0051] Indicates the mission data size of the mission vehicle at the lth frame;

[0052] Indicates the part of the task processing delay that exceeds the l-1 frame;

[0053] represents the V2I channel gain of the mission vehicle at the lth frame;

[0054] represents the V2V channel gain of the mission vehicle at the lth frame;

[0055] Indicates the computing power of the target idle vehicle corresponding to the task vehicle performing V2V unloading in the lth frame;

[0056] Represents the computing power of the mission vehicle at the lth frame.

[0057] (2) Action space: the agent’s action a l It is the partial offloading decision of tasks and the joint allocation of computing resources and communication resources, which can be expressed as

[0058]

[0059] (3) Reward function: transform action a l After being applied to the environment, the agent receives corresponding rewards or penalties. Considering that the goal of reinforcement learning is to maximize the reward, this paper uses the cost function consisting of frame length ξ minus delay and energy consumption to express the reward after the task vehicle h performs the corresponding partial task unloading and resource allocation decision

[0060]

[0061] In the lth frame, state s l and action a l The corresponding reward r(s l ,a l )for

[0062]

[0063] 3. Introduce LSTM neural network into the actor network and critic network of DDPG algorithm. In the actor network, add LSTM unit before the fully connected layer. In the critic network, part of the input is obtained by the input state through LSTM unit, and the other part is output by the action of the actor network.

[0064] In order to learn the partial task offloading and resource allocation strategy, the algorithm needs to be trained. First, the agent observes the environmental state s of the MEC network in the partial task offloading mode of this paper. l , the actor network generates an action a according to the current state l Action decision a l It includes task data allocation, transmission power allocation, and MEC server CPU resource allocation among local, MEC server, and target idle vehicles. Based on the actions performed in the environment, the environment provides the agent with a reward r(s l ,a l ) and evolves to the state s at the next moment l+1 . The data group (current state s l 、Action a l , reward r(s l ,a l ), next frame state s l+1) is stored in the experience replay buffer. When enough experience data sets are stored in the experience replay buffer, a small batch of data sets are randomly sampled from the experience replay buffer for training to update the parameters of the actor and critic networks.

[0065] Compared with the prior art, the present invention adopts the above technical solution and has the following technical effects:

[0066] In the MEC-enabled Internet of Vehicles scenario, compared with the existing DRL-based algorithm, the LSTM-DDPG-based mobile edge computing method of the present invention solves the task offloading and resource allocation problems, can effectively achieve the trade-off between latency and energy consumption, and has significant performance improvements in both average task latency and average task energy consumption. Description of the drawings:

[0067] Figure 1 It is a schematic diagram of the vehicle networking MEC network model of the present invention;

[0068] Figure 2 is the average task delay under different energy consumption weight coefficients;

[0069] Figure 3 is the average task energy consumption under different energy consumption weight coefficients. Specific implementation method:

[0070] The preferred embodiments of the present invention are described in detail below in conjunction with the accompanying drawings so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more definite definition of the protection scope of the present invention.

[0071] Figure 1 is the network model of the present invention, considering a vehicle network MEC network model that performs partial task offloading and resource allocation. The network model contains a base station connected to all RSUs in the scene, which can collect status information of vehicles and RSUs. The RSUs deployed at the intersection are equipped with MEC servers, which can provide computing power for the task vehicles within the coverage area. The set of task vehicles is denoted as C H ={1,2,...,H}. The system operation process is divided into a series of frames with a time length of ξ, and the frame number is l∈N + At the beginning of each frame, the base station will perceive the network environment status from a global perspective and select appropriate partial task offloading and resource allocation strategies.

[0072] In order to make full use of the available computing resources in the surrounding area, the task vehicle can divide the task into three parts, which are processed locally, offloaded to the surrounding idle vehicles through V2V communication, and offloaded to the MEC server through V2I communication for processing. Assuming that the MEC server contains multiple CPU cores, when the vehicle offloads the task data to the MEC server through V2I communication, the MEC server will provide a separate CPU core for data processing for each task.

[0073] Since vehicles may be blocked by buildings during V2V communication, the influence of line-of-sight (LOS) communication and non-line-of-sight (NLOS) communication is considered in the network model. Considering the strict requirements of applications on latency and reliability in the Internet of Vehicles scenario, partial offloading and resource allocation of tasks are subject to the latency and reliability constraints of URLLC. When the total latency of task processing exceeds one frame, it means that the task has failed, and the excess latency will have a significant impact on the task processing of the next frame.

[0074] Figure 2 and Figure 3 They are the average mission delay and average mission energy consumption under different energy consumption weight coefficients. The simulation scenario includes two blocks with line-of-sight and non-line-of-sight channels, and two RSUs are deployed at the corresponding intersections. The number of mission vehicles H is set to 4, and the average speed of the vehicle is set to 36km / h. There are 5 idle vehicles randomly distributed around each mission vehicle, and the positions of all vehicles are updated every 100ms. The maximum transmission power P of the mission vehicle max Set to 23dBm. During simulation, the task vehicle will select the nearest idle vehicle as the target idle vehicle for V2V task offloading. The V2V and V2I channel models of the task vehicle refer to the channel settings of the urban road scene in the protocol 3GPP TR 36.885.

[0075] In order to better learn the parameters of the actor network and the critic network, the learning rates are set to 0.0001 and 0.001, respectively. The soft update rate of the target network is set to 0.005. The capacity of the experience replay buffer is set to 20,000, and the batch sampling size is set to 128. In addition, the training phase went through 80 rounds, each containing 2,000 frames. The simulation compares the task offloading and resource allocation schemes based on the following four algorithms: full offloading to MEC server (ALL-MEC), deep Q learning (DQN), DQN with LSTM network (LSTM-DQN), and DDPG.

[0076] The simulation results show that there is an obvious trade-off between the latency and energy consumption of task processing. As the energy consumption weight coefficient β increases, the latency increases while the energy consumption decreases. In addition, under different energy consumption weight coefficients, the performance of the DDPG-based algorithm is always higher than that of the DQN-based algorithm in terms of both latency and energy consumption. The LSTM-DDPG algorithm that introduces the LSTM neural network has significant performance improvements in latency and energy consumption compared to other comparison algorithms.

[0077] The above-mentioned embodiments only express several implementation methods of the present invention, and the description is relatively specific and detailed, but it cannot be understood as limiting the scope of the invention patent. It should be pointed out that for ordinary technicians in this field, several modifications and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention.

Claims

1. A partial task offloading and resource allocation method based on LSTM-DDPG, characterized in that: The steps include: First, create a vehicle network MEC network model for partial task offloading and resource allocation; the network model includes a base station connected to all RSUs in the scene, which can collect status information of vehicles and RSUs; RSUs deployed at intersections are equipped with MEC servers, which can provide computing power for task vehicles within the coverage area; the set of task vehicles is recorded as ; Divide the system operation process into a series of time lengths The frame number is Indicates that at the beginning of each frame, the base station will perceive the network environment status from a global perspective and select appropriate partial task offloading and resource allocation strategies; In step 1, the task vehicle can divide the task into three parts, which are processed locally, offloaded to surrounding idle vehicles through V2V communication, and offloaded to the MEC server through V2I communication for processing; The specific offloading to the MEC server through V2I communication includes: No. Frame, mission vehicle When performing V2I task offloading, the signal-to-interference-to-noise ratio of the uplink It can be expressed as , in, Indicated in Frame time mission vehicle Transmit power during V2I communication, Indicates Frame time mission vehicle The V2I channel gain, represents the noise power, Indicates Other mission vehicles in different RSUs in the frame time are different from the mission vehicle Interference caused by V2I communication; When the mission vehicle Mission Vehicles When interference occurs in V2I communication, is 1, otherwise it is 0; considering the influence of the finite code length (FBL) mechanism, the Frame time mission vehicle The transmission rate when V2I task offloading is , in, represents the channel dispersion function, represents the Gaussian Q function, represents the probability of data transmission decoding error, Indicates the code length for uplink transmission; Unloading to surrounding idle vehicles through V2V communication specifically includes: No. Frame time, mission vehicle When the V2V task is offloaded, the signal-to-interference-to-noise ratio of the uplink It can be expressed as , in, Indicates the transmission power during V2V communication, represents the channel gain of the V2V channel, Indicates the interference caused by other mission vehicles to the V2V communication of this vehicle. Indicates other mission vehicles in different RSUs to the mission vehicle Interference caused by V2V communication; When the mission vehicle Mission Vehicles When interference occurs in V2V communication, is 1, otherwise it is 0; Frame time, mission vehicle The transmission rate when offloading V2V tasks is , No. Frame time, mission vehicle The calculation of task data can be divided into three parts: local calculation, MEC server calculation and target idle vehicle calculation; (1) The latency of local computing is , in, Indicates Frame time mission vehicle The amount of task data assigned to the vehicle for calculation, Indicates the number of CPU cycles required to calculate each bit of data. Indicates Frame time mission vehicle The computing power of The energy consumption of local computing is , in, Indicates the chip architecture coefficient, which depends on the chip architecture of the CPU; (2) The latency calculated by the MEC server is , in, Indicates Frame Mission Vehicle The amount of task data allocated for calculation in the MEC server, Indicates Frame time MEC server is the mission vehicle The CPU computing resources allocated to the offloaded task data, The energy consumption of MEC server calculation is ; (3) The calculation delay of the target idle vehicle is , in, Indicates Frame time mission vehicle The amount of task data allocated to be calculated in the target idle vehicle, Indicates Frame time mission vehicle The computing power of the target idle vehicle selected for V2V task offloading, The energy consumption of the target idle vehicle is calculated as , Ignoring the delay and energy consumption in the process of transmitting the task results back to the task vehicle, the overall delay can be expressed as , in, Indicates Frame time mission vehicle The part of the task processing delay that exceeds the limit; When the delay of task processing exceeds one frame, the task will be discarded. Frame time mission vehicle The overall energy consumption of task processing can be expressed as , Considering latency and energy consumption as optimization targets for task offloading decisions and resource allocation, in order to minimize the system cost, the optimization problem is expressed as , in, Represents the weight coefficient of energy consumption; Second, transform partial task offloading and resource allocation problems into reinforcement learning models, including the design of state space, action space, and reward function; (1) State space: Before the agent makes an action decision, it must first obtain the current state input from the environment. ,in , indicating the The mission data size of the mission vehicle in the frame time; , indicating the The part of the frame time task processing delay that exceeds the limit; , indicating the V2I channel gain of the mission vehicle at frame time; , indicating the V2V channel gain of the mission vehicle at frame time; , indicating the The computing power of the target idle vehicle corresponding to the task vehicle performing V2V unloading during the frame time; , indicating the The computing power of the mission vehicle during the frame time; (2) Action space: the actions of the agent It is the partial offloading decision of tasks and the joint allocation of computing resources and communication resources, which can be expressed as ; (3) Reward function: transform the action After being applied to the environment, the agent receives corresponding rewards or penalties; the frame length is used The mission vehicle is described by subtracting the cost function consisting of delay and energy consumption Rewards after executing the corresponding partial task offloading and resource allocation decision , No. In the frame, the status and actions Corresponding rewards for ; 3. Introduce LSTM neural network into the actor network and critic network of DDPG algorithm; in the actor network, add LSTM unit before the fully connected layer; in the critic network, part of the input is obtained by the input state through LSTM unit, and the other part is output by the action of the actor network.

2. The method for partial task offloading and resource allocation based on LSTM-DDPG according to claim 1, characterized in that: In step three, the algorithm needs to be trained; The environmental state of the MEC network in the partial task offloading mode observed by the agent , the actor network generates an action based on the current state ; Action decision Including task data allocation locally, MEC server and target idle vehicles, transmission power allocation and MEC server CPU resource allocation; based on the actions performed in the environment, the environment provides rewards to the agent , and evolves to the state of the next moment ; The current state of the data group ,action ,award , next frame status Stored in the experience replay buffer; when enough experience data sets are stored in the experience replay buffer, a small batch of data sets are randomly sampled from the experience replay buffer for training to update the parameters of the actor and critic networks.

Citation Information

Patent Citations

  • Joint optimization method for task unloading calculation cost and time delay in Internet of Vehicles

    CN110650457A

  • Vehicle-mounted task unloading method based on deep reinforcement learning in edge computing environment

    CN114615265A