Task unloading strategy determination and resource allocation method for vehicle edge computing network

By building a digital twin model in the vehicle edge computing network and using deep learning technology to train the Actor network and Critic network, the problem of high latency of the Internet of Vehicles is solved, and flexible allocation of resources and improved computing efficiency is achieved.

CN120075227APending Publication Date: 2025-05-30HEBEI UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510313707.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When facing multiple tasks, existing vehicle edge computing networks cannot flexibly allocate computing resources according to the size of each task, resulting in a high delay in the Internet of Vehicles.

Method used

By building a digital twin model, using the Actor network and Critic network for training, the optimal task offloading strategy and resource allocation scheme are determined to minimize total latency and resource costs.

Benefits of technology

It realizes the rational and flexible allocation of computing resources in multi-vehicle and multi-task scenarios, reduces vehicle network delays, and improves the data computing efficiency and intelligence level of the Internet of Vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005315310940000031
    Figure BDA0005315310940000031
  • Figure BDA0005315310940000038
    Figure BDA0005315310940000038
  • Figure BDA0005315310940000058
    Figure BDA0005315310940000058
Patent Text Reader

Abstract

The invention provides a task unloading strategy determination and resource allocation method for a vehicle edge computing network. The method comprises the following steps: S1, constructing a digital twin model; s2, determining a state space and an action space; s3, determining a reward function and a discount utility; s4, initializing a discount factor, an experience playback pool and parameters in the Actor network and the Critic network; s5, the intelligent agent executes an action according to the current strategy and the current state, and the environment returns a next state and reward and stores the state and the reward into an experience playback pool; s6, repeating the step S5 to reach a target round; s7, randomly sampling a batch from the experience playback pool, and training an Actor network and a Critic network; s8, repeating the steps S5-S7 until a stop condition is met; s9, according to the to-be-calculated base station and the vehicle information received by the to-be-calculated base station, establishing an edge Internet of Vehicles environment based on digital twinning; and determining a current state, and inputting the trained Actor network and Critic network to obtain a current unloading ratio and a current estimated computing resource. According to the invention, the vehicle task processing time delay is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a task offloading method, specifically a method for determining a task offloading strategy and resource allocation in a vehicular edge computing network. Background Art

[0002] With the development of the Internet of Things, more and more in-vehicle applications such as autonomous driving, navigation, and video have been installed in modern vehicles, resulting in an increase in the number of vehicle computing tasks that need to be solved. However, due to the insufficient computing power of the vehicle itself, it may not be able to process these computing tasks generated by in-vehicle applications in a timely manner. Although cloud computing can solve the problem of insufficient vehicle resources, its long-distance deployment will lead to a large delay. Vehicular Edge Computing (VEC), as a promising technology, can offload computing tasks to nearby VEC servers or roadside units and assist in processing computing tasks by using the computing resources of the VEC servers themselves. However, VEC also faces new problems such as vehicle mobility and environmental dynamics.

[0003] Digital Twin (DT), as an emerging technology, promotes the mapping, iterative operation, and interaction between virtual models and physical objects. Based on historical data and real-time status, it uses machine learning and Internet of Things technologies to establish digital copies for physical entities, and can enhance real-time interaction and achieve close monitoring through reliable communication between the digital space and the physical system. In terms of vehicle task processing, applying digital twin technology to the VEC task processing scenario can enhance data mining, simulation, and analysis of the vehicle network.

[0004] The existing research on the combination of DT and vehicular edge computing network is still in its infancy and there is still much room for improvement. First of all, the research on using DT for prediction in VEC to optimize offloading decisions is still lacking. Using DT for prediction in VEC to optimize offloading decisions cannot flexibly allocate computing resources according to the size of each task when facing multiple tasks, resulting in a high delay in the vehicle network. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for determining a task offloading strategy and resource allocation in a vehicular edge computing network to solve the problem of high delay in the vehicle network.

[0006] The present invention is implemented as follows:

[0007] A method for determining a task offloading strategy and resource allocation in a vehicular edge computing network includes the following steps:

[0008] S1. Build a digital twin model, where the digital twin model includes a base station equipped with a VEC server and N vehicles as users covered by the base station. Each vehicle is an agent, and the states of the agents and the state of the digital twin model are the environment; initialize the time and environmental state in the digital twin model;

[0009] S2. Determine the computing requests, channel gains, task offloading methods, delay calculation methods, estimation errors of estimated computing resources and actual computing resources of the vehicles according to the digital twin model; calculate the estimated computing resources of the tasks according to the task offloading methods; determine the computing requests, channel gains and estimation errors of the vehicles as the state space; determine the task offloading methods and estimated computing resources as the action space;

[0010] S3. Determine the total delay for completing the tasks according to the channel gains, task offloading methods, delay calculation methods, estimated computing resources of the tasks, and the estimation errors of the estimated computing resources and actual computing resources; determine the resource cost required for offloading the tasks according to the task offloading methods; determine the task offloading decision problem according to the total delay for completing the tasks and the resource cost required for offloading the tasks, and use the task offloading decision problem as the reward function; calculate the discounted utility according to the reward function and discount factor at the previous moment;

[0011] S4. Establish an Actor network and a Critic network, and initialize the parameters in the policy, discount factor, experience replay pool, Actor network and Critic network respectively;

[0012] S5. The agents execute actions according to the policy and the current state, the environment returns the next state and the reward, and store the current state, action, reward and next state in the experience replay pool;

[0013] S6. Repeat step S5 until the target round is reached;

[0014] S7. Randomly sample a batch from the experience replay pool to train the Actor network and the Critic network;

[0015] S8. Repeat steps S5 - S7 until the stopping condition is reached;

[0016] S9. Based on the base station to be calculated and the vehicle information received by the base station to be calculated, establish an edge vehicle networking environment based on digital twin; according to the edge vehicle networking environment of digital twin, determine the current task and the current position of the vehicle; according to the current position of the vehicle, calculate the current channel gain; randomly generate the current estimation error of the estimated computing resources and the actual computing resources; use the current task, the current position of the vehicle, the current channel gain, and the current estimation error as the current state, and input them into the trained Actor network and Critic network to obtain the current offloading method and the current estimated computing resources.

[0017] Further, the N vehicles are represented by {1, 2,..., N}; the time is discretized into several time slots, and the set of time slots is represented by t ∈ {1, 2,..., T}; within a single time slot, the vehicle user completes a task offloading decision and a position movement.

[0018] Further, the computing request generated by the vehicle within time slot t is denoted, where is the computing request generated by vehicle n, and are the K tasks generated by vehicle n.

[0019] Further, the task offloading method is represented by the offloading ratio If then the task is calculated locally; if ∈(0, 1), then the task is partially offloaded to the base station for calculation; if then the task is calculated at the base station.

[0020] Further, the delay calculation method includes the local calculation delay method and the edge calculation delay method;

[0021] The total delay for completing each task is the maximum of the local calculation delay and the edge calculation delay.

[0022] Further, the calculation formula for the edge calculation delay is:

[0023]

[0024] where is the offloading ratio, is the size of the k-th task generated by vehicle n, is the information transmission rate between vehicle n and the base station, is the information transmission rate between vehicle n and the base station, is the estimation error of the actual computing resources and the estimated computing resources of the k-th task of vehicle n, is the time difference between the estimated completion time and the actual completion time.

[0025] Furthermore, the calculation formula for the channel gain between the vehicle and the base station is:

[0026]

[0027] where represents small-scale fading, represents the sum of large-scale path fading and shadowing.

[0028] Furthermore, the task offloading decision problem is:

[0029]

[0030] where is the total delay for vehicle n to process the k-th task, is the resource cost for vehicle n to process the k-th task, is the estimated error between the estimated computing resource and the actual computing resource of the k tasks of vehicle n, is the estimated computing resource of the k tasks of vehicle n, is the maximum computing resource of the VEC server, represents the maximum processing time limit of the k-th task generated by vehicle n, is the signal transmission power between vehicle n and the base station, is the maximum transmission power of vehicle n, is the task offloading method.

[0031] Furthermore, the calculation formula for determining the discount factor is:

[0032]

[0033] where t 0 is the previous moment, the discount factor γ ∈ (0, 1), and rn is the reward function.

[0034] The present invention utilizes digital twin technology and mobile edge computing technology, improves the real-time performance and accuracy of data, enhances the data computing efficiency and intelligent level of the vehicle network, and alleviates the high latency problem of the vehicle network. The present invention trains the Actor network and the Critic network through a digital twin model, with the goal of minimizing the total delay and the minimum payment cost, and trains with multiple agents at the same moment and multiple computing tasks of each agent as the state space. Considering partial offloading of tasks, during the training process, the minimum total delay and the minimum payment cost are used as the reward function until the offloading ratio and the estimated computing resource corresponding to the minimum total delay and the minimum payment cost are obtained. When facing multiple vehicles and multiple tasks for each vehicle, the computing resources are reasonably and flexibly allocated to reduce the latency of the vehicle network.

[0035] The task offloading and resource allocation method for a vehicle edge computing network constructed based on the maximum vehicle delay and payment price proposed by the present invention can reasonably deploy vehicle digital twins in edge servers, reasonably allocate edge computing resources, and effectively reduce the task processing delay and cost of vehicles. The present invention allows agents to consider the strategies of other agents during training, is more applicable to complex interaction environments, and helps to make better decisions in environments with other agents. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 is a flowchart of the method of the present invention.

[0037] Figure 2 is a system model diagram of an ultra-dense mobile edge computing network. DETAILED DESCRIPTION OF THE INVENTION

[0038] The present invention will be further described in detail below with reference to the drawings and embodiments.

[0039] A method for determining a task offloading strategy and allocating resources for a vehicle edge computing network provided by the present invention specifically includes the following steps:

[0040] S1. Construct a digital twin model.

[0041] As Figure 1 shown, the digital twin model includes a base station provided with a VEC server and N vehicles as users covered by the base station. Each vehicle is an agent, and the state of the agent and the state of the digital twin model are the environment; initialize the time and environmental state in the digital twin model.

[0042] The N vehicles are represented by {1, 2,..., N}; the time is discretized into several time slots, and the set of time slots is represented by t ∈ {1, 2,..., T}; within a single time slot, the vehicle user completes a task offloading decision and a location movement.

[0043] S2. Determine the computing request, channel gain, task offloading method, delay calculation method, estimation error of estimated computing resources and actual computing resources of the vehicle according to the digital twin model; calculate the estimated computing resources of the task according to the task offloading method; determine the computing request, channel gain and estimation error of the vehicle as the state space; determine the task offloading method and estimated computing resources as the action space.

[0044] The computing request generated by the vehicle within time slot t is represented by , where is the computing request generated by vehicle n, is the K tasks generated by vehicle n, is the position of vehicle n. Each vehicle is an agent, and the external world where the agent is located is the environment; initialize the time and environmental state in the digital twin model.

[0045] Among them, Among them, represents the size of the k-th task generated by vehicle n, represents the number of CPU cycles required to calculate the k-th task per unit bit of vehicle n, represents the maximum processing time limit of the k-th task generated by vehicle n.

[0046] The vehicle also needs to send its own computing resource F n and speed

[0047] For the position of the vehicle, the position of the vehicle at the initial moment is sent by the vehicle. At subsequent moments, the position of the vehicle can be directly sent by the vehicle or calculated.

[0048] The specific calculation method is as follows: Establish a spatial orthogonal coordinate system with the base station as the origin. The positive direction of the x-axis is the direction in which the vehicle travels along the lane, that is, east. The positive direction of the y-axis is south, and the positive direction of the z-axis is along the direction of the base station antenna. Vehicle n located in lane j at time slot t can be expressed as Among them, X i,j (t) represents the horizontal ordinate of vehicle n traveling on lane j in time slot t, and Y i,j (t) represents the ordinate of vehicle n. When the value of the size τ of each time slot is small enough, the position of the vehicle in each time slot can be approximately considered constant. At the same time, the position of the vehicle in the current time slot t is related to the position in the previous time slot t - 1. Therefore, the abscissa X n,j (t) of vehicle n traveling on lane j in time slot t can be further expressed as: Among them represents the driving speed of vehicle n.

[0049] Randomly generate the estimation error between the estimated computing resource and the actual computing resource of the k tasks of vehicle n Estimation error is generally set between -0.5 and 0.5.

[0050] In actual calculation, the edge vehicle networking environment based on the digital twin is deployed on the cloud server. Ignoring the communication time between the edge vehicle networking environment and the cloud server and the base station respectively, the vehicle can communicate with the base station through wireless connection. Without considering the influence of environmental factors on path loss, the calculation formula for the channel gain between the vehicle and the base station is:

[0051]

[0052] Among them, is small-scale fading, is the sum of large-scale path fading and shadowing, is large-scale path fading, is shadowing, L is the decorrelation distance, is noise with a mean of 0 and a standard deviation of 8.

[0053] The task offloading method is represented by the offloading ratio When , the task is computed locally; when , the task is partially offloaded to the base station for computing; when , the task is computed at the base station.

[0054] The delay calculation method includes the local computing delay method and the edge computing delay method.

[0055] Define the state space s n (t) as:

[0056] s n (t) = {D t , Δf est-t , υ pos-t , g t}

[0057] Among them, D t is the current tasks of N vehicles received at the current moment, is the K current tasks of vehicle n, represents the size of the k-th task generated by vehicle n, represents the number of CPU cycles required to compute the k-th task generated by vehicle n per unit bit, represents the maximum processing time limit of the k-th task generated by vehicle n; Δf est-t is the current estimation error between the estimated computing resources and the actual computing resources of all tasks at the current moment, is the current estimation error between the estimated computing resources and the actual computing resources of the k-th task randomly generated by vehicle n at the current moment, υ po-s is the positions of all vehicles at the current moment, is the current position of vehicle n, is the current channel gain between vehicle n and the base station.

[0058] Define the action space as a n(t) is as follows: Among them, w t is the offloading ratio of all tasks of all vehicles at time t, is the offloading ratio of the k-th task of vehicle n, is the estimation error of all tasks of all vehicles at time t, is the estimation error of the k-th task of vehicle n at time t.

[0059] S3. Determine the total delay to complete the task according to the channel gain, task offloading method, delay calculation method, estimated computing resources of the task, and the estimation error between the estimated computing resources and the actual computing resources. Determine the resource cost required for offloading tasks according to the task offloading method; determine the task offloading decision problem according to the total delay to complete the task and the resource cost required for offloading tasks, and use the task offloading decision problem as the reward function; calculate the discounted utility according to the reward function and the discount factor at the previous moment.

[0060] For different task offloading methods, the total delay to complete the computing task is different.

[0061] The present invention considers applying Orthogonal Frequency Division Multiple Access (OFDMA) to N V2I links. During the information transmission process, there will be various interferences. As a general noise model, Gaussian white noise can be used to describe the noise interference in the channel. Therefore, the information transmission rate between vehicle n and the base station in time slot t can be expressed as:

[0062]

[0063] Among them, B is the available bandwidth allocated by the base station to vehicle n, is the signal transmission power between vehicle n and the base station in time slot t, and σ 2 is the power of Gaussian white noise.

[0064] Local computing delay The calculation formula of is:

[0065]

[0066] Among them, F n represents the computing resources of vehicle n itself.

[0067] Since there is a mapping relationship between the vehicle and the edge vehicle networking environment, there may be positive or negative estimation errors between the estimated computing resources and the actual situation The time difference between the estimated completion time and the true completion time The calculation formula is:

[0068]

[0069] Among them, represents the estimated computing resources of vehicle n for task D n,k .

[0070] Edge computing delay The calculation formula is:

[0071] For task D n,k , if there is partial offloading processing and local processing, it can be considered that the two are executed in parallel. Therefore, the total delay n,k required to complete D can be expressed as:

[0072] When the k-th task of vehicle n decides to offload, it needs to pay the resource cost (i.e., price) to the VEC server. The calculation formula for the price to be paid for the k-th task of vehicle n is:

[0073]

[0074] Among them, ε n,k is the unit resource cost to be paid to the VEC server.

[0075] Taking the problem of minimizing the total consumption of K tasks generated by each vehicle as the task offloading decision problem, it is expressed as:

[0076]

[0077] Among them, is the maximum computing resource of the VEC server, is the maximum transmission power of vehicle n.

[0078] Determine the reward function r n (t), among which,

[0079] Determine the discount factor as:

[0080]

[0081] Among them, t 0 represents the previous moment, and γ ∈ (0, 1) represents the discount factor in the process.

[0082] By maximizing the long-term cumulative reward of each Agent Each agent can obtain the optimal execution strategy and resource allocation plan.

[0083] S4. Establish the Actor network and the Critic network, and initialize the parameters in the strategy, discount factor, experience replay pool, Actor network, and Critic network respectively.

[0084] Establish the Actor network and the Critic network. The Actor network is responsible for deciding what action to take in a given state, and the Critic network evaluates the quality of the action taken by the Actor network. Assign the initial values to the discount factor γ, the strategy and the parameters in the Actor network and the Critic network. The initialization of the experience replay pool is to determine a certain storage size and create an empty list to store information such as the current state, action, reward, and next state.

[0085] S5. The agent executes an action according to the strategy and the current state, the environment returns the next state and the reward, and the current state, action, reward, and next state are stored in the experience replay pool.

[0086] As Figure 2 shown, the initial state of the agent can be set by the technician. The agent executes action a n according to the strategy and the state, and the environment returns the next state s n ′ and the reward r n , and stores the current state s n , action a n , reward r n , and the next state s n ′ in the experience replay pool.

[0087] S6. Repeat step S5 until the target number of rounds is reached.

[0088] Repeat step S5 according to the set target number of rounds. At a certain moment, the current state s n , action a n , reward r n , and the next state s n ′ can be obtained in the experience replay pool until the target number of rounds or until the storage size of the experience replay pool is reached, and the current state s n , action a n , reward r n , and the next state s n ′ corresponding to multiple moments are obtained.

[0089] S7. Randomly sample a batch from the experience replay pool to train the Actor network and the Critic network.

[0090] Randomly sample a batch from the experience replay pool, which is the current state, action, reward, and next state at at least two moments.

[0091] The Critic network can access the information of all agents, including the state and action, which allows it to accurately evaluate the expected return of each action. In the execution phase, the Actor network of each agent can only make decisions based on its own local observations. The Q-value function Q(s,a) describes the expected sum of rewards that an agent can obtain in a series of future interactions after executing action a when it is in state s. The sum of rewards discounts future rewards according to a certain discount factor to reflect the characteristic that recent rewards are more important than distant rewards.

[0092] The estimating Actor network calculates the current state of an agent in the experience replay pool and outputs the current action. The target Actor network calculates the next state of an agent in the experience replay pool and outputs the next action; the estimating Critic network calculates the Q-values by sequentially calculating the current states and current actions of all agents; for a single agent n, the calculation formula is:

[0093]

[0094] where is for calculating the expected value, are the parameters used to estimate the Critic network.

[0095] The target Critic network calculates the Q'-values by sequentially calculating the next states and next actions of all agents. For a single agent n, the calculation formula is:

[0096]

[0097] where is the Q-value output by the target Critic network Q' at state s' and action μ'(s'), is the target Q-value used to train the target Critic network, are the parameters used for the target Critic network.

[0098] To update the parameters, the calculation formula of the loss function is:

[0099]

[0100] where

[0101] To minimize the loss function, the Critic network parameters are updated by stochastic gradient descent The calculation formula is:

[0102]

[0103] Due to the distributed execution of the Actor network deployed on each vehicle, Agent n takes actions by observing the local state. For the parameter update method of the Actor network, the gradient descent method is selected:

[0104]

[0105] where, represents the action policy taken by Agent n based on the current state.

[0106] During the training process, to ensure the stability of the algorithm, a soft update method is adopted for the parameters of the target Actor network and the target Critic network:

[0107]

[0108] where η ∈ (0, 1) represents the parameter update rate, are the parameters of the estimated Critic network, are the parameters of the target Critic network, are the updated parameters, As the updated parameters of the target Critic network, are the parameters of the estimated Actor network, are the parameters of the target Actor network, As the updated parameters of the target Critic network.

[0109] S8. Repeat steps S5 - S7 until the stop condition is reached.

[0110] The stop conditions include: the loss function and the reward start to converge, and the target round is reached, then the training can be stopped. When repeating steps S5 - S7, the experience values of this round can replace the experience values of the previous round, that is, the current state s n at different moments of this round, the action a n the reward r n and the next state s n ′ replace the current state s n at different moments of the previous round, the action a n the reward r n and the next state s n ′.

[0111] S9. Based on the base station to be calculated and the vehicle information received by the base station to be calculated, establish an edge vehicle networking environment based on digital twin; according to the edge vehicle networking environment of digital twin, determine the current task and the current position of the vehicle; according to the current position of the vehicle, calculate the current channel gain; randomly generate the current estimation error of the estimated computing resource and the actual computing resource; use the current task, the current position of the vehicle, the current channel gain, and the current estimation error as the current state, and input them into the trained Actor network and Critic network to obtain the current offloading method and the current estimated computing resource.

[0112] The vehicle information received by the base station to be calculated is the current task, the current position, and the current computing resource of the vehicle itself that sends a computing request to the base station to be calculated. The edge vehicle networking environment based on digital twin is deployed on the cloud server.

[0113] Use the current task, the current estimated computing resource and the current estimation error between the current actual computing resource, the current position of the vehicle, and the current channel gain between the vehicle and the base station received by the base station to be calculated as the current state, and the current state s n (t) is expressed as: s n (t) = {D t , Δf est-t , υ pos-t , g t} where D t is the current task of N vehicles received at the current moment, is the K current tasks of vehicle n, represents the size of the kth task generated by vehicle n, represents the number of CPU cycles required to calculate the kth task generated by vehicle n per unit bit, represents the maximum processing time limit of the kth task generated by vehicle n; Δf est-t is the current estimation error between the estimated computing resource and the actual computing resource of all tasks at the current moment, is the current estimation error between the estimated computing resource and the actual computing resource of the kth task of vehicle n randomly generated at the current moment, υ pos-t is the position of all vehicles with tasks at the current moment, is the current position of vehicle n, is the current channel gain between vehicle n and the base station.

[0114] The vehicle also needs to send its own computing resource Fn and speed

[0115] The calculation method of the parameters in the current state is the same as that of the parameters in step S2.

[0116] The edge vehicle Internet of Things environment deployment based on digital twins is on the cloud server. Ignoring the communication time between the edge vehicle Internet of Things environment and the cloud server and the base station respectively, the vehicle can communicate with the base station through wireless connection. Without considering the influence of environmental factors on path loss, the calculation formula for the channel gain between the vehicle and the base station is:

[0117]

[0118] where is small-scale fading, is the sum of large-scale path fading and shadow, is large-scale path fading, is shadow, L is the decorrelation distance, is noise with a mean of 0 and a standard deviation of 8. The small-scale fading at the initial moment is randomly generated, v n is the speed of vehicle n.

[0119] Input the current state into the Actor network and the Critic network to obtain the current action The current task offloading method and the current estimated computing resources can be obtained, and the base station to be calculated uses the current estimated computing resources to calculate the tasks of the vehicle.

Claims

1. A method for determining a task offloading strategy and allocating resources for a vehicle edge computing network, characterized in that: The following steps are involved: S1. Construct a digital twin model, wherein the digital twin model includes a base station provided with a VEC server and N vehicles covered by the base station as users, each vehicle as an intelligent agent, and the state of the intelligent agent and the state of the digital twin model as the environment; Initialize the time and environment states in the digital twin model; S2. Determine the vehicle's computing request, channel gain, task offloading method and delay calculation method, estimated computing resources and estimated error of actual computing resources based on the digital twin model; Calculate the estimated computing resources of the task according to the task offloading method; determine the computing request, channel gain and estimation error of the vehicle as a state space; Determine the task offloading methods and estimated computing resources as an action space; S3. Determine the total delay to complete the task based on the channel gain, task offloading method, delay calculation method, estimated computing resources of the task, and estimated error between the estimated computing resources and the actual computing resources; According to the task offloading method, determine the resource cost required for offloading the task; according to the total delay to complete the task and the resource cost required for offloading the task, determine the task offloading decision problem, and use the task offloading decision problem as the reward function; according to the reward function and discount factor at the previous moment, calculate the discounted utility; S4. Establish the Actor network and the Critic network, and initialize the strategy, discount factor, experience replay pool, and parameters in the Actor network and the Critic network respectively; S5. The agent performs actions according to the strategy and current state, and the environment returns the next state and reward. The current state, action, reward, and next state are stored in the experience replay pool. S6. Repeat step S5 until the target round is reached; S7. Randomly sample a batch from the experience replay pool and train the Actor network and the Critic network. S8. Repeat steps S5-S7 until the stop condition is reached; S9. Establish an edge vehicle network environment based on digital twins according to the base station to be calculated and the vehicle information received by the base station to be calculated; determine the current task and the current position of the vehicle according to the edge vehicle network environment of the digital twin; calculate the current channel gain according to the current position of the vehicle; randomly generate the current estimated computing resources and the current estimated error of the actual computing resources; take the current task, the current position of the vehicle, the current channel gain and the current estimated error as the current state, input them into the trained Actor network and Critic network, and obtain the current unloading mode and the current estimated computing resources.

2. The method for determining task offloading strategy and allocating resources for a vehicle edge computing network according to claim 1 is characterized in that: The N vehicles are represented by {1,2,...N,}; time is discretized into a number of time slots, and the set of time slots is represented by t∈{1,2,...T,}; within a single time slot, the vehicle user completes a task offloading decision and position movement.

3. The method for determining task offloading strategy and allocating resources for a vehicle edge computing network according to claim 2 is characterized in that: The computation request generated by the vehicle in time slot t is given by It indicates that, The computation request generated for vehicle n, are the K tasks generated by vehicle n.

4. The method for determining task offloading strategy and allocating resources for a vehicle edge computing network according to claim 1, characterized in that: The task offloading method is based on the offloading ratio Indicates that if The task is calculated locally; if Then the task is partially offloaded to the base station for calculation; if The task is calculated at the base station.

5. The method for determining task offloading strategy and allocating resources for a vehicle edge computing network according to claim 1, characterized in that: The delay calculation methods include local delay calculation method and edge delay calculation method; The total delay to complete each task is the maximum delay between the local computing delay and the edge computing delay.

6. The method for determining task offloading strategy and resource allocation of a vehicle edge computing network according to claim 5 is characterized in that the edge The calculation formula for calculating the delay is: in, is the unloading ratio, The size of the kth task generated for vehicle n, is the information transmission rate between vehicle n and the base station, is the information transmission rate between vehicle n and the base station, is the estimated error between the actual computing resources and the estimated computing resources of the kth task of vehicle n, The difference between the estimated completion time and the actual completion time.

7. The method for determining task offloading strategy and allocating resources for a vehicle edge computing network according to claim 1, characterized in that: The calculation formula of the channel gain between the vehicle and the base station is: in, represents small-scale fading, Represents the sum of large-scale path fading and shadowing.

8. The method for determining task offloading strategy and allocating resources for a vehicle edge computing network according to claim 1 is characterized in that: The unloading decision problem is: in, The total delay for vehicle n to process the kth task, The resource cost of processing the kth task for vehicle n, is the estimated error between the estimated computing resources and the actual computing resources of the k tasks of vehicle n, is the estimated computational resources for k tasks of vehicle n, is the maximum computing resource of the VEC server, represents the maximum processing time limit of the kth task generated by vehicle n, is the signal transmission power between vehicle n and the base station, is the maximum transmission power of vehicle n, This is the task offloading mode.

9. The method for determining task offloading strategy and allocating resources for a vehicle edge computing network according to claim 1, characterized in that: Determine the discount factor The calculation formula is: Among them, t0 is the previous moment, the discount factor γ∈(0,1), and rn is the reward function.

Citation Information

Patent Citations

  • Multi-target joint optimization task unloading strategy based on deep reinforcement learning in Internet of Vehicles

    CN116321298A

  • Vehicle-mounted task unloading and resource allocation method and device, equipment and storage medium

    CN118972804A

  • Digital twin-based edge-end collaborative scheduling method for heterogeneous tasks and resources

    US20250086005A1