Vehicle-to-Everything (V2X) Joint Task Offloading Method Based on V2V and V2I Communication

By constructing a joint task offloading method for vehicle-to-everything (V2V) and V2I communication, and utilizing a dual-delay deep deterministic strategy gradient algorithm to optimize resource allocation, the problem of unbalanced computing resource scheduling in V2V is solved, achieving efficient task offloading and energy consumption optimization, and improving user experience.

CN117768860BActive Publication Date: 2025-10-28CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311700446.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-12
Publication Date
2025-10-28
Estimated Expiration
2043-12-12

AI Technical Summary

Technical Problem

Existing vehicle-to-everything (V2X) task offloading methods mainly consider V2I or V2V standalone modes, which fail to effectively schedule and coordinate idle vehicle computing resources, resulting in low computing resource utilization, high service latency, and high energy consumption.

Method used

A joint task offloading method for vehicle-to-everything (V2V) communication based on V2I communication is constructed. By building a task offloading model, a user experience quality model and a reinforcement learning model, a dual-delay deep deterministic policy gradient algorithm is adopted to optimize resource allocation, thereby achieving dynamic scheduling of vehicle roles and efficient utilization of computing resources.

Benefits of technology

It improves resource utilization and task service rate in the vehicle-to-everything (V2X) environment, reduces system service latency and user-side energy consumption, and enhances user experience quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117768860B_ABST
    Figure CN117768860B_ABST
Patent Text Reader

Abstract

This invention relates to a joint task offloading method for vehicle-to-everything (V2V) networks based on V2V and V2I communication, belonging to the field of wireless communication. The method includes: S1: constructing a task offloading model for the system, including a local computing model and an edge offloading model; S2: calculating the latency and energy consumption of task execution based on the system's task offloading model, and constructing a user experience quality model by comprehensively considering the total latency, total energy consumption, and system service rate; S3: constructing a reinforcement learning model including a state space, action space, and reward function based on the system's task offloading model and user experience quality model; S4: solving the reinforcement learning model using a dual-latency deep deterministic policy gradient algorithm to obtain the optimal solution. This invention improves user experience quality by fully utilizing the idle computing resources of vehicles in the scenario, reducing energy consumption on the user side while ensuring system service rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of wireless communication and relates to a joint task offloading method for vehicle-to-everything (V2V) communication based on V2V (vehicle-to-vehicle) and V2I (vehicle-to-infrastructure) communication. Background Technology

[0002] With the rapid increase in the number of vehicles, users' demands for vehicle intelligence are also rising. Computationally complex and time-sensitive tasks, such as high-precision maps, autonomous driving, and VR / AR, place higher demands on vehicle computing capabilities. However, the limited computing resources of vehicles restrict the further development of vehicle applications. Vehicle Edge Computing (VEC) effectively alleviates the shortage of vehicle computing resources and reduces the execution latency and energy consumption of vehicle computing tasks by deploying edge servers carrying a large number of computing units at the roadside. However, when the number of task requests is large, it still leads to high service latency. With the advancement of data processing capabilities of in-vehicle chips and the increase in vehicle integration, vehicles themselves carry abundant computing resources, and vehicles with some idle computing resources can also be regarded as edge service nodes. Therefore, fully exploring and utilizing the idle resources of vehicles in the Internet of Vehicles is of great significance for reducing service latency and improving user experience quality.

[0003] However, existing research mainly focuses on the offloading modes of V2I or V2V separately. There are fewer studies that consider the task offloading scenario based on cloud-edge-device collaboration in the Internet of Vehicles. However, most of these studies treat vehicles as auxiliary nodes or fixed vehicles as task vehicles or service vehicles, without considering that vehicles may provide or request computing resources at different times, and cannot effectively schedule and coordinate the idle computing resources of vehicles. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide a joint task offloading method for vehicle-to-everything (V2V) communication based on V2V and V2I communication, which can improve the utilization rate of resources and system service rate in the V2V environment, and reduce application service latency and user-side energy consumption.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] A method for offloading joint tasks in vehicle-to-everything (V2V) communication based on V2V and V2I communication, specifically including the following steps:

[0007] S1: Construct the system's task offloading model, including the local computing model and the edge offloading model;

[0008] S2: Calculate the latency and energy consumption of task execution based on the system's task offloading model, and construct a user experience quality model by comprehensively considering the total latency, total energy consumption of the task, and system service rate.

[0009] S3: Based on the system's task offloading model and user experience quality model, construct a reinforcement learning model including state space, action space, and reward function;

[0010] S4: Use the dual-delay deep deterministic strategy gradient algorithm to solve the reinforcement learning model and obtain the optimal solution of the model.

[0011] Further, in step S1, the task offloading model of the system is constructed, specifically including: considering a typical vehicle-to-everything (V2X) scenario with two bidirectional intersections, this scenario includes an edge server carrying computing units and m mobile vehicles, denoted as N = {1, 2, ..., m, m+1}, where element m+1 represents the edge server; to fully explore and utilize the idle computing resources of the vehicles, the vehicles on the road are divided into a set of task vehicles (denoted as N1) that generate computing tasks and a set of service vehicles (denoted as N2) that provide computing services; at different times, the role of a vehicle as a service vehicle or a task vehicle is not fixed; the computing tasks arriving at the task vehicle can be executed locally, offloaded to the edge server for execution, or offloaded to a service vehicle for execution. Based on the V2X scenario, the task offloading model of the system includes a local computing model and an edge offloading model;

[0012] The local computing model refers to executing the vehicle arrival task locally, with task processing latency... and energy consumption They can be represented as follows:

[0013]

[0014]

[0015] in, This represents the computation delay of the task generated by vehicle i at time t when it is executed locally. This represents the CPU computing frequency allocated by the system at time t; the computing resources allocated to vehicle i for its local computing task cannot exceed its computing capacity f. i max ; Indicates the size of the computation task; Indicates the computational complexity of the task; κ represents the energy consumption of the task generated by vehicle i at time t during local execution, and is a constant related to the device chip architecture.

[0016] The edge offloading model refers to offloading tasks to edge service nodes with available computing resources, including service vehicles and edge servers, for execution, thus reducing task processing latency. and energy consumption They can be represented as follows:

[0017]

[0018]

[0019] in, This represents the computational delay at time t for the task generated by vehicle i to be unloaded and executed on service node j. This represents the transmission delay at time t from when the task generated by vehicle i is unloaded to when it is executed by service node j. This indicates the energy consumption for task data transmission. Indicates the energy consumption for task execution; This represents the CPU computing frequency allocated to the task by service node j at time t; the total computing resources allocated to service node j cannot exceed the computing power of service node j. Let represent the uplink data transmission rate from task vehicle i to service node j at time t, and its calculation formula is:

[0020]

[0021] Where B represents the system bandwidth; This represents the percentage of bandwidth allocated to task vehicle i at time t; The signal-to-noise ratio (SNR) of the link between task vehicle i and service node j is expressed as... and Let Ni represent the transmit power and channel gain of vehicle i at time t, respectively, and N0 represent additive white Gaussian noise.

[0022] When the task of vehicle i is offloaded to service vehicle j for execution, the communication distance d between vehicle i and vehicle j is... i,j The maximum V2V communication distance L cannot be exceeded. The results of the task execution can be returned to vehicle i via V2V or V2I relay communication. Since the task result data volume is very small, the transmission delay of the task result returning from the computing node can be ignored.

[0023] Furthermore, in step S2, the constructed user experience quality model is as follows:

[0024]

[0025]

[0026]

[0027]

[0028]

[0029]

[0030]

[0031] Where Ψ represents the system utility function, λ1 and λ2 represent the user's preference for latency and energy consumption, and λ1 + λ2 = 1; Υ t The system service rate is the ratio of the number of tasks successfully served by the system at time t to the total number of computing tasks arriving, and ξ represents the penalty coefficient. and This indicates the time and energy consumed by the task during local processing; f represents the task unloading decision for vehicle i at time t; i,j This indicates that service node j represents the CPU computing frequency allocated to the task; f i,i This indicates the CPU computing frequency allocated to the system.

[0032] Constraint C1 ensures that the total computing resources allocated to service nodes cannot exceed the computing capacity of the service nodes; Constraint C2 ensures that the computing resources allocated to a task when it is computed locally cannot exceed the computing capacity of the vehicle; Constraint C3 ensures that the total bandwidth resources allocated to the system cannot exceed its system bandwidth capacity; Constraints C4 and C5 ensure that a task can only be executed on one of the local or edge service nodes; Constraint C6 ensures that the V2V communication distance cannot exceed its maximum communication distance.

[0033] Furthermore, in step S3, the reinforcement learning model constructed includes:

[0034] (1) State space

[0035] Define the local observation value of the system on vehicle i at time t as in, This indicates the type of vehicle i, where 1 represents a service vehicle and 0 represents a mission vehicle. This indicates the location information of vehicle i; This indicates the task information for vehicle i; This represents the available computing resources for vehicle i; the system's state space at time t can be represented as... in, This indicates the available computing resources for edge nodes;

[0036] (2) Action space

[0037] At the beginning of each time slot, the edge server makes a corresponding decision based on the system state information; the system's decision for vehicle i at time t is defined as follows: Where N SN = N - N1 + {0}, where element 0 indicates that the task is executed locally; the action space of the system at time t can be represented as

[0038] (3) Reward function

[0039] The system's goal is to improve the user experience quality by maximizing the utility function Ψ; therefore, the reward function is defined as:

[0040]

[0041] In the reinforcement learning model, the edge server is responsible for collecting vehicle information within its coverage area at the beginning of each time slot to obtain the global state observations s of the environment. t Then, based on the current state, a decision is made to unload the task. t And execute, at which point the new state s of the environment is obtained. t+1 With reward r t Then, the empirical data tuples (s) t ,a t ,s t+1 ,r t Place it in the experience replay area.

[0042] Furthermore, in step S4, the reinforcement learning model is solved using a dual-delay deep deterministic policy gradient algorithm. Specifically, the problem to be optimized is transformed into a Markov decision process and an actor network and a critic network are constructed. The agent updates the network parameters by continuously interacting with the environment to obtain rewards, thereby approximating the optimal solution of the objective function. The algorithm mainly includes a state space s, an action space a, a reward function r, an actor network, a critic network, and an experience replay region.

[0043] The beneficial effects of this invention are as follows: The method of this invention comprehensively considers the roadside MEC server and the service vehicles on the road that can provide computing services. By constructing a task offloading model and a user experience quality model, and optimizing the model based on the dual-latency deep deterministic policy gradient algorithm, it provides a task offloading scheme for the computationally intensive and time-sensitive tasks of vehicles, thereby improving the utilization rate of resources and the task service rate in the vehicle network environment, and reducing the system's service latency and user-side energy consumption.

[0044] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description

[0045] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:

[0046] Figure 1 A flowchart of the vehicle-to-everything (V2V) joint task offloading method based on V2V and V2I communication provided by the present invention;

[0047] Figure 2 This is a structural diagram of the vehicle networking system based on V2V and V2I communication in this invention;

[0048] Figure 3 This is a flowchart illustrating the gradient algorithm of the dual-delay deep deterministic strategy in this invention. Detailed Implementation

[0049] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0050] Please see Figures 1-3 , Figure 1 A flowchart of a method for offloading joint computing tasks based on V2V and V2I in a vehicle-to-everything (V2X) network, provided as an embodiment of the present invention, is included in this method, comprising the following steps:

[0051] Step 1: Build a task offloading model for the system that allows computational tasks arriving at a vehicle to be computed locally, offloaded to a service vehicle for computation, or offloaded to an edge server for computation.

[0052] In step 1, consider a vehicle-to-everything (V2X) scenario involving two bidirectional traffic intersections, such as... Figure 2As shown in the diagram, this scenario includes an edge server carrying computing units and m mobile vehicles, denoted as N = {1, 2, ..., m, m+1}, where element m+1 represents the edge server. To fully utilize the idle computing resources of the vehicles, the vehicles on the road are divided into a task vehicle set (denoted as N1) that generates computing tasks and a service vehicle set (denoted as N2) that provides computing services. The role of a vehicle as a service vehicle or a task vehicle is not constant at different times. In this system model, time is discretized into equal-length time slots, denoted as Δt. Within different time slots, the channel conditions and workload of the vehicles are different. The computing tasks arriving at time t are... The computation can be performed locally, offloaded to a service vehicle, or computed on an edge server. Indicates the size of the calculation task. T represents the computational complexity of the task. max Indicates the deadline for the task.

[0053] The computational model includes a local computational model and an edge offloading model.

[0054] The local computing model refers to performing the vehicle arrival calculation task locally, expressed as:

[0055]

[0056]

[0057] in, This represents the computation delay of the task generated by vehicle i at time t when it is executed locally. Indicates the CPU computing frequency allocated by the system; Indicates the size of the computation task; This indicates the computational complexity of the task. Let κ represent the energy consumption of the task generated by vehicle i at time t during local execution, where κ is a constant related to the device chip architecture.

[0058] The edge offloading model refers to offloading tasks to edge service nodes (including service vehicles and edge servers) with available computing resources for execution, represented as:

[0059]

[0060]

[0061] in, This represents the computational delay at time t for the task generated by vehicle i to be unloaded and executed on service node j. This represents the transmission delay at time t from when the task generated by vehicle i is unloaded to when it is executed on service node j. This indicates the energy consumption for task data transmission. This indicates the energy consumption for task execution. Let represent the uplink data transmission rate from task vehicle i to service node j at time t. The formula for calculating the uplink data transmission rate is:

[0062]

[0063] Where B represents the system bandwidth; This represents the percentage of bandwidth allocated to task vehicle i at time t; The signal-to-noise ratio (SNR) of the communication link between task vehicle i and service node j is expressed as... and Let Ni represent the transmit power and channel gain of vehicle i at time t, respectively, and N0 represent additive white Gaussian noise.

[0064] Step 2: Calculate the latency and energy consumption of task execution based on the system's task offloading model, and construct a user experience quality model by comprehensively considering the total latency and energy consumption of the task as well as the system service rate.

[0065] In step 2, the user experience quality model is as follows:

[0066]

[0067]

[0068]

[0069]

[0070]

[0071]

[0072]

[0073] Where ψ represents the system utility function, λ1 and λ2 represent the user's preference for latency and energy consumption, and λ1 + λ2 = 1; Υ t Let ξ be the system service rate, representing the ratio of the number of tasks successfully served by the system at time t to the total number of computational tasks arriving, and let ξ be the penalty coefficient.

[0074] Constraint C1 ensures that the total computing resources allocated to service nodes cannot exceed the computing capacity of the service nodes; constraint C2 ensures that the computing resources allocated to tasks when they are computed locally cannot exceed the computing capacity of the vehicle; constraint C3 ensures that the total bandwidth resources allocated to the system cannot exceed its system bandwidth capacity; constraints C4 and C5 ensure that tasks can only be executed on one of the local or edge service nodes; constraint C6 ensures that the V2V communication distance cannot exceed its maximum communication distance.

[0075] Step 3: Based on the system's task unloading model and user experience quality model, construct a reinforcement learning model, including state space s, action space a, and reward function r.

[0076] (1) State space

[0077] Define the local observation value of the system on vehicle i at time t as in, This indicates the type of vehicle i, where 1 represents a service vehicle and 0 represents a mission vehicle. This indicates the location information of vehicle i; This indicates the task information for vehicle i; Let represent the available computing resources for vehicle i. The state space of the system at time t can be represented as: in, This indicates the available computing resources for edge nodes.

[0078] (2) Action space

[0079] At the beginning of each time slot, the edge server makes a corresponding decision based on the system state information. Let the system's decision regarding vehicle i at time t be defined as follows: Where N SN = N - N1 + {0}, where element 0 indicates that the task is executed locally. The action space of the system at time t can be represented as:

[0080] (3) Reward function

[0081] The system's goal is to improve the user's experience quality by maximizing the utility function ψ. Therefore, the reward function is defined as:

[0082]

[0083] In the reinforcement learning model, the edge server is responsible for collecting vehicle information within its coverage area at the beginning of each time slot to obtain the global state observations s of the environment. t Then, based on the current state, a decision is made to unload the task. t And execute, at which point the new state s of the environment is obtained. t+1 With reward r tThen, the empirical data tuples (s) t ,a t ,s t+1 ,r t Place it in the experience replay area.

[0084] Step 4: Use the dual-delay deep deterministic strategy gradient algorithm to solve the reinforcement learning model to obtain the optimal solution of the model.

[0085] In step 4, the dual-delay deep deterministic policy gradient algorithm is used to solve the optimization problem in the reinforcement learning model, combined with... Figure 3 This can be further expressed as:

[0086] Step 4-1: Initialize the actor network π(s; θ) with parameters θ, and the critic networks Q1(s, a; ω1) and Q2(s, a; ω2) with parameters ω1 and ω2.

[0087] Step 4-2: Initialize the actor target network π′(s;θ′) and the critic target network Q′1(s,a;ω′1) and Q′2(s,a;ω′2) with parameters θ′←θ, ω′1←ω1, and ω′2←ω2.

[0088] Step 4-3: Initialize the experience replay area Initialize the reward discount factor γ and reset the system environment state;

[0089] Step 4-4: The actor network determines the current environment state based on s. t Make a decision t ;

[0090] Steps 4-5: Perform action a t And receive a reward r t The environmental state becomes s t+1 ;

[0091] Steps 4-6: Transform the empirical data tuples (s) t ,a t ,s t+1 ,r t Add to experience replay area

[0092] Steps 4-7: From the experience replay area Medium-sampled small batch data N;

[0093] Steps 4-8: Obtain from the critic network Based on the actor target network Based on the critic target network and Target Q value

[0094] Steps 4-8: Update the critic network parameters using the mean squared error loss function

[0095] Steps 4-9: Update the actor network parameters θ using the deterministic policy gradient function, where

[0096] Step 4-10: Update all target networks using soft updates, represented as: θ′←τθ+(1-τ)θ′, ω′ i ←τω i +(1-τ)ω′ i , where τ is the soft update coefficient.

[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for joint task offloading in vehicle-to-everything (V2V) communication based on V2V and V2I communication, characterized in that, The method specifically includes the following steps: S1: Construct the system's task offloading model, including the local computing model and the edge offloading model; S2: Calculate the latency and energy consumption of task execution based on the system's task offloading model, and construct a user experience quality model by comprehensively considering the total latency, total energy consumption of the task, and system service rate. S3: Based on the system's task offloading model and user experience quality model, construct a reinforcement learning model including state space, action space, and reward function; S4: Use the dual-delay deep deterministic strategy gradient algorithm to solve the reinforcement learning model to obtain the optimal solution of the model; In step S1, the task offloading model of the system is constructed, specifically including: considering a typical vehicle-to-everything (V2X) scenario with two bidirectional intersections, this scenario includes an edge server carrying computing units and One moving vehicle, denoted as , of which elements This refers to edge servers; to fully utilize the idle computing resources of vehicles, vehicles on the road are divided into task vehicles that generate computing tasks. and a collection of service vehicles that provide computing services At different times, the role of a vehicle as a service vehicle or a task vehicle is not fixed; the computing tasks arriving at the task vehicle can be executed locally, offloaded to an edge server, or offloaded to a service vehicle; based on the aforementioned vehicle-to-everything (V2X) scenario, the system's task offloading model includes a local computing model and an edge offloading model. The local computing model refers to executing the vehicle arrival task locally, with task processing latency... and energy consumption They are represented as follows: in, express t Vehicles on duty i The computational latency of the generated tasks when executed locally; express t The CPU computing frequency allocated by the time system; vehicle i The computing resources allocated to a local computing task cannot exceed its computing capacity. ; Indicates the size of the computation task; Indicates the computational complexity of the task; express t Vehicles on duty i The energy consumption of executing tasks locally. It is a constant related to the device chip architecture; The edge offloading model refers to offloading tasks to edge service nodes with available computing resources, including service vehicles and edge servers, for execution, thus reducing task processing latency. and energy consumption They are represented as follows: in, express t Vehicles on duty i The generated tasks are unloaded to the service node. j Execution computation latency, express t Vehicles on duty i The generated tasks are unloaded to the service node. j Execution transmission delay; Indicates the energy consumption for task data transmission. Indicates the energy consumption for task execution; express t Real-time service nodes j CPU computing frequency allocated to the task; service node j The total allocated computing resources cannot exceed the service node's capacity. j computing power ; express t From the mission vehicle at all times i To service node j The uplink data transmission rate is calculated using the following formula: in, Indicates the system bandwidth; express t Time allocated to mission vehicles i The percentage of bandwidth; Indicates the mission vehicle i With service nodes j The signal-to-noise ratio of the link between them is expressed as , They represent t Vehicles on duty i Transmit power and channel gain, This represents additive white Gaussian noise; When the vehicle i Task unloaded to service vehicle j During execution, the vehicle i With vehicles j Communication distance between The maximum V2V communication distance L cannot be exceeded; the result of the task execution is returned to the vehicle via V2V or V2I relay communication. i The transmission delay of task results returning from the computing node is negligible. The constructed user experience quality model is as follows: in, Represents the system utility function. , This indicates the user's preference for latency and energy consumption. ; For the system's service rate, t This represents the ratio of the number of tasks successfully served by the system at any given time to the total number of computational tasks that arrive. Indicates the penalty coefficient; and These represent the time and energy consumed by the task during local processing, respectively. express t Time vehicle i Task unloading decision; Indicates service node j Calculate the CPU frequency allocated to the task; This indicates the CPU computing frequency allocated to the system; constraint C1 ensures that the total computing resources allocated to service nodes cannot exceed the computing capacity of the service nodes; constraint C2 ensures that the computing resources allocated to tasks when computing locally cannot exceed the computing capacity of the vehicle; constraint C3 ensures that the total bandwidth resources allocated to the system cannot exceed its system bandwidth capacity; constraints C4 and C5 ensure that tasks can only be executed on one of the local or edge service nodes; constraint C6 ensures that the V2V communication distance cannot exceed its maximum communication distance.

2. The vehicle-to-everything (V2X) joint task offloading method according to claim 1, characterized in that, In step S3, the reinforcement learning model constructed includes: (1) State space definition t Timetable system for vehicles i The local observation value is , in, Indicates vehicle i The type is 1 for service vehicles and 0 for mission vehicles; Indicates vehicle i Location information; Indicates vehicle i Task information; Indicates vehicle i Available computing resources; the system in t The state space representation at time t is as follows ,in, Indicates the available computing resources for edge nodes; (2) Action space At the beginning of each time slot, the edge server makes a corresponding decision based on system status information; define t Timetable system for vehicles i The decision is ,in Element 0 indicates that the task is executed locally; t The action space representation of the time-space system is as follows ; (3) Reward function The system's objective is to maximize the utility function. To improve the quality of the user experience, the reward function is defined as follows: In reinforcement learning models, edge servers are responsible for collecting vehicle information within their coverage area at the beginning of each time slot to obtain global state observations of the environment. Then, based on the current state, a task unloading decision is made. And execute, at which point the new state of the environment is obtained. With rewards Then, the empirical data tuples Place it in the experience replay area.

3. The vehicle-to-everything (V2X) joint task offloading method according to claim 1, characterized in that, In step S4, the reinforcement learning model is solved using a dual-delay deep deterministic policy gradient algorithm. Specifically, the problem to be optimized is transformed into a Markov decision process and an actor network and a critic network are constructed. The agent updates the network parameters by continuously interacting with the environment to obtain rewards, thereby approximating the optimal solution of the objective function.

Citation Information

Patent Citations

  • Task unloading method and device, electronic equipment and storage medium

    CN112988285A

  • Computing task unloading method based on deep reinforcement learning in electric power internet of things

    CN114065963A