Task scheduling method in Internet of Vehicles scene based on deep reinforcement learning

By adopting a task scheduling method based on deep reinforcement learning in the Internet of Vehicles scenario, it is transformed into Markov decision-making problems and using the Dueling DQN network to make scheduling decisions, the task delay problem caused by limited computing resources of edge servers is solved, and more efficient task scheduling and energy management are achieved.

CN120066740AInactive Publication Date: 2025-05-30SHANDONG ZHENGYUN INFORMATION TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510534464.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the Internet of Vehicles scenario, due to limited computing resources, edge servers are difficult to allocate sufficient computing resources to a large number of computing tasks at the same time, resulting in increased task waiting and execution delays, affecting the real-time and overall performance of the task.

Method used

The task scheduling method based on deep reinforcement learning is adopted to transform the task scheduling problems into Markov decision-making problems, and a scheduling decision model is built through the Dueling DQN network, action evaluation and parameter optimization are carried out to realize task scheduling.

Benefits of technology

Through this method, the task scheduling time is shortened, energy consumption is reduced, and the efficiency and accuracy of task scheduling is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066740A_ABST
    Figure CN120066740A_ABST
Patent Text Reader

Abstract

The invention discloses a task scheduling method in an Internet of Vehicles scene based on deep reinforcement learning, belongs to the technical field of deep learning, and aims to solve the technical problems of how to realize task scheduling in the Internet of Vehicles scene and reduce task scheduling completion time and energy consumption. Comprising the following steps: for a task initiated by a vehicle, converting a task scheduling problem of the task between the vehicle and an edge server into a Markov decision problem; a scheduling decision model is constructed based on the DQN, the scheduling decision model comprises a main network and a target network which are the same in structure, in the action evaluation stage, the main network is used for predicting and outputting the Q value of each action in the current state, and in the parameter optimization stage, the main network is used for calculating the Q value according to the current state and the action corresponding to the current state. The target network is used for predicting the Q value of each action in the next state and outputting the maximum Q value; and for the Markov decision problem, action evaluation and parameter optimization are carried out based on the scheduling decision model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep reinforcement learning, and more specifically, to a task scheduling method in the Internet of Vehicles (IoV) scenario based on deep reinforcement learning. Background Art

[0002] The widespread application of the Internet of Things (IoT), especially in fields such as smart home and intelligent transportation, has led to an explosive growth in the amount of data generated. To effectively address the challenges of data processing in the intelligent transportation field and promote the development of smart cities, the Internet of Vehicles (IoV) has become one of the important research scenarios in the IoT field. However, due to the limited battery life, storage capacity, and computing power of vehicles, it is challenging or even infeasible to implement computationally intensive applications only within the vehicle. Although cloud computing excels in resource-intensive data processing and artificial intelligence training, due to the latency issues introduced by data transmission, data needs to be continuously transmitted to the data center and then returned, which limits its application in the intelligent transportation field. Edge cloud is considered one of the potential solutions to this problem.

[0003] Edge cloud extends the convenience of cloud computing to the edge network hosted by micro data centers. Compared with connecting to the central data center, it can achieve faster storage, analysis, and data processing speeds. Through edge cloud, nearby vehicles can execute applications on edge servers through computing offloading, thus greatly reducing the round-trip transmission latency and alleviating the burden on the backhaul network. However, with the development of 5G and autonomous driving technologies, more and more tasks in the IoV are latency-sensitive, which poses a huge pressure on edge servers.

[0004] To address this challenge, edge servers need to cooperate with vehicles to jointly complete computing tasks. In fact, some vehicles may be relatively idle during certain time periods, resulting in underutilized additional computing and storage resources and potentially causing resource waste. Therefore, cooperating with vehicles is a more effective and practical solution. In addition, in the IoV, the continuous durability of vehicles and the predictability of their trajectories ensure the feasibility of this architecture.

[0005] In a wireless fading environment, the time - interval channel state has a significant impact on the optimal offloading decision of vehicle systems. Currently, research on computational offloading mainly focuses on offloading decisions and task scheduling, aiming to reduce task execution latency and the energy consumption of the task offloading system. However, existing research has not fully addressed the multi - task scheduling problem within edge servers. Due to the limited computational resources of edge servers, when a large number of computational tasks are offloaded for execution, edge servers often cannot allocate sufficient computational resources to all tasks simultaneously. This may lead to increased task waiting and execution latency. Such resource limitations and task competition may affect the real - time performance and overall performance of tasks, especially for those with strict real - time requirements. Therefore, it is crucial to establish a collaborative framework between the edge and the vehicle and effectively address the multi - task scheduling challenges on edge servers.

[0006] How to achieve task scheduling in the vehicle - to - everything (V2X) scenario and reduce the completion time and energy consumption of task scheduling is a technical problem that needs to be solved. Summary of the Invention

[0007] The technical task of the present invention is to address the above - mentioned deficiencies and provide a task scheduling method in the V2X scenario based on deep reinforcement learning to solve the technical problem of how to achieve task scheduling in the V2X scenario and reduce the completion time and energy consumption of task scheduling.

[0008] A task scheduling method in the V2X scenario based on deep reinforcement learning according to the present invention is applied among roadside units, vehicles, and edge servers. Each roadside unit is configured with an edge server. The method includes the following steps: The vehicle, as a device, sends its device status information to the edge server connected to it. The edge server, as a device, sends its device status information and the device status information of the vehicles connected to it to other edge servers. For the tasks initiated by the vehicle, the task scheduling problem between the vehicle and the edge server is transformed into a Markov decision problem. Among them, the device status information of the vehicle and the edge server is used as the state, and the device executing the task is used as the action. Based on the Dueling DQN network, a scheduling decision model is constructed. The scheduling decision model includes a main network and a target network with the same structure. In the action evaluation stage, the main network is used to take the current state as input and predict and output the Q - value of each action in the current state. In the parameter optimization stage, the main network is used to calculate the Q - value according to the current state and the action corresponding to the current state. The target network is used to predict the Q - value of each action in the next state according to the next state and output the maximum Q - value. For the Markov decision problem, action evaluation and parameter optimization are carried out based on the scheduling decision model to achieve task scheduling. Among them, the action evaluation includes: taking the current state as the input, predicting the Q value of each action in the current state through the main network, selecting and executing an action based on the Q value of each action by the intelligent agent, the environment returning the action reward and the next state according to the executed action, and using the current state, the action corresponding to the current state, the action reward, and the next state as the sample data of the quadruple, and storing the sample data in the experience replay pool; The parameter optimization includes: taking the sample data selected from the experience replay pool as the input, predicting the Q value of each action in the next state through the target network and selecting one Q value as the output, calculating based on the current state and the action corresponding to the current state, predicting the Q value of the action through the main network, constructing a loss function based on the Q value output by the current state and the maximum Q value output by the target network, optimizing the network parameters of the main network by minimizing the loss function, and during the parameter optimization process, when the number of iterations reaches the predetermined step size, periodically copying the network parameters of the main network to the target network until the change value of the network parameters of the main network is less than the threshold.

[0009] Preferably, during action evaluation, based on the Q value of each action, the intelligent agent selects and executes an action based on the algorithm.

[0010] Preferably, the main network and the target network are used to perform the following operations: Taking the state as the input, performing feature extraction through the convolutional layer to obtain the feature representation of the input state; Converting the feature representation of the multi-dimensional tensor into a one-dimensional vector through the Flat operation; Performing a fully connected layer process on the state value function and the advantage function, and combining the state value function based on the state and the advantage function based on the action to generate the final Q value. The Q value function is expressed as: , where represents the state, represents the action, represents the network parameters of the common part in the main network and the target network, and represent the network parameters of the two fully connected layers.

[0011] Preferably, during action evaluation, for the Q value function, centered on the advantage function, the Q value function is equivalent to: , , Converting the Q value into a probability distribution through the softmax function. The probability distribution is expressed as: , Based on the probability distribution, the intelligent agent takes the probability Execute the action with the highest current probability. The Q-value calculation formula is as follows: .

[0012] Preferably, for the Markov decision process, the reward function is defined as: , where, represents the weight coefficient, represents the processing time of task and represents the energy consumption required to execute task After the agent selects and executes an action, the environment returns an action reward for the action based on the reward function.

[0013] Preferably, the parameter optimization includes the following steps: L100. Set the step size N; L200. The state at time , the action corresponding to the state at time and the action reward , and the state at time are stored in the experience replay pool as sample data. Then, read the sample data from the experience replay pool, use the state at time as the input, predict the Q-value of each action through the target network, and select the maximum Q-value as the output of the target network; L300. Take as the value function of the main network, take as the value function of the target network, and construct a loss function based on and . The loss function is expressed as: , , where, represents the value function of the main network, represents the state at time represents the state selected action, represents the network parameters of the main network, represents the discount factor, represents the network parameters of the target network; L400. Update the network parameters of the main network based on the loss function through the gradient descent method. The formula is as follows: , where, represents the learning rate; L500. Execute steps L200 - L400 for multiple iterations. When the number of iterations reaches the step size N, copy the network parameters of the main network to the network parameters of the target network. Based on the optimized target network, execute steps L200 - L500 for multiple iterations until the change in the network parameters of the main network is less than the threshold.

[0014] Preferably, during parameter optimization, sample data is read from the experience replay pool as input based on the empirically assigned priority. The empirically assigned priority has the following calculation formula: , where, represents a small positive number, represents a positive hyperparameter used to adjust the distribution of priorities; Convert the empirically assigned priority to a probability distribution through proportional sampling. The probability distribution is expressed as: , where, represents the total number of experiences, is a hyperparameter that controls the shape of the distribution.

[0015] A task scheduling method based on deep reinforcement learning in a vehicle - to - everything (V2X) scenario of the present invention has the following advantages: 1. Convert the task scheduling problem in the V2X scenario into a Markov decision process problem. For the Markov decision problem, perform action evaluation and parameter optimization based on the scheduling decision model to achieve task scheduling, thereby shortening the task scheduling time and reducing energy consumption; 2. During the parameter optimization process, assign priorities to each sample data in the experience replay pool, randomly sample a batch of sample data as experiences. Then, use the main network to calculate the Q - value estimate of the current state, use the target network to calculate the target Q - value of the next state, and then calculate the error. Use the error to update the parameters of the main network, thereby improving the efficiency of network parameter optimization and the accuracy of prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0017] The present invention will be further described below with reference to the drawings.

[0018] Figure 1 It is a flowchart of a task scheduling method in a vehicle networking scenario based on deep reinforcement learning for an embodiment. Specific embodiments

[0019] The present invention will be further described below with reference to the drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it. However, the embodiments cited are not intended to limit the present invention. Without conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.

[0020] The embodiments of the present invention provide a task scheduling method in a vehicle networking scenario based on deep reinforcement learning, which is used to solve the technical problem of how to implement task scheduling in a vehicle networking scenario and reduce the completion time and energy consumption of task scheduling.

[0021] Embodiment:

[0022] A task scheduling method in a vehicle networking scenario based on deep reinforcement learning of the present invention is applied among roadside units, vehicles, and edge servers. Each roadside unit is configured with an edge server. As Figure 1 shown, the method includes the following steps: S100. The vehicle, as a device, sends its device status information to the edge server connected to it. The edge server, as a device, sends its device status information and the device status information of the vehicles connected to it to other edge servers; S200. For the tasks initiated by the vehicle, convert the task scheduling problem between the vehicle and the edge server into a Markov decision problem. Among them, the device status information of the vehicle and the edge server is used as the state, and the device executing the task is used as the action; S300. Build a scheduling decision model based on the Dueling DQN network. The scheduling decision model includes a main network and a target network with the same structure. In the action evaluation stage, the main network is used to take the current state as the input and predict and output the Q value of each action in the current state. In the parameter optimization stage, the main network is used to calculate the Q value according to the current state and the action corresponding to the current state. The target network is used to predict the Q value of each action in the next state according to the next state and output the maximum Q value; S400. For the Markov decision problem, action evaluation and parameter optimization are performed based on the scheduling decision model to achieve task scheduling.

[0023] Among them, the action evaluation in step S400 includes: taking the current state as the input, predicting the Q-value of each action of the current state through the main network, selecting an action based on the Q-value of each action by the agent and executing it, the environment returns the action reward and the next state according to the executed action, taking the current state, the action corresponding to the current state, the action reward, and the next state as the sample data of the quadruple, and storing the sample data in the experience replay pool.

[0024] The parameter optimization in step S400 includes: taking the sample data selected from the experience replay pool as the input, predicting the Q-value of each action of the next state through the target network and selecting a Q-value as the output, calculating the Q-value of the action predicted through the main network based on the current state and the action corresponding to the current state, constructing a loss function based on the Q-value output by the current state and the maximum Q-value output by the target network, optimizing the network parameters of the main network by minimizing the loss function. During the parameter optimization process, when the number of iterations reaches the predetermined step size, the network parameters of the main network are periodically copied to the target network until the change value of the network parameters of the main network is less than the threshold.

[0025] In this embodiment, the scenario where vehicles generate continuous computing tasks in the edge cloud architecture of the urban vehicle-to-everything (V2X) network is considered. A large number of RSUs (roadside units) are deployed at different locations in the city, including near traffic lights, street lights, public billboards, and building tops, to achieve extensive urban coverage. To provide faster data processing and computing services for vehicles, an edge server is deployed on each RSU, enabling the RSUs to communicate with each other through a local area network connection.

[0026] Vehicles can be either task creators, generating multiple tasks, or task executors, completing various tasks. The RSUs act as schedulers and executors of V2X tasks, regularly collecting status information from connected vehicle terminals and sharing this information with other RSUs.

[0027] As the task initiator, the vehicle assigns all tasks except the locally executable tasks to the RSUs for scheduling. Subsequently, at the RSU level, a comprehensive evaluation of the edge servers and available resources of idle vehicles is carried out to formulate an effective task scheduling strategy. Based on this strategy, tasks are assigned to different devices for execution. Finally, after the tasks are completed, all execution devices send the results to the task creator.

[0028] In step S200 of this embodiment, the task scheduling task is modeled into a Markov decision process (MDP) model, and the MDP model represents the interaction process between the agent and its environment. The agent is a decision-making and learning agent that performs operations by observing information such as the current traffic situation. The environment includes elements related to the task scheduling of the entire road network, such as vehicles and edge servers. The agent endeavors to explore the optimal policy by taking different actions in different states in order to maximize the reward in the long run. The MDP model consists of the following elements: State space: State Task The device status information of each vehicle and each edge server at the current moment; Action space: The action space encompasses all potential operations that the agent may perform. In the model of this embodiment, the action of the agent is defined as selecting a device to execute the task, which means that the action space can be described as: , Reward function: By analyzing the environmental state and the actions taken by the agent, the reward function provides a feedback signal to measure the superiority of the decisions made by the agent in a given environment. To minimize time and energy consumption, the reward function is defined as: , where, represents the weight coefficient. represents the task 's processing time, represents the execution of the task 's required energy consumption.

[0029] In step S300 of this embodiment, a DQN network is selected to construct a scheduling decision model, and the main network and the target network in this model are used to perform the following operations: (1) Taking the state as the input and performing feature extraction through the convolutional layer to obtain the feature representation of the input state; (2) Converting the feature representation of the multi-dimensional tensor into a one-dimensional vector through the Flat operation; (3) Performing a fully connected layer processing on the state value function and the advantage function, and merging the state value function based on the state and the advantage function based on the action to generate the final Q value.

[0030] The dueling DQN architecture divides the network into two parts. The first part is the state value function , which is only associated with the state, and the second component is the advantage function , which is associated with both the state and the action. Therefore, the value function ( value) can be expressed as: , Among them, represents the state, represents the action, represents the network parameters of the common part in the main network and the target network, and represents the network parameters of two fully connected layers.

[0031] In order to improve the recognizability of the model, with the advantage function as the center, the mathematical form of the dueling network actually used is: , Then, in this embodiment, the softmax function is used to convert the value into a probability distribution: , so as to maximize the knowledge already obtained currently.

[0032] In step S400 of this embodiment, during action evaluation, based on the probability distribution, the agent selects the action with the highest current probability with probability to execute, and the Q-value calculation formula is: , The parameter optimization in step S400 of this embodiment includes the following steps: L100. Set the step size N; L200. The state at time , the action corresponding to the state at time and the action reward and, as well as the state at time are stored in the experience replay pool as sample data. After reading the sample data from the experience replay pool, using the state at time as the input, predict the Q-value of each action through the target network, and select the maximum Q-value as the output of the target network; L300. Take as the value function of the main network, take as the value function of the target network, and based on and construct a loss function, and the loss function is expressed as: , , Among them, represents the value function of the main network, Represents The state at a moment Represents the state The selected action Represents the network parameters of the main network Represents the discount factor Represents the network parameters of the target network; L400. Update the network parameters of the main network based on the loss function and by the gradient descent method. The formula is as follows: , Wherein, Represents the learning rate; L500. Execute steps L200 - L400 for multiple iterations. When the number of iterations reaches the step size N, copy the network parameters of the main network to the network parameters of the target network. Based on the optimized target network, execute steps L200 - L500 for multiple iterations until the change in the network parameters of the main network is less than the threshold.

[0033] Wherein, when optimizing the parameters, read the sample data from the experience replay pool as input based on the experience - assigned priority. The experience - assigned priority The calculation formula is as follows: , Wherein, Represents a small positive number, used to avoid the situation of zero priority and ensure that each experience has a certain chance of being selected, Represents a positive hyperparameter, used to adjust the distribution of priorities. A larger Enhances the influence of experiences with higher priorities. In this embodiment, the experience - assigned priority is converted into a probability distribution through proportional sampling. The probability distribution is expressed as: ; Wherein, Represents the total number of experiences, Is a hyperparameter that controls the shape of the distribution. This indicates that when selecting samples, the higher - priority samples have a higher probability of being selected. The hyperparameter Controls the shape of this probability distribution, making the difference between priorities more obvious.

[0034] In this embodiment, during the execution of action evaluation and parameter optimization, the main network outputs the Q - value of each action according to the current state The agent adopts The policy to select an action And execute it. The environment returns a reward value And the next state , and store the current state, the executed action, the obtained reward, and the next state in the prioritized experience replay pool for subsequent learning. During training, assign priorities to each experience in the experience replay pool and randomly sample a batch of experiences. Then, use the main network to calculate the Q-value estimate of the current state, use the target network to calculate the target Q-value of the next state, and then calculate the error. Finally, use the error to update the parameters of the main network.

[0035] The above has detailedly demonstrated and described the present invention through the drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above-mentioned multiple embodiments, those skilled in the art can know that more embodiments of the present invention can be obtained by combining the means in the above different embodiments, and these embodiments are also within the protection scope of the present invention.

Claims

1. A task scheduling method in a vehicle networking scenario based on deep reinforcement learning, characterized in that: Applied between a roadside unit, a vehicle and an edge server, each roadside unit is configured with an edge server, and the method comprises the following steps: The vehicle, as a device, sends its device status information to the edge server connected to it, and the edge server, as a device, sends its device status information and the device status information of the vehicle connected to it to other edge servers; For tasks initiated by vehicles, the task scheduling problem between vehicles and edge servers is transformed into a Markov decision problem, where the device status information of the vehicle and edge server is used as the state, and the device that executes the task is used as the action; A scheduling decision model is constructed based on the Dueling DQN network. The scheduling decision model includes a main network and a target network with the same structure. In the action evaluation stage, the main network is used to take the current state as input and predict the Q value of each action in the current state. In the parameter optimization stage, the main network is used to calculate the Q value according to the current state and the action corresponding to the current state. The target network is used to predict the Q value of each action in the next state according to the next state and output the maximum Q value. For Markov decision problems, action evaluation and parameter optimization are performed based on the scheduling decision model to achieve task scheduling; Among them, action evaluation includes: taking the current state as input, predicting the Q value of each action in the current state through the main network, based on the Q value of each action, the agent selects an action and executes it, the environment returns the action reward and the next state according to the executed action, and the current state, the action and action reward corresponding to the current state, and the next state are used as four-element sample data, and the sample data is stored in the experience replay pool; Parameter optimization includes: taking sample data selected from the experience replay pool as input, predicting the Q value of each action in the next state through the target network, and selecting a Q value as output, calculating based on the current state and the action corresponding to the current state, predicting the Q value of the action through the main network, constructing a loss function based on the Q value output of the current state and the maximum Q value output of the target network, optimizing the network parameters of the main network by minimizing the loss function, and during the parameter optimization process, when the number of iterations reaches a predetermined step size, periodically copying the network parameters of the main network to the target network until the change value of the network parameters of the main network is less than a threshold.

2. The task scheduling method in the vehicle networking scenario based on deep reinforcement learning according to claim 1 is characterized in that: When evaluating actions, the agent evaluates the Q value of each action based on The algorithm picks an action and executes it.

3. The task scheduling method in the vehicle networking scenario based on deep reinforcement learning according to claim 1 is characterized in that: The primary network and the target network are used to perform the following operations: Taking the state as input, feature extraction is performed through the convolution layer to obtain the feature representation of the input state; The feature representation of a multi-dimensional tensor is converted into a one-dimensional vector through the Flat operation; The state value function and advantage function are processed by the fully connected layer, and the state-based state value function and the action-based advantage function are merged to generate the final Q value. The Q value function is expressed as: , in, Indicates the status, Indicates action, The network parameters representing the common part of the main network and the target network, and Represents the network parameters of the two fully connected layers.

4. The task scheduling method in the vehicle networking scenario based on deep reinforcement learning according to claim 3 is characterized in that: When evaluating actions, for the Q value function, centered on the advantage function, the Q value function is equivalent to the following: , , The Q value is converted into a probability distribution through the softmax function, and the probability distribution is expressed as: , Based on the probability distribution, the agent Select the action with the highest current probability to execute. The Q value calculation formula is: 。 5. The task scheduling method in the vehicle networking scenario based on deep reinforcement learning according to any one of claims 1 to 4, characterized in that: For the Markov decision process, the reward function is defined as: , in, represents the weight coefficient, Indicates the task processing time, Indicates execution of tasks The energy consumption required; After the agent selects an action and executes it, the environment returns an action reward for the action based on the reward function.

6. The task scheduling method in the vehicle networking scenario based on deep reinforcement learning according to any one of claims 1 to 4, characterized in that: Parameter optimization includes the following steps: L100, set step length N; L200, Status at all times , Actions corresponding to the state at the moment and action rewards ,as well as Status at all times After being stored in the experience replay pool as sample data, the sample data is read from the experience replay pool. Status at all times For input, predict the Q value of each action through the target network, and select the maximum Q value As the output of the target network; L300, As the value function of the main network, As the value function of the target network, based on and Construct loss function, loss function It is expressed as: , , in, represents the value function of the main network, express Moment status, Indicates status Select the action, represents the network parameters of the main network, represents the discount factor, represents the network parameters of the target network; L400, based on the loss function, updates the network parameters of the main network through the gradient descent method. The formula is as follows: , in, represents the learning rate; L500, execute steps L200-L400 for multiple iterations. When the number of iterations reaches the step length N, copy the network parameters of the main network to the network parameters of the target network. Based on the optimized target network, execute steps L200-L500 for multiple iterations until the change in the network parameters of the main network is less than the threshold.

7. The task scheduling method in the vehicle networking scenario based on deep reinforcement learning according to claim 6 is characterized in that: During parameter optimization, sample data is read from the experience replay pool as input based on the experience allocation priority. The calculation formula is as follows: , in, represents a small positive number, represents a positive hyperparameter used to adjust the distribution of priorities; The empirical allocation priority is converted into a probability distribution through proportional sampling, and the probability distribution is expressed as: , in, Indicates the total amount of experience, is a hyperparameter that controls the shape of the distribution.

Citation Information

Patent Citations

  • 3D medical image detection method and system

    CN114792311A

  • Edge computing task unloading method based on deep reinforcement learning in ultra-dense network

    CN115499441A

  • Fully distributed routing method and system based on deep reinforcement learning

    CN116248164A

  • Edge under-cloud task scheduling method and system based on deep reinforcement learning

    CN116954866A

  • Multi-AGV load balancing and task scheduling method based on Dueling DQN algorithm

    CN117474295A