A time-rule recognition intelligent unloading optimization method based on reinforcement learning
By adopting a time law recognition intelligent unloading optimization method based on reinforcement learning in the vehicle edge computing environment, combining Transformer and LSTM models to optimize the unloading strategy of computing tasks, the problem of unloading management and optimization of computing tasks in the vehicle edge computing environment is solved, and efficient computing resource utilization and system performance improvement is achieved.
Patent Information
- Application Number
- CN202411135927.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-19
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-08-19
AI Technical Summary
In vehicle edge computing environments, it is difficult for the prior art to effectively manage and optimize the unloading of computing tasks, especially in consideration of the diversity of vehicle resources, differences in parking behaviors, and time regularity of tasks.
The intelligent unloading optimization method based on reinforcement learning is adopted to obtain vehicle, equipment and task status information in the parking lot, build a Markov process, and combine Transformer and LSTM models to optimize the unloading strategy of computing tasks.
It significantly improves the utilization efficiency of computing resources and system performance, reduces the total service cost of the system, optimizes the calculation and uninstallation process, and enhances the reliability of task uninstallation decisions.
Smart Images

Figure CN119127333B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of mobile edge computing and computing offloading optimization, and in particular, relates to an intelligent offloading optimization method based on time law recognition based on reinforcement learning. Background Art
[0002] With the rapid development of mobile networks and 5G technology, more and more electronic devices, such as smartphones, tablets, street lights and cameras in smart cities, are connected to the network. These devices often require a lot of computing resources when performing various tasks that require low latency, such as AI algorithms, augmented reality (AR) and virtual reality (VR). If all data is offloaded to the cloud server, it may not only cause overload of the cloud server, increase transmission delay and bandwidth cost, but also cause data security issues. Due to size and design limitations, the computing power of IoT devices is usually insufficient to support complex computing tasks. In this context, Mobile Edge Computing (MEC) came into being, which provides a solution to support IoT devices to offload complex tasks by utilizing proximal computing resources. Especially in the field of Vehicle Edge Computing (VEC), with the popularization of new energy vehicles, it becomes feasible to use vehicles in parking lots as computing nodes. Doing so can offload the computing tasks of nearby devices to these vehicles, which not only significantly improves the computing performance of IoT devices, but also reduces dependence on cloud services, thereby reducing latency and enhancing data security. However, the high mobility of vehicles and the single computing environment of on-board systems pose challenges to task offloading, requiring more precise management and optimization of computing tasks in this environment.
[0003] In order to adapt to the dynamic vehicle edge computing environment, computation offloading based on deep reinforcement learning has become an effective solution strategy. However, many research works in the field of deep reinforcement learning (DRL) have ignored the importance of container scheduling when formulating computation offloading strategies, failed to fully consider the diversity of different vehicle resources and the differences in parking behaviors, and the temporal regularity of vehicle processing tasks, all of which may affect the effectiveness and efficiency of classic deep reinforcement learning algorithms such as Deep Q-Network (DQN) and Proximal Policy Optimization (PPO). In view of these problems, classic deep learning models such as Transformer and LSTM have attracted widespread attention from academia and industry. As a powerful natural language processing model, the Transformer structure is good at capturing dependency information between units and generating sequences. At the same time, the long short-term memory network is good at identifying task arrival sequences with temporal patterns. Based on these advantages, a new sequence recognition task scheduling algorithm (Sequence-Aware Task Scheduling, SATS) is proposed. Based on the deep reinforcement learning model, the algorithm innovatively integrates the above two structures with memory functions. This fusion significantly improves the understanding of dynamic management and optimization of vehicle network resources and helps to accurately identify task sequences with temporal patterns, thereby improving the efficiency and energy saving of task scheduling.
[0004] Among the existing solutions to the problem of computational offloading of stopped vehicles, research work can be divided into two categories. One is the computational offloading scheme based on traditional heuristic algorithms, and the other is the online learning computational offloading scheme based on deep learning. One of the challenges faced by computational offloading schemes based on heuristic algorithms is that there are many assumptions, and the effect will be better in specific scenarios, but the portability and robustness of the algorithm are poor. In the era of MEC and 5G, the wireless communication environment and computing tasks have become more complex. It is very challenging to design a computational offloading optimization algorithm that can effectively improve system efficiency and meet system requirements. The computational offloading scheme based on DRL can learn the future direction from the data, so it can effectively solve the offloading strategy in some complex systems. Summary of the invention
[0005] In order to solve the above technical problems, the present invention proposes a time law recognition intelligent unloading optimization method based on reinforcement learning, which realizes the learning and extraction of computing task characteristics, thereby effectively utilizing idle computing resources in the vehicle, reducing the total service cost of the system, and optimizing the computing unloading process to improve the utilization efficiency of computing resources and system performance.
[0006] To achieve the above object, the present invention provides a time law recognition intelligent unloading optimization method based on reinforcement learning, comprising:
[0007] Obtain information about vehicles, equipment and task status parked in the parking lot;
[0008] Construct a Markov process based on the parked vehicles, equipment and task status information;
[0009] Based on the Markov process, intelligent calculation and unloading optimization are performed on the vehicles parked in the parking lot.
[0010] Furthermore, the information of the status of vehicles, equipment and tasks parked in the parking lot is obtained, and the Markov process method is constructed including:
[0011] Define a state space, and divide the state space into node state, mirror state, task state, and historical trajectory;
[0012] defining an action space to assign tasks to parked vehicles or local equipment, taking into account the configuration of containers therein, where the tasks are either processed locally or unloaded to parked vehicles;
[0013] Define a reward function where the reward includes the expected and actual cost of the task.
[0014] Furthermore, the node status is represented as:
[0015]
[0016] Among them, μ1(t) to μ M (t) represents the memory of each node at time t, d1(t) to d M (t) represents the remaining storage space of each node at time t, x1(t) to x M (t) and y1(t) to y M (t) represents the x-coordinate and y-coordinate of the node at time t, arrive and arrive Respectively represent the transmission power and computing power of each node, B1 to B M It indicates the bandwidth allocated to the node, F1 to F M Represents the CPU frequency of each node;
[0017] The mirror state is represented as:
[0018]
[0019] The task status is expressed as:
[0020]
[0021] Among them, d k is the size of task k, m k is the memory required by task k, f k is the total number of CPU cycles required to complete the computational task, i k is the specific image required for the task, is the arrival time of the task, is the maximum tolerable deadline for completing this task; K(t) refers to the set of all tasks generated at time t. It is to gather together the information of all tasks generated at time t;
[0022] The historical trajectory is expressed as:
[0023] Furthermore, the action space is expressed as:
[0024]
[0025] Among them, a t is the action space at time t, The numbers of all available parking vehicles in the parking lot. For all devices that need task scheduling, K t The set of tasks generated by all devices at time t.
[0026] Furthermore, the reward function is expressed as:
[0027]
[0028] Among them, r t is the reward at time t, w t is the weight parameter, is the ideal total delay of task execution in time slot t, is the total delay actually consumed, K(t) is the set of all tasks generated at time t, is the ideal total energy cost, is the actual total energy cost.
[0029] Furthermore, the method for intelligently calculating and optimizing the unloading of vehicles parked in a parking lot based on the Markov process includes:
[0030] Initialize the policy network parameters, value network parameters and sampling strategy;
[0031] The policy network model in reinforcement learning is updated to a network model that combines Transformer and LSTM, where the environment state is obtained by observing in the simulation environment at each time slot t of each round. Processing through Transformer get Processed by fully connected layers get Processing via long short-term memory networks get The final merged state is
[0032] Initialized policy network parameters, value network parameters, and rewards for sampling strategy calculation tasks;
[0033] Training a policy gradient network based on the reward of the task to obtain the relative advantage of each action, introducing a restricted alternative objective function to replace the original update function based on the relative advantage of each action, and completing the policy gradient network training;
[0034] The method to obtain the relative advantage of each action is:
[0035] A π (s′ t , a t )=Q π (s′ t , a t )-V π (s′ t )
[0036] Among them, A π (s′ t , a t ) is the relative advantage of each action, Q π (s′ t , a t ) is the state-action value function, which means that in state s′ t Take action a t And following the value of strategy π, V π (s′ t ) means in state s′ t The expected return of following strategy π;
[0037] Based on the trained policy gradient network, the energy consumption and delay of the system are calculated to obtain the total service cost, wherein the delay includes transmission delay, download delay and calculation delay, and also includes local calculation delay and local download delay, and the energy consumption includes unloading energy consumption and local energy consumption;
[0038] An offloading policy is updated based on the total service cost.
[0039] Further, the transmission delay is expressed as:
[0040]
[0041] in, is the transmission delay of task k, d k is the size of the task, The speed from the corresponding device to the designated parked vehicle, is an indicator symbol;
[0042] The download delay is expressed as:
[0043]
[0044] in, is the download delay of task k, is the size of the image, is the transmission delay from the base station to the stopped vehicle n, is the queuing delay of task k on parked vehicle n when it arrives, is an indicator symbol, indicating that the task is assigned to stop vehicle n, It is an indicator symbol of the mirror information;
[0045] The computation delay is expressed as:
[0046]
[0047] Here f k is the number of CPU cycles required for task k, and F n is the CPU calculation frequency of parked vehicle n.
[0048] Furthermore, the local calculation delay is expressed as:
[0049]
[0050] in, is the local computation delay of task k, f k is the number of CPU cycles required for task k, is the CPU frequency of the device that generates task k, β k (t) is a local indicator symbol;
[0051] The local download delay is expressed as:
[0052]
[0053] where β k (t) represents the indicator symbol of local execution, when β k When (t) = 1, task k is assigned to local execution and task q is assigned to local execution. k .
[0054] Furthermore, the unloading energy consumption is expressed as:
[0055]
[0056] in, is the calculated power of parking vehicle n, is the transmission power of parked vehicle n.
[0057] The local energy consumption is expressed as:
[0058]
[0059] in, is the computational energy cost of executing task k locally, is the computational latency cost of executing task k locally, is the computing power of the corresponding local device, β k (t) is the indicator symbol for local execution, is the transmission energy cost of local execution, is the transmission power for the local implementation.
[0060] Furthermore, the total service cost is expressed as:
[0061]
[0062] Among them, w t is the weight parameter, is the total delay cost required to execute task k, is the total energy cost required to execute task k.
[0063] Technical effect of the present invention: The present invention discloses an intelligent unloading optimization method for identifying time rules based on reinforcement learning, which uses containerization technology to enhance the computing and storage capabilities in the MEC environment, by assembling containers in parked vehicles and providing lightweight and extensive task offloading services inside parked vehicles. The parked vehicle downloads the appropriate container from the base station to execute the task unloaded from the requesting device. In addition, the arrival of tasks in the present invention follows clear rules, such as traffic peaks in specific time periods, which is different from the assumption that tasks arrive randomly in existing solutions. This design is more in line with the actual situation in the real world and highlights the efficiency of the strategy. In order to solve the problem that existing models ignore task patterns and fail to recognize and adapt to the complex and evolving relationships between parked vehicles, the present invention integrates LSTM and Transformer models into our system. This integration captures the historical patterns of task arrival and deepens the understanding of the interconnection between parked vehicles. In this way, the performance of the system is significantly improved, especially in reducing delay costs and enhancing the reliability of task offloading decisions. This paper proposes a policy gradient-based reinforcement learning algorithm named SATS, which is designed for VEC networks. It focuses on solving the delay and limited computing power problems of PVs and establishes an optimization problem to reduce both energy and time costs in the network. In order to control the service cost of the overall system, a weight factor is set for the system, which is set by the system administrator according to the actual application scenario. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0065] Figure 1 The present invention provides an architecture for optimizing computational offloading by using a SATS algorithm in a scenario where a vehicle is used as a node in an embodiment of the present invention;
[0066] Figure 2 The SATS algorithm model is combined with the identification timing component according to the embodiment of the present invention. DETAILED DESCRIPTION
[0067] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0068] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0069] like Figure 1-2 As shown, this embodiment provides a time law recognition intelligent unloading optimization method based on reinforcement learning, including:
[0070] Obtain information about vehicles, equipment and task status parked in the parking lot;
[0071] Construct a Markov process based on the parked vehicles, equipment and task status information;
[0072] Based on the Markov process, intelligent calculation and unloading optimization are performed on the vehicles parked in the parking lot.
[0073] Among the existing computation offloading optimization schemes, the heuristic algorithm-based schemes have many assumptions and their algorithm portability and robustness are poor; the deep learning-based online learning schemes fail to effectively adapt to the dynamic environment. Therefore, an intelligent online optimization method based on reinforcement learning is proposed to solve the above problems. First, a vehicle task offloading system model assisted by nearby base stations is constructed. A container-based vehicle edge computing framework is introduced, which is customized for parking scenarios, enabling parked vehicles to download necessary images to offload tasks in the VEC environment. In this scenario, a base station (BS), as a central node, provides service coverage to N parked vehicles and devices numbered from N to M. The system time is divided into discrete time frames of equal length. Each parked vehicle is equipped with a computing unit, which is defined by bandwidth, maximum storage capacity, CPU frequency, maximum memory and remaining memory in the time slot, and the coordinates of the parked vehicle. For task processing, PV must start a container, which relies on locally accessible images. To ensure uninterrupted task execution, the necessary images must be pre-obtained and locally stored. The image library is represented by I, and each image i corresponds to a different container. Assume that the requested container is equivalent to the image associated with the request, and the images are stored in the BS.
[0074] The algorithm first collects information about the parked vehicles, equipment, and task status in the environment. This information is analyzed through three different processing branches. The Transformer branch processes the features of the parked vehicles and equipment, as well as the related mirror features, aiming to capture the complex dependencies between these elements and output a vector processed by the Transformer. The fully connected branch focuses on the representation of the current state. It accepts task features as input and outputs encoded task features. At the same time, the long short-term memory network (LSTM) branch is responsible for identifying temporal patterns, processing historical data of tasks, and generating encoded task history features. The outputs of these three branches are connected with a copy of the vehicle and equipment features, merged into total features, and then input into the policy network to determine the best scheduling action. The actions taken by the policy network generate rewards, which in turn serve as feedback for subsequent decisions. The core of the SATS algorithm lies in the use of policy gradient methods to promote policy optimization. This method directly optimizes the policy to maximize the expected return. The policy gradient method is adopted because of its good balance between effect and computational efficiency, and improves the stability and performance of learning by fine-tuning the policy, making it very suitable for the SATS algorithm.
[0075] In the modeling of this system, the duration of each time slot is carefully adjusted to ensure that the state transition probability remains almost unchanged over a long period of time based on the task resource requirements and does not follow a uniform distribution. In addition, the way tasks arrive and the way the environment is updated exhibit a property that past events do not affect future results, which is called the memoryless property. Given these characteristics, this scenario is effectively represented using the Markov Decision Process (MDP) framework.
[0076] Step 1): Define the state space. The state space includes the complete framework of the VEC environment, which consists of three key elements: computing nodes (including parked vehicles and equipment), user-requested tasks, and the history of task arrival. In order to analyze the dependencies of each node and its computing resources in more detail, the state is further subdivided into four parts: node state, mirror state, task state, and historical trajectory. The node state is expressed by the following formula:
[0077]
[0078] where μ1(t) to μ M (t) represents the memory of each node at time t, d1(t) to d M (t) represents the remaining storage space of each node at time t (used to store the image of the execution task), x1(t) to x M (t) and y1(t) to y M(t) represents the x-coordinate and y-coordinate of the node at time t, arrive and arrive Respectively represent the transmission power and computing power of each node, B1 to B M It indicates the bandwidth allocated to the node, F1 to F M Indicates the CPU frequency of each node.
[0079] In order to represent the image storage information and take into account the limited number of image types, it is necessary to track whether each node stores these images. Therefore, a matrix of 0s and 1s is constructed as the image state:
[0080]
[0081] In this matrix, each column represents whether all devices contain a specific image i, and each row represents the situation of different image libraries I contained in a single device. 0 in the matrix means that the corresponding device does not store the corresponding image, and 1 means that the device has stored the image.
[0082] The task status includes the image number required for execution, as well as the resources and constraints requested by the task. Therefore, the task status is represented as follows:
[0083]
[0084] where d k is the size of task k, m k is the memory required by task k, f k is the total number of CPU cycles required to complete the computational task, i k It is a specific image required for the task, selected from a limited number of available images. is the arrival time of the task, and is the maximum tolerable deadline for completing this task. K(t) refers to the set of all tasks generated at time t. It is to bring together the information of all tasks generated at time t.
[0085] The historical state of time slot t is given by It indicates that the number of tasks arriving in the previous time slots is recorded. By looking back at the past 5 time slots and using the LSTM model based on these historical backgrounds, the model analyzes and predicts the data.
[0086] Step 2): Define the action space. The policy assigns tasks to parked vehicles or local devices, taking into account the container configuration. It determines whether the task should be processed locally or offloaded to the parked vehicle. Therefore, the action space includes all parked vehicles and devices, represented as the union of the set of parked vehicles and the requesting device, as follows:
[0087]
[0088] Step 3): Define the reward function. It is crucial to define the right reward in a reinforcement learning algorithm. Since different tasks require different amounts of computing power, considering only the total latency may lead to an unstable training process. Therefore, the reward includes both the expected and actual cost of the task and is defined as follows:
[0089]
[0090] where w t is the weight parameter, defines the ideal delay for task execution in time slot t, and F max Represents the maximum CPU frequency among all parked vehicles and devices. The task completion time is close to , which reflects the efficiency of time utilization. Similarly, the ideal energy cost use That is, the lowest value calculated among all PVs and devices is used for calculation. This reduces the incentives when the energy usage is minimized, emphasizing the importance of minimizing energy usage to improve overall efficiency.
[0091] Specific intelligent computing offloading optimization method:
[0092] Step 1): Parameter initialization. Initialize the strategy network parameters θ, value network parameters φ, and initialize the sampling strategy θ π , and θ old Set to θ. Determine the rectified linear unit Relu(x) = max(0, x) as the activation function
[0093] Step 2): Calculate the reward of the task. In each time slot t of each round, the environment state is obtained by observing the simulation environment Processing through Transformer get Processed by fully connected layers get Processing via long short-term memory networks get Calculate the merged state
[0094] Step 3): Train the policy gradient network. Policy gradient methods aim to explicitly optimize the policy by evaluating the gradient of the expected reward with respect to the policy parameters. In particular, Actor-Critic algorithms, as a subset of policy gradient techniques, combine the principle of policy enhancement (performed by the "actor") with the principle of value function approximation (performed by the "critic"). In this structure, the role of the critic involves approximating the value function and guiding the actor in policy updates. This process uses an advantage function to accurately estimate the policy gradient, which quantifies the relative advantage of each action by evaluating how an action compares to the average action in a specific state, and its formula is as follows:
[0095] A π (s′ t , a t )=Q π (s′ t , a t )- V π(s′ t );
[0096] Here the value function V π (s′ t ) means that in state s′ t The expected return of following strategy π is Q π (s′ t , a t ) is the state-action value function, which means that in state s′ t Take action a t And follow the value of strategy π. In order to make the strategy update more stable, a more sophisticated advantage estimator, the Generalized Advantage Estimator (GAE), is used. The calculation method is as follows:
[0097]
[0098] In this equation, λ is the GAE parameter, which plays a key role in balancing bias and variance, and is used to reduce variance and accelerate the learning process. t represents the temporal difference (TD) error at time step t. Therefore, the update function of the value network, i.e. the critic, is expressed as:
[0099]
[0100] In addition, in order to avoid falling into the local optimum during training and sampling, Proximal Policy Optimization (PPO) is adopted. This method introduces a restricted alternative objective function to replace the original update function, which is expressed as
[0101]
[0102] Step 4): Calculate the system energy consumption and delay. The delay in the system can be divided into transmission delay, download delay and calculation delay. First, define the transmission rate between the two devices at time t as:
[0103]
[0104] When task k is assigned to PV n, the transmission delay can be expressed as:
[0105]
[0106] where q k represents the device that originally generated task k, so It refers to the speed from the corresponding device to the designated parked vehicle. As an indicator symbol, , which indicates that task k is assigned to parking vehicle n.
[0107] Each PV node is assigned a specific storage capacity, and this capacity and memory constraint ensures the total size and memory used by all containers hosted on a single PV. To effectively manage this constraint, each PV manages its container storage queue using a Least Frequently Used (LFU) strategy. When an incoming task requires a container that is not currently in the queue and the existing storage or memory is insufficient, the protocol requires the removal of the container with the least number of accesses in the queue. This measure is to reclaim the necessary space. This iterative process of replacement and reallocation continues until the PV can reasonably accommodate the requirements of new containers for newly incoming tasks. When task k is assigned to PVn, the download delay is expressed as:
[0108]
[0109] Here It is the indicator symbol of the mirror information. When it indicates mirror image k (This refers to the image number required for task k) stored on parked vehicle n. Conversely, when it is equal to 0, it means that n lacks the corresponding image. represents the queuing delay at parked vehicle n when task k arrives. Refers to the size of the image, N cRefers to BS. This formula is understood as the time to download the image corresponding to the task to be executed from the BS plus the queuing time.
[0110] The computation time refers to the processing time of task k on a specified PV, expressed as:
[0111]
[0112] Here f k is the number of CPU cycles required for task k, and F n is the CPU calculation frequency of parked vehicle n.
[0113] The calculated power of parked vehicle n is The transmission power is The corresponding computational energy consumption and transmission energy consumption are expressed as:
[0114]
[0115] When the task is predicted by the algorithm and assigned to the local execution, the download delay of the local device is expressed as:
[0116]
[0117] where β k (t) Correspondence Indicates the indicator symbol for local execution. When β k When (t) = 1, task k is assigned to local execution, that is, assigned to q k .
[0118] The computation delay of the local device is expressed as:
[0119]
[0120] The corresponding energy consumption is expressed as:
[0121]
[0122] The total service cost is expressed as:
[0123]
[0124] The total delay and total energy consumption are expressed as:
[0125]
[0126] Step 5): Update the unloading strategy. After collecting enough samples in the environment, update the policy network. The above steps are summarized as Algorithm 1, as shown in Table 1.
[0127] Table 1
[0128]
[0129] The present invention discloses an intelligent unloading optimization method for time regularity recognition based on reinforcement learning, which uses containerization technology to enhance the computing and storage capabilities in the MEC environment, by assembling containers in parked vehicles and providing lightweight and extensive task offloading services inside the parked vehicles. The parked vehicles download appropriate containers from the base station to execute the tasks unloaded from the requesting device. In addition, the arrival of tasks in the present invention follows clear rules, such as traffic peaks in specific time periods, which is different from the assumption that tasks arrive randomly in existing solutions. This design is more in line with the actual situation in the real world and highlights the efficiency of the strategy. In order to solve the problem that the existing models ignore task patterns and fail to recognize and adapt to the complex and evolving relationships between parked vehicles, the present invention integrates LSTM and Transformer models into our system. This integration captures the historical patterns of task arrival and deepens the understanding of the interconnected relationships between parked vehicles. In this way, the performance of the system is significantly improved, especially in reducing delay costs and enhancing the reliability of task offloading decisions. This paper proposes a policy gradient-based reinforcement learning algorithm named SATS, which is designed for VEC networks. It focuses on solving the delay and limited computing power problems of PVs and establishes an optimization problem to reduce both energy and time costs in the network. In order to control the service cost of the overall system, a weight factor is set for the system, which is set by the system administrator according to the actual application scenario.
[0130] The above are only preferred specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A time law recognition intelligent unloading optimization method based on reinforcement learning, characterized in that: The following steps are involved: Obtain information about vehicles, equipment and task status parked in the parking lot; Construct a Markov process based on the parked vehicles, equipment and task status information; Based on the Markov process, intelligent calculation and unloading optimization of the vehicles parked in the parking lot are performed; Obtaining the status information of vehicles, equipment and tasks parked in the parking lot, and constructing the Markov process method include: Define a state space, and divide the state space into node state, mirror state, task state, and historical trajectory; defining an action space to assign tasks to parked vehicles or local equipment, taking into account the configuration of containers therein, where the tasks are either processed locally or unloaded to parked vehicles; Define a reward function where the reward includes the expected and actual cost of the task; The node status is represented as: ; in, arrive represents the memory of each node at time t, arrive represents the remaining storage space of each node at time t, arrive and arrive Respectively represent the x-coordinate and y-coordinate of the node at time t, arrive and arrive Respectively represent the transmission power and computing power of each node, arrive It indicates the bandwidth allocated to the node. arrive Represents the CPU frequency of each node; The mirror state is represented as: ; The task status is expressed as: ; in, is the size of task k, is the memory required for task k, is the total number of CPU cycles required to complete the computational task, is the specific image required for the task, is the arrival time of the task, is the maximum tolerable deadline for completing this task; Refers to the set of all tasks generated at time t, It is to gather together the information of all tasks generated at time t; The historical trajectory is expressed as: .
2. The time law recognition intelligent unloading optimization method based on reinforcement learning as claimed in claim 1 is characterized in that: The action space is expressed as: ; in, is the action space at time t, The numbers of all available parking vehicles in the parking lot. For all devices that need task scheduling, The set of tasks generated by all devices at time t.
3. The time law recognition intelligent unloading optimization method based on reinforcement learning as claimed in claim 1 is characterized in that: The reward function is expressed as: ; in, is the reward at time t, is the weight parameter, is the ideal total delay of task execution in time slot t, is the total delay actually consumed, is the set of all tasks generated at time t, is the ideal total energy cost, is the actual total energy cost.
4. The time law recognition intelligent unloading optimization method based on reinforcement learning as claimed in claim 1 is characterized in that: The method for intelligently calculating and optimizing the unloading of vehicles parked in a parking lot based on the Markov process includes: Initialize the policy network parameters, value network parameters and sampling strategy; Update the policy network model in reinforcement learning to a network model that combines Transformer and LSTM, where in each time slot of each round Obtaining environmental status by observing in a simulated environment , processed by Transformer get ; Processed by fully connected layers get ; Processed by long short-term memory networks get The final merged state is ; Initialized policy network parameters, value network parameters, and rewards for sampling strategy calculation tasks; Training a policy gradient network based on the reward of the task to obtain the relative advantage of each action, introducing a restricted alternative objective function to replace the original update function based on the relative advantage of each action, and completing the policy gradient network training; The method to obtain the relative advantage of each action is: ; in, For the relative advantages of each action, is the state-action value function, which means that in the state Take action And follow the strategy The value of Indicates in status Compliance Strategy Expected return; Based on the trained policy gradient network, the energy consumption and delay of the system are calculated to obtain the total service cost, wherein the delay includes transmission delay, download delay and calculation delay, and also includes local calculation delay and local download delay, and the energy consumption includes unloading energy consumption and local energy consumption; An offloading policy is updated based on the total service cost.
5. The time law recognition intelligent unloading optimization method based on reinforcement learning as claimed in claim 4 is characterized in that: The transmission delay is expressed as: ; in, is the transmission delay of task k, is the size of the task, is the speed from the corresponding device to the designated parked vehicle, is an indicator symbol; The download delay is expressed as: ; in, is the download delay of task k, is the size of the image, is the transmission delay from the base station to the stopped vehicle n, is the queuing delay of task k on parked vehicle n when it arrives, is an indicator symbol, indicating that the task is assigned to stop vehicle n, It is an indicator symbol of the mirror information; The computation delay is expressed as: ; Here is the number of CPU cycles required for task k, and is the CPU calculation frequency of parked vehicle n.
6. The time law recognition intelligent unloading optimization method based on reinforcement learning as claimed in claim 4 is characterized in that: The local computation delay is expressed as: ; in, is the local computation delay of task k, is the number of CPU cycles required for task k, is the CPU frequency of the device that generates task k, It is a local indicator symbol; The local download delay is expressed as: ; in Indicates the local execution indicator. When , it means that task k is assigned to local execution, and assigned to .
7. The time law recognition intelligent unloading optimization method based on reinforcement learning as claimed in claim 4 is characterized in that: The unloading energy consumption is expressed as: ; ; in, is the calculated power of parking vehicle n, is the transmission power of parked vehicle n; The local energy consumption is expressed as: ; ; in, is the computational energy cost of executing task k locally, is the computational latency cost of executing task k locally, is the computing power corresponding to the local device, Indicator for local execution, is the transmission energy cost of local execution, is the transmission power for the local implementation.
8. The time law recognition intelligent unloading optimization method based on reinforcement learning as claimed in claim 4 is characterized in that: The total service cost is expressed as: ; in, is the weight parameter, is the total delay cost required to execute task k, is the total energy cost required to execute task k.
Citation Information
Patent Citations
Internet of vehicles task unloading method and system and electronic equipment
CN116112525A
Multi-agent reinforcement learning Internet of Vehicles calculation unloading method
CN116633936A