Vehicle-mounted edge computing network delay optimization and perception task unloading joint optimization method

By optimizing task offloading decisions between vehicles through MAPPO and RPM mechanisms, the problems of resource coordination and cross-scenario adaptability in vehicular networks are solved, achieving more efficient task offloading and latency optimization, which is suitable for vehicular edge computing networks in intelligent transportation systems.

CN120812618APending Publication Date: 2025-10-17NORTHEASTERN UNIV CHINA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511081543.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing technologies fail to fully utilize the communication resources between vehicles in vehicle networks, and the designed task offloading strategies lack cross-scenario generalization capabilities and are difficult to adapt to the dynamic and changing vehicle network environment.

Method used

The multi-agent proximal policy optimization algorithm (MAPPO) is combined with deep reinforcement learning. Through a learnable communication graph structure and ranking policy memory (RPM) mechanism, the task offloading decision between vehicles is optimized, and a Markov decision process (MDP) model is established to realize the learning and policy transfer of intelligent agents in different scenarios.

Benefits of technology

It improves the coordination capabilities between vehicles, reduces task latency, enhances the generalization and adaptability of the algorithm, adapts to the complex and ever-changing vehicle network environment, and optimizes task unloading efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120812618A_ABST
    Figure CN120812618A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data processing and transmission of vehicle intelligent driving, and relates to a vehicle-mounted edge computing network delay optimization and perception task unloading joint optimization method based on deep reinforcement learning. According to the method, on the basis of a multi-agent near-end strategy optimization algorithm (MAPPO), complex task unloading is modeled into a Markov decision process (MDP), so that each agent can continuously learn and optimize an unloading strategy in a dynamic environment. In order to enhance the cooperative capability between vehicles, a learnable communication graph structure is introduced, so that the vehicles can autonomously establish V2V communication connection based on perception information, thereby realizing more efficient task sharing and resource utilization. Besides, in order to improve the generalization ability of the algorithm, a ranking strategy memory (RPM) mechanism is designed and is used for enhancing the learning stability and strategy migration ability of multiple agents in different scenes.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of data processing and transmission of intelligent driving of vehicles, and relates to a vehicle-mounted edge computing network delay optimization and perception task offloading joint optimization method based on deep reinforcement learning. BACKGROUND

[0002] Internet of Vehicles (IOV) is entering a new stage of development as a promising application technology. By realizing interconnection between vehicles, IOV can provide high-reliability vehicle-mounted multimedia services, cooperative cruise control, and precise route navigation functions. This lays a solid foundation for the construction of intelligent transportation systems (ITSs) and significantly improves the driving safety and travel experience of passengers. For related literature, please refer to:

[0003] Under this development trend, vehicles will continuously generate a large number of delay-sensitive and computation-intensive tasks, especially in critical areas such as road safety, assisted driving, and autonomous driving. For details, see the following literature:

[0004] These advanced applications have high demands on communication and computing resources, requiring low-delay response and high-speed transmission, which leads to an exponential increase in load in the transportation network. Therefore, in order to support ITSs applications while reducing system cost and improving network efficiency, a reasonable task offloading mechanism is particularly important.

[0005] For details, see the following literature:

[0006] Cloud computing is considered an effective means to solve the above problems. However, in the process of transmitting content to remote cloud servers, due to the transmission distance and backhaul link bandwidth, cloud-based processing architecture is not always feasible. As the transmission distance increases, system delay and unreliability increase significantly, which has become a major challenge in large-scale content distribution. Although the capacity of the backhaul link is continuously improving, the utilization of wireless spectrum has approached the theoretical limit, so relying solely on cloud computing cannot completely solve this problem. For details, see the following literature:

[0007] In recent years, vehicle edge computing (VEC) has been proposed as a highly promising technical solution to address the above challenges. Vehicle edge computing (VEC) offloads tasks to mobile edge computing (MEC) servers deployed on base stations (BS), which not only significantly reduces the execution and transmission delay of tasks, but also effectively alleviates the congestion pressure of the core network. At the same time, this approach also reduces the computing and storage burden of vehicles and edge base stations. This architecture optimizes the use efficiency of network resources and improves the overall performance and reliability of the system, especially suitable for high-density user scenarios. However, if the task offloading strategy is not appropriate, it may introduce additional queuing delay and system overhead. Therefore, it is crucial to implement an intelligent task offloading mechanism. Please refer to

[0008] To further address the computing and communication challenges in high-density environments, researchers have begun to focus on how to efficiently implement task offloading and resource allocation in VEC architectures. See the following references:

[0009] Reference 1: DAI Y, XU D, MAHARJAN S, et al. Joint Load Balancing and Offloading in Vehicular Edge Computing and Networks [J / OL]. IEEE Internet of Things Journal, 2019: 4377-4387. http: / / dx.doi.org / 10.1109 / jiot.2018.2876298. DOI: 10.1109 / jiot.2018.2876298.

[0010] Reference 2: HOU Y, WANG C, ZHU M, et al. Joint allocation of wireless resource and computing capability in MEC-enabled vehicular network [J / OL]. China Communications, 2021: 64-76. http: / / dx.doi.org / 10.23919 / jcc.2021.06.006. DOI: 10.23919 / jcc.2021.06.006.

[0011] Reference 3: LIU Y, WANG S, HUANG J, et al. A Computation Offloading Algorithm Based on Game Theory for Vehicular Edge Networks [C / OL] / / 2018 IEEE International Conference on Communications (ICC), Kansas City, MO. 2018. http: / / dx.doi.org / 10.1109 / icc.2018.8422240. DOI: 10.1109 / icc.2018.8422240.

[0012] Document 4: ZHANG J, GUO H, LIU J, et al. Task Offloading in Vehicular Edge Computing Networks: A Load-Balancing Solution [J / OL]. IEEE Transactions on Vehicular Technology, 2020: 2092-2104.

[0013] http: / / dx.doi.org / 10.1109 / tvt.2019.2959410. DOI: 10.1109 / tvt.2019.2959410.

[0014] In document 1, load balancing is combined with task offloading, and a joint optimization problem is proposed to maximize system utility. To reduce the total network delay and guarantee the service reliability of vehicle user equipment (VUE), Hou et al. proposed an algorithm for joint wireless and edge computing resource allocation (JAWC) in document 2. Liu et al. designed a game theory-based offloading strategy in document 3 to improve the offloading efficiency in vehicular edge networks. In document 4, researchers proposed a load balancing offloading scheme based on software-defined networks (SDN) for FiWi-enhanced VEC networks, which realizes centralized network and vehicle information management by introducing SDN. However, these algorithms for computing offloading and resource allocation often have high computational complexity, making it difficult to achieve the expected performance in dynamic VEC environments.

[0015] In recent years, due to the strong ability of reinforcement learning in decision optimization, it has been widely applied in VEC networks to achieve efficient and flexible intelligent offloading decisions. Please refer to the following documents:

[0016] Document 5: Ning Z, Zhang K, Wang X, et al. Intelligent Edge Computing in Internet of Vehicles: A Joint Computation Offloading and Caching Solution [J / OL]. IEEE Transactions on Intelligent Transportation Systems, 2021: 2212-2225. http: / / dx.doi.org / 10.1109 / tits.2020.2997832. DOI: 10.1109 / tits.2020.2997832.

[0017] Document 6: SHI J, DU J, SHEN Y, et al. DRL-Based V2V Computation Offloading for Blockchain-Enabled Vehicular Networks [J].

[0018] Document 7: WU Z, YAN D. Deep reinforcement learning-based computation offloading for 5G vehicle-aware multi-access edge computing network [J / OL]. China Communications, 2021: 26-41. http: / / dx.doi.org / 10.23919 / jcc.2021.11.003. DOI: 10.23919 / jcc.2021.11.003.

[0019] In view of the complexity of task offloading and content caching in vehicular networks, Ning et al. document 5 proposes an efficient online algorithm OMEN, aiming to minimize the total network delay under the condition of limited energy of RSU. Shi et al. document 6 proposes a smart contract-oriented vehicle task allocation scheme to ensure that the computation offloading process between vehicles is both secure and reliable. In addition, Wu et al. document 7 proposes a joint computation offloading and task migration algorithm (JCOTM) based on deep reinforcement learning for the optimization of computation offloading and task migration in multi-user vehicle-aware multi-access edge computing networks, to achieve dual optimization of system delay and energy consumption.

[0020] The existing technology has the following problems:

[0021] On the one hand, some work only focuses on computation offloading between vehicles and edge servers (such as base stations or RSUs), ignoring the potential of vehicle-to-vehicle (V2V) communication in resource coordination and task sharing, and failing to fully exploit the computing power of vehicular networks themselves.

[0022] On the other hand, some researches, although showing good offloading effect in specific scenarios, lack the cross-scene generalization ability of the designed strategies or agents, making it difficult to adapt to dynamic and changing actual vehicular network environments, such as different vehicle densities, network topology changes or heterogeneous computing resources. Therefore, when designing task offloading schemes, how to fully integrate the characteristics of V2V communication and improve the generalization and adaptation ability of agents in various complex scenarios is still an important problem to be solved. SUMMARY

[0023] To solve the above problems, the application provides a vehicle-mounted edge computing network delay optimization and perception task offloading joint optimization method, which models the complex task offloading as a Markov decision process (MDP) based on a multi-agent proximal policy optimization algorithm (MAPPO), so that each agent can continuously learn and optimize the offloading strategy in a dynamic environment.

[0024] To enhance the cooperation ability between vehicles, the application introduces a learnable communication graph structure, so that vehicles can establish V2V communication links based on perception information, thereby realizing more efficient task sharing and resource utilization. In addition, to improve the generalization ability of the algorithm, a ranking policy memory (RPM) mechanism is designed to enhance the learning stability and strategy migration ability of multi-agent in different scenarios.

[0025] The technical solution of the application comprises the following steps:

[0026] Step 1: Scene modeling

[0027] In the vehicle edge computing (VEC) scenario of the application, it is necessary to establish a key performance indicator model for evaluating task offloading decisions, especially upload and download delays. It should be noted that the subsequent methods of the application are based on these delay indicators (specifically represented by six types of delay values represented by formulas 1-6) to optimize offloading decisions, rather than relying on specific delay calculation formulas themselves. To clearly describe the scenario and subsequent optimization objectives, a set of exemplary calculation formulas (formulas 1-6) will be used in the following to quantify the transmission delay under different offloading paths. These formulas aim to specify the sources and influencing factors of delay (such as data volume, transmission coefficient, offloading ratio), but the application is inclusive of the actual calculation method of delay. Any calculation model or method that can provide the six types of delay values defined in formulas 1-6 can be applied to the optimization framework proposed in the application.

[0028] The VEC scenario includes a local layer, an edge computing layer, and a cloud computing layer, as shown in Figure 1

[0029] The local layer has multiple vehicles with different speeds, represented by the set V = {1, 2, 3,..., v,..., V}, where V represents the number of vehicles.

[0030] ​The edge computing layer is composed of multiple MEC servers. Let M = {1, 2, 3,..., m,..., M} be the set of MEC servers, and M represents the number of MEC servers. Assume that the tasks in the environment can be perceived by vehicles, and the tasks need to be partially offloaded to other nodes of the network. Assume that a task is perceived at time Ts t, denoted as D(t) = {Di(t), β(t), Do(t), h(t)}, where Di(t) represents the task input data size, β(t) represents the computational difficulty coefficient, Do(t) represents the task output data size, and h(t) is the maximum delay that the task can tolerate. The computational difficulty coefficient β(t) varies with the consumed computing resources. For example, processing simple text data transmission (such as vehicle status reporting) is obviously less difficult than processing emergency braking control in autonomous driving.

[0031] Since the vehicles and MEC servers have limited computing and transmission resources, they will maintain some task buffers, and the arrival tasks of these task buffers include spontaneous tasks and offloaded tasks. Assume that the arrival rate of tasks generated by vehicle v and MEC m at time slot τ is and These arrival rates are independent and identically distributed random processes. The corresponding processing rates are and The distribution of the tasks generated by the nodes and the environment state, including the available computing resources of the vehicles and MECs and their task buffers, are all unknown in advance. In this environment, the processing capacity of the cloud is unlimited. The vehicle perceives the task and partially offloads it to MEC m Vehicle v and cloud p C (τ), where satisfies the requirement.

[0032] In the network model, three types of transmissions are mainly considered: transmissions between the cloud and MEC servers, between the cloud and vehicles, between MEC servers and vehicles, and between vehicles. Assume that the communication cost between vehicles is only related to λ and the data volume. The tasks perceived by vehicles are offloaded to different network nodes for processing, and the results are returned to the vehicles after processing. To standardize the delay representation of different transmission technologies, assume that the transmission delay is proportional to the transmission data volume and is independent of the transmission distance and channel conditions. Specifically, use transmission coefficient δ M_C to represent the delay between the cloud and MEC nodes. Considering the dynamic characteristics of the channel, the model can also be extended to real-world transmission scenarios, such as scenarios based on Shannon's theorem or cellular network channel models. Similarly, transmission coefficients δ M_V and δ C_V are defined for transmissions between MEC servers and vehicles and between the cloud and vehicles, respectively. When a task is offloaded to a MEC server, the transmission delay is δ Partially offloaded to vehicle v, the transmission delay is where and are the upload delay in equation (1) and the download delay in equation (2), respectively.

[0033]

[0034] From the above equations, it can be inferred that if the task is executed on the local vehicle, the transmission delay is zero. However, if the task is offloaded to the communicating vehicle, the transmission delay is determined by the product of the data size and the transmission rate λ.

[0035] Similarly, the proportion of tasks offloaded to the MEC server involves a transmission delay including and where is the upload delay defined in equation (3), is the download delay defined in equation (4).

[0036]

[0037] Equation (3) represents the uploading of data to the MEC server for processing, while equation (4) represents the transmission of the processed results from the MEC server to all vehicles. Finally, the transmission delay related to the offloading of tasks to the cloud includes and Specifically, the in equation (5) is divided into two consecutive phases: initially offloading the task to the MEC server, and subsequently transmitting it from the MEC server to the cloud. Conversely, the in equation (6) represents the delay involved in disseminating the processed results from the cloud to all vehicles.

[0038]

[0039] Step two: Establishing the task

[0040] In this scenario, it is assumed that all queuing dynamics follow the First-In-First-Out (FIFO) principle. This ensures that tasks are processed strictly in the order of their arrival, without prioritization or differential treatment based on the specific attributes of the tasks. Furthermore, the present invention operates on the premise that each task is of equal importance, thereby ensuring that no task is favored or prioritized during the scheduling and processing phases.

[0041] Let represent the computing capacity of vehicle v, and B(v, τ) represent the buffer queue of vehicle v at time slot τ. The proportion of tasks processed in vehicle v is The required queuing time is given by equation (7),

[0042]

[0043] where, is the nth task in the buffer queue B(v, t) of vehicle v. The delay of vehicle v to process the task is denoted in equation (8) as It is assumed that the processing delay is proportional to the amount of data processed and inversely proportional to the processing coefficient.

[0044]

[0045] At this time, the buffer queue length on vehicle v is denoted as

[0046]

[0047] where, is the length of the buffer queue at time slot t.

[0048] Similarly, when the task is executed on the MEC server with a proportion of the processing capacity, the queue delay and the processing delay are given by equations (10) and (11)

[0049]

[0050] where B(m, t) is the buffer queue of MEC server m at time slot t. Meanwhile, the buffer queue length of MEC server m is denoted as

[0051]

[0052] It is assumed that the processing capacity of the cloud is infinite, thus both the queue delay and the processing delay are considered as zero.

[0053] Based on the analysis of the system model above, the primary goal is to optimize the task computation offloading allocation. Given that the computation resource of vehicles is limited in this system model, the system cost is minimized with respect to the task delay within the VEC network. The total delay of the system includes the local computation layer delay the MEC layer computation delay and the cloud computation layer delay t C (τ), which can be denoted as s

[0054]

[0055] Since the task is partially offloaded to multiple nodes for processing and the processing result is transmitted to the vehicle, the total delay t(τ) in equation (16) depends on the path with the largest delay among multiple paths, without considering the assembling time.

[0056]

[0057] Considering the limited computational resources available to vehicles in this system model, the optimization scheme focuses on reducing the system cost related to task delay within the VEC network. This approach specifically optimizes online operations in each time slot without relying on future information about the environment, such as the task arrival distribution at nodes and channel dynamics. The problem is formulated as a multi-stage stochastic problem (P1),

[0058]

[0059] t(τ)<h(τ),(C3)

[0060]

[0061] where constraint C1 indicates that the sum of offloading ratios equals 1. Constraint C2 specifies that the offloading ratio of any node should be between 0 and 1. Constraint C3 ensures that the offloading and execution time of each task does not exceed the tolerable time. Constraint C4 guarantees that the length of the vehicle's buffer queue should be less than the maximum vehicle buffer queue length at any time. Constraint C5 indicates that the length of the MEC server's buffer queue should be less than the maximum MEC server buffer queue length at any time. Additionally, since the cloud computing layer has abundant storage and computing resources, the computing resource limitations of this layer are not considered.

[0062] Step three: Alternative method using deep reinforcement learning (DRL) to obtain the optimal solution of P1

[0063] When optimizing the continuous task offloading problem P1, traditional optimization methods such as convex optimization and numerical optimization are proven to be ineffective due to the complex expressions involved in the piecewise functions and objective functions in the delay model. Since the problem is non-linear and non-convex, P1 becomes a typical mixed-integer non-linear programming (MINLP) problem, which is generally considered NP-hard. Especially in the context of vehicle mobility and the randomness of computing task arrivals, using traditional methods such as exhaustive search to find the optimal solution often leads to unacceptable computational complexity.

[0064] As shown in Figure 2 , the agent samples a policy in RPM and then makes offloading decisions based on the state and communication information. Considering the unpredictable channel conditions and random arrival of computing tasks, a Markov Decision Process (MDP) is adopted to solve the P1 problem, which is defined as follows:

[0065] 1) State space: In this system, vehicles are partially observable. Therefore, the state s v (τ) of a single vehicle consists of the state of the vehicle itself and the state of the MEC server it observes. The state sv (τ)∈S can be defined as

[0066]

[0067] 2) Action space: At time slot τ, each agent chooses its action a v (τ) according to its state s i (τ). Then the actions of all agents are combined to represent the collective action of all agents on the environment at time slot τ. To satisfy the first requirement in P1, the softmax operation is used to modify the action. Thus, the action of a single agent is defined by equation (18), and the combined action of all agents is defined by equation (19).

[0068]

[0069] 3) Reward function: At state s (τ), an agent chooses an action a

[0070] (τ) and receives a reward r (τ) immediately after performing a (τ). The goal of the present invention is to maximize the cumulative reward under given constraints, and the reward function is usually related to the objective function. Assume that the computational task has a maximum delay constraint h

[0071] (τ). If the system delay exceeds the constraint h (τ), the task is considered incomplete and should incur a penalty. Therefore, the reward function can be defined as follows:

[0072]

[0073] The working mechanism is as follows: upon receiving the state s i (τ), the actor network generates a probability distribution of all possible actions. Then, through a sampling process, an action a i is selected based on these probabilities. Next, the selected action a i is applied to the environment, resulting in the observation of the next state s i+1 and the corresponding reward r i . Subsequently, the experience sequence {s i , a i , r i , s i+1} is stored in the replay buffer for later use in training the neural network.

[0072] To update the network parameters, a batch of experiences B = {(s i , a i , r i , s i+1 ), i = 1, …, N} is randomly sampled, where N is the number of experiences. Then, the state value V(s i ) can be given by equations (21) and (22),

[0073]

[0074] where a new is the action generated by the policy network at input s i R i (t) is the discounted expected future reward over time slot t. Finally, the parameters of the value network are updated using the loss function in equation (23).

[0075]

[0076] In the above equation, y i is given by equation (24), where γ is the discount factor used to compute the present value of future rewards, V ψ (s i+1 ) is the value prediction of the Critic network for the next state s i+1 .

[0077] y i = r i + γV ψ (s i+1 ) (24)

[0078] To update the parameters of the policy network, a clipped objective function is optimized. The loss function is given by equation (25):

[0079] L(θ) = E i [min(r i (θ)A i , clip(r i (θ), 1 - ∈, 1 + ∈)A i )] (25)

[0080] where is the probability ratio between the new and old policy, A i = Q ψ (s i , a i ) - V ψ (s i ) is the advantage function, and ∈ is the clipping range used to limit the policy update.

[0081] ACPO not only solves the multi-agent cooperation problem through the communication graph and shared information, allowing each agent to make more global and optimized decisions based on local observations, but also improves the reward and generalization ability of the algorithm by incorporating the RPM method.

[0082] First, the critic network parameters ψ, actor network parameters θ, and communication graph parameters The experience replay buffer B and the RPM are initialized. Then, the algorithm runs for t_eps episodes. At the beginning of each episode, the environment is reset and the parameters of the vehicles and the MEC servers are initialized.

[0083] Subsequently, the algorithm iterates from time slot 1 to time slot T in each episode. The state s ′ (τ) is obtained through the environment state s(τ) and the communication graph. Then, the agent samples a policy from the RPM and takes an action based on the local observation and the communication state. Next, the reward and the next state s(τ+1) are obtained according to the action of each agent. A small batch of samples is randomly selected to update the parameters: the Critic network is updated according to formula (23), the Actor network is updated according to formula (25), and the communication graph is jointly updated according to formulas (23) and (25). Finally, the return of the training episode is used to update the RPM. The process details of the ACPO algorithm are shown in Algorithm 1.

[0084] The beneficial effects of the present application are:

[0085] The present application builds a multi-agent vehicle simulation environment based on the Python platform, and compares the performance with six kinds of baseline schemes such as Shortest Queue (LQ), Deep Q-Network (DQN), Deep Deterministic Policy Gradient (DDPG), Soft Actor-Critic (SAC), MAPPO and MAPPO algorithm with communication graph mechanism. Under the same network configuration and task arrival conditions, by setting multiple typical test scenes, the change trend of the task completion rate, average unloading delay and training reward of the system under different algorithms is observed. The ACPO algorithm proposed in the present application performs superiorly in terms of task completion rate, with an average improvement rate of 5.23% to 10.36%, effectively reducing the perception task discard caused by task overtime or communication failure. At the same time, in terms of average delay, the ACPO algorithm can reduce the system delay by 10.22% to 20.33%, significantly improving the task response speed of the overall network. Due to the communication graph structure design and ranking strategy memory mechanism, the ACPO algorithm has higher reward and stronger generalization ability in a new environment, and can quickly complete policy learning and updating in both static road conditions and dynamic traffic flow environments. The method proposed in the present application is suitable for distributed deployment of Internet of Vehicles environment, without relying on central controller or additional auxiliary information, and has good adaptability and scalability, which can effectively cope with the complex and variable unloading decision-making requirements in future intelligent transportation systems. BRIEF DESCRIPTION OF DRAWINGS

[0086] Figure 1 is a system network architecture diagram.

[0087] Figure 2 ACPO structure diagram.

[0088] Figure 3 ACPO reward curve is shown with each algorithm, and ACPO reward is obviously better than other algorithms.

[0089] Figure 4 The task completion rate of each algorithm is compared respectively.

[0090] Figure 5 The average delay of each algorithm is compared respectively. DETAILED DESCRIPTION

[0091] The embodiments of the application will be described in detail below with reference to the technical solutions and accompanying drawings.

[0092] The task is simulated in PyTorch, and the simulation scenario is set as follows: the number of vehicles in the perception area ranges from 6 to 10 vehicles, the moving speed of each vehicle ranges from 10 to 20 m / s, and the initial queue backlog data volume ranges from 3 to 9 MB. It is assumed that the tasks generated in the scene can be perceived and processed by any MEC node, which enables the system to dynamically adjust its relative workload according to the processing capacity of each MEC node. The initial queue backlog data volume of the MEC server is set to range from 25 to 56 MB, which is much larger than the backlog of the vehicle, which reflects that random events that need to be forwarded and processed by the MEC usually have a higher arrival rate.

[0093] The transmission delay coefficient is represented by three parameters: δ M_C : represents the delay coefficient between the cloud server and the MEC server. δ M_V : represents the delay coefficient between the MEC server and the vehicle. δ C_V : represents the delay coefficient between the cloud server and the vehicle.

[0094] Where δ M_V = 3.6 < δ M_C = 4.2 < δ C_V = 7.8, from the modeling, if this inequality is not satisfied, such as δ M_V ≥ δ C_V , then there is no need to offload to the MEC server, and it can be directly offloaded to the cloud.

[0095] The task is represented using a four-dimensional vector D(t) = {Di(t), 2β(t), Do(t), h(t)} (Di(2t) represents the task input data size [1, 3] MB, β(t) represents the computation difficulty coefficient [0.7, 1.0], Do(t) represents the task output data size [0.06, 0.16] MB, and h(t) is the maximum delay that the task can tolerate [3, 7] s). The model assumes that at the beginning of each time slot, the vehicle perceives a new task and needs to offload it partially to other nodes in the network (such as other MECs or the cloud). The proposed deep reinforcement learning method employs a fully connected multi-layer perceptron (MLP) structure in its Actor module, which contains 1 input layer, 2 hidden layers (each with 256 neurons), and 1 output layer.

Claims

1. A joint optimization method for vehicle-mounted edge computing network latency optimization and perception task offloading, characterized in that: The steps include: Step 1: Scene Modeling The VEC scenario includes a local layer, an edge computing layer, and a cloud computing layer. In the local layer, there are multiple vehicles with different speeds, represented by the set V = {1, 2, 3, ..., v, ..., V}, where V represents the number of vehicles. The edge computing layer consists of multiple MEC servers. Let M = {1, 2, 3, ..., m, ..., M} be the set of MEC servers, where M represents the number of MEC servers. Assume that the tasks in the environment can be perceived by the vehicles and that some of the tasks need to be offloaded to other nodes in the network. Assume that the task is perceived at time TS t, represented by D(t) = {Di(t), β(t), Do(t), h(t)}, where Di(t) represents the task input data size, β(t) represents the computational difficulty coefficient, Do(t) represents the task output data size, and h(t) is the maximum tolerable delay of the task; Assume that the arrival rates of vehicle v and MECm at the node generation task in time slot τ are and These arrival rates are independent and identically distributed random processes; the corresponding processing rates are and Cloud computing layer: The distribution of node generation tasks and environment states, including the available computing resources of vehicles and MEC and their task buffers, are unknown in advance; in this environment, the processing capacity of the cloud is unlimited; the vehicle perception tasks are partially offloaded to vehicle Heyunρ C (τ), where meet the requirements; Step 2: Create a task Assumptions represents the computing power of vehicle v, B(v,τ) represents the buffer queue of vehicle v at time slot τ; the processing ratio in vehicle v is Mission The required queuing time is given by formula (7): in, is the nth task in the buffer queue B(v,τ) of vehicle v; the delay of vehicle v processing the task is expressed in formula (8) as For simplicity, it is assumed that the processing delay is proportional to the amount of data processed and inversely proportional to the processing coefficient; At this time, the buffer queue length on vehicle v is expressed as in, is the length of the buffer queue at time slot τ; Similarly, when the task is in the ratio When executing on the MEC server, the queue delay and processing delays Given by formulas (10) and (11) Among them, B(m,τ) is the buffer queue of MEC server m in time slot τ; at the same time, the buffer queue length of MEC server m is expressed as Step 3: Alternative method to obtain the optimal solution of P1 using deep reinforcement learning (DRL) The agent samples the policies in the RPM and then makes offloading decisions based on the state and communication information; considering the unpredictable channel conditions and the random arrival of computing tasks, a Markov decision process (MDP) is adopted to solve the P1 problem.

2. A joint optimization method for vehicle-mounted edge computing network latency optimization and perception task offloading according to claim 1, characterized in that: In step 3, a Markov decision process (MDP) is used to solve the P1 problem, which is defined as follows: 1) State space: In this system, vehicles are partially observable; therefore, the state s of a single vehicle v (τ) consists of the state of the vehicle itself and the state of the MEC server it observes; the state s at time slot τ v (τ)∈S can be defined as 2) Action space: At time slot τ, each agent chooses its action α according to its state v (τ); the actions of all agents are then combined to represent the collective action of all agents on the environment at time slot τ; to satisfy the first requirement in P1, the actions are modified using a softmax operation; thus, the action of a single agent is defined by equation (18), and the combined action of all agents is defined by equation (19). 3) Reward function: In state s(τ), the agent chooses an action A(τ) and receives a reward r(τ) immediately after executing A(τ). The goal of this invention is to maximize the cumulative reward under given constraints. The reward function is usually related to the objective function. Assume that the computational task has a maximum delay limit h(τ). If the system delay exceeds the constraint h(τ), the task is considered incomplete and should incur a penalty. Therefore, the reward function can be defined as follows: Its working mechanism is as follows: After receiving the status s i When , the actor network generates a probability distribution over all possible actions; then, through a sampling process, it selects an action a based on these probabilities. i ; Next, the selected action a i applied to the environment, resulting in the observation of the next state s i+1 and the corresponding reward r i ; then , the experience sequence {s i ,a i ,r i ,s i+1 } is stored in the replay buffer for later use in training the neural network; In order to update the network parameters, a batch of experience B = {(s i ,a i ,r i ,s i+1 ), i=1,…,N}, where N is the number of experiences; then, the state value V(s i ) can be given by equations (21) and (22), Among them, a new is the policy network at input s i The action generated by R i (t) is the discounted expected future reward at time slot τ; finally, the loss function in formula (23) is used to update the parameters ψ of the value network; In the above equation, y i Given by formula (24), where γ is the discount factor used to calculate the present value of future rewards, V ψ (s i+1 ) is the Critic network's response to the next state s i+1 Value prediction; yes i =r i +γV ψ (s i+1 ) (24) In order to update the parameters of the policy network, a clipped objective function is optimized; the loss function is shown in the following formula (25): L(θ)=E i [min(r i (i)A i ,clip(r i (θ),1-∈,1+∈)A i )](25) in, is the probability ratio between the new strategy and the old strategy, A i =Q ψ (s i ,a i )-V ψ (s i ) is the advantage function, and ∈ is the clipping range used to restrict policy updates.

3. A joint optimization method for vehicle-mounted edge computing network delay optimization and perception task offloading according to claim 1 or 2, characterized in that: In step 1, when modeling the scenario, the network model primarily considers three types of transmission: transmission between the cloud and MEC servers, between the cloud and vehicles, between MEC servers and vehicles, and between vehicles. It is assumed that the communication cost between vehicles is only related to λ and the amount of data. The tasks perceived by the vehicle are offloaded to different network nodes for processing. After processing, the results are returned to the vehicle. To standardize the latency representation of different transmission technologies, it is assumed that the transmission delay is proportional to the amount of transmitted data and is independent of the transmission distance and channel conditions. Using the transmission coefficient δ M_C To represent the delay between the cloud and MEC nodes; considering the dynamic characteristics of the channel, the model can also be extended to real-world transmission scenarios, such as those based on Shannon’s theorem or cellular network channel models; similarly, the transmission coefficient δ is defined for the transmission between the MEC server and the vehicle and between the cloud and the vehicle. M_V and δ C_V ; When the task is at a ratio When part of the unloading is done to vehicle v, the transmission delay is in and are the upload delay in formula (1) and the download delay in formula (2), respectively; According to the above formula, it can be inferred that if the task is executed on the local vehicle, the transmission delay is zero; however, if the task is offloaded to the communication vehicle, the transmission delay is determined by the product of the data size and the transmission rate λ; Likewise, proportionally The transmission delay involved in offloading tasks to MEC servers includes and in is the upload rate defined in formula (3), is the download rate defined in formula (4); Formula (3) represents uploading data to the MEC server for processing, while Formula (4) represents transmitting the processing results from the MEC server to all vehicles; finally, the transmission delay associated with task offloading to the cloud includes and Specifically, in formula (5) It is divided into two consecutive stages: initially offloading the task to the MEC server and then transferring it from the MEC server to the cloud; in contrast, the represents the delay involved in propagating the processed results from the cloud to all vehicles; 4. A joint optimization method for vehicle-mounted edge computing network latency optimization and perception task offloading according to claim 3, characterized in that: Step 2: When creating a task, the total system delay includes the local computing layer delay MEC layer calculation delay and cloud computing layer delay t C (τ), which can be expressed as s Since the task is partially offloaded to multiple nodes for processing and the processing results are transmitted to the vehicle, the total delay t(τ) in formula (16) depends on the path with the largest delay among multiple paths without considering the assembly time; The problem is formulated as a multi-stage stochastic problem (P1), t(τ) <h(τ),(C3) Among them, constraint C1 indicates that the sum of the offloading ratios is equal to 1; constraint C2 stipulates that the offloading ratio of any node should be between 0 and 1; constraint C3 ensures that the offloading and execution time of each task does not exceed the tolerable time; constraint C4 ensures that the length of the vehicle buffer queue should be less than the maximum vehicle buffer queue length at any time, and C5 indicates that the length of the MEC server buffer queue should be less than the maximum MEC server buffer queue length at any time.

5. A joint optimization method for vehicle-mounted edge computing network delay optimization and perception task offloading according to claim 1, 2 or 4, characterized in that: In step 3, first, initialize the critic network parameters ψ, actor network parameters θ and communication graph parameters The experience replay buffer B and RPM are initialized; then, the algorithm runs t_eps rounds; at the beginning of each round, the environment is reset and the parameters of the vehicle and MEC server are initialized; Subsequently, the algorithm iterates from time slot 1 to time slot T in each round; the state s′(τ) is obtained through the environment state s(τ) and the communication graph; then, the agent samples the policy from the RPM and takes an action based on the local observation and communication state; next, the reward and the next state s(τ+1) are obtained according to the action of each agent; a small batch of samples is randomly selected to update the parameters: the Critic network is updated according to formula (23), the Actor network is updated according to formula (25), and the communication graph is jointly updated according to formulas (23) and (25); finally, the reward of the training round is used to update the RPM.

6. A joint optimization method for vehicle-mounted edge computing network latency optimization and perception task offloading according to claim 3, characterized in that: In step 3, first, initialize the critic network parameters ψ, actor network parameters θ and communication graph parameters The experience replay buffer B and RPM are initialized; then, the algorithm runs t_eps rounds; at the beginning of each round, the environment is reset and the parameters of the vehicle and MEC server are initialized; Subsequently, the algorithm iterates from time slot 1 to time slot T in each round; the state s′(τ) is obtained through the environment state s(τ) and the communication graph; then, the agent samples the policy from the RPM and takes an action based on the local observation and communication state; next, the reward and the next state s(τ+1) are obtained according to the action of each agent; a small batch of samples is randomly selected to update the parameters: the Critic network is updated according to formula (23), the Actor network is updated according to formula (25), and the communication graph is jointly updated according to formulas (23) and (25); finally, the reward of the training round is used to update the RPM.

Citation Information

Cited By

  • Complex task calculation unloading method, device and equipment for industrial internet of things and medium

    CN121387573A