Multi-RSU vehicle-road cooperation task unloading optimization method based on MCW-TD3 algorithm
By employing the DDCPN algorithm in the perception of vehicle-road cooperative infrastructure, the problem of balancing high perception accuracy and latency is solved by optimizing task offloading and resource allocation. This enables efficient resource utilization and fast inference under dynamic communication links, reducing latency and improving system efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies have failed to effectively balance the need for high perception accuracy with strict latency requirements in vehicle-road cooperative infrastructure perception, and have not fully considered dynamic communication link quality and resource allocation, resulting in insufficient utilization of computing resources and bandwidth.
A latency-accuracy balance optimization method based on DDCPN is adopted. By constructing Markov Decision Process (MDP) and Discrete Driven Continuous Policy Network (DDCPN) algorithms, task offloading and resource allocation are optimized. Combined with the computing resources of vehicles and RSUs, partial offloading and dynamic resource allocation are achieved to reduce task execution latency and ensure fast inference accuracy.
Under the constraints of fast inference accuracy and energy consumption, it significantly reduces task execution latency, improves resource utilization, adapts to dynamic and complex vehicular network environments, and achieves faster convergence speed and higher system efficiency.
Smart Images

Figure CN121793084A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of vehicle networking and relates to the research on vehicle matching and resource allocation strategies for optimizing the end-to-end latency problem in the scenario of vehicle-road cooperative infrastructure perception. Background Technology
[0002] While vehicle-to-infrastructure (V2I) perception holds great promise, it faces significant challenges. Transmitting large amounts of raw sensor data from intelligent vehicles to a central server consumes enormous bandwidth and introduces substantial latency, especially under poor wireless conditions. Conversely, processing such data locally on the vehicle is often constrained by limited onboard computing resources. Balancing the demands for high perception accuracy with stringent latency requirements remains a key obstacle for Advanced Driver Assistance Systems (ADAS). Therefore, this paper proposes a novel parallel perception framework that incorporates an adaptive partial computation offloading strategy, enabling the task vehicle to offload some of its computational tasks to adjacent vehicles or roadside units (RSUs), effectively alleviating resource constraints and reducing latency. We formulate a joint optimization problem that minimizes task execution latency while ensuring high-speed inference accuracy and efficient resource utilization. By modeling this as a Markov Decision Process (MDP) and employing an algorithm based on a Discrete-Driven Continuous Policy Network (DCPN), our solution exhibits superior performance in balancing accuracy and latency, enabling robust environmental perception.
[0003] A search revealed application publication number CN119136257A, entitled "Link Topology Adaptive Offloading Method for Edge Computing in Vehicle-to-Everything (V2I) Systems." This invention constructs a dynamic V2I and V2V joint system, refining the latency minimization problem of sequential subtasks within the joint system as a path optimization problem. It defines the state space, state at each time step, action space, and reward of a Markov Decision Process (DDQN), modeling the path optimization problem as a Markov Decision Process. The Markov Decision Process utilizes a Generative Network (GCN) to adaptively learn hidden information from the states at different time steps based on different link topologies, providing this information to a Direct Decision-Making Network (DDQN) to select the action that yields the highest reward, thus minimizing the total latency. This invention develops a fine-grained partial offloading model for sequential subtasks in dynamic V2I and V2V joint systems with link topologies, optimizing collaborative offloading strategies. Based on GCN and DDQN adaptive offloading methods, it adapts to different link topologies and minimizes total latency.
[0004] However, this invention still has some shortcomings. Because it focuses solely on global latency research, it doesn't consider the effect of task processing or the accuracy of inference. Therefore, this invention addresses this issue by considering whether the accuracy of the processed task meets the requirements when processing inference tasks. If not, it considers proceeding to the next calculation. This paper adopts a global inference approach. Furthermore, the invention in CN119136257A only considers the offloading strategy as the variable to be optimized, failing to consider that the communication link quality is constantly changing during vehicle dynamics, and resource allocation is also insufficiently considered. However, this invention considers the Uu interface and PC5 interface for V2I and V2V links respectively, with OFDMA access under the Uu interface. This is closer to the road conditions in real-world vehicle networking scenarios. Moreover, this paper provides a more granular division of task offloading, fully considering the computing resources between nodes, and also performs a more granular division of communication resources, improving communication rates under adverse channel conditions by allocating sub-channels. Summary of the Invention
[0005] This invention aims to solve the problems of the prior art. It proposes a time-delay-accuracy balance optimization method based on DDCPN in vehicle-road cooperative infrastructure perception. The technical solution of this invention is as follows:
[0006] A time-delay-accuracy balance optimization method based on DDCPN in vehicle-road cooperative infrastructure perception includes the following steps:
[0007] Construct a vehicle-to-everything (V2X) system, which consists of multiple vehicles and a single RSU;
[0008] Establish to minimize all mission vehicles in The problem involves optimizing the total delay of completing a perception task within a given timeframe; the optimization problem is modeled as a Markov Decision Process (MDP).
[0009] Task offloading and resource allocation optimization are performed based on the Discrete-Driven Continuous Policy Network (DDCPN) algorithm.
[0010] Furthermore, the construction of the vehicle networking system specifically includes:
[0011] The RSU is equipped with sensors connected to an edge computing server; vehicles within the RSU's coverage area are divided into two categories: mission vehicles, consisting of a set of... Service vehicles, by collection The task vehicle generates environmental perception tasks to support vehicle applications, while nearby service vehicles utilize their idle computing resources to help execute these tasks; the system's time is divided into multiple discrete time slots. Assuming the vehicular network topology is static within each time slot but changes between time slots, each task vehicle generates periodic environment-aware vehicular machine learning (VML) tasks, comprising two stages: feature extraction and fast inference. Each VML task can be processed locally or partially offloaded to a service vehicle or RSU. The RSU performs feature extraction and fast inference by processing two data streams: data offloaded from the associated task vehicle and perception data from its own assisting vehicles. Once the processing results from the task vehicle and potential service vehicles are uploaded to the RSU, they are fused with the RSU's own results. If the fused output is below a predefined perception accuracy threshold, the RSU activates its full inference module to generate the final output and then sends it back to the task. Otherwise, the fused fast inference result is used as the final perception output.
[0012] Furthermore, the establishment is designed to minimize the impact on all mission vehicles. The optimization problem involves addressing the total latency of completing the perception task within a given timeframe.
[0013]
[0014] These represent the matching strategy between task vehicles and service vehicles in the V2V link, the offloading ratio of different nodes, the transmission rate of the V2V link, and the allocation decision of sub-channels, respectively. This represents the total global latency. , , These represent the percentages of tasks that are offloaded to the local machine, RSU, and neighboring vehicles, respectively. The binary decision variable is used to represent whether there is a match between vehicles. The transmission rate for each packet. This refers to the packet transmission frequency. This represents the inference accuracy after fast inference for the task. , These represent the energy consumption generated during the entire task processing process and the energy consumption generated by the service vehicle's auxiliary task calculations. , These represent the channel allocation decision between the service vehicle and the RSU, and the channel allocation decision between the mission vehicle and the RSU, respectively. This indicates the maximum number of sub-channels between the service vehicle and the RSU.
[0015] in Indicates vehicle Ideal Shannon V2V channel capacity This is the minimum processing accuracy requirement. They represent the first time. During the time slot, in the vehicle The maximum tolerable processing delay for the generation task. and Representing the mission vehicles Service vehicles Energy constraints.
[0016] Furthermore, the modeling of the optimization problem as a Markov Decision Process (MDP) specifically includes:
[0017] State space: ,in These represent the channel gain between the mission vehicle and the RSU, and the channel gain between the service vehicle and the RSU, respectively. This indicates the probability of successful transmission in a V2V link. These represent the native perception data generated by each mission vehicle and the perception data generated by the RSU for collaborating vehicles, respectively.
[0018] Action space: ,in This indicates the vehicle matching strategy for the V2V link. Represented as packet transmission frequency, These represent the sub-channel allocation strategies for the task vehicle and the service vehicle in the V2I link, respectively. This indicates the unloading ratio of each node.
[0019] Reward function: If the constraints of the above formula are not satisfied, a penalty factor is added. Then the reward function is A penalty factor is used to prevent the agent from providing invalid decision variables.
[0020] We employ DQN to optimize resource allocation and V2V pairing strategies for the first... Discrete action space of each time slot and state space The optimal Q value can be calculated using the recursive Bellman equation.
[0021]
[0022] in It is represented as a discount factor. These are represented as the next state value and action value, respectively.
[0023] Resource allocation and V2V pairing decisions obtained by the DQN module are fed into the DDPG as state-action inputs. This facilitates computational unloading decisions. Joint optimization. In DDPG, deterministic actions are generated by the actuator network. ,in This represents the weight parameters of the Actor network, and Indicates attenuation variance Annealing Gaussian exploration noise. The Actor processes environmental conditions. Actions generated by DQN To calculate the unloading ratio In the agent's execution action Then it receives the reward value. Similar to DQN, DDPG also utilizes an empirical replay mechanism to mitigate sample correlation. The Q-value function in DDPG follows the Bellman optimality principle:
[0024]
[0025] Furthermore, the optimization of task offloading and resource allocation based on the Discrete-Driven Continuous Policy Network (DDCPN) algorithm specifically includes:
[0026] The Discrete-Driven Continuous Policy Network (DDCPN) integrates discrete action decision-making with continuous action generation. Specifically, through a cascaded structure, DDCPN tightly couples the output of the Deep Q-Network (DQN), i.e., discrete actions, with the input of the Deep Deterministic Policy Gradient (DDPG), to achieve continuous action generation, thereby promoting more efficient policy optimization. For each vehicle, the global controller at the RSU coordinates the DQN and DDPG to determine the optimal task unloading rate and resource allocation parameters.
[0027] like Figure 3 As shown, the environmental state variables required by the DQN network are first obtained from the environment and then input into the network to obtain partial actions. Then the action Together with the current environment state variables, these are inputs into the DDPG network, which then outputs the remaining actions. By and The values are fed into the environment together to obtain the reward value and the next state value. This is used to correct the weight values of the DDPG and DQN networks. This process is repeated iteratively until a converged network is obtained.
[0028] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements a delay-accuracy balance optimization method based on DDCPN in vehicle-road cooperative infrastructure perception as described in any one of the claims.
[0029] A non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements a delay-accuracy balance optimization method based on DDCPN in vehicle-road cooperative infrastructure perception as described in any one of the claims.
[0030] The advantages and beneficial effects of this invention are as follows:
[0031] This invention proposes a novel vehicle-road cooperative environmental perception framework. Each task vehicle processes its perception data in parallel with the Resource Utility Unit (RSU). Vehicles can process data locally or offload some data to nearby service vehicles or RSUs. The RSU aggregates the processed data and performs full inference when needed, improving the efficiency of environmental perception and decision-making by leveraging the computing power of vehicles and infrastructure. A joint optimization problem combining partial offloading and dynamic resource allocation is proposed to achieve a latency-accuracy balance optimization. The goal is to minimize task execution latency while ensuring fast inference accuracy, and simultaneously satisfying constraints on energy consumption and computing resource usage. This provides important guidance for selecting the optimal offloading target. The optimization problem is modeled as a Markov Decision Process (MDP). Addressing the problem of continuous and discrete mixed action spaces in task offloading and resource allocation decisions, a task offloading and resource allocation optimization algorithm based on Discrete Driven Continuous Policy Network (DDCPN) is proposed, which exhibits good convergence efficiency. Simulation results show that the algorithm has a faster convergence speed, effectively reduces task execution latency, and better adapts to the dynamic and complex vehicular network environment while satisfying constraints such as fast inference accuracy, energy consumption, and computing resource utilization. Attached Figure Description
[0032] Figure 1 This is a schematic diagram of the system model structure of a preferred embodiment of the present invention;
[0033] Figure 2 Task processing flowchart;
[0034] Figure 3 : DDCPN algorithm structure diagram;
[0035] Figure 4 Reward convergence curves for different algorithms;
[0036] Figure 5 Convergence performance curves of the DDCPN-based algorithm under different numbers of vehicles;
[0037] Figure 6 System latency curves for fast inference at different average accuracy thresholds;
[0038] Figure 7 System latency curves under different average computational intensities for different tasks;
[0039] Figure 8 System delay curves at different vehicle speeds;
[0040] Figure 9 System delay curves for different numbers of V2I sub-channels. Detailed Implementation
[0041] The technical solutions of the embodiments of the present invention will be clearly and thoroughly described below with reference to the accompanying drawings. The described embodiments are merely some embodiments of the present invention.
[0042] The technical solution of the present invention to solve the above-mentioned technical problems is:
[0043] The main content of this invention is to propose a joint optimization problem that minimizes task execution latency while ensuring high-speed inference accuracy and efficient resource utilization. By modeling it as a Markov Decision Process (MDP) and employing a Discrete-Continuous Policy Network-based (DDCPN) driven algorithm, our solution exhibits superior performance in balancing accuracy and latency, enabling robust environmental awareness.
[0044] The technical solution of this invention is as follows: the system model structure is as follows Figure 1 As shown, this is a vehicle-to-everything (V2X) system consisting of multiple vehicles and a single RSU. The RSU is equipped with sensors (e.g., cameras and millimeter-wave radar) and connects to an edge computing server. Vehicles within the RSU's coverage area are divided into two categories: mission vehicles, which consist of a collection of... Service vehicles, by collection Task vehicles generate environment-aware tasks to support vehicle applications, while nearby service vehicles utilize their idle computing resources to help perform these tasks.
[0045] The system's time is divided into multiple discrete time slots. Assuming the vehicular network topology is static within each time slot but changes between slots, each task vehicle generates a periodic vehicular machine learning (VML) task, comprising two phases: feature extraction and fast inference. Each VML task can be processed locally or partially offloaded to a service vehicle or RSU. The RSU performs feature extraction and fast inference by processing two data streams: data offloaded from the associated task vehicle and perception data from its own assisting vehicles. Once the processing results from the task vehicle and potential service vehicles are uploaded to the RSU, they are fused with the RSU's own results. If the fused output is below a predefined perception accuracy threshold, the RSU activates its full inference module to generate the final output and sends it back to the task; otherwise, the fused fast inference result is used as the final perception output. The task processing flow is as follows: Figure 2 As shown.
[0046] Specifically, in the first The VML tasks generated in each time slot are composed of a quintuple. It means that, among them Indicates by vehicle In the The size of the raw data generated per time slot This indicates that the vehicle is located at the RSU. The size of the generated auxiliary data, This represents the number of CPU cycles required to process one bit of VML data. It is the maximum allowed latency for processing a one-bit VML task. This represents the minimum fast inference accuracy for the VML task.
[0047] Communication rate calculations can be divided into V2I and V2V links. For V2I (Vehicle-to-Infrastructure) links, the Third Generation Partnership Project (3GPP) introduced a new radio V2X (Vehicle-to-Everything) communication in Release 16. Here, we consider a wireless network supported by the NR-V2X standard. Orthogonal Frequency Division Multiple Access (OFDMA) Uu interfaces are used to allocate bandwidth to the mission vehicle and service vehicle respectively. There are orthogonal sub-channels. This represents the sub-channel assigned to the mission vehicle. This represents the sub-channel assigned to the service vehicle. and These represent the sub-channel allocations for service vehicles and mission vehicles, respectively. Wherein:
[0048]
[0049]
[0050] Therefore, the mission vehicle The transmission rate with RSU can be expressed as:
[0051]
[0052] Service vehicles The transmission rate with RSU can be expressed as:
[0053]
[0054] in , Ensure that each sub-channel is in the first... Each time slot is allocated to only one vehicle for the V2I link. Channel gain between RSU and ,in This represents the distance at the t-th time slot. Vehicle The path loss between the RSU and the RSU. Its calculation expression is... ,in, This represents the carrier frequency in GHz. Similarly, the calculation of service vehicles... Channel gain.
[0055] For a V2V communication link, the packet transmission rate of the mission vehicle through the PC5 interface can be expressed as: Given that wireless transmission operates in half-duplex mode, a transmission error occurs if two vehicles select the same subframe to send their packets. This is due to the half-duplex effect in the first... The probability of packet loss in a time slot can be expressed as: From the vehicle To the vehicle In V2V communication, because the vehicle is in the t-th time slot Location (at a distance) The received signal power (above) drops below the sensing power threshold. The resulting error rate can be expressed as:
[0056]
[0057] vehicle To the vehicle The probability of a successful transmission can be: Therefore, the data transfer rate between them is: ,in Size of each data packet.
[0058] Latency calculation can be divided into local calculation, service vehicle calculation, and RSU calculation. First, let's look at the local latency calculation. The local task latency calculation expression is:
[0059]
[0060] The energy consumption generated by task processing is:
[0061]
[0062] in It depends on the effective energy coefficient of the chip architecture.
[0063] The upload latency for the locally calculated results is:
[0064]
[0065] in Indicates the size of the processing result.
[0066] The energy consumption of uploading data is:
[0067]
[0068] The total latency consumption of the local computing method is:
[0069]
[0070] If a portion of the tasks needs to be offloaded to one of the adjacent service vehicles, the offload ratio is denoted as... Introduce a binary variable. Indicates whether the vehicle has been selected as a service vehicle. Vehicle Unload raw sensing data to the vehicle The transmission delay is expressed as:
[0071]
[0072] vehicle To the vehicle The energy consumption for transmitting data is expressed as:
[0073]
[0074] Service vehicles The processing delay at that point is represented as:
[0075]
[0076] Calculated energy consumption at the service vehicle location Represented as:
[0077]
[0078] Subsequently, from the vehicle The upload latency to the RSU processing result is given by the following formula:
[0079] .
[0080] The energy consumed in transmission is:
[0081]
[0082] Therefore, if the vehicle Uninstall in Each time slot generates part of its task to serve vehicles. The total latency (including communication and computation latency) used to complete this part of the task is expressed as:
[0083]
[0084] In some cases, to leverage the powerful computing capabilities of the RSU, improve computational accuracy, and reduce latency, a portion of the task is offloaded to the RSU. Vehicle offloading rate. Represented as ,in Therefore, the upload delay for this part of the data is expressed as:
[0085]
[0086] The energy consumption for transmitting this part of the data is expressed as:
[0087]
[0088] in, Indicates the vehicles that upload data to the RSU. The transmit power. We assume that the RSU itself will generate power related to the vehicle. The relevant sensor data, its size is Therefore, the RSU needs to process data to quickly infer vehicle information. The total amount of data is The computational delay can be expressed as:
[0089]
[0090] Therefore, if the vehicle If some of its tasks are offloaded to RSU, the total latency for completing these tasks is expressed as:
[0091]
[0092] Since RSUs are typically connected to a stable power source via cables, their energy consumption is not within the scope of this study; this study primarily focuses on vehicle energy consumption. Indicates vehicle The achievable inference accuracy, Indicates service vehicles The achievable inference accuracy, and This indicates the achievable inference accuracy of the RSU's fast inference module. (Based on the vehicle...) In the The fast inference accuracy of a task generated in a time slot is given by the following formula:
[0093]
[0094] The mission vehicle and auxiliary service vehicles send the processed results to the RSU for result fusion and optional full inference, and then send the final perception results back to the mission vehicle. Let... For binary parameters, when the fast inference accuracy is lower than the required perception accuracy... hour ,otherwise Because the amount of data processed and transmitted is relatively small, the time spent on final result fusion is ignored and broadcast, and the vehicles... In the The total processing latency for tasks generated in each time slot is:
[0095]
[0096] in, This indicates the size of the processing result output by the fast inference module at RSU.
[0097] vehicle and its paired service vehicles In the In each time slot, targeting vehicles The total energy consumption of the generated task is given by the following formula:
[0098]
[0099] We have established an optimization problem to minimize the time required for all mission vehicles to complete the task. Total delay in completing the perception task over the duration:
[0100]
[0101] in Indicates vehicle Ideal Shannon V2V channel capacity This is the minimum processing accuracy requirement. They represent the first time. During the time slot, in the vehicle The maximum tolerable processing latency for the generated task. Furthermore, and Representing the mission vehicles Service vehicles Energy constraints.
[0102] The inherent complexity of the above target problem stems from the coexistence of continuous variables (unloading ratio). ) and discrete variables (V2 V pairwise decision) Sub-channel allocation decision and V2V transmission rate This problem is exacerbated by their close interdependence. To address this issue, we leverage Deep Reinforcement Learning (DRL), which natively supports hybrid discrete-continuous optimization. For example... Figure 3 As shown, we employ a Discrete-Driven Continuous Policy Network (DDCPN), which integrates discrete action decision-making with continuous action generation. Specifically, through a cascaded structure, DDCPN tightly couples the output (discrete actions) of a Deep Q-Network (DQN) with the input of a Deep Deterministic Policy Gradient (DDPG) to achieve continuous action generation, thereby facilitating more efficient policy optimization. For each vehicle, the global controller at the RSU coordinates the DQN and DDPG to determine the optimal task unloading rate and resource allocation parameters.
[0103] We first implemented our proposed DDCPN-based partial offloading and resource allocation optimization algorithm (considering the length of the full name, we...). Figure 4 We refer to this as "Proposed" and evaluate its convergence speed. We evaluate and compare its performance with baseline algorithms in various scenarios. Our proposed algorithm considers serving vehicles and RSUs for comparison with the following three schemes:
[0104] Option 1: Collaborative Awareness and Binary Offloading: In this approach, tasks from the vehicle are processed in a binary manner, with two options: local processing or offloading to the RSU. The processing results are then aggregated at the RSU for dataset aggregation.
[0105] Option 2: Collaborative Awareness with Partial Offloading to RSU: Tasks generated by the task vehicle are either processed locally or partially offloaded to the RSU. The processing results from the task vehicle are transmitted to the RSU for multi-source data fusion.
[0106] Option 3: Collaborative Perception with Partial Offloading to Service Vehicles: Tasks generated by task vehicles are processed locally or partially offloaded to service vehicles. The service vehicles then transmit the processed data to the task vehicles for information fusion.
[0107] The system simulation parameters are set as follows: communication bandwidth In the V2I link, the signal transmission power of the mission vehicle and the service vehicle to the RSU is respectively... , Signal transmission power between the mission vehicle and the service vehicle in the V2V link Energy consumption threshold of the mission vehicle. The vehicle's speed range is The CPU cycle count of RSU is carrier frequency The calculation frequencies of mission vehicles and service vehicles both obey... The uniform distribution. The precision calculation of the mission vehicle. The accuracy of service vehicle calculations Precision of calculation at RSU CPU revolutions required per bit task .
[0108] Figure 4 The convergence performance of three algorithms (DDCPN, DDPG, and DQN) at 5000 iterations is shown. The cumulative reward of all algorithms increases with the number of iterations. DQN converges earliest (around 200 iterations) but ultimately achieves the lowest reward, indicating it is a suboptimal strategy. This stems from the discretization of the action space in DQN, which generates a large number of state-action pairs. This complexity hinders accurate Q-value updates, causing it to converge to a suboptimal solution. In contrast, DDPG achieves a higher final reward due to its actor-critic architecture: the actor network generates continuous actions, and the critic network evaluates their quality. This allows DDPG to efficiently handle continuous action spaces, especially in high-dimensional environments. However, its convergence is slower, requiring nearly 1800 iterations. This delay is because DDPG must jointly optimize continuous and discrete action selection strategies, leading to more complex training dynamics. The proposed DDCPN algorithm outperforms both by integrating DQN and DDPG into a cascaded structure, where the output of DQN guides the action selection of DDPG, thus reducing the search space. This hierarchical approach simplifies the exploration process and enables faster and more stable convergence.
[0109] Figure 5The convergence of our DDCPN-based optimization algorithm is demonstrated under different numbers of vehicles in the task (e.g., the marginal difference between 8 and 12 vehicles is 0.002), but the average reward converges within 500 rounds across different cases. This robustness stems from the algorithm's hybrid action space decomposition, which independently optimizes continuous and discrete actions, maintaining stable convergence despite expanding the action space at higher vehicle intensities.
[0110] Figure 6 The system latency of different schemes was compared at different fast inference accuracy thresholds. All schemes showed increased latency with stricter accuracy requirements, but key differences emerged. Specifically, when the threshold exceeded 0.82, the scheme partially offloading to the service vehicle suffered a sharp latency spike. The limited computational power of the service vehicle (relative to the RSU) forced extended processing times to meet the accuracy threshold. The binary offloading scheme processed tasks locally or by fully offloading to the RSU. It maintained a consistently high latency level across different accuracy thresholds. This occurred because full local computation resulted in high processing latency, while full offloading to external nodes resulted in significant data transfer latency. In contrast, at relatively relaxed accuracy thresholds, the scheme partially offloading to the RSU and the proposed method maintained a similar trend toward better performance, both dynamically optimizing the offloading ratio for different tasks. However, when the accuracy threshold exceeded 0.86, both schemes experienced latency spikes. This is because the task vehicle must prioritize offloading a larger proportion of its tasks to the RSU equipped with higher-level computational resources to meet the higher accuracy constraints of fast inference, inevitably introducing longer uplink latency. Compared to schemes that partially offload to the RSU, our proposed method shows a slight advantage because it integrates service vehicles as supplementary offloading targets to mitigate the longer selection time of service vehicles as the accuracy threshold increases further (e.g., 0.88), since they cannot process tasks with such high precision. This fact forces task vehicles to perform almost complete offloading to the RSU at the cost of increased transmission latency, resulting in the latency performance of the proposed scheme almost coinciding with the latency performance convergence curve of the binary offloading baseline.
[0111] We define the computational intensity of a task as the number of CPU cycles (cycles per bit) required to process one bit of task data. Different tasks have different computational intensities. Figure 7The system latency of four schemes was compared under different average computational intensities across all tasks. While all schemes demonstrated that latency increased with increasing intensity, the scheme partially offloading to the service vehicle outperformed the binary offloading scheme below 1500 cycles / bit, but suffered from a rapid performance degradation due to the service vehicle processor (whose computational power is much lower than that of the RSU) after exceeding this threshold, making it difficult to handle computationally intensive tasks. In contrast, the binary offloading scheme maintained lower latency (125 ms at 3000 cycles / bit) at high computational intensities by fully utilizing the upper-level computational resources of the RSU. The scheme partially offloading to the RSU performed similarly to the proposed method below 1000 cycles / bit; this performance degradation stemmed from the system's inherent dependence on the RSU, which is more powerful in handling tasks with higher computational intensities, unfortunately triggering escalating communication latency, especially when the amount of offloaded data exceeded the channel capacity. Our method overcomes this limitation by adaptively distributing the workload among tasks. Optimizations were made for the vehicle, service vehicle, and RSU based on computational intensity and real-time network conditions. Compared with the scheme partially offloading to the RSU ( Figure 7 Compared to the best-performing baseline, our method achieves a 26% reduction in peak latency. This improvement is achieved by dynamically assigning portions of the task to service vehicles, which alleviates the coupled compute-communication latency bottleneck.
[0112] Figure 8The latency comparison of four schemes is presented at different vehicle speeds. The scheme of partially offloading to the service vehicle maintains stable latency because it focuses on the speed of the service vehicle relative to each mission vehicle, rather than the speed of an individual vehicle, to achieve efficient offloading. The binary offloading scheme exhibits an initial increase in latency, driven by offloading to the RSU under worse V2I link conditions caused by higher vehicle speeds. When the latency exceeds a specified threshold, further increases are suppressed because mission vehicles prefer to keep computation local rather than offloading to the RSU, thus mitigating the burden of transmission latency. At this stage, only the computational resources equipped on each mission vehicle affect the latency. Unlike the binary offloading scheme, the scheme of partially offloading to the RSU still relies on the RSU, making the latency more sensitive to speed changes and exhibiting an upward trend. When the vehicle speed increases to 65-80 km / h, uplink instability caused by high vehicle maneuverability makes offloading to the RSU non-competitive, prompting the mission vehicle to switch to fully local processing. This shift to local processing causes the latency performance of the partially offloading to the RSU scheme to converge with that of the binary offloading scheme. The proposed scheme maintains minimal system latency at different aircraft speeds through two key mechanisms: (a) different unloading targets and adaptive unloading ratios between different entities effectively mitigate the system latency degradation caused by velocity-dependent channel impairments; (b) the superior convergence efficiency of our hybrid reinforcement learning framework enables rapid optimization of unloading decisions under dynamic conditions (vehicle speed: 65-80 km / h, no unloading to RSU), with latency consistent with the scheme of partially unloading to service vehicles.
[0113] Figure 9 The system latency of four schemes is demonstrated as the number of V2I sub-channels changes. The scheme that partially offloads to serving vehicles is unaffected because it does not perform uplink transmissions to the RSU. The other three schemes show that the overall system latency decreases as the number of V2I sub-channels increases, thanks to the increased communication resources. Notably, when the number of V2I sub-channels reaches 14, the latency of the binary offloading scheme drops sharply—this is because before reaching this threshold, due to the high uplink latency, this scheme defaults to a fully local processing mode. The overall decreasing trend of the three schemes also exhibits diminishing marginal returns: even with more sub-channels, performance improvements are limited by the computational capabilities of the roadside units (RSUs). Nevertheless, our proposed scheme performs best by flexibly and fully utilizing the communication and computational resources of surrounding vehicles and infrastructure, demonstrating its potential in addressing the dynamics and complexity of future intelligent transportation systems.
[0114] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions.
[0115] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0116] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0117] The above embodiments should be understood as illustrative only and not as limiting the scope of protection of the present invention. After reading the description of the present invention, those skilled in the art can make various alterations or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.
Claims
1. A multi-RSU vehicle-road cooperative task offloading optimization method based on the MCW-TD3 algorithm, characterized in that, Includes the following steps: Establish a single-slot, multi-user, multi-objective optimization problem, where the optimization problem is a utility function of latency and cost; The optimization problem is established as a Markov decision process, and the improved dual-delay deep deterministic policy gradient algorithm TD3 is used to solve the problem. The improved dual-delay deep deterministic policy gradient algorithm TD3 introduces the "dual Q network" strategy, that is, using two independent Q networks Q1 and Q2, and selecting the smaller value of these two networks when calculating the target Q value. Then, a weighted minimum value strategy is adopted, and a weighting mechanism is proposed to allow the outputs of multiple Q networks to be weighted and optimized according to the solution.
2. The multi-RSU vehicle-road cooperative task offloading optimization method based on the MCW-TD3 algorithm according to claim 1, characterized in that, The establishment of a single-slot multi-user multi-objective optimization problem specifically includes: in, They represent At any moment Signal transmission power, RSU At any moment Allocated computing resources, RSU Sub-channel allocation decision factor For vehicles The proportion of data unloaded locally; , All are weighting factors. ; Global system latency handling for all VUs The maximum processing delay threshold is set at 100%. This represents the total cost of all VUs. The maximum cost threshold, This is the maximum transmit power threshold for the vehicle. The set of RSUs, i.e. , Indicates the first One RSU, Represented as time At that time, the threshold for calculating the sum of frequencies allocated to vehicles under the current RSU. Each RSU has a set of vehicles, represented by a set. ,in , , For vehicles The proportion of data unloaded to RSU Represented as the threshold for the number of vehicles allocated to the same sub-channel, RSU The sub-channel set is represented as follows: , , These represent vehicle upload latency and vehicle upload latency threshold, respectively. , Indicates in Tasks generated in real time, RSU VUs The total processing latency and latency threshold, , These represent the processing latency of the data at the RSU and the processing latency threshold of the data itself, respectively. , , These represent the energy consumption for local vehicle processing tasks, the energy consumption for uploading data, and the threshold for vehicle energy consumption, respectively. , In respectively During the time slice, the vehicle The data processing accuracy and the data processing accuracy threshold.
3. The multi-RSU vehicle-road cooperative task offloading optimization method based on the MCW-TD3 algorithm according to claim 1, characterized in that, The optimization problem is established as a Markov decision process, and the improved dual-delay deep deterministic policy gradient algorithm TD3 is used to solve the problem. Specifically, the process includes: obtaining an initial state based on the initial environment, which is then used as the input to the MCW-TD3 algorithm neural network. The neural network outputs a reasonable action based on the input, and then inputs this action back into the environment to obtain the state value and reward value for the next time step. If the output action does not meet the constraints, a penalty is added to reduce the reward value, and the new state value is then input into the neural network to obtain the action value for the next time step. This mechanism is repeated continuously to update the network weights. The process stops when the loss function tends to stabilize and the number of iterations reaches a certain number, resulting in a better set of neural network weights. The MCW algorithm uses a weighted minimum method to estimate the Q-value, combining maximum and minimum value operations to limit the estimated range of the Q-value and reduce overestimation and underestimation biases. The target Q-value network is: It is a noisy target policy action. and It is the target Critic network. and There are two weight values that satisfy... ,and , > 0; Critic Network loss function It is its predicted value With target value Mean square error: Critic Network loss function It is its predicted value With target value Mean square error: 。 4. An electronic device, characterized in that, The system includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the multi-RSU vehicle-road cooperative task offloading optimization method based on the MCW-TD3 algorithm as described in any one of claims 1 to 3.
5. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the multi-RSU vehicle-road cooperative task offloading optimization method based on the MCW-TD3 algorithm as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Link topology adaptive unloading method for edge calculation of Internet of Vehicles
CN119136257A