V2V task unloading and collaborative reasoning method based on model segmentation

By employing a vehicle trajectory prediction method based on spatiotemporal feature fusion and multi-agent deep reinforcement learning, the complexity of global trajectory prediction and V2V task offloading in the Internet of Vehicles is solved, achieving efficient task offloading and collaborative reasoning, thereby improving the system's service quality and computational efficiency.

CN121056838APending Publication Date: 2025-12-02NANJING TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511148373.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-17
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

In the Internet of Vehicles (IoV), existing technologies struggle to effectively utilize global trajectory prediction to improve V2V task offloading performance, and fail to adequately consider collaborative reasoning between vehicles, leading to complexity and latency issues in task offloading.

Method used

A vehicle trajectory prediction method based on spatiotemporal feature fusion and multi-agent deep reinforcement learning are adopted, combined with model segmentation technology, to achieve global vehicle trajectory prediction and task scheduling optimization. Task offloading and collaborative inference are performed through the cooperation of MEC controller and base station.

Benefits of technology

It significantly reduces task processing latency and incompleteness, and improves the system's service quality and computing efficiency, especially in high-speed and dynamic traffic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121056838A_ABST
    Figure CN121056838A_ABST
Patent Text Reader

Abstract

In order to relieve the high load of edge computing equipment, the invention provides a V2V task unloading and collaborative reasoning method based on model segmentation, and in the V2V task unloading and collaborative reasoning-oriented Internet of Vehicles, firstly, a vehicle trajectory prediction model based on spatial-temporal feature fusion is adopted, and interaction features between vehicles are captured, so that the V2V task unloading and collaborative reasoning is realized. Accurate global trajectory prediction information is provided for a task unloading decision; and then, multi-agent deep reinforcement learning is adopted, and an optimal global task scheduling strategy is output through joint rewards, so that the classification precision and the reasoning delay of the vehicle-mounted model are doubly optimized. A simulation result shows that the method can effectively reduce the service delay and the unfinished rate of the visual task in a dynamic traffic environment, so that the service quality of the system is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the application of artificial intelligence technology in the Internet of Vehicles, specifically a V2V (vehicle-to-vehicle) task offloading and collaborative reasoning method based on model segmentation. Background Technology

[0002] Vehicle-to-vehicle (V2V) communication is achieved through Dedicated Short-Range Communications (DSRC). [1] Or Cellular Vehicle-to-Everything (C-V2X) [2] Wireless communication enables low-latency and highly reliable data transmission even at good line-of-sight (LOS) distances. V2V communication does not rely entirely on fixed infrastructure such as roadside units or base stations, and can maintain efficient operation even in remote road sections without 4G or 5G network coverage. [3] Vehicle-to-vehicle (V2V) communication can extend the perception range by sharing sensor data and offload computationally intensive tasks to nearby vehicles to alleviate the burden of mobile edge computing (MEC). [4] The server faces resource consumption pressure. Compared to the high latency of vehicle-to-infrastructure (V2I) task offloading, numerous studies have verified that complete V2V offloading can reduce task processing latency. [5] However, this offloading method increases the computational burden on neighboring service vehicles. Model segmentation offers a feasible solution. Customer and service vehicles deploy segmented shallow and deep networks respectively, and collaboratively compute the task results. Constructing a globally optimal task offloading and collaborative reasoning framework for segmented shallow and deep networks presents some new challenges:

[0003] (1) V2V task offloading for global trajectory prediction perception.

[0004] The high-speed movement of vehicles causes continuous reconfiguration of the vehicle-to-everything (V2X) topology, making the establishment of stable and reliable communication links more complex. Furthermore, sudden changes in vehicle speed, direction, and density can cause intermittent network interruptions, affecting the integrity and latency of offloaded data transmission.

[0005] To this end, reference [6] designed a vehicle trajectory prediction method based on temporal convolutional networks and used reinforcement learning that relies on trajectory prediction to schedule tasks, thereby significantly reducing task completion delay. Reference [7] proposed a task offloading scheme that combines trajectory prediction. This scheme combines long short-term memory and convolutional neural networks to predict the base stations and arrival times that vehicles will pass through in the future, and offloads tasks accordingly to support high-intensity demands. Men et al. [8] By capturing elements that can improve prediction accuracy through a perceptual gated recurrent unit network, and designing a near-end policy optimization algorithm to obtain the optimal one-to-many offloading decision, the task execution cost is reduced.

[0006] The aforementioned literature uses local trajectory prediction to assist V2I task offloading, but does not utilize global trajectory prediction to improve local task offloading performance. Furthermore, the continuous changes in V2V network topology increase the complexity of task offloading. Therefore, how to construct a linkage paradigm between global trajectory prediction and local V2V task offloading based on existing V2I task offloading methods is an unsolved problem.

[0007] (2) V2V task offloading and collaborative reasoning based on model segmentation.

[0008] Vehicles process tasks and make decisions in real time by interacting with surrounding infrastructure. Most in-vehicle applications are computationally intensive and latency-sensitive, so offloading all tasks to infrastructure can cause long response times.

[0009] To improve processing efficiency, researchers have considered task offloading schemes based on model segmentation. Reference [9] designs an adaptive model splitting algorithm and splits the task into independent working blocks to minimize inference latency. Reference

[10] proposes a collaborative optimization method for model segmentation and task offloading, using a greedy algorithm to obtain offloading decisions to complete collaborative inference. Reference

[11] designs an asynchronous advantage Actor-Critic method based on multi-task learning to output a model segmentation strategy that reduces inference latency. This method reduces overall, marginal, and local inference latency by 4.76%, 10.04%, and 8.03%, respectively.

[0010] Existing model-segmentation-based task offloading methods mostly utilize the computing power of edge devices to complete collaborative reasoning, rarely considering task offloading and collaborative reasoning between vehicles. Therefore, how to jointly design model segmentation and task offloading in a V2V model is a problem worth exploring.

[0011] (3) Joint optimization of trajectory prediction and task unloading.

[0012] There is a high degree of coupling between vehicle trajectory prediction and task unloading decisions. Accurate trajectory prediction can effectively improve the reliability and efficiency of task unloading, while accurate task unloading strategies can provide beneficial feedback to trajectory prediction.

[0013] Researchers have proposed many effective task offloading schemes by taking advantage of the strong coupling between the two. Reference

[12] proposes a multi-agent deep reinforcement learning method based on mobile perception collaboration, which uses the multi-step trajectory prediction module based on Informer to optimize the task offloading strategy of the multi-agent. Reference

[13] proposes a task offloading algorithm called LSTM-MADDPG, which models the task relationship into a directed acyclic graph so that the multi-agent can formulate a more refined offloading strategy according to the task relationship, thereby reducing the global processing cost. Reference

[14] proposes a delay optimization scheduling strategy based on vehicle trajectory prediction, which backs up the data required for task execution to the candidate server according to the trajectory predicted by Bi-LSTM, thereby reducing the impact of prediction deviation on task offloading.

[0014] Although existing methods design scheduling strategies based on the coupling between trajectory prediction and task offloading, most of them simply combine pre-trained trajectory prediction and task offloading models, rarely performing joint training and optimization, thus limiting the improvement of task scheduling performance. In particular, the bidirectional feedback between trajectory prediction and task offloading in V2V mode is more valuable for selecting the most suitable service vehicle.

[0015] Therefore, how to construct a two-way closed loop of trajectory prediction and task unloading is an unsolved problem. Summary of the Invention

[0016] To address the aforementioned problems, this invention proposes a V2V task offloading and collaborative reasoning method based on model segmentation, aiming to accelerate task processing speed and improve Quality of Service (QoS). Specifically, this invention includes:

[0017] A model-segmentation-based V2V task offloading and collaborative reasoning method for vehicle-to-everything (V2V) networks:

[0018] (I) First, a vehicle trajectory prediction method based on spatiotemporal feature fusion is adopted to predict the global vehicle trajectory;

[0019] (ii) Then, multi-agent deep reinforcement learning based on model segmentation is used to obtain the future position of the vehicle from the global vehicle trajectory and obtain the optimal global task scheduling strategy.

[0020] The proposed vehicle-to-everything (V2V) architecture for V2V task offloading and collaborative reasoning consists of an MEC controller, base stations, and vehicles. Let the sets of base stations and vehicles be... and

[0021] The base station schedules tasks based on the predicted vehicle trajectory, computing power, and queue load, and sends scheduling information; the MEC controller connects to the base station and undertakes the task of predicting vehicle trajectories.

[0022] Each vehicle is equipped with a vehicle model for handling visual tasks. The vehicle model takes external environment image information collected by the vehicle's vision sensors as input and obtains inference results. The network structure of the vehicle model, i.e., the vehicle network, is a deep neural network containing multiple separable modules. The vehicle network is divided into a multi-exit network. At a certain exit, the vehicle network is divided into a shallow network and a deep network to obtain a collaborative network for collaborative inference. The vehicle that initiates the collaboration request is the customer vehicle, and the vehicle that provides the collaboration service is the service vehicle. The roles of the customer vehicle and the service vehicle are dynamically changing. In the collaboration state, the customer vehicle offloads the task to the service vehicle, which performs the corresponding inference work. After completing the inference work, the service vehicle transmits the inference results back to the customer vehicle.

[0023] In step (1), a trajectory prediction model based on spatiotemporal feature fusion is used for trajectory prediction. In the trajectory prediction model, firstly, the recurrent neural network (RNN) trajectory encoder extracts the dynamic motion features of the vehicle through historical trajectory data. Then, the spatiotemporal interactive encoder captures and fuses the spatiotemporal features of the vehicle trajectory through parallel temporal and spatial graph attention. Finally, the LSTM predictive decoder accurately predicts the vehicle's driving trajectory based on the fused spatiotemporal features.

[0024] In step (ii), each base station is abstracted as an agent; the MEC controller uses multi-agent deep reinforcement learning (MADRL) to solve the multi-agent Markov decision process (MMDP) problem based on the global trajectory prediction results.

[0025] The main contributions of this invention are as follows:

[0026] First, to address problem (1), this paper proposes a vehicle trajectory prediction method based on spatiotemporal feature fusion. This method uses a graph neural network (GNN) and an attention mechanism to capture the dynamic interaction features between vehicles, and uses this to accurately predict the global vehicle trajectory. This trajectory information provides a basis for making decisions on task offloading and collaborative reasoning.

[0027] Second, in response to problem (2), this paper proposes a multi-agent deep reinforcement learning based on model segmentation, which improves the global task processing performance by jointly optimizing the V2V task offloading and collaborative reasoning strategies of multiple agents.

[0028] Third, in response to problem (3), this paper designs a joint training framework for trajectory prediction and task unloading to optimize both trajectory prediction accuracy and collaborative reasoning efficiency.

[0029] Simulation results show that the proposed scheme outperforms existing benchmark methods in terms of trajectory prediction and service latency. Attached Figure Description

[0030] Figure 1 This represents a vehicle-to-everything (V2V) architecture oriented towards V2V task offloading and collaborative reasoning.

[0031] Figure 2 This represents an example of vehicle trajectory prediction.

[0032] Figure 3 This represents vehicle trajectory prediction based on spatiotemporal feature fusion.

[0033] Figure 4 This represents a task scheduling framework based on multi-agent reinforcement learning.

[0034] Figures 5(a) and 5(b) illustrate the convergence of task scheduling, where...

[0035] Figure 5(a) shows the convergence curve of average service latency.

[0036] Figure 5(b) shows the probability density distribution of average service delay.

[0037] Figures 6(a) and 6(b) respectively illustrate the impact of the number of tasks on system service performance.

[0038] Figure 6(a) shows the average task service latency for different numbers of tasks.

[0039] Figure 6(b) shows the average classification accuracy for different numbers of tasks.

[0040] Figures 7(a) and 7(b) show the impact of vehicle speed on system service performance, respectively.

[0041] Figure 7(a) shows the average task service latency at different speeds.

[0042] Figure 7(b) shows the average task completion rate at different speeds.

[0043] Figures 8(a) and 8(b) respectively visualize the task unloading strategies under two completion delay constraints.

[0044] Figure 8(a) shows low values ​​indicating delay constraints.

[0045] Figure 8(b) illustrates the high-delay constraint. Detailed Implementation

[0046] 1 Overview

[0047] To alleviate the high load on edge computing devices, this paper proposes a V2V task offloading and collaborative inference method based on model segmentation. First, a vehicle trajectory prediction model based on spatiotemporal feature fusion captures the interaction features between vehicles to provide accurate global trajectory prediction information for task offloading decisions. Then, multi-agent deep reinforcement learning uses joint rewards to output the optimal global task scheduling strategy, thus doubly optimizing the classification accuracy and inference latency of the onboard model.

[0048] Simulation results show that the proposed method can effectively reduce the service latency and incomplete rate of visual tasks in dynamic traffic environments, thereby significantly improving the service quality of the system.

[0049] 2 System Model

[0050] Figure 1 This paper presents a software-defined vehicular cooperative network (SVCN) architecture for V2V task offloading and cooperative inference, consisting of an MEC controller, base stations, and vehicles. The set of base stations and vehicles is as follows: and Suppose the in-vehicle network is a deep neural network containing L modules. Given a set of exits... The vehicular network is segmented into a multi-egress network. (Vehicle) In export The in-vehicle network is segmented into shallow networks. and deep networks φ i,l A collaborative network is generated for collaborative reasoning. Vehicles requesting collaboration are labeled as Client Vehicles (CVs); otherwise, they are labeled as Server Vehicles (SVs). Vehicle roles change dynamically. The base station schedules tasks based on the vehicle's predicted trajectory, computing power, and queue load, and uses the Beacon protocol.

[15] Sending scheduling information. The MEC controller connects to the base station via a wired network and is responsible for vehicle trajectory prediction.

[0051] 2.1 Vehicle Trajectory Prediction Model

[0052] like Figure 2 As shown, the position coordinates of vehicle i at time t are Then it is in the time window The historical trajectory is

[0053]

[0054] in Indicates the length of the backtracking horizon. The MEC controller will... The input trajectory prediction model produces a predicted trajectory.

[0055]

[0056] This trajectory can provide the base station with the vehicle's future location to optimize unloading decisions. For example, Figure 1 Vehicle a can offload tasks to vehicle b that it will encounter in the future, not just vehicle c within the current communication range. Vehicles d and e are close to each other but traveling in opposite directions. Relative motion can cause connection instability or even interruption, so vehicle d cannot offload tasks to vehicle e and must process them locally.

[0057] 2.2 Communication model

[0058] High-speed vehicle movement causes dynamic changes in the V2V network topology. When the V2V link is interrupted, customer vehicles may not be able to directly obtain the inference results from the serving vehicle. This result needs to be relayed to the customer vehicle via the base station. The following details the two communication models: V2V and V2I2V.

[0059] 1) V2V communication: Based on the IEEE 802.11p protocol

[16] , V2V communication adopts a multi-hop self-organizing network architecture, and its channel follows independent and identically distributed Rayleigh fading

[17] . Let the set of vehicles serving customer vehicle i be . Customer vehicles and service vehicles The path loss function between them is κ i,j =63.3 + 17.7log 10 (d i,j )

[18] , where d i,j This represents the distance between vehicle i and vehicle j. V2V communication is affected by path loss, Rayleigh fading, and neighbor interference, so the channel signal-to-noise ratio can be calculated as...

[0060]

[0061] Where p i and p j Both represent transmission power, g i,j Indicates Rayleigh's loss, This represents Gaussian white noise. According to Shannon's formula, the V2V data transmission rate is...

[0062] μ i,j =B V2V log2(1+SINR i,j (4)

[0063] Among them B V2V This indicates the bandwidth of V2V communication.

[0064] 2) V2I2V Communication: Due to deviation from the driving direction or sudden acceleration, customer vehicle i may leave the communication range of service vehicle j before receiving the final inference result. In this case, the inference result needs to be forwarded to customer vehicle i via base station m. The signal-to-noise ratio of the channel between service vehicle j and base station m is...

[0065]

[0066] Where d j,m (d j,m′ ) represents the distance between vehicle j and base station m (m′), α represents the path loss exponent, and g j,m Indicates Rayleigh's loss, This represents Gaussian white noise. According to Shannon's formula, the uplink V2I data transfer rate is...

[0067] μ j,m =B V2I log2(1+SINR j,m (6)

[0068] Among them B V2I This represents the V2I communication bandwidth. Next, base station m sends the received inference results to customer vehicle i. Similar to equations (5) and (6), the channel signal-to-noise ratio between base station m and customer vehicle i is...

[0069]

[0070] Where d i,m g represents the distance between customer vehicle i and base station m. i,m This represents Rayleigh loss. Therefore, the downlink I2V data transfer rate is...

[0071] μ m,i =B I2V log2(1+SINR m,i (8)

[0072] Among them B I2V This indicates the I2V communication bandwidth.

[0073] 2.3 Task completion delay

[0074] The system divides the time domain into a series of time windows of equal length. Each window contains a set of scheduling slots of equal length. The duration of the time window is f (w) Let the sets of computing units configured for customer vehicle i and service vehicle j be respectively... and And the base station's sub-channel set is Both vehicles and base stations follow a First-Come, First-Served (FCFS) model.

[19] The principle of handling visual tasks. When the window... Initially, the MEC controller jointly fine-tunes the trajectory prediction model and task offloading model based on the vehicle's historical driving trajectories in window w-1 to adapt to continuous changes in the traffic environment. In the scheduling slot... The base station uses a Multi-Agent Deep Reinforcement Learning (MADRL) algorithm to generate optimal V2V task offloading and collaborative inference strategies based on task requirements, in order to efficiently achieve distributed vision task processing. The vision task h of customer vehicle i is defined by the triple {x...} h ,τ h ,ε h} constitutes, where x h ,τ h ,ε h These represent the computational cost required for perceiving the image, the computational cost for reasoning, and the completion delay constraint, respectively.

[0075] The following details the task completion delays, including delays in task unloading, processing, and result return.

[0076] 1) Offloading Delay: Offloading delay is the latency required for the customer vehicle to perform shallow network inference tasks and transmit intermediate features to the service vehicle. The completion latency of in-vehicle vision tasks is often limited to milliseconds, so vehicles must initiate and process tasks quickly. To simplify problem modeling, this paper does not consider queuing delay when calculating offloading delay. Based on the discussion of the communication model above, offloading delay is calculated using both V2V and V2I2V methods.

[0077] Under the scheduling slot t of window w, the set of tasks initiated by customer vehicle i is as follows: Its base is Therefore, the set of customer vehicle tasks received by base station m is Its base is Suppose that the number of computing units equipped in customer vehicle i is O. i And the maximum GPU cycle time for each computing unit is Set shallow network The extracted intermediate features are Its data size and the computational cost required for inference are respectively and Define a binary variable to represent partial unloading (1) or local processing (0).

[0078]

[0079] Task h process The V2V offloading latency required for inference and result transmission to service vehicle j is

[0080]

[0081] The two terms in equation (10) represent computation and transmission delays, respectively. Additionally, the customer vehicle offloads the task to the service vehicle via a V2I2V mechanism, delivering the final inference result when the two vehicles approach each other. Let the binary variable b( m 1,) i,h,l =1 and b( m 1,) i,h,l =0 represents unloading of V2V and V2I2V respectively, then equation (10) can be changed to

[0082]

[0083] 2) Processing latency: Processing latency is the time required for the service vehicle j to perform deep network inference on intermediate features. Let the number of computing units equipped in service vehicle j be... And the maximum GPU cycle time for each computing unit is Let the deep network φ j,l Continue reasoning about intermediate features The required computation is The expression for handling delays is:

[0084]

[0085] 3) Return Delay: Return delay is the time required for the final inference result to be transmitted to the customer vehicle via V2V or V2I2V communication. Let the final inference result of task h be... Its data size is Let there be two variables. If the service vehicle j transmits the visual processing results via V2V (V2I2V) communication, then the return delay can be expressed as:

[0086]

[0087] 3. Problem Modeling

[0088] Customer vehicles utilize the computing resources of service vehicles to alleviate the computational pressure caused by high-intensity tasks, thereby improving real-time processing efficiency. A challenge of model-based task offloading is jointly optimizing trajectory prediction and task scheduling to achieve global load balancing.

[0089] Under the scheduling slot t of window w, the task completion delay is the sum of equations (11)(12)(13), that is...

[0090]

[0091] Furthermore, the system's average service latency can be expressed as:

[0092]

[0093] in Represents a set of vehicles The cardinality. To evaluate the effectiveness of the task scheduling strategy, a binary variable is defined.

[0094]

[0095] Where ε h This indicates the time constraint for task h. These represent the processing results of task h in ε. h Successful and unsuccessful deliveries are considered. Let the scheduling strategy for base station m in scheduling slot t be: The set of scheduling strategies within the time window w is: Corresponding to these scheduling strategies, the allocation strategies for vehicle computing resources and base station spectrum resources are as follows: and in Record the incomplete rate of tasks within window w.

[0096]

[0097] According to equation (17), the system's task non-completion rate can be expressed as:

[0098]

[0099] Given a set of time windows The optimization objective of the system's V2V offloading is to minimize the average task service latency. Over a long period, this optimization problem has been modeled as follows:

[0100]

[0101] question The goal is to improve the ability and efficiency of customer vehicles in handling vision tasks through V2V collaboration, thereby optimizing long-term quality of service. Constraint (19a) limits the computing resources used by vehicles to perform tasks to no more than the total available resources. Constraint (19b) limits the total spectrum resources allocated by base stations to no more than the total available resources. Constraint (19c) stipulates that task offloading and result return are both binary decisions. Constraint (19d) requires the task non-completion rate to be less than a predetermined value to meet the quality of service standard.

[0102] 4. Problem Decoupling and Algorithm Design

[0103] This invention addresses the problem The problem is decoupled into two sub-problems: trajectory prediction and task scheduling. The trajectory prediction sub-problem is to infer future travel paths based on historical vehicle trajectories, providing global vehicle location information for task unloading decisions. The task scheduling sub-problem is to solve for the optimal scheduling strategy based on the predicted location relationships between customer vehicles and potential service vehicles, thereby achieving system load balancing.

[0104] 4.1 Trajectory Prediction Based on Spatiotemporal Feature Fusion

[0105] question Minimize long-term average service latency by accurately predicting vehicle trajectories.

[0106]

[0107] refer to Figure 3 This invention designs a trajectory prediction model based on spatiotemporal feature fusion, wherein the recurrent neural network (RNN)

[20] trajectory encoder extracts dynamic motion features from vehicle historical trajectory data; the spatiotemporal interactive encoder captures and fuses the spatiotemporal features of vehicle trajectory through parallel time and space graph attention; and the LSTM predictive decoder accurately predicts vehicle driving trajectory based on the fused spatiotemporal features. The three sub-modules are described in detail below.

[0108] 1) RNN Trajectory Encoder: The encoder predicts the trajectories of multiple vehicles in parallel and introduces a masking mechanism to adapt to dynamic changes in the number of vehicles. Let the set of historical trajectories of I vehicles be denoted as . And the linear transformation of embedding low-dimensional planar coordinates into a high-dimensional vector space is Emb(·). The vehicle motion features extracted by the input trajectory encoder are

[0109]

[0110] in This represents the motion characteristics of vehicle i in the scheduling slot t.

[0111] 2) The spatiotemporal interactive encoder system is modeled as a directed graph G = (U, V), where U is the set of vehicles and V is the set of edges. Motion characteristics The embedded node i. The spatiotemporal interaction encoder adopts a dual-branch structure: the multi-head attention of the temporal branch captures the temporal dependence of the vehicle trajectory, while the graph neural network (GNN)

[21] of the spatial branch captures the spatial interaction relationship of the vehicle trajectory. Specifically, The temporal features of the vehicle trajectory are obtained by inputting a temporal attention layer.

[0112]

[0113] Among them W t B t These are learnable weight matrices and bias terms. Additionally, the Spatial Graph Attention Mechanism (S-GAT) is used.

[22] Spatial features of vehicle trajectories are extracted by adaptively adjusting the adjacency matrix.

[0114]

[0115] S-GAT excels at capturing spatial correlations between zones within the global spatial information of vehicle trajectories, providing high-precision vehicle driving information to base stations in different zones. Finally, the spatiotemporal interactive encoder stitches together temporal and spatial features to generate spatiotemporal interactive fusion features.

[0116]

[0117] Concat(,) represents the concatenation operation.

[0118] 3) LSTM predictive decoder: spatiotemporal fusion features The input to the LSTM predictive decoder yields the predicted trajectories of I vehicles.

[0119]

[0120] 4.2 MADRL-based task scheduling

[0121] The task scheduling subproblem minimizes the system's long-term average task completion latency by optimizing the global offloading decision.

[0122]

[0123] question This can be transformed into a Multi-agent Markov Decision Process (MMDP). Each base station is abstracted as an agent. The MEC controller uses Multi-agent Deep Reinforcement Learning (MADRL) to solve the MMDP based on the global trajectory prediction results. MADRL utilizes multi-agent interactions to improve global task offloading and collaborative reasoning performance. The action space, state space, and reward function of task offloading are described below.

[0124] State space: Task offloading at base station m requires comprehensive consideration of task attributes, resource allocation, and candidate service vehicle trajectories. In the scheduling slot... Under these conditions, the environmental state of base station m can be represented as follows:

[0125]

[0126] in Let m represent the predicted trajectory of the vehicle belonging to base station m. Therefore, the global environment state of the M base stations is:

[0127]

[0128] Action Space: The MEC controller, in conjunction with the offloading actions of M base stations, generates a global scheduling policy. Therefore, the action decision of MADRL under scheduling slot t can be expressed as:

[0129]

[0130] Reward Function: Task service latency reflects the performance and efficiency of V2V collaborative inference. The system defines a joint reward function by minimizing the average service latency.

[0131]

[0132] in This represents the average service latency for all vehicles in scheduling slot t. P represents the time interval between two scheduling slots. t This indicates a penalty for failing to meet service delay standards.

[0133] Figure 4 This paper demonstrates the task scheduling process involving M agents, where the key to task offloading is selecting the most suitable service vehicle. To achieve this, each agent employs a Transformer layer and a Pointer Network layer. The former aggregates the state information of candidate service vehicles, while the latter uses a pointer mechanism to precisely select the service vehicle.

[0134] Suppose that the state characteristics of vehicle i under scheduling slot t include motion characteristics, predicted trajectory, and resources required for inference, i.e. Therefore, the set of state features of the vehicles belonging to base station m is: agent m will Input local DRL to obtain embedding vector Then, the Transformer is used to calculate the correlation between vehicles and output the weighted features.

[0135]

[0136] A context vector is calculated using equation (29).

[0137]

[0138] Among them I m This indicates the number of vehicles belonging to base station m. and The query vector set is obtained by passing each fully connected layer. and pointer vector Finally, the service vehicle selection probability based on feature similarity can be calculated as follows:

[0139]

[0140] The dimension η of the query vector is used to prevent training instability caused by an excessively large inner product; the activation function tanh(·) is used to ensure gradient stability. According to equation (31), the Pointer Network selects the vehicle with the highest probability as the service vehicle.

[0141] MADRL extends the decision-making process of agent m to other M-1 agents and employs centralized training and distributed execution to improve global task scheduling capabilities. A joint reward representing global value guides the policy updates of the M Actor networks. Let the stochastic policy of the Actor network in agent m be π. m and the probability of its action is The loss function can then be defined as:

[0142]

[0143] in Let represent the joint reward at iteration t. The agent updates the parameters using a gradient descent strategy. The gradient of equation (32) can be calculated as:

[0144]

[0145] in Let represent the joint reward at iteration t-1. Therefore, the update formula for the Actor network parameters is:

[0146]

[0147] The update rate β1 is usually set to a small constant to ensure the stability of convergence. As can be seen from equation (32), all M Actor networks use local state information for decision-making, but they rely on a global scheduling strategy optimized by joint reward output.

[0148] 4.3 Joint Optimization of Trajectory Prediction and Unloading Decision

[0149] Trajectory prediction can capture dynamic changes in network topology in real time, providing accurate spatiotemporal information for task offloading decisions; simultaneously, the execution effect of task offloading decisions can provide effective feedback for optimizing trajectory prediction. This coupling relationship prompts the system to construct a joint optimization of trajectory prediction and task offloading. At the beginning of window w, the MEC controller uses the historical trajectories of window w-1... Calculate the predicted trajectory Next, in scheduling slot t, agent m predicts the trajectory based on local predictions. Resource status and task attributes Generate uninstallation strategy This strategy is applied to local V2V or V2I2V collaborative reasoning and earns rewards. The joint reward of M agents This is used to jointly guide the parameter updates of the M Actor networks. Simultaneously, base station m collects and records the set of actual driving trajectories of vehicles belonging to its respective window w. And calculate the local root mean square prediction error.

[0150]

[0151] Let the set of real driving trajectories recorded by M base stations be . After receiving error values ​​from M base stations, the MEC controller calculates the global root mean square prediction error.

[0152]

[0153] To ensure that the offloading decision results improve trajectory prediction performance, the MEC controller calculates the global average policy loss at the end of window w using the local policy losses transmitted by M base stations.

[0154]

[0155] And establish a joint loss function

[0156]

[0157] Where β2 represents the weight. According to equation (38), the MEC controller uses gradient descent to fine-tune trajectory prediction in order to improve the system's environmental adaptability.

[0158] As shown above, the joint optimization of trajectory prediction and task scheduling forms a complete closed loop of "prediction-decision-feedback," which can provide a stable and efficient task scheduling strategy for vehicle-to-everything (V2X) networks. The proposed joint training process is summarized as Algorithm 1.

[0159]

[0160]

[0161] 5. Simulation Experiments and Performance Analysis

[0162] This section uses simulation experiments to comprehensively evaluate the effectiveness and efficiency of the proposed method.

[0163] A four-lane urban road with an intersection was constructed within a 5 square kilometer area. Four base stations with a communication range of 500 meters were evenly deployed along both sides of the road. The number of vehicles served by each base station ranged from [20, 100]. The average vehicle speed was 50 km / h. Vehicle trajectory data were obtained from the public dataset NGSIM US-101.

[23] The vehicle randomly generates tasks on each scheduling slot, and the task sizes follow a uniform distribution. The simulation experiment was deployed on the Ubuntu 24.04 operating system, with an Intel(R) Core(TM) i9-14900 CPU, a GeForce RTX 4090 GPU, and 24GB of video memory. The RNN encoder used a single-layer gated recurrent unit. The LSTM decoder used a two-layer LSTM, with each layer's hidden state set to 64 dimensions. The Adam optimizer was used to train the GNN, with the learning rate and number of training epochs set to 0.001 and 50 epochs, respectively. Other simulation parameters are detailed in Table 1.

[0164] Table 1 Simulation Experiment Parameters

[0165]

[0166] The five baseline methods selected for the experiment all include two modules: trajectory prediction and task scheduling. Table 2 lists the implementation details of each baseline method. Baseline method-2 uses the proposed trajectory prediction and employs DDQN for global unloading decision-making; baseline method-4 uses a multi-module behavior perception model to predict vehicle trajectories; baseline method-5 uses spatiotemporal tensor extraction and fusion of interactive features to predict trajectories.

[0167] Table 2 Implementation of 5 benchmark methods

[0168]

[0169]

[0170] 5.1 Convergence Analysis of Task Scheduling

[0171] To verify the stability of MADRL training, our experiments analyzed the convergence of the task scheduling algorithm. As shown in Figure 5(a), under the premise of using the trajectory prediction method presented in this paper, the average task service latency of the proposed method converges to a lower value faster than the baseline method-2. During the convergence phase, when the average task service latency of the baseline method-2 is around 0.32 seconds, the proposed method has stabilized at around 0.19 seconds, with an overall convergence performance improvement of about 40%. This performance difference is because the centralized DDQN of the baseline method-2 requires the state information of all vehicles as input. As the number of vehicles increases, the input dimension increases rapidly, leading to increased training complexity of DDQN, which in turn affects the convergence speed. The proposed method adopts multi-agent distributed training, where each agent only needs to perform local task scheduling, while the MEC controller uses joint rewards to optimize global task scheduling. This avoids input dimension explosion and utilizes the parallel computation of multiple agents to accelerate the overall training speed. The probability density distribution of the average service latency in Figure 5(b) also verifies this result. The peak position of the service latency probability density of the proposed method is much smaller than that of the baseline method-2, meaning that the average service latency of MADRL is smaller. Clearly, the proposed method is more suitable for handling task scheduling problems in large-scale vehicle-to-everything (V2X) networks.

[0172] 5.2 Impact of the number of tasks on performance

[0173] The second set of experiments evaluated the impact of task quantity on performance. In each scheduling slot across 20 time windows, the number of tasks generated by vehicles was divided into three ranges: 10–20, 20–30, and 30–40. Figures 6(a) and 6(b) show that the proposed method outperforms the baseline method-1 in both average task service latency and average task classification accuracy under different task quantities, with an average service latency reduction of approximately 18%–37% and an average classification accuracy increase of approximately 5%–15%. In the baseline method-1, customer vehicles only perform local processing and cannot utilize the computing resources of service vehicles for collaborative inference, thus increasing the average task service latency. Especially when task intensity is high (30–40), the proposed method achieves higher classification accuracy with lower processing latency. This is due to the model-segmentation-based V2V or V2I2V collaborative inference mechanism. Customer vehicles can fully utilize the computing resources of service vehicles to dual optimize inference accuracy and latency.

[0174] 5.3 The impact of vehicle speed on performance

[0175] The third set of experiments examined the impact of vehicle speed on performance. In Figure 7(a), when the vehicle speed increased from 30 km / h to 70 km / h, the average task service latency of all five methods increased, but the proposed method showed strong stability due to the smallest increase. Even when the vehicle speed reached 70 km / h, the average service latency of the proposed method was only 0.39 seconds, unaffected by high-speed movement. Corresponding to Figure 7(a), Figure 7(b) shows the average task completion rate of the five methods. The MADRL of the baseline method-3, due to the deviation in task offloading strategy caused by the lack of vehicle trajectory prediction, saw its average task service latency rise sharply from 0.24 seconds to 0.77 seconds, an increase of 221%. This highlights the crucial role of trajectory prediction in V2V task offloading, especially in high-mobility environments. Baseline methods-4 and-5 also used trajectory prediction to assist MADRL in obtaining more accurate offloading decisions, but their decision performance showed a significant downward trend with increasing vehicle speed. For example, when the vehicle speed is 70 km / h, the task completion rate decreases by 31.2% and 36.8% compared to the proposed method, while the average task service latency increases by 136.9% and 148%. The baseline method-4 has high behavioral perception prediction complexity and is difficult to adapt to highly dynamic scenarios; while the proposed prediction method is simple and efficient, better balancing prediction accuracy and real-time performance. Furthermore, compared to the DDQN of the baseline method-2, the proposed method's MADRL achieves locally optimal decisions for each agent through collaborative optimization of multi-agent decision-making behavior. When the vehicle speed is 70 km / h, the proposed method reduces the average task service latency by 20.4% and improves the task completion rate by 16.1% compared to the baseline method-2. These results demonstrate that the proposed method fully leverages the synergistic advantages of trajectory prediction and MADRL to achieve globally optimal task offloading and collaborative inference performance.

[0176] 5.4 Impact of Task Unloading Strategies on Performance

[0177] A key metric for task offloading strategies is the completion delay constraint. Failure to execute tasks within the delay constraint will severely impact the system's task completion rate. Therefore, the fourth set of experiments analyzes the impact of task offloading strategies on system service performance under different delay constraints. Task delay constraints are set as low-latency scenarios (0.2s-0.4s) and high-latency scenarios (1.0s-1.2s). Figures 8(a) and 8(b) respectively statistically analyze task offloading behavior within one hour under the two scenarios, where the horizontal axis represents the task sequence number, the vertical axis represents the offloading exit sequence number, and exit L indicates local execution. Under the low-latency constraint, customer vehicles tend to offload tasks to service vehicles with shallow entry points to offload most of the inference work and reduce their own computational burden; while under the high-latency constraint, customer vehicles tend to offload tasks to service vehicles with deep entry points to increase local inference work and reduce communication overhead. This demonstrates that a reasonable task offloading strategy can balance computational and communication overhead while strengthening information sharing and reducing the risk of processing failure.

[0178] 6. Summary

[0179] This invention proposes a model-segmentation-based V2V task offloading and collaborative inference method to improve system service quality in high-speed mobile environments. First, a trajectory prediction method based on spatiotemporal feature fusion is used to accurately predict the global vehicle trajectory, providing reliable location information for task offloading decisions. Then, MADRL, aided by the predicted trajectory, outputs the optimal global task offloading and collaborative inference strategy, achieving dual optimization of classification accuracy and inference efficiency. Simulation results show that the proposed method can improve system service quality with a high completion rate. Compared to other benchmark methods, the proposed method is more practical and superior in dynamic traffic environments.

[0180] References

[0181] [1]Nita-Rotaru C,Curtmola R.Security and privacy aspects in thededicated short-range communications(DSRC)protocol[M] / / Encyclopedia ofCryptography,Security and Privacy.Cham:Springer Nature Switzerland,2025:2279-2282.

[0182] [2]Liang Z,Han J,Li X,et al.Testing cellular vehicle-to-everythingcommunication performance and feasibility in automated vehicles[C] / / 2024IEEEIntelligent Vehicles Symposium(IV).IEEE,2024:2917-2922.

[0183] [3]Ngo H,Fang H,Wang H.Cooperative perception with V2V communicationfor autonomous vehicles[J].IEEE Transactions on Vehicular Technology,2023,72(9):11122-11131.

[0184] [4]Wang X,Li J,Ning Z,et al.Wireless powered mobile edge computingnetworks:A survey[J].ACM Computing Surveys,2023,55(13):1-37.

[0185] [5]Kumari A,Kumar S,Raw R S.Investigating the reliability of vehicle-to-vehicle and vehicle-to-infrastructure communication[C] / / InternationalConference on Innovations in Computational Intelligence and ComputerVision.Singapore:Springer Nature Singapore,2024:483-500.

[0186] [6]Wu X,Dong J,Bao W,et al.Augmented intelligence of things foremergency vehicle secure trajectory prediction and task offloading[J].IEEEInternet of Things Journal,2024,11(22):36030-36043.

[0187] [7]Zeng J,Gou F,Wu J.Task offloading scheme combining deepreinforcement learning and convolutional neural networks for vehicletrajectory prediction in smart cities[J].ComputerCommunications,2023,208:29-43.

[0188] [8]Men R,Fan X,Yan J,et al.A communication link lifetime prediction-supported V2V partial computation offloading scheme for autonomous driving[J].Journal of Intelligent&Fuzzy Systems,2024,46(3):6355-6368.

[0189] [9]Yan G,Liu C,Liu K.ASPM:Reliability-oriented DNN inferencepartition and offloading in vehicular edge computing[C] / / 2023IEEE 26thinternational conference on intelligent transportation systems(ITSC).IEEE,2023:3298-3303.

[0190]

[10] Wang L,Wang J,Dai C.DNN-based task partitioning and offloadingwith reliability guarantees in multi-UAV-assisted MEC system[C] / / 2024IEEEInternational Symposium on Product Compliance Engineering-Asia(ISPCE-ASIA).IEEE,2024:1-8.

[0191]

[11] Li H,Li X,Fan Q,et al.Distributed DNN inference with fine-grainedmodel partitioning in mobile edge computing networks[J].IEEE Transactions onMobile Computing,2024,23(10):9060-9074.

[0192]

[12] Zhang X,Wang C,Zhu Y,et al.Multi-agent deep reinforcementlearning with trajectory prediction for task migration-assisted computationoffloading[J].IEEE Transactions on Mobile Computing,2025,24(7):5839-5856.

[0193]

[13] Ji S,Li J,Jin H,et al.Resource aware multi-user task offloadingin mobile edge computing[C] / / 2024 IEEE International Conference on WebServices(ICWS).IEEE,2024:665-675.

[0194]

[14] Zeng F,Zhang Z,Wu J.Task offloading delay minimization invehicular edge computing based on vehicle trajectory prediction[J].DigitalCommunications and Networks,2024,11(2):537-546.

[0195]

[15] Manasreh D,Swaleh S,Nazzal M D.Evaluation of BLE beacontechnology for time critical I2V communication to support CAV deployment onurban roadways[J].Internet of Things,2023,24:100932.

[0196]

[16] Katragadda J.Performance comparison of IEEE 802.11 P and IEEE802.11 BD for advanced V2X applications[C] / / 2024 First InternationalConference on Data,Computation and Communication(ICDCC).IEEE,2024:593-597.

[0197]

[17] Fang Y,Li M,Yu F R,et al.Parallel offloading and resourceoptimization for multi-hop adhoc network-enabled CBTC with mobile edgecomputing[J].IEEE Transactions on Vehicular Technology,2023,73(2):2684-2698.

[0198]

[18] Wang K,Wang X,Liu X,et al.Task offloading strategy based onreinforcement learning computing in edge computing architecture of internetof vehicles[J].IEEE Access,2020,8:173779-173789.

[0199]

[19] Veeramanickam M R M,Venkatesh B,Bewoor L A,et al.IoT based smartparking model using arduino UNO with FCFS priority scheduling[J].Measurement:Sensors,2022,24:100524.

[0200]

[20] Al-Aql N,Al-Shammari A.Hybrid RNN-LSTM networks for enhancedintrusion detection in vehicle can systems[J].Journal of Electrical Systems,2024,20(6):3019-3031.

[0201]

[21] Han K,Wang Y,Guo J,et al.Vision GNN:An image is worth graph ofnodes[J].Advances in Beural Information Processing Systems,2022,35:8291-8303.

[0202]

[22] Sang H,Li S,Wang J,et al.Group vehicle trajectory predictionmodel based on multi-graph fusion[J].Computers and Electrical Engineering,2025,123:110053.

[0203]

[23] Coifman B,Li L.A critical evaluation of the next generationsimulation(ngsim)vehicle trajectory dataset[J].Transportation Research PartB:Methodological,2017,105:362-377.

[0204]

[24] Liao H,Li Z,Shen H,et al.Bat:Behavior-aware human-like trajectoryprediction for autonomous driving[C] / / Proceedings of the AAAI Conference onArtificial Intelligence.2024,38(9):10332-10340.

[0205]

[25] Wang Y,Zhao S,Zhang R,et al.Multi-vehicle collaborative learningfor trajectory prediction with spatio-temporal tensor fusion[J].IEEETransactions on Intelligent Transportation Systems,2020,23(1):236-248.

Claims

1. A V2V task offloading and collaborative reasoning method based on model segmentation, characterized in that... In the Internet of Vehicles (IoV) for V2V task offloading and collaborative reasoning: (I) First, a vehicle trajectory prediction method based on spatiotemporal feature fusion is adopted to predict the global vehicle trajectory; (ii) Then, multi-agent deep reinforcement learning based on model segmentation is used to obtain the future position of the vehicle from the global vehicle trajectory and obtain the optimal global task scheduling strategy. The proposed vehicle-to-everything (V2V) architecture for V2V task offloading and collaborative reasoning consists of an MEC controller, base stations, and vehicles. Let the sets of base stations and vehicles be... and The base station schedules tasks based on the predicted vehicle trajectory, computing power, and queue load, and sends scheduling information; the MEC controller connects to the base station and undertakes the task of predicting vehicle trajectories. Each vehicle is equipped with a vehicle model for handling visual tasks. The vehicle model takes external environment image information collected by the vehicle's vision sensors as input and obtains inference results. The network structure of the vehicle model, i.e., the vehicle network, is a deep neural network containing multiple separable modules. The vehicle network is divided into a multi-exit network. At a certain exit, the vehicle network is divided into a shallow network and a deep network to obtain a collaborative network for collaborative inference. The vehicle that initiates the collaboration request is the customer vehicle, and the vehicle that provides the collaboration service is the service vehicle. The roles of the customer vehicle and the service vehicle are dynamically changing. In the collaboration state, the customer vehicle offloads the task to the service vehicle, which performs the corresponding inference work. After completing the inference work, the service vehicle transmits the inference results back to the customer vehicle. In step (1), a trajectory prediction model based on spatiotemporal feature fusion is used for trajectory prediction. In the trajectory prediction model, firstly, the recurrent neural network (RNN) trajectory encoder extracts the dynamic motion features of the vehicle through historical trajectory data. Then, the spatiotemporal interactive encoder captures and fuses the spatiotemporal features of the vehicle trajectory through parallel temporal and spatial graph attention. Finally, the LSTM predictive decoder accurately predicts the vehicle's driving trajectory based on the fused spatiotemporal features. In step (ii), each base station is abstracted as an agent; the MEC controller uses multi-agent deep reinforcement learning (MADRL) to solve the multi-agent Markov decision process (MMDP) problem based on the global trajectory prediction results.

2. The V2V task offloading and collaborative reasoning method based on model segmentation according to claim 1, characterized in that: It also includes (iii) using a joint training method of trajectory prediction and task unloading to dually optimize trajectory prediction accuracy and collaborative reasoning efficiency.

3. The V2V task offloading and collaborative reasoning method based on model segmentation according to claim 1 or 2, characterized in that: In a vehicle cooperative network, the in-vehicle network is a deep neural network containing multiple separable L modules; given a set of exits... The vehicular network is segmented into a multi-exit network; vehicle In export The in-vehicle network is segmented into shallow networks. and deep networks φ i,l To generate a collaborative network for collaborative reasoning; vehicles that propose collaborative needs are labeled as customer vehicles (CV), and vehicles that provide collaborative services are labeled as service vehicles (SV); The base station schedules tasks based on the vehicle's predicted trajectory, computing power, and task queue load, and sends scheduling information. The MEC controller connects to the base station and is responsible for vehicle trajectory prediction. Vehicle trajectory prediction model: The position coordinates of vehicle i at time t are Then it is in the time window The historical trajectory is in Indicates the length of the backtracking horizon; MEC controller will The input trajectory prediction model produces a predicted trajectory. This trajectory provides the base station with the vehicle's future location to optimize unloading decisions.

4. The V2V task offloading and collaborative reasoning method based on model segmentation according to claim 3, characterized in that it is for vehicle networking. Communication in: Communication between customer vehicles and service vehicles includes two types: vehicle-to-vehicle (V2V) communication and vehicle-to-infrastructure (V2I2V) communication. 1) V2V communication: V2V communication adopts a multi-hop self-organizing network architecture, and its channel follows independent and identically distributed Rayleigh fading; Let the set of vehicles serving customer vehicle i be . Customer vehicles and service vehicles The path loss function between them is κ i,j =63.3 + 17.7log 10 (d i,j ), where d i,j Let represent the distance between vehicle i and vehicle j; V2V communication is affected by path loss, Rayleigh fading, and neighbor interference, so the channel signal-to-noise ratio is calculated as follows: Where p i and p j Both represent transmission power, g i,j Indicates Rayleigh's loss, Indicates Gaussian white noise; According to Shannon's formula, the V2V data transfer rate is... μ i,j =B V2V log2(1+SINR i,j ) (4) Among them B V2V Indicates the bandwidth of V2V communication; 2) Vehicle-to-Infrastructure (V2I2V) Communication: If customer vehicle i leaves the communication range of service vehicle j before receiving the final inference result from service vehicle j, the inference result is forwarded to customer vehicle i via base station m; the signal-to-noise ratio of the channel between service vehicle j and base station m is... Where d j,m (d j,m′ ) represents the distance between vehicle j and base station m (m′), α represents the path loss exponent, and g j,m Indicates Rayleigh's loss, Indicates Gaussian white noise; According to Shannon's formula, the uplink vehicle-to-infrastructure (V2I) data transfer rate is... μ j,m =B V2I log2(1+SINR j,m ) (6) Among them B V2I This represents the V2I communication bandwidth; then, base station m sends the received inference results to customer vehicle i; the channel signal-to-noise ratio between base station m and customer vehicle i is... Where d i,m g represents the distance between customer vehicle i and base station m. i,m Indicates Rayleigh's loss; Therefore, the downlink infrastructure-to-vehicle I2V data transmission rate is μ m,i =B I2V log2(1+SINR m,i ) (8) Among them B I2V This indicates the I2V communication bandwidth.

5. The V2V task offloading and collaborative reasoning method based on model segmentation according to claim 4, characterized in that: Task completion delay in the Internet of Vehicles: Divide the time domain into a series of time windows of equal length. Each window contains a set of scheduling slots of equal length. The duration of the time window is f (w) Let the sets of computing units configured for customer vehicle i and service vehicle j be respectively... and The base station's sub-channel set is m Both vehicles and base stations follow the First-Come, First-Served (FCFS) principle when handling visual tasks. When window Initially, the MEC controller jointly fine-tunes the trajectory prediction model and the task offloading model based on the vehicle's historical driving trajectory in window w-1. In the scheduling slot Based on task requirements, the base station uses the MADRL algorithm to generate the optimal V2V task offloading and collaborative reasoning strategy; The visual task h of customer vehicle i is composed of the triple {x h ,τ h ,ε h } constitutes, where x h ,τ h ,ε h These represent the computational cost required for perceiving the image, inference, and the completion delay constraint, respectively. Task completion delay includes task unloading delay, processing delay, and result return delay; 1) Uninstallation delay Offload latency is the latency required for a customer vehicle to perform shallow network inference tasks and transmit intermediate features to the service vehicle; offload latency is calculated in two ways: V2V and V2I2V. Under the scheduling slot t of window w, the set of tasks initiated by customer vehicle i is as follows: Its base is Then the set of customer vehicle tasks received by base station m is Its base is Suppose that the number of computing units equipped in customer vehicle i is O. i And the maximum GPU cycle time for each computing unit is Set shallow network The extracted intermediate features are Its data size and the computational cost required for inference are respectively and Define a binary variable to represent partial unloading (1) or local processing (0). Task h process The V2V offloading latency required for inference and result transmission to service vehicle j is The two terms in equation (10) represent the computation delay and the transmission delay, respectively; The customer vehicle offloads the task to the service vehicle in a V2I2V manner and delivers the final inference result when the two vehicles approach each other; Let there be two variables. and If V2V and V2I2V are unloaded respectively, then equation (10) is changed to 2) Handling delays Processing latency is the time required for deep network inference of intermediate features for serving vehicles; Let the number of computing units equipped in service vehicle j be . And the maximum GPU cycle time for each computing unit is Let the deep network φ j,l Continue reasoning about intermediate features The required computation is The expression for handling delays is: 3) Return delay Return latency is the time required for the final inference result to be transmitted to the customer's vehicle via V2V or V2I2V communication. Let the final reasoning result of task h be Its data size is Let there be two variables. and Let represent the visual processing results (i.e., inference results) transmitted by service vehicle j via V2V and V2I2V communication, respectively. Then the return latency is expressed as... Modeling problems based on delay Under the scheduling slot t of window w, the task completion delay is the sum of the unloading delay, processing delay, and result return delay, i.e., equations (11), (12), and (13). Then, the average service latency is expressed as in Represents a set of vehicles The cardinality; To evaluate the effectiveness of the task scheduling strategy, a binary variable is defined. Where ε h This indicates the time constraint for completing task h; and These represent the processing results of task h in ε. h Successful and unsuccessful deliveries; Let the scheduling strategy of base station m in scheduling slot t be: The set of scheduling strategies within the time window w is: The corresponding scheduling strategy, and the allocation strategy for vehicle computing resources and base station spectrum resources are as follows: and in Record the incomplete rate of tasks within window w. According to equation (17), the task non-completion rate is expressed as: Given a set of time windows The optimization objective of V2V offloading is to minimize the average task service latency. Over a long period, the optimization problem is modeled as follows: question The goal is to improve the ability and efficiency of customer vehicles in handling vision tasks through V2V collaboration, thereby optimizing long-term service quality. Constraint (19a) limits the computing resources used by the vehicle to perform the task to no more than the total amount; Constraint (19b) limits the sum of spectrum resources allocated to base stations to no more than the total amount; Constraint (19c) stipulates that task unloading and result return are both binary decisions; Constraint (19d) requires the task non-completion rate to be less than the predetermined value in order to meet the service quality standard; Problem decoupling The problem The problem is decoupled into two sub-problems: trajectory prediction and task scheduling. Trajectory prediction subproblem It infers the future driving path of a vehicle based on its historical trajectory, providing global vehicle location information for task unloading decisions; question Minimize long-term average service latency by accurately predicting vehicle trajectories. Task scheduling subproblem It is to solve for the optimal scheduling strategy based on the predicted location relationship between customer vehicles and potential service vehicles in order to achieve load balancing of the system; question Minimize the system's long-term average task completion latency by optimizing global offloading decisions.

6. The V2V task offloading and collaborative reasoning method based on model segmentation according to claim 5, characterized in that... In the trajectory prediction model based on spatiotemporal feature fusion in step (one): 1) RNN trajectory encoder The encoder predicts the driving trajectories of multiple vehicles in parallel and introduces a masking mechanism to adapt to the dynamic changes in the number of vehicles. Let the set of historical trajectories of I vehicles be . And the linear transformation of embedding low-dimensional planar coordinates into a high-dimensional vector space is Emb(·); The vehicle motion features extracted by the input trajectory encoder are in This represents the motion characteristics of vehicle i in the scheduling slot t; 2) Spatiotemporal interactive encoder Construct a directed graph G = (U, V) to represent the vehicle cooperation network, where U is the set of vehicles and V is the set of edges; Motion characteristics Embedded node i; The spatiotemporal interactive encoder adopts a dual-branch structure: Multi-head attention in time-branching captures the temporal dependencies of vehicle trajectories; Spatial branching graph neural networks (GNNs) capture the spatial interaction relationships of vehicle trajectories; The temporal features of the vehicle trajectory are obtained by inputting a temporal attention layer. Among them W t B t It consists of learnable weight matrices and bias terms; Spatial Graph Attention Network (S-GAT) extracts spatial features of vehicle trajectories by adaptively adjusting the adjacency matrix. S-GAT captures the spatial correlation between zones in the global spatial information of vehicle trajectories, providing high-precision vehicle driving information for base stations in different zones; Finally, the spatiotemporal interaction encoder concatenates temporal and spatial features to generate fused features of spatiotemporal interaction. Where Concat(·,·) represents the concatenation operation; 3) LSTM predictive decoder Spatiotemporal fusion features The input to the LSTM predictive decoder yields the predicted trajectories of I vehicles.

7. The V2V task offloading and collaborative reasoning method based on model segmentation according to claim 6, characterized in that: In step (two), task scheduling based on MADRL: question The problem is transformed into a multi-agent Markov decision process (MMDP); each base station is abstracted as an agent; the MEC controller uses multi-agent deep reinforcement learning (MADRL) to solve the MMDP problem based on the global trajectory prediction results. The action space, state space, and reward function for task unloading are described below: State space: Task offloading at base station m requires comprehensive consideration of task attributes, resource allocation, and candidate service vehicle trajectory information; in the scheduling slot Under these conditions, the environmental state of base station m is represented as follows: in Let m represent the predicted trajectory of the vehicle belonging to base station m; therefore, the global environment state of M base stations is: Action space: The MEC controller, in conjunction with the offloading actions of M base stations, generates a global scheduling strategy; Therefore, the action decision of MADRL under scheduling slot t is represented as: Reward function: Task service latency reflects the performance and efficiency of V2V collaborative inference; the joint reward function is defined by minimizing the average service latency. in Δf represents the average service latency of all vehicles in scheduling slot t. t (w) P represents the time interval between two scheduling slots. t This indicates the penalty for failing to meet service delay standards; In the task scheduling process of M agents, the key to task unloading is to select the most suitable service vehicle; therefore, each agent is equipped with a Transformer layer and a Pointer Network layer, which are responsible for aggregating the state information of candidate service vehicles and using the pointer mechanism to accurately select the service vehicle, respectively. Suppose that the state characteristics of vehicle i under scheduling slot t include motion characteristics, predicted trajectory, and resources required for inference, i.e. Therefore, the set of state features of the vehicle to which base station m belongs is: The decision-making process of an agent m is as follows: agent m will... Inputting local reinforcement learning DRL to obtain embedding vectors Then, the Transformer is used to calculate the correlation between vehicles and output the weighted features. A context vector is calculated using equation (29). Among them I m This indicates the number of vehicles belonging to base station m; and The query vector set is obtained by passing each fully connected layer. and pointer vector Finally, the service vehicle selection probability based on feature similarity can be calculated as follows: The dimension η of the query vector is used to prevent training instability caused by an excessively large inner product; the activation function tanh(·) is used to ensure gradient stability. According to equation (31), the vehicle with the highest probability is selected as the service vehicle; MADRL extends the decision-making process of one agent m to other M-1 agents, and uses centralized training and distributed execution to improve global task scheduling capabilities; the joint reward, representing global value, guides the policy updates of the M Actor networks. Let the stochastic policy of the Actor network in agent m be π. m and the probability of its action is The loss function is defined as follows: Where r t (w) Denotes the joint reward at iteration t; The agent m updates its parameters using a gradient descent strategy; the gradient of equation (32) is calculated as follows: in This represents the joint reward for the (t-1)th iteration; Therefore, the update formula for the Actor network parameters is: Where β1 represents the update rate; As shown in equation (32), all M Actor networks use local state information to make decisions, and they rely on a global scheduling strategy optimized by joint reward output.

8. The V2V task offloading and collaborative reasoning method based on model segmentation according to claim 7, characterized in that: In step (iii), the joint optimization method for trajectory prediction and unloading decision is as follows: When window w starts, the MEC controller uses the historical trajectory of window w-1. Calculate the predicted trajectory Next, in scheduling slot t, agent m predicts the trajectory based on local predictions. Resource status and task attributes Generate uninstallation strategy This strategy is applied to local V2V or V2I2V collaborative reasoning and earns rewards. The joint reward of M agents Used to jointly guide the parameter updates of M Actor networks; simultaneously, base station m collects and records the set of actual driving trajectories of vehicles belonging to its respective window w. And calculate the local root mean square prediction error. Let the set of real driving trajectories recorded by M base stations be . After receiving error values ​​from M base stations, the MEC controller calculates the global root mean square prediction error. At the end of window w, the MEC controller calculates the global average policy loss using the local policy losses transmitted by M base stations. And establish a joint loss function Where β2 represents the weight; According to Equation (38), the MEC controller uses gradient descent to fine-tune trajectory prediction in order to improve the system’s environmental adaptability; The joint optimization of trajectory prediction and task scheduling forms a complete closed loop of "prediction-decision-feedback", providing stable and efficient task scheduling strategies for the Internet of Vehicles.