High expressway multi-tunnel group V2X communication time delay compensation method based on dynamic game-reinforcement learning double-layer framework
Through the V2X communication delay compensation method of dynamic game-reinforced learning dual-layer framework, the problem of high delay and bit error rates caused by signal occlusion and multipath effect in tunnel group scenarios is solved, efficient spectrum and power distribution is achieved, and communication efficiency and system reliability are improved.
Patent Information
- Application Number
- CN202510750963.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-08-29
AI Technical Summary
In continuous tunnel group scenarios such as highways and expressways, V2X communication faces problems such as signal occlusion and high delay and bit error rates caused by multipath effect, poor resource allocation, lack of self-healing mechanisms and policy transmission discontinuity, which affects communication efficiency and reliability.
Using a dual-layer framework based on dynamic game-reinforced learning, we optimize spectrum and power distribution, design a self-healing mechanism and policy delivery mechanism to realize communication delay compensation through tunnel entrance precompensation, continuous compensation within the tunnel, tunnel exit recovery and cross-tunnel cloud collaboration strategies.
Improve the adaptability of resource allocation in dynamic environments, optimize spectrum and power allocation in real time, reduce bit error rate and delay, ensure communication efficiency and system reliability, quickly recover communication failures, and maintain policy continuity and consistency.
Smart Images

Figure CN120568451A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence and V2X communication technology, and specifically relates to a method for compensating V2X communication delay in high-speed road multi-tunnel groups based on a dynamic game-reinforcement learning dual-layer framework. Background Art
[0002] V2X (vehicle-to-everything) communication technology has a multifaceted application in the transportation sector. Through various communication methods, including vehicle-to-vehicle (V2V), vehicle-to-infrastructure (V2I), vehicle-to-pedestrian (V2P), and vehicle-to-network (V2N), it significantly improves traffic safety, optimizes traffic efficiency, supports and enhances the functionality of intelligent transportation systems, promotes environmental protection, and enhances emergency response capabilities. These applications enhance road safety and traffic efficiency, driving the intelligent and sustainable development of the transportation industry. With continued technological advancement, the application of V2X communication technology in the transportation sector is expected to expand and optimize.
[0003] V2X communication faces a series of unique challenges in scenarios with continuous tunnels, such as those on highways and expressways. First, the complex tunnel environment means signals are easily blocked and attenuated by the tunnel structure, resulting in degraded communication quality. Second, tunnels experience significant multipath effects, with reflections and refractions increasing signal latency and bit error rates. Furthermore, the speed and position of vehicles in tunnels constantly change, dynamically adapting communication needs and the environment. Finally, spectrum resources in tunnels are limited and need to be rationally allocated to meet the communication needs of different vehicles, placing higher demands on resource management.
[0004] Existing V2X communication technology still has shortcomings in scenarios involving continuous tunnel clusters. First, existing resource allocation strategies are poorly adaptable in dynamic environments and cannot optimize spectrum and power allocation in real time, resulting in low communication efficiency and significantly increased bit error rates and latency. Furthermore, existing systems lack effective self-healing mechanisms in the face of communication failures, preventing rapid communication restoration. Finally, existing technologies lack effective policy transmission mechanisms in multi-tunnel scenarios, resulting in insufficient policy continuity and consistency, impacting the overall performance and reliability of the system. Summary of the Invention
[0005] To address the shortcomings of the existing technology, the present invention provides a high-speed road multi-tunnel group V2X communication delay compensation method based on a dynamic game-reinforcement learning dual-layer framework, which can optimize spectrum and power allocation in real time, improve communication efficiency, and reduce bit error rate and delay.
[0006] In order to achieve the above object, the technical solution adopted by the present invention is: In a first aspect, a method for compensating V2X communication delay in a high-speed road tunnel group is provided, comprising: in response to a vehicle about to enter a tunnel, executing a set tunnel entrance communication delay pre-compensation strategy; in response to the vehicle traveling in the tunnel, executing a set in-tunnel continuous compensation strategy; in response to the vehicle exiting the tunnel, executing a set tunnel exit delay recovery strategy; in response to the vehicle entering a tunnel group, executing a set cross-tunnel cloud-based collaboration strategy and tunnel group closed-loop optimization strategy.
[0007] Furthermore, the tunnel entrance communication delay pre-compensation strategy is implemented through a hybrid game-reinforcement learning two-layer framework, including: at the lower layer, using the DDPG model to make spectrum resource allocation decisions based on the current channel state; at the upper layer, updating and optimizing the vehicle alliance's power strategy through game theory methods to obtain the optimal power allocation decision; applying the decision results of the upper and lower layers to the actual system, executing the pre-compensation strategy, and measuring the actual end-to-end delay; calculating the reward value based on the actual measured delay, and updating the DDPG model through an online learning mechanism, thereby achieving effective pre-compensation for the tunnel entrance communication delay.
[0008] Furthermore, at the lower level, the DDPG model optimizes the spectrum allocation strategy based on the predicted delay value to maximize the reward value R: , in, is the weight of the delay; is the predicted delay value; is the optimal state of vehicle i, i.e., the spectrum allocation strategy; is the optimal action of vehicle i, i.e., the power allocation strategy; It is the whole system strategy; is the power consumption of the ith vehicle, is the power consumption weight of the i-th vehicle, is the mathematical expectation; By reconstructing the reward function and dynamically adjusting the reward value, the model can be optimized towards reducing latency. , , , in, is the delay-related reward value; is the actual measured delay; It is a custom delay tolerance threshold; At the upper layer, a dynamic asymmetric game modeling tunnel group communication is introduced, and the participants are defined as each vehicle i in the vehicle alliance; the strategy space is the power strategy of each vehicle and spectrum allocation strategies ;The data transmission rate is ; The utility function of each vehicle is : , in, is the weight of the data rate; is the end-to-end delay of vehicle i; and are the power and spectrum strategies of vehicle i, respectively; and are the power and spectrum strategies of other vehicles respectively; for each vehicle i, in the strategies of other vehicles When fixed, is the optimal strategy: , in, is the optimal power strategy set of other vehicles at equilibrium, is the optimal spectrum strategy set of other vehicles in equilibrium; each vehicle updates its strategy through gradient ascent: , , in, is the learning rate, is the updated power strategy of vehicle i at the k+1th iteration, is the updated spectrum strategy of vehicle i in the k+1th iteration, 、 They represent the power strategy set and spectrum strategy set of all other vehicles at the kth iteration respectively.
[0009] Furthermore, the continuous compensation strategy in the tunnel includes: building a hybrid network architecture of VANET and fixed relay, designing a link selection algorithm based on multi-attribute decision-making; using entropy weight method to determine the weight of each attribute , , Where S is the total number of attributes considered in link selection, is the entropy value of the j-th attribute, , Where m is the total number of vehicles, is the normalized value of vehicle i on attribute j, is the normalized value of other vehicle k on attribute j; Calculate the comprehensive evaluation value of each plan , , in, is the normalized maximum value of all vehicles on attribute j, is the normalized minimum value of all vehicles on attribute j; The link with the highest score is selected through comprehensive evaluation value, and then distributed ADMM is used for control and continuous compensation of delay, and the global optimal power strategy is adopted in real time. and global optimal spectrum strategy Optimize vehicle power strategy and spectrum allocation strategy, the objective function is: , , , in, It is delay The weight coefficient of It is power The weight coefficient of 、 is the regularization coefficient, balancing the deviation between the current strategy and the optimal strategy, is the maximum transmit power of the vehicle.
[0010] Furthermore, the tunnel exit delay recovery strategy includes: proposing a coordinate transformation method based on the Lie group in the time-space reference unification algorithm, defining the SE(3) transformation matrix, using nonlinear Kalman filtering to compensate for the cumulative error of INS, and improving the positioning accuracy by predicting the update formula of the covariance matrix to ensure the high-priority transmission of emergency messages and the accurate unification of the time-space reference.
[0011] Furthermore, to ensure the accurate and unified time and space reference of the vehicle at the tunnel exit, the local coordinate system of the vehicle is converted to the WGS84 coordinate system; then the strategy starting distance is determined, based on the tunnel length L and the maximum speed of the vehicle. , vehicle i's current delay , calculate the starting distance : , Dynamically adjust real-time strategies based on current communication and positioning status , , so that it gradually approaches the optimal strategy , , so as to minimize the communication delay, , in, It is at the moment Real-time strategy when It is at the moment Real-time strategy when 、 It is at the moment Real-time power strategy and spectrum allocation strategy; 、 It is the optimal power strategy and spectrum allocation strategy; 、 is the learning rate, which controls the speed of policy updates.
[0012] Furthermore, the cross-tunnel cloud collaboration strategy includes: first, designing a distributed update framework based on federated learning in the iterative learning of digital twins; each tunnel entrance and exit edge node maintains the optimal power strategy and spectrum allocation strategies and delay , and perform model training based on local data; the cloud is responsible for performing model aggregation and calculating the global model , coordinate cloud and edge nodes; secondly, in the dynamic resource allocation mechanism, a resource optimization model of end-edge-cloud collaboration is established; by adopting the Lyapunov optimization framework, long-term constraints are converted into real-time decisions to achieve joint optimization of communication and computing resources, while reducing energy consumption and meeting real-time requirements.
[0013] Furthermore, the global model for: , in, is the data volume of tunnel exit / entry node i, n is the total data volume of all nodes, are the local model parameters of node i, The resource optimization model for device-edge-cloud collaboration has the following optimization goals: , The constraints are: , in, is the communication energy consumption of the i-th node; is the computational energy consumption of the i-th node; is the total delay; is the maximum allowed delay.
[0014] Furthermore, the tunnel group closed-loop optimization strategy includes: first, establishing adjacent tunnel strategy transfer rules, updating the strategy to the optimal strategy after exiting the tunnel, ending the single tunnel process, and setting the optimal power strategy of the current tunnel exit to and spectrum allocation strategies Pass it to the next tunnel entrance; secondly, establish a fault self-healing mechanism, monitor the system status in real time through the anomaly detection model, and trigger the reconstruction of the compensation strategy when the Mahalanobis distance threshold is reached. After the anomaly is detected, it is automatically adjusted to ensure that the current delay status is always maintained at the optimal strategy.
[0015] In a second aspect, a high-speed road tunnel group V2X communication delay compensation system is provided, comprising a storage medium and a processor; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the high-speed road tunnel group V2X communication delay compensation method described in the first aspect.
[0016] Compared with the prior art, the present invention has the following beneficial effects: (1) The present invention implements a pre-compensation strategy for tunnel entrance communication delay when a vehicle is about to enter a tunnel; implements a continuous compensation strategy within the tunnel; implements a tunnel exit delay recovery strategy upon exiting the tunnel; and implements a cross-tunnel cloud collaboration strategy and tunnel group closed-loop optimization strategy within the tunnel group. This improves the adaptability of resource allocation strategies in dynamic environments, enables real-time optimization of spectrum and power allocation, improves communication efficiency, and reduces bit error rate and delay. (2) The present invention designs an effective self-healing mechanism that can quickly restore communication in the event of a communication failure; (3) The present invention designs an effective policy delivery mechanism, which maintains the continuity and consistency of the policy in multi-tunnel scenarios and improves the overall performance and reliability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a schematic diagram of the main process of a method for compensating V2X communication delay in a high-speed road multi-tunnel group based on a dynamic game-reinforcement learning dual-layer framework provided by an embodiment of the present invention; Figure 2 Schematic diagram of a hybrid game-reinforcement learning two-layer framework in an embodiment of the present invention. DETAILED DESCRIPTION
[0018] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention.
[0019] Example 1 like Figure 1 、 Figure 2 As shown, a method for compensating V2X communication delay in a high-speed highway multi-tunnel group based on a dynamic game-reinforcement learning dual-layer framework includes: in response to a vehicle about to enter a tunnel, executing a set tunnel entrance communication delay pre-compensation strategy; in response to the vehicle traveling in the tunnel, executing a set tunnel continuous compensation strategy; in response to the vehicle exiting the tunnel, executing a set tunnel exit delay recovery strategy; in response to the vehicle entering the tunnel group, executing a set cross-tunnel cloud-based collaboration strategy and tunnel group closed-loop optimization strategy.
[0020] Step 1: Complete tunnel entrance communication delay pre-compensation through a hybrid game-reinforcement learning two-layer framework.
[0021] First, at the lower layer, a deep deterministic policy gradient (DDPG) model is used to make spectrum resource allocation decisions based on the current channel state. Next, at the upper layer, a game-theoretic approach is used to update and optimize the vehicle alliance's power policy, resulting in the optimal power allocation strategy. The decision results from both layers are then applied to the actual system, pre-compensation strategies are implemented, and actual end-to-end latency is measured. Finally, a reward is calculated based on the measured latency, and the DDPG model is updated through an online learning mechanism, effectively pre-compensating tunnel entrance communication latency.
[0022] The delay control mechanism of the upper-layer game. Delay modeling is the foundation of the two-layer pre-compensation architecture, used to analyze and optimize end-to-end delay. End-to-end delay is composed of queuing delay, transmission delay, and processing delay.
[0023] End-to-end delay consists of: , in, is the end-to-end delay; is the queuing delay, the time data waits in the queue for processing; is the transmission delay, the time it takes for data to be transmitted in the channel; It is the processing delay, which is the time it takes for data to be processed at the receiving end.
[0024] The goal of precompensation is to minimize the weighted sum of squares of delay variations, where the weights are determined by the time decay factor.
[0025] Precompensation objective function : , Where K is the number of vehicles or data flows; is the timestamp of the k-th vehicle or the k-th data stream; It is the time decay factor, which is used to reflect the timeliness of the delay; is the delay tolerance threshold, the maximum delay that the system expects to achieve; is the actual end-to-end delay of the kth vehicle or data flow.
[0026] The delay control mechanism of the upper-level game introduces a dynamic asymmetric game to more comprehensively model the complex interactive relationship in multi-tunnel group communication. Game definition: the participants are each vehicle i in the vehicle alliance; the strategy space is the power strategy of each vehicle and spectrum allocation strategies ;The data transmission rate is ; The utility function of each vehicle is : , in, is the weight of the data rate; is the weight of power consumption; is the weight of the delay; is the end-to-end delay of vehicle i; and are the power and spectrum strategies of vehicle i, respectively; and are the power and spectrum strategies of other vehicles respectively. For each vehicle i, in the strategies of other vehicles When fixed, is the optimal strategy: , in, is the optimal power strategy set of other vehicles at equilibrium, is the optimal spectrum strategy set of other vehicles in equilibrium; each vehicle updates its strategy through gradient ascent: , , in, is the learning rate, is the updated power strategy of vehicle i at the k+1th iteration, is the updated spectrum strategy of vehicle i in the k+1th iteration, 、 They represent the power strategy set and spectrum strategy set of all other vehicles at the kth iteration respectively.
[0027] The latency prediction mechanism of the lower-level DDPG layer. Delay gradient information includes the partial derivatives of latency with respect to various variables, such as channel state, vehicle speed, and interference. Based on this information, an LSTM latency predictor is constructed to effectively capture dynamic latency changes. Furthermore, by reconstructing the reward function, the DDPG model is further incentivized to optimize latency performance.
[0028] , , in, is the input time-delay gradient sequence; LSTM It is a long short-term memory network, used to process time series data; regressor It is a linear regression layer that outputs the predicted delay value; is the output sequence after LSTM processing, is the delay prediction value.
[0029] The lower-layer DDPG model enhances the state space and introduces delay gradient information to improve the accuracy of delay prediction.
[0030] State Space Augmentation: , in, is the time-delayed gradient information; is the channel state variable; is the vehicle speed; is the interference level. The state-action space is defined as shown in Table 1.
[0031] Table 1 Definition of state-action space
[0032] DDPG optimizes the spectrum allocation strategy by predicting the delay value to maximize the reward value R: , in, is the weight of the delay; f is the predicted delay prediction function; is the optimal state of vehicle i (spectrum allocation strategy); is the optimal action (power strategy) for vehicle i; It is the whole system strategy; is the power consumption of the ith vehicle, is the power consumption weight of the i-th vehicle, is the mathematical expectation.
[0033] To further incentivize the DDPG model to optimize latency performance, the reward function was restructured. This new reward function dynamically adjusts the reward value based on the relationship between actual latency and target latency, guiding the model toward latency reduction.
[0034] Reward function reconstruction: , , , in, is the delay-related reward value; is the actual measured delay; It is a custom delay tolerance threshold.
[0035] The upper and lower layer strategies achieve overall performance improvement through joint optimization. The upper and lower layers share a joint reward function that comprehensively considers latency, power consumption, and data rate: , in, It is the objective function value of the joint optimization of the upper and lower layer strategies.
[0036] The upper and lower layer strategies are collaboratively updated to gradually optimize the overall system performance. The upper layer updates the power policy based on the joint reward function, and the lower layer updates the action based on the joint reward function. The upper layer transmits the power policy downstream, and the lower layer feeds back latency predictions to the upper layer. Convergence of the joint reward function is checked. If so, iterations are terminated; otherwise, optimization continues.
[0037] This paper uses a hybrid game-reinforcement learning two-layer framework to pre-compensate tunnel entrance communication delay and obtain an optimal power allocation strategy. The decision results of the upper and lower layers are then applied to the actual system, the pre-compensation strategy is executed, and the actual end-to-end delay is measured. Finally, a reward value is calculated based on the measured delay, and the DDPG (Deep Deterministic Policy Gradient) model is updated through an online learning mechanism.
[0038] Step 2: Continuous compensation in the tunnel.
[0039] First, a hybrid network architecture of the Internet of Vehicles and fixed relays is constructed. Then, a link selection algorithm based on multi-attribute decision-making is designed. The quality of different links is evaluated by defining a decision matrix, and the entropy weight method is used to determine the weight of each attribute. , and calculate the comprehensive evaluation value of each plan , in order to select the optimal link and design and solve the joint optimization strategy to achieve real-time compensation.
[0040] Construct a hybrid network architecture of VANET (Vehicular ad-hoc network) and fixed relay, and design a link selection algorithm based on multiple attribute decision making (MADM). Define the decision matrix , where the attribute set includes received signal strength (RSSI), link delay , bit error rate BER and other indicators.
[0041] Optimal Power Strategy and optimal spectrum allocation strategy Initialize RSSI. , in, is the channel gain from node i to j, is the error term, is the received signal strength from node i to node j.
[0042] , in, is the comprehensive evaluation value of the link from node i to node j in multi-attribute decision making, is the link delay from node i to node j, is the bit error rate from node i to node j.
[0043] Latency and bit error rate Through the optimal spectrum allocation strategy Mapped to specific channel characteristics.
[0044] The entropy weight method is used to determine the weight of each attribute , , Where S is the total number of attributes considered in link selection, is the entropy value of the j-th attribute, and the calculation formula is: , Where m is the total number of vehicles, is the normalized value of vehicle i on attribute j, is the normalized value of other vehicle k on attribute j.
[0045] Calculate the comprehensive evaluation value of each plan , , in, is the normalized maximum value of all vehicles on attribute j, is the normalized minimum value of all vehicles on attribute j.
[0046] The link with the highest score is selected through comprehensive evaluation value, and then distributed ADMM is used for control and continuous compensation of delay, and the global optimal power strategy is adopted in real time. and global optimal spectrum strategy Optimize vehicle power strategy and spectrum allocation strategy to achieve efficient delay compensation and resource allocation.
[0047] The objective function is: , , , in, is the time delay of vehicle i in this period; is the action of vehicle i during this period; It is delay The weight coefficient of It is power The weight coefficient of 、 is the regularization coefficient, balancing the deviation between the current strategy and the optimal strategy, is the maximum transmit power of the vehicle.
[0048] Step 3: Tunnel exit delay recovery.
[0049] First, in the time-space benchmark unification algorithm, a coordinate transformation method based on Lie group is proposed, the SE(3) transformation matrix is defined, and the nonlinear Kalman filter is used to compensate for the cumulative error of INS. The update formula of the predicted covariance matrix is used to improve the positioning accuracy, accurately locate the vehicle status and communication delay, ensure the high-priority transmission of emergency messages and the accurate unification of time-space benchmarks, and optimize the delay recovery strategy.
[0050] First, to ensure the accurate and unified time and space reference of the vehicle at the tunnel exit, the local coordinate system of the vehicle is converted to the WGS84 coordinate system. Then the strategy start distance is determined based on the tunnel length L and the maximum speed (speed limit) of the vehicle. , vehicle i's current delay , calculate the starting distance : , Through real-time compensation in the tunnel, the current channel quality, delay and bit error rate communication indicators are mastered, and the communication status at the tunnel exit is evaluated. According to the current communication and positioning status, the real-time strategy is dynamically adjusted. , , so that it gradually approaches the optimal strategy , , so as to minimize the communication delay.
[0051] , in, It is at the moment Real-time strategy when It is at the moment Real-time strategy when 、 It is at the moment Real-time power strategy and spectrum allocation strategy; 、 It is the optimal power strategy and spectrum allocation strategy; 、 is the learning rate, which controls the speed of policy updates.
[0052] Step 4: Cross-tunnel cloud collaboration.
[0053] First, in the iterative learning of digital twins, a distributed update framework based on federated learning is designed. Each tunnel entrance and exit edge node maintains the optimal power strategy and spectrum allocation strategies and delay , and train the model based on local data. The cloud is responsible for performing model aggregation and calculating the global model , coordinating cloud and edge nodes. Secondly, within the dynamic resource allocation mechanism, a resource optimization model for device-edge-cloud collaboration is established. By employing the Lyapunov optimization framework, long-term constraints are transformed into real-time decisions, enabling joint optimization of communication and computing resources while reducing energy consumption and meeting real-time requirements.
[0054] A distributed update framework based on federated learning. In the iterative learning of digital twins, a distributed update framework based on federated learning is designed. Each tunnel entrance and exit edge node maintains the optimal power strategy , spectrum allocation strategy And the corresponding delay , and perform model training based on local data.
[0055] , in, is the loss function; are the federated learning model parameters, are the local model parameters of node i.
[0056] The cloud is responsible for performing model aggregation and calculating the global model for: , in, is the data volume of tunnel exit / entrance node i, n is the total data volume of all nodes, n is the total data volume of all nodes, is the local model parameter of node i, and N is the total number of tunnel entrance and exit nodes of the expressway tunnel group.
[0057] The cloud will be the global model Distribute to each edge node for the next round of training.
[0058] Dynamic resource allocation mechanism. Lyapunov optimization framework is used to transform long-term constraints into real-time decisions, achieving joint optimization of communication and computing resources. A resource optimization model for device-edge-cloud collaboration is established: Optimization goal: , The constraints are: , in, is the communication energy consumption of the i-th node; is the computational energy consumption of the i-th node; is the total delay; is the maximum allowed delay.
[0059] Through the Lyapunov optimization framework, long-term constraints are transformed into real-time decisions, achieving joint optimization of communication and computing resources.
[0060] Step 5: Closed-loop optimization of multiple tunnel groups.
[0061] First, establish the adjacent tunnel policy transmission rules. Second, establish a fault self-healing mechanism. Through the anomaly detection model, the system status is monitored in real time. When the Mahalanobis distance threshold is reached, the compensation strategy is reconfigured. After the anomaly is detected, the strategy is automatically adjusted to ensure that the current delay status is always maintained at the optimal strategy.
[0062] Using the global model formed in step 4 , establish policy transfer rules to ensure that the policy of each tunnel entrance is initialized to the optimal policy of the previous tunnel exit, maintaining the continuity and consistency of the policy.
[0063] After exiting the tunnel, the strategy is updated to the optimal strategy, the single tunnel process ends, and the optimal strategy of the current tunnel exit is updated. and spectrum allocation strategies Pass to next tunnel entrance: , , And implement the fault self-healing mechanism in it, monitor the system status in real time through the anomaly detection model, and use the global model to detect an anomaly. and real-time communication status x , dynamically adjust system parameters.
[0064] , Here, T is the transpose symbol.
[0065] Real-time monitoring of system status x , calculate the Mahalanobis distance ,when When an anomaly occurs, compensation strategy reconstruction is triggered: system parameters (power allocation, spectrum allocation, routing selection, etc.) are dynamically adjusted according to the anomaly type and severity to restore the normal operation of the system and ultimately achieve closed-loop optimization of multiple tunnel groups.
[0066] , in, is the learning rate, which is used to control the speed of adjustment; is the system parameter that needs to be adjusted at time t, is the system parameter that needs to be adjusted at time t+1.
[0067] Example 2 Based on the high-speed road multi-tunnel group V2X communication delay compensation method based on the dynamic game-reinforcement learning dual-layer framework described in Example 1, this embodiment provides a high-speed road multi-tunnel group V2X communication delay compensation system based on the dynamic game-reinforcement learning dual-layer framework, including a storage medium and a processor; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the high-speed road multi-tunnel group V2X communication delay compensation method based on the dynamic game-reinforcement learning dual-layer framework described in Example 1.
[0068] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for compensating V2X communication delay in a high-speed road tunnel group, characterized in that: include: In response to the vehicle about to enter a tunnel, executing a set tunnel entrance communication delay pre-compensation strategy; In response to the vehicle traveling in the tunnel, executing a set tunnel continuous compensation strategy; In response to the vehicle exiting the tunnel, executing the set tunnel exit delay recovery strategy; In response to a vehicle entering a tunnel group, the set cross-tunnel cloud-based collaborative strategy and tunnel group closed-loop optimization strategy are executed.
2. The V2X communication delay compensation method for a high-speed road tunnel group according to claim 1 is characterized in that: The tunnel entrance communication delay pre-compensation strategy is implemented through a hybrid game-reinforcement learning two-layer framework, including: At the lower layer, the DDPG model is used to make spectrum resource allocation decisions based on the current channel status; At the upper level, the power strategy of the vehicle alliance is updated and optimized through game theory methods to obtain the optimal power allocation decision; Apply the upper and lower layer decision results to the actual system, execute the pre-compensation strategy, and measure the actual end-to-end delay; The reward value is calculated based on the actual measured delay, and the DDPG model is updated through the online learning mechanism, thereby achieving effective pre-compensation for the tunnel entrance communication delay.
3. The V2X communication delay compensation method for a high-speed road tunnel group according to claim 2 is characterized in that: At the lower level, the DDPG model optimizes the spectrum allocation strategy by predicting the delay value to maximize the reward value R: , in, is the weight of the delay; is the predicted delay value; is the optimal state of vehicle i, i.e., the spectrum allocation strategy; is the optimal action of vehicle i, i.e., the power allocation strategy; It is the whole system strategy; is the power consumption of the ith vehicle, is the power consumption weight of the i-th vehicle, is the mathematical expectation; By reconstructing the reward function and dynamically adjusting the reward value, the model can be optimized towards reducing latency. , , , in, is the delay-related reward value; is the actual measured delay; It is a custom delay tolerance threshold; At the upper layer, a dynamic asymmetric game modeling tunnel group communication is introduced, and the participants are defined as each vehicle i in the vehicle alliance; the strategy space is the power strategy of each vehicle and spectrum allocation strategies ;The data transmission rate is ; The utility function of each vehicle is : , in, is the weight of the data rate; is the end-to-end delay of vehicle i; and are the power and spectrum strategies of vehicle i, respectively; and are the power and spectrum strategies of other vehicles respectively; for each vehicle i, in the strategies of other vehicles When fixed, is the optimal strategy: , in, is the optimal power strategy set of other vehicles at equilibrium, is the optimal spectrum strategy set of other vehicles in equilibrium; each vehicle updates its strategy through gradient ascent: , , in, is the learning rate, is the updated power strategy of vehicle i at the k+1th iteration, is the updated spectrum strategy of vehicle i in the k+1th iteration, 、 They represent the power strategy set and spectrum strategy set of all other vehicles at the kth iteration respectively.
4. The V2X communication delay compensation method for a high-speed road tunnel group according to claim 3 is characterized in that: The continuous compensation strategy in the tunnel includes: Build a hybrid network architecture of VANET and fixed relay, and design a link selection algorithm based on multi-attribute decision-making; The entropy weight method is used to determine the weight of each attribute , , Where S is the total number of attributes considered in link selection, is the entropy value of the j-th attribute, , Where m is the total number of vehicles, is the normalized value of vehicle i on attribute j, is the normalized value of other vehicle k on attribute j; Calculate the comprehensive evaluation value of each plan , , in, is the normalized maximum value of all vehicles on attribute j, is the normalized minimum value of all vehicles on attribute j; The link with the highest score is selected through comprehensive evaluation value, and then distributed ADMM is used for control and continuous compensation of delay, and the global optimal power strategy is adopted in real time. and global optimal spectrum strategy Optimize vehicle power strategy and spectrum allocation strategy, the objective function is: , , , in, It is delay The weight coefficient of It is power The weight coefficient of 、 is the regularization coefficient, balancing the deviation between the current strategy and the optimal strategy, is the maximum transmit power of the vehicle.
5. The V2X communication delay compensation method for a high-speed road tunnel group according to claim 4 is characterized in that: The tunnel exit delay recovery strategy includes: proposing a coordinate transformation method based on the Lie group in the space-time benchmark unification algorithm, defining the SE(3) transformation matrix, using nonlinear Kalman filtering to compensate for the cumulative error of the INS, improving the positioning accuracy by predicting the update formula of the covariance matrix, and ensuring the high-priority transmission of emergency messages and the accurate unification of the space-time benchmark.
6. The V2X communication delay compensation method for a high-speed road tunnel group according to claim 5 is characterized in that: To ensure the accurate and unified time and space reference of the vehicle at the tunnel exit, the local coordinate system of the vehicle is converted to the WGS84 coordinate system; then the strategy starting distance is determined, based on the tunnel length L and the maximum speed of the vehicle. , vehicle i's current delay , calculate the starting distance : , Dynamically adjust real-time strategies based on current communication and positioning status , , so that it gradually approaches the optimal strategy , , so as to minimize the communication delay, , in, It is at the moment Real-time strategy when It is at the moment Real-time strategy when 、 It is at the moment Real-time power strategy and spectrum allocation strategy; 、 It is the optimal power strategy and spectrum allocation strategy; 、 is the learning rate, which controls the speed of policy updates.
7. The V2X communication delay compensation method for a high-speed road tunnel group according to claim 6 is characterized in that: The cross-tunnel cloud collaboration strategy includes: First, in the iterative learning of digital twins, a distributed update framework based on federated learning is designed; each tunnel entrance and exit edge node maintains the optimal power strategy and spectrum allocation strategies and delay , and perform model training based on local data; the cloud is responsible for performing model aggregation and calculating the global model ,coordinate cloud and edge nodes; Secondly, in the dynamic resource allocation mechanism, a resource optimization model for end-edge-cloud collaboration is established; by adopting the Lyapunov optimization framework, long-term constraints are converted into real-time decisions to achieve joint optimization of communication and computing resources, while reducing energy consumption and meeting real-time requirements.
8. The V2X communication delay compensation method for a high-speed road tunnel group according to claim 7 is characterized in that: Global Model for: , in, is the data volume of tunnel exit / entry node i, n is the total data volume of all nodes, are the local model parameters of node i, The resource optimization model for device-edge-cloud collaboration has the following optimization goals: , The constraints are: , in, is the communication energy consumption of the i-th node; is the computational energy consumption of the i-th node; is the total delay; is the maximum allowed delay.
9. The V2X communication delay compensation method for a high-speed road tunnel group according to claim 8, characterized in that: The tunnel group closed-loop optimization strategy includes: First, establish the adjacent tunnel strategy transfer rule, update the strategy to the optimal strategy after exiting the tunnel, end the single tunnel process, and set the optimal power strategy of the current tunnel exit to and spectrum allocation strategies Pass to the next tunnel entrance; Secondly, a fault self-healing mechanism is established to monitor the system status in real time through the anomaly detection model. When the Mahalanobis distance threshold is reached, the compensation strategy reconstruction is triggered and automatically adjusted after an anomaly is detected, so that the current delay status is always maintained at the optimal strategy.
10. A V2X communication delay compensation system for a high-speed road tunnel group, characterized in that: including storage media and processors; The storage medium is used to store instructions; The processor is used to operate according to the instruction to execute the high-speed road tunnel group V2X communication delay compensation method described in any one of claims 1 to 9.
Citation Information
Patent Citations
Internet of vehicles calculation unloading and power optimization method based on potential game
CN115052262A
Automatic driving vehicle predictive safety control method and system based on vehicle infrastructure cooperation
CN116714579A
V2X terminal self-compensation time synchronization method
CN118764943A
Method of optimizing dependent task offloading in internet of vehicles using deep reinforcement learning
US20250165311A1