Unmanned aerial vehicle assisted Internet of Vehicles congestion control method based on branch dual-depth Q network
By using a branched dual-deep Q-network architecture and drone assistance, combined with traditional and deep learning algorithms, adaptive congestion control of vehicle-to-everything (V2X) networks in highly dynamic environments was achieved. This solved the problems of poor adaptability of traditional algorithms and high computational cost of DRL, and improved network resource utilization and transmission stability.
Patent Information
- Application Number
- CN202511356259.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2026-01-06
AI Technical Summary
Existing vehicle-to-everything (V2X) congestion control technologies are ill-suited to the challenges of highly dynamic topologies and sudden traffic surges. Traditional heuristic algorithms struggle to adapt, while DRL-based algorithms suffer from high computational costs and low decision-making efficiency. Unmanned aerial vehicle (UAV)-assisted solutions lack rapid response and traffic management capabilities, making it difficult to meet the demands of V2X transmission stability, real-time performance, and large-scale application requirements.
A branched dual-depth Q network (BD3QN) architecture is adopted, combining the traditional TCP Cubic algorithm with the BD3QN algorithm. The network status is perceived in real time through the Monitor module, a progressive weight migration strategy is used to switch the algorithm, and UAVs are introduced as mobile relay nodes to build temporary communication links for traffic splitting, so as to realize parallel processing and optimization decision-making of vehicle transmission power and associated nodes.
It significantly improves the adaptability and control precision of vehicle-to-everything (V2X) networks in highly dynamic environments, enables on-demand allocation of computing resources and precise management of network status, reduces transmission latency and packet loss rate, improves network resource utilization, and meets the real-time and stability requirements of V2X networks.
Smart Images

Figure CN121284635A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of vehicle-to-everything (V2X) communication technology, specifically relating to a UAV-assisted V2X congestion control method based on a branched dual-depth Q-network. Background Technology
[0002] In intelligent transportation systems, the Internet of Vehicles (IoV) serves as a core technology, building communication bridges between vehicles (V2V) and between vehicles and infrastructure (V2I), providing crucial support for improving traffic efficiency and ensuring driving safety. However, with the continuous improvement of vehicle intelligence, the communication demands of IoV are experiencing explosive growth, and data transmission volume is surging dramatically. Network congestion is gradually becoming a core obstacle restricting its performance optimization and large-scale application. In the IoV environment, the high dynamism of vehicle movement leads to rapidly changing network topology, making it difficult to maintain the stability of communication links and causing frequent interruptions and failures in data transmission paths. Existing technologies, unable to effectively cope with these complex scenarios, are seriously hindering the further development and widespread implementation of IoV.
[0003] To provide more stable and efficient transmission services in vehicular network (V2N) scenarios, a feasible solution is to deploy traditional heuristic congestion control algorithms (such as TCPNew Reno, Bic, and Cubic) at base stations, enabling centralized decision-making. While heuristic algorithms are easy to implement and deploy, they suffer from problems like misjudgment of network resources and low resource utilization in the highly dynamic V2N environment. For example, packet loss caused by non-congestion factors such as signal blocking or link interruption may be misjudged as network congestion, leading to inappropriate adjustments to transmission strategies and reduced resource utilization. Other research employs deep reinforcement learning (DRL) techniques (such as Vivace and ACC algorithms) for transmission rate adjustment. While these methods can respond to dynamic network environments in real time, they suffer from high computational costs and slow model convergence in V2N scenarios, easily leading to decision lag and limiting their large-scale application. Both of the aforementioned algorithms can alleviate congestion to some extent in normal scenarios, but they lack the ability to quickly manage sudden flows (such as a sudden surge of more than 5 times in traffic flow in a local area caused by a traffic accident). This has become a core technical shortcoming that restricts the Internet of Vehicles from responding to sudden communication needs.
[0004] To address link congestion caused by sudden surges in traffic in localized areas, technologies have introduced drone-assisted communication solutions to divert traffic. In recent years, drones have been widely used in the field of connected vehicles (V2V) due to their maneuverability, rapid deployment, and ability to establish line-of-sight communication links. When network congestion occurs, drones equipped with communication modules can flexibly fly over congested areas, dynamically adjust their hovering positions, and quickly establish temporary communication links to divert localized data traffic, thereby effectively alleviating congestion. However, it is important to note that both vehicle-mounted terminal devices and drones in V2V generally face limitations in computing power, storage, and energy consumption. Therefore, designing congestion control methods that balance transmission performance and resource consumption to adapt to the dynamically changing network environment of V2V is crucial for ensuring end-to-end transmission performance in drone-assisted V2V.
[0005] Chinese invention patent CN118283700A, entitled "A Method for Channel Congestion Control in Vehicle-to-Everything (V2X) Based on Multi-Agent Deep Reinforcement Learning," addresses the channel congestion problem in V2X by constructing an LTE-V2X Mode 4 communication and channel model and training agents using the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm. The agents dynamically adjust transmission power and transmission intervals based on channel busy rate, service priority, and packet delivery rate, optimizing the strategy through a reward function with weighted and penalized terms to ultimately improve packet delivery rate during channel congestion and ensure decision-making fairness. However, this solution has the following shortcomings: First, it relies solely on a multi-agent deep reinforcement learning model, requiring significant computational resources and time for training, while V2X onboard devices have limited computing power and energy consumption, making model deployment and adaptation difficult. Second, it cannot accurately handle sudden traffic surges, lacking rapid response and mitigation capabilities when faced with sudden increases in local link load (such as traffic spikes caused by traffic accidents).
[0006] Chinese invention patent CN116260773B, entitled "A Congestion Control Method and Related Equipment," achieves control through congestion information feedback between the sending end, network device, and receiving end. The sending end sends a message containing initial congestion information; the network device determines the congestion status of the current hop and integrates the information; the receiving end feeds back the integrated congestion information to the sending end, which then adjusts its transmission strategy accordingly to reduce bandwidth usage and assist in path optimization. However, this solution has two main drawbacks: First, it has poor adaptability. When vehicles rapidly switch links, the congestion information from the previous hop has not yet been fed back, and the sending end has already used outdated information for route selection, easily leading to link interruption or secondary congestion. Second, it is prone to misjudgment. When vehicles enter tunnels or densely populated areas with tall buildings, wireless channel fading causes a sharp increase in latency, which the network device may misjudge as congestion and accumulate the congestion hop count. The sending end may therefore reduce its transmission rate, while the actual network load still has remaining bandwidth, resulting in resource waste.
[0007] In summary, among existing vehicle-to-everything (V2X) congestion control technologies, traditional heuristic algorithms, while easy to deploy, struggle to adapt to highly dynamic topologies and sudden traffic surges. DRL-based algorithms can respond to dynamic environments but are limited by computational costs and decision-making efficiency. Meanwhile, drone-assisted solutions have not yet achieved collaborative adaptation across "algorithm-device-scenario." All three have their own technical shortcomings and lack effective integration, making it difficult to address the comprehensive challenges of the dynamic complexity, frequent sudden traffic surges, and resource constraints inherent in V2X. Therefore, there is an urgent need for an integrated congestion control method that can perceive network status in real time, flexibly adapt algorithm strategies, and coordinate with drones for traffic diversion. This method aims to overcome existing technological bottlenecks and meet the V2X requirements for transmission stability, real-time performance, and large-scale application. Summary of the Invention
[0008] The purpose of this invention is to provide a UAV-assisted vehicle-to-everything (V2X) congestion control method based on a branched dual-depth Q-network, which improves the network's adaptability to high-speed vehicle movement and frequent topology changes, ensures the continuous stability and low latency of the communication link, and improves network resource utilization and decision-making efficiency, so as to maintain high-quality data transmission services even in the event of a sudden surge in traffic.
[0009] To achieve the above objectives, the technical solution adopted by this invention is: a UAV-assisted vehicle-to-everything (V2X) congestion control method based on branched dual-depth Q-networks, comprising the following steps:
[0010] S1. Deploy a Monitor module at the base station of the vehicle network. This module collects network status data in the vehicle network in real time. The network status data includes at least packet loss rate and round-trip time (RTT). Calculate the congestion threshold based on the network status data to determine whether the network is in a congested state.
[0011] S2. Construct a branched dual-depth Q network (BD3QN) architecture. This architecture includes a shared module and a branch module. The shared module extracts features from the network state information collected and preprocessed by the Monitor module. The branch module decomposes the action space into sub-action branches that are independent of power regulation and association decision-making, thereby realizing parallel processing of vehicle transmission power adjustment and vehicle-base station or UAV association decision-making.
[0012] S3. Based on the network congestion status determined by the Monitor module, adaptively switch between the TCP Cubic algorithm and the BD3QN algorithm: When the network is in a non-congested or lightly loaded state, the TCP Cubic algorithm is called to adapt to the network state by adjusting the congestion window size; when the network is in a congested state, a gradual weight migration strategy is adopted to complete the smooth switch from the TCP Cubic algorithm to the BD3QN algorithm in stages, and the BD3QN algorithm performs joint optimization decisions on the vehicle's transmission power and communication associated nodes.
[0013] S4. Introduce drones as mobile relay nodes. When the Monitor module detects network congestion in a local area, it dispatches drones to fly over the congested area to establish temporary communication links. Based on the output of the BD3QN algorithm, it dynamically adjusts the drone's position and communication links to divert local data traffic.
[0014] Further, in step S1, the condition for determining the congestion threshold is: when the Monitor module detects that the network packet loss rate exceeds the first set threshold, or the RTT exceeds the second set threshold, the network is determined to be in a congested state.
[0015] Further, in step S2, the sharing module includes three fully connected layers fc1, fc2, and fc3. The input layer receives the preprocessed network state information and extracts 128-dimensional abstract features through fc1, fc2, and fc3 in sequence. Each sub-action branch of the branch module is a multilayer perceptron structure, which receives the 128-dimensional features output by the sharing module as input, first maps them to 256-dimensional features, and then maps them to the action space corresponding to each action dimension.
[0016] Furthermore, in step S3, the progressive weight migration strategy includes three stages: a trial period, an optimization period, and a decision period. During the trial period, the TCP Cubic algorithm is used as the main control logic, retaining its congestion window adjustment mechanism, while the BD3QN algorithm only optimizes the vehicle transmission power parameters. During the optimization period, the decision weight of the BD3QN algorithm is increased compared to the trial period, while the decision weight of the TCP Cubic algorithm is decreased. The BD3QN algorithm simultaneously analyzes the rate and packet loss information of the transport layer, constructs a two-dimensional decision space of transmission power and network congestion, and dynamically adjusts the vehicle transmission power and associated nodes. During the decision period, when the BD3QN algorithm learns enough network state-action feedback samples and the decision stability meets the preset requirements, it switches to the BD3QN fully autonomous decision-making state.
[0017] Further, in step S3, the BD3QN algorithm includes the following sub-steps:
[0018] S41. Decompose the action space into power regulation branch and correlation decision branch;
[0019] S42. Extract environmental state features through the shared module;
[0020] S43. Each branch network outputs the optimal power adjustment action and associated node selection action respectively;
[0021] S44. Select the optimal action by fusing the advantage function and the state value function.
[0022] Furthermore, the reward function of the BD3QN algorithm is designed to minimize the total system transmission delay and satisfies the following two constraints:
[0023] C1:
[0024] C2:
[0025] Where C1 represents the vehicle association uniqueness constraint, ensuring that each vehicle can only be uniquely associated with one communication node at any given time; C2 represents the vehicle power control factor constraint, whereby the vehicle's power control factor must be within the interval [0,1] and discretized into K equally spaced values; Let be the indicator function; if vehicle v is associated with the nth communication node at time t, then otherwise N is the total number of communication nodes; Represents a set of vehicles; β v (t) represents the power control factor of vehicle v at time t.
[0026] Furthermore, when the vehicle communicates with the base station, the transmission delay T between the two is... v r The formula for calculating (t) is:
[0027]
[0028] In the formula, D represents the transmission delay when data is transmitted between vehicle v and base station r at time t; v (t) is the amount of data that vehicle v needs to transmit at time t; It is the spectral bandwidth of the communication channel between vehicle v and base station r; It is the signal-to-interference-plus-noise ratio (SINR) when vehicle v communicates with base station r at time t. Indicates a collection of vehicles;
[0029] Signal-to-interference-plus-noise ratio when vehicle v communicates with base station r The calculation formula is:
[0030]
[0031] In the formula, Let be the channel fading coefficient from vehicle v to base station r; This represents the channel state between vehicle v and base station r, mainly considering wireless network path loss factors, and is calculated based on the free space path loss model. β is the maximum transmission power of vehicle v; v (t) is the power control factor for the signal transmitted by vehicle v; σ R (t) represents the interference caused by other vehicles within the same base station to the communication link between vehicle v and base station r; N0 is the power spectral density of additive white Gaussian noise, which is determined according to the system environment.
[0032] Furthermore, in step S4, when the drone acts as a relay node, the formula for calculating the transmission delay between it and the vehicle is as follows:
[0033]
[0034] In the formula, This represents the transmission delay when vehicle v transmits data to drone u at time t; D is the spectral bandwidth of the communication channel between vehicle v and drone u; v (t) is the amount of data that vehicle v needs to transmit at time t; Let V be the signal-to-interference-plus-noise ratio (SIR) when vehicle v communicates with drone u at time t. Indicates a collection of vehicles. Indicates a collection of drones;
[0035] Signal-to-interference-plus-noise ratio when vehicle v communicates with drone u The calculation formula is:
[0036]
[0037] In the formula, This represents the channel state between vehicle v and drone u, which is calculated based on the free space path loss model. β is the channel fading coefficient from vehicle v to drone u; v (t) is the power control factor for the signal transmitted by vehicle v; σ is the maximum transmission power of vehicle v; U (t) represents the interference caused to drone u by other vehicles within the same drone u; σ UO (t) represents the interference caused by other drones to the communication link between vehicle v and drone u; N0 is the power spectral density of additive white Gaussian noise.
[0038] Furthermore, the training and optimization process of the BD3QN algorithm includes: storing the state transition samples generated by the interaction between the vehicle and the communication node into the experience replay buffer according to priority; the state transition samples include the current network state, the action performed, the reward obtained, and the next network state; at each preset fixed training step interval, the parameters of the main network are synchronously copied to the target network, and a loss function is designed based on the output difference between the main network and the target network, and the parameters of the main network are updated through the backpropagation algorithm.
[0039] Furthermore, the method also includes a network recovery phase: when the Monitor detects that the network state has recovered to a lightly loaded or stable state, it first stops the BD3QN exploration process and freezes the current parameters of the BD3QN model; then, it adopts a hybrid strategy linear transition method to gradually switch the control logic back to TCP Cubic, and the decision weight of TCP Cubic is gradually increased to 100%, thus completing the network recovery.
[0040] The beneficial effects of the above scheme are as follows:
[0041] 1. This invention significantly improves the adaptive capability and control precision of vehicle-to-everything (V2X) networks in highly dynamic environments through the organic integration of multiple technological innovations. By using a Monitor module to perceive network status in real time and employing a three-stage progressive weight migration strategy—"exploration phase - optimization phase - decision phase"—it achieves a smooth and stable switch between the traditional TCP Cubic algorithm and the intelligent BD3QN algorithm. This method fully utilizes the advantages of traditional algorithms—low computational overhead and fast response—under light loads, avoiding decision oscillations, while precisely activating deep learning models for intervention when congestion risks occur. This achieves on-demand allocation of computing resources and precise control of network status.
[0042] 2. This invention successfully addresses the challenges of decision complexity and real-time performance in high-dimensional state spaces through an innovative Branched Dual-Deep Q-Network (BD3QN) architecture. This architecture decouples complex joint decision-making actions, such as vehicle transmit power control and base station / UAV association selection, into multiple parallel sub-action branches. This transforms the action space from an exponentially growing size with the number of vehicles to a linear one, fundamentally overcoming the "curse of dimensionality." This enables the model to make decisions with lower computational overhead and shorter latency, significantly improving decision efficiency and fully meeting the stringent millisecond-level response requirements of vehicle-to-everything (V2X) networks.
[0043] 3. To address the challenge of sudden traffic surges that traditional networks struggle to handle, this invention introduces drones as mobile relay nodes, significantly enhancing the network's congestion resilience and coverage. When traffic surges in a localized area due to a traffic accident, drones can quickly reach the hotspot, establishing high-quality line-of-sight communication links and effectively diverting overloaded traffic from ground base stations. This not only alleviates localized congestion but also significantly reduces transmission latency for connected vehicles due to its stable link quality, and can cover communication blind spots such as tunnel exits, ensuring the continuity of network services.
[0044] 4. This invention optimizes the overall system performance through the deep synergy of an "intelligent switching mechanism, efficient BD3QN algorithm, and UAV dynamic assistance." This method effectively reduces network conflicts and packet loss through precise power control, link selection, and resource allocation, maximizing the utilization of spectrum resources. This results in an overall increase in network throughput, reduced average data transmission latency and jitter, and maximizes the utilization of vehicle network resources, demonstrating high practical value and comprehensive advantages. Attached Figure Description
[0045] Figure 1 This is a diagram of the congestion control architecture of the present invention;
[0046] Figure 2 This describes the workflow of the monitor in an embodiment of the present invention. Detailed Implementation
[0047] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0048] It should be noted that, unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.
[0049] This embodiment proposes a UAV-assisted vehicular network congestion control method based on branching double deep Q network (BD3QN-CC). Addressing the problem of data transmission congestion caused by the frequent changes in network topology and the need for continuous link reconnection due to the high dynamism of vehicles in vehicular networks, traditional congestion control algorithms are unable to adjust in a timely manner. This invention utilizes an intelligent monitor to perceive network load in real time and make policy decisions. It leverages the lightweight nature of traditional congestion control methods and the flexible adjustment advantages of deep reinforcement learning (DRL) congestion control methods, adaptively switching between the BD3QN algorithm and traditional heuristic congestion control algorithms. Simultaneously, it introduces UAV-assisted construction of temporary communication links for traffic diversion, achieving intelligent congestion adjustment.
[0050] A UAV-assisted vehicle-to-everything (V2X) congestion control method based on branched dual-depth Q-networks includes the following steps:
[0051] S1. Deploy a Monitor module at the base station of the vehicle network. This module collects network status data in the vehicle network in real time. The network status data includes at least packet loss rate and round-trip time (RTT). Calculate the congestion threshold based on the network status data to determine whether the network is in a congested state.
[0052] S2. Construct a branched dual-depth Q network (BD3QN) architecture. This architecture includes a shared module and a branch module. The shared module extracts features from the network state information collected and preprocessed by the Monitor module. The branch module decomposes the action space into sub-action branches that are independent of power regulation and association decision-making, thereby realizing parallel processing of vehicle transmission power adjustment and vehicle-base station or UAV association decision-making.
[0053] S3. Based on the network congestion status determined by the Monitor module, adaptively switch between the TCP Cubic algorithm and the BD3QN algorithm: When the network is in a non-congested or lightly loaded state, the TCP Cubic algorithm is called to adapt to the network state by adjusting the congestion window size; when the network is in a congested state, a gradual weight migration strategy is adopted to complete the smooth switch from the TCP Cubic algorithm to the BD3QN algorithm in stages, and the BD3QN algorithm performs joint optimization decisions on the vehicle's transmission power and communication associated nodes.
[0054] S4. Introduce drones as mobile relay nodes. When the Monitor module detects network congestion in a local area, it dispatches drones to fly over the congested area to establish temporary communication links. Based on the output of the BD3QN algorithm, it dynamically adjusts the drone's position and communication links to divert local data traffic.
[0055] The implementation process of the present invention will be described in detail below with reference to the accompanying drawings:
[0056] like Figure 1 The diagram shows the architecture of the BD3QN-CC method, which mainly achieves congestion control through two-stage collaboration: the first stage is a load awareness and intelligent policy switching mechanism based on a smart monitor; the second stage is a transmission power adaptive adjustment mechanism based on a hybrid policy.
[0057] Figure 1 The upper section presents the load awareness and intelligent policy switching mechanism based on the intelligent monitor. As a key module for awareness and intelligent policy switching, the intelligent monitor can collect multi-dimensional network information such as packet loss count and round-trip time (RTT) in real time. By calculating the congestion threshold, it can accurately assess the degree of network congestion, thereby driving a gradual and smooth switching between the TCP Cubic algorithm and the BD3QN algorithm.
[0058] Figure 1The following section demonstrates a hybrid strategy-based adaptive power adjustment mechanism, primarily composed of the TCPCubic and BD3QN algorithms. During system operation, the intelligent monitor selects a strategy based on collected network status information: If the TCPCubic algorithm is enabled, the system initially increases the congestion window exponentially during the slow start phase, then switches to linear growth once a threshold is reached; when congestion occurs, the congestion window is reduced to decrease the data packet transmission rate, thereby reducing the amount of data injected into the network. If the BD3QN algorithm is enabled, after the monitor collects network status information, it preprocesses this information through a "shared module," then inputs it into a branch module to calculate multiple sub-branch dimensions. At each sub-branch dimension, the optimal power adjustment action is selected through an argmax operation, thus achieving adaptive adjustment of the vehicle's transmission power.
[0059] The specific design details are as follows:
[0060] (I) Monitor-based load awareness and intelligent policy switching mechanism
[0061] In the Internet of Vehicles (IoV) environment, base stations, as critical infrastructure deployed along roads, serve as hubs for information exchange between vehicles (V2V) and between vehicles and infrastructure (V2I). In this invention, the Monitor primarily operates on the base station. Facing the complex scenario of frequent and dynamic vehicle access, the Monitor first collects multi-dimensional network load data in real time, such as packet loss count and RTT (Round-Trip Time), and constructs a state feature database after normalization. Then, it calculates a congestion threshold to determine if the network is congested. If the network is not congested, the TCP Cubic algorithm is invoked by default to ensure basic data transmission. When the packet loss rate reaches a set threshold, or the RTT exceeds a specific time threshold, the BD3QN algorithm module is activated based on the accumulated state data. Using a "gradual weight migration" strategy, a smooth transition from the TCP Cubic algorithm to the BD3QN algorithm is completed in stages, thereby optimizing data transmission efficiency. Its workflow is as follows: Figure 2 As shown.
[0062] The specific workflow of Monitor is as follows:
[0063] 1. Cold start phase (initial vehicle connection)
[0064] In the context of connected vehicles, when a vehicle enters a network service area from a communication blind spot and completes initial access, the intelligent decision-making model (BD3QN algorithm) lacks local network state data accumulation. If the deep reinforcement learning (DRL) algorithm is directly activated at this time, it is prone to decision oscillations due to the "experience gap". Therefore, the system first defaults to calling the mature TCP Cubic congestion control algorithm, using its classic "slow start-congestion avoidance-fast recovery" logical framework to ensure basic transmission stability at the moment the vehicle connects.
[0065] Simultaneously, the Monitor perception module is activated. This module relies on the base station servers deployed in the vehicle-to-everything (V2X) network to acquire data such as packet loss count and RTT in real time with a collection cycle of T1, and calculates congestion thresholds to provide a basis for subsequent network status assessment.
[0066] 2. Decision-making stage (dynamic network adjustment)
[0067] (1) Congestion triggering conditions
[0068] The Monitor module continuously monitors the network status in two dimensions. When it detects that the packet loss rate reaches the set threshold for the total number of data packets sent, or the RTT exceeds a specific time threshold, it determines that the network has entered a "local / global congestion risk state" and triggers the intervention of the BD3QN intelligent decision module.
[0069] (2) Smooth switching strategy
[0070] To avoid network oscillations caused by traditional instantaneous switching, this invention adopts a smooth switching strategy of "gradual weight migration," completing the control logic transition in three stages: a trial period, an optimization period, and a decision period. Here, two different weights τ1 and τ2 are set (τ1+τ2=1 and τ1<τ2).
[0071] Trial Period: In the initial switchover phase, the period is the same as the Monitor's data collection cycle. The core objective during this phase is to ensure the "continuity" of network transmission, with TCP Cubic still serving as the main control logic (retaining its mature congestion window adjustment mechanism). Simultaneously, BD3QN's "light intervention" mode is introduced—only allowing the intelligent model to optimize transmission power parameters (BD3QN decision weight τ1). During this phase, BD3QN, based on the initial state library accumulated during the cold start phase, combined with real-time packet loss and RTT data, outputs a "power adjustment suggestion value," which is then weighted and fused with TCP Cubic's native congestion window control logic (e.g., TCP Cubic weight allocation is τ1, BD3QN weight allocation is τ2).
[0072] Optimization Phase: As network status data increases, within a set period of 2T1 after the trial period ends, the collaborative optimization phase begins, and the BD3QN decision weights are increased to τ2. At this point, the intelligent model not only analyzes the transmission layer's rate and packet loss information but also constructs a two-dimensional decision space of "transmission power - network congestion." For example, in areas with dense traffic but good channel quality, the transmission power is appropriately increased to maintain high-speed transmission; in areas with severe channel interference, power is preferentially reduced to ensure transmission stability.
[0073] Decision-making phase: Once BD3QN has learned enough network state-action feedback samples and determined that the intelligent model possesses "autonomous decision-making stability," it enters the decision-making phase. At this point, it fully switches to the BD3QN fully autonomous decision-making state. The model dynamically adapts to complex scenarios in the Internet of Vehicles (IoV) such as high-speed vehicle movement, frequent topology changes, and sudden traffic surges, based on the "state perception-value assessment-action output" closed loop of deep reinforcement learning. It adjusts multi-dimensional parameters such as transmission power in real time, enabling congestion control strategies to accurately match dynamic network changes.
[0074] 3. Recovery Phase (Network Status Stabilizes)
[0075] (1) Conditions for recovery
[0076] The Monitor module uses "continuous achievement of dual indicators" as the basis for network stabilization. Through continuous monitoring over period T1, if the packet loss rate is consistently lower than the set threshold or the RTT is lower than a specific time threshold, the network is determined to have recovered from the congested state to the "light load / stable state".
[0077] (2) Strategy rollback and optimization process:
[0078] Intelligent model freezing: To prevent BD3QN from introducing new transmission fluctuations due to continuous exploration (random action attempts) after the network state stabilizes, the exploration process of BD3QN is stopped first, and the current parameters of the model are frozen (the policy network weights at the time of stabilization are saved).
[0079] Gradual rollback to traditional algorithm: A "hybrid strategy linear transition" approach is adopted to gradually switch the control logic back to TCPCubic. Within cycle T1, a "TCP Cubic + BD3QN weighted hybrid strategy" is constructed, with the decision weight of TCP Cubic gradually increasing to 100%, smoothly completing the strategy rollback.
[0080] (II) Transmit Power Adaptive Adjustment Mechanism Based on Hybrid Strategy
[0081] This section focuses on the adaptive transmission power adjustment mechanism based on a hybrid strategy. After the Monitor completes strategy formulation and executes gradual switching, it enters the specific network adjustment phase, which is jointly implemented by the TCP Cubic algorithm and the BD3QN algorithm. During the initial vehicle access phase, the BD3QN algorithm lacks local network state data accumulation, and directly enabling DRL would lead to inaccurate decisions. Therefore, the TCP Cubic algorithm is used by default in the initial phase. The core of this algorithm is to adapt to the network state by adjusting the congestion window. During the slow start phase, the transmission window growth follows the formula:
[0082] W Cubic (t)=C·(TL) 3 +W max .
[0083] Where C is the proportionality coefficient, determined by network parameters such as the bandwidth-delay product; T is the time variable, reflecting the adjustment period of the congestion window; and L is the base time point for window growth, relative to the maximum value W of the previous congestion window. max Related.
[0084] When no network congestion is detected, the TCP Cubic algorithm, using the aforementioned formula, gradually increases the sending window according to a cubic function, ensuring a smooth increase in data transmission volume and adapting to the current network state. However, when the Monitor detects that the packet loss rate reaches a threshold for total transmitted data packets, or that the RTT reaches a specific threshold, it determines that the network has entered a "local / global congestion risk state," triggering the intervention of the BD3QN intelligent decision-making module. At this point, the Monitor dynamically allocates weight ratios based on the network state, gradually switching the TCP Cubic algorithm to the BD3QN algorithm. The BD3QN algorithm, relying on DRL, outputs the vehicle's current transmission power and associated decision-making action vectors, enabling flexible and precise adjustment of the vehicle's transmission power.
[0085] The focus of this invention is the design of the BD3QN algorithm, and the congestion control process of the BD3QN algorithm will be described in detail below.
[0086] 1. Problem Modeling Process
[0087] (1) Network Architecture
[0088] In the BD3QN-CC method of this invention, a vehicle-to-everything (V2X) system consisting of one base station, several drones, and V vehicles is considered. For a system with multiple base stations, it is divided into several subsystems, each containing only one base station, and each subsystem applies the strategy of this invention to solve the congestion control problem. The set of vehicles is represented as... Base stations are represented by r; the set of drones is represented by That is, there are a total of U-frames of drones that can participate in the service in the system; the hybrid set of base stations and drones is represented as The total number of elements in this set is N.
[0089] Vehicle association constraint: Each vehicle can only be associated with one of the following sets within a given time period t. A node, using binary variables This indicates the association between vehicle v and a base station or drone, satisfying the following condition:
[0090]
[0091] (2) Delay Model
[0092] Vehicle-to-everything (V2X) networks are characterized by high-speed vehicle movement and frequent changes in network topology. Node movement alters channel fading characteristics, and as the number of vehicles increases, competition for channel resources intensifies. These factors directly impact transmission time. Transmission delay directly and instantly reflects these dynamic changes, allowing for rapid assessment of the communication link status in a dynamic V2X environment. Transmission delay is the most direct indicator for evaluating communication performance in this dynamic environment. Therefore, this invention primarily uses transmission delay as the core evaluation metric.
[0093] Due to the high speed of vehicle movement, the channel state between the vehicle and the base station / vehicle and the drone changes frequently. While the drone moves, its speed and topology change frequency are much lower than that of the vehicle, and the drone and base station are relatively static, resulting in low transmission delay fluctuations. Therefore, we mainly consider the transmission delay between the vehicle and the base station, and between the vehicle and the drone. Communication between the vehicle and the base station uses Dedicated Short Range Communication (DSRC) technology; communication between the vehicle and the drone is based on a line-of-sight channel, assuming no surrounding buildings obstructing the view. When vehicle v accesses the network, it periodically sends a fixed amount of data to base station r or drone u, sending a fixed amount of D data within each time t. v (t) bytes of data.
[0094] Transmission latency when the vehicle communicates with the base station:
[0095] This mode refers to communication between a vehicle and a base station. When vehicle v sends data to base station r in fixed byte counts, the uplink signal-to-interference-plus-noise ratio (SINNR) received by base station r is... Represented as:
[0096]
[0097] in, Let be the channel fading coefficient from vehicle v to base station r; This represents the channel state between vehicle v and base station r, mainly considering the wireless network path loss factor, and is calculated based on the Free Space Path Loss (FSPL) model. β is the maximum transmission power of vehicle v; v (t) is the power control factor for the signal transmitted by vehicle v; σ R (t) represents the interference caused by other vehicles within the same base station to the communication link between vehicle v and base station r; N0 is the power spectral density of additive white Gaussian noise, which is determined according to the system environment.
[0098] Interference from other vehicles within the same base station σ R The formula for calculating (t) is as follows:
[0099]
[0100] Among them, P i r (t) represents the transmission power of other vehicles i within the same base station to this base station r.
[0101] According to Shannon's theorem, the transmission delay between the vehicle and the base station...
[0102]
[0103] in, It is the spectral bandwidth of the communication channel between vehicle v and base station r.
[0104] Transmission latency when vehicles communicate with drones:
[0105] In a vehicle-to-drone (V2D) communication scenario, when vehicle v sends data to drone u in fixed byte counts, the uplink signal-to-interference-plus-noise ratio (SINNR) received by drone u is... Represented as:
[0106]
[0107] in, This represents the channel state between vehicle v and drone u, which is calculated based on the free space path loss model. β is the channel fading coefficient from vehicle v to drone u; v (t) is the power control factor for the signal transmitted by vehicle v; σ is the maximum transmission power of vehicle v; U (t) represents the interference caused to drone u by other vehicles within the same drone u; σ UO (t) represents the interference caused by other drones to the communication link between vehicle v and drone u; N0 is the power spectral density of additive white Gaussian noise.
[0108] Within the same drone, besides vehicle v, the interference σ generated by other vehicles i on the drone. U The formula for calculating (t) is as follows:
[0109]
[0110] Among them, P i u (t) represents the transmission power of other vehicles i within the same UAV u to the UAV u.
[0111] Interference σ from other drones on the communication link between vehicle v and drone u UO The formula for calculating (t) is as follows:
[0112]
[0113] in, It is the transmission power of vehicle v inside other drones j to drone j.
[0114] Transmission latency between vehicles and drones
[0115]
[0116] in, D is the spectral bandwidth of the communication channel between vehicle v and drone u; v (t) is the amount of data that vehicle v needs to transmit at time t.
[0117] In summary, the total transmission delay for all vehicles communicating with the UAV within the system at time t is...
[0118]
[0119] (3) Problem transformation
[0120] In this invention, the optimization problem is transformed into minimizing the total system delay, and the mathematical expression is:
[0121] min T total (t).
[0122] At the same time, the following restrictions must be met:
[0123] C1: Vehicle association uniqueness constraint
[0124] To ensure that each vehicle can only uniquely associate with one communication node (such as a base station or drone) at any given time, the mathematical expression is:
[0125]
[0126] C2: Vehicle power control factor constraint
[0127] The vehicle's power control factor must be within the interval [0,1] and discretized into K equally spaced values, as expressed mathematically:
[0128]
[0129] 2. BD3QN Algorithm Design
[0130] (1) Markov Modeling
[0131] 1) State Space
[0132] In this invention, all algorithms are deployed on base stations, which need to determine the communication node associations and power control factors for the vehicles they serve. To accurately represent the state of the vehicle-to-everything (V2X) system, the state space is designed to contain three key pieces of information, based on network operation data collected by the Monitor module:
[0133] The packet loss situation of each vehicle at the previous moment, using This indicates the packet loss status of vehicle v at time t;
[0134] The round-trip time (RTT) of each vehicle at the previous moment is used as This represents the RTT state of vehicle v at time t;
[0135] The association decisions between each vehicle and communication node (base station or drone) at the previous moment, using This represents the associated decision state of vehicle v at time t.
[0136] In summary, the state space of a vehicle-to-everything (V2X) system can be represented as:
[0137]
[0138] Where V represents the total number of vehicles served by the base station. For vehicle assembly.
[0139] 2) Action Space
[0140] In this vehicle-to-everything (V2X) system, each vehicle has two executable sub-actions: power adjustment and correlation decision. The overall action space of the system is represented as A(t):
[0141] A(t) = {A1(t),...,A v (t),...,A V (t)}.
[0142] in, A v (t) contains two sub-actions. This represents the index value that establishes an association between the v-th vehicle and the base station r or the drone u. This represents the power control factor when vehicle v sends data to communication node n.
[0143] 3) Reward function
[0144] In this vehicle-to-everything (V2X) system, the optimization objective is to minimize the total transmission latency. Therefore, at time t, the reward function is expressed as:
[0145] R(t) = T total (t).
[0146] Among them, T total (t) represents the total transmission delay of the system at time t. Using this reward function, the algorithm can prioritize actions that minimize the total transmission delay during the decision-making process.
[0147] (2) BD3QN training process
[0148] 1) BD3QN structure
[0149] The core framework of this invention is BD3QN, which mainly consists of a shared module and branch modules. The shared module constructs a shared feature extraction network containing multiple fully connected layers (fc1, fc2, fc3). The input layer receives the preprocessed environmental state and extracts 128-dimensional abstract features step by step through the hidden layers (fc1→fc2→fc3). The branch module establishes sub-action branches for tasks such as transmission power and association decision-making in the system. Each sub-action branch is a multilayer perceptron (MLP) structure, which receives the 128-dimensional features output from the shared module as input, first maps them to 256-dimensional features, and finally maps them to the action space corresponding to each action dimension.
[0150] Based on the design principles of the BD3ON action space, the branching logic for vehicle-related aspects in this invention is as follows: Assuming that each vehicle's power control corresponds to an independent sub-branch, the system needs to construct V power control sub-branches. The output dimension of each sub-branch consists of K discrete power values to adapt to the power factor adjustment requirements. Simultaneously, considering that the base station needs to synchronously decide the association relationship between each vehicle and the base station or drone node, an association decision sub-branch needs to be added for each vehicle. Each association decision sub-branch consists of N possible base station and drone association nodes, resulting in a final number of action branches of 2V.
[0151] For this vehicle-to-everything (V2X) system, if a classic DRL algorithm (such as DQN or DDPG) is used, the action space size for base station decision-making is (N... V ·K VThis would significantly consume base station computing resources and affect decision-making efficiency. BD3QN, however, utilizes a branch network architecture, treating each vehicle's power regulation and related decisions as a sub-branch. Within time t, the base station selects the combination of related decisions and the optimal power value from each action sub-branch, reducing the action space to V·(K+N). Therefore, introducing a branch network effectively improves decision-making efficiency and reduces computing resource consumption in a vehicle-to-everything (V2X) environment. The calculation process for vehicle sub-branch actions will be described below:
[0152] Dominance function calculation:
[0153] In a multi-vehicle cooperative network environment, each vehicle sub-branch calculates its advantage function through a shared network. The core function of the advantage function is to measure the "advantage of choosing a certain action compared to the average action." For vehicle v, the simplified formula for calculating its advantage function is as follows:
[0154]
[0155] Among them, Q θ (S v (t),A v (t) represents the vehicle v in state S. v (t) Select action A v The action value of (t), V θ (S v (t) represents the state of vehicle v in state S. v The state value under (t) is provided by the State Value call of the shared module.
[0156] Q-value fusion and action selection:
[0157] To obtain a more accurate action value estimate, the advantage function and the State Value are fused together, and the fusion calculation formula is as follows:
[0158]
[0159] Determining the optimal action of vehicle v (such as the transmission power and communication association decisions of vehicle v) is achieved by maximizing the fused Q value (argmax operation), as shown in the formula:
[0160]
[0161] Among them, A v * (t) represents the optimal action of vehicle v in the current state. Each vehicle v selects an optimal action value, and the optimal action values of all vehicles together form a 2V-dimensional action vector to complete the final action output, thereby realizing the collaborative control of multiple vehicles.
[0162] Target Q value calculation:
[0163] To ensure the stability of the target during reinforcement learning, a target network is introduced to calculate the target Q-value (denoted as y). v The formula for calculating the target Q value is as follows:
[0164]
[0165] Among them, R v (t) represents the action A performed by vehicle v. v The immediate reward obtained after (t) reflects the contribution of the action to the system goal at the current moment; γ∈(0,1] is a discount factor used to weigh the importance of immediate rewards versus future rewards. This indicates that under the main network θ, state S v The optimal action corresponding to ′(t), i.e., from the next state S v All possible A'(t) v In the action ′(t), select the action that maximizes the Q value of the main network θ output.
[0166] Loss function calculation:
[0167] The training objective of the main network is to make the predicted Q-value as close as possible to the target Q-value, achieved by minimizing the mean squared error (MSE) loss function. Loss function The definition is as follows:
[0168]
[0169] in, This represents the expectation of the dataset D, used to depict the overall error between the predicted Q value and the target Q value.
[0170] Main network parameter update:
[0171] The Adam optimizer, employing stochastic gradient descent and backpropagation, is used to optimize the loss function. The negative gradient direction updates the parameters of the main network θ, enabling the main network to continuously learn better Q-value prediction capabilities to adapt to the dynamically changing network environment of the vehicle network.
[0172] 2) Training process
[0173] This section details the training process of the BD3QN algorithm. Its core objective is to enable the base station agent to learn the optimal decision-making strategy adapted to the vehicle-to-everything (V2X) scenario through reinforcement learning, so as to achieve effective control of network congestion and optimized resource scheduling.
[0174]
[0175] In implementing this invention, the following steps and details can be followed:
[0176] 1. System Initialization Configuration: Configure wireless communication equipment, assign IP addresses to participating vehicles and infrastructure, deploy the AODV routing protocol to ensure dynamic adaptability of network routing, and deploy the TCP protocol to provide a reliable transmission foundation for data transmission. Set up a vehicle simulation scenario, define the vehicle and road information within the scenario, and load the pre-trained BD3QN model parameters θ.
[0177] 2. TCP Cubic Parameter Initialization: Initialize the TCP Cubic congestion window (cwnd) and slow start threshold (ssthresh) to meet the resource allocation requirements of the initial communication state of the vehicle network, laying the foundation for congestion control in the subsequent normal network state.
[0178] 3. Monitor decision-making and strategy execution
[0179] Initial strategy execution: In the initial stage, the TCP Cubic congestion control algorithm is used to increase the congestion limit (cwnd) to adapt to data transmission under normal network load.
[0180] Congestion Detection and Strategy Switching: When the Monitor module detects that the vehicle density exceeds a set threshold or the packet loss rate is higher than a preset percentage, it determines that network congestion has occurred. The BD3QN module is then activated, and through a gradual switching strategy, it progressively adjusts parameters such as vehicle transmission power to more accurately respond to highly dynamic and congested network environments.
[0181] Network Recovery and Policy Switchback: During congestion control using the BD3QN strategy, if the Monitor module detects a packet loss rate below a certain percentage and a vehicle density below a set threshold for a continuous period, the network is considered to have recovered. At this point, the system gradually switches back to the TCP Cubic algorithm, leveraging its advantages of low latency and ease of deployment under normal network conditions to continue data transmission.
[0182] 5. Status Acquisition and Feature Processing: The system acquires the current network operating status in real time. The collected status data includes, but is not limited to, network packet loss rate, vehicle density, link transmission latency, and bandwidth utilization. The collected status data S(t) is input into the shared module, where it undergoes data cleaning and feature extraction to generate a 128-dimensional feature vector for subsequent action decisions by the BD3QN algorithm.
[0183] 6. Action Decision: Based on the generated 128-dimensional feature vector, the action branch module calculates the Q-values of the corresponding action dimensions such as transmission power, and then combines the Q-values of each dimension to form the final action vector, thereby determining the optimal action of the vehicle in the current network state.
[0184] 7. Feedback Optimization and Parameter Update: Continuously record the network state transition sequence and store these sequences in the replay buffer. At regular time intervals, use the data in the replay buffer to update the BD3QN model parameters, thereby optimizing and improving the model. This allows the BD3QN model to better adapt to the dynamically changing network environment of vehicular networks and improve congestion control performance.
[0185] In summary, this invention addresses the congestion control challenge in highly dynamic scenarios of vehicle-to-everything (V2X) communication by proposing a solution of "Monitor module perception and intelligent switching - hybrid control strategy of TCP Cubic and BD3QN algorithms - UAV relay-assisted communication". First, relying on the Monitor module deployed at the base station, multi-dimensional network load data such as packet loss rate and round-trip time (RTT) are collected in real time. The network state (non-congested / lightly loaded, congested) is accurately determined by calculating the congestion threshold. Based on this determination, the Monitor module dynamically decides on the control strategy: when the network is in a non-congested or lightly loaded state, the TCP Cubic algorithm is invoked, leveraging its lightweight and fast local feedback characteristics to quickly adapt to the conventional network by adjusting the congestion window; when the network enters a congested state, a smooth switching mechanism is triggered, gradually activating the BD3QN algorithm. Utilizing its decoupled branch network architecture, precise control of the highly dynamic environment is achieved, fully leveraging the complementary advantages of the two algorithms. To address sudden traffic spikes in localized areas, drones are introduced as mobile relay nodes. When the Monitor module detects localized congestion, the drone is dispatched to quickly fly over the congested area, establishing a line-of-sight communication link to divert localized data traffic. This directly reduces the transmission pressure on ground links, shortens data packet transmission latency, and further optimizes network transmission performance. This invention integrates the Monitor's perception and intelligent switching capabilities, the collaborative adaptive adjustment capabilities of TCP Cubic and BD3QN algorithms, and the flexible assistance capabilities of drones to construct an innovative technical solution. These three elements are mutually supportive and inseparable.
[0186] Finally, it should be noted that any parts of this invention not described in detail are prior art. Those skilled in the art will understand that the above descriptions are merely preferred embodiments of the invention and are not intended to limit the invention. Although the invention has been described in detail with reference to the foregoing examples, those skilled in the art can still modify the technical solutions described in the foregoing examples or make equivalent substitutions for some of the technical features. All modifications and equivalent substitutions made within the spirit and principles of the invention should be included within the scope of protection of the invention.
Claims
1. A UAV-assisted vehicular networking congestion control method based on branch double-deep Q network, characterized in that, The method comprises the following steps: S1, deploying a Monitor module in the base station of the Internet of Vehicles, which collects network state data in the Internet of Vehicles in real time, wherein the network state data at least includes a packet loss rate and a round-trip time (RTT), and a congestion threshold is calculated according to the network state data to determine whether the network is in a congestion state; S2, constructing a branch double deep Q network (BD3QN) architecture, which comprises a shared module and a branch module, wherein the shared module extracts features from the network state information collected and preprocessed by the Monitor module, and the branch module decomposes an action space into power control and association decision independent sub-action branches to realize parallel processing of vehicle transmission power adjustment and vehicle and base station or unmanned vehicle association relationship decision; S3, based on the network congestion state determined by the Monitor module, adaptively switching a TCP Cubic algorithm and a BD3QN algorithm: when the network is in a non-congestion or light load state, the TCP Cubic algorithm is called to adapt to the network state by adjusting the size of the congestion window; when the network is in a congestion state, a gradual weight migration strategy is adopted to complete the smooth switching from the TCP Cubic algorithm to the BD3QN algorithm in stages, and the BD3QN algorithm is used to make joint optimization decisions on the transmission power of the vehicle and the communication association node; S4, introducing an unmanned vehicle as a mobile relay node, when the Monitor module detects that network congestion occurs in a local area, the unmanned vehicle is dispatched to fly over the congested area to construct a temporary communication link, and the position of the unmanned vehicle and the communication link are dynamically adjusted according to the output decision of the BD3QN algorithm to offload local data traffic.
2. The branch double deep Q network based UAV-assisted V2X congestion control method of claim 1, wherein, In step S1, the judgment condition of the congestion threshold is that when the Monitor module detects that the network packet loss rate exceeds a first set threshold or the RTT exceeds a second set threshold, it is determined that the network enters a congestion state. 3.The UAV-assisted V2V congestion control method based on branch double-deep Q network according to claim 1, wherein, In step S2, the shared module comprises three fully connected layers fc1, fc2 and fc3, the input layer receives the preprocessed network state information, and 128-dimensional abstract features are extracted through fc1, fc2 and fc3 in sequence; Each sub-action branch of the branch module is a multi-layer perceptron structure, which receives the 128-dimensional features output by the shared module as input, maps them to 256-dimensional features, and then maps them to the action space corresponding to each action dimension.
4. The branch double deep Q network based UAV-assisted V2X congestion control method of claim 1, wherein, In step S3, the gradual weight migration strategy includes three stages: an exploration stage, an optimization stage and a decision stage; In the exploration stage, the TCP Cubic algorithm is used as the main control logic, the congestion window adjustment mechanism is retained, and the BD3QN algorithm only optimizes the vehicle transmission power parameter; in the optimization stage, the decision weight of the BD3QN algorithm is increased and the decision weight of the TCP Cubic algorithm is decreased compared with the exploration stage, the BD3QN algorithm analyzes the rate and packet loss information of the transmission layer to construct a two-dimensional decision space of transmission power and network congestion degree, and dynamically adjusts the vehicle transmission power and the association node; in the decision stage, when the BD3QN algorithm learns enough network state-action feedback samples and the decision stability meets the preset requirements, the BD3QN algorithm is switched to a fully autonomous decision state.
5. The branch double deep Q network based UAV-assisted V2X congestion control method of claim 1, wherein, In step S3, the BD3QN algorithm comprises the following sub-steps: S41, decomposing the action space into a power regulation branch and an associated decision branch; S42, extracting environmental state features through a sharing module; S43, each branch network outputs an optimal power adjustment action and an associated node selection action respectively; S44, fusing an advantage function and a state value function to select an optimal action.
6. The branch double deep Q network based UAV-assisted V2X congestion control method of claim 5, wherein, The reward function of the BD3QN algorithm is designed to minimize the total transmission delay of the system, and satisfies the following two constraint conditions: C1: C2: wherein C1 represents a vehicle-associated uniqueness constraint, ensuring that each vehicle can only be uniquely associated to one communication node at any time; C2 represents a vehicle power control factor constraint, the power control factor of a vehicle needs to be within the interval [0, 1] and be discretized into K equidistant values; wherein is an indicator function, if vehicle v is associated to the nth communication node at time t, then otherwise N is the total number of communication nodes; denotes a set of vehicles; β v (t) denotes the power control factor of vehicle v at time t.
7. The branch double deep Q network based UAV-assisted V2X congestion control method of claim 6, wherein, Transmission latency between a vehicle and a base station when they communicate The calculation formula is: wherein, denotes the transmission delay when data is transmitted between the vehicle v and the base station r at time t; D v (t) is the amount of data to be transmitted by the vehicle v at time t; is the spectral bandwidth of the communication channel between the vehicle v and the base station r; is the signal-to-interference-plus-noise ratio when the vehicle v communicates with the base station r at time t; denotes the set of vehicles; Signal-to-interference-and-noise ratio when a vehicle v communicates with a base station r The calculation formula is: wherein is the channel fading coefficient from the vehicle v to the base station r; denotes the channel state between the vehicle v and the base station r, mainly considering the wireless network path loss factor; is the maximum transmission power of the vehicle v; β v (t) is the power control factor of the vehicle v transmitting signals; σ R (t) is the interference from other vehicles in the same base station to the communication link between the vehicle v and the base station r; N0is the power spectral density of additive white Gaussian noise.
8. The branch double deep Q network based UAV-assisted V2X congestion control method of claim 1, wherein, In step S4, when the unmanned aerial vehicle is used as a relay node, the transmission delay calculation formula between the unmanned aerial vehicle and the vehicle is: wherein, denotes the transmission delay of the vehicle v when transmitting data to the drone u at time t; is the spectral bandwidth of the communication channel between the vehicle v and the drone u; D v (t) is the amount of data to be transmitted by the vehicle v at time t; is the signal-to-interference-plus-noise ratio of the communication between the vehicle v and the drone u at time t; denotes the set of vehicles, denotes the set of drones; Signal to interference and noise ratio when a vehicle v communicates with a drone u The calculation formula is: wherein Hv,u(t) represents the channel state between vehicle v and UAV u, which is calculated according to the free space path loss model; βv,u(t) is the channel fading coefficient from vehicle v to UAV u; v Pv(t) is the power control factor for the signal transmitted by vehicle v; Pmax is the maximum transmit power of vehicle v; U Iv,u(t) is the interference caused by other vehicles within UAV u to UAV u; UO Iu,v(t) is the interference caused by other UAVs to the communication link between vehicle v and UAV u; and N0is the power spectral density of additive white Gaussian noise.
9. The branch double deep Q network based UAV-assisted V2X congestion control method of claim 1, wherein, The training optimization process of the BD3QN algorithm comprises: storing state transition samples generated by the vehicle and the communication node interaction into an experience replay buffer according to priority, wherein the state transition samples comprise a current network state, an executed action, an obtained reward and a next network state; every interval of a preset fixed training step, synchronously copying parameters of a main network to a target network, and designing a loss function based on an output difference value of the main network and the target network, and updating the parameters of the main network through a back propagation algorithm.
10. The branch double deep Q network based UAV-assisted V2X congestion control method of claim 1, wherein, The method further comprises a network recovery phase: when the Monitor detects that the network state recovers to a light load or a stable state, first, stopping the exploration process of the BD3QN, and freezing the current parameters of the BD3QN model; then, using a hybrid policy linear transition mode, gradually switching the control logic back to TCP Cubic, gradually increasing the decision weight of TCP Cubic to 100%, and completing the network recovery.
Citation Information
Patent Citations
A congestion control method and related equipment
CN116260773B
Internet of vehicles channel congestion control method based on multi-agent deep reinforcement learning
CN118283700A