A gnn-d3qn-based channel and power resource allocation method for internet of vehicles
By combining global interference information and local vehicle observations with the GNN-D3QN model, efficient joint decision-making on vehicle-to-everything (V2I) channels and power is achieved, solving the computational complexity and dynamic adaptability issues of resource allocation in existing technologies, and improving the capacity of V2I systems and the success rate of V2V communication.
Patent Information
- Application Number
- CN202610715783.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-22
- Publication Date
- 2026-08-25
AI Technical Summary
Existing V2X resource allocation methods suffer from NP-hard computational complexity, making it difficult to achieve globally optimal resource allocation in high-speed vehicle scenarios. Furthermore, they fail to fully utilize local vehicle observation information, resulting in an overly strong reliance on channel state information and an inability to adapt to dynamically changing requirements.
A channel and power resource allocation method based on GNN-D3QN is adopted. Global interference information is extracted by graph neural network and combined with local vehicle observations to construct a deep reinforcement learning model, realize joint decision-making on channel and power, and optimize resource allocation strategy.
It improves resource allocation efficiency, reduces spectrum conflicts, increases V2I system capacity and V2V communication success rate, and adapts to the dynamic changes of complex vehicle networking environments.
Smart Images

Figure CN122640841A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of vehicle-to-everything (V2X) communication technology, specifically relating to a method for allocating V2X channels and power resources based on GNN-D3QN. Background Technology
[0002] With the continuous advancement of smart city construction, the strategic importance of Vehicle-to-Everything (V2X) technology in the field of intelligent transportation is becoming increasingly prominent. This technology encompasses two core communication modes: Vehicle-to-Infrastructure (V2I) communication and Vehicle-to-Vehicle (V2V) communication, which complement and coordinate with each other in intelligent transportation systems. Through the deep integration of these two communication methods, comprehensive information interaction and sharing between vehicles, road infrastructure, pedestrians, and cloud networks can be achieved. As the automotive industry accelerates its evolution towards cutting-edge technologies such as autonomous driving, intelligent navigation, and automatic parking, the core value and application prospects of V2X are attracting increasing attention. However, V2X technology still faces critical bottlenecks in communication performance and system security that urgently need to be addressed. To address these challenges, both academia and industry have proposed a series of V2X solutions to optimize vehicle communication performance from multiple dimensions. Representative examples include Cellular Vehicle-to-Everything (C-V2X) and New Radio Vehicle-to-Everything (NR-V2X) developed under the 3GPP framework. However, the introduction of these emerging technologies also brings new technical challenges. To meet the inherent wireless communication requirements of V2X scenarios, resource allocation has become one of the core research topics in vehicular networks. Given that this problem is NP-hard in terms of computational complexity, ensuring both the data transmission success rate of V2V links and the throughput requirements of V2I links in a vehicular network environment presents considerable technical challenges.
[0003] Currently, traditional V2X vehicle-to-everything (V2X) resource allocation solutions are mainly divided into two categories: centralized and distributed, forming the core of the current state of traditional resource allocation. Centralized solutions, represented by the Mode 3 mode of 3GPP R14 LTE-V2X, involve unified scheduling and management of radio resources by cellular network base stations or roadside units (RSUs). This allows for global coordination of resource demands for both V2I and V2V links, reducing the probability of resource conflicts to some extent. However, it suffers from high scheduling latency and strong dependence on base stations, making it prone to resource allocation lag in scenarios with dense vehicle traffic and frequent link switching, and unable to meet the low-latency requirements of safety-related services. Distributed solutions, typically represented by the perceptual semi-static scheduling (SPS) algorithm in Mode 4, allow vehicles to autonomously perceive the surrounding channel conditions and select resources, reducing dependence on infrastructure and offering better latency. However, these solutions generally suffer from half-duplex errors, hidden terminal errors, and resource block conflicts, resulting in low packet reception rates and difficulty adapting to the dynamic changes in complex vehicular scenarios. In addition, traditional resource allocation methods include greedy algorithms, game theory algorithms, and heuristic algorithms. These methods are mostly based on fixed mathematical models and constraints for static or semi-static resource allocation. Although the computational logic is simple and easy to implement, they lack the ability to adapt to the dynamics of vehicle scenarios. They cannot respond in real time to the resource allocation adjustment needs caused by vehicle movement, channel time-varying, and fluctuations in service demands. Furthermore, they are difficult to achieve global optimization in multi-objective optimization (such as the coordinated optimization of latency, reliability, and throughput), and it is difficult to overcome the computational complexity bottleneck caused by NP-hard problems.
[0004] With the rapid development of artificial intelligence technology, Deep Reinforcement Learning (DRL), with its powerful dynamic decision-making and environmental adaptation capabilities, has gradually become a research hotspot in the field of V2X vehicle-to-everything (V2X) resource allocation, forming a new state of technological research. Deep Reinforcement Learning combines the feature extraction capabilities of deep neural networks with the sequential decision-making capabilities of reinforcement learning. It does not rely on precise system mathematical models and can autonomously learn resource allocation strategies through continuous interaction between the agent and the dynamic onboard environment, perfectly adapting to the dynamic and uncertain requirements of V2X scenarios. Currently, V2X resource allocation schemes based on deep reinforcement learning have been explored in multiple directions: In network slicing scenarios, by designing reasonable state spaces, action spaces, and reward functions, agents are trained to achieve dynamic resource allocation between low-latency V2V links and high-throughput V2I links, improving the performance of non-safety services while ensuring latency constraints for safety-related services; In complex scenarios such as urban canyons, based on classic deep reinforcement learning algorithms such as Q-learning, communication channels and transmit power are dynamically adjusted to effectively solve channel congestion problems, significantly improving channel utilization and reducing packet loss rates to a low level. Meanwhile, researchers have further improved the algorithm's decision-making efficiency and emergency response capabilities by optimizing the reward function design (such as dynamically adjusting the reward weights for different business types) and introducing Priority Experience Replay (PER), enabling the agent to quickly respond to critical events such as collision warnings. However, the current application of deep reinforcement learning in V2X resource allocation still has shortcomings: some algorithms have complex training processes and require a large number of samples, resulting in high training latency; some solutions do not fully consider the collaborative decision-making problem of multiple vehicles and multiple links, and are prone to local optima; in addition, in scenarios where the environment changes rapidly due to high-speed vehicle movement, the real-time performance and stability of the algorithm still need further optimization. These issues have become the main research challenges and improvement directions in this field.
[0005] In summary, traditional methods suffer from problems such as excessive reliance on channel state information, difficulty in achieving global optimization, and inability to perform end-to-end learning and adaptation. Furthermore, existing DRL methods focus on improving network architecture to achieve more efficient resource allocation strategies, lacking consideration of information derived from local vehicle observations.
[0006] Therefore, the GNN-D3QN resource allocation method of the present invention can make the most of global interference information and combine it with vehicle local information to achieve efficient resource allocation, overcoming the shortcomings of existing artificial intelligence technologies.
[0007] A prior art method for V2X resource allocation based on graph neural networks and reinforcement learning (publication number CN119277443A) is disclosed. This method introduces the concepts of graph neural networks and reinforcement learning, modeling the vehicle-to-everything (V2X) communication link as graph nodes and interference relationships as graph edges. It uses a graph neural network (GNN) to extract global features and a direct-access QN (DQN) to make resource decisions, achieving dynamic allocation of V2X channels and power. However, this approach does not integrate and jointly train the graph neural network and DQN, failing to fully utilize vehicle-local observation information and global interference information for collaborative optimization. It also cannot achieve precise decoupling decisions for combined channel and power actions, making it difficult to balance V2I throughput and V2V transmission reliability in high-speed mobile scenarios. Summary of the Invention
[0008] The main objective of this invention is to overcome the shortcomings of existing technologies and propose a channel and power resource allocation method based on GNN-D3QN. The system constructed by this invention can extract global interference information of vehicle-to-everything (V2I) networks, improve the efficiency of sub-channel and resource allocation, ultimately reduce spectrum conflicts, maximize V2I system capacity, and improve the success rate of V2V information transmission.
[0009] The technical solution adopted by this invention to achieve the above objectives is: a method for allocating channel and power resources in vehicle-to-everything (V2X) networks based on GNN-D3QN, comprising the following steps: Step 1: Establish the basic framework of the communication system, define the vehicle networking environment and conditions, and provide a foundation for the subsequent GNN-D3QN deep reinforcement learning algorithm; Step 2: Construct a GNN-D3QN deep reinforcement learning model, using the low-dimensional features of the graph neural network and local vehicle observations as inputs, and outputting channel and power decisions to improve the overall performance of the system. Step 3: Train the model and output the joint channel and power allocation decision. After the deep reinforcement learning model is trained, all network parameters (including weights and biases) are determined and saved as the model. Inputting the model into the vehicle-to-everything (V2X) environment will enable resource allocation.
[0010] Preferably, the communication system model includes V2I links and V2V links. Assuming there are M orthogonal sub-bands, the channel power gain of the Kth V2V link in the mth sub-band is set as follows:
[0011] in This indicates small-scale fading. This represents large-scale fading consisting of shadowing effects and path loss. The link signal-to-interference-plus-noise ratio (SINR) and transmission rate are calculated based on this gain.
[0012] Preferably, based on the channel power gain, the signal-to-interference-plus-noise ratio (SINR) of the V2I link on the m-th subband and the V2V link on the k-th subband at the receiver is constructed as follows: .
[0013] Preferably, based on the channel power gain, the signal-to-interference-plus-noise ratio (SINR) of the V2V link on the m-th subband of the receiver is constructed as follows:
[0014] in
[0015] The transmit power of the m-th V2I and the k-th V2V on the subband is set to and , It is the interference channel gain of the Kth V2V link to the mth subband. It is the channel gain from the transmitter to the base station of the m-th V2I link. Indicates noise power. =1 indicates that the nth V2V link is used to transmit messages using the mth sub-channel; otherwise, it is 0. This can be considered a simplified interference power. For example, if a V2V link between vehicle A and vehicle B communicates via spectrum a, and vehicle C also uses the same frequency subband for communication, both transmitters will interfere with each other. The same applies to V2I links.
[0016] Preferably, the transmission rate of the m-th V2I link in the m-th subband is calculated based on the signal-to-interference-plus-noise ratio (SINR) of the V2I link: .
[0017] Preferably, the V2V link transmission rate is calculated based on the signal-to-interference-plus-noise ratio (SINR) of the V2V link:
[0018] B is the bandwidth of each subband, representing the k-th V2V link in the m-th subband.
[0019] Preferably, the vehicle network is modeled as an undirected weighted graph. Each V2V communication link is defined as a node in the graph, and the set of nodes is denoted as V={v1,v2,…,vk}. The node initialization vector includes the channel power gain, sub-channel power gain, interference signal strength, and the resource selection result of the link in the previous time slot; and The edge weights between them are represented as follows: .in, Let p be the Euclidean distance between the p-th and q-th V2V links. For square array The maximum distance value in the range.
[0020] Preferably, a graph convolutional aggregation method based on a message-passing mechanism is adopted, and the node feature is iteratively updated as follows:
[0021] in, Let be the trainable weight matrix and bias term in the i-th iteration. Nodes share weights to reduce model complexity. The activation function is a linear rectifier, and its expression is: .
[0022] Preferably, based on the collected and observed state information, the D3QN network selects actions according to the policy. Since the agent needs to select both sub-channels and power levels simultaneously, we combine these two types of actions into a composite action. By combining sub-channels and power into a composite action, the agent's selected composite action is mapped to two dimensions, representing the selection of sub-channels and power levels, respectively:
[0023]
[0024] Here, % represents modulo operation, and / represents division that rounds down to the nearest integer. This represents the sub-channel selection action of the decomposition, and This indicates the selection of the power level.
[0025] Preferably, in step 2, the model training reward function is:
[0026] in, Indicates the weight of the V2I link. This represents the weight of the transmission time utilized. Tt represents the remaining time, and T0 represents the transmission delay limit. Therefore, ( -Tt) represents the time used for transmission.
[0027] (1) Construct a single-cell vehicle-to-everything (V2X) communication system model and define the graph structure of the V2X. This modeling approach clarifies the V2X communication standards, extracts global features of the V2X, provides a theoretical basis for resource allocation algorithms, and thus improves the communication performance of the V2X system.
[0028] (2) A deep reinforcement learning algorithm is constructed, and a graph neural network is introduced into it to improve the learning efficiency of the model by extracting global features. Based on the constructed system model, experimental data is generated for training, thereby improving the algorithm's adaptability and performance in complex environments. This invention achieves the following effects.
[0029] The final optimized resource allocation scheme is obtained based on the deep learning algorithm trained above.
[0030] The beneficial effects of this invention are as follows: This invention combines deep reinforcement learning and graph neural networks to construct a channel and power resource allocation scheme for vehicle-to-everything (V2X) networks. This scheme can extract global features, improve the efficiency of resource allocation, thereby reducing spectrum interference and increasing the system's V2I capacity and V2V communication success rate. Attached Figure Description
[0031] Figure 1 This is a schematic diagram of a vehicle-to-everything (V2X) system. Figure 2 This is a schematic diagram of the GNN-D3QN algorithm. Detailed Implementation
[0032] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0033] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the present invention or its application or use. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0034] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0035] Example 1: A method for allocating channel and power resources in vehicle-to-everything (V2X) networks based on GNN-D3QN, comprising the following steps: Step 1: Establish the basic framework of the communication system, define the vehicle networking environment and conditions, and provide a foundation for the subsequent GNN-D3QN deep reinforcement learning algorithm; Step 2: Construct a GNN-D3QN deep reinforcement learning model, using the low-dimensional features of the graph neural network and local vehicle observations as inputs, and outputting channel and power decisions to improve the overall performance of the system. Step 3: Train the model and output the joint channel and power allocation decision. After the deep reinforcement learning model is trained, all network parameters (including weights and biases) are determined and saved as the model. Inputting the model into the vehicle-to-everything (V2X) environment will enable resource allocation.
[0036] Example 2: This example is based on the previous example and further elaborates on the present invention.
[0037] Step 1: Establish the basic framework of the communication system, define the vehicle networking environment and conditions, and provide a foundation for the subsequent GNN-D3QN deep reinforcement learning algorithm.
[0038] Communication modeling: Assuming there is a V2I link, it is represented as And K pairs of V2V links, represented as .
[0039] Set the channel power gain of the Kth V2V link in the mth sub-band to...
[0040] The SINR of the m-th subband V2I link and the k-th subband V2V link at the receiver can be constructed as follows:
[0041] and
[0042] in
[0043] The transmit power of the m-th V2I and the k-th V2V on the subband is set to and , It is the interference channel gain of the Kth V2V link to the mth subband. It is the channel gain from the transmitter to the base station of the m-th V2I link. Indicates noise power. =1 indicates that the nth V2V link is used to transmit messages using the mth sub-channel; otherwise, it is 0. It can be considered as a simplified interference power.
[0044] Transmission rate of the m-th V2I link in the m-th subband
[0045] Similarly, we can obtain
[0046] B is the bandwidth of each subband, representing the k-th V2V link in the m-th subband.
[0047] Graph modeling: Model the Internet of Vehicles as an undirected weighted graph .
[0048] Node set: Each V2V communication link is defined as a node in the graph, and the node set is denoted as . The node initialization vector includes the channel power gain, sub-channel power gain, interference signal strength, and the resource selection result of the link in the previous time slot.
[0049] Edge set and edge weight: The interference relationship between V2V links is defined as an edge in the graph. An undirected edge is established between the corresponding nodes only when the distance between two V2V links is less than a preset neighbor distance threshold. All edges constitute a set. To quantify the interference intensity between links, each edge is assigned a weight: the closer the links are, the stronger the mutual interference, and the greater the weight of the corresponding edge. Node and The edge weights between them can be expressed as: .in, Let p be the Euclidean distance between the p-th and q-th V2V links. For square array The maximum distance value in the range.
[0050] Neighbor set: For node v, its neighbor set This refers to all nodes in the graph that are connected to this node by an edge, i.e., other V2V link nodes with strong interference associations within the distance threshold.
[0051] Graph aggregation mechanism: A graph convolution aggregation method based on message passing is adopted. The iterative update process of node features is as follows:
[0052] in, Let be the trainable weight matrix and bias term in the i-th iteration. Nodes share weights to reduce model complexity. The activation function is a linear rectifier, and its expression is: .
[0053] Step 2: Construct a GNN-D3QN deep reinforcement learning model, using low-dimensional features of the graph neural network and local vehicle observations as inputs, and outputting channel and power decisions to improve the overall performance of the system.
[0054] The details of the states, actions, and rewards in the D3QN network are as follows.
[0055] State space: For a V2X environment, state information mainly includes the vehicle's interaction with the environment. The observations and low-dimensional features extracted by the GNN model To help the agents make better decisions, information on the channel and power selection of each agent's neighbors is collected. And calculate the ratio of the remaining bits that the vehicle needs to send to the total number of bits that need to be sent. And the remaining transmission time under delay constraints. .
[0056] Action Space: Based on the collected and observed state information, the D3QN network selects actions according to a policy. Since the agent needs to select both sub-channels and power levels simultaneously, we combine these two types of actions into a composite action. The composite action selected by the agent is mapped to two dimensions, representing the selection of sub-channels and power levels, respectively.
[0057]
[0058]
[0059] Here, % represents modulo operation, and / represents division that rounds down to the nearest integer. This represents the sub-channel selection action of the decomposition, and This indicates the selection of the power level.
[0060] Reward function:
[0061] in, Indicates the weight of the V2I link. This represents the weight of the transmission time utilized. Tt represents the remaining time, and T0 represents the transmission delay limit. Therefore, ( -Tt) represents the time used for transmission.
[0062] The GNN-D3QN model consists of two layers: 1. Graph Neural Network Layer: A two-layer graph convolutional network with 60 input dimensions and 20 output dimensions, extracting global interference features. 2. D3QN Decision Layer: A three-layer fully connected layer (300→200→120), taking GNN features and local observations as input, and outputting 60-dimensional action Q-values. Algorithm Workflow: 1. Local observations are generated for each vehicle's V2V link. 2. The GNN aggregates global information based on interference and outputs low-dimensional features. 3. The features are input to the D3QN, which outputs channel and power decisions. 4. Decisions are executed and rewards are obtained, updating the GNN and D3QN parameters.
[0063] Step 3: After the deep reinforcement learning model is trained, all parameters of the network (including weights and biases) are determined and saved as the model. Inputting the model into the vehicle network environment will enable resource allocation.
[0064] Through the above steps, we have constructed a vehicle-to-everything (V2X) channel and power allocation system. This system can reduce spectrum interference, increase V2I capacity, and improve V2V communication success rate in complex V2X communication scenarios.
Claims
1. A method for allocating channel and power resources in vehicular networks based on GNN-D3QN, characterized in that, Includes the following steps: Step 1: Establish the basic framework of the communication system and define the vehicle-to-everything (V2X) environment and conditions; Step 2: Construct a GNN-D3QN deep reinforcement learning model, using the low-dimensional features of the graph neural network and local vehicle observations as inputs, and outputting channel and power decisions; Step 3: Train the model and output the joint channel and power allocation decision.
2. The method according to claim 1, characterized in that, The communication system model includes V2I links and V2V links. It is assumed that there are M orthogonal subbands, and the channel power gain of the Kth V2V link in the m-th subband is set as follows: in, This indicates small-scale fading. This represents large-scale fading consisting of shadowing effects and path loss.
3. The method according to claim 2, characterized in that, Based on the channel power gain, the signal-to-interference-plus-noise ratio (SINR) of the V2I link in the m-th subband and the V2V link in the k-th subband at the receiver is constructed as follows: 。 4. The method according to claim 3, characterized in that, Based on the channel power gain, the signal-to-interference-plus-noise ratio (SINR) of the m-th subband V2V link at the receiver is constructed as follows: in The transmit power of the m-th V2I and the k-th V2V on the subband is set to and , It is the interference channel gain of the Kth V2V link to the mth subband. It is the channel gain from the transmitter to the base station of the m-th V2I link. Indicates noise power. =1 indicates that the nth V2V link is used to transmit messages using the mth sub-channel; otherwise, it is 0.
5. The method according to claim 3, characterized in that, Based on the signal-to-interference-plus-noise ratio (SINR) of the V2I link, calculate the transmission rate of the m-th V2I link in the m-th subband: 。 6. The method according to claim 4, characterized in that, Calculate the V2V link transmission rate based on the signal-to-interference-plus-noise ratio (SINR) of the V2V link: B is the bandwidth of each subband, representing the k-th V2V link in the m-th subband.
7. The method according to claim 1, characterized in that, Model the Internet of Vehicles as an undirected weighted graph Each V2V communication link is defined as a node in the graph, and the set of nodes is denoted as V={v1,v2,…,vk}. The node initialization vector includes the channel power gain, sub-channel power gain, interference signal strength, and the resource selection result of the link in the previous time slot; and The edge weights between them are represented as follows: ;in, Let p be the Euclidean distance between the p-th and q-th V2V links. For square array The maximum distance value in the range.
8. The method according to claim 7, characterized in that, Employing a graph convolutional aggregation method based on a message-passing mechanism, the node features are iteratively updated as follows: in, Let be the trainable weight matrix and bias term in the i-th iteration. Nodes share weights to reduce model complexity. The activation function is a linear rectifier, and its expression is: .
9. The method according to claim 1, characterized in that, Based on the collected and observed state information, the D3QN network selects actions according to the policy. Since the agent needs to select both sub-channels and power levels simultaneously, we combine these two types of actions into a composite action. The composite action selected by the agent is mapped to two dimensions, representing the selection of sub-channels and power levels, respectively: Where % represents modulo operation, and / represents division that rounds down to the nearest integer; This represents the sub-channel selection action of the decomposition, and This indicates the selection of the power level.
10. The method according to claim 1, characterized in that, In step 2, the model training reward function is: Among them, Indicates the weight of the V2I link. The weight of the transmission time used is indicated; Tt represents the remaining time, and T0 represents the transmission delay limit. Therefore, ( -Tt) represents the time used for transmission.
Citation Information
Patent Citations
V2X resource allocation method based on graph neural network and reinforcement learning
CN119277443A