Channel resource allocation method and system based on hierarchical reinforcement learning

Through the channel resource allocation method based on hierarchical reinforcement learning, the problems of communication interference and uneven resource allocation in cellular vehicle network are solved, fair and efficient allocation of spectrum resources are achieved, communication reliability and overall performance are improved, and latency is reduced.

CN119967616BActive Publication Date: 2025-09-02SHANDONG UNIV OF SCI & TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510098848.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-09-02
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

Communication interference problems between vehicles and between vehicles and infrastructure in cellular vehicle networking are prominent, resulting in a decline in communication quality and an increase in data transmission delay. The existing research is insufficient in terms of the fairness and efficiency of resource allocation, especially in high-density traffic scenarios, how to achieve efficient utilization and fair distribution of spectrum resources is a key issue.

Method used

The channel resource allocation method based on hierarchical reinforcement learning is adopted, and by building an environmental model, fairness model, reliability model and delay model of cellular vehicle networking, the channel resource allocation problem is decomposed into the upper-level clustering problem and the channel resource optimization problem in the lower-level clustering. The deep reinforcement learning of multiple agents is used to optimize spectrum resources and transmission power allocation, the upper-level controller is designed to make clustering decisions, and the lower-level controller is optimized.

Benefits of technology

It realizes fair and efficient allocation of spectrum resources and transmission power in complex environments, improves communication reliability and overall performance, reduces latency, and improves resource utilization efficiency and system stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119967616B_ABST
    Figure CN119967616B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of intelligent transportation technology, and specifically discloses a channel resource allocation method and system based on hierarchical reinforcement learning. The present invention proposes a fairness model based on Jain's fairness index, which is used to measure the fairness of resource allocation and ensure that the resource allocation between different vehicles is relatively balanced; a reliability model is established, which measures the probability of successful transmission of information within a specified time by introducing the concept of interruption probability, thereby ensuring the communication needs of the V2V link; a delay model is designed based on the information age AOI model, which is used to evaluate the time it takes for information to be generated and received, so as to ensure timely transmission of information. The present invention also constructs a cellular vehicle network channel resource allocation model based on hierarchical reinforcement learning. Through hierarchical design, the model decomposes the cellular vehicle network channel resource allocation problem into an upper-layer clustering problem and a lower-layer intra-cluster channel resource optimization problem, thereby ensuring fair and efficient resource allocation in a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent transportation and relates to a channel resource allocation method and system based on hierarchical reinforcement learning. Background Art

[0002] In recent years, with the rapid advancement of technology and the increasing complexity of urban transportation, intelligent transportation systems have gradually become an important means of improving urban traffic conditions. In today's wave of intelligent transportation system development, cellular vehicle-to-everything networks (V2X networks), as one of the key technologies, are driving innovation in the transportation sector. As an emerging communication technology, cellular vehicle-to-everything networks (V2X networks) enable real-time communication between vehicles and their surroundings (including other vehicles, infrastructure, pedestrians, etc.) through cellular networks, greatly ensuring traffic safety and improving travel efficiency.

[0003] Cellular vehicle-to-vehicle (V2V) and vehicle-to-infrastructure (V2I) communications are enabled through cellular networks. These communications not only improve traffic safety and efficiency but also provide a solid foundation for autonomous driving and intelligent traffic management. V2V communication, in particular, enables the sharing of real-time traffic information through direct vehicle-to-vehicle information exchange, thereby improving the overall efficiency of road traffic. With the rapid development of intelligent transportation systems and autonomous driving technologies, building efficient and reliable V2V communication networks has become a research priority.

[0004] However, in the practical application of cellular vehicle-to-vehicle (V2V) networks, communication interference between vehicles and between vehicles and infrastructure is becoming increasingly prominent. In particular, mutual interference between V2V and V2I links can lead to reduced communication quality, increased data transmission latency, and even information loss. This not only impacts vehicle operation but also poses potential safety risks. To mitigate communication interference, the proper allocation of spectrum resources and transmission power is crucial. Spectrum allocation involves properly allocating limited spectrum resources to different communication links to minimize interference and maximize system capacity. Power allocation, on the other hand, involves properly adjusting the transmission power of each link while ensuring communication quality.

[0005] However, unfairness often occurs during spectrum and power allocation, where some links may occupy more resources while others are resource-starved, resulting in a decrease in overall system performance. Currently, the stability and reliability of communication links under high-speed vehicle movement and complex environments remain a challenge. Furthermore, the fairness and efficiency of resource allocation require further study, particularly in high-density traffic scenarios, where efficient utilization and fair allocation of spectrum resources are key issues. Furthermore, latency is particularly prominent in V2V communications, as real-time performance is crucial for ensuring traffic safety and efficiency. Existing research still has many shortcomings in reducing latency and improving communication reliability and fairness. Summary of the Invention

[0006] The purpose of this paper is to propose a channel resource allocation method based on hierarchical reinforcement learning. This method first constructs a network model for vehicle-to-vehicle links in a cellular vehicle-to-network (V2V) network, detailing the channel model, interference model, fairness model, reliability model, and delay model. Based on this model, a hierarchical reinforcement learning-based channel resource allocation model for cellular V2V networks is proposed. This model incorporates a two-layer control strategy with different time scales to rationally allocate spectrum resources and transmission power, ensuring fair and efficient resource allocation in complex environments.

[0007] In order to achieve the above-mentioned goal, the present invention adopts the following technical solutions:

[0008] The channel resource allocation method based on hierarchical reinforcement learning includes the following steps:

[0009] Step 1. Build an environmental model for the cellular vehicle network, including vehicle mobility patterns, communication links, channels, and interference models, and design fairness, reliability, and latency models for the cellular vehicle network.

[0010] Step 2. Based on the cellular vehicle network environment model, fairness model, reliability model, and delay model established in Step 1, a cellular vehicle network channel resource allocation model based on hierarchical reinforcement learning is constructed.

[0011] Through hierarchical design, the model decomposes the cellular vehicle network channel resource allocation problem into the upper-layer clustering problem and the lower-layer intra-cluster channel resource optimization problem, and designs the upper-layer controller and the lower-layer controller accordingly.

[0012] The upper-layer controller performs clustering based on the interference graph, assigning V2V links with minimal interference to the same cluster to reduce intra-cluster interference and provide a stable environment for the lower layers to further optimize channel resource allocation. The upper-layer controller uses the DQN method to make clustering decisions, and through continuous learning and adjustment, the clustering scheme is gradually optimized.

[0013] The lower-layer controller optimizes the allocation of channel resources through multi-agent deep reinforcement learning (MARL), which is used to assign appropriate transmission power to each V2V link to maximize link transmission performance while reducing interference. The lower-layer controller adopts the DQN method to continuously optimize the transmission rate and interference management by learning the optimal resource allocation strategy in a dynamic environment.

[0014] In addition, based on the above-mentioned channel resource allocation method based on hierarchical reinforcement learning, the present invention also proposes a corresponding channel resource allocation system based on hierarchical reinforcement learning, which adopts the following technical solutions:

[0015] The channel resource allocation system based on hierarchical reinforcement learning includes the following modules:

[0016] The cellular vehicle network model building module is used to build the cellular vehicle network environment model, including vehicle mobility patterns, communication links, channels and interference models, and to design the fairness model, reliability model and delay model of the cellular vehicle network;

[0017] The cellular vehicle network channel resource allocation module builds a cellular vehicle network channel resource allocation model based on hierarchical reinforcement learning based on the established cellular vehicle network environment model, fairness model, reliability model, and delay model;

[0018] Through hierarchical design, the model decomposes the cellular vehicle network channel resource allocation problem into the upper-layer clustering problem and the lower-layer intra-cluster channel resource optimization problem, and designs the upper-layer controller and the lower-layer controller accordingly.

[0019] The upper-layer controller performs clustering based on the interference graph, assigning V2V links with minimal interference to the same cluster to reduce intra-cluster interference and provide a stable environment for the lower layers to further optimize channel resource allocation. The upper-layer controller uses the DQN method to make clustering decisions, and through continuous learning and adjustment, the clustering scheme is gradually optimized.

[0020] The lower-layer controller optimizes the allocation of channel resources through multi-agent deep reinforcement learning (MARL), which is used to assign appropriate transmission power to each V2V link to maximize link transmission performance while reducing interference. The lower-layer controller adopts the DQN method to continuously optimize the transmission rate and interference management by learning the optimal resource allocation strategy in a dynamic environment.

[0021] The present invention has the following advantages:

[0022] As described above, the present invention describes a channel resource allocation method and system based on hierarchical reinforcement learning. Specifically, the present invention constructs a cellular vehicle-to-vehicle (V2V) link network model. Specifically, the present invention proposes a fairness model based on Jain's fairness index to measure the fairness of resource allocation and ensure a relatively balanced resource allocation between different vehicles. The present invention also establishes a reliability model that, by introducing the concept of interruption probability, measures the probability of successful information transmission within a specified time, thereby ensuring the communication requirements of the V2V link. The present invention also designs a delay model based on the age of information (AOI) model to evaluate the time elapsed from information generation to reception, thereby ensuring timely information delivery. Furthermore, the present invention constructs a cellular vehicle-to-vehicle (V2V) channel resource allocation model based on hierarchical reinforcement learning. Through a hierarchical design, this model decomposes the cellular vehicle-to-vehicle (V2V) channel resource allocation problem into a clustering problem for the upper-level controller and an intra-cluster channel resource optimization problem for the lower-level controller. In the upper-level algorithm, an interference graph-based clustering algorithm is proposed. By constructing an interference graph, V2V links with less interference are allocated to the same cluster to reduce intra-cluster interference. Clustering decisions are made using a multi-agent deep reinforcement learning method to optimize interference management. In the lower-level algorithm, a channel resource optimization algorithm based on multi-agent reinforcement learning is designed. For the V2V links within each cluster, the transmission rate is maximized and the interference is minimized by optimizing the allocation of sub-channels and transmission power. The present invention is based on a cellular vehicle network channel resource allocation model based on hierarchical multi-agent deep reinforcement learning, which reasonably allocates spectrum resources and transmission power, ensuring fair and efficient allocation of resources in complex environments. In addition, the present invention also verifies the performance of the proposed channel resource allocation method based on hierarchical reinforcement learning in different scenarios through simulation experiments, and evaluates its performance in terms of resource utilization efficiency, communication delay, system robustness and adaptability. By comparing and analyzing with traditional methods, it is demonstrated that the method of the present invention significantly improves the fairness, reliability and overall performance of resource allocation. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 This is a basic network architecture diagram of a cellular vehicle network in an embodiment of the present invention;

[0024] Figure 2 This is a multi-time scale control strategy diagram for cellular vehicle network channel resource allocation in an embodiment of the present invention;

[0025] Figure 3 This is a diagram of the hierarchical architecture of cellular vehicle network channel resource allocation in an embodiment of the present invention;

[0026] Figure 4 Schematic diagram of the rewards obtained from V2V link training as the number of iterations increases in a specific example of the present invention;

[0027] Figure 5Schematic diagram comparing the sum of V2I link rates for different V2V payload sizes in a specific example of the present invention;

[0028] Figure 6 FIG4 is a schematic diagram showing a comparison of fairness of V2V links with different V2V load sizes in a specific example of the present invention;

[0029] Figure 7 Schematic diagram of transmission delay of V2V links with different V2V payload sizes in a specific example of the present invention. DETAILED DESCRIPTION

[0030] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:

[0031] Example 1

[0032] With the development of intelligent transportation systems, cellular vehicle-to-vehicle (V2V) technology has gradually become an important means of improving traffic safety and efficiency. However, due to the large number of vehicles, fast movement speeds, and complex interference in cellular V2V networks, efficient and reliable allocation of wireless resources has become an urgent problem. This paper proposes a channel resource allocation method based on hierarchical reinforcement learning to ensure fair and efficient resource allocation in complex environments. First, a network model of vehicle-to-vehicle links in cellular V2V networks is constructed, and the channel model, interference model, fairness model, reliability model, and delay model are described in detail. Based on this, a cellular V2V channel resource allocation model based on hierarchical reinforcement learning is proposed, which adopts a two-level control strategy with different time scales. To implement the two-level control strategy, the channel resource allocation problem is decomposed into two sub-problems: minimizing the interference of vehicle-to-vehicle links and maximizing the transmission rate of inter-vehicle links. At the upper level, a clustering algorithm for vehicle-to-vehicle links based on an interference graph is proposed. The inter-vehicle links are clustered using a multi-agent deep reinforcement learning method to optimize interference management. A channel resource optimization algorithm based on multi-agent reinforcement learning is designed at the lower layer. For the vehicle-to-vehicle links within each cluster, the transmission rate is maximized and the interference is minimized by optimizing the distribution of transmission power.

[0033] The channel resource allocation method based on hierarchical reinforcement learning in this embodiment specifically includes the following steps:

[0034] Step 1. Build an environmental model for the cellular vehicle network, including vehicle mobility patterns, communication links, channels, and interference models, and design the fairness model, reliability model, and delay model for the cellular vehicle network.

[0035] Specifically, they first established a cellular vehicle network environment model, encompassing vehicle mobility patterns, communication links, channels, and interference models. This model considered vehicle motion characteristics, path loss, shadowing and multipath effects, as well as co-channel and adjacent channel interference in different scenarios. Next, they designed a fairness model to ensure that different vehicles receive equitable treatment in spectrum resource allocation. They then constructed a communication link reliability model to evaluate link stability and data transmission success rates, ensuring reliable communication under varying traffic conditions. Finally, they designed a latency model to analyze data transmission delays, particularly in high-density traffic scenarios, and to investigate how to effectively reduce latency to meet the demands of real-time communication.

[0036] Through detailed analysis of the environmental model, fairness model, reliability model, and delay model, we aim to build a comprehensive model that can more accurately simulate the actual scenarios and requirements of cellular vehicle networks. This will not only provide an important basis for the design and implementation of subsequent optimization algorithms, but also lay a solid foundation for the channel resource allocation algorithm based on hierarchical reinforcement learning.

[0037] In order to design and evaluate resource allocation methods in cellular vehicle networks, the present invention constructs a simulation environment that can simulate actual cellular vehicle network scenarios, including vehicle mobility patterns, communication links, channels, and interference models.

[0038] The following introduces the basic architecture of the cellular vehicle network, which mainly includes the V2I (Vehicle-to-Infrastructure) link and the V2V (Vehicle-to-Vehicle) link. Figure 1 As shown, the V2I link connects each vehicle to a base station (BS) or a BS-type roadside unit, while the V2V link provides direct communication between neighboring vehicles.

[0039] Each V2I link is allocated an independent spectrum sub-band to avoid interference between V2I links.

[0040] The V2I link is mainly used for high-data-rate entertainment services. Specifically, vehicles communicate with base stations (BSs) through cellular networks to transmit high-data-rate entertainment content such as video streaming and online games.

[0041] The set of V2I links is represented as M is the number of spectrum subbands.

[0042] The V2V link is primarily used for periodically generated safety message transmissions, such as direct communication between vehicles to exchange location information, speed, and acceleration, to improve road safety and traffic efficiency.

[0043] The set of V2V links is represented as K is the number of V2V links.

[0044] In this cellular vehicle network architecture, in order to effectively manage and optimize the use of spectrum resources, a spectrum sharing mechanism is adopted, allowing the V2V link to reuse the spectrum sub-bands of the V2I link, thereby improving the spectrum utilization of the overall system under limited spectrum resources.

[0045] However, due to the reuse of V2V links, interference may be introduced, which requires the design of effective interference management strategies to ensure the communication quality and reliability between different links.

[0046] The channel model is used to describe the signal propagation characteristics in wireless communications, including path loss, small-scale fading, and large-scale fading factors. Since the signal attenuates as the propagation distance increases, the path loss model is defined as:

[0047]

[0048] Where PL(d) and PL(d0) are the path losses at distance d and reference distance d0, respectively, and n is the path loss exponent.

[0049] The path loss model is used in this design to calculate signal strength, manage interference, perform link budget analysis, and optimize resource allocation. For example, in an urban environment, buildings and other obstacles can significantly increase path loss.

[0050] The signal fluctuates rapidly due to multipath effects, so the Rayleigh fading model is used. The channel gain h k [m] follows an exponential distribution with a mean of 1. Small-scale fading primarily describes rapid variations due to multipath propagation and phase interference. For example, when a vehicle travels under an overpass or in a tunnel, the signal will experience severe small-scale fading.

[0051] The signal slowly varies due to obstruction and terrain changes, including path loss and shadowing. Large-scale fading describes the slow changes in signal strength, such as shadowing caused by buildings, trees, and other large obstacles.

[0052] In this environment, the channel power gain g of the kth V2V link on the mth subband is k [m] is represented by:

[0053] g k [m] = α k h k [m](2)

[0054] where α k represents large-scale fading (including path loss and shadowing effects), h k[m] represents small-scale fading, which obeys an exponential distribution with a mean of 1 (i.e., it is assumed to be an exponential distribution with a mean of 1).

[0055] The interference model is used to describe the impact of co-frequency interference on the communication link. The V2V link interference is the interference caused by other V2V links reusing the same spectrum and the interference caused by the V2I link. k [m] is calculated as follows:

[0056]

[0057] in, is the transmit power of the mth V2I link, is the interference channel gain of the mth V2I link to the kth V2V link, ρ k′ [m] is an indicator variable indicating whether the k′th V2V link uses the mth spectrum resource block. is the transmission power of the kth V2V link, g k′,k [m] is the interference channel gain of the k′th V2V link to the kth V2V link. For example, when multiple vehicles communicate on the same spectrum resource block, serious interference will occur, affecting the communication quality.

[0058] There is no interference between V2I links, but there will be interference caused by the spectrum reused by V2V links.

[0059] The received signal-to-interference-and-noise ratio (SINR) of the V2I link is:

[0060]

[0061] in, is the SINR of the mth V2I link, and are the transmission powers of the mth V2I link and the kth V2V link on the mth subchannel, σ 2 is the noise power, g k,B [m] is the interference from the kth V2V link to the base station on the mth subchannel, is the interference from the mth V2I link to the base station on the mth subchannel.

[0062] ρ k [m] indicates whether the k-th V2V link uses the m-th spectrum resource block, ρ k [m] is a binary variable. k [m] = 1 means that the kth V2V link selects the mth subband, and ρ k [m]=0 means the opposite.

[0063] When building a cellular vehicle network environment model, the evaluation of channel rate and link performance is a key part.

[0064] Among them, the V2I link rate Expressed as:

[0065] Where W is the bandwidth of the spectrum subband; is the SINR of the mth V2I link. A high SINR means a higher link rate, which is suitable for high-data-rate entertainment services.

[0066] The V2V link rate R k Expressed as:

[0067] in, is the SINR of the kth V2V link.

[0068] In cellular vehicle-to-vehicle networks (V2Vs), the fairness of channel resource allocation is crucial for improving network performance and user experience. This section designs a fairness model for V2V links to ensure fair channel resource allocation across vehicles. To ensure fairness in V2V link resource allocation, the Jain's fairness index is used as a fairness metric.

[0069] Jain's fairness index f(R k ) is defined as follows:

[0070]

[0071] Among them, R k is the rate of the kth V2V link, K is the total number of V2V links; Jain's fairness index ranges from 0 to 1, and the closer the value is to 1, the fairer the resource allocation. In order to achieve fairness in resource allocation in cellular vehicle networks, this paper mathematically models the fairness model. Assume represents the set of all V2V links, and the goal of the fairness model is to maximize Jain's fairness index f(R k ), described as follows:

[0072]

[0073] At the same time, when designing the fairness model, that is, Equation (8), the following assumptions and constraints are made:

[0074]

[0075] Where M represents the number of spectrum sub-bands, and formula (9) indicates that each V2V link can only use one spectrum sub-band;

[0076]

[0077] Formula (10) indicates that the transmission power of each V2V link is limited by the maximum power P max .

[0078] The fairness model of the cellular vehicle network constructed by the present invention takes into account the interference between V2V links and between V2V and V2I links, so that the interference power should be kept within an acceptable range.

[0079] In cellular vehicle-to-vehicle networks, reliability refers to the probability that information can be successfully transmitted within a specified timeframe. This is crucial for secure message delivery. To ensure that the communication requirements of each V2V link are met, a reliability model will be designed. Reliability can be measured by the transmission success rate of the V2V link, while reliability will be described using the outage probability, which is defined as the probability that data transmission fails within a specified timeframe.

[0080] Accordingly, reliability can be defined as the complement of outage probability. Next, the outage probability will be derived.

[0081] First, assume that the SINR threshold of the kth V2V link on the mth subband is γ th , then the interruption probability P out Defined as:

[0082]

[0083] in, represents the interruption probability of the kth V2V link, Pr is a probability symbol; represents the SINR of the k-th V2V link on the m-th subband, which is defined as:

[0084]

[0085] in, is the transmit power of the kth V2V link on the mth subband, g k [m] is the channel gain, σ 2 is the noise power, I k [m] is the interference power; Substituting into formula (11), we have:

[0086]

[0087] Formula (13) can be transformed into:

[0088]

[0089] Assuming the channel gain g k [m] follows an exponential distribution with unit mean, and its probability density function for:

[0090]

[0091] where λ is the distribution parameter; for an exponential distribution with unit mean, λ = 1; therefore, the cumulative distribution function (CDF) is:

[0092]

[0093] Among them, F gk (x) represents the cumulative distribution function; from this, the outage probability is:

[0094]

[0095] Assumptions represents the set of all V2V links. The objectives of the reliability model are described as follows:

[0096]

[0097] By minimizing the sum of outage probabilities, the reliability of the entire cellular vehicle network can be improved.

[0098] Latency is a key performance metric for cellular vehicle-to-vehicle (V2V) networks, particularly in vehicle-to-vehicle (V2V) communications. Low latency is crucial for improving driving safety and information transmission efficiency. This section designs a latency model for V2V links based on the Age of Information (AOI) model to ensure timely information delivery.

[0099] Age of Information (AOI) measures the time from when information is generated to when it is received. Unlike traditional latency metrics, AOI not only takes into account transmission time, but also the frequency of information updates. It is defined as follows:

[0100] A k (t) = tU k (t) (19)

[0101] Among them, A k (t) is the information age of the kth V2V link at time t, U k (t) is the time when the latest information is updated.

[0102] The update rule of information age is defined as follows:

[0103]

[0104] Among them, A k(t) is the information age of the kth V2V link at time t, Δt is the time interval, R k is the transmission rate of the kth V2V link, R min Is the minimum transmission rate requirement.

[0105] At the same time, in order to minimize the average information age of the V2V link, the following objective function is defined:

[0106]

[0107] Where K is the number of V2V links, T is the time period, and A k (t) is the information age of the kth V2V link at time t.

[0108] Step 1 above established a cellular V2V link network model, including the design of fairness, reliability, and latency models. These models laid the foundation for achieving fair allocation of spectrum and power resources, improving communication reliability, and reducing latency. These models will be applied to the design of specific channel resource allocation methods in subsequent steps.

[0109] Step 2. Based on the cellular vehicle network environment model, fairness model, reliability model, and delay model built in step 1, a cellular vehicle network channel resource allocation model based on hierarchical reinforcement learning is constructed.

[0110] When dealing with a large number of agents, traditional multi-agent deep reinforcement learning (MARL) suffers from the exponential growth of the state and action spaces, as each agent must consider the strategies and possible actions of the others. Furthermore, in MARL, an agent's learning environment is composed of other learning agents, which leads to instability in the learning environment, affecting the convergence and stability of the training process, resulting in the so-called "environmental non-stationarity" problem. Furthermore, because agents in MARL must simultaneously learn how to interact with other agents, this can lead to low learning efficiency.

[0111] Therefore, this paper proposes a cellular V2X channel resource allocation model based on hierarchical reinforcement learning to address the challenges of limited spectrum resources and interference management. Through a hierarchical design, this model decomposes the cellular V2X channel resource allocation problem into a clustering problem at the top level and an intra-cluster channel resource optimization problem at the bottom level. This model aims to ensure the reliability of vehicle-to-vehicle (V2V) and vehicle-to-infrastructure (V2I) communications while maximizing overall network performance.

[0112] Figure 2This paper presents an architecture for handling vehicle-to-vehicle (V2V) communications at different timescales. At large timescales, a high-level control strategy is responsible for system-level V2V link clustering. At small timescales, a low-level control strategy is responsible for link-level intra-cluster channel resource optimization, ensuring communication continuity and reliability while improving system performance.

[0113] Figure 3 The process of a vehicle channel resource allocation method based on hierarchical reinforcement learning is demonstrated. This hierarchical reinforcement learning-based cellular vehicle network channel resource allocation model decomposes the cellular vehicle network channel resource allocation problem into a clustering problem of the upper-level controller and an intra-cluster channel resource optimization problem of the lower-level controller through a hierarchical design.

[0114] The upper-layer controller performs clustering based on the interference graph, assigning V2V links with less interference to the same cluster to reduce intra-cluster interference. This provides a stable environment for the lower layers, allowing them to further optimize channel resource allocation. The upper-layer algorithm uses a deep Q-network (DQN) to make clustering decisions, and through continuous learning and adjustment, the clustering scheme is gradually optimized.

[0115] The lower-layer controller optimizes the allocation of channel resources through multi-agent deep reinforcement learning (MAR L). The goal of the lower layer is to allocate appropriate transmission power to each V2V link to maximize link transmission performance while reducing interference. The lower-layer algorithm also uses the deep Q network (DQN) method to continuously optimize the transmission rate and interference management by learning the optimal resource allocation strategy in a dynamic environment.

[0116] The following details the design of the upper-level controller for the interference graph-based clustering approach. The upper-level controller's goal is to assign V2V links with minimal interference to the same cluster, specifically, to the same subchannel to reduce intra-cluster interference. This provides a good initial condition for channel resource optimization decisions in the lower-level controller. The specific design process is as follows:

[0117] Suppose a network has several V2V links, each with varying degrees of interference. By constructing an interference graph, we can visually represent the interference relationships between links. In the interference graph, nodes represent V2V links, and edge weights represent the channel gains between links. A reward function penalizes the sum of interference within each cluster, guiding the agent to select the optimal clustering scheme that minimizes intra-cluster interference.

[0118] In the upper-level controller, the action space is defined as possible clustering schemes; each action corresponds to a scheme for allocating V2V links to different clusters and selecting an independent subchannel for each cluster.

[0119] set up is the action space, then the action represents a clustering scheme; specifically, the goal of the clustering scheme is to assign V2V links with less interference to the same cluster in order to better allocate channel resources.

[0120] Assume there are K V2V links in the network, and the action space of the upper controller is represented as a finite set, where each action represents a specific clustering scheme:

[0121]

[0122] Where L is the number of clusters, C i represents the set of V2V links contained in the i-th cluster. The V2V links within each cluster have minimal interference. Since clustering is equivalent to selecting subchannels, the M disjoint subchannels of the cellular vehicle network wireless spectrum resources constitute the action space of the upper-layer controller; each subchannel is occupied by a V2I link.

[0123] The state space includes the state information of all V2V links in the current network, such as the channel gain, interference situation and current cluster allocation of each link.

[0124] set up is the state space, then the state Represents the current state of the network. The design of the state space is as follows:

[0125]

[0126] Where K is the number of V2V links, M is the number of subbands, and g k [m] represents the channel gain of the k-th V2V link on the m-th subband, I k [m] represents the interference suffered by the k-th V2V link on the m-th subband, C k Indicates the current cluster allocation status of the k-th V2V link.

[0127] At each coherent time step t, the states of all V2V links are again observed Z(t), defined by the observation function O:

[0128] Z(t)=O(s t ) (twenty four)

[0129] Among them, s t Represents the state at time t.

[0130] The reward function of the upper-layer controller will be used to evaluate the pros and cons of each clustering scheme. To ensure the effectiveness of clustering and the convergence of training, the reward function should take into account the interference between links and resource utilization.

[0131] Let r be the reward function, then it is as follows:

[0132]

[0133] in, represents the current set of all clusters, C represents one of the clusters, (i, j) represents two V2V links belonging to the same cluster C, g i,j Denotes the channel gain between V2V links i and j. By minimizing the total interference within each cluster, effective interference management can be achieved and the reward function can be ensured to gradually converge during training.

[0134] After the upper-layer controller completes its decision-making, the V2V links in the cellular vehicle network are properly allocated to different clusters, that is, different subchannels are selected. Next, the goal of the lower-layer controller design is to select the appropriate transmission power for each V2V link to maximize link transmission performance while minimizing interference.

[0135] The design process of the lower-level controller design is introduced in detail below.

[0136] In the channel resource optimization decision in the lower-layer controller, the action space is defined as the combination of the transmit power of each V2V link; let is the action space of the kth agent, then the action represents the power allocation scheme of the kth V2V link.

[0137] Therefore, the action space is designed as follows:

[0138]

[0139] Define P i represents the optional transmit power level, i∈[1,N], each action a k It only consists of one transmission power p; the present invention limits the power control to four levels, namely [23dBm, 10dBm, 5dBm, -100dBm].

[0140] Among them, -100dBm means that the V2V transmission power is zero.

[0141] The state space includes the channel state information and environment information of each V2V link; is the state space of the kth agent, then the state represents the current state of the th V2V link. The state space is designed as follows:

[0142]

[0143] Among them, g k[m] represents the gain of the current channel, I k [m] represents the interference from other links, R k Indicates the current transmission rate, A k represents the current information age, P out represents the interruption probability; at each coherent time step t, the state of each V2V agent k is subject to observation Z k (t), is defined by the observation function O as:

[0144] Z k (t)=O(s k (t),k) (28)

[0145] Among them, s k (t) represents the state of V2V link k at time t.

[0146] The reward function in the lower-layer controller is used to evaluate the pros and cons of each channel resource allocation scheme. To ensure the effectiveness of channel resource allocation, the reward function should comprehensively consider factors such as transmission rate, interference, delay, reliability, and fairness.

[0147] Assume r k is the reward function of the kth agent, then the reward function r k (s k ,a k )as follows:

[0148] r k (s k ,a k )=αR k -βI k -γA k -δP out +ηf(R k ) (29)

[0149] Among them, R k Indicates the current transmission rate, I k Indicates the interference from other links, A k represents the current information age, P out represents the interruption probability, f(R k ) is the fairness indicator, and α, β, γ, δ, and η are weight parameters used to balance the impact of transmission rate, interference, delay, reliability, and fairness on the reward.

[0150] The following section details a cellular vehicle network channel resource allocation decision-making method based on hierarchical reinforcement learning. This method combines a two-layer structure of upper and lower controllers to improve the communication performance of the cellular vehicle network by optimizing channel resource allocation.

[0151] The Q model and loss function model will be described in detail below, and the specific steps of the method will be given.

[0152] The Q model is the core model in reinforcement learning, which uses Q values ​​to evaluate the value of taking an action in a given state. The hierarchical reinforcement learning model in this invention includes an upper-layer controller and a lower-layer controller, each of which has its own Q model.

[0153] For the upper-level controller, its Q-value function is expressed as:

[0154]

[0155] Among them, Q m (s m ,a m ) represents the Q value function of the upper controller, s m is the state of the upper controller, a m is the action to be performed by the upper controller, r m is the return of the upper controller, γ is the discount factor; A m Represents a set of actions; (s m )′ means in state s m Next, perform action a m The state reached later, a′ represents the state (s m )′ to perform the action.

[0156] For the lower-layer controller, its Q-value function is expressed as:

[0157]

[0158] Among them, Q c (s c ,a c ) represents the Q value function of the lower controller, s c is the state of the lower controller, a c is the action to be performed by the lower-level controller, r c is the return of the lower controller, γ is the discount factor; A c Represents a set of actions; (s c )′ means in state s c Next, perform action a c The state reached later, a′ represents the state (s c )′ to perform the action.

[0159] When training the Q network, the loss function is used to measure the gap between the current Q value and the target Q value, which is defined as follows:

[0160]

[0161] Among them, L(θ) represents the loss function, θ is the parameter of the current Q network, θ - are the parameters of the target Q network.

[0162] The specific process of cellular vehicle network channel resource allocation decision-making based on hierarchical reinforcement learning is given below.

[0163] First, the Q networks of all agents are randomly initialized; each V2V link acts as an agent and independently maintains its Q network for selecting actions and evaluating rewards; the initialization of the Q network is crucial to the effectiveness of the algorithm, and reasonable initialization can accelerate the training process.

[0164] In the upper controller, each V2V link is initialized as a single cluster C k = {k}, and construct the interference graph G(V,E); the nodes in the graph represent V2V links, and the edge weights are the channel gains g between the links i,j .

[0165] In each iteration, the vehicle positions and the large-scale fading parameter α are updated to simulate vehicle movement and environmental changes. For each time slice t, all V2V agents observe the state Z(t), select an action A(t) based on the ∈-greedy strategy, and execute the action, or clustering operation. Through continuous iteration and updates, the agents gradually learn the optimal clustering strategy.

[0166] After all agents take action, link pairs with interference less than the threshold θ are merged and the cluster allocation scheme {C k Each cluster is allocated an independent resource block, or subchannel, to ensure there is no interference between clusters. This process minimizes interference within each cluster by gradually optimizing the clustering scheme, thereby improving overall network performance.

[0167] After executing the clustering scheme, the reward r(t) (i.e., the upper-level reward) of each agent is calculated; the reward function takes into account the interference within the cluster and evaluates the pros and cons of the clustering scheme by minimizing the interference within the cluster.

[0168] After a small-scale fading of the update channel, all agents observe the new state Z(t+1) and store the state transition (Z(t), A(t), r(t), Z(t+1)) in the replay memory D. By storing and replaying historical state transitions, agents can repeatedly utilize past experience during training, accelerating the learning process.

[0169] In the lower controller, for each time slice t, all V2V agents independently observe the state Z k (t), select action a according to the ∈-greedy strategy k (t), and performs the action i.e. power allocation.

[0170] Each V2V agent selects action a k (t) Power distribution operation is performed; action a k (t) represents the power allocation decision of agent k in the current state; after executing the power allocation plan, the local reward r of each agent is calculated k (t) (i.e., the lower-level reward), the reward function comprehensively considers transmission rate, interference, delay, reliability, and fairness, and adjusts the impact of each factor on the reward through weight parameters α, β, γ, δ, and η.

[0171] Each agent observes a new state Z k (t+1), and transfer the state (Z k (t),a k (t),r k (t),Z k (t+1)) is stored in the playback memory D k From playback memory D k Small batches of data are uniformly sampled in the network, and stochastic gradient descent is used to optimize the error between the Q network and the learning target; by continuously updating the Q network, the intelligent agent gradually learns the optimal power allocation strategy.

[0172] By jointly optimizing the Q-value functions of the upper-layer controller and the lower-layer controller, the present invention can improve the overall communication performance of the cellular vehicle network while ensuring fairness and reliability.

[0173] In the upper-level controller, efficient subchannel allocation is achieved by constructing an interference graph and implementing a clustering strategy based on interference minimization. In the lower-level controller, a multi-agent reinforcement learning approach optimizes transmit power allocation, maximizing transmission rate while minimizing interference and latency. This hierarchical design not only improves resource utilization efficiency but also enhances robustness and adaptability, providing a strong foundation for efficient and reliable cellular vehicle-to-vehicle communication.

[0174] The present invention designs and implements a fair spectrum and power allocation model, which can significantly reduce the interference between V2V and V2I links, thereby improving the reliability of communication. This is of great significance for ensuring real-time information exchange between vehicles and improving traffic safety. By establishing an optimization model for fair resource allocation, it is possible to balance the resource allocation of each communication link while ensuring the overall performance of the system. This not only improves the fairness of the system, but also ensures the stability and reliability of different links, laying a solid foundation for the practical application of cellular vehicle networks. By introducing a multi-agent hierarchical deep reinforcement learning method (which has strong adaptability and efficient computing power, and can provide a new solution for resource allocation in cellular vehicle networks), and designing a spectrum and power allocation algorithm, efficient resource management can be achieved in a dynamic and complex communication environment.

[0175] Furthermore, to validate the effectiveness of the proposed method, a simulation experiment was conducted on the proposed model to verify the performance of the established resource allocation method based on multi-agent collaborative learning. The simulation scenario was constructed based on the Manhattan traffic scenario specified in the 3GPP TR36.885 standard, which includes nine city blocks. Vehicles, base stations, V2I communication links, and V2V communication lines were set up in this traffic simulation scenario. The specific experimental parameter settings are shown in Table 1.

[0176] Table 1 Experimental parameter settings

[0177]

[0178] The hierarchical structure of the present invention may lead to fluctuations in returns, especially in the early stages, where high-level strategies may frequently change sub-goals, leading to unstable returns. Figure 4 As can be seen in the figure, the rewards fluctuate significantly during the first 250,000 iterations. These peaks and valleys likely reflect the effectiveness of the subgoals selected by the upper-level control policy. For example, in some cases, the lower-level control policy may have selected a very effective subgoal, resulting in a rapid increase in rewards; while in other cases, it may have selected a less appropriate subgoal, causing a decrease in rewards. The overall upward trend indicates that the HDRL system is continuously refining its strategy and gradually learning more efficient task allocation and action execution, indicating that both the upper-level and lower-level control policies are gradually optimizing during the learning process.

[0179] Overall, the cumulative reward obtained by the V2V link as the number of training iterations increases is shown, verifying the convergence behavior of the proposed hierarchical reinforcement learning method. As can be seen from the figure, the reward per iteration increases with the number of training iterations, demonstrating the effectiveness of the proposed training algorithm. When the total training set reaches approximately 300,000, the reward performance gradually converges, despite some fluctuations in channel fading caused by mobility in the vehicle environment. Based on this conclusion, when evaluating the performance of the V2V link, each agent DQN network was trained 350,000 times to ensure the convergence of the learning algorithm.

[0180] The experimental results of the V2V link channel resource allocation algorithm based on hierarchical reinforcement learning are analyzed and discussed in detail below.

[0181] The performance of our proposed algorithm (Hierarchical Deep Reinforcement Learning, or HDRL) was evaluated in various scenarios by comparing it with traditional methods: multi-agent deep reinforcement learning (MARL), single-agent deep reinforcement learning (SARL), and randomized methods. The experiments focused on evaluating transmission rate, reliability, communication latency, fairness, system robustness, and adaptability.

[0182] from Figure 5 As can be seen, while the proposed method achieves a lower V2I link rate compared to MARL and SARL, this is because the reward function designs of MARL and SARL focus solely on transmission rate and reliability, while the proposed method considers more factors. However, as the load increases, the transmission rates of all methods decrease. This is because the increased load causes more resources to be occupied by the V2V link, reducing the resources allocated to the V2I link. Overall, the proposed HDRL performs relatively stably. As the load increases, the sum of the V2I link rates gradually decreases, but the decrease is small. This demonstrates that the HDRL method effectively balances resource allocation between V2V and V2I links under varying load conditions, effectively improving the overall resource utilization efficiency of the system. As the load gradually increases, the HDRL method's sum of rates remains high, demonstrating strong robustness and adaptability. It can also be seen that the overall decrease in the HDRL method with increasing load is minimal.

[0183] Figure 6Figure 2 shows the fairness indices of V2V links for various algorithms under different V2V load sizes. The figure shows significant differences in the fairness performance of different algorithms under different load conditions. HDRL's fairness index remains high under all load conditions, demonstrating its fairness in resource allocation. As the load increases, the HDRL method's fairness index increases slightly, remaining consistently between 0.25 and 0.3. The HDRL method's fairness index is significantly higher than that of other methods, demonstrating its ability to effectively allocate resources, minimizing resource utilization differences between V2V links. The SARL method's fairness index remains low under all load conditions, reaching almost zero and even falling short of the random method overall. The MARL method performs slightly better. The HDRL method consistently maintains a high level of fairness under varying load conditions, ensuring balanced resource allocation between V2V links. This demonstrates the HDRL method's robustness and adaptability in resource allocation.

[0184] from Figure 7 As can be seen in the figure, the V2V link transmission delay of various algorithms shows an upward trend as the V2V payload increases. This is because as the payload increases, channel resources become more scarce, resulting in an increase in data transmission time. When the payload is small (e.g., 2×1060 bytes), the HDRL method achieves the lowest transmission delay, approximately 6 seconds. This demonstrates that the HDRL method effectively utilizes resources and reduces data transmission delay under light load conditions. As the payload increases, the HDRL method's transmission delay gradually increases but remains low. For example, when the payload is 6×1060 bytes, the HDRL method achieves a transmission delay of approximately 15 seconds, significantly lower than other methods. Analysis of V2V link transmission delays under different V2V payload sizes shows that the HDRL method exhibits the lowest transmission delay under all load conditions, demonstrating its efficiency and adaptability in resource allocation. Other methods, however, generally have higher delays because they do not consider latency.

[0185] These experimental results demonstrate that the HDRL approach demonstrates superior performance under varying load conditions, particularly in terms of latency and fairness. Through hierarchical reinforcement learning, the HDRL approach effectively balances resource allocation requirements across various V2V links, improving overall system performance. Compared to traditional multi-agent and single-agent reinforcement learning approaches, the HDRL approach not only maintains high transmission efficiency and reliability under high load conditions, but also significantly improves resource allocation fairness.

[0186] Experiments validated the performance of the proposed channel resource allocation algorithm based on hierarchical multi-agent deep reinforcement learning in various scenarios, evaluating its performance in terms of resource utilization efficiency, communication latency, system robustness, and adaptability. Comparative analysis with traditional methods demonstrated significant improvements in resource allocation fairness, reliability, and overall performance, thus validating the superiority of the proposed method in improving resource utilization efficiency and communication performance.

[0187] Example 2

[0188] This embodiment 2 describes a channel resource allocation system based on hierarchical reinforcement learning. This system is based on the same inventive concept as the channel resource allocation method based on hierarchical reinforcement learning described in the above embodiment 1.

[0189] The channel resource allocation system based on hierarchical reinforcement learning in this embodiment includes the following modules:

[0190] The cellular vehicle network model building module is used to build the cellular vehicle network environment model, including vehicle mobility patterns, communication links, channels and interference models, and to design the fairness model, reliability model and delay model of the cellular vehicle network;

[0191] The cellular vehicle network channel resource allocation module builds a cellular vehicle network channel resource allocation model based on hierarchical reinforcement learning based on the established cellular vehicle network environment model, fairness model, reliability model, and delay model;

[0192] Through hierarchical design, the model decomposes the cellular vehicle network channel resource allocation problem into the upper-layer clustering problem and the lower-layer intra-cluster channel resource optimization problem, and designs the upper-layer controller and the lower-layer controller accordingly.

[0193] The upper-layer controller performs clustering based on the interference graph, assigning V2V links with minimal interference to the same cluster to reduce intra-cluster interference and provide a stable environment for the lower layers to further optimize channel resource allocation. The upper-layer controller uses the DQN method to make clustering decisions, and through continuous learning and adjustment, the clustering scheme is gradually optimized.

[0194] The lower-layer controller optimizes the allocation of channel resources through multi-agent deep reinforcement learning (MARL), which is used to assign appropriate transmission power to each V2V link to maximize link transmission performance while reducing interference. The lower-layer controller adopts the DQN method to continuously optimize the transmission rate and interference management by learning the optimal resource allocation strategy in a dynamic environment.

[0195] It should be noted that the implementation process of the functions and roles of each functional module in the channel resource allocation system based on hierarchical reinforcement learning is specifically described in the implementation process of the corresponding steps in the channel resource allocation method based on hierarchical reinforcement learning in the above embodiment 1, and will not be repeated here.

[0196] Of course, the above description is only a preferred embodiment of the present invention, and the present invention is not limited to the above-mentioned embodiments. It should be noted that all equivalent substitutions and obvious deformation forms made by any technician familiar with this field under the guidance of this specification fall within the substantive scope of this specification and should be protected by the present invention.

Claims

1. A channel resource allocation method based on hierarchical reinforcement learning, characterized in that: The steps include: Step 1. Build an environmental model for the cellular vehicle network, including vehicle mobility patterns, communication links, channels, and interference models, and design fairness, reliability, and latency models for the cellular vehicle network. The construction process of the fairness model is as follows: To ensure the fairness of V2V link resource allocation, Jain's fairness index is used as a fairness indicator; The process of building a reliability model is as follows: Reliability is measured by the transmission success rate of the V2V link, and reliability is described by the outage probability. The outage probability is defined as the probability of failure to successfully transmit data within a specified time, and reliability is defined as the complement of the outage probability. The delay model is constructed as follows: Based on the Age of Information (AOI) model, a latency model for V2V links is designed. AOI measures the time from information generation to reception, taking into account not only transmission time but also the frequency of information updates. Step 2. Based on the cellular vehicle network environment model, fairness model, reliability model, and delay model established in Step 1, a cellular vehicle network channel resource allocation model based on hierarchical reinforcement learning is constructed. Through hierarchical design, the model decomposes the cellular vehicle network channel resource allocation problem into the upper-layer clustering problem and the lower-layer intra-cluster channel resource optimization problem, and designs the upper-layer controller and the lower-layer controller accordingly. The upper-layer controller performs clustering based on the interference graph, assigning V2V links with minimal interference to the same cluster to reduce intra-cluster interference and provide a stable environment for the lower layers to further optimize channel resource allocation. The upper-layer controller uses the DQN method to make clustering decisions, and through continuous learning and adjustment, the clustering scheme is gradually optimized. The lower-level controller optimizes the allocation of channel resources through multi-agent deep reinforcement learning (MARL), assigning appropriate transmission power to each V2V link to maximize link transmission performance while reducing interference. The lower-level controller uses the DQN method to continuously optimize the transmission rate and interference management by learning the best resource allocation strategy in a dynamic environment; In step 2, the design process of the lower-layer controller is as follows: In the channel resource optimization decision in the lower-layer controller, the action space is defined as the combination of the transmit power of each V2V link; let is the action space of the kth agent, then the action represents the power allocation scheme of the kth V2V link; Therefore, the action space is designed as follows: Define P i represents the optional transmit power level, i∈[1,N], each action a k It consists of only one transmission power p; The state space includes the channel state information and environment information of each V2V link; is the state space of the kth agent, then the state represents the current state of the th V2V link. The state space is designed as follows: Among them, g k [m] represents the gain of the current channel, I k [m] represents the interference from other links, R k Indicates the current transmission rate, A k represents the current information age, P out represents the probability of interruption; At each coherent time step t, the state of each V2V agent k is subject to observation Z k (t), is defined by the observation function O as: Z k (t)=O(s k (t),k) (28) Among them, s k (t) represents the state observed by the kth agent at time t; Assume r k is the reward function of the kth agent, then the reward function r k (s k ,a k )as follows: r k (s k ,a k )=αR k -βI k -γA k -δP out +ηf(R k ) (29) Among them, R k Indicates the current transmission rate, I k Indicates the interference from other links, A k represents the current information age, P out represents the interruption probability, f(R k ) is the fairness index, and α, β, γ, δ and η are weight parameters.

2. The channel resource allocation method based on hierarchical reinforcement learning according to claim 1, characterized in that: In step 1, the process of constructing the environment model of the cellular vehicle network is as follows: The network architecture of cellular vehicle-to-vehicle connectivity includes V2I links and V2V links. The V2I link connects each vehicle to a base station (BS) or a BS-type roadside unit, while the V2V link provides direct communication between adjacent vehicles. Each V2I link is assigned an independent spectrum sub-band to avoid interference between V2I links; The set of V2I links is represented as The set of V2V links is represented as Where K is the number of V2V links and M is the number of spectrum sub-bands; The channel model is used to describe the signal propagation characteristics in wireless communications, including path loss, small-scale fading, and large-scale fading factors. Since the signal attenuates as the propagation distance increases, the path loss model is defined as: Where PL(d) and PL(d0) are the path losses at distance d and reference distance d0, respectively, and n is the path loss exponent. Small-scale fading describes rapid changes due to multipath propagation and phase interference; large-scale fading describes slowly varying signal strength, i.e., slow changes due to obstacles and terrain changes, including path loss and shadowing effects; The channel power gain g of the kth V2V link on the mth subband k [m] is represented by: g k [m]=α k h k [m](2) where α k represents large-scale fading, h k [m] represents small-scale fading, which obeys exponential distribution and has a mean value of 1; The interference model is used to describe the impact of co-frequency interference on the communication link. The V2V link interference is the interference caused by other V2V links reusing the same spectrum and the interference caused by the V2I link. k [m] is calculated as follows: in, is the transmit power of the mth V2I link, is the interference channel gain of the mth V2I link to the kth V2V link, ρ k′ [m] is an indicator variable indicating whether the k′th V2V link uses the mth spectrum resource block. is the transmission power of the kth V2V link, g k′,k [m] is the interference channel gain of the k′th V2V link to the kth V2V link; The received signal-to-interference-and-noise ratio (SINR) of the V2I link is: in, is the SINR of the mth V2I link, and are the transmission powers of the mth V2I link and the kth V2V link on the mth subchannel, σ 2 is the noise power, g k,B [m] is the interference from the kth V2V link to the base station on the mth subchannel, is the interference from the mth V2I link to the base station on the mth subchannel; ρ k [m] is used to describe whether the k-th V2V link uses the m-th spectrum resource block; Among them, the V2I link rate Expressed as: Where W is the bandwidth of the spectrum subband; The V2V link rate R k Expressed as: in, is the SINR of the kth V2V link.

3. The channel resource allocation method based on hierarchical reinforcement learning according to claim 2, characterized in that: In step 1, the process of constructing the fairness model of the cellular vehicle network is as follows: The definition is as follows: Among them, R k is the rate of the kth V2V link, K is the total number of V2V links; Jain's fairness index ranges from 0 to 1, and the closer the value is to 1, the fairer the resource allocation is; assuming represents the set of all V2V links, and the goal of the fairness model is to maximize Jain's fairness index f(R k ), described as follows: At the same time, when designing the fairness model, that is, formula (8), the following assumptions and constraints are made: Formula (9) indicates that each V2V link can only use one spectrum subband; Formula (10) indicates that the transmission power of each V2V link is limited by the maximum power P max .

4. The channel resource allocation method based on hierarchical reinforcement learning according to claim 2, characterized in that: In step 1, the process of constructing the reliability model of the cellular vehicle network is as follows: First, assume that the SINR threshold of the kth V2V link on the mth subband is γ th , then the interruption probability P out Defined as: in, represents the interruption probability of the kth V2V link, Pr is a probability symbol; represents the SINR of the k-th V2V link on the m-th subband, which is defined as: in, is the transmit power of the kth V2V link on the mth subband, g k [m] is the channel gain, σ 2 is the noise power, I k [m] is the interference power; Substituting into formula (11), we have: Formula (13) is further transformed into: Assuming the channel gain g k [m] follows an exponential distribution with unit mean, and its probability density function for: where λ is the distribution parameter; for an exponential distribution with unit mean, λ = 1; therefore, the cumulative distribution function (CDF) is: Among them, F gk (x) represents the cumulative distribution function; from this, the outage probability is: Assumptions represents the set of all V2V links. The objectives of the reliability model are described as follows: By minimizing the sum of outage probabilities, the reliability of the entire cellular vehicle network can be improved.

5. The channel resource allocation method based on hierarchical reinforcement learning according to claim 2, characterized in that: In step 1, the process of constructing the delay model of the cellular vehicle network is as follows: The definition is as follows: A k (t)=t-U k (t) (19) Among them, A k (t) is the information age of the kth V2V link at time t, U k (t) is the time when the latest information is updated; The update rule of information age is defined as follows: Among them, A k (t) is the information age of the kth V2V link at time t, Δt is the time interval, R k is the transmission rate of the kth V2V link, R min is the minimum transmission rate requirement; At the same time, in order to minimize the average information age of the V2V link, the following objective function is defined: Where K is the number of V2V links, T is the time period, and A k (t) is the information age of the kth V2V link at time t.

6. The channel resource allocation method based on hierarchical reinforcement learning according to claim 1, characterized in that: In step 2, the design process of the upper controller is as follows: In the upper-layer controller, the action space is defined as possible clustering schemes; each action corresponds to a scheme for allocating V2V links to different clusters and selecting an independent subchannel for each cluster; set up is the action space, then the action represents a clustering scheme; specifically, the goal of the clustering scheme is to assign V2V links with less interference to the same cluster in order to better allocate channel resources; Assume there are K V2V links in the network, and the action space of the upper controller is represented as a finite set; Each action represents a specific clustering scheme: Where L is the number of clusters, C i represents the set of V2V links contained in the i-th cluster; Since the clustering process is equivalent to the sub-channel selection process, the M disjoint sub-channels of the cellular vehicle network wireless spectrum resources are the action space of the upper-layer controller; each sub-channel is occupied by a V2I link; The state space includes the state information of all V2V links in the current network; set up is the state space, then the state Represents the current state of the network. The design of the state space is as follows: Where K is the number of V2V links, M is the number of spectrum subbands, and g k [m] represents the channel gain of the k-th V2V link on the m-th subband, I k [m] represents the interference suffered by the k-th V2V link on the m-th subband, C k Indicates the current cluster allocation status of the k-th V2V link; At each coherent time step t, the states of all V2V links are again observed Z(t), defined by the observation function O: Z(t)=O(s t ) (24) Among them, s t represents the state observed at time t; The reward function of the upper-layer controller will be used to evaluate the pros and cons of each clustering scheme. To ensure the effectiveness of clustering and the convergence of training, the reward function should take into account the interference between links and resource utilization. Let r be the reward function, then it is as follows: in, represents the current set of all clusters, C represents one of the clusters, (i, j) represents two V2V links belonging to the same cluster C, g i,j represents the channel gain between V2V links i and j.

7. The channel resource allocation method based on hierarchical reinforcement learning according to claim 1, characterized in that: For the upper-level controller, its Q-value function is expressed as: Among them, Q m (s m ,a m ) represents the Q value function of the upper controller, s m is the state of the upper controller, a m is the action to be performed by the upper controller, r m is the return of the upper controller, γ is the discount factor; A m Represents a set of actions; (s m )′ means in state s m Next, perform action a m The state reached later, a′ represents the state (s m )′ the action to be performed; For the lower-layer controller, its Q-value function is expressed as: Among them, Q c (s c ,a c ) represents the Q value function of the lower controller, s c is the state of the lower controller, a c is the action to be performed by the lower-level controller, r c is the return of the lower controller, γ is the discount factor; A c Represents a set of actions; (s c )′ means in state s c Next, perform action a c The state reached later, a′ represents the state (s c )′ the action to be performed; When training the Q network, the loss function is used to measure the gap between the current Q value and the target Q value, which is defined as follows: Among them, L(θ) represents the loss function, θ is the parameter of the current Q network, θ - are the parameters of the target Q network.

8. The channel resource allocation method based on hierarchical reinforcement learning according to claim 7, characterized in that: In step 2, the process of cellular vehicle network channel resource allocation decision-making based on hierarchical reinforcement learning is as follows: First, randomly initialize the Q networks of all agents; Each V2V link acts as an intelligent agent and independently maintains its Q network for selecting actions and evaluating rewards; In the upper controller, each V2V link is initialized as a single cluster C k = {k}, and construct the interference graph G(V,E), where the nodes in the interference graph G(V,E) represent V2V links and the edge weights are the channel gains g between the links. i,j ; Define the interference graph G(V,E) where V represents the set of nodes and E represents the set of edges; In each iteration, the vehicle positions and large-scale fading parameter α are updated to simulate vehicle movement and environmental changes. For each time slice t, all V2V agents observe the state Z(t), select an action A(t) according to the ∈-greedy strategy, and execute the action. Through continuous iteration and updating, the agents gradually learn the optimal clustering strategy. After all agents take action, they merge link pairs with interference less than a threshold θ and update the cluster allocation scheme. Each cluster is assigned a separate resource block, or subchannel. After executing the clustering scheme, each agent's reward r(t) is calculated. The reward function considers intra-cluster interference and evaluates the clustering scheme by minimizing intra-cluster interference. After a small-scale fading of the update channel, all agents observe the new state Z(t+1) and store the state transition (Z(t), A(t), r(t), Z(t+1)) in the replay memory; In the lower controller, for each time slice t, all V2V agents independently observe the state Z k (t), select action a according to the ∈-greedy strategy k (t), and perform the action, i.e., power allocation; Each V2V agent selects action a k (t) Power distribution operation is performed; action a k (t) represents the power allocation decision of agent k in the current state; after executing the power allocation plan, the local reward r of each agent is calculated k (t); adjust the impact of each factor on the return through weight parameters α, β, γ, δ and η; each agent observes the new state Z k (t+1), and transfer the state (Z k (t),a k (t),r k (t),Z k (t+1)) is stored in the replay memory; Small batches of data are uniformly sampled from the replay memory, and stochastic gradient descent is used to optimize the error between the Q network and the learning target; by continuously updating the Q network, the intelligent agent gradually learns the optimal power allocation strategy.

9. A channel resource allocation system based on hierarchical reinforcement learning, for implementing the channel resource allocation method based on hierarchical reinforcement learning as claimed in claim 1, characterized in that: The channel resource allocation system based on hierarchical reinforcement learning includes the following modules: The cellular vehicle network model building module is used to build the cellular vehicle network environment model, including vehicle mobility patterns, communication links, channels and interference models, and to design the fairness model, reliability model and delay model of the cellular vehicle network; The cellular vehicle network channel resource allocation module builds a cellular vehicle network channel resource allocation model based on hierarchical reinforcement learning based on the established cellular vehicle network environment model, fairness model, reliability model, and delay model; Through hierarchical design, the model decomposes the cellular vehicle network channel resource allocation problem into the upper-layer clustering problem and the lower-layer intra-cluster channel resource optimization problem, and designs the upper-layer controller and the lower-layer controller accordingly. The upper-layer controller performs clustering based on the interference graph, assigning V2V links with minimal interference to the same cluster to reduce intra-cluster interference and provide a stable environment for the lower layers to further optimize channel resource allocation. The upper-layer controller uses the DQN method to make clustering decisions, and through continuous learning and adjustment, the clustering scheme is gradually optimized. The lower-level controller optimizes the allocation of channel resources through multi-agent deep reinforcement learning (MARL), assigning appropriate transmission power to each V2V link to maximize link transmission performance while reducing interference. The lower-layer controller adopts the DQN method to continuously optimize the transmission rate and interference management by learning the best resource allocation strategy in a dynamic environment.