Channel resource allocation method and system based on hierarchical reinforcement learning
By adopting a channel resource allocation method based on hierarchical reinforcement learning in cellular vehicle networking, the problems of communication interference and unfair resource allocation are solved, efficient and fair resource allocation is achieved, and communication performance and traffic safety are improved.
Patent Information
- Application Number
- CN202510098848.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-22
AI Technical Summary
In the cellular network of vehicles, communication interference between vehicles and between vehicles and infrastructure is serious, resulting in a decline in communication quality, an increase in data transmission delay, and may even lead to information loss, affecting the normal operation of vehicles and traffic safety.
The channel resource allocation method based on hierarchical reinforcement learning is adopted, and the fairness, reliability and delay model is designed by building an environmental model of cellular vehicle networking, and the channel resource allocation problem is decomposed by multi-agent deep reinforcement learning to optimize channel resource allocation problems to upper-level clustering problems and channel resource optimization problems in the lower-level clustering, and spectrum resources and transmission power allocation are optimized.
It realizes fair and efficient allocation of resources in complex environments, reduces communication delays, improves the robustness and adaptability of the system, and significantly improves the communication performance and user experience of cellular vehicle networking.
Smart Images

Figure CN119967616A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent transportation and relates to a channel resource allocation method and system based on hierarchical reinforcement learning. Background Art
[0002] In recent years, with the rapid advancement of science and technology and the increasing complexity of urban traffic, intelligent transportation systems have gradually become an important means to improve urban traffic conditions. In the current wave of development of intelligent transportation systems, cellular vehicle-to-everything networks, as one of the key technologies, are promoting innovation in the field of transportation. Cellular vehicle-to-everything networks (Cellular V2X Networks), as an emerging communication technology, realizes real-time communication between vehicles and the surrounding environment (including other vehicles, infrastructure, pedestrians, etc.) through cellular networks, greatly ensuring traffic safety and improving travel efficiency.
[0003] Cellular vehicle networking enables all-round communication between vehicles (V2V) and vehicles and infrastructure (V2I) through cellular networks. These communication methods not only improve traffic safety and efficiency, but also provide a solid foundation for autonomous driving and intelligent traffic management. Especially in V2V communication, direct information exchange between vehicles can realize the sharing of real-time traffic information, thereby improving the overall efficiency of road traffic. With the rapid development of intelligent transportation systems and autonomous driving technologies, building an efficient and reliable V2V communication network has become a research focus.
[0004] However, in the actual application of cellular vehicle networks, the problem of communication interference between vehicles and between vehicles and infrastructure has become increasingly prominent. In particular, the mutual interference between V2V and V2I links may lead to a decline in communication quality, increased data transmission delays, and even information loss, which will not only affect the normal operation of the vehicle, but may also bring certain safety hazards. In order to alleviate the problem of communication interference, it is particularly important to reasonably allocate spectrum resources and transmission power. Spectrum allocation involves how to reasonably allocate limited spectrum resources to different communication links to minimize interference and increase system capacity. Power allocation requires reasonable adjustment of the transmission power of each link while ensuring communication quality.
[0005] However, in the process of spectrum and power allocation, unfairness often occurs. Some links may occupy more resources, while other links are resource-poor, resulting in a decline in the overall performance of the system. At present, the stability and reliability of communication links under high-speed vehicle movement and complex environments remain a challenge. At the same time, the fairness and efficiency of resource allocation also need to be studied in depth, especially in high-density traffic scenarios. How to achieve efficient utilization and fair allocation of spectrum resources is a key issue. In addition, the delay problem is particularly prominent in V2V communication, because real-time is an important factor in ensuring traffic safety and efficiency, and existing research still has many shortcomings in how to reduce delays and improve communication reliability and fairness. Summary of the invention
[0006] The purpose of the present invention is to propose a channel resource allocation method based on hierarchical reinforcement learning. The method first constructs a network model of vehicle-to-vehicle links in cellular vehicle networks, and describes the channel model, interference model, fairness model, reliability model and delay model in detail. On this basis, a cellular vehicle network channel resource allocation model based on hierarchical reinforcement learning is further proposed. The cellular vehicle network channel resource allocation model has two-layer control strategies with different time scales to reasonably allocate spectrum resources and transmission power, ensuring fair and efficient allocation of resources in complex environments.
[0007] In order to achieve the above-mentioned problem, the present invention adopts the following technical solution:
[0008] The channel resource allocation method based on hierarchical reinforcement learning includes the following steps:
[0009] Step 1. Build an environmental model of the cellular vehicle network, including vehicle mobility patterns, communication links, channels, and interference models, and design a fairness model, reliability model, and delay model for the cellular vehicle network.
[0010] Step 2. Based on the cellular vehicle network environment model, fairness model, reliability model and delay model built in step 1, a cellular vehicle network channel resource allocation model based on hierarchical reinforcement learning is constructed;
[0011] Through hierarchical design, the model decomposes the cellular vehicle network channel resource allocation problem into the upper-layer clustering problem and the lower-layer intra-cluster channel resource optimization problem, and designs the upper-layer controller and the lower-layer controller accordingly.
[0012] The upper-layer controller performs clustering based on the interference graph and allocates V2V links with less interference to the same cluster to reduce intra-cluster interference and provide a stable environment for the lower layer, which facilitates the lower layer to further optimize channel resource allocation. The upper-layer controller uses the DQN method to make clustering decisions and gradually optimizes the clustering scheme through continuous learning and adjustment.
[0013] The lower-level controller optimizes the allocation of channel resources through multi-agent deep reinforcement learning (MARL) to allocate appropriate transmission power to each V2V link to maximize link transmission performance while reducing interference. The lower-level controller adopts the DQN method to continuously optimize the transmission rate and interference management by learning the optimal resource allocation strategy in a dynamic environment.
[0014] In addition, based on the above-mentioned channel resource allocation method based on hierarchical reinforcement learning, the present invention also proposes a corresponding channel resource allocation system based on hierarchical reinforcement learning, which adopts the following technical solutions:
[0015] The channel resource allocation system based on hierarchical reinforcement learning includes the following modules:
[0016] The cellular vehicle network model building module is used to build the environmental model of the cellular vehicle network, including vehicle mobility mode, communication link, channel and interference model, and design the fairness model, reliability model and delay model of the cellular vehicle network;
[0017] As well as the cellular vehicle network channel resource allocation module, based on the established cellular vehicle network environment model, fairness model, reliability model and delay model, a cellular vehicle network channel resource allocation model based on hierarchical reinforcement learning is constructed;
[0018] Through hierarchical design, the model decomposes the cellular vehicle network channel resource allocation problem into the upper-layer clustering problem and the lower-layer intra-cluster channel resource optimization problem, and designs the upper-layer controller and the lower-layer controller accordingly.
[0019] The upper-layer controller performs clustering based on the interference graph and allocates V2V links with less interference to the same cluster to reduce intra-cluster interference and provide a stable environment for the lower layer, which facilitates the lower layer to further optimize channel resource allocation. The upper-layer controller uses the DQN method to make clustering decisions and gradually optimizes the clustering scheme through continuous learning and adjustment.
[0020] The lower-level controller optimizes the allocation of channel resources through multi-agent deep reinforcement learning (MARL) to allocate appropriate transmission power to each V2V link to maximize link transmission performance while reducing interference. The lower-level controller adopts the DQN method to continuously optimize the transmission rate and interference management by learning the optimal resource allocation strategy in a dynamic environment.
[0021] The present invention has the following advantages:
[0022] As described above, the present invention relates to a channel resource allocation method and system based on hierarchical reinforcement learning. Among them, the present invention constructs a cellular vehicle network V2V link network model. Specifically, the present invention proposes a fairness model based on Jain's fairness index to measure the fairness of resource allocation and ensure that the resource allocation between different vehicles is relatively balanced; the present invention establishes a reliability model, and by introducing the concept of interruption probability, measures the probability of successful transmission of information within a specified time, thereby ensuring the communication needs of the V2V link; the present invention designs a delay model based on the information age (AOI) model to evaluate the time from information generation to reception, so as to ensure the timely transmission of information. In addition, the present invention also constructs a cellular vehicle network channel resource allocation model based on hierarchical reinforcement learning. The model decomposes the cellular vehicle network channel resource allocation problem into the clustering problem of the upper controller and the intra-cluster channel resource optimization problem of the lower controller through hierarchical design. Among them, in the upper algorithm, a clustering algorithm based on interference graph is proposed. By constructing the interference graph, the V2V links with less interference are allocated to the same cluster to reduce the intra-cluster interference. Clustering decision-making is performed through a multi-agent deep reinforcement learning method to optimize interference management. In the lower-level algorithm, a channel resource optimization algorithm based on multi-agent reinforcement learning is designed. For the V2V link within each cluster, the transmission rate is maximized and the interference is minimized by optimizing the allocation of sub-channels and transmission power. The present invention is based on a cellular vehicle network channel resource allocation model based on hierarchical multi-agent deep reinforcement learning, which reasonably allocates spectrum resources and transmission power, ensuring fair and efficient allocation of resources in complex environments. In addition, the present invention also verifies the performance of the proposed channel resource allocation method based on hierarchical reinforcement learning in different scenarios through simulation experiments, and evaluates its performance in terms of resource utilization efficiency, communication delay, system robustness and adaptability. By comparing and analyzing with traditional methods, it is demonstrated that the method of the present invention significantly improves the fairness, reliability and overall performance of resource allocation. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 This is a basic network architecture diagram of a cellular vehicle network in an embodiment of the present invention;
[0024] Figure 2 A multi-time scale control strategy diagram for allocating channel resources of a cellular vehicle network in an embodiment of the present invention;
[0025] Figure 3 This is a hierarchical architecture diagram of cellular vehicle network channel resource allocation in an embodiment of the present invention;
[0026] Figure 4 It is a schematic diagram of the rewards obtained from the V2V link training as the number of iterations increases in a specific example of the present invention;
[0027] Figure 5It is a schematic diagram for comparing the sum of V2I link rates of different V2V load sizes in a specific example of the present invention;
[0028] Figure 6 It is a schematic diagram of fairness comparison of V2V links with different V2V load sizes in a specific example of the present invention;
[0029] Figure 7 Schematic diagram of transmission delay of V2V links with different V2V load sizes in a specific example of the present invention. DETAILED DESCRIPTION
[0030] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:
[0031] Example 1
[0032] With the development of intelligent transportation systems, cellular vehicle networking technology has gradually become an important means to improve traffic safety and efficiency. However, due to the large number of vehicles, fast moving speed and complex interference in cellular vehicle networking, how to efficiently and reliably allocate wireless resources has become an urgent problem to be solved. The present invention proposes a channel resource allocation method based on hierarchical reinforcement learning to ensure fair and efficient allocation of resources in a complex environment. First, a network model of vehicle-to-vehicle links in cellular vehicle networking is constructed, and the channel model, interference model, fairness model, reliability model and delay model are described in detail. On this basis, a cellular vehicle networking channel resource allocation model based on hierarchical reinforcement learning is proposed, which has a two-layer control strategy with different time scales. In order to implement the two-layer control strategy, the channel resource allocation problem is decomposed into two sub-problems of minimizing the interference of vehicle-to-vehicle links and maximizing the transmission rate of links between vehicles. A clustering algorithm for vehicle-to-vehicle links based on interference graphs is proposed at the upper level. The links between vehicles are clustered by a multi-agent deep reinforcement learning method to optimize interference management. A channel resource optimization algorithm based on multi-agent reinforcement learning is designed at the lower layer. For the vehicle-to-vehicle links within each cluster, the transmission rate is maximized and the interference is minimized by optimizing the allocation of transmission power.
[0033] The channel resource allocation method based on hierarchical reinforcement learning in this embodiment specifically includes the following steps:
[0034] Step 1. Build an environmental model of the cellular vehicle network, including vehicle mobility patterns, communication links, channels, and interference models, and design the fairness model, reliability model, and delay model of the cellular vehicle network.
[0035] Specifically, we first established a cellular vehicle network environment model, including vehicle mobility patterns, communication links, channels, and interference models, taking into account vehicle motion characteristics, path loss, shadow effects, and multipath effects in different scenarios, as well as co-channel interference and adjacent channel interference. Next, we designed a fairness model to ensure that different vehicles are treated fairly in the allocation of spectrum resources. Then, we constructed a reliability model for the communication link to evaluate the stability of the communication link and the success rate of data transmission, and to ensure the reliability of communication under different traffic conditions. Finally, we designed a delay model to analyze the delay of data transmission, especially in high-density traffic scenarios, and to study how to effectively reduce the delay to meet the needs of real-time communication.
[0036] Through detailed analysis of the environmental model, fairness model, reliability model and delay model, we aim to build a comprehensive model that can more accurately simulate the actual scenarios and needs of cellular vehicle networks. This will not only provide an important basis for the design and implementation of subsequent optimization algorithms, but also lay a solid foundation for the channel resource allocation algorithm based on hierarchical reinforcement learning.
[0037] In order to design and evaluate the resource allocation method in cellular vehicle networking, the present invention constructs a simulation environment that can simulate actual cellular vehicle networking scenarios, including vehicle movement patterns, communication links, channels, and interference models.
[0038] The following introduces the basic architecture of cellular vehicle networking, which mainly includes V2I (Vehicle-to-Infrastructure) link and V2V (Vehicle-to-Vehicle) link. Figure 1 As shown, the V2I link connects each vehicle to a base station (BS) or BS-type roadside unit, while the V2V link provides direct communication between neighboring vehicles.
[0039] Each V2I link is assigned an independent spectrum sub-band to avoid interference between V2I links.
[0040] The V2I link is mainly used for high data rate entertainment services. Specifically, vehicles communicate with base stations (BS) through cellular networks to transmit high data rate entertainment content such as video streaming and online games.
[0041] The set of V2I links is represented as M is the number of spectral subbands.
[0042] The V2V link is mainly used for the transmission of periodically generated safety messages, such as direct communication between vehicles, exchanging location information, speed and acceleration, etc., to improve road safety and traffic efficiency.
[0043] The set of V2V links is represented as K is the number of V2V links.
[0044] In this cellular vehicle network architecture, in order to effectively manage and optimize the use of spectrum resources, a spectrum sharing mechanism is adopted, allowing the V2V link to reuse the spectrum sub-band of the V2I link, thereby improving the spectrum utilization of the overall system under limited spectrum resources.
[0045] However, due to the reuse of V2V links, interference may be introduced, which requires the design of effective interference management strategies to ensure the communication quality and reliability between different links.
[0046] The channel model is used to describe the signal propagation characteristics in wireless communications, including path loss, small-scale fading, and large-scale fading factors. The signal will attenuate as the propagation distance increases, so the path loss model is defined as:
[0047]
[0048] Where PL(d) and PL(d0) are the path losses at distance d and reference distance d0 respectively, and n is the path loss exponent.
[0049] The path loss model is used in this design to calculate signal strength, manage interference, perform link budget analysis, and optimize resource allocation. For example, in an urban environment, buildings and other obstacles can significantly increase path loss.
[0050] The signal fluctuates rapidly due to multipath effects, so the Rayleigh fading model is used. The channel gain h k [m] follows an exponential distribution with a mean of 1. Small-scale fading mainly describes rapid changes caused by multipath propagation and phase interference. For example, when a vehicle travels under an overpass or in a tunnel, the signal will experience severe small-scale fading.
[0051] Signals vary slowly due to obstruction and terrain changes, including path loss and shadowing. Large-scale fading describes slow changes in signal strength, such as shadowing caused by buildings, trees, and other large obstacles.
[0052] In this environment, the channel power gain g of the kth V2V link on the mth subband is k [m] is represented by:
[0053] g k [m] = α k h k [m](2)
[0054] where α k represents large-scale fading (including path loss and shadowing effects), h k[m] represents small-scale fading, which follows an exponential distribution with a mean of 1 (i.e., it is assumed to be an exponential distribution with a mean of 1).
[0055] The interference model is used to describe the impact of co-frequency interference on the communication link. The V2V link interference is the interference caused by other V2V links reusing the same spectrum and the interference caused by the V2I link. k [m] is calculated as follows:
[0056]
[0057] in, is the transmit power of the mth V2I link, is the interference channel gain of the mth V2I link to the kth V2V link, ρ k′ [m] is an indicator variable indicating whether the k′th V2V link uses the mth spectrum resource block. is the transmission power of the kth V2V link, g k′,k [m] is the interference channel gain of the k′th V2V link to the kth V2V link. For example, when multiple vehicles communicate on the same spectrum resource block, serious interference will occur, affecting the communication quality.
[0058] There is no interference between V2I links, but there will be interference caused by the spectrum reused by V2V links.
[0059] The received signal-to-interference-to-noise ratio (SINR) of the V2I link is:
[0060]
[0061] in, is the SINR of the mth V2I link, and are the transmission powers of the mth V2I link and the kth V2V link on the mth subchannel, σ 2 is the noise power, g k,B [m] is the interference from the kth V2V link to the base station on the mth subchannel, is the interference from the mth V2I link to the base station on the mth subchannel.
[0062] ρ k [m] indicates whether the kth V2V link uses the mth spectrum resource block, ρ k [m] is a binary variable. k [m] = 1 means that the kth V2V link selects the mth subband, and ρ k [m]=0 means the opposite.
[0063] When building a cellular vehicle network environment model, the evaluation of channel rate and link performance is a key part.
[0064] Among them, the V2I link rate It is expressed as:
[0065] Where W is the bandwidth of the spectrum subband; is the SINR of the mth V2I link. A high SINR means a higher link rate, which is suitable for high data rate entertainment services.
[0066] The V2V link rate R k It is expressed as:
[0067] in, is the SINR of the kth V2V link.
[0068] In cellular vehicle networks, the fairness of channel resource allocation is of great significance for improving network performance and user experience. This section will design a fairness model for V2V links to ensure that channel resource allocation is fair between different vehicles. In order to ensure the fairness of V2V link resource allocation, Jain's fairness index is used as a fairness indicator.
[0069] Jain's Fairness Index f(R k ) is defined as follows:
[0070]
[0071] Among them, R k is the rate of the kth V2V link, K is the total number of V2V links; Jain's fairness index ranges from 0 to 1, and the closer the value is to 1, the fairer the resource allocation is. In order to achieve fairness in resource allocation in cellular vehicle networks, this paper mathematically models the fairness model. Assume represents the set of all V2V links. The goal of the fairness model is to maximize Jain's fairness index f(R k ), described as follows:
[0072]
[0073] At the same time, when designing the fairness model, that is, formula (8), the following assumptions and constraints are made:
[0074]
[0075] Where M represents the number of spectrum sub-bands, and formula (9) indicates that each V2V link can only use one spectrum sub-band;
[0076]
[0077] Formula (10) indicates that the transmission power of each V2V link is limited by the maximum power P max .
[0078] The fairness model of the cellular vehicle network constructed by the present invention takes into account the interference between V2V links and between V2V and V2I links, so that the interference power should be kept within an acceptable range.
[0079] In cellular vehicle networks, reliability refers to the probability that information can be successfully transmitted within a specified time. This is very important for the delivery of safe messages. In order to ensure that the communication requirements of each V2V link can be met, a reliability model will be designed. Reliability can be measured by the transmission success rate of the V2V link, and the interruption probability will be used to describe reliability. The interruption probability can be defined as the probability of failing to successfully transmit data within a specified time.
[0080] Accordingly, reliability can be defined as the complement of outage probability. Next, the outage probability will be derived.
[0081] First, assume that the SINR threshold of the kth V2V link on the mth subband is γ th , then the interruption probability P out Defined as:
[0082]
[0083] in, represents the interruption probability of the kth V2V link, Pr is a probability symbol; represents the SINR of the k-th V2V link on the m-th subband and is defined as:
[0084]
[0085] in, is the transmit power of the kth V2V link on the mth subband, g k [m] is the channel gain, σ 2 is the noise power, I k [m] is the interference power; Substituting into formula (11), we have:
[0086]
[0087] Formula (13) can be transformed into:
[0088]
[0089] Assuming the channel gain g k[m] follows an exponential distribution with unit mean, and its probability density function for:
[0090]
[0091] where λ is the distribution parameter; for an exponential distribution with unit mean, λ = 1; therefore the cumulative distribution function CDF is:
[0092]
[0093] Among them, F gk (x) represents the cumulative distribution function; hence, the outage probability is:
[0094]
[0095] Assumptions represents the set of all V2V links. The goal of the reliability model is described as follows:
[0096]
[0097] By minimizing the sum of outage probabilities, the reliability of the entire cellular vehicle network can be improved.
[0098] Latency is an important indicator of cellular vehicle network performance, especially in vehicle-to-vehicle communication, where low latency is crucial to improving driving safety and information transmission efficiency. This section will design a delay model for V2V links based on the Age of Information (AOI) model to ensure timely information transmission.
[0099] Age of Information (AOI) measures the time from when information is generated to when it is received. Unlike traditional latency metrics, AOI not only takes into account the transmission time, but also the frequency of information updates, and is defined as follows:
[0100] A k (t) = tU k (t) (19)
[0101] Among them, A k (t) is the information age of the kth V2V link at time t, U k (t) is the time when the latest information is updated.
[0102] The update rule of information age is defined as follows:
[0103]
[0104] Among them, A k (t) is the information age of the kth V2V link at time t, Δt is the time interval, Rk is the transmission rate of the kth V2V link, R min Is the minimum transmission rate requirement.
[0105] At the same time, in order to minimize the average information age of the V2V link, the following objective function is defined:
[0106]
[0107] Where K is the number of V2V links, T is the time period, and A k (t) is the information age of the kth V2V link at time t.
[0108] The above step 1 constructs the cellular vehicle network V2V link network model, including the design of fairness model, reliability model and delay model. Through these models, the foundation is laid for achieving fair allocation of spectrum and power resources, improving communication reliability and reducing delay. These models will be applied to the design of specific channel resource allocation methods in subsequent steps.
[0109] Step 2. Based on the cellular vehicle network environment model, fairness model, reliability model and delay model built in step 1, a cellular vehicle network channel resource allocation model based on hierarchical reinforcement learning is constructed.
[0110] When dealing with a large number of agents, traditional multi-agent deep reinforcement learning (MARL) causes the state space and action space to grow exponentially because each agent must consider the strategies and possible actions of other agents. In addition, in MARL, the learning environment of the agent is composed of other agents that are learning, which makes the learning environment unstable, thus affecting the convergence and stability of the training process, resulting in the so-called "environmental non-stationarity" problem. At the same time, since agents in MARL must simultaneously learn how to interact with other agents, this may lead to low learning efficiency.
[0111] Therefore, this paper proposes a cellular vehicle network channel resource allocation model based on hierarchical reinforcement learning to cope with the challenges of limited spectrum resources and interference management. Through hierarchical design, this model decomposes the cellular vehicle network channel resource allocation problem into the upper-layer clustering problem and the lower-layer intra-cluster channel resource optimization problem, aiming to ensure the reliability of vehicle-to-vehicle communication (V2V) and vehicle-to-infrastructure (V2I) communication while maximizing the overall network performance.
[0112] Figure 2 The architecture for handling vehicle-to-vehicle (V2V) communications at different time scales is presented. At large time scales, the high-level control strategy is responsible for system-level V2V link clustering. At small time scales, the low-level control strategy is responsible for link-level intra-cluster channel resource optimization to ensure communication continuity and reliability while improving system performance.
[0113] Figure 3 The process of the vehicle channel resource allocation method based on hierarchical reinforcement learning is demonstrated. The cellular vehicle network channel resource allocation model based on hierarchical reinforcement learning decomposes the cellular vehicle network channel resource allocation problem into the clustering problem of the upper-level controller and the intra-cluster channel resource optimization problem of the lower-level controller through hierarchical design.
[0114] The upper-layer controller performs clustering based on the interference graph and assigns V2V links with less interference to the same cluster to reduce intra-cluster interference, providing a stable environment for the lower layer to further optimize channel resource allocation. The upper-layer algorithm uses the deep Q network (DQN) to make clustering decisions, and through continuous learning and adjustment, the clustering scheme is gradually optimized.
[0115] The lower-level controller optimizes the allocation of channel resources through multi-agent deep reinforcement learning (MAR L). The goal of the lower layer is to allocate appropriate transmission power to each V2V link to maximize the link transmission performance while reducing interference. The lower-level algorithm also uses the deep Q network (DQN) method to continuously optimize the transmission rate and interference management by learning the optimal resource allocation strategy in a dynamic environment.
[0116] The following is a detailed introduction to the design of the upper-layer controller based on the clustering method based on the interference graph. The goal of the upper-layer controller design is to assign V2V links with less interference to the same cluster, that is, to select the same subchannel to reduce intra-cluster interference, thereby providing good initial conditions for the channel resource optimization decision in the lower-layer controller. The specific design process is as follows:
[0117] Assume that there are several V2V links in a network, and the degree of interference between each link is different. By constructing an interference graph, the interference relationship between links can be intuitively represented. In the interference graph, the nodes represent V2V links, and the weights of the edges represent the channel gains between the links. The reward function penalizes the sum of the interference within each cluster, thereby guiding the agent to select the optimal clustering scheme to minimize the interference within the cluster.
[0118] In the upper-level controller, the action space is defined as possible clustering schemes; each action corresponds to a scheme for allocating V2V links to different clusters and selecting an independent subchannel for each cluster.
[0119] set up is the action space, then the action represents a clustering scheme; specifically, the goal of the clustering scheme is to allocate V2V links with less interference to the same cluster so as to better allocate channel resources.
[0120] Assume that there are K V2V links in the network, and the action space of the upper controller is represented as a finite set, where each action represents a specific clustering scheme:
[0121]
[0122] Where L is the number of clusters, C i represents the set of V2V links contained in the i-th cluster. The V2V links in each cluster have a small interference degree. Since the clustering process is equivalent to the process of selecting subchannels, the M non-overlapping subchannels of the cellular vehicle network wireless spectrum resources are the action space of the upper controller; each subchannel is occupied by a V2I link.
[0123] The state space includes the state information of all V2V links in the current network, such as the channel gain, interference and current cluster allocation of each link.
[0124] set up is the state space, then the state Represents the current state of the network. The design of the state space is as follows:
[0125]
[0126] Where K is the number of V2V links, M is the number of subbands, and g k [m] represents the channel gain of the kth V2V link on the mth subband, I k [m] represents the interference suffered by the kth V2V link on the mth subband, C k Indicates the current cluster allocation status of the kth V2V link.
[0127] At each coherent time step t, the states of all V2V links are again observed Z(t), defined by the observation function O:
[0128] Z(t)=O(s t ) (twenty four)
[0129] Among them, s t Represents the state at time t.
[0130] The reward function of the upper-layer controller will be used to evaluate the pros and cons of each clustering scheme. To ensure the effectiveness of clustering and the convergence of training, the reward function should take into account the interference between links and the utilization of resources.
[0131] Let r be the reward function, then it is as follows:
[0132]
[0133] in, represents the current set of all clusters, C represents one of the clusters, (i, j) represents two V2V links belonging to the same cluster C, and g i,j represents the channel gain between V2V links i and j. By minimizing the total interference within each cluster, effective management of interference can be achieved and the reward function can be ensured to gradually converge during training.
[0134] After the upper-layer controller makes the decision, the V2V links in the cellular vehicle network have been reasonably allocated to different clusters, that is, different sub-channels have been selected. Next, the goal of the lower-layer controller design is to select the appropriate transmission power for each V2V link to maximize the link transmission performance while minimizing interference.
[0135] The design process of the lower-level controller design is introduced in detail below.
[0136] In the channel resource optimization decision in the lower-layer controller, the action space is defined as the combination of the transmit power of each V2V link; let is the action space of the kth agent, then the action represents the power allocation scheme of the kth V2V link.
[0137] Therefore, the action space is designed as follows:
[0138]
[0139] Define P i represents the optional transmission power level, i∈[1,N], each action a k It only consists of one transmission power p; the present invention limits the power control to four levels, namely [23dBm, 10dBm, 5dBm, -100dBm].
[0140] Among them, -100dBm means that the V2V transmission power is zero.
[0141] The state space includes the channel state information and environment information of each V2V link; is the state space of the kth agent, then the state Represents the current state of the th V2V link, and the state space is designed as follows:
[0142]
[0143] Among them, g k [m] represents the gain of the current channel, I k [m] represents the interference from other links, R k Indicates the current transmission rate, A k represents the current information age, Pout represents the interruption probability; at each coherent time step t, the state of each V2V agent k is observed again by Z k (t), is defined by the observation function O as:
[0144] Z k (t) = O(s k (t),k) (28)
[0145] Among them, s k (t) represents the state of V2V link k at time t.
[0146] The reward function in the lower-layer controller is used to evaluate the pros and cons of each channel resource allocation scheme. In order to ensure the effectiveness of channel resource allocation, the reward function should comprehensively consider factors such as transmission rate, interference, delay, reliability and fairness.
[0147] Assume r k is the reward function of the kth agent, then the reward function r k (s k ,a k )as follows:
[0148] r k (s k ,a k )=αR k -βI k -γA k -δP out +ηf(R k ) (29)
[0149] Among them, R k Indicates the current transmission rate, I k Indicates the interference from other links, A k represents the current information age, P out represents the interruption probability, f(R k ) is the fairness index, and α, β, γ, δ, and η are weight parameters used to balance the impact of transmission rate, interference, delay, reliability, and fairness on the return.
[0150] The following is a detailed introduction to the cellular vehicle network channel resource allocation decision method based on hierarchical reinforcement learning. This method combines the upper controller and the lower controller to improve the communication performance of the cellular vehicle network by optimizing the channel resource allocation.
[0151] The Q model and loss function model will be described in detail below, and the specific steps of the method will be given.
[0152] The Q model is the core model in reinforcement learning, which evaluates the value of taking a certain action in a given state through the Q value. The hierarchical reinforcement learning model in the present invention includes an upper controller and a lower controller, and each controller has its own Q model.
[0153] For the upper-level controller, its Q-value function is expressed as:
[0154]
[0155] Among them, Q m (s m ,a m ) represents the Q value function of the upper controller, s m is the state of the upper controller, a m is the action to be performed by the upper controller, r m is the return of the upper controller, γ is the discount factor; A m Represents a collection of actions; (s m )′ means in state s m Next, perform action a m The state reached later, a′ represents the state in which m )' to perform the action.
[0156] For the lower-layer controller, its Q-value function is expressed as:
[0157]
[0158] Among them, Q c (s c ,a c ) represents the Q value function of the lower controller, s c is the state of the lower controller, a c is the action to be performed by the lower-level controller, r c is the return of the lower controller, γ is the discount factor; A c Represents a collection of actions; (s c )′ means in state s c Next, perform action a c The state reached later, a′ represents the state in which c )' to perform the action.
[0159] When training the Q network, the loss function is used to measure the gap between the current Q value and the target Q value, which is defined as follows:
[0160]
[0161] Among them, L(θ) represents the loss function, θ is the parameter of the current Q network, and θ - are the parameters of the target Q network.
[0162] The specific process of cellular vehicle network channel resource allocation decision-making based on hierarchical reinforcement learning is given below.
[0163] First, the Q networks of all agents are randomly initialized; each V2V link acts as an agent and independently maintains its Q network for selecting actions and evaluating rewards; the initialization of the Q network is crucial to the effectiveness of the algorithm, and reasonable initialization can accelerate the training process.
[0164] In the upper controller, each V2V link is initialized as a single cluster C k = {k}, and construct an interference graph G(V,E); the nodes in the graph represent V2V links, and the edge weights are the channel gains g between the links i,j .
[0165] In each iteration, the vehicle position and large-scale fading parameter α are updated to simulate vehicle movement and environmental changes; for each time slice t, all V2V agents observe the state Z(t), select action A(t) according to the ∈-greedy strategy, and perform the action, i.e., clustering operation. Through continuous iteration and updating, the agent gradually learns the optimal clustering strategy.
[0166] After all agents take action, link pairs with interference less than the threshold θ are merged and the cluster allocation scheme {C k}; Each cluster is allocated an independent resource block, i.e. a subchannel, to ensure that there is no interference between clusters; this process gradually optimizes the clustering scheme to minimize the interference within each cluster to improve the overall network performance.
[0167] After executing the clustering scheme, the reward r(t) (i.e., the upper-level reward) of each agent is calculated; the reward function takes into account the interference within the cluster and evaluates the pros and cons of the clustering scheme by minimizing the interference within the cluster.
[0168] After updating the small-scale fading of the channel, all agents observe the new state Z(t+1) and store the state transition (Z(t), A(t), r(t), Z(t+1)) in the replay memory D. By storing and replaying historical state transitions, agents can repeatedly use past experience during training and accelerate the learning process.
[0169] In the lower controller, for each time slice t, all V2V agents independently observe the state Z k (t), select action a according to the ∈-greedy strategy k (t), and performs the action, i.e., power allocation.
[0170] Each V2V agent selects action a k (t) Power distribution operation is performed; action a k(t) represents the power allocation decision of agent k in the current state; after executing the power allocation plan, the local reward r of each agent is calculated k (t) (i.e., the lower-level reward), the reward function comprehensively considers the transmission rate, interference, delay, reliability, and fairness, and adjusts the impact of each factor on the reward through the weight parameters α, β, γ, δ, and η.
[0171] Each agent observes a new state Z k (t+1), and transfer the state to (Z k (t),a k (t),r k (t),Z k (t+1)) is stored in the playback memory D k From playback memory D k Small batches of data are uniformly sampled in the algorithm, and the error between the Q network and the learning target is optimized using stochastic gradient descent. By continuously updating the Q network, the agent gradually learns the optimal power allocation strategy.
[0172] The present invention can improve the overall communication performance of the cellular vehicle network while ensuring fairness and reliability by jointly optimizing the Q value functions of the upper-layer controller and the lower-layer controller.
[0173] In the upper-level controller, effective allocation of subchannels is achieved by constructing an interference graph and a clustering strategy based on interference minimization; in the lower-level controller, the multi-agent reinforcement learning method is used to optimize the allocation of transmission power, maximize the transmission rate and minimize interference and delay. Through the hierarchical design, the entire system not only improves the resource utilization efficiency, but also enhances the robustness and adaptability of the system, providing a strong guarantee for the realization of efficient and reliable cellular vehicle network communication.
[0174] The present invention designs and implements a fair spectrum and power allocation model, which can significantly reduce the interference between V2V and V2I links, thereby improving the reliability of communication. This is of great significance for ensuring real-time information exchange between vehicles and improving traffic safety. By establishing an optimization model for fair resource allocation, it is possible to balance the resource allocation of each communication link while ensuring the overall performance of the system. This not only improves the fairness of the system, but also ensures the stability and reliability of different links, laying a solid foundation for the practical application of cellular vehicle networks. By introducing a multi-agent hierarchical deep reinforcement learning method (which has strong adaptability and efficient computing power, and can provide a new solution for resource allocation in cellular vehicle networks), and designing a spectrum and power allocation algorithm, efficient resource management can be achieved in a dynamic and complex communication environment.
[0175] In addition, in order to verify the effectiveness of the method proposed in the present invention, a simulation experiment is conducted on the model proposed in the present invention to verify the performance of the resource allocation method based on multi-agent collaborative learning. The simulation scenario is constructed based on the Manhattan traffic scenario specified in the 3GPP TR36.885 standard, which contains 9 blocks. Vehicles, base stations, V2I communication links and V2V communication lines are set in the traffic simulation scenario, and the specific experimental parameter settings are shown in Table 1.
[0176] Table 1 Experimental parameter settings
[0177]
[0178] The hierarchical structure of the present invention may cause fluctuations in returns, especially in the early stages, when high-level strategies may frequently change sub-goals, resulting in unstable returns. Figure 4 As can be seen in the graph, the returns fluctuate greatly in the first 250,000 iterations. The peaks and valleys may reflect the effectiveness of the sub-goals selected by the upper-level control strategy. For example, in some cases, the lower-level control strategy may have selected a very effective sub-goal, causing the return to rise rapidly; while in other cases, it may have selected a less appropriate sub-goal, causing the return to fall. The overall upward trend shows that the HDRL system is constantly improving its strategy and gradually learning more effective task allocation and action execution, which shows that both the upper-level and lower-level control strategies are gradually optimized during the learning process.
[0179] Overall, the cumulative rewards obtained by the V2V link as the number of training iterations increases are shown, verifying the convergence behavior of the proposed hierarchical reinforcement learning method. As can be seen from the figure, the reward for each iteration increases with the number of training iterations, proving the effectiveness of the proposed training algorithm. When the total training set reaches approximately 300,000, the reward performance gradually converges despite some fluctuations due to channel fading caused by mobility in the vehicle environment. Based on this conclusion, when evaluating the performance of the V2V link, each agent DQN network is trained 350,000 times, thus ensuring the convergence performance of the learning algorithm.
[0180] The experimental results of the V2V link channel resource allocation algorithm based on hierarchical reinforcement learning are analyzed and discussed in detail below.
[0181] The performance of the proposed algorithm (Hierarchical Deep Reinforcement Learning, HDRL for short) in different scenarios is evaluated by comparing it with traditional methods (i.e. multi-agent deep reinforcement learning (MARL), single-agent deep reinforcement learning (SARL) and random method RANDOM). The experiment mainly evaluates the following aspects: transmission rate, reliability, communication delay, fairness, system robustness and adaptability.
[0182] from Figure 5 It can be seen that although the V2I link rate of the method of the present invention is lower than that of MARL and SARL, it is because the reward function design of MARL and SARL only focuses on transmission rate and reliability, while the method of the present invention considers more factors, which leads to this situation. However, as the load increases, the transmission rate of all methods decreases. This is because the increase in load will cause more resources to be occupied by the V2V link, thereby reducing the resources allocated to the V2I link. On the whole, the performance of the HDRL proposed in the present invention is relatively stable. As the load increases, the sum of the V2I link rates gradually decreases, but the decrease is small. This shows that the HDRL method can better balance the resource allocation of V2V and V2I links under different load conditions, and effectively improve the overall resource utilization efficiency of the system. In the process of gradually increasing load, the sum of the rates of the HDRL method remains at a high level, showing strong robustness and adaptability. At the same time, it can also be seen that as the load gradually increases, the overall decrease of the HDRL method is the smallest.
[0183] Figure 6 The fairness index of V2V links of various algorithms under different V2V load sizes is shown in Figure 2. As can be seen from the figure, there are significant differences in the fairness performance of different algorithms under different load conditions. The fairness index of HDRL remains at a high level under all load conditions, showing its fairness in resource allocation. As the load increases, the fairness index of the HDRL method increases slightly and always remains between 0.25 and 0.3. It can be seen that the fairness index of the HDRL method is significantly higher than that of other methods, indicating that it can effectively allocate resources, so that the resource utilization differences between various V2V links are small. The fairness index of the SARL method is low under all load conditions and is almost 0, and is even worse than the random method overall, while the MARL method has a slightly better effect. The HDRL method can always maintain a high level of fairness under different load conditions to ensure balanced resource allocation between V2V links, reflecting the robustness and adaptability of the HDRL method in resource allocation.
[0184] from Figure 7As can be seen from the figure, as the V2V load increases, the V2V link transmission delay of various algorithms shows an upward trend. This is because as the load increases, channel resources become more scarce, resulting in an increase in the time required for data transmission. When the load is small (such as 2×1060Bytes), the HDRL method has the smallest transmission delay, which is about 6 seconds. This shows that the HDRL method can effectively utilize resources and reduce the delay of data transmission under light load conditions. As the load increases, the transmission delay of the HDRL method gradually increases, but always remains at a low level. For example, when the load is 6×1060Bytes, the transmission delay of the HDRL method is about 15 seconds, which is significantly lower than other methods. By analyzing the transmission delay of the V2V link under different V2V load sizes, it can be seen that the HDRL method shows the lowest transmission delay under all load conditions, showing its high efficiency and adaptability in resource allocation. Since other methods do not consider the delay factor, they generally have higher delays.
[0185] Through these experimental results, it can be seen that the HDRL method shows superior performance under different load conditions, especially in terms of latency and fairness. The HDRL method effectively balances the resource allocation requirements between various V2V links through hierarchical reinforcement learning, improving the overall performance of the system. Compared with traditional multi-agent and single-agent reinforcement learning methods, the HDRL method not only maintains high transmission efficiency and reliability under high load conditions, but also makes significant improvements in the fairness of resource allocation.
[0186] The performance of the proposed channel resource allocation algorithm based on hierarchical multi-agent deep reinforcement learning in different scenarios was verified through experiments, and its performance in resource utilization efficiency, communication delay, system robustness and adaptability was evaluated. A comparative analysis with traditional methods showed that the fairness, reliability and overall performance of resource allocation of the method of the present invention were significantly improved, thus verifying the superiority of the method of the present invention in improving resource utilization efficiency and communication performance.
[0187] Example 2
[0188] This embodiment 2 describes a channel resource allocation system based on hierarchical reinforcement learning, which is based on the same inventive concept as the channel resource allocation method based on hierarchical reinforcement learning described in the above embodiment 1.
[0189] The channel resource allocation system based on hierarchical reinforcement learning in this embodiment includes the following modules:
[0190] The cellular vehicle network model building module is used to build the environmental model of the cellular vehicle network, including vehicle mobility mode, communication link, channel and interference model, and design the fairness model, reliability model and delay model of the cellular vehicle network;
[0191] As well as the cellular vehicle network channel resource allocation module, based on the established cellular vehicle network environment model, fairness model, reliability model and delay model, a cellular vehicle network channel resource allocation model based on hierarchical reinforcement learning is constructed;
[0192] Through hierarchical design, the model decomposes the cellular vehicle network channel resource allocation problem into the upper-layer clustering problem and the lower-layer intra-cluster channel resource optimization problem, and designs the upper-layer controller and the lower-layer controller accordingly.
[0193] The upper-layer controller performs clustering based on the interference graph and allocates V2V links with less interference to the same cluster to reduce intra-cluster interference and provide a stable environment for the lower layer, which facilitates the lower layer to further optimize channel resource allocation. The upper-layer controller uses the DQN method to make clustering decisions and gradually optimizes the clustering scheme through continuous learning and adjustment.
[0194] The lower-level controller optimizes the allocation of channel resources through multi-agent deep reinforcement learning (MARL) to allocate appropriate transmission power to each V2V link to maximize link transmission performance while reducing interference. The lower-level controller adopts the DQN method to continuously optimize the transmission rate and interference management by learning the optimal resource allocation strategy in a dynamic environment.
[0195] It should be noted that the implementation process of the functions and effects of each functional module in the channel resource allocation system based on hierarchical reinforcement learning is specifically described in the implementation process of the corresponding steps in the channel resource allocation method based on hierarchical reinforcement learning in the above-mentioned Example 1, and will not be repeated here.
[0196] Of course, the above description is only a preferred embodiment of the present invention, and the present invention is not limited to the above embodiments. It should be noted that all equivalent substitutions and obvious deformation forms made by any technician familiar with the field under the guidance of this specification fall within the essential scope of this specification and should be protected by the present invention.
Claims
1. A channel resource allocation method based on hierarchical reinforcement learning, characterized in that: The steps include: Step 1. Build an environmental model of the cellular vehicle network, including vehicle mobility patterns, communication links, channels, and interference models, and design a fairness model, reliability model, and delay model for the cellular vehicle network. Step 2. Based on the cellular vehicle network environment model, fairness model, reliability model and delay model built in step 1, a cellular vehicle network channel resource allocation model based on hierarchical reinforcement learning is constructed; Through hierarchical design, the model decomposes the cellular vehicle network channel resource allocation problem into the upper-layer clustering problem and the lower-layer intra-cluster channel resource optimization problem, and designs the upper-layer controller and the lower-layer controller accordingly. The upper-layer controller performs clustering based on the interference graph and allocates V2V links with less interference to the same cluster to reduce intra-cluster interference and provide a stable environment for the lower layer, which facilitates the lower layer to further optimize channel resource allocation. The upper-layer controller uses the DQN method to make clustering decisions and gradually optimizes the clustering scheme through continuous learning and adjustment. The lower-level controller optimizes the allocation of channel resources through multi-agent deep reinforcement learning (MARL) to allocate appropriate transmission power to each V2V link to maximize link transmission performance while reducing interference. The lower-level controller adopts the DQN method to continuously optimize the transmission rate and interference management by learning the best resource allocation strategy in a dynamic environment.
2. The channel resource allocation method based on hierarchical reinforcement learning according to claim 1, characterized in that: In step 1, the process of constructing the environment model of the cellular vehicle network is as follows: The network architecture of cellular vehicle networks includes V2I links and V2V links. The V2I link connects each vehicle to a base station BS or a BS-type roadside unit, while the V2V link provides direct communication between adjacent vehicles. Each V2I link is assigned an independent spectrum sub-band to avoid interference between V2I links; The set of V2I links is represented as The set of V2V links is represented as Where K is the number of V2V links and M is the number of spectrum sub-bands; The channel model is used to describe the signal propagation characteristics in wireless communications, including path loss, small-scale fading, and large-scale fading factors. The signal will attenuate as the propagation distance increases, so the path loss model is defined as: Where PL(d) and PL(d0) are the path losses at distance d and reference distance d0, respectively, and n is the path loss exponent; Small-scale fading describes rapid changes due to multipath propagation and phase interference; large-scale fading describes slowly changing signals in signal strength, i.e., slowly changing signals due to obstacles and terrain changes, including path loss and shadowing effects; The channel power gain g of the kth V2V link on the mth subband k [m] is represented by: g k [m]=α k h k [m](2) where α k represents large-scale fading, h k [m] represents small-scale fading, which follows an exponential distribution with a mean value of 1; The interference model is used to describe the impact of co-frequency interference on the communication link. The V2V link interference is the interference caused by other V2V links reusing the same spectrum and the interference caused by the V2I link. k [m] is calculated as follows: in, is the transmit power of the mth V2I link, is the interference channel gain of the mth V2I link to the kth V2V link, ρ k′ [m] is an indicator variable indicating whether the k′th V2V link uses the mth spectrum resource block. is the transmission power of the kth V2V link, g k′,k [m] is the interference channel gain of the k′th V2V link to the kth V2V link; The received signal-to-interference-to-noise ratio (SINR) of the V2I link is: in, is the SINR of the mth V2I link, and are the transmission powers of the mth V2I link and the kth V2V link on the mth subchannel, σ 2 is the noise power, g k,B [m] is the interference from the kth V2V link to the base station on the mth subchannel, is the interference from the mth V2I link to the base station on the mth subchannel; ρ k [m] is used to describe whether the kth V2V link uses the mth spectrum resource block; Among them, the V2I link rate It is expressed as: Where W is the bandwidth of the spectrum subband; The V2V link rate R k It is expressed as: in, is the SINR of the kth V2V link.
3. The channel resource allocation method based on hierarchical reinforcement learning according to claim 2, characterized in that: In step 1, the construction process of the fairness model of the cellular vehicle network is as follows: To ensure the fairness of V2V link resource allocation, Jain's fairness index is used as a fairness indicator and is defined as follows: Among them, R k is the rate of the kth V2V link, K is the total number of V2V links; Jain's fairness index ranges from 0 to 1, and the closer the value is to 1, the fairer the resource allocation is; assuming represents the set of all V2V links. The goal of the fairness model is to maximize Jain's fairness index f(R k ), described as follows: At the same time, when designing the fairness model, that is, formula (8), the following assumptions and constraints are made: Formula (9) indicates that each V2V link can only use one spectrum subband; Formula (10) indicates that the transmission power of each V2V link is limited by the maximum power P max .
4. The channel resource allocation method based on hierarchical reinforcement learning according to claim 2, characterized in that: In step 1, the process of constructing the reliability model of the cellular vehicle network is as follows: Reliability is measured by the transmission success rate of the V2V link, and the interruption probability will be used to describe reliability; the interruption probability is defined as the probability of failing to successfully transmit data within a specified time, and reliability is defined as the complement of the interruption probability; First, assume that the SINR threshold of the kth V2V link on the mth subband is γ th , then the interruption probability P out Defined as: in, represents the interruption probability of the kth V2V link, Pr is a probability symbol; represents the SINR of the k-th V2V link on the m-th subband and is defined as: in, is the transmit power of the kth V2V link on the mth subband, g k [m] is the channel gain, σ 2 is the noise power, I k [m] is the interference power; Substituting into formula (11), we have: Formula (13) is further transformed into: Assuming the channel gain g k [m] follows an exponential distribution with unit mean, and its probability density function for: where λ is the distribution parameter; for an exponential distribution with unit mean, λ = 1; therefore the cumulative distribution function CDF is: Among them, F gk (x) represents the cumulative distribution function; hence, the outage probability is: Assumptions represents the set of all V2V links. The goal of the reliability model is described as follows: By minimizing the sum of outage probabilities, the reliability of the entire cellular vehicle network can be improved.
5. The channel resource allocation method based on hierarchical reinforcement learning according to claim 2, characterized in that: In step 1, the construction process of the delay model of the cellular vehicle network is as follows: Based on the information age AOI model, a delay model for V2V links is designed. AOI measures the time from information generation to reception. AOI not only considers the transmission time, but also the frequency of information update. It is defined as follows: A k (t)=t-U k (t) (19) Among them, A k (t) is the information age of the kth V2V link at time t, U k (t) is the time when the latest information is updated; The update rule of information age is defined as follows: Among them, A k (t) is the information age of the kth V2V link at time t, Δt is the time interval, R k is the transmission rate of the kth V2V link, R min is the minimum transmission rate requirement; At the same time, in order to minimize the average information age of the V2V link, the following objective function is defined: Where K is the number of V2V links, T is the time period, and A k (t) is the information age of the kth V2V link at time t.
6. The channel resource allocation method based on hierarchical reinforcement learning according to claim 1, characterized in that: In step 2, the design process of the upper controller is as follows: In the upper controller, the action space is defined as possible clustering schemes; each action corresponds to a scheme for allocating V2V links to different clusters and selecting an independent subchannel for each cluster; set up is the action space, then the action represents a clustering scheme; Specifically, the goal of the clustering scheme is to assign V2V links with less interference between each other to the same cluster in order to better allocate channel resources; Assume there are K V2V links in the network, and the action space of the upper controller is represented as a finite set; Each action represents a specific clustering scheme: Where L is the number of clusters, C i represents the set of V2V links contained in the i-th cluster; Since the clustering process is equivalent to the process of selecting sub-channels, the M disjoint sub-channels of the cellular vehicle network wireless spectrum resources are the action space of the upper-layer controller; each sub-channel is occupied by a V2I link; The state space includes the state information of all V2V links in the current network; set up is the state space, then the state Represents the current state of the network. The design of the state space is as follows: Where K is the number of V2V links, M is the number of spectrum subbands, and g k [m] represents the channel gain of the kth V2V link on the mth subband, I k [m] represents the interference suffered by the kth V2V link on the mth subband, C k Indicates the current cluster allocation status of the kth V2V link; At each coherent time step t, the states of all V2V links are again observed Z(t), defined by the observation function O: Z(t)=O(s t ) (24) Among them, S t represents the state observed at time t; The reward function of the upper-layer controller will be used to evaluate the pros and cons of each clustering scheme. To ensure the effectiveness of clustering and the convergence of training, the reward function should take into account the interference between links and the utilization of resources. Let r be the reward function, as shown below: in, represents the current set of all clusters, C represents one of the clusters, (i, j) represents two V2V links belonging to the same cluster C, and g i,j represents the channel gain between V2V links i and j.
7. The channel resource allocation method based on hierarchical reinforcement learning according to claim 6, characterized in that: In step 2, the design process of the lower-layer controller is as follows: In the channel resource optimization decision in the lower-layer controller, the action space is defined as the combination of the transmit power of each V2V link; let is the action space of the kth agent, then the action represents the power allocation scheme of the kth V2V link; Therefore, the action space is designed as follows: Define P i represents the optional transmission power level, i∈[1,N], each action a k It consists of only one transmission power p; The state space includes the channel state information and environment information of each V2V link; is the state space of the kth agent, then the state Represents the current state of the th V2V link, and the state space is designed as follows: Among them, g k [m] represents the gain of the current channel, I k [m] represents the interference from other links, R k Indicates the current transmission rate, A k represents the current information age, P out represents the interruption probability; At each coherent time step t, the state of each V2V agent k is observed Z k (t), is defined by the observation function O as: Z k (t)=O(s k (t),k) (28) Among them, s k (t) represents the state observed by the kth agent at time t; Assume r k is the reward function of the kth agent, then the reward function r k (s k ,a k )as follows: r k (s k ,a k )=αR k -βI k -γA k -δP out +ηf(R k ) (29) Among them, R k Indicates the current transmission rate, I k Indicates the interference from other links, A k represents the current information age, P out represents the interruption probability, f(R k ) is the fairness index, and α, β, γ, δ and η are weight parameters.
8. The channel resource allocation method based on hierarchical reinforcement learning according to claim 7, characterized in that: For the upper-level controller, its Q-value function is expressed as: Among them, Q m (s m ,a m ) represents the Q value function of the upper controller, s m is the state of the upper controller, a m is the action to be performed by the upper controller, r m is the return of the upper controller, γ is the discount factor; A m Represents a collection of actions; (s m )′ means in state s m Next, perform action a m The state reached later, a′ represents the state in which m )′ the action to be performed; For the lower-layer controller, its Q-value function is expressed as: Among them, Q c (s c ,a c ) represents the Q value function of the lower controller, s c is the state of the lower controller, a c is the action to be performed by the lower-level controller, r c is the return of the lower controller, γ is the discount factor; A c Represents a collection of actions; (s c )′ means in state s c Next, perform action a c The state reached later, a′ represents the state in which c )′ the action to be performed; When training the Q network, the loss function is used to measure the gap between the current Q value and the target Q value, which is defined as follows: Among them, L(θ) represents the loss function, θ is the parameter of the current Q network, and θ - are the parameters of the target Q network.
9. The channel resource allocation method based on hierarchical reinforcement learning according to claim 8, characterized in that: In step 2, the process of cellular vehicle network channel resource allocation decision based on hierarchical reinforcement learning is as follows: First, randomly initialize the Q networks of all agents; Each V2V link acts as an intelligent agent and independently maintains its Q network for selecting actions and evaluating rewards; In the upper controller, each V2V link is initialized as a single cluster C k = {k}, and construct an interference graph G(V,E), where the nodes in the interference graph G(V,E) represent V2V links and the edge weights are the channel gains g between the links. i,j ; Define the interference graph G(V,E) where V represents the set of nodes and E represents the set of edges; In each iteration, the vehicle position and large-scale fading parameter α are updated to simulate vehicle movement and environmental changes. For each time slice t, all V2V agents observe the state Z(t), select actions A(t) according to the ∈-greedy strategy, and execute the actions. Through continuous iteration and updating, the agents gradually learn the optimal clustering strategy. After all agents take action, link pairs with interference less than the threshold θ are merged and the cluster allocation scheme is updated. An independent resource block, i.e., a subchannel, is allocated to each cluster. After the clustering scheme is executed, the reward r(t) of each agent is calculated. The reward function takes into account the interference within the cluster and evaluates the pros and cons of the clustering scheme by minimizing the interference within the cluster. After the update channel fades in a small range, all agents observe the new state Z(t+1) and store the state transition (Z(t), A(t), r(t), Z(t+1)) in the replay memory; In the lower controller, for each time slice t, all V2V agents independently observe the state Z k (t), select action a according to the ∈-greedy strategy k (t), and perform actions, namely, power allocation; Each V2V agent selects action a k (t) Power distribution operation is performed; action a k (t) represents the power allocation decision of agent k in the current state; after executing the power allocation plan, the local reward r of each agent is calculated k (t); adjust the impact of each factor on the return through weight parameters α, β, γ, δ and η; each agent observes the new state Z k (t+1), and transfer the state to (Z k (t),a k (t),r k (t),Z k (t+1)) is stored in the playback memory; Small batches of data are uniformly sampled from the replay memory, and stochastic gradient descent is used to optimize the error between the Q network and the learning target; by continuously updating the Q network, the agent gradually learns the optimal power allocation strategy.
10. A channel resource allocation system based on hierarchical reinforcement learning, characterized in that: Includes the following modules: The cellular vehicle network model building module is used to build the environmental model of the cellular vehicle network, including vehicle mobility mode, communication link, channel and interference model, and design the fairness model, reliability model and delay model of the cellular vehicle network; As well as the cellular vehicle network channel resource allocation module, based on the established cellular vehicle network environment model, fairness model, reliability model and delay model, a cellular vehicle network channel resource allocation model based on hierarchical reinforcement learning is constructed; Through hierarchical design, the model decomposes the cellular vehicle network channel resource allocation problem into the upper-layer clustering problem and the lower-layer intra-cluster channel resource optimization problem, and designs the upper-layer controller and the lower-layer controller accordingly. The upper-layer controller performs clustering based on the interference graph and allocates V2V links with less interference to the same cluster to reduce intra-cluster interference and provide a stable environment for the lower layer, which facilitates the lower layer to further optimize channel resource allocation. The upper-layer controller uses the DQN method to make clustering decisions and gradually optimizes the clustering scheme through continuous learning and adjustment. The lower-level controller optimizes the allocation of channel resources through multi-agent deep reinforcement learning (MARL) to allocate appropriate transmission power to each V2V link to maximize link transmission performance while reducing interference. The lower-level controller adopts the DQN method to continuously optimize the transmission rate and interference management by learning the best resource allocation strategy in a dynamic environment.
Citation Information
Patent Citations
Spectrum resource allocation method and device for vehicle network, equipment and storage medium
CN115297550A
Multi-agent air-ground network resource allocation method based on federated learning
CN116546462A
Cognitive Internet of Things resource allocation method based on multi-agent reinforcement learning algorithm
CN117354833A
Multi-agent system networking and resource optimization method
CN118921712A
Resource allocation management for co-channel co-existence in intelligent transport systems
US20220287083A1