Adaptive resource allocation method for ultra-dense Internet of Vehicles based on deep reinforcement learning

By establishing a weighted interference model and Markov decision process problem in ultra-dense vehicle networks, and using the FedAvg-AC algorithm to optimize resource allocation, the problem of differential interference between vehicles is solved, efficient channel and power allocation is achieved, and network throughput and resource reuse rate are improved.

CN119865914BActive Publication Date: 2025-09-30CHONGQING HUAQING AUTOMOBILE PARTS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510010094.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-09-30
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

In ultra-dense Internet of Vehicles, the overlapping communication ranges between vehicles lead to differential interference, resulting in data packet loss, reduced transmission rate and increased transmission delay. The existing resource allocation scheme has high computational overhead, low efficiency and slow response, and cannot meet real-time dynamic resource requirements.

Method used

A weighted interference model is established, and the resource allocation model is constructed as a Markov decision process problem. The FedAvg-AC algorithm is used to solve it. The degree of interference is evaluated through the interference weight graph and correlation matrix. A resource allocation method based on deep reinforcement learning is designed to optimize channel and power allocation to maximize network throughput and resource reuse.

Benefits of technology

It effectively reduces differential interference, improves network throughput and resource reuse, ensures reliable V2V communication, and improves the efficiency and response speed of resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119865914B_ABST
    Figure CN119865914B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for adaptive resource allocation in an ultra-dense vehicle network based on deep reinforcement learning. The method comprises: establishing a communication model for the ultra-dense vehicle network, wherein the communication model includes a base station, a transmitting vehicle, and a receiving vehicle. The base station exchanges data with the vehicle via a vehicle-to-infrastructure (V2I) link, and the vehicles communicate with each other via a vehicle-to-vehicle (V2V) link; determining a weighted interference model centered on the receiving vehicle based on the interference relationship between the vehicles in the communication model; establishing a resource allocation model for the communication model based on the communication model and the weighted interference model, wherein the resource allocation model is an optimization problem with the goal of maximizing resource reuse and ensuring that the allocated resources do not conflict under certain constraints; converting the optimization problem into a Markov decision process problem; and solving it using the FedAvg‑AC algorithm to obtain a resource allocation strategy for the ultra-dense vehicle network. The present invention has the beneficial effect of improving the efficiency of resource allocation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of vehicle networks, and specifically relates to an adaptive resource allocation method for ultra-dense vehicle networks based on deep reinforcement learning. Background Art

[0002] With the rapid development of intelligent transportation, Internet of Vehicles (IoV) technology is becoming an increasingly important component of modern urban traffic management. As more and more vehicles are equipped with advanced communication and sensing capabilities, these intelligent vehicles can exchange and collaborate in real time through vehicle-to-vehicle (V2V) and vehicle-to-infrastructure (V2I) communications, thereby forming an ultra-dense Internet of Vehicles (IoV). In this ultra-dense IoV network, overlapping communication ranges inevitably lead to interference between vehicles. Furthermore, the varying interference capabilities between different vehicles create a phenomenon known as differential interference (or differential interference) in dense IoV networks. This differential interference leads to data packet loss, reduced transmission rates, and increased transmission delays, significantly complicating channel allocation and power distribution in IoV networks. Therefore, identifying an effective solution to mitigate the differential interference caused by overlapping vehicles and ensure reliable V2V communication in ultra-dense IoV deployments is a pressing and critical challenge.

[0003] In recent years, the conflict between communication demand and communication resources in the Internet of Vehicles (IoV) has intensified. Numerous researchers have proposed various resource allocation optimization schemes to address the network resource allocation problem in IoVs. A game-theory-based spectrum access scheme has been proposed for opportunistic spectrum access in cognitive radio vehicular ad hoc networks, enabling vehicles to access licensed frequency bands in a distributed manner. To address the challenges of large-scale data transmission in current network architectures, a long short-term memory (LSTM) network has been employed to predict dynamic throughput, and a genetic algorithm-based resource allocation algorithm has been developed, significantly improving resource utilization efficiency. For high-speed convoys, existing techniques have proposed a joint optimization problem that aims to simultaneously optimize task offloading decisions, communication, and computational resource allocation for both the convoy and base stations to maximize utility and network throughput. A multi-agent neighboring strategy optimization algorithm, Ly-MAPPO, has been proposed for efficiently managing cooperation overhead. This algorithm relies on local observations and utilizes Lyapunov optimization techniques for IoV resource management. Furthermore, other communication infrastructure can also assist in resource management in IoV networks. For example, an optimization algorithm based on a dynamic digital twin model employs a two-stage incentive mechanism derived from the Stackelberg game and the Alternating Direction Method of Multipliers (ADMM) distributed algorithm. This algorithm aims to optimize resource allocation in drone-assisted connected vehicles (IoVs) to improve vehicle user satisfaction and overall energy efficiency. However, these solutions suffer from high computational overhead, low efficiency, and slow response times, making them unable to meet the real-time dynamic resource requirements of IoVs. Therefore, a new robust computational approach is urgently needed to address resource allocation in IoVs.

[0004] With the advancement of machine learning, reinforcement learning (RL) technology has become an effective means of managing and controlling vehicular network resources by making intelligent decisions. Existing technologies use deep reinforcement learning for hybrid task offloading, focusing on vehicle-to-edge (V2E) and vehicle-to-vehicle (V2V) offloading to optimize resource utilization and management under strict latency constraints and dynamic task characteristics. Existing technologies integrate a Double Deep Q-Network (DDQN) and deep deterministic policy gradient to recommend actions with mixed discrete-continuous variables, aiming to reduce computational costs. To address the challenges posed by dynamic environments, technologies applying reinforcement learning to handle high-dimensional and continuous state-action spaces have emerged. However, existing reinforcement learning (RL) models in these studies are centralized and fail to account for the unreliable communication connections typical of ultra-dense and dynamic vehicular networks. A distributed RL framework based on MultiIntelligent Deep Deterministic Policy Gradient (MADDPG) is being developed to manage wireless resources in connected vehicular networks. To improve RL performance in the Internet of Things, a semi-synchronous federated learning protocol has emerged. Later, an algorithm based on federated learning emerged, which balances resource consumption by considering node characteristics and resource load. For resource management of air-ground integrated networks in 6G vehicular networks, a resource allocation method using federated learning and dueling dual-depth Q networks (FedD3QN) has emerged in the prior art, which can reduce the link delay from vehicle to infrastructure. However, these studies did not consider channel and power allocation under differential interference in ultra-dense deployments. In ultra-dense vehicular networks, the overlap of vehicle communication ranges leads to different interference capabilities of the same vehicle on different vehicles. This differential interference leads to a decrease in data transmission rate, packet loss rate and network throughput. These challenges seriously affect the reliability of vehicle communications and complicate resource allocation in ultra-dense vehicular networks. Summary of the Invention

[0005] In view of this, the object of the present invention is to provide an adaptive resource allocation method for ultra-dense Internet of Vehicles based on deep reinforcement learning.

[0006] The purpose of the present invention is achieved through the following technical solutions:

[0007] A method for adaptive resource allocation in ultra-dense Internet of Vehicles (IoV) based on deep reinforcement learning, comprising:

[0008] Establish a communication model for an ultra-dense Internet of Vehicles (IoV). The communication model includes a base station and K vehicles, where the K vehicles include F sending vehicles and R receiving vehicles. The base station exchanges data with the vehicles via a V2I link, and the vehicles communicate with each other via a V2V link, where the sending vehicle sends signals to the receiving vehicle. The number of V2V links is N.

[0009] According to the interference relationship between vehicles in the communication model, a weighted interference model centered on the receiving vehicle is determined;

[0010] Establishing a resource allocation model for the communication model based on the communication model and the weighted interference model, wherein the resource allocation model is an optimization problem with the goal of maximizing resource reuse and ensuring that the allocated resources do not conflict under certain constraints;

[0011] Converting the optimization problem into a Markov decision process problem;

[0012] The FedAvg-AC algorithm is used to solve the Markov decision process problem and obtain the resource allocation strategy for ultra-dense Internet of Vehicles.

[0013] Furthermore, a weighted interference model centered on the receiving vehicle is determined, specifically including:

[0014] For each receiving vehicle, determining the interference weight of each sending vehicle on the receiving vehicle and generating a weight matrix;

[0015] According to the plurality of sending vehicles, the plurality of receiving vehicles and the weight matrix, an interference weight graph G={F, R, W} is established, wherein F is a vertex set representing the set of sending vehicles; R is an edge set representing the set of receiving vehicles; and W represents the weight matrix;

[0016] The interference weight map is converted into the correlation matrix H g Represents, where the incidence matrix H g A row represents the interference weight of a sending vehicle to each receiving vehicle, and the correlation matrix H g A column represents the interference weight of a receiving vehicle from each sending vehicle;

[0017] According to the correlation matrix and resource allocation scheme, the interference level v of each receiving vehicle is determined. r , where r = 1, 2, ..., R, and R represents the total number of received vehicles;

[0018] Determine the overall interference level I of ultra-dense vehicular networks.

[0019] Furthermore, the interference level of each receiving vehicle is expressed as:

[0020]

[0021] Among them, ν r represents the interference level of the rth receiving vehicle, c represents the resource allocation of the sending vehicle; h r Represents the incidence matrix H g The combination of all elements in the rth column of r Indicates c and h r The result of multiplying and summing the elements at the corresponding positions in c; c[i] represents the resources allocated to the i-th node in c, that is, the resources allocated to the i-th sending vehicle, ν r = 0 means there is no interference in the rth receiving vehicle, otherwise, ν r ≠0.

[0022] Furthermore, the optimization problem of the resource allocation model is expressed as:

[0023]

[0024] C3:l n,m ∈(0, 1)

[0025]

[0026] C5: I = 0,

[0027] Where λ1,λ2∈(0,1), M represents the set of channels, M={1,2,…,M}, M represents the total number of channels, η r represents the resource reuse rate, P represents the power set, P={p1,p2,…,p F}, p f represents the signal transmission power selected by the f-th transmitting vehicle when transmitting the signal; N = {1, 2, ..., N} represents the set of V2V links; η0 represents the resource reuse rate of the entire network;

[0028] represents the data transmission rate of the nth V2V link at time t, R min represents the minimum transmission rate of the channel, t∈{1,2,...,T}, T represents the total time;

[0029] represents the signal-to-noise ratio of the nth V2V link at time t; γ0 represents the signal-to-noise ratio threshold;

[0030] l n,m Indicates the selection of the mth channel by the nth V2V link. When the nth V2V link selects the mth channel, l n,m =1; otherwise, when the nth V2V link does not select the mth channel, l n,m =0.

[0031] Furthermore, converting the optimization problem into a Markov decision process problem includes:

[0032] Treat each V2V link as an intelligent agent;

[0033] According to the vehicle network information at time t, the state space at time t is represented as S = {M F ,γ,M,P},M F represents the channel selection of F sending vehicles in the communication model, γ represents the set of signal-to-noise ratios of all channels, M represents the total number of channels, and P represents the set of transmission powers selected by all sending vehicles;

[0034] The set of action spaces is represented as: A={a1,a2,…,a T}, a t ={l n,m ,p f,m} represents the action at time t, which includes channel selection l n,m and power selection p f,m , p f,m represents the signal transmission power when the f-th transmitting vehicle selects the m-th channel, f∈{1,2,...,F}, n∈{1,2,...,N}, m∈{1,2,...,M};

[0035] The reward value obtained by each agent based on the action at time t is:

[0036] The action-value function is: Among them, E represents the expectation, the network state space S0 represents the initialization value of the state space; π represents the strategy; ρ represents the attenuation factor; Q π represents the cumulative reward when starting the network state space S0 under strategy π.

[0037] Furthermore, the FedAvg-AC algorithm specifically includes:

[0038] Construct a global model and a local AC model. Each agent corresponds to a local AC model. Each agent uses the global model to construct a local AC model. During the global model training process, each agent replays the local experience from the corresponding local experience pool D. n Randomly extract a preset number of data sets B n To update the local AC model corresponding to the agent; after a round of learning, the parameters of the global model are obtained by weighted averaging the parameters of each local AC model.

[0039] Furthermore, the method specifically includes:

[0040] The parameter w of the global model at time t-1 ist-1 As the parameter of each local AC model at time t-1

[0041] n∈{1,2,...,N};

[0042] According to the parameters of each local AC model at time t-1 Determine the loss function of each local AC model at time t-1

[0043] Calculating gradients Get the parameters of the local AC model at time t Update local AC model;

[0044] Parameters of each local AC model Perform weighted averaging to obtain the global model parameter w at time t t ;

[0045] The global model parameter w at time t t Sent to each local AC model for the next round of training.

[0046] Furthermore, the parameters of each local AC model are Perform weighted averaging to obtain the global model parameter w at time t t , which can be expressed as: Among them, |B n | represents dataset B n The amount of data in .

[0047] Furthermore, the local AC model includes an actor network π(s,a;θ) and a critic network q(s,a;ψ), where the actor network is used to identify the action policy in the vehicle network and its parameter is θ, and the critic network is used to evaluate the quality of the policy and its parameter is ψ.

[0048] The parameters of the AC model are updated by updating θ and ψ through the gradient descent method. Specifically,

[0049]

[0050] Among them, δ t =q t -(r t +ρ·q t+1 ), α and β represent the learning rate, q t =q(s t ,a t ,ψ t ), r t represents the reward value at time t, q t represents the critic network at time t.

[0051] The beneficial effects of the present invention are:

[0052] This paper first evaluates the communication interference between vehicles in V2V scenarios and constructs a weighted interference model centered on the receiver (i.e., the interference-receiving vehicle) to visualize the interference relationship between various transmitters in a dense vehicle network environment.

[0053] The present invention formulates a resource allocation problem in the presence of differential interference in dense IoVs and defines a combinatorial optimization problem with network throughput and resource reuse rate as objective functions, with the goal of maximizing network throughput and resource reuse rate. Furthermore, to quantify the degree of conflict in ultra-dense IoV networks, a conflict degree is defined. Finally, conflict-free resource allocation is achieved by establishing constraints that the conflict degree must be zero, the signal-to-noise ratio must be greater than a threshold, and the transmission rate must be greater than a threshold.

[0054] The present invention transforms the combinatorial optimization problem of conflict-free resource allocation in dense vehicular networks into a Markov Decision Process (MDP) model, where the reward function is complexly designed to comply with constraints such as maximizing network throughput and optimizing resource reuse. Within this MDP framework, the present invention uses the Actor-Critic (AC) algorithm to solve the combinatorial optimization problem of resource allocation in dense vehicular networks.

[0055] In order to accelerate the convergence speed of the reinforcement learning algorithm and improve the learning efficiency, the present invention proposes a reinforcement learning algorithm based on the FedAvg-AC (Federated Actor Critic) framework.

[0056] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be further described in detail below with reference to the accompanying drawings, in which:

[0058] Figure 1 is the vehicle network scene graph;

[0059] Figure 2 is a schematic diagram of the interference model;

[0060] Figure 3 It is a resource management framework based on FedAvg-AC;

[0061] Figure 4 It is a graph of the relationship between convergence and different learning rates;

[0062] Figure 5 This is a SINR comparison chart;

[0063] Figure 6 This is a comparison chart of network throughput;

[0064] Figure 7 This is a resource reuse rate comparison chart. DETAILED DESCRIPTION

[0065] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the preferred embodiments are only for illustrating the present invention, and are not intended to limit the scope of protection of the present invention.

[0066] The present invention proposes a method for adaptive resource allocation in ultra-dense vehicle networks based on deep reinforcement learning, which includes:

[0067] Step 1: Establish a communication model for ultra-dense vehicle networks;

[0068] Step 2: Based on the interference relationship between vehicles in the communication model, determine the weighted interference model centered on the receiving vehicle;

[0069] Step 3: Based on the communication model and the weighted interference model, a resource allocation model for the communication model is established. The resource allocation model is an optimization problem with the goal of maximizing resource reuse and ensuring that the allocated resources do not conflict under certain constraints.

[0070] Step 4: Convert the optimization problem into a Markov decision process problem;

[0071] Step 5: Use the FedAvg-AC algorithm to solve the Markov decision process problem and obtain the resource allocation strategy for ultra-dense Internet of Vehicles.

[0072] Figure 1 is the scene graph of the vehicle network. Figure 1As shown, the communication model includes a base station and K vehicles, and the K vehicles include F sending vehicles and R receiving vehicles. The sending vehicle (also called the sending end, sender, sending end or sending vehicle, etc.) sends a signal to the receiving vehicle (also called the receiving end, receiver, receiving end or receiving vehicle). In the present invention, only the interference of the sender to the receiver is considered. The base station exchanges data with the vehicle through the V2I link, and the vehicles communicate with each other through the V2V link, where the number of V2V links is N; N = {1,2,...,N} represents the set of all V2V links; it should be noted that the present invention considers the time slot communication system, and each V2V link can only select one channel to transmit the signal in one time slot, and for the power of the interfering sender (i.e., the sending vehicle), each interfering sender can only select one transmission power. Define M = {1,2,…,M} to represent the set of channels, and M represents the total number of channels. P = {p1,p2,…,p F} represents the power set, p f represents the signal transmission power selected by the f-th sending vehicle when sending a signal, f∈{1,2,...,F}.

[0073]

[0074] Among them, l n,m Indicates the selection of the mth channel by the nth V2V link. When the nth V2V link selects the mth channel, l n,m =1; otherwise, when the nth V2V link does not select the mth channel, l n,m =0.

[0075] The V2V wireless channel follows the free space loss mathematical model, so the channel gain of the nth V2V link in the mth channel in time slot t (also called “time t”) is It can be expressed as:

[0076]

[0077] in, represents the small-scale fading power component of the nth V2V link in the mth channel at time t, and assumes that it obeys an exponential unit mean distribution, It represents the large-scale fading power component of the n-th V2V link in the m-th channel at time t that is not affected by frequency.

[0078] According to the Rayleigh fading channel model, the signal-to-interference-plus-noise ratio (SINR) of the mth channel at time t is Can be defined as:

[0079]

[0080] in, represents the signal transmission power of the nth V2V link in the mth channel at time t, σ 2 represents the power of additive Gaussian white noise in the channel; and They represent the transmit power and channel gain of the nth V2V link at time t respectively.

[0081] SINR is the ratio of the signal power to the combined interference and noise power in the system, as shown in Equation (3). A higher SINR indicates that the receiver can more clearly distinguish the desired signal amidst reduced interference and noise levels.

[0082] The data transmission rate of the nth V2V link is:

[0083]

[0084] In vehicular networks, SINR is a key quality of service (QoS) parameter to evaluate the reliability of data transmission. When the instantaneous SINR of the V2V link drops to the threshold γ o When the SINR exceeds the specified threshold, the receiver has difficulty in accurately decoding the transmitted information. Consequently, this leads to a breakdown in communication between vehicles. Therefore, in order to ensure reliable V2V transmission, it is crucial that the SINR exceeds a specified threshold, so the following equation needs to be satisfied:

[0085]

[0086] Next, the interference types in dense IOV networks are analyzed. This paper establishes an interference model centered on interference receivers to address the resource allocation problem in ultra-dense IOV networks.

[0087] Figure 2 is a schematic diagram of the interference model. Figure 2 As shown, the interference model includes 5 receivers (i.e., receiving vehicles, such as Figure 2 R1, R2, ..., R5) and 8 transmitters (i.e., sending vehicles, such as Figure 2 S1, S2, ..., S8) in IoV, the interference weight w is defined at the receiver side:

[0088]

[0089] Among them, P S Represents the transmitter's transmission power, and multiple w can form a weight matrix (i.e. Figure 6 interference matrix in ).

[0090] According to the interference weight of the receiver, interference can be defined as two types, strong interference and weak interference.

[0091] Strong interference: When the interference weight of a single transmitter on a receiver is greater than or equal to the threshold w0, that is, the interference of a single transmitter on a receiver will affect the normal communication of the receiver, this interference phenomenon is called strong interference, w≥w0=1. For example, Figure 2 The interference of S1 to R1 is strong interference.

[0092] Weak interference: When the interference weight of a single transmitter on the receiver is less than the threshold, it is called weak interference, w<w0=1, for example Figure 2 The interference from S8 to R5 is weak interference.

[0093] For the receiver, the weak interference caused by a single transmitter often does not affect the normal communication between vehicles. However, when the receiver is interfered with by multiple receivers, it will cause the normal communication between vehicles to be impossible. This interference phenomenon is called cumulative interference. For example, Figure 2 S5 and S6 interfere with R2. Whether it is cumulative interference or strong interference, it will cause the receiving end to be unable to receive data normally, reduce the data transmission rate, and seriously affect the normal communication between vehicles.

[0094] In order to quantify the interference degree in the interference model, an interference matrix is ​​established and the interference degree is set to measure the interference degree in the in-vehicle Internet.

[0095] Determine a weighted interference model centered on the interference receiving vehicle, including:

[0096] For each interfering receiving vehicle, determining the interference weight of each sending vehicle on the receiving vehicle and generating a weight matrix;

[0097] According to the plurality of sending vehicles, the plurality of receiving vehicles and the weight matrix, an interference weight graph G={F, R, W} is established, wherein F is a vertex set representing the set of sending vehicles; R is an edge set representing the set of receiving vehicles; and W represents the weight matrix;

[0098] Next, the interference weight map is transformed into the correlation matrix H g To express it, its matrix element h(a,e) can be expressed as:

[0099]

[0100] Among them, h(a,e) represents H g The element in the ath row and eth column in , h(a,e)=w means that vertex a is on edge e, and w represents the weight of a on edge e, where a represents the element in F, that is, the sending end; e represents the element in R, that is, the receiving end.

[0101] Figure 1The interference weight map shown can be represented by the matrix H (i.e., the correlation matrix H g )express:

[0102]

[0103] The interference weight graph directly displays the interference relationship in the IoV. When allocating resources, transmitters with strong or cumulative interference cannot be allocated the same resource blocks. For example, S1 on R1 in H cannot be divided into the same resource blocks as S2 and S4.

[0104] To measure the interference in the receiver, we have the correlation matrix H g and resource allocation scheme to determine the interference level v of each interfering receiving vehicle r , where r = 1, 2, ..., R, and R represents the total number of interference receiving vehicles;

[0105] The present invention only considers the interference from the signal transmitter to the signal receiver. A receiving end may be interfered with by multiple transmitting vehicles. The interference level of each receiving vehicle is expressed as:

[0106]

[0107] Among them, ν r represents the interference level of the rth receiving vehicle, c represents the resource allocation of the sending vehicle; h r Represents the incidence matrix H g The combination of all elements in the rth column of r Indicates c and h r The result of multiplying and summing the elements at the corresponding positions in c; c[i] represents the resources allocated to the i-th node in c, that is, the resources allocated to the i-th sending vehicle, The coefficient of the same resource can be extracted, ν r = 0 means there is no interference in the rth interference receiving vehicle, otherwise, ν r ≠0.

[0108] For example, if the resources allocated in the rth column (e.g., column 1) are (A, A, B, B, C, C, D, D), that is, resource A is allocated to S1 and S2, resource B is allocated to S3 and S4, resource C is allocated to S5 and S6, and resource D is allocated to S7 and S8, that is, c = [A, A, B, B, C, C, D, D] in formula (9), hr = [1, 0.3, 0, 0.5, 0, 0, 0, 0]T in formula (9), then c·h r =1A+0.3A+0.5B=1.3A+0.5B, c[1]=A, then In this way, the coefficients of the resource A that is also allocated will be extracted.

[0109] Then, the interference level I of the entire ultra-dense vehicle network can be determined:

[0110]

[0111] The resource allocation problem in vehicular networks considering differential interference can be transformed into a combinatorial optimization problem. The main goal of this paper is to optimize resource allocation in ultra-densely deployed vehicular networks, avoid resource conflicts, and improve the network throughput of the entire vehicular network. Therefore, the conflict-free resource allocation problem in densely deployed vehicular networks can be described as follows:

[0112]

[0113] Where λ1,λ2∈(0,1), M represents the set of channels, M={1,2,…,M}, M represents the total number of channels, η r represents the resource reuse rate, P represents the power set, P={p1,p2,…,p F}, p f represents the signal transmission power selected by the f-th transmitting vehicle when transmitting the signal; N = {1, 2, ..., N} represents the set of V2V links; η r represents the resource reuse rate of the rth interference receiving vehicle;

[0114] represents the data transmission rate of the nth V2V link at time t, R min represents the minimum transmission rate of the channel, t∈{1,2,...,T}, T represents the total time;

[0115] represents the signal-to-noise ratio of the nth V2V link at time t; γ0 represents the signal-to-noise ratio threshold;

[0116] l n,m Indicates the selection of the mth channel by the nth V2V link. When the nth V2V link selects the mth channel, l n,m =1; otherwise, when the nth V2V link does not select the mth channel, l n,m =0;

[0117] For the above constraints, constraint C1 ensures that the minimum network throughput is not less than the minimum network throughput in the IoV network. Constraint C2 ensures that V2V links can communicate with each other, that is, the signal-to-noise ratio is greater than the minimum threshold. Constraints C3 and C4 represent the constraints on channel and power selection in IoV resource allocation. Constraint C5 ensures that IoV network resource allocation should satisfy conflict-free resource allocation. The problem represented by Equation (11) is a time-averaged combinatorial optimization problem, which is mathematically challenging and computationally difficult.

[0118] To this end, this paper develops an MDP model to solve the resource management problem, namely the combinatorial optimization problem. It also proposes an action-critic (AC) resource allocation method based on reinforcement learning. To accelerate learning, a federated reinforcement learning method, the FedAvg-AC algorithm, is proposed.

[0119] The above combinatorial optimization problem in the ultra-dense IoV setting can be formulated as an MDP problem, which is Markov and fully observable.

[0120] Converting the optimization problem into a Markov decision process problem involves:

[0121] Treat each V2V link as an intelligent agent;

[0122] According to the vehicle network information at time t, the state space at time t is represented as S = {M F ,γ,M,P},M F represents the channel selection of F sending vehicles in the communication model, γ represents the set of signal-to-noise ratios of all channels, M represents the total number of channels, and P represents the set of transmission powers (or transmit powers) selected by all sending vehicles (or senders);

[0123] The set of action spaces is represented as: A={a1,a2,…,a T}, a t ={l n,m ,p k,m} represents the action at time t, which includes channel selection l n,m and power selection p k,m , p k,m represents the transmission power when the k-th vehicle selects the m-th channel, k∈{1,2,...,K}, n∈{1,2,...,N}, m∈{1,2,...,M};

[0124] The reward value obtained by each agent based on the action at time t is:

[0125]

[0126] That is, the objective function is used as the reward value function;

[0127] The action-value function is:

[0128]

[0129] Where E represents the expectation, the network state space S0 represents the initial state space (i.e., the initialization value of the state space); π represents the strategy; ρ represents the attenuation factor; Q π represents the cumulative reward when starting the network state space S0 under strategy π.

[0130] During training, if the resource management scheme is optimal—that is, if the allocated resources are non-conflicting and resource reuse is maximized—then the reward value is maximized. Conversely, if the allocated resources are conflicting and resource reuse is low, the reward value is minimized and negative. In the reinforcement learning network constructed by the present invention, each agent is constantly searching for a strategy that maximizes the reward value.

[0131] MDP models can be solved using Q-learning, policy gradient methods, and deep Q-learning. However, the slow convergence of the Q function or its approximation in Q-learning often leads to suboptimal performance and hinders the discovery of the optimal policy within a reasonable number of iterations. In addition, Q-learning is inefficient at learning stochastic policies, especially in environments with continuous action spaces, such as those found in IoV networks where channel conditions and transmit power vary.

[0132] In contrast, policy gradient methods manipulate policies directly in policy space and can typically converge to local optimal policies faster than Q-learning. To address the challenge of learning optimal policies for intelligent resource management in continuous state and action spaces, actor-critic (AC) learning algorithms integrate policy and value learning, aiming for efficient convergence.

[0133] In the AC framework, the actor (i.e., actor) and critic (i.e., critic) components are crucial. The actor determines a strategy for selecting actions by observing the network state, while the critic evaluates the strategy using rewards obtained from environmental feedback. In a vehicle network environment where V2V connections serve as intelligent agents, each connection autonomously perceives the current network conditions and takes actions based on its learned strategy, operating in a decentralized manner. After these actions, the IoV environment provides the new state and immediate feedback to the agents, enabling them to update their strategies for subsequent iterations.

[0134] The AC model (i.e., AC framework) consists of an actor network π(s, a; θ) and a critic network q(s, a; ψ). The actor network is used to identify the action strategy in the vehicle network, and its parameter is θ. The critic network is used to evaluate the quality of the strategy, and its parameter is ψ.

[0135] The gradient descent method is used to update θ, which can be expressed as follows:

[0136]

[0137] Temporal difference learning scheme is also used in AC reinforcement learning to calculate the temporal difference (TD) error between the estimated value and the true value, which can be expressed as:

[0138] δ t =q t -(rt +ρ·q t+1 ), (15)

[0139] When using the critic network to estimate the state value function, ψ can be updated by gradient descent, which can be expressed as:

[0140]

[0141] Among them, α and β represent the learning rate, q t =q(s t ,a t ,ψ t ), r t represents the reward value at time t; q t represents the critic network at time t.

[0142] However, the AC algorithm converges slowly. To solve this problem, the present invention utilizes a federated learning method to improve the convergence speed of the AC algorithm. In the considered ultra-dense IoV network, all agents use the global model of the FedAvg-AC server to build a local AC network. During the global model training, each agent randomly samples a small amount of data set B from the local experience replay pool D to update its own local AC model. The local update of the nth RL agent minimizes the target Q-network L(w t Specifically, each agent replays the local experience pool D corresponding to the agent. k Randomly extract a preset number of data sets B k To update the local AC model corresponding to the agent; after a round of learning, the parameters of the global model are obtained by weighted averaging the parameters of each local AC model, thereby deriving the FedAvg-AC global network.

[0143] The loss function L(w t ) can be expressed as follows:

[0144]

[0145] Among them, L(w t ) represents the global model loss function, represents the local AC model loss function, |B n | represents dataset B n The amount of data in .

[0146] Specifically, the FedAvg-AC algorithm includes:

[0147] Build a global model and a local AC model. Each agent corresponds to a local AC model, and each agent uses the global model to build a local AC model.

[0148] During global model training, perform the following operations:

[0149] The parameter w of the global model at time t-1 is t-1 As the parameter of each local AC model at time t-1 n∈{1,2,...,N};

[0150] According to the parameters of each local AC model at time t-1 Determine the loss function of each local AC model at time t-1

[0151] Calculating gradients Get the parameters of the local AC model at time t Update the local AC model, that is:

[0152]

[0153] Next, the parameters of each local AC model are Perform weighted averaging to obtain the global model parameter w at time t t ,Right now:

[0154]

[0155] Then, the server sets the global model parameter w at time t t Sent to each local AC model (e.g., by mass email) for the next round of training.

[0156] Algorithm 1 summarizes the training process. The algorithm flow is as follows:

[0157]

[0158] The effectiveness of the proposed algorithm is verified through simulation experiments using the urban scenario in 3GPP TR 36.885. The FedAvg-AC algorithm is implemented in PyTorch. The parameters of the proposed FedAvg-AC are detailed in Table 1.

[0159] Table 1 Simulation experiment parameters

[0160]

[0161] In order to verify the effectiveness of the algorithm in this paper, a simulation performance comparison was carried out between it and other algorithms such as random network resource allocation (RM), network resource allocation based on greedy algorithm (GA), and network resource allocation based on maximum node degree algorithm (MND).

[0162] Figure 4This figure illustrates the convergence behavior of the FedAvgAC-based network resource management method across various learning rates. Throughout the training phase, a fixed number of 20 V2V links is used, and the reward value remains stable as training progresses. The learning rate plays a crucial role in controlling the learning dynamics of the network model, influencing the speed at which the algorithm converges to the global optimum. A learning rate that is too low can cause the algorithm to become trapped in a local optimum. As shown in the figure, setting the learning rate to 0.0001 shows a significant improvement in convergence performance. Therefore, a learning rate of 0.0001 was selected for subsequent experiments. Figure 4 The effectiveness of the FedAvg-AC network resource management algorithm is emphasized.

[0163] The advantages of the present invention are described below through analysis of three indicators (Signal to Interference and Noise Ratio SINR, network throughput, and resource reuse rate).

[0164] SINR is the ratio of the signal power to the combined interference and noise power in the system, as shown in Equation (3). A higher SINR indicates that the receiver can more clearly distinguish the desired signal amidst reduced interference and noise levels.

[0165] Network throughput: This performance metric evaluates the network throughput of the Internet of Vehicles after the resource allocation algorithm has allocated all communication link resources. It can be expressed as:

[0166]

[0167] in, Represents the average SINR. The higher the network throughput, the higher the maximum rate the network can accept, and the better the network performance.

[0168] Resource reuse rate: This performance metric indicates the extent to which network resources are effectively utilized within a specific time period. A high network resource reuse rate means that network resources are fully utilized, reducing idle waste and improving overall network performance and efficiency. It can be expressed as:

[0169]

[0170] Where M represents the total number of resources (i.e. the total number of channels), M k Indicates the total number of resources utilized.

[0171] Figure 5 The relationship between the number of V2V links and SINR is shown. A higher SINR value means that the network has lower interference. Figure 5The SINR values ​​of this algorithm under different transmission links were compared with those of three other algorithms. Compared with the other three algorithms, this algorithm achieves a higher signal-to-interference-and-noise ratio. This is because it adaptively selects the optimal resource allocation scheme, effectively avoiding interference, improving the SINR, and ensuring transmission rates. Simulation results demonstrate that this algorithm effectively avoids interference and improves the SINR.

[0172] Figure 6 The network throughput achieved using various algorithms in a connected vehicle environment is shown. It is clear that as the number of V2V connections increases, the network throughput also increases. This can be attributed to the enhanced transmission links, which enable a larger amount of data to be transmitted simultaneously, thereby improving the overall network throughput. In addition, the FedAvg-AC algorithm proposed in this invention also significantly outperforms the other algorithms. This superiority can be attributed to the strong adaptability of the algorithm. Therefore, Figure 6 It is verified that the proposed algorithm significantly improves the network throughput of the system.

[0173] Figure 7 The relationship between different numbers of transmitters and resource reuse rate in dense vehicle networking is shown. A higher resource reuse rate indicates that the system effectively utilizes the available communication resources. Figure 7 The resource reuse rate of the proposed algorithm is compared with that of other three comparison algorithms under different numbers of reuse times. The results show that the proposed algorithm outperforms similar algorithms in terms of resource reuse rate. This improvement is due to the ability of the algorithm to adaptively select the best resource allocation scheme based on the current network status. Therefore, Figure 7 It is verified that the proposed algorithm significantly improves the resource reuse rate.

[0174] In summary, the present invention addresses the resource allocation problem in ultra-dense overlapping interference scenarios, establishes a weighted interference model, and proposes a resource allocation algorithm based on FedAvg-AC to allocate channel and power resources. By studying the characteristics of ultra-dense vehicle communication distance overlap, the relationship between ultra-dense vehicle communication distance overlap and interference is analyzed. A weighted interference model is used to evaluate the interference level of the entire vehicle network, and the resource allocation problem is redefined as a combinatorial optimization problem. To address this problem, a method for establishing an MDP model for an ultra-dense vehicle network based on a weighted interference model is explored. Then, an AC-based resource allocation algorithm is designed for the MDP model. In order to accelerate the convergence speed of the AC algorithm and improve the computing power of the network, a FedAvg-AC resource allocation algorithm based on federated learning is proposed. Finally, computer simulation is applied to verify the effectiveness of the algorithm. This study provides valuable insights for resource management in ultra-dense vehicle networks.

[0175] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.

Claims

1. A method for adaptive resource allocation in ultra-dense vehicle networks based on deep reinforcement learning, characterized in that: include: Establish a communication model for an ultra-dense Internet of Vehicles (IoV). The communication model includes a base station and K vehicles, where the K vehicles include F sending vehicles and R receiving vehicles. The base station exchanges data with the vehicles via a V2I link, and the vehicles communicate with each other via a V2V link, where the sending vehicle sends signals to the receiving vehicle. The number of V2V links is N. According to the interference relationship between vehicles in the communication model, a weighted interference model centered on the receiving vehicle is determined; Establishing a resource allocation model for the communication model based on the communication model and the weighted interference model, wherein the resource allocation model is an optimization problem with the goal of maximizing resource reuse and ensuring that the allocated resources do not conflict under certain constraints; Converting the optimization problem into a Markov decision process problem; The FedAvg-AC algorithm is used to solve the Markov decision process problem and obtain the resource allocation strategy for ultra-dense Internet of Vehicles. The FedAvg-AC algorithm specifically includes: Construct a global model and a local AC model, treat each V2V link as an agent, and each agent corresponds to a local AC model. Each agent uses the global model to build a local AC model. During the global model training process, each agent replays the local experience pool D corresponding to the agent. n Randomly extract a preset number of data sets B n To update the local AC model corresponding to the agent; after a round of learning, the parameters of the global model are obtained by weighted averaging the parameters of each local AC model.

2. The method for adaptive resource allocation in ultra-dense vehicle networks based on deep reinforcement learning according to claim 1, characterized in that: Determine the weighted interference model centered on the receiving vehicle, including: For each receiving vehicle, determining the interference weight of each sending vehicle on the receiving vehicle and generating a weight matrix; According to the plurality of sending vehicles, the plurality of receiving vehicles and the weight matrix, an interference weight graph G={F, R, W} is established, wherein F is a vertex set representing the set of sending vehicles; R is an edge set representing the set of receiving vehicles; and W represents the weight matrix; The interference weight map is converted into the correlation matrix H g Represents, where the incidence matrix H g A row represents the interference weight of a sending vehicle to each receiving vehicle, and the correlation matrix H g A column represents the interference weight of a receiving vehicle from each sending vehicle; According to the correlation matrix and resource allocation scheme, the interference level v of each receiving vehicle is determined. r , where r = 1, 2, ..., R, and R represents the total number of received vehicles; Determine the overall interference level I of ultra-dense vehicular networks.

3. The method for adaptive resource allocation in ultra-dense vehicle networks based on deep reinforcement learning according to claim 2, characterized in that: The interference level of each receiving vehicle is expressed as: Among them, ν r represents the interference level of the rth receiving vehicle, c represents the resource allocation of the sending vehicle; h r Represents the incidence matrix H g The combination of all elements in the rth column of r Indicates c and h r The result of multiplying and summing the elements at the corresponding positions in c; c[i] represents the resources allocated to the i-th node in c, that is, the resources allocated to the i-th sending vehicle, ν r = 0 means there is no interference in the rth receiving vehicle, otherwise, ν r ≠0.

4. The method for adaptive resource allocation in ultra-dense vehicle networks based on deep reinforcement learning according to claim 1, characterized in that: The optimization problem of the resource allocation model is expressed as: C5:I=0, Where λ1,λ2∈(0,1), M represents the set of channels, M={1,2,…,M}, M represents the total number of channels, η r represents the resource reuse rate, P represents the power set, P={p1,p2,…,p F }, p f represents the signal transmission power selected by the f-th transmitting vehicle when transmitting the signal; N = {1, 2, ..., N} represents the set of V2V links; η0 represents the resource reuse rate of the entire network; represents the data transmission rate of the nth V2V link at time t, R min represents the minimum transmission rate of the channel, t∈{1,2,...,T}, T represents the total time; represents the signal-to-noise ratio of the nth V2V link at time t; γ0 represents the signal-to-noise ratio threshold; Indicates the selection of the mth channel by the nth V2V link. When the nth V2V link selects the mth channel, Otherwise, when the nth V2V link does not select the mth channel, 5. The method for adaptive resource allocation in ultra-dense vehicle networks based on deep reinforcement learning according to claim 4 is characterized in that: Converting the optimization problem into a Markov decision process problem involves: According to the vehicle network information at time t, the state space at time t is represented as S = {M F ,γ,M,P},M F represents the channel selection of F sending vehicles in the communication model, γ represents the set of signal-to-noise ratios of all channels, M represents the total number of channels, and P represents the set of transmission powers selected by all sending vehicles; The set of action spaces is represented as: A={a1,a2,…,a T }, Indicates the action at time t, which includes channel selection and power selection p f,m , p f,m represents the signal transmission power when the f-th transmitting vehicle selects the m-th channel, f∈{1,2,...,F}, n∈{1,2,...,N}, m∈{1,2,...,M}; The reward value obtained by each agent based on the action at time t is: The action-value function is: Among them, E represents the expectation, the network state space S0 represents the initialization value of the state space; π represents the strategy; ρ represents the attenuation factor; Q π represents the cumulative reward when starting the network state space S0 under strategy π.

6. The method for adaptive resource allocation in ultra-dense vehicle networks based on deep reinforcement learning according to claim 1, characterized in that: The FedAvg-AC algorithm specifically includes: The parameter w of the global model at time t-1 is t-1 As the parameter of each local AC model at time t-1 n∈{1,2,...,N}; According to the parameters of each local AC model at time t-1 Determine the loss function of each local AC model at time t-1 Calculating gradients Get the parameters of the local AC model at time t Update local AC model; Parameters of each local AC model Perform weighted averaging to obtain the global model parameter w at time t t ; The global model parameter w at time t t Sent to each local AC model for the next round of training.

7. The method for adaptive resource allocation in ultra-dense vehicle networks based on deep reinforcement learning according to claim 1, characterized in that: Parameters of each local AC model Perform weighted averaging to obtain the global model parameter w at time t t , which can be expressed as: Among them, |B n | represents dataset B n The amount of data in .

8. The method for adaptive resource allocation in ultra-dense vehicle networks based on deep reinforcement learning according to claim 1, characterized in that: The local AC model consists of an actor network π(s,a;θ) and a critic network q(s,a;ψ). The actor network is used to identify the action strategy in the vehicle network, and its parameter is θ. The critic network is used to evaluate the quality of the strategy, and its parameter is ψ. The parameters of the AC model are updated by updating θ and ψ through the gradient descent method. Specifically, Among them, δ t =q t -(r t +ρ·q t+1 ), α and β represent the learning rate, q t =q(s t ,a t ,ψ t ), r t represents the reward value at time t, q t represents the critic network at time t, and ρ represents the decay factor.