Internet of vehicles communication resource allocation method based on multi-agent deep reinforcement learning
Through the combination of multi-agent deep reinforcement learning and meta-learning framework, the quantitative error and environmental adaptability problems in the allocation of Internet of Vehicles resources are solved, the joint optimization of spectrum and power is achieved, and the resource allocation efficiency and collaboration of Internet of Vehicles communications are improved.
Patent Information
- Application Number
- CN202510481616.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-25
AI Technical Summary
The existing DRL-based Internet of Vehicle resource allocation algorithms have quantization errors in continuous power allocation and are difficult to adapt to the ever-changing environmental challenges in V2X communication, resulting in inefficient resource allocation.
The multi-agent deep reinforcement learning algorithm is used to combine the meta-learning framework, and the discrete spectrum subband selection is processed through the DDQN network and the DDPG network to process continuous power control, realize the joint optimization of spectrum and power, and use the LSTM network to extract timing information, and introduce the meta-learning mechanism to improve the algorithm's adaptability in a dynamic environment.
Effectively handling continuous action space improves the accuracy and adaptability of resource allocation, improves the collaboration and algorithm performance of Internet of Vehicles communication, and meets the needs of low latency and high reliability.
Smart Images

Figure CN120378836A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of wireless communication and vehicle networking, and particularly to a communication resource allocation method for vehicle networking based on multi-agent deep reinforcement learning. Background Art
[0002] In recent years, the automotive industry has shown a development trend of "new four modernizations": electrification, intelligence, networking, and sharing. The intelligent transportation system field has put forward development directions such as digitalization, networking, intelligence, and automation, all of which urgently require vehicle networking to provide basic communication and connection support capabilities. With the accelerating integration of the automotive industry with the information and communication industry and the transportation industry, the intelligent transportation system (ITS) and high-order autonomous driving technologies are experiencing revolutionary breakthroughs. As the core technology to achieve efficient collaboration between vehicle-to-vehicle (V2V), vehicle-to-infrastructure (V2I), vehicle-to-pedestrian (V2P), and vehicle-to-network (V2N), cellular vehicle-to-everything (C-V2X) has become a research hotspot in the global communication and transportation fields and the key to seizing the strategic high ground of intelligent transportation for various countries.
[0003] Currently, the evolution of the international mainstream vehicle networking wireless communication technology mainly divides into two major routes: dedicated short-range communication (DSRC) technology and cellular vehicle-to-everything (C-V2X) technology. As the application of cellular communication technology in vehicle networking, C-V2X avoids duplicate construction by leveraging existing cellular infrastructure, reducing deployment costs. With the support of LTE and 5G technologies, the 3rd Generation Partnership Project (3GPP) divides C-V2X into LTE-V2X based on 4G long-term evolution (LTE) and NR-V2X based on 5G new radio (NR). Relying on existing base stations, C-V2X provides vehicles with greater bandwidth, higher transmission rates, and larger coverage areas, and has higher scalability than DSRC. Moreover, due to the better link budget and controlled quality of service (QoS) of C-V2X communication, it can better meet the requirements of various advanced in-vehicle applications for high reliability, high data rate, and low latency.
[0004] 3GPP standardized LTE-V2X in its Release 1.14, which introduced two modes of C-V2X, Mode 3 and Mode 4. Both modes enable vehicles to perform D2D communication through licensed cellular bands, but they differ in radio resource allocation methods. In Mode 3, the cellular network (usually the BS) manages the radio resources that vehicles use for their direct vehicle communication, which is equivalent to centralized resource allocation at the BS. In Mode 4, vehicles autonomously select radio resources, greatly enhancing vehicle autonomy and enabling vehicles to operate with or without a cellular network.
[0005] There have been many studies on distributed resource allocation methods in the Internet of Vehicles. Due to its characteristics in environmental interaction, the reinforcement learning method has outstanding performance in wireless resource management. In "Deep reinforcement learning based resource allocation for V2v communications" (H. Ye, G. Y. Li and B.-H. F. Juang, IEEE Trans. Veh. Technol., vol. 68, no. 4, pp. 3163-3173, Apr. 2019.), a new resource allocation strategy for the heterogeneous cellular vehicle-to-everything (C-V2X) scenario that meets the actual application is proposed by finding the best subbands and transmission power levels for each V2V user in V2V communications based on the DQN algorithm while meeting the requirement of low latency. "Distributed deep deterministic policy gradient for power allocation control in D2D-based V2V communications" (K. K. Nguyen, T. Q. Duong, N. A. Vien, N.-A. Le-Khac and L. D. Nguyen, IEEE Access, vol. 7, pp. 164533-164543, 2019) proposes two new methods based on the deep deterministic policy gradient algorithm for the power allocation problem in D2D-based V2V communications with the energy efficiency as the optimization goal. "Meta-Reinforcement Learning Based Resource Allocation for Dynamic V2X Communications" (Y. Yuan, G. Zheng, K.-K. Wong and K. B. Letaief, IEEE Transactions on Vehicular Technology, vol. 70, no. 9, pp. 8964-8977, Sept. 2021.) proposes a new DRL algorithm based on meta-learning that can achieve fast adaptation in a new environment with limited experience.
[0006] The existing DRL-based resource allocation algorithms for the Internet of Vehicles still have deficiencies, mainly including the following two points: First, most of the existing work discretizes the continuous power, inevitably resulting in quantization errors, thus reducing the algorithm performance; Second, many works have preset weights in the training, so they solve for a specific pair of weight values and it is difficult to adapt to the challenges brought by the constantly changing environment in V2X communications. Summary of the Invention
[0007] Aiming at the defects existing in the prior art, the present invention provides a communication resource allocation method for vehicle-to-everything (V2X) networks based on multi-agent deep reinforcement learning, which can implement a dynamic spectrum and power joint allocation method, is applicable to the scenario where V2V links reuse V2I orthogonal spectrum resources, and realizes adaptive resource allocation in a high-dynamic environment through a multi-agent deep reinforcement learning algorithm combined with a meta-learning framework.
[0008] The purpose of the present invention is achieved by at least one of the following technical solutions.
[0009] A communication resource allocation method for vehicle-to-everything (V2X) networks based on multi-agent deep reinforcement learning includes the following steps:
[0010] S1: In the cellular vehicle-to-everything (C-V2X) scenario, initialize the number, position, speed and direction of vehicles, the size of the transmission payload and the transmission time limit; construct a wireless channel model according to the 3GPP standard, and calculate the signal-to-interference-plus-noise ratio (SINR) of V2V and V2I links.
[0011] S2: According to the SINR of V2V and V2I links, construct a resource allocation problem model for vehicle-to-everything (V2X) networks.
[0012] S3: Model the resource allocation problem for vehicle-to-everything (V2X) networks as a Markov decision process, define the state space, action space and reward function, and build the basis of intelligent decision-making logic.
[0013] S4: Use the LSTM network to extract the temporal information in the state space, deeply mine the dynamic change law of the communication environment, and convert the temporal information into the dimension suitable for the input of the DDQN network and DDPG network of each agent in the multi-agent network for the Markov decision process.
[0014] S5: Adopt a divide-and-conquer strategy to execute resource allocation. Use the DDQN network to handle the selection of discrete spectrum sub-bands, and utilize its efficient decision-making characteristics for the discrete action space to complete the accurate allocation of spectrum sub-bands. With the help of the DDPG network to handle continuous power control, aiming at the refined adjustment requirements of the continuous action space, realize the optimal regulation of the transmission power, and finally realize the joint optimization of continuous power allocation and discrete spectrum sub-band allocation in the multi-agent network.
[0015] S6: Update the parameters of the DDQN network and DDPG network of each agent based on the meta-learning mechanism. Through multi-task training, make the DDQN network and DDPG network learn the common characteristics and difference laws of different resource allocation tasks, strengthen the adaptability of the algorithm in the dynamic vehicle-to-everything (V2X) network environment, and complete the communication resource allocation for vehicle-to-everything (V2X) networks.
[0016] Furthermore, it is characterized in that in step S1, a wireless channel model is constructed according to the 3GPP standard. The large-scale channel fading is composed of path loss and shadow effect, and is divided into two scenarios: Line of Sight (LOS) and Non-Line of Sight (NLOS). The path loss between V2V links in the LOS case is calculated by the formula:
[0017]
[0018] where f c is the system frequency, h′ BS and h′ MS are the effective antenna heights of the base station and the mobile user respectively; d′ BP represents the break point distance greater than 3 meters, which is jointly determined by the system frequency and the effective antenna heights of the base station and the mobile user. d a and d b represent the horizontal and vertical distances between two vehicles communicating via V2V links respectively, and d represents the Euclidean distance between two vehicles communicating via V2V links;
[0019] The path loss between V2V links in the NLOS case is calculated by the formula:
[0020]
[0021] n j = max(2.8 - 0.0024d b , 1.84);
[0022] where f c is the system frequency; and represent the calculation results of swapping the two dimensions of horizontal distance and vertical distance respectively, and finally take the minimum value among them. n j represents the occlusion correction factor of the jth occluder encountered by the vehicle in the communication environment;
[0023] The path loss of the V2I link is within the LOS range, and its calculation formula is:
[0024]
[0025] where h′ BS and h′ MS are the effective antenna heights of the base station and the mobile user respectively, and d' represents the Euclidean distance between the vehicle communicating via V2I link and the base station;
[0026] Shadow fading can be simplified as:
[0027]
[0028] where S * (n) is an independent and identically distributed random vector in the nth time slot, which is a normal distribution with a mean of 0 and a standard deviation of 0.02; D describes the driving distance of a vehicle in a time slot, and D dEco represents the decorrelation distance, which is different in V2I and V2V links; e is the base of the natural logarithm, S(n) represents the shadow fading component in the nth time slot, and S(n - 1) represents the shadow fading component in the (n - 1)th time slot;
[0029] Assume that the transceivers of all vehicles use one antenna. There are M V2I links and K V2 links in the vehicle network, which are represented by the sets M' = {1, …, m, …, M} and K' = {1, …, k, …, K} respectively;
[0030] Assume that the V2I links have been pre - allocated orthogonal spectrum sub - bands, that is, the mth V2I link occupies the mth spectrum sub - band, and the transmit power of the V2I link is a fixed value of 23 dBm;
[0031] The V2V links achieve spectrum sharing through frequency band selection and power control. ρ k [m] is a boolean spectrum sub - band selection variable, indicating whether the kth V2V link transmits on the mth spectrum sub - band. If the kth V2V link transmits on the mth spectrum sub - band, then ρ k [m] = 1, otherwise ρ k [m] = 0;
[0032] Assume that the channel fading is the same within a spectrum sub - band and independent between different spectrum sub - bands. h k [m] is the power component of the small - scale fading of the kth V2V link on the mth spectrum sub - band, which follows an exponential distribution;
[0033] α k represents the large - scale fading of the kth V2V link that is independent of frequency. Then, within a coherence time, the channel power gain of the kth V2V link on the mth spectrum sub - band is expressed as
[0034] Assume that the interference channel power gain from the transmitter of the k'th V2V link through the mth spectrum sub - band to the receiver of the kth V2V link is expressed as The interference channel power gain from the transmitter of the kth V2V link to the mth V2I link on the mth spectrum sub - band is expressed as The channel gain from the transmitter of the m-th V2I link to the base station in the m-th spectral subband is denoted as The interference channel power gain from the transmitter of the m-th V2I link to the receiver of the k-th V2V link in the m-th spectral subband is denoted as The power of the transmitted signal at the transmitter of the m-th V2I link is The magnitude of the transmission power at the transmitter of the k-th V2V link in the m-th spectral subband is The noise power is σ 2 ;
[0035] The signal-to-interference-plus-noise ratio (SINR) of the m-th V2I link in the m-th spectral subband The calculation expression is:
[0036]
[0037] The signal-to-interference-plus-noise ratio (SINR) of the k-th V2V link in the m-th spectral subband The calculation expression is:
[0038]
[0039] Furthermore, in step S2, a vehicle-to-everything (V2X) resource allocation problem model is constructed as follows:
[0040] Assume that each V2V link can access at most one orthogonal spectral subband, i.e., Σ m ρ k [m] ≤ 1, where W is the bandwidth of each spectral subband. Then, according to the Shannon formula, the channel capacity of the m-th V2I link in the m-th spectral subband is denoted as:
[0041]
[0042] Similarly, the channel capacity of the k-th V2V link in the m-th spectral subband is denoted as:
[0043]
[0044] Since V2I links are usually used to support high-throughput entertainment services, this can be expressed as the maximization of the sum rate of V2I links ; V2V links mainly focus on the reliable transmission of safety-critical messages, aiming to achieve a high successful transmission probability for all V2V users during information transmission; compared with V2I links, the objective function of V2V links is more complex; therefore, the successful transmission of each V2V link is first defined by the following conditions:
[0045]
[0046] Among them, T represents the transmission time limit, and Δ T represents the duration of one transmission time slot, and t k represents the time slot when the k-th V2V link starts to transmit information, and B represents the size of the transmitted information payload;
[0047] Introduce the binary indicator ω k,τ as a quantization index for successful transmission, representing the success status of the k-th V2V link in the τ-th payload transmission:
[0048]
[0049] The success probability η of all V2V links within the limited transmission time period is:
[0050]
[0051] Among them, O k represents the actual number of transmission attempts of this V2V link during the transmission period. The specific rule is: if the cumulative data volume reaches or exceeds B during the τ-th attempt, then O k = τ; if the transmission is still not completed after T / Δ T attempts, then O k = T / Δ T , and at this time ω k,τ = 0.
[0052] Furthermore, in step S2, the definition of successful V2V transmission is measured by the data accumulation within the time limit, emphasizing the co-optimization of low latency and high reliability; this definition directly affects the design of the resource allocation strategy, requiring the algorithm to consider both rate and latency when selecting spectrum sub-bands and allocating power to meet the transmission requirements of safety messages. The vehicle-to-everything (V2X) resource allocation problem model is as follows:
[0053]
[0054] Among them, P max represents the maximum transmission power of the vehicle transmitter for V2V communication, and P I represents the transmission power of the vehicle transmitter for vehicle-to-infrastructure (V2I) communication.
[0055] Furthermore, in step S3, the vehicle-to-everything (V2X) resource allocation problem is modeled as a Markov decision process, specifically as follows:
[0056] Each V2V link acts as an agent, interacting with an unknown environment to obtain experience and then learning the optimal policy from the experience; each agent has its own DDQN and DDPG networks; K V2V links form a multi-agent network to jointly explore an environment and optimize the power control policy according to the changes in the environmental state;
[0057] At time slot t, given the environmental state of the current agent k Then the action is output according to the policy The combined state space of multiple agents is S t , when each agent obtains its respective action, multiple agents form a combined action A t , the agents jointly execute the action and obtain a reward R t , where all agents obtain the same reward in the system to promote their cooperative behavior, and the environmental state where agent k is located With probability Transfer to the environmental state at the next moment For a single agent k, it can only obtain the local environmental state of the environment where it is located And action A t , while the global channel condition and the actions of other agents are unknown;
[0058] All channel information observed by a single agent k on the m-th V2I link at time slot t Includes the channel gain of agent k itself on the m-th V2I link at time slot t And the interference channel power gain from the transmitter of other V2V link k' to agent k itself on the m-th V2I link at time slot t For all m ∈ M, the interference channel power gain from the transmitter of agent k to the base station is And the interference channel from the transmitter of the m-th V2I link to agent k itself on the m-th V2I link at time slot t is The interference of agent k on the m-th V2I link at the previous time slot t - 1 The number of spectrum sub-bands occupied by agent k on the m-th V2I link at the previous time slot t - 1 The remaining payload of agent k at the previous time slot t And the remaining transmission time Therefore, the state of agent k at time slot t Is:
[0059]
[0060] Each agent autonomously selects the transmission frequency band and controls the transmission power according to the policy to achieve good system performance; there are a total of M optional spectrum sub-bands, and each spectrum sub-band has been occupied by a V2I link. The agent needs to select to share the spectrum resources with a V2I link to improve the utilization rate of the precious spectrum resources; the transmission power of the V2V link takes values in the continuous interval [-100, 23] dBm, rather than discretizing the action space through quantization, which improves the performance of the existing resource allocation algorithms with discrete action spaces; the combination of transmission frequency band selection and transmission power control forms the hybrid action space of the agent as follows:
[0061]
[0062] wherein, represents the spectrum sub-band selection of agent k at time slot t, represents the power selection of agent k at time slot t.
[0063] Furthermore, in step S3, in the Markov decision process, when reinforcement learning solves the target problem that is difficult to optimize in high-dimensional complex scenarios, the key lies in the design of its reward function; the reward function is set as a trade-off between the total capacity of the V2I link and the transmission success probability of the V2V link load, and a time penalty term is added. The reward r(t) at time slot t is:
[0064]
[0065] wherein, represents the channel capacity of the m-th V2I link on the m-th spectrum sub-band at time slot t, represents the channel capacity of the k-th V2V link on the m-th spectrum sub-band at time slot t, T represents the transmission time limit, represents the remaining transmission time of the k-th V2V link at time slot t, λ I 、λ V 、λ are hyperparameters selected according to experience.
[0066] Furthermore, in step S4, a Long Short-Term Memory (LSTM) network is introduced to extract the temporal information in the state space of the Markov decision process, and the temporal information extracted by the Long Short-Term Memory network is mapped through a fully connected layer to convert it into a dimension suitable for the input of the DDQN network and the DDPG network of each agent in the multi-agent network.
[0067] Furthermore, in step S5, the discrete spectrum sub-band selection is processed based on the DDQN network, and the continuous power control is processed based on the DDPG network to realize the joint optimization of continuous power allocation and discrete spectrum sub-band allocation in the multi-agent network, specifically as follows:
[0068] The action selection of spectrum sub-bands based on the DDQN network adopts the ε-greedy strategy to balance exploration and exploitation, as follows:
[0069]
[0070] Among them, θ represents the parameters of the main network of the DDQN network, a(t) and s(t + 1) respectively represent the possible actions corresponding to time slot t and the state of the next time slot t + 1 after selecting an action. The agent randomly selects an action with probability ε and selects the action that maximizes the Q value Q(s(t + 1), a(t); θ) of the main network Q with probability 1 - ε;
[0071] The target Q network is utilized in the DDQN network to calculate the target Q value y D (t) to suppress the problem of overestimating the Q value:
[0072]
[0073] γ D represents the discount factor of the DDQN network, θ target is the parameter of the target network, Q(s(t + 1), a(t); θ) is the Q value of the main network, and is the Q value of the target network;
[0074] The loss function L D (θ) of the DDQN network is calculated using the mean square error:
[0075] L D (θ) = E[(Q(s(t), a(t); θ) - y(t)) 2 ;
[0076] Among them, E[] represents the expectation;
[0077] The DDPG network includes an Actor network and a Critic network, and also adopts the structure of the main network and the target network. Among them, the policy-based Actor network, based on the environmental state s(t) collected at the current time slot t, then selects the corresponding power action a p (t) according to the policy π(s(t), φ); when the Actor makes an action decision for each observation, each agent aims to maximize the cumulative discounted reward, and the loss function L A (φ) of the Actor network is:
[0078]
[0079] Among them, φ is the parameter of the Actor network, is the parameter of the Critic network, is the Q-value of the Critic network;
[0080] The value-based critic Critic network uses the state-action Q-value function to evaluate the quality of the actions selected by the Actor network and is updated by randomly sampling a small batch of sample experiences from the experience pool; the Critic network adjusts the parameters of the evaluation network by minimizing the loss, and the Critic network utilizes the target Critic network Calculate the target Q-value y C (t) as:
[0081]
[0082] where γ C represents the discount factor of the Critic network, is the Q-value of the target Critic, and φ target are the parameters of the target Actor network, are the parameters of the target Critic network. The loss function of the Critic network is:
[0083]
[0084] where a p (t) is the power selection action of the Actor network, is the Q-value of the Critic network.
[0085] Furthermore, in step S6, the training of the Actor networks and Critic networks of the DDQN network and the DDPG network is regarded as different tasks of meta-learning, and the parameters of the DDQN network, Actor network, and Critic network are updated based on the meta-learning mechanism. For each task, individual-level and global-level updates are performed as follows:
[0086] The individual-level update is the process of learning the policy selection for each task. For the individual-level parameter θ task , the task parameters are updated through the loss L task of each task:
[0087]
[0088] where α represents the individual-level update rate, and L task represents the individual task loss;
[0089] The global-level update is performed on the basis of the individual-level update. First, the expectation L meta of the meta-loss of all individual-level updates is calculated:
[0090]
[0091] Among them, A represents that A rounds of individual-level updates have been carried out. represents the loss of the i-th task.
[0092] Finally, aggregate the gradients to update the global parameter θ global :
[0093]
[0094] Among them, β represents the global-level update rate.
[0095] Furthermore, in step S6, according to the obtained global parameters of the DDQN network and DDPG network of each agent, using the final spectrum sub-band selection and power selection of each agent as the optimal resource allocation parameters, dynamically adjust the transmission parameters of each vehicle according to the optimal resource allocation parameters to achieve the optimal resource allocation effect.
[0096] Compared with the prior art, the advantages of the present invention are as follows:
[0097] The present invention proposes a vehicle-to-everything (V2X) communication resource allocation method based on multi-agent deep reinforcement learning to achieve cooperation between vehicle user equipment (VUE). The present invention uses distributed DDPG agents to improve the error caused by quantifying power in previous studies and realizes effective processing of the continuous action space. By improving the policy gradient method, the algorithm of the present invention is more in line with the resource allocation scenario in the real communication environment. The present invention introduces long short-term memory (LSTM) to handle the time-dependent relationship of sequence data, make reasonable decisions using historical information, and thus improve the learning efficiency. By introducing meta-learning and training on multiple tasks, the model can learn the commonalities and differences between different tasks, thereby improving the stability and generalization ability of the learning process. It significantly improves the cooperation between agents and further optimizes the performance of the algorithm for the dynamic V2X environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0098] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0099] Figure 1 It is a schematic flowchart of a vehicle-to-everything (V2X) communication resource allocation method based on multi-agent deep reinforcement learning in an embodiment of the present invention.
[0100] Figure 2 It is a schematic diagram showing the change of the transmission rate of the vehicle-to-infrastructure (V2I) link with the transmission payload in the test environment disclosed in an embodiment of the present invention.
[0101] Figure 3 Schematic diagram of the change of the transmission success rate of the V2V link with the transmission payload in the test environment disclosed in the embodiment of the present invention.
[0102] Figure 4 Schematic diagram of the change of the remaining transmission payload of each V2V link with time in the test environment disclosed in the embodiment of the present invention. Detailed implementation manners
[0103] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0104] In one embodiment, a communication resource allocation method for a vehicle-to-everything (V2X) network based on multi-agent deep reinforcement learning, as Figure 1 shown, includes the following steps:
[0105] S1: In the cellular vehicle-to-everything (C-V2X) scenario, initialize the number, position, speed and direction of vehicles, the size of the transmission payload, and the transmission time limit; construct a wireless channel model according to the 3GPP standard, and calculate the signal-to-interference-plus-noise ratio (SINR) of the V2V and V2I links;
[0106] According to the 3GPP standard, the large-scale channel fading of the wireless channel model is composed of path loss and shadow effect, and is divided into two scenarios: line-of-sight (LOS) and non-line-of-sight (NLOS); the path loss between V2V links in the LOS case is calculated by the formula:
[0107]
[0108] where f c is the system frequency, h′ BS , h′ MS are the effective antenna heights of the base station and the mobile user respectively; in one embodiment, d′ BP represents the break point distance greater than 3 meters, which is jointly determined by the system frequency and the effective antenna heights of the base station and the mobile user, d a and d b represent the horizontal and vertical distances between the two vehicles performing V2V link communication respectively, and d represents the Euclidean distance between the two vehicles performing V2V link communication;
[0109] The path loss between V2V links in the NLOS case is calculated by the formula:
[0110]
[0111] n j = max(2.8 - 0.0024d b , 1.84);
[0112] where f c is the system frequency; and respectively represent the calculation results in two dimensions of the horizontal and vertical distances of the handover, and finally take the minimum value among them. n j represents the occlusion correction factor of the jth occluder encountered by the vehicle in the communication environment. In one embodiment, j = 1;
[0113] The path loss of the V2I link is within the line-of-sight range, and its calculation formula is:
[0114]
[0115] where h' BS and h' MS are the effective antenna heights of the base station and the mobile user respectively, and d' represents the Euclidean distance between the vehicle performing V2I link communication and the base station;
[0116] Shadow fading can be simplified as:
[0117]
[0118] where S * (n) is an independent and identically distributed random vector in the nth time slot, which is a normal distribution with a mean of 0 and a standard deviation of 0.02; D describes the driving distance of the vehicle in a time slot, and D dEco represents the decorrelation distance, which is different in V2I and V2V links. In one embodiment, it is 50m in the V2I link and 10m in the V2V link; e is the base of the natural logarithm, S(n) represents the shadow fading component in the nth time slot, and S(n - 1) represents the shadow fading component in the (n - 1)th time slot;
[0119] Assume that all vehicle transceivers use one antenna. The vehicle network includes M V2I links and K V2 links, which are represented by the sets M' = {1,..., m,..., M} and K' = {1,..., k,..., K} respectively;
[0120] Assume that the V2I links have been pre-allocated orthogonal spectrum sub-bands, that is, the mth V2I link occupies the mth spectrum sub-band, and the transmit power of the V2I link is a fixed value of 23 dBm;
[0121] The V2V link realizes spectrum sharing through band selection and power control, where ρ k [m] is a spectrum sub - band selection variable of boolean value, indicating whether the k - th V2V link transmits on the m - th spectrum sub - band. If the k - th V2V link transmits on the m - th spectrum sub - band, then ρ k [m]=1; otherwise ρ k [m]=0;
[0122] Assume that the channel fading is the same within a spectrum sub - band and independent between different spectrum sub - bands. h k [m] is the power component of the small - scale fading of the k - th V2V link on the m - th spectrum sub - band, and it follows an exponential distribution;
[0123] α k represents the large - scale fading of the k - th V2V link independent of frequency. Then, within a coherence time, the channel power gain of the k - th V2V link on the m - th spectrum sub - band is expressed as
[0124] Assume that the interference channel power gain from the transmitter of the k'- th V2V link through the m - th spectrum sub - band to the receiver of the k - th V2V link is expressed as The interference channel power gain from the transmitter of the k - th V2V link to the m - th V2I link on the m - th spectrum sub - band is expressed as The channel gain from the transmitter of the m - th V2I link to the base station on the m - th spectrum sub - band is expressed as The interference channel power gain from the transmitter of the m - th V2I link to the receiver of the k - th V2V link on the m - th spectrum sub - band is expressed as The power of the transmitted signal of the transmitter of the m - th V2I link is The transmit power magnitude of the transmitter of the k - th V2V link on the m - th spectrum sub - band is The noise power is σ 2 ;
[0125] The signal - to - interference - plus - noise ratio (SINR) of the m - th V2I link on the m - th spectrum sub - band is calculated as follows:
[0126]
[0127] The signal - to - interference - plus - noise ratio (SINR) of the k - th V2V link on the m - th spectrum sub - band is calculated as follows:
[0128]
[0129] S2: Construct a vehicle - to - everything (V2X) network resource allocation problem model according to the signal - to - interference - plus - noise ratio (SINR) of vehicle - to - vehicle (V2V) and vehicle - to - infrastructure (V2I) links, as follows:
[0130] Assume that each V2V link can access at most one orthogonal spectrum sub - band, i.e., ∑ m ρ k [m] ≤ 1. Given that W is the bandwidth of each spectrum sub - band, according to the Shannon formula, the channel capacity of the m - th V2I link on the m - th spectrum sub - band is expressed as:
[0131]
[0132] Similarly, the channel capacity of the k - th V2V link on the m - th spectrum sub - band is expressed as:
[0133]
[0134] Since V2I links are usually used to support high - transmission - rate entertainment services, this can be represented as the maximization of the sum - rate of V2I links; V2V links mainly focus on the reliable transmission of safety - critical messages, aiming to achieve a high successful transmission probability for all V2V users during information transmission; compared with V2I links, the objective function of V2V links is more complex; therefore, first define the successful transmission of each V2V link through the following conditions:
[0135]
[0136] where T represents the transmission time limit, Δ T represents the duration of a transmission time slot, t k represents the time slot when the k - th V2V link starts to transmit information, and B represents the size of the transmitted information payload;
[0137] Introduce a binary indicator ω k,τ as a quantization index for successful transmission, representing the successful state of the k - th V2V link in the τ - th payload transmission:
[0138]
[0139] The success probability η of all V2V links within the limited transmission time period is:
[0140]
[0141] where O k represents the actual number of attempts to transmit by this V2V link during the transmission period. The specific rule is: if the cumulative data volume reaches or exceeds B at the τ - th attempt, then O k = τ; if at T / ΔT If the transmission is still not completed after the attempt, then O k = T / Δ T , at this time ω k,τ = 0.
[0142] The definition of successful V2V transmission is measured by the data accumulation within the time limit, emphasizing the collaborative optimization of low latency and high reliability; this definition directly affects the design of the resource allocation strategy, requiring the algorithm to consider both rate and latency when selecting spectrum subbands and allocating power to meet the transmission requirements of safety messages. The vehicle-to-everything (V2X) resource allocation problem model is as follows:
[0143]
[0144] In one embodiment, P max represents the maximum transmit power of the vehicle transmitter for V2V communication, with a value of 23 dBm. P I represents the transmit power of the vehicle transmitter for V2I communication, with a fixed value of 23 dBm.
[0145] S3: Model the V2X resource allocation problem as a Markov decision process, define the state space, action space, and reward function, and build the basis for intelligent decision-making logic, as follows:
[0146] Each V2V link acts as an agent, interacting with the unknown environment to obtain experience, and then learning the optimal policy from the experience; each agent has its own DDQN and DDPG networks; K V2V links form a multi-agent network to jointly explore an environment and optimize the power control strategy according to the changes in the environmental state;
[0147] At time slot t, given the environmental state of the current agent k Then output the action according to the policy The joint state space of multiple agents is S t , when each agent obtains its own action, multiple agents form a joint action A t , and the agents jointly execute the action to obtain the reward R t , where all agents obtain the same reward in the system to promote their cooperative behavior. The environmental state of agent k transfers to the environmental state at the next moment with probability For a single agent k, it can only obtain the local environmental state where it is located and the action A t , while the global channel condition and the actions of other agents are unknown;
[0148] All channel information observed by a single agent k on the m-th V2I link at time slot t including the channel gain of agent k itself on the m-th V2I link at time slot t and the interference channel power gain from the transmitters of other V2V links k' to agent k itself on the m-th V2I link at time slot t For all m ∈ M, the interference channel power gain from the transmitter of agent k to the base station is and the interference channel from the transmitter of the m-th V2I link to agent k itself on the m-th V2I link at time slot t is the interference of agent k on the m-th V2I link in the previous time slot t - 1 the number of times the spectrum sub-band is occupied by agent k on the m-th V2I link in the previous time slot t - 1 the remaining payload of agent k in the previous time slot t and the remaining transmission time Therefore, the state of agent k at time slot t is:
[0149]
[0150] Each agent independently selects the transmission frequency band and controls the transmission power according to the policy to achieve good system performance; there are a total of M optional spectrum sub-bands, and each spectrum sub-band has been occupied by a V2I link. The agent needs to select to share the spectrum resources with a V2I link to improve the utilization rate of the precious spectrum resources; the transmission power of the V2V link takes values in the continuous interval [-100, 23] dBm, rather than discretizing the action space through quantization, which improves the performance of the existing resource allocation algorithms with discrete action spaces; the combination of transmission frequency band selection and transmission power control forms the hybrid action space of agent k as follows:
[0151]
[0152] where represents the spectrum sub-band selection of agent k at time slot t, represents the power selection of agent k at time slot t.
[0153] Furthermore, in step S3, in the Markov decision process, when reinforcement learning solves the problem of difficult-to-optimize objectives in high-dimensional complex scenarios, the key lies in the design of its reward function; the reward function is set as a trade-off between the total capacity of the V2I link and the transmission success probability of the V2V link load, and a time penalty term is added. The reward r(t) at time slot t is:
[0154]
[0155] Among them, represents the channel capacity of the m-th V2I link in time slot t on the m-th spectrum sub-band, represents the channel capacity of the k-th V2V link in time slot t on the m-th spectrum sub-band, T represents the transmission time limit, represents the remaining transmission time of the k-th V2V link in time slot t, λ I and λ V and λ are hyperparameters selected according to experience. In one embodiment, λ I = 0.1, λ V = 0.9, and λ = 1.
[0156] S4: Introduce a Long Short-Term Memory (LSTM) network to extract the temporal information in the state space of the Markov decision process, and map the temporal information extracted by the LSTM network through a fully connected layer to convert it into a dimension suitable for the input of the DDQN network and the DDPG network for each agent in the multi-agent network (Sherstinsky A.Fundamentals of recurrent neural network(RNN)and long short-term memory(LSTM)network[J].Physica D:Nonlinear Phenomena,2020,404:132306.).
[0157] S5: Adopt a divide-and-conquer strategy to perform resource allocation. Use the DDQN network to handle the selection of discrete spectrum sub-bands, and utilize its efficient decision-making characteristics for the discrete action space to complete the precise allocation of spectrum sub-bands. With the help of the DDPG network to handle continuous power control, aiming at the refined adjustment requirements of the continuous action space, achieve the optimal regulation of the transmission power, and finally realize the joint optimization of continuous power allocation and discrete spectrum sub-band allocation in the multi-agent network, specifically as follows:
[0158] The selection of spectrum sub-band actions based on the DDQN network adopts the ε-greedy strategy to balance exploration and exploitation, as follows:
[0159]
[0160] Among them, θ represents the main network parameters of the DDQN network, a(t) and s(t + 1) respectively represent the possible actions corresponding to time slot t and the state of the next time slot t + 1 after selecting an action. The agent randomly selects an action with probability ε and selects the action that maximizes the Q value Q(s(t + 1), a(t); θ) of the main network Q with probability 1 - ε;
[0161] Utilizing the target Q-network in the DDQN network Calculate the target Q-value y D (t) to suppress the problem of overestimating the Q-value:
[0162]
[0163] γ D represents the discount factor of the DDQN network, θ target is the target network parameter, Q(s(t+1),a(t);θ) is the main network Q-value, is the target network Q-value;
[0164] The loss function L D (θ) of the DDQN network is calculated using the mean squared error:
[0165] L D (θ) = E[(Q(s(t),a(t);θ) - y(t)) 2 ;
[0166] where E[] represents the expectation;
[0167] The DDPG network includes an Actor network and a Critic network, and also adopts the structure of a main network and a target network. Among them, the policy-based Actor network, based on the environmental state s(t) collected at the current time slot t, then selects the corresponding power action a p (t); when the Actor makes an action decision for each observation, each agent aims to maximize the cumulative discounted reward. The loss function L A (φ) of the Actor network is:
[0168]
[0169] where φ is the parameter of the Actor network, is the parameter of the Critic network, is the Q-value of the Critic network;
[0170] The value-based critic Critic network uses the state-action Q-value function to evaluate the quality of the actions selected by the Actor network and is updated by randomly sampling a small batch of sample experiences from the experience pool; the Critic network adjusts the parameters of the evaluation network by minimizing the loss. The Critic network utilizes the target Critic network Calculate the target Q-value y C (t) as:
[0171]
[0172] where γC Denotes the discount factor of the Critic network, is the Q-value of the target Critic, φ target are the parameters of the target Actor network, are the parameters of the target Critic network. The loss function of the Critic network is:
[0173]
[0174] where a p (t) is the power selection action of the Actor network, is the Q-value of the Critic network.
[0175] S6: Update the parameters of the DDQN network and the DDPG network of each agent based on the meta-learning mechanism. Through multi-task training, the DDQN network and the DDPG network learn the common features and different rules of different resource allocation tasks, strengthen the adaptability of the algorithm in the dynamic vehicle networking environment, and complete the communication resource allocation in the vehicle networking.
[0176] Regarding the training of the Actor network and the Critic network of the DDQN network and the DDPG network as different tasks of meta-learning, update the parameters of the DDQN network, the Actor network, and the Critic network based on the meta-learning mechanism. For each task, both individual-level and global-level updates are performed, as follows:
[0177] The individual-level update is the process of learning the policy selection for each task. For the individual-level parameter θ task , update the task parameters through the loss L task of each task:
[0178]
[0179] where α represents the individual-level update rate, and L task represents the individual task loss;
[0180] The global-level update is performed on the basis of the individual-level update. First, calculate the expectation L meta of the meta-loss of all individual-level updates:
[0181]
[0182] where A represents that a total of A rounds of individual-level updates have been performed, represents the loss of the i-th task.
[0183] Finally, aggregate the gradients to update the global parameter θ global :
[0184]
[0185] Among them, β represents the global-level update rate.
[0186] According to the global parameters obtained for the DDQN network and DDPG network of each agent, taking the final spectrum sub-band selection and power selection of each agent as the optimal resource allocation parameters, dynamically adjusting the transmission parameters of each vehicle according to the optimal resource allocation parameters to achieve the optimal resource allocation effect.
[0187] In one embodiment, a simulation experiment is conducted. In the simulation experiment, the main parameters are shown in Table 1.
[0188] Table 1
[0189] Parameter Value Number M of V2I links 4 Number K of V2I links 4 Transmission power of the transmitting end of V2I links 23 dBm Transmission power of the transmitting end of V2V links [-100, 23] dBm Maximum transmission power of the transmitting end of V2V links 23 dBm Carrier frequency 2 GHz Bandwidth 4 MHz Height of the base station antenna 25m Gain of the base station antenna 8 dBi Noise figure of the base station receiver 5 dB Height of the vehicle antenna 1.5m Gain of the vehicle antenna 3 dBi Noise figure of the vehicle receiver 9 dB Vehicle speed 36 km / h <![CDATA[Noise factor σ 2 > 9 dB Maximum time limit T 100 ms
[0190] In one embodiment, as Figure 2 shown, the variation of the sum rate of the V2I link with the V2V link load size for four different algorithms is presented, namely the exhaustive algorithm, the proposed algorithm MADRL-MAML of the present invention, the ordinary multi-agent deep reinforcement learning algorithm MADRL, and the random algorithm. The proposed algorithm of the present invention is close to the performance of the exhaustive algorithm and is significantly superior to the ordinary multi-agent deep reinforcement learning algorithm and the random algorithm.
[0191] In one embodiment, as Figure 3 shown, the variation of the transmission success rate of the V2V link with the V2V link load size for four different algorithms is presented, namely the exhaustive algorithm, the proposed algorithm MADRL-MAML of the present invention, the ordinary multi-agent deep reinforcement learning algorithm MADRL, and the random algorithm. The proposed algorithm of the present invention is second only to the performance of the exhaustive algorithm and is still superior to the ordinary multi-agent deep reinforcement learning algorithm and the random algorithm. Especially when the load increases, it can still maintain a higher transmission success rate of the V2V link.
[0192] In one embodiment, as Figure 4 shown, the variation of the remaining transmission load of the proposed algorithm of the present invention with the transmission time when the total transmission load size is 2 * 1060 Bytes is presented. It can be seen that the remaining loads of the 4 V2V links all drop to 0 within 30 ms, demonstrating that the proposed algorithm of the present invention can complete reliable transmission within the maximum time limit.
[0193] The method provided by the foregoing embodiments of the present invention builds a multi-agent deep reinforcement learning model. By training multiple DDPG agents, in the scenario of cooperative C-V2X, spectrum allocation and continuous power control are achieved through deep reinforcement learning, solving the quantization error problem caused by the existing resource allocation algorithms with discrete action spaces, and ensuring good V2V link reliability and vehicle networking system performance. By introducing LSTM and meta-learning, the method of the present invention addresses the problems of sequential data and dynamic environmental changes.
[0194] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A method for allocating communication resources in a vehicle network based on multi-agent deep reinforcement learning, characterized in that, It includes the following steps: S1: In the cellular vehicle-to-everything (C-V2X) scenario, initialize the number, location, speed, direction, transmission payload size, and transmission time limit of vehicles; construct a wireless channel model according to the 3GPP standard, and calculate the signal-to-interference-plus-noise ratio (SINR) of vehicle-to-vehicle (V2V) and vehicle-to-infrastructure (V2I) links; S2: Construct a vehicle network resource allocation problem model according to the SINR of V2V and V2I links; S3: Model the vehicle network resource allocation problem as a Markov decision process (MDP), define the state space, action space, and reward function, and build the basis of intelligent decision-making logic; S4: Use the long short-term memory (LSTM) network to extract the temporal information in the state space, deeply mine the dynamic change law of the communication environment, and convert the temporal information into the dimension suitable for the input of the deep double Q-network (DDQN) and deep deterministic policy gradient (DDPG) networks of each agent in the multi-agent network of the MDP; S5: Adopt a divide-and-conquer strategy to execute resource allocation. Use the DDQN network to handle the selection of discrete spectrum sub-bands to complete the allocation of spectrum sub-bands, and use the DDPG network to handle continuous power control to achieve the optimal regulation of the transmission power, and finally realize the joint optimization of continuous power allocation and discrete spectrum sub-band allocation in the multi-agent network; S6: Update the parameters of the DDQN and DDPG networks of each agent based on the meta-learning mechanism. Through multi-task training, make the DDQN and DDPG networks learn the common characteristics and difference laws of different resource allocation tasks, strengthen the adaptability of the algorithm in the dynamic vehicle network environment, and complete the allocation of vehicle network communication resources.
2. The method for allocating communication resources in a vehicle network based on multi-agent deep reinforcement learning according to claim 1, wherein In step S1, a wireless channel model is constructed according to the 3GPP standard. The large-scale channel fading is jointly composed of path loss and shadowing effect, and is divided into two scenarios: Line of Sight (LOS) and Non-Line of Sight (NLOS). The path loss between V2V links in the case of LOS is calculated by the formula: where f c is the system frequency, h′ BS , h′ MS are the effective antenna heights of the base station and the mobile user respectively; d′ BP represents the breakpoint distance, d a and d b represent the horizontal and vertical distances between two vehicles communicating via a V2V link respectively, and d represents the Euclidean distance between two vehicles communicating via a V2V link; Path Loss between V2V Links in Non-Line-of-Sight Conditions The calculation formula is as follows: where f c is the system frequency; and respectively represent the calculation results in two dimensions of the horizontal and vertical distances of the handover, and finally take the minimum value among them. n j represents the occlusion correction factor of the j-th occluder encountered by the vehicle in the communication environment; Path Loss of the V2I Link All are within the line-of-sight range, and its calculation formula is as follows: where h′ BS and h′ MS are the effective antenna heights of the base station and the mobile user, respectively, and d' represents the Euclidean distance between the vehicle performing V2I link communication and the base station; The shadow fading can be simplified as: Among them, S * (n) is an independent and identically distributed random vector in the n-th time slot, following a normal distribution with a mean of 0 and a standard deviation of 0.02; D describes the driving distance of the vehicle in one time slot, and D dEco represents the decorrelation distance, which is different in V2I and V2V links; e is the base of the natural logarithm, S(n) represents the shadow fading component in the n-th time slot, and S(n - 1) represents the shadow fading component in the (n - 1)-th time slot; Assume that the transceivers of all vehicles use one antenna. There are M V2I links and K V2V links in the vehicle network, which are represented by sets M' = {1, …, m, …, M} and K' = {1, …, k, …, K} respectively; Assume that orthogonal spectrum sub-bands have been pre-allocated to V2I links, that is, the m-th V2I link occupies the m-th spectrum sub-band, and the transmission power of the V2I link is a fixed value of 23 dBm; The V2V link realizes spectrum sharing through band selection and power control, and ρ k [m] is a spectrum sub-band selection variable of boolean value, indicating whether the k-th V2V link transmits on the m-th spectrum sub-band. If the k-th V2V link transmits on the m-th spectrum sub-band, then ρ k [m] = 1, otherwise ρ k [m] = 0; Assume that the channel fading is the same within a spectral sub - band and independent between different spectral sub - bands, and \(h k [m]\) is the power component of the small - scale fading of the \(k\) - th V2V link on the \(m\) - th spectral sub - band, which follows an exponential distribution; α k represents the large-scale fading of the k-th V2V link independent of frequency. Then, within a coherence time, the channel power gain of the k-th V2V link on the m-th spectral subband is expressed as Suppose the interference channel power gain from the transmitter of the k'-th V2V link to the receiver of the k-th V2V link through the m-th spectrum subband is denoted as The interference channel power gain from the transmitter of the k-th V2V link to the m-th V2I link in the m-th spectrum subband is denoted as The channel gain from the transmitter of the m-th V2I link to the base station in the m-th spectrum subband is denoted as The interference channel power gain from the transmitter of the m-th V2I link to the receiver of the k-th V2V link in the m-th spectrum subband is denoted as The power of the transmitted signal of the transmitter of the m-th V2I link is The magnitude of the transmission power of the transmitter of the k-th V2V link in the m-th spectrum subband is The noise power is σ 2 ; The signal-to-interference-plus-noise ratio (SINR) of the m-th V2I link on the m-th spectrum sub-band is calculated by the following expression: The signal-to-interference-plus-noise ratio (SINR) of the k-th V2V link on the m-th spectral subband is calculated by the following expression:
3. The method for allocating communication resources in a vehicle network based on multi-agent deep reinforcement learning according to claim 2, wherein, In step S2, construct a vehicle network resource allocation problem model as follows: Assume that each V2V link can access at most one orthogonal spectrum sub-band, i.e., Σ m ρ k [m] ≤ 1, where W is the bandwidth of each spectrum sub-band. Then, according to the Shannon formula, the channel capacity of the m-th V2I link on the m-th spectrum sub-band is expressed as: Similarly, the channel capacity of the k-th V2V link on the m-th spectrum sub-band is expressed as: Define the successful transmission of each V2V link through the following conditions: where T represents the transmission time limit, Δ T represents the duration of a transmission time slot, t k represents the time slot when the k-th V2V link starts transmitting information, and B represents the size of the transmission information payload; Introduce the binary indicator ω k,τ As a quantization metric for successful transmission, it represents the success status of the k-th V2V link in the τ-th payload transmission: The success probability η of all V2V links within the limited transmission time period is: Among them, O k represents the actual number of attempts to transmit during the transmission cycle of the V2V link. The specific rule is as follows: If the cumulative data volume reaches or exceeds B at the τ-th attempt, then O k = τ; If the transmission is not completed after T / Δ T attempts, then O k = T / Δ T , and at this time ω k,τ = 0.
4. The method for allocating communication resources in a vehicle network based on multi-agent deep reinforcement learning according to claim 3, characterized in that In step S2, the vehicle network resource allocation problem model is as follows: Among them, P max represents the maximum transmission power of the vehicle transmitter for V2V communication, and P I represents the transmission power of the vehicle transmitter for V2I communication.
5. The method for allocating communication resources in a vehicle networking according to claim 4, wherein In step S3, model the vehicle network resource allocation problem as an MDP as follows: Each V2V link acts as an agent, interacts with the unknown environment to obtain experience, and then learns the optimal policy from the experience; each agent has its own DDQN and DDPG networks; K V2V links form a multi-agent network to jointly explore an environment and optimize the power control strategy according to the change of the environmental state; At time slot t, given the environmental state of the current agent k Then, according to the policy, the action is output The joint state space of multiple agents is S t , after each agent obtains its respective action, multiple agents form a joint action A t , after the agents jointly execute the action, they obtain a reward R t , where all agents obtain the same reward in the system to promote their cooperative behavior, and the environmental state where agent k is located With probability Transfers to the environmental state at the next moment For a single agent k, it can only obtain the local environmental state of its own location And action A t , while the global channel condition and the actions of other agents are unknown; All channel information observed by a single agent k on the m-th V2I link at time slot t including the channel gain of agent k itself on the m-th V2I link at time slot t and the interference channel power gain from the transmitter of other V2V link k' to agent k itself on the m-th V2I link at time slot t For all m ∈ M, the interference channel power gain from the transmitter of agent k to the base station is and the interference channel from the transmitter of the m-th V2I link to agent k itself on the m-th V2I link at time slot t is the interference of agent k on the m-th V2I link at the previous time slot t - 1 the number of spectrum sub-bands occupied by agent k on the m-th V2I link at the previous time slot t - 1 the remaining payload of agent k at the previous time slot t and the remaining transmission time Therefore, the state of agent k at time slot t is as follows: Each agent independently performs transmission band selection and transmit power control according to the policy; there are a total of M optional spectrum sub-bands, and each spectrum sub-band has been occupied by a V2I link. The agent needs to select to share the spectrum resources with a V2I link to improve the utilization rate of the precious spectrum resources; the transmit power of the V2V link takes values in the continuous interval [-100, 23] dBm, rather than discretizing the action space through quantization, which improves the performance of the existing resource allocation algorithms with discrete action spaces; the combination of transmission band selection and transmit power control forms the hybrid action space of the agent as follows: Among them, represents the spectrum sub-band selection of agent k at time slot t, represents the power selection of agent k at time slot t.
6. The method for allocating communication resources in a vehicle network based on multi-agent deep reinforcement learning according to claim 5, wherein In step S3, in the Markov decision process, the reward function is set as a trade-off between the total capacity of the V2I link and the transmission success probability of the V2V link load, and a time penalty term is added. The reward r(t) at time slot t is: Among them, represents the channel capacity of the m-th V2I link in the time slot t on the m-th spectrum sub-band, represents the channel capacity of the k-th V2V link in the time slot t on the m-th spectrum sub-band, and T represents the transmission time limit, represents the remaining transmission time of the k-th V2V link in the time slot t, and λ I and λ V are hyperparameters selected according to experience.
7. The method for allocating communication resources in a vehicle network based on multi-agent deep reinforcement learning according to claim 6, wherein In step S4, a long short-term memory network is introduced to extract the temporal information in the state space of the Markov decision process, and the temporal information extracted by the long short-term memory network is mapped through a fully connected layer and converted into a dimension suitable for the input of the DDQN network and the DDPG network of each agent in the multi-agent network.
8. The method for allocating communication resources in a vehicle networking according to claim 7, wherein In step S5, the DDQN network is used to handle the discrete spectrum sub-band selection, and the DDPG network is used to handle the continuous power control to achieve the joint optimization of continuous power allocation and discrete spectrum sub-band allocation in the multi-agent network, as follows: The ε-greedy strategy is adopted for the spectrum sub-band action selection based on the DDQN network to balance exploration and exploitation, as follows: where θ represents the parameters of the main network of the DDQN network, a(t) and s(t + 1) respectively represent the possible actions corresponding to time slot t and the state of the next time slot t + 1 after selecting the action. The agent randomly selects an action with probability ε and selects the action that maximizes the Q value Q(s(t + 1), a(t); θ) of the main network Q with probability 1 - ε; Using the target Q-network in the DDQN network Calculate the target Q-value y D (t) to suppress the problem of overestimating the Q-value: γ D represents the discount factor of the DDQN network, θ target is the target network parameter, Q(s(t + 1), a(t); θ) is the main network Q value, is the target network Q value; The loss function L of the DDQN network D (θ) is calculated using the mean squared error: L D (θ) = E[(Q(s(t), a(t); θ) - y(t)) 2 ; where E[] represents the expectation; The DDPG network consists of an Actor network and a Critic network, also adopting the structure of a main network and a target network. Among them, the policy-based Actor network, based on the environmental state s(t) of the current time slot t collected, then selects the corresponding power action a according to the policy π(s(t), φ). p (t); when the Actor makes an action decision for each observation, each agent aims to maximize the cumulative discounted reward, and the loss function L A (φ) is as follows: where φ are the parameters of the Actor network, are the parameters of the Critic network, is the Q value of the Critic network; The value-based Critic network uses the state-action Q-value function to evaluate the quality of the actions selected by the Actor network and is updated by randomly sampling experiences from the experience pool; the Critic network adjusts the parameters of the evaluation network by minimizing the loss, and the Critic network utilizes the target Critic network Calculate the target Q-value y C (t) as follows: Among them, γ C represents the discount factor of the Critic network, is the Q value of the target Critic, φ target are the parameters of the target Actor network, are the parameters of the target Critic network; Loss function of the Critic network is as follows: Among them, a p (t) is the power selection action of the Actor network, is the Q value of the Critic network.
9. The method for allocating communication resources in a vehicle network based on multi-agent deep reinforcement learning according to claim 8, wherein In step S6, the training of the Actor network and the Critic network of the DDQN network and the DDPG network is regarded as different tasks of meta-learning, and the parameters of the DDQN network, the Actor network, and the Critic network are updated based on the meta-learning mechanism. For each task, both individual-level and global-level updates are performed, as follows: The individual-level update is the process of learning policy selection for each task, for the individual-level parameter θ task , through the loss L of each task task Update the task parameters: Among them, α represents the individual-level update rate, and L task represents the individual task loss; The global-level update is performed on the basis of the individual-level updates. First, the expected value L of the meta-loss of all individual-level updates is calculated meta : Among them, A represents that a total of A rounds of individual-level updates have been performed, represents the loss of the i-th task; Finally, aggregate the gradients to update the global parameter θ global : where β represents the global-level update rate.
10. The method for allocating communication resources in a vehicle networking according to any one of claims 1-9, characterized in that, In step S6, according to the global parameters of the DDQN network and the DDPG network of each agent, the final spectrum sub-band selection and power selection of each agent are used as the optimal resource allocation parameters, and the transmit parameters of each vehicle are dynamically adjusted according to the optimal resource allocation parameters to achieve the optimal resource allocation effect.
Citation Information
Cited By
Resource allocation method and system of cellular Internet of Vehicles based on graph reinforcement learning
CN121692412A
Multi-agent reinforcement learning Internet of Vehicles resource allocation method, device and system, and storage medium
CN121888210A
D2D communication resource allocation method
CN122227419A