Internet of vehicles communication resource allocation method based on federal multi-agent deep reinforcement learning
Through the deep reinforcement learning method of federated multi-agents, combined with asynchronous federated learning and dynamic weight adjustment, the problems of privacy leakage, low communication efficiency and insufficient dynamic adaptability in the Internet of Vehicles are solved, and efficient and secure resource allocation optimization is achieved.
Patent Information
- Application Number
- CN202510566259.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art has problems such as privacy leakage risks, low communication efficiency, insufficient dynamic adaptability and limited scalability in the Internet of Vehicles. Especially in the resource allocation of V2I and V2V links, it is difficult to achieve efficient and privacy-protected dynamic resource allocation.
A method based on federated multi-agent deep reinforcement learning is adopted, and a method based on asynchronous federated learning and multi-agent deep deterministic strategy gradient algorithm is combined with a dynamic weight adjustment mechanism to achieve privacy protection and resource allocation optimization among vehicles.
It significantly improves spectrum efficiency, transmission success rate and system adaptability, ensures data security, and supports flexible expansion and resource allocation efficiency in large-scale Internet of Vehicles scenarios.
Smart Images

Figure CN120390236A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the deep cross - technical field of intelligent transportation, wireless communication, and distributed machine learning, and is related to the communication optimization of the vehicle - to - everything network driven by artificial intelligence. Specifically, it relates to a method for allocating communication resources in a vehicle - to - everything network based on federated multi - agent deep reinforcement learning. Background Art
[0002] With the rapid development of intelligent transportation systems and vehicle - to - everything (V2X) networks, efficient communication between vehicles and between vehicles and infrastructure has become an important support for realizing intelligent driving, vehicle - road coordination, and traffic optimization. Vehicle - to - infrastructure (V2I) communication and vehicle - to - vehicle (V2V) communication are two core modes in V2X communication. Among them, the V2I link mainly serves high - data - rate applications (such as high - definition video transmission), while the V2V link needs to transmit critical safety messages (such as collision warnings) within a specified time delay. To make full use of limited spectrum resources, the V2V link needs to dynamically reuse the channel resources allocated to the V2I link, which leads to interference and resource competition between the links. In such a vehicle - to - everything network with shared channels, how to reasonably allocate communication resources (such as spectrum resources and transmission power) has become an urgent problem to be solved. On the one hand, the efficient allocation of spectrum resources can reduce conflicts and interference between links, improve spectrum efficiency and network transmission reliability; on the other hand, the dynamic optimization of power resources helps to balance the link signal quality and interference level, thereby improving the transmission success rate of the V2V link and the total capacity of the V2I link.
[0003] The IoV communication resource allocation method based on federated multi-agent deep reinforcement learning (AFL-MADDPG) is an innovative technology designed to address the highly dynamic, privacy-sensitive, and distributed collaboration requirements of IoV scenarios. The high-speed movement of vehicles in IoV scenarios leads to frequent changes in channel states and network topology. Traditional static resource allocation methods based on mathematical programming or game theory rely on global information and have high computational complexity, making them difficult to respond to dynamic environments in real time. Centralized multi-agent deep reinforcement learning (such as MADDPG) can model multi-vehicle collaboration, but requires uploading global state data, which poses privacy risks and incurs significant communication overhead. Among existing similar technologies, non-federated distributed deep reinforcement learning (such as Distri-DDPG) protects privacy through local training, but suffers from data silos and heterogeneous vehicle data, leading to model convergence difficulties and low collaboration efficiency. Combinations of federated learning and single-agent reinforcement learning (such as FedDDPG) achieve privacy protection but neglect the competitive and cooperative relationships among multiple agents, making it difficult to optimize global resource utilization. In addition, existing methods generally face core defects such as low communication efficiency (frequent parameter synchronization increases network load), insufficient dynamic adaptability (fixed update cycle cannot cope with sudden channel changes) and limited scalability (model dimension explodes when the number of vehicles surges). Summary of the Invention
[0004] In order to solve the technical problems existing in the prior art, the present invention provides a method for allocating communication resources in the Internet of Vehicles based on federated multi-agent deep reinforcement learning. This method integrates the distributed privacy protection framework of federated learning with the collaborative decision-making mechanism of multi-agent deep reinforcement learning, uses MADDPG to model the interaction relationship between vehicles in local training, and dynamically optimizes the global strategy through federated aggregation. This not only avoids the sharing of sensitive data, but also improves resource allocation efficiency through distributed collaboration. At the same time, it introduces an adaptive dynamic weight adjustment mechanism to reduce network overhead and accelerate model convergence in heterogeneous data environments, providing a new paradigm of "privacy-efficiency-real-time" collaborative optimization for dynamic resource allocation in the Internet of Vehicles.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A method for allocating communication resources in the Internet of Vehicles (IoV) based on federated multi-agent deep reinforcement learning. The steps are as follows:
[0007] S1. In the IoV scenario, each vehicle is modeled as an agent and locally trained using a multi-agent deep deterministic policy gradient algorithm. Each agent selects actions based on local observations to perform spectrum access, transmit power control, and bandwidth allocation.
[0008] S2. Adopt the asynchronous federated learning method. Each agent uploads its local model parameters to the global server, and the model parameters are weighted and aggregated through a dynamic weight adjustment mechanism to generate a global model.
[0009] S3. The dynamic weight adjustment mechanism calculates weights based on the model update quality, communication quality, and update frequency of the agents to optimize the global model aggregation.
[0010] S4. The updated global model is sent to each agent to achieve distributed privacy protection and dynamic optimization of vehicle networking communication resources.
[0011] Furthermore, in step S1, the local observation state includes: the vehicle's own channel gain, the interference of other V2V links, the interference of the V2I link to the vehicle, the interference of the base station, the bandwidth requirement of the vehicle's current task, the task priority, the global network load, the global interference level, the communication quality, the model update quality, and the model update frequency.
[0012] Furthermore, in step S1, the designed action space includes the following variables:
[0013] A discrete spectrum sub-band selection variable for dynamically accessing the orthogonal sub-bands or idle sub-bands of the V2I link;
[0014] A continuous transmit power control variable with a value range from the predefined minimum transmit power to the maximum transmit power;
[0015] A continuous bandwidth allocation variable with a value range from the minimum bandwidth requirement to the maximum available bandwidth to adapt to the resource requirements of different priority tasks.
[0016] Furthermore, in step S1, the reward function of the multi-agent deep deterministic policy gradient algorithm consists of the following parameters:
[0017] The total capacity of all V2I links;
[0018] The V2V link payload transmission success rate, dynamically calculated through the actual transmission capacity and the penalty mechanism;
[0019] The task priority weight, used to reflect the importance of the vehicle task;
[0020] The communication quality index, based on the channel strength and delay;
[0021] The global model optimization goal, associated with the global performance metrics of asynchronous federated learning; and the weights of the V2I link capacity and the V2V link success rate are balanced through dynamic weight parameters.
[0022] Furthermore, in step S2, the dynamic weight adjustment mechanism calculates the weights of each agent through the following formula:
[0023] w k = α·Q k + β·C k + γ·F k (1)
[0024] Where: Q k represents the model update quality of the vehicle, C k represents the communication quality of the vehicle, F k represents the model update frequency of the vehicle, and α, β, γ respectively represent weight adjustment coefficients for dynamically optimizing the global model aggregation effect.
[0025] Furthermore, in step S2, the multi-agent deep deterministic policy gradient algorithm adopts a centralized Critic network and a distributed Actor network architecture, where:
[0026] The centralized Critic network updates parameters by minimizing the Q-value error and evaluates the global effect of the joint actions of all agents;
[0027] The distributed Actor network generates action policies based on local observation states and optimizes local resource allocation by maximizing the cumulative reward;
[0028] Each agent stores interaction experiences through an experience replay buffer and updates network parameters in a mini-batch sampling manner.
[0029] Furthermore, the weight adjustment coefficients α, β, γ are dynamically adjusted according to network load, global interference level, and task requirements to optimize the adaptability of the global model to different communication environments and heterogeneous tasks.
[0030] Furthermore, in step S4, each V2V link accesses an orthogonal sub-band at the same time and preferentially reuses the idle sub-band resources of the V2I link to reduce resource conflicts and cross-link interference.
[0031] Advantages of the present invention:
[0032] Compared with the prior art, the vehicle-to-everything (V2X) communication resource allocation method based on federated multi-agent deep reinforcement learning of the present invention has the following technical features and advantages:
[0033] (1) Privacy protection and data security: Through the federated learning framework, vehicles only perform model training locally and upload encrypted model parameters without sharing original communication data, effectively avoiding the leakage of sensitive information and meeting the strict requirements of V2X for privacy protection.
[0034] (2) Dynamic environment adaptation ability: By integrating the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm, the agents can perceive the channel state, network topology changes, and task requirements in real time, and dynamically adjust the spectrum access, power control, and bandwidth allocation strategies, so as to maintain the efficiency and stability of resource allocation even in scenarios of high-speed vehicle movement and sudden interference.
[0035] (3) Communication efficiency and low overhead: Adopting the Asynchronous Federated Learning (AFL) mechanism, which supports the agents to upload model parameters on demand instead of synchronous updates, significantly reducing the load pressure on the communication network; at the same time, the dynamic weight adjustment mechanism reduces the transmission of redundant parameters by optimizing the model aggregation weights, further improving the communication efficiency.
[0036] (4) Balance between global optimization and local cooperation: By evaluating the global joint action effect through a centralized Critic network and combining with a distributed Actor network to execute local policies, an effective modeling of the cooperation and competition relationships among multiple agents is achieved. While ensuring the flexibility of local resource allocation, the overall spectrum efficiency of the system and the performance of V2I / V2V links are optimized.
[0037] (5) Efficient utilization of spectrum resources: The V2V link intelligently reuses the orthogonal sub-band resources of the V2I link, and selects the optimal sub-band through the dynamic spectrum access strategy to reduce cross-link interference; combined with power control and bandwidth allocation, the total capacity of the V2I link is maximized, while ensuring a high transmission success rate for the V2V link.
[0038] (6) Scalability and robustness: The distributed architecture avoids the single-point failure risk of centralized processing, and the model dimension is decoupled from the number of vehicles, supporting flexible expansion in large-scale vehicle networking scenarios; the dynamic weight adjustment mechanism alleviates the model convergence problem brought by the heterogeneity of vehicle data through adaptive parameter aggregation, improving the robustness of the algorithm in complex environments.
[0039] (7) Significant practical application value: It provides key technical support for 6G-V2X standardization, promotes the cost reduction and efficiency improvement of vehicle networking services, helps to achieve the goals of intelligent driving, vehicle-road coordination, and traffic safety, and at the same time indirectly promotes the achievement of the carbon neutrality goal by optimizing energy consumption and interference control.
[0040] Through the innovative integration of "federated learning + multi-agent deep reinforcement learning", the present invention solves the privacy-efficiency-real-time triangle contradiction faced by dynamic resource allocation in vehicle networking. On the premise of ensuring data security, it significantly improves the spectrum efficiency, transmission success rate, and system adaptability, and has broad application prospects and commercial value. Description of the Drawings
[0041] To more clearly illustrate the technical solutions of the embodiments of the present invention, the present invention will be described in detail below in conjunction with the drawings and detailed embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0042] Among them:
[0043] Figure 1 is the vehicle networking system model diagram of the present invention;
[0044] Figure 2 is the AFL-MADDPG framework diagram of the present invention;
[0045] Figure 3 is the impact diagram of the number of vehicles of the present invention on the network transmission capacity;
[0046] Figure 4 is the comparison diagram of the system spectral efficiency of each algorithm under different numbers of vehicles of the present invention. Specific Embodiments
[0047] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. The following combines the attached Figures 1-4 to further illustrate the vehicle networking communication resource allocation method based on federated multi-agent deep reinforcement learning.
[0048] Embodiment 1
[0049] Existing multi-agent reinforcement learning algorithms usually rely on local cooperation among agents and lack a global optimization perspective. In complex scenarios, it may be difficult to effectively coordinate the strategies among agents. In addition, information sharing among agents may involve sensitive data (such as location, channel state), posing a risk of privacy leakage. To address the above challenges, the present invention proposes a vehicle networking communication resource allocation method based on federated multi-agent deep reinforcement learning. Asynchronous federated learning realizes global experience sharing in a distributed learning framework by dynamically aggregating model parameters and uploading the model parameters with lower transmission costs to the global server, while protecting the privacy of vehicles. The multi-agent deep deterministic policy gradient (MADDPG) algorithm can effectively learn the cooperation strategies among vehicles and dynamically adjust the communication resource allocation to adapt to the real-time changes of the vehicle network by combining the advantages of a centralized Critic network and a distributed Actor network. In addition, the dynamic weight adjustment mechanism proposed by the present invention can optimize the parameter aggregation weights of the global model according to the changes in the communication environment of each vehicle, thereby improving the global and local performance.
[0050] The vehicle-to-everything (V2X) communication scenario considered in this invention consists of a single base station and multiple vehicle users. This scenario consists of M vehicle-to-infrastructure (V2I) links, K vehicle-to-vehicle (V2V) links, and some interference links. The sets of V2I links and V2V links are denoted as M = {1, 2, …, M} and K = {1, 2, …, K}, respectively. It is assumed that the orthogonal spectral sub-bands have been pre-allocated to the V2I links for uploading data, i.e., the m-th V2I link occupies the m-th sub-band. The transmit power of the V2I link is a fixed value, and the V2V links need to opportunistically access the V2I channels and use appropriate transmit powers. Therefore, the objective of this invention is to design a continuous-action-based spectrum access, power transmission, and bandwidth allocation scheme for V2V links to maximize the total capacity of the V2I links, improve the transmission success rate of the V2V links, reduce interference, and thus enhance the overall spectrum efficiency.
[0051] Assume that the channel fading is approximately the same within a sub-band and independent between different sub-bands. h k [m] is the power component of the small-scale fading, and it is assumed to follow an exponential distribution. α k characterizes the large-scale fading that is independent of frequency. Then, within a coherence time, the channel power gain of the k-th V2V link on the m-th sub-band is expressed as
[0052] g k [m] = α k h k [m] (2)
[0053] Express the signal-to-interference-plus-noise ratio (SINR) of the m-th V2I link and the k-th V2V link on the m-th channel as
[0054]
[0055] In the formula: and represent the transmit powers of the m-th V2I link and the k-th V2V link, respectively; σ 2 is the noise power; g k [m] represents the channel gain of the k-th V2V link on the m-th channel; g k,B [m] represents the interference channel gain from the k-th V2V link to the base station on the m-th channel; g k′,k [m] represents the interference channel gain from the k'-th V2V link to the k-th V2V link on the m-th channel; represents the channel gain of the m-th V2I link on the m-th channel, represents the interference channel gain from the m-th V2I link to the k-th V2V link on the m-th channel; the interference power Ik [m] is
[0056]
[0057] where: ρ k [m] is a spectrum selection variable of a Boolean value, which indicates whether the k-th V2V link transmits on the m-th subband. If so, ρ k [m] = 1, otherwise ρ k [m] = 0. It is assumed that each V2V link can access at most one orthogonal subband, that is
[0058] Then, according to Shannon's formula, the capacity of the m-th V2I link on the m-th subband can be expressed as
[0059]
[0060] Similarly, the calculation formula for the channel capacity of the k-th V2V link on the m-th subband is
[0061]
[0062] where: B m is the bandwidth of the m-th subband.
[0063] By accumulating the capacities of V2I and V2V links in all subbands, the total system capacity is obtained as
[0064]
[0065] The overall spectral efficiency of the system is defined as the ratio of the total capacity to the total bandwidth
[0066]
[0067] where: B total is the total system bandwidth.
[0068] As described above, the objective of the present invention is to meet the requirements of low-latency and highly reliable real-time data transmission of V2V links while increasing the total capacity of V2I links. For this purpose, the probability of successfully transmitting the payload within a certain time limit is defined herein as:
[0069]
[0070] where: L represents the size of the V2V link transmission payload generated in each period T, with the unit of bit; Δ T represents the channel coherence time.
[0071] In summary, the resource allocation problem studied in the present invention in the vehicle-to-everything (V2X) network can be described as follows: in the vehicle-to-vehicle (V2V) link, how to intelligently reuse the sub-channels of vehicle-to-infrastructure (V2I), and select an appropriate transmission power for data transmission to reduce resource conflicts and at the same time reduce its interference to the V2I link, that is, while pursuing the maximization of the total capacity of the V2I link to improve the payload successful transmission rate per unit time of the V2V link, thereby enhancing the overall spectral efficiency of the system shown in formula (8).
[0072] The present invention models the V2X communication resource allocation problem as a task based on multi-agent reinforcement learning, and adopts a scheme combining asynchronous federated learning (AFL) and multi-agent deep deterministic policy gradient (MADDPG) algorithm. Each vehicle is regarded as an agent, and in the process of interacting with the environment, the resource allocation strategy is optimized through reinforcement learning. Multiple agents jointly explore a dynamically changing network environment, and continuously adjust the resource allocation strategy by sharing global information and local experience to improve the performance of the entire V2X system. To solve the problem of privacy leakage in traditional methods, an asynchronous federated learning mechanism is adopted, so that the learning process of each vehicle is carried out locally, and only the model update parameters are aggregated with the global model, thus ensuring the protection of data privacy. To further improve the adaptability and resource allocation efficiency of the system, the present invention also introduces a dynamic weight adjustment mechanism, which flexibly adjusts the update strategy of the global model according to the real-time communication environment and task requirements of each vehicle. The entire algorithm design is divided into two stages: the local training stage and the global aggregation stage. In the local training stage, each agent independently trains, optimizes the strategy based on the local environment, updates the model and uploads it. In the global aggregation stage, the global server receives the model updates, weighted aggregates the models through the dynamic weight adjustment mechanism, and distributes the new global model. Next, the design idea of the V2X communication resource allocation algorithm based on the combination of AFL and MADDPG will be described in detail.
[0073] To meet the requirements of multi-agent cooperation, dynamic weight adjustment and global model optimization in the V2X network, the present invention designs a state space that includes local observation information, task requirement information, and global statistical information. Each vehicle, as an agent, can only observe the local communication environment information related to it, including: its own channel gain g k [m]: the channel quality of vehicle k on channel m; the interference of other V2V links g k,k' [m]: the interference of other V2V link vehicles to vehicle k on channel m; the interference of the V2I link to the vehicle g m,k [m]: the interference of the V2I link on channel m to vehicle k; the interference of the base station g k,B[m]: Interference of the base station's transmitted signal on vehicle k. In the vehicle-to-everything (V2X) network, the heterogeneity of vehicle task requirements has an important impact on the resource allocation strategy. Therefore, the following task-related information is incorporated into the state space: bandwidth requirement b k : Bandwidth requirement of vehicle k for the current task; task priority p k : Task priority of vehicle k, which determines the importance of the task. To better utilize the global model aggregation ability in asynchronous federated learning, global statistical information is introduced to reflect the overall network state, including: global network load L global : Overall load condition of the current network (such as channel occupancy rate); global interference level I global : Overall interference intensity of all communication links in the network. To support the dynamic weight adjustment mechanism, metrics that affect the global model optimization are further added to the state space: communication quality C k : Current communication quality of vehicle k (such as channel strength, latency, etc.); model update quality Q k : Contribution degree of the local model update of vehicle k to the global model; model update frequency F k : Frequency of model update of vehicle k.
[0074] Combining the above information, the state space of vehicle k at time t is designed as:
[0075]
[0076] In the vehicle-to-everything (V2X) communication resource allocation scheme proposed in the present invention, the design of the action space directly determines the decision-making ability of the agent (vehicle) in the multi-agent reinforcement learning framework. To adapt to the complexity of the dynamic environment and resource allocation requirements in the V2X network, the action space design of the present invention covers three core aspects: spectrum access, transmit power control, and bandwidth allocation.
[0077] The action of each agent (vehicle) k at time step t is defined as:
[0078]
[0079] where: f k : Represents the spectrum sub-band selected by the vehicle, which is a discrete variable. The vehicle can dynamically select the current optimal spectrum resource (for example, occupy unused spectrum or share the V2I frequency band).
[0080] represents the transmit power of vehicle k on the V2V link, which is a continuous variable, and its value range is:
[0081] represents the bandwidth allocated to the current task of vehicle k, which is a continuous variable, and the range is: Used to meet the mission requirements of the vehicle and ensure the communication performance of tasks of different priorities.
[0082] The vehicle determines the communication resource access strategy by selecting the spectrum sub-band. The goal of spectrum selection is to maximize the transmission success rate and reduce system interference. k The discrete selection range includes multiple predefined spectrum sub-bands (such as V2I sub-bands, idle sub-bands, etc.). The intelligent agent dynamically adjusts the transmission power based on the channel quality and interference situation to balance energy consumption and communication quality. is a continuous variable calculated by the agent's Actor network based on the current state. To adapt to the heterogeneity of multiple tasks, the agent can dynamically adjust bandwidth allocation. High-priority tasks will be allocated more bandwidth to ensure fairness and efficiency in resource allocation. is a continuous variable determined by the agent based on task requirements and global bandwidth constraints.
[0083] By simultaneously controlling spectrum access, transmit power, and bandwidth allocation, intelligent agents can optimize resource allocation strategies across multiple dimensions to adapt to complex IoV communication scenarios. Spectrum access and transmit power control are primarily based on local observations, while bandwidth allocation incorporates information such as task priorities provided by the global model, achieving a unified approach of local optimization and global collaboration. The action space design leverages MADDPG's ability to process continuous action spaces while also aligning with the distributed execution characteristics of multi-agent reinforcement learning. The action space allows for flexible adjustments to the agent's decision variables to improve resource allocation efficiency in the face of dynamic changes in vehicle density, task requirements, or interference intensity.
[0084] Reinforcement learning, when solving difficult optimization problems in complex, high-dimensional scenarios, focuses on the design of its reward function. The reward function in this paper aims to simultaneously optimize the total capacity of the V2I link and the success rate of payload transmission in the V2V link to improve overall system performance. Furthermore, to adapt to the dynamic IoV environment and the requirements of global model optimization, the reward function incorporates task priority, communication quality, and global performance-related metrics.
[0085] Combining the V2I link capacity, V2V link success rate, task priority, and global optimization goals, the agent’s reward function is designed as follows:
[0086]
[0087] Among them: V2I link capacity: It represents the total capacity of all V2I links and is used to measure the communication performance between vehicles and base stations.
[0088] V2V link load transmission success rate:
[0089]
[0090] When the payload is successfully delivered, the actual transmission capacity is used as a reward; when it fails, a constant penalty is imposed.
[0091] Task priority p k : Used to represent the importance of the current task of vehicle k. Higher-priority tasks receive higher weights.
[0092] Communication quality C k : Reflects the communication conditions of vehicle k (such as channel quality, latency, etc.), and is used to adjust its reward value.
[0093] Global model optimization objective Q global : Represents the performance metrics of the global model in the asynchronous federated learning framework (such as the magnitude of global loss reduction, global throughput, etc.).
[0094] Dynamic weight λ(t): A dynamically adjusted weight parameter used to balance the V2I link capacity and the V2V link success rate:
[0095]
[0096] By dynamically adjusting the weight parameter λ(t), the focus is adaptively optimized according to the current network state (such as the proportion of V2I link capacity and V2V link success rate). Introducing the task priority p k , ensuring that high-priority tasks can receive more guarantees in resource allocation. Through the global performance metric Q global , making the local optimization of the agent consistent with the global goal and improving the global convergence of the asynchronous federated learning framework. Combining link capacity, communication quality, and task requirements, the reward function can dynamically adapt to the highly variable network environment in the vehicle network.
[0097] The AFL-MADDPG method can technically achieve a balance of high resource utilization, low latency, and strong privacy protection, economically promote cost reduction and efficiency improvement of vehicle network services, and contribute to traffic safety and carbon neutrality goals at the social level. Its core value lies in solving the triangular contradiction of vehicle network dynamics, privacy, and scalability through the architectural innovation of "federated learning + multi-agent collaboration", providing key technical reserves for 6G-V2X standardization (3GPP R19 research direction).
[0098] Example 2
[0099] Asynchronous Federated Learning (AFL) is a distributed machine learning method that ensures data privacy while avoiding computational bottlenecks in centralized methods when training models among multiple agents (such as vehicles). Each vehicle trains the model locally using its own communication data and periodically uploads the updated model parameters to the global server for aggregation without the need to transmit the original data to a central location. Different from traditional synchronous federated learning, AFL allows agents to upload updates at different time points, which can reduce communication latency and improve computational efficiency.
[0100] In the Internet of Vehicles, asynchronous federated learning can effectively address vehicle mobility, network instability, and data privacy issues. For example, vehicles train resource allocation models based on local channel status, bandwidth requirements, etc., and pass the local model updates to the global server through AFL. The global model aggregates based on these updates to optimize communication resource allocation between vehicle-to-vehicle (V2V) and vehicle-to-infrastructure (V2I). In this way, each agent in the Internet of Vehicles can collaboratively learn the optimal communication resource allocation strategy while maintaining data privacy.
[0101] Multi-Agent Deep Deterministic Policy Gradient (MADDPG) is a reinforcement learning algorithm proposed for cooperation and competition problems in multi-agent environments. It is based on the Deterministic Policy Gradient (DDPG) method of deep reinforcement learning and enables multiple agents to cooperate in a shared environment through centralized training and distributed execution. MADDPG allows each agent to independently execute a policy (Actor network) and interact with the environment, while using a shared Critic network to evaluate the joint actions of all agents.
[0102] In the communication resource allocation problem in the Internet of Vehicles, each vehicle is regarded as an agent and needs to make decisions (such as choosing a spectrum, adjusting power, etc.) based on the current communication environment (such as channel quality, bandwidth, power requirements, etc.). The core principle of MADDPG lies in the use of the Actor-Critic architecture of each agent:
[0103] Policy-based Actor network: Selects the optimal action (such as power control, spectrum access) based on the current state.
[0104] Value-based Critic network: Evaluates the effect of the joint actions of all agents and calculates the Q value to guide policy updates.
[0105] MADDPG is optimized through centralized training and distributed execution. All agents share a global Q-function for training, but each agent executes its own policy, avoiding overly complex global optimization. In the vehicle networking scenario, vehicles collaborate through MADDPG to learn resource allocation strategies, optimizing the total capacity and success rate of V2I and V2V links.
[0106] Specifically, the optimization objective of MADDPG is to maximize the expected cumulative reward:
[0107]
[0108] where R i (t) is the reward obtained by agent i at time t, and γ is the discount factor. The Critic network evaluates the joint actions of all agents by minimizing the error, while the Actor network optimizes its respective policy through the gradient ascent method. In this way, agents can optimize resource allocation in the vehicle network based on global information and local decisions.
[0109] The communication resource allocation problem in the vehicle network is a highly dynamic multi-agent problem, involving multiple factors such as spectrum access, power control, and bandwidth allocation. Traditional centralized methods have computational bottlenecks and risks of privacy leakage. Especially in a large-scale distributed system like the vehicle network, the dynamics and data privacy of vehicles make centralized processing difficult to apply effectively. While a single multi-agent deep reinforcement learning (such as MADDPG) can optimize resource allocation strategies well, it still faces challenges in data privacy and communication efficiency.
[0110] To solve these problems, the present invention proposes a fusion scheme combining asynchronous federated learning (AFL) with the multi-agent deep deterministic policy gradient (MADDPG) algorithm. Through asynchronous federated learning, each vehicle (agent) can perform model training locally and upload parameter updates, thus ensuring data privacy and reducing communication overhead. At the same time, combined with the MADDPG algorithm, vehicles can optimize the allocation of communication resources in a multi-agent collaboration framework, fully considering the cooperation and competition relationships among vehicles in the vehicle network.
[0111] In addition, in asynchronous federated learning, due to different communication environments, task requirements, and model update qualities of different vehicles, simple average weighting may cause some agents to contribute too much or too little to the global model. To improve the performance and adaptability of the global model, we introduce a dynamic weight adjustment mechanism. This mechanism dynamically adjusts the weight of each agent in the aggregation of the global model according to factors such as the update quality, communication quality, and update frequency of each vehicle. Thereby improving the efficiency of resource allocation and the adaptability of the system.
[0112] This algorithm combines the advantages of asynchronous federated learning and MADDPG, and is divided into a local training stage and a global aggregation stage, and a dynamic weight adjustment mechanism is introduced. Its basic process is as follows:
[0113] Local training stage (vehicle side):
[0114] Each vehicle k obtains its local state at time step t Based on the current state The vehicle selects an action through its Actor network The vehicle performs communication operations according to the selected action and interacts with the environment (other vehicles, base stations). Observe the environmental feedback and obtain the reward and the next state Store the current interaction experience into the experience replay buffer. The vehicle locally trains its Actor network and Critic network using the data in the experience replay buffer:
[0115] Critic network update: Calculate the target value using the target network:
[0116]
[0117] Minimize the error to update the Critic network:
[0118]
[0119] Actor network update: Update the policy network parameters so that the actions output by the policy can maximize the Q value of the Critic network.
[0120] After the vehicle completes a round of local training, it uploads its model update (parameters θ k ) to the global server.
[0121] Global aggregation stage (global server side)
[0122] The global server asynchronously receives the local model update parameters θ uploaded by each vehicle k , and stores these updates. The server calculates the dynamic weights according to the communication quality, model update quality, and update frequency of each vehicle:
[0123] w k =α·Q k +β·C k +γ·F k (19)
[0124] Where: Q k : The model update quality of the vehicle. C k : The communication quality of the vehicle. F k: The model update frequency of the vehicle. α, β, γ: Weight adjustment coefficients used to balance the importance of each indicator.
[0125] The model update quality of the vehicle directly affects the effect of the global model. High-quality model updates (such as a smaller decrease in the loss function or higher accuracy) should be given higher weights. The communication quality of the vehicle (such as signal strength, bandwidth, latency, etc.) affects the frequency and timeliness of its model updates. Vehicles with better communication quality should make more contributions. In asynchronous updates, vehicles with a higher update frequency contribute more to the global model. Vehicles that update frequently should be given higher weights.
[0126] According to the dynamic weight w k Perform weighted averaging on the model parameters of all vehicles to update the global model:
[0127]
[0128] The updated global model is stored as the latest version and sent to all vehicles.
[0129] In the local training stage, each vehicle trains a reinforcement learning model based on the local observation state to optimize the resource allocation strategy while ensuring data privacy. In the global aggregation stage, the server integrates the model updates uploaded by the vehicles through a dynamic weight adjustment mechanism to optimize the global model and ensure the overall performance of the system. Through the cooperation of these two stages, the algorithm can achieve efficient resource allocation in a dynamic vehicle networking scenario while taking into account privacy protection and global performance optimization.
[0130] The present invention proposes a method for allocating communication resources in a vehicle networking by combining asynchronous federated learning (AFL) and multi-agent deep deterministic policy gradient (MADDPG) algorithms, and introduces a dynamic weight adjustment mechanism to solve the problem of high-dynamic communication resource allocation in vehicle networking. By combining AFL and MADDPG, the present invention can not only achieve efficient resource allocation on the premise of ensuring data privacy, but also utilize the cooperation and competition mechanisms of multi-agent reinforcement learning to optimize communication performance in a dynamic network environment. The simulation experiment results show that the proposed method has shown significant advantages in multiple experimental scenarios. Compared with traditional centralized methods and the single MADDPG algorithm, the scheme combined with AFL has achieved better performance in terms of spectral efficiency, total communication capacity, transmission success rate, system convergence speed, etc. Especially after introducing the dynamic weight adjustment mechanism, the algorithm can dynamically optimize the update of the global model according to the communication environment and task requirements of each vehicle, improving the adaptability and resource allocation efficiency of the system. In addition, the introduction of asynchronous federated learning not only effectively guarantees privacy protection, but also significantly reduces communication overhead and improves the spectral efficiency of the system.
[0131] The pseudo-code of the vehicle-to-everything (V2X) communication resource allocation algorithm based on AFL-MADDPG is as follows:
[0132]
[0133]
[0134] As mentioned above, the above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A vehicle-to-everything (V2X) communication resource allocation method based on federated multi-agent deep reinforcement learning, characterized in that The steps are as follows: S1. In the vehicle-to-everything (V2X) scenario, each vehicle is modeled as an agent and locally trained using the multi-agent deep deterministic policy gradient (MADDPG) algorithm. Each agent selects actions based on the local observation state and performs spectrum access, transmit power control, and bandwidth allocation. S2. An asynchronous federated learning method is adopted. Each agent uploads its local model parameters to the global server, and the model parameters are weighted and aggregated through a dynamic weight adjustment mechanism to generate a global model. S3. The dynamic weight adjustment mechanism calculates weights based on the model update quality, communication quality, and update frequency of the agents to optimize the aggregation of the global model. S4. The updated global model is sent to each agent to achieve distributed privacy protection and dynamic optimization of V2X communication resources.
2. The method for allocating communication resources in a vehicle networking as claimed in claim 1, wherein In step S1, the local observation state includes: the vehicle's own channel gain, interference from other vehicle-to-vehicle (V2V) links, interference from vehicle-to-infrastructure (V2I) links to the vehicle, interference from the base station, the bandwidth requirement of the vehicle's current task, task priority, global network load, global interference level, communication quality, model update quality, and model update frequency.
3. The method for allocating communication resources in a vehicle network based on federated multi-agent deep reinforcement learning according to claim 1, characterized in that In step S1, the action space is designed to include the following variables: A discrete spectrum subband selection variable for dynamically accessing orthogonal subbands or idle subbands of the V2I link. A continuous transmit power control variable with a value range from the predefined minimum transmit power to the maximum transmit power. A continuous bandwidth allocation variable with a value range from the minimum bandwidth requirement to the maximum available bandwidth to adapt to the resource requirements of different priority tasks.
4. The method for allocating communication resources in a vehicle network based on federated multi-agent deep reinforcement learning according to claim 1, characterized in that In step S1, the reward function of the MADDPG algorithm consists of the following parameters: The total capacity of all V2I links. The V2V link payload transmission success rate, dynamically calculated through the actual transmission capacity and a penalty mechanism. The task priority weight, used to reflect the importance of the vehicle's task. A communication quality metric based on channel strength and latency. The global model optimization objective, associated with the global performance metric of asynchronous federated learning; and the weights of the V2I link capacity and V2V link success rate are balanced by dynamically adjusting the weight parameters.
5. The method for allocating communication resources in a vehicle network based on federated multi-agent deep reinforcement learning according to claim 1, wherein In step S2, the dynamic weight adjustment mechanism calculates the weights of each agent through the following formula: w k = α·Q k + β·C k + γ·F k (18) Where: Q k represents the model update quality of the vehicle, C k represents the communication quality of the vehicle, F k represents the model update frequency of the vehicle, and α, β, γ respectively represent weight adjustment coefficients for dynamically optimizing the global model aggregation effect.
6. The method for allocating communication resources in a vehicle network based on federated multi-agent deep reinforcement learning according to claim 1, characterized in that, In step S2, the MADDPG algorithm adopts a centralized Critic network and a distributed Actor network architecture, where: The centralized Critic network updates parameters by minimizing the Q-value error and evaluates the global effect of the joint actions of all agents. The distributed Actor network generates action policies based on the local observation state and optimizes local resource allocation by maximizing the cumulative reward. Each agent stores interaction experiences in an experience replay buffer and updates network parameters in a mini-batch sampling manner.
7. The method for allocating communication resources in a vehicle networking as claimed in claim 5, wherein, The weight adjustment coefficients α, β, γ are dynamically adjusted according to the network load, global interference level, and task requirements to optimize the adaptability of the global model to different communication environments and heterogeneous tasks.
8. The method for allocating communication resources in a vehicle network based on federated multi-agent deep reinforcement learning according to claim 1, characterized in that In step S4, each V2V link accesses an orthogonal subband at the same time and preferentially reuses the idle subband resources of the V2I link to reduce resource conflicts and cross-link interference.
Citation Information
Cited By
Federal reinforcement learning task allocation method and system for heterogeneous unmanned aerial vehicle cluster
CN120722931A
A Federated Reinforcement Learning Task Allocation Method and System for Heterogeneous Unmanned Aerial Vehicle Swarms
CN120722931B
Distributed scheduling method for whole vehicle manufacturing stamping resources under cloud side end cooperation
CN120952497A
Multi-network multi-granularity resource optimization method based on federal deep reinforcement learning
CN121099375A
Multi-network-connected hybrid electric vehicle heat energy comprehensive management method based on federal reinforcement learning
CN121256954A