Internet of vehicles spectrum sharing method based on multi-agent deep reinforcement learning

By constructing a spectrum sharing method based on multi-agent deep reinforcement learning, the problem of low spectrum resource utilization efficiency and difficulty in balancing communication needs in vehicle-to-everything (V2X) is solved. Joint optimization of spectrum and power is achieved, improving the system's communication performance and stability. This method is applicable to intelligent vehicle communication systems in urban roads and autonomous driving scenarios.

CN121013085APending Publication Date: 2025-11-25CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511187360.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing vehicular network spectrum management methods are inefficient in resource utilization under high-dynamic and high-density traffic environments, struggle to meet both V2V and V2I communication needs, lack stability in optimization strategies, and present challenges in joint spectrum and power regulation and multi-objective performance trade-offs.

Method used

A spectrum sharing method based on multi-agent deep reinforcement learning is constructed. It adopts a centralized training-distributed execution architecture and achieves joint optimization of subband selection and transmit power through joint state modeling, action control and reward design. Combined with the near-end policy optimization algorithm and generalized advantage estimation, the policy stability and convergence speed are improved.

Benefits of technology

It enables joint optimization decisions of spectrum and power without relying on global information, improves spectrum utilization, balances V2V communication reliability and V2I link throughput performance, reduces system energy consumption and scheduling complexity, and adapts to complex and ever-changing traffic and channel environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121013085A_ABST
    Figure CN121013085A_ABST
Patent Text Reader

Abstract

The invention relates to the field of wireless communication and resource optimization of the Internet of Vehicles, and discloses a spectrum sharing method based on multi-agent deep reinforcement learning, which realizes combined control of sub-bands and power by constructing each V2V link as a collaborative learning environment of an agent and training a strategy network by adopting a near-end strategy optimization algorithm. The spectrum utilization rate and the communication success rate are improved. According to the method, under a centralized training and distributed execution framework, the V2I communication requirement is considered while the V2V communication performance is optimized through joint design of the state, the action and the reward function, and the overall service quality of the system is remarkably improved. The method has good convergence stability and strategy generalization ability, is suitable for dynamically changing and densely interfering vehicle networking scenes, has the advantages of low communication overhead, high real-time performance, strong expansibility and the like, and can be widely applied to vehicle-mounted communication resource management tasks in intelligent traffic environments such as urban roads and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of wireless communication resource management technology in the Internet of Vehicles (IoV), specifically relating to an IoV spectrum sharing method based on multi-agent deep reinforcement learning. Background Technology

[0002] With the rapid development of intelligent connected vehicles, the Internet of Vehicles (IoV) has become a crucial infrastructure for next-generation intelligent transportation systems. IoV enables wireless communication between vehicles (V2V) and between vehicles and infrastructure (V2I), building a dynamic and collaborative information exchange system to support key applications such as autonomous driving, path planning, and collaborative perception. This is of great significance for improving traffic safety and efficiency. However, IoV communication heavily relies on limited wireless spectrum resources. Faced with ever-increasing communication demands, the scarcity of spectrum resources has become a core bottleneck restricting system scalability and service quality.

[0003] In existing technologies, traditional spectrum resource allocation strategies mainly rely on centralized optimization or rule-based static configuration, which are difficult to adapt to topology changes and dynamic link evolution in high-speed mobile environments. Furthermore, V2I and V2V communications have significantly different quality of service requirements: the former typically targets high-bandwidth, high-throughput in-vehicle entertainment and information services, while the latter emphasizes low latency and high reliability to ensure the real-time transmission of security messages. This heterogeneous communication demand exacerbates the complexity of resource allocation and the challenges of dynamic scheduling.

[0004] While existing research has attempted to incorporate game theory, graph theory, and heuristic algorithms for spectrum sharing modeling, these methods often rely heavily on environmental information, struggle to adapt to non-stationary multi-agent environments, and are prone to getting trapped in local optima in high-dimensional state spaces, thus limiting system performance. Deep Reinforcement Learning (DRL) has demonstrated excellent generalization and approximation capabilities in large-scale dynamic optimization tasks in recent years and is considered an emerging solution for vehicular network spectrum management. In particular, the Multi-Agent Reinforcement Learning (MARL) architecture can achieve autonomous decision-making by each agent relying only on local state observations, possessing distributed scalability and strong environmental adaptability.

[0005] However, existing deep reinforcement learning-based methods for application in vehicle-to-everything (V2X) networks still suffer from problems such as unstable policy training, insufficient inter-agent collaboration, and simplistic reward function design. Challenges remain, particularly in joint spectrum and power regulation, multi-objective performance trade-offs, and model deployment overhead control. Therefore, there is an urgent need to design a deep reinforcement learning spectrum sharing method that is suitable for dynamic V2X environments, possesses centralized training and distributed execution capabilities, and integrates high-dimensional state inputs and joint decision outputs to improve spectrum utilization efficiency and system communication performance in real-world traffic scenarios. Summary of the Invention

[0006] To address the problems of low resource utilization efficiency, difficulty in balancing V2V and V2I communication needs, and insufficient stability of optimization strategies in existing vehicle-to-everything (V2X) spectrum management methods under high-dynamic and high-density traffic environments, this invention proposes a spectrum sharing method based on multi-agent deep reinforcement learning. It constructs an agent collaborative learning mechanism under a centralized training-distributed execution architecture, and achieves joint optimization of subband selection and transmit power through joint state modeling, action control, and reward design. This improves spectrum utilization while simultaneously ensuring the reliability of V2V communication and the throughput performance of V2I links, thereby enhancing the overall service quality of the system.

[0007] The system model constructed in this invention is based on a cellular network architecture, assuming that each V2I link occupies a dedicated orthogonal sub-band, and each V2V link can share these spectrum resources under the premise of satisfying interference constraints. Each V2V link is modeled as an agent, whose local state includes information such as current channel gain, received interference power, remaining transmission load, and time; its actions are composed of a spectrum sub-band number and a transmit power level; its reward function integrates the V2V communication success rate and the V2I capacity index, adjusting the optimal trade-off between the two types of performance through a weighted mechanism. During training, a multi-agent near-end policy optimization algorithm is used to construct an Actor-Critic structure, using generalized advantage estimation (GAE) to improve the stability and convergence speed of policy updates, and using probability ratio pruning and entropy regularization techniques to enhance policy exploration capabilities and avoid getting trapped in local suboptimal conditions.

[0008] During the execution phase, each V2V link agent autonomously makes decisions based on its local state without central coordination, outputting spectrum access actions and power control strategies to ensure efficient resource sharing and interference management with limited communication overhead. This invention not only possesses good scalability and generalization capabilities but also adapts to complex and ever-changing traffic and channel environments, supports collaborative optimization of heterogeneous communication tasks, and reduces system energy consumption and scheduling complexity while ensuring communication service quality. It is particularly suitable for deploying intelligent vehicle communication systems in urban roads, highways, and autonomous driving scenarios.

[0009] This invention offers the following advantages: By introducing a multi-agent deep reinforcement learning framework, it enables joint optimization decisions on spectrum and power without relying on global information, significantly reducing communication and computational overhead compared to traditional centralized optimization methods; by employing a near-end policy optimization algorithm combined with generalized advantage estimation and policy pruning mechanisms, it effectively improves the stability and convergence speed of policy training, enhancing policy robustness in dynamic traffic environments; in terms of communication performance, the proposed method can simultaneously improve V2I link throughput and V2V task transmission success rate, achieving system-level optimization under heterogeneous communication requirements.

[0010] Furthermore, this method is applicable to complex vehicle-to-everything (V2X) scenarios such as high density, multiple disturbances, and time-varying channels. It possesses excellent real-time performance, scalability, and deployment flexibility. It not only meets the core requirements of spectrum sharing in intelligent connected vehicles but can also be extended to other resource allocation tasks in ubiquitous mobile environments, demonstrating high engineering application value and broad prospects for promotion. Attached Figure Description

[0011] Figure 1 This is a vehicle-to-everything (V2X) communication scenario system model for V2X spectrum sharing in an embodiment of the present invention;

[0012] Figure 2 This is a partial observable Markov decision process diagram in an embodiment of the present invention;

[0013] Figure 3 This is a flowchart illustrating the steps of an embodiment of the present invention;

[0014] Figure 4 This is a cumulative reward graph for agent training in an embodiment of the present invention;

[0015] Figure 5 This is a graph showing the variation of total V2I user throughput with load size in an embodiment of the present invention;

[0016] Figure 6 This is a graph showing the change in V2V user success rate with load size in an embodiment of the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] This invention provides a spectrum sharing method for vehicle-to-everything (V2X) networks based on multi-agent deep reinforcement learning. In V2X communication systems that include vehicle-to-infrastructure (V2I) links and vehicle-to-vehicle (V2V) links, it achieves efficient spectrum sharing and power control by constructing a collaborative learning environment and employing a near-end policy optimization algorithm. Specifically, it includes the following steps:

[0019] Initialize the spectrum sharing system environment and establish a vehicle network model that includes multiple V2I and V2V communication links;

[0020] Design a joint optimization objective function and construct a reward mechanism that balances V2V task success rate and V2I link throughput.

[0021] Under the centralized training and distributed execution framework, the policy network of the agents is trained using the multi-agent proximal policy optimization algorithm;

[0022] During the execution phase, each V2V agent independently selects spectrum resources and transmit power based on its current local state information, ultimately achieving efficient spectrum sharing and ensuring system service quality.

[0023] In the embodiments provided by this invention, considering resource allocation in the context of IOV (Internet of Vehicles) systems, a system model for Internet of Vehicles communication scenarios is constructed, such as... Figure 1 As shown, the vehicle-to-everything (V2X) spectrum sharing system model constructed in this embodiment of the invention includes a roadside base station (BS), several V2I communication links, and multiple V2V communication links. The BS has the ability to allocate fixed spectrum resources to ensure the quality of V2I communication; the V2V links reuse spectrum resources under interference constraints to complete the rapid transmission of local information.

[0024] To achieve optimal joint allocation of spectrum resources and power levels by intelligent agents, this invention constructs a multi-agent interactive learning environment and models the local observation and policy learning behavior of each V2V link using POMDP, such as... Figure 2 As shown.

[0025] In this example, the following is provided: Figure 3 The diagram shows a flowchart of a multi-agent deep reinforcement learning-based spectrum sharing method for vehicle-to-everything (V2X) networks. The system initialization and environment construction phase includes the following general steps:

[0026] S1. Initialize the spectrum sharing system environment and establish a vehicle-to-everything (V2I) and V2V communication model containing multiple V2I and V2V communication links. By constructing the communication topology, setting the spectrum resource pool and environmental parameters, and introducing a partially observable Markov decision process (POMDP) ​​to model the V2V links in the system, a multi-agent interaction environment is established, laying the foundation for subsequent reinforcement learning training.

[0027] S2. Design a joint optimization objective function and construct a reward mechanism that balances V2V task transmission success rate and V2I link throughput. Through dynamic evaluation of V2I and V2V link performance, design a local reward function suitable for multi-objective scenarios and formulate policy evaluation criteria based on system QoS requirements.

[0028] S3. In a multi-agent architecture with centralized training and distributed execution, the policy network of each V2V agent is trained using a proximal policy optimization algorithm. During training, an Actor-Critic structure, pruning update, and generalized advantage estimation mechanism are employed to improve the learning stability and convergence performance of the multi-agent system in non-stationary environments.

[0029] S4. After entering the execution phase, each V2V agent independently selects spectrum resources and transmit power based on the current local state information, and completes low-latency, high-response collaborative resource scheduling without the need for global information interaction, thereby achieving joint assurance of spectrum reuse and quality of service.

[0030] like Figure 1 As shown, the spectrum sharing system environment described in this embodiment includes a central base station (BS), several V2I communication links, and multiple V2V communication links. The BS allocates a portion of the spectrum sub-band resources in an orthogonal manner to ensure data transmission on the V2I links; the V2V links reuse spectrum resources under controlled interference conditions to achieve efficient information exchange between multiple vehicles.

[0031] The goal of this phase is to complete communication topology modeling, spectrum resource initialization, physical channel modeling, and the construction of the reinforcement learning environment. Specifically, this includes the following steps:

[0032] S101. Set the total number of vehicles to M, including N V2I links, with the V2I set M = {1,2,...,m} and the V2V set N = {1,2,...,n}. Each link consists of a terminal vehicle and a base station, occupying one subband resource. Simultaneously, construct multiple V2V communication pairs, each consisting of a transmitting vehicle and a receiving vehicle, forming a candidate set for multiplexing subband resources. All V2V links can share spectrum in a non-orthogonal manner, but are subject to the interference threshold constraints of the V2I links.

[0033] S102. Orthogonal Frequency Division Multiplexing (OFDM) technology achieves parallel transmission by decomposing a frequency-selectively fading wireless channel into multiple flat-fading subcarriers. Adjacent subcarriers are divided into spectrum sub-bands, assuming that the channel fading characteristics within each sub-band are basically the same and that the sub-bands are independent of each other. The selectable transmit power set for the V2V link is set to [23, 15, 5, -100] dBm, where -100 dBm represents the silent transmission state. The system also sets the maximum allowable transmit power P. maxSubband bandwidth W, system noise power σ 2 Low-level communication parameters, etc.

[0034] S103. The channel gain of each communication link includes the large-scale path loss a. n Log-shading fading and Rayleigh fast decay n [m]. During the coherence time, the channel power gain g of the nth V2V link corresponding to the mth subband (corresponding to the mth V2I link). n [m] = a n h n [m], where h n The [m] power component belongs to small-scale fading and is assumed to follow an exponential distribution, a n This represents large-scale fading independent of frequency. Similarly, the interference channel from the transmitter of the k-th V2V link to the receiver of the n-th V2V link in subband m is denoted by g. k,n [m], the interference channel from the transmitter of the nth V2V link to the base station in subband m is g. n,B [m], the interference channel in subband m from the transmitter of the m-th V2I link to the receiver of the n-th V2V link is...

[0035] S104 and V2I links are primarily used for high-data-rate entertainment services, therefore their design goal is to maximize total capacity. V2V links are primarily used for periodic secure information transmission, and their rate is modeled as the probability of successfully transmitting the payload of a packet of size B within time T. For example... Figure 2 As shown, this embodiment models each V2V link as a reinforcement learning agent, forming a multi-agent system. Each agent's observation space consists of its accessible local states, including current channel gain, receiver interference power, remaining data volume, and remaining time; it cannot observe the states of other agents. The system satisfies a partially observable Markov decision process (POMDP) ​​structure, supporting each agent in learning resource allocation behavior through a policy network.

[0036] To achieve coordinated optimization of spectrum resources and power control, it is necessary to model the observation state, available actions, and optimization objectives of each agent. This invention designs a distributed sensing mechanism based on local states without relying on global information, combined with a joint reward function to coordinate different types of communication needs. Specifically, it includes the following steps:

[0037] S201, such as Figure 2 As shown, this invention models each V2V link as an independent agent based on the POMDP structure, and constructs a state space considering its observable information. The state space includes local CSI, received interference power, remaining V2V payload, and transmission time. Status information is obtained through link measurement and local cache statistics, without relying on information from the central controller or other intelligent agents, thus satisfying the conditions for distributed execution.

[0038] S202. In this invention, the V2V link can only select its transmit power from four discrete values: [23, 15, 5, -100] dBm. This can also be extended to a continuous space. Here, each action is a combination of spectrum and transmit power selection, with a dimension of M×4. By constructing a Cartesian product of the action space, each agent has a finite set of discrete actions of a fixed dimension, facilitating output modeling using a policy network.

[0039] S203. This invention simultaneously optimizes two types of communication objectives: the total rate of V2I links and the transmission success rate of V2V links. To this end, the reward function comprehensively balances V2I link capacity and V2V load transmission success rate, achieving synergistic optimization of both. Where, λ l and λ d This represents the weights used to balance the transmission rate targets for V2I and V2V users. Through this modeling, the agent possesses the ability to make joint resource control decisions based on local observation states and can learn stable long-term policies under a centralized training architecture.

[0040] This invention employs a centralized training and distributed execution framework to achieve efficient learning of multi-agent collaborative strategies. The training phase is centrally conducted on a central server, allowing all agents to share environmental information and global reward signals. During the execution phase, each V2V link agent independently executes its strategy based on local observations, meeting the requirements of low communication overhead and practical deployment. The training framework is an extension of single-agent proximal policy optimization to a multi-agent proximal policy optimization structure. Each agent consists of a policy network (Actor) and a value function network (Critic), specifically including the following steps:

[0041] S301. In this embodiment, each agent is modeled using two independent neural networks: the policy network takes the current state s as input. i (t), the probability distribution of the output in the action space Value function networks, given the same input state, output state value estimates.

[0042] S302. Each agent has an Actor-Critic network simulating a V2V link. The agent optimizes its final policy by maximizing the cumulative reward. The cumulative reward function is defined as G. t The state-value function and state-action-value function in the evaluation network of the nth agent are respectively and If a certain state-action pair G tTo generate a positive impact, the probability of its occurrence should be increased; conversely, to generate a negative impact, its probability should be decreased. The advantage function of the nth agent is defined as: It measures how well or poorly a current action is performed relative to the average level.

[0043] S303. The GAE method is used to estimate the strategy advantage function, and the time difference error and long-term discounted return are combined to alleviate the variance fluctuation problem in strategy iteration. Generalized Advantage Estimation (GAE) function λ∈[0,1] is a hyperparameter used to balance the trade-off between variance and bias in the estimation. Based on this, the update strategy of the Actor network can be obtained. It can be obtained The pruning function represents the probability ratio of importance sampling, limiting the probability ratio of the new and old policies to the interval (1-ε, 1+ε), thereby suppressing training instability caused by excessively large policy update magnitudes. The accuracy of the actions selected by the Actor network largely depends on the Critic network's evaluation of the state value. Temporal difference error is defined as... This error can be used to guide the parameter updates of the Critic network, which is then updated to... ω* represents the parameters of the k-th agent Critic after network update, where ω is the parameter before update, and lr c It is the learning rate. This update method minimizes the temporal difference error through gradient descent, thereby improving the estimation accuracy of the state-value function.

[0044] S304. All agents interact in parallel within the simulated environment, collecting trajectory data, including state, action, reward, and next state. The server uniformly performs advantage estimation, gradient backpropagation, and network updates, and periodically synchronizes network parameters to each agent. In this embodiment, all V2V link agents share the same set of network parameters, ensuring policy generalization and training efficiency. The system obtains a converged policy network with good spectrum reuse coordination capabilities and cross-link robustness, providing assurance for distributed policy execution in the actual deployment phase.

[0045] All V2V agents have completed centralized training in the previous stage, and the policy network parameters have converged and been uploaded to each vehicle node or edge computing platform. This stage does not rely on central coordination or global information; resource allocation decisions are made entirely based on local observations, achieving distributed spectrum sharing with low communication overhead. The core processes of the execution phase include the following steps:

[0046] S401. Each V2V communication pair loads its trained policy network parameters, including the Actor network and necessary normalization factors, into its local cache. The system resets the environment state, clearing any residual task and link interference information from the previous cycle.

[0047] S402. At the beginning of each scheduling cycle, the V2V transmitting vehicle constructs a current local state vector based on the current subband channel gain, received interference power, remaining data volume, and remaining transmission time, which serves as the input to the policy network.

[0048] S403. Each agent inputs its state into its local policy network, outputs an action probability distribution, and selects a spectrum sub-band number and transmission power according to the maximum probability (or sampling) to form a joint action. The action output does not need to be negotiated with other agents, thus meeting the requirements of distributed operation.

[0049] S404: The system schedules all V2V links to transmit in parallel according to their selected actions, calculating the received SINR, transmission rate, and successful data transmission volume for each link pair. If a task is completed or times out within the current period, the system records the task status and resets the link load. The V2I links simultaneously monitor whether interference caused by V2V multiplexing exceeds a threshold.

[0050] S405. The system enters the next scheduling cycle, and each agent repeats steps S401 to S404 until all tasks are completed or the global time limit is reached. In long-term operation, the transmission success rate of each link, the total system throughput, and the V2I service quality can be periodically evaluated to support subsequent system adaptation and model updates.

[0051] To verify the effectiveness of the proposed multi-agent deep reinforcement learning-based spectrum sharing method for vehicle-to-everything (V2X) communication, system-level performance testing was conducted using simulation. The simulation environment in this embodiment was constructed based on the urban macrocell scenario defined in Annex A of 3GPP TR 36.885. The environment comprehensively considers key factors such as vehicle speed, channel fading characteristics, road structure, traffic density, and V2V communication service load, effectively simulating spectrum resource competition and interference in real-world traffic communication scenarios.

[0052] In this embodiment, the simulation platform is implemented using Python 3.8, and the reinforcement learning training framework uses PyTorch 2.4.1. During the agent's policy learning process, a policy network and a value network are constructed for each V2V link. Both types of networks adopt a fully connected neural network structure, containing three hidden layers with neuron sizes of 500, 250, and 120 respectively. ReLU (Rectified Linear Unit) activation functions are used between layers to enhance nonlinear modeling capabilities, and the Adam algorithm is used to optimize the network parameters.

[0053] Table 1 Environmental Parameters

[0054]

[0055] Table 1 shows the system simulation parameter settings used in this embodiment, covering underlying communication parameters such as spectrum allocation, path loss model, subband bandwidth, and maximum transmit power. The algorithm tuning parameters used during training include discount factors, GAE smoothing factors, pruning thresholds, number of training epochs, and batch size, as detailed in Table 2.

[0056] Table 2. Optimization of Hyperparameters

[0057]

[0058] exist Figure 4 The figure illustrates the cumulative reward per training set as the number of training iterations increases, reflecting the convergence behavior of the proposed multi-agent reinforcement learning method. The figure shows that the cumulative reward per set gradually increases as the training progresses, validating the effectiveness of the proposed training algorithm. The spectrum sharing strategy proposed in this invention exhibits excellent system performance in various simulation environments. For example, under dense traffic conditions, agents can adaptively allocate spectrum and control power, maintaining a V2V task completion rate above 95% and V2I link throughput fluctuations within 5%. Compared with traditional greedy selection or centralized scheduling methods, this method possesses higher policy stability and collaborative efficiency. Furthermore, the convergence rounds during training remain below 3000, enabling rapid deployment and making it suitable for low-cost implementation on RSUs or edge nodes.

[0059] like Figure 5 The figure shows the performance curves of the total throughput of the V2I link as a function of the V2V payload size B. In this embodiment, multiple V2V communication pairs are set up, and their initial task data volume is gradually increased to simulate the dynamic situation of load growth. Simulation results show that as the V2V link load increases, all strategies in the system exhibit a decreasing trend in V2I throughput. This phenomenon is mainly due to the fact that a larger V2V task volume leads to a longer continuous transmission time and an increased transmission power, thereby interfering with the V2I link over a longer period of time, affecting its signal-to-noise ratio and reducing capacity.

[0060] The distributed spectrum sharing method based on multi-agent reinforcement learning proposed in this invention maintains stable performance under various load levels, with overall throughput significantly outperforming traditional strategies and approaching the optimal performance of centralized maxV2V methods. Notably, despite using a fixed task load (2×1060 bytes) during the training phase, this method still maintains good policy response in different load tests, demonstrating good generalization ability and environmental robustness.

[0061] like Figure 6The figure shows the trend of V2V link task transmission success rate as the payload size B changes. As the task data volume increases, all algorithms show a gradual decrease in success rate. This phenomenon is mainly attributed to the fact that large data packets require more time to transmit, making them more likely to fail before the task deadline and thus be judged as failed tasks. Nevertheless, the centralized maxV2V strategy maintained a 100% success rate throughout the entire test range. This is primarily achieved by increasing the transmission rate to quickly complete data tasks and avoid subsequent interference to the system.

[0062] The distributed strategy proposed in this invention significantly outperforms traditional Q-learning and greedy strategies in terms of V2V task success rate, and maintains stable performance similar to the centralized strategy under various load intensities. These results further validate the feasibility and robustness of this method in task-driven dynamic environments, making it suitable for real-time spectrum scheduling in typical V2X scenarios.

[0063] This invention addresses the problems of spectrum resource conflicts in heterogeneous links, diversified communication requirements, and difficulties in collaborative optimization in vehicle-to-everything (V2X) networks by designing a spectrum sharing method based on multi-agent deep reinforcement learning. First, a collaborative shared communication system model incorporating V2I and V2V communication links is constructed. A multi-agent interaction environment is designed, incorporating inter-vehicle interference constraints, task-driven requirements, and spectrum allocation strategies. Then, each V2V communication link is modeled as an independent agent, and a partially observable Markov decision process (POMDP) ​​is introduced to describe its local state space. A joint subband selection and power control action structure is designed, enabling agents to make independent decisions under distributed conditions. Next, a centralized training-distributed execution mechanism based on a multi-agent proximal optimization strategy is proposed, combining a pruning strategy and a generalized advantage estimation method to improve policy stability and convergence speed in non-stationary environments. Furthermore, a unified reward mechanism balances V2V task completion rate and V2I link throughput performance, guiding agents to form collaborative behaviors during spectrum reuse. Finally, a simulation platform is built based on the TR 36.885 standard, and comparative experiments and system evaluations are completed, verifying the comprehensive performance advantages of this invention in terms of spectrum utilization efficiency, communication success rate, and adaptability.

[0064] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for spectrum sharing in vehicle-to-everything (V2X) networks based on multi-agent deep reinforcement learning, characterized in that, Includes the following steps: S1: Initialize the spectrum sharing system environment, establish a vehicle network model containing multiple V2I and V2V communication links, model each V2V communication pair as an independent intelligent agent, and define the state space including sub-band channel gain, interference power, remaining task load and time. S2: Design a joint optimization objective function, construct a reward mechanism that balances V2V task success rate and V2I link throughput, and define an action space that includes a combination of available subband number and power level to support joint control decisions of spectrum and power. S3: Under the centralized training and distributed execution framework, the multi-agent proximal policy optimization algorithm is used to train the agent's policy network, and mechanisms such as generalized advantage estimation and policy pruning are introduced to improve learning efficiency and stability. S4: During the execution phase, each V2V agent independently selects spectrum resources and transmit power based on the current local state information, ultimately achieving efficient spectrum sharing and system service quality assurance.

2. The method according to claim 1, characterized in that, The channel modeling described above assumes that V2I users require large link capacity to support infotainment applications, while V2V users require high-reliability links to transmit secure information. The V2I link set is M = {1, 2, ..., m}, and the V2V link set is N = {1, 2, ..., n}. It is also assumed that the M V2I links have been pre-allocated with M orthogonal sub-channels with fixed transmit power, establishing a strict one-to-one mapping relationship. That is, the m-th V2I link occupies the m-th sub-channel, and there is no interference between the sub-channels. For the purpose of improving spectrum utilization, these sub-channels can be reused by V2V links. Orthogonal Frequency Division Multiplexing (OFDM) technology achieves parallel transmission by decomposing a frequency-selectively fading wireless channel into multiple flat-fading subcarriers. Adjacent subcarriers are divided into spectrum sub-bands. It is assumed that the channel fading characteristics within each sub-band are essentially the same and that the sub-bands are independent of each other. Whether multiplexing is used is represented by a Boolean variable when calculating link capacity. Within a coherent time period, the channel power gain g of the nth V2V link on the mth sub-band (occupied by the mth V2I link) is... n [m], the interference channel from the transmitter of the k-th V2V link through the m-th subband to the receiver of the n-th V2V link is represented as g. k,n [m], the interference channel from the transmitter of n V2V links to the base station in subband m is represented as g. n,B [m], the interference channel from the transmitters of m V2I links to the receivers of n V2V links in subband m is represented as:

3. The method according to claim 1, the capacity of the m-th V2I link and the n-th V2V link Defined by the following formula: in Let n be the transmission power of the V2V transmitter. Let σ be the transmit power of the m-th V2I transmitter, and both transmit powers are on the m-th subband. 2 It is noise power, ρ n [m] is a Boolean value, indicating whether the nth V2V link reuses the mth V2I link, W is the channel bandwidth, and the interference power is expressed as:

4. The method according to claim 1, characterized in that, The spectrum resource management problem can be modeled as the following optimization objective: C3:∑mρ n [m]≤1,m∈M,n∈N C6: 0 ≤ T ≤ 100 (ms) C1 and C2 ensure that the SINR of the V2I and V2V links is greater than the threshold, guaranteeing that the quality of the currently received signal meets the minimum reception requirements of the system, enabling reliable demodulation and decoding. C3 ensures that each pair of V2V links can only reuse the channel resources of one pair of V2I links. C4 ensures that the V2I links communicate using a fixed transmission power of 23 dBm, which is the commonly used selected power for V2V transmitters and is clearly defined in 3GPP. C5 indicates that the V2V link selects a transmission power of [23, 15, 5, -100] dBm, where -10... 0dBm means the transmission power is 0; C6 indicates that the V2V link needs to complete the transmission of the payload within the time constraint T; C7 indicates that the transmission rate of the V2V link in each coherent time slot cannot be less than B / T. The resource management problem described above is a non-convex optimization problem with high solution complexity. Furthermore, acquiring global information leads to increased signaling overhead, further limiting the practical application of traditional central optimization algorithms. To address this issue, this study transforms the problem into a partially observable Markov process, as shown in Figure 2. At time t, the V2V link n receives the environmental state... And generate actions independently according to the established strategy. Combining actions a t The environment then evolves to the next state based on the state transition probability. Each V2V link will receive a corresponding reward. t n Furthermore, an algorithm based on multi-agent reinforcement learning is proposed to realize V2V link subband selection and transmitter power control.

5. The method according to claim 1, characterized in that, The state space Includes: Current channel gain {G} n [m]} m∈M Received interference power {I n [m]} m∈M Remaining load B n With remaining time T n ,in 6. The method according to claim 1, characterized in that, The action space is a two-dimensional combination, including the spectrum number M = {1,2,...,m} and the transmit power.

7. The method according to claim 1, characterized in that, The reward function is: Where, λ l and λ d This represents the weight used to balance the transmission rate targets for V2I and V2V users. in Let m be the capacity of the m-th V2I link. For the capacity of the nth V2V link, for each agent, the reward is set to the effective V2V transmission capacity until the payload is delivered. After that, the reward is set to a constant β, the value of which is larger than the maximum possible V2V transmission capacity. In order to obtain more rewards, the agent will deliver the payload task faster.

8. The method according to claim 1, characterized in that, Each agent has an Actor-Critic network simulating a V2V link. In the reinforcement learning algorithm, the agent optimizes its final policy by maximizing the cumulative reward, which is defined as G. t : Where γ∈(0,1) is the discount factor, R i+t For the reward function, the state-value function and state-action-value function in the evaluation network of the nth agent are respectively: in Describes the policy of agent n, if a certain state-action pair G t To generate a positive impact, the probability of its occurrence should be increased; conversely, to generate a negative impact, its probability should be decreased. The advantage function of the nth agent is defined as follows: It measures how well or poorly the current action is performed relative to the average level; To overcome the non-convergence problem during the training process of multi-agent deep reinforcement learning, a generalized advantage estimation technique is introduced to reduce variance and bias. Combined with the advantage function of GAE, it is expressed as: Where λ∈[0,1] is a hyperparameter used to balance the trade-off between variance and bias in the estimation. Based on this, the update strategy of the Actor network can be obtained, as shown in the formula: in The probability ratio of importance sampling is represented by the clipping function, which restricts the probability ratio of the new and old policies to the interval (1-ε, 1+ε), thereby suppressing training instability caused by excessive policy update magnitude. The accuracy of the actions selected by the Actor network largely depends on the Critic network's evaluation of the state value. Temporal difference error is defined as... This error can be used to guide the parameter updates of the Critic network: Where ω* represents the parameters of the k-th agent Critic after network update, ω is the parameter before update, and lr c It is the learning rate. This update method minimizes the temporal difference error through gradient descent, thereby improving the estimation accuracy of the state value function.

9. The method according to claim 1, characterized in that, During the learning phase, each agent can receive a system reward. Then, the agents use centralized learning to train the Critic and Actor networks. During the execution phase, each agent selects the action to be executed by its trained Actor network based on the received information. All V2V agents share a joint experience pool. In each training session, gradient descent is performed through random sampling to update their respective policies and value networks.