A multi-agent spectrum sharing method based on NOMA-V2X

By combining the block coordinate descent framework and the MAPPO framework for optimization, the problem of spectrum reuse and power control coupling in NOMA-V2X scenarios is solved, achieving efficient spectrum resource allocation, improving system throughput, reliability and robustness, and adapting to the complex interference environment of high-density vehicle-to-everything (V2X) networks.

CN122458032APending Publication Date: 2026-07-24CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHONGQING UNIV OF POSTS & TELECOMM
Filing Date
2026-05-14
Publication Date
2026-07-24

Smart Images

  • Figure CN122458032A_ABST
    Figure CN122458032A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of vehicle-to-everything (V2X) communication, and particularly relates to a multi-agent spectrum sharing method for non-orthogonal multiple access based vehicle-to-everything (NOMA-V2X) communication. Firstly, a NOMA-V2X system model is established, and a mixed integer non-convex resource optimization problem is decomposed into two sub-problems of multiplexing structure optimization and power control. Secondly, a NOMA sensing exchange matching algorithm is used to optimize the spectrum multiplexing relationship, suppress the aggregate interference, and ensure the feasibility of successive interference cancellation (SIC) decoding. Then, based on a heterogeneous multi-agent proximal policy optimization (MAPPO), continuous power fine regulation is realized, and a centralized training distributed execution framework is used to complete adaptive decision-making. Finally, joint optimization is realized through block coordinate descent framework alternating iteration. The application effectively reduces system interference and signaling overhead, significantly improves V2I throughput, spectrum efficiency and energy efficiency, while ensuring high-reliability V2V transmission, and is suitable for high-density and high-dynamic vehicle-to-everything (V2X) scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of vehicle-to-everything (V2X) communication technology, specifically relating to a multi-agent spectrum sharing method for vehicle-to-everything (NOMA-V2X) communication based on non-orthogonal multiple access. Background Technology

[0002] With the rapid development of intelligent connected transportation and autonomous driving technologies, vehicle-to-everything (V2X) has become a core communication technology supporting smart transportation and road safety. In typical scenarios such as high-density urban areas and high-speed traffic, the contradiction between large-scale vehicle node access, concurrent transmission of service data, and scarce spectrum resources is becoming increasingly prominent. The traditional Orthogonal Multiple Access (OMA) mechanism requires each vehicle-to-infrastructure (V2I) link to exclusively occupy subband resources, resulting in low spectrum utilization and difficulty in simultaneously meeting the needs of massive vehicle terminal access and high-throughput transmission, becoming a key bottleneck restricting the large-scale deployment of V2X.

[0003] Non-Orthogonal Multiple Access (NOMA), through power domain multiplexing and Successive Interference Cancellation (SIC) technology, enables concurrent transmission of multiple links within the same subband, significantly improving system capacity and spectral efficiency, and is considered the next-generation core access solution for vehicle-to-everything (V2X) networks. However, when NOMA is introduced into V2X scenarios, the reuse of V2I licensed spectrum at the bottom layer of the vehicle-to-vehicle (V2V) link can cause complex cross-layer and same-layer interference. The aggregated interference of multiple V2V links directly affects the stability of base station SIC decoding, V2I reception performance, and V2V transmission reliability. Spectrum reuse decisions and power control are strongly coupled, transforming the resource allocation problem into a mixed-integer non-convex optimization problem with discrete and continuous variables. Traditional convex optimization and heuristic algorithms are difficult to solve quickly in highly dynamic vehicular environments.

[0004] Existing vehicle-to-everything (V2X) resource allocation methods are mostly based on the OMA architecture, which is insufficient for interference coordination optimization in NOMA-V2X. Some solutions adopt centralized decision-making, requiring global real-time channel state information, resulting in high signaling overhead and latency, and are unsuitable for the high-speed movement characteristics of vehicles. While multi-agent reinforcement learning-based methods can achieve distributed decision-making, they do not address the layered decoupling of the coupling characteristics between the multiplexing structure and power control. Under strong interference environments, training suffers from severe oscillations and slow convergence, and it is difficult to guarantee the high reliability requirements of V2V services. Furthermore, existing matching algorithms do not fully consider the externalities of aggregated interference, easily leading to subband interference overload, SIC decoding failure, and a significant degrade in system performance.

[0005] In summary, for high-density, high-dynamic NOMA-V2X scenarios, there is an urgent need for a spectrum sharing resource allocation technology that can decouple the multiplexing structure and power control, balance system throughput and service reliability, and adapt to distributed online deployment, so as to break through the spectrum resource bottleneck of vehicle-to-everything (V2X) and improve the robustness and practicality of the system in complex interference environments. Summary of the Invention

[0006] To address the aforementioned technical problems, this invention proposes a multi-agent spectrum sharing method based on NOMA-V2X, comprising the following steps:

[0007] Construct a NOMA-V2X communication system with a single base station and multiple vehicle terminals, divide the V2I uplink pairs and V2V links, establish channel gain, signal-to-interference-plus-noise ratio and transmission rate models, and determine spectrum reuse constraints and power constraints.

[0008] The mixed integer non-convex optimization problem, which aims to maximize the total V2I rate while taking into account the reliability of V2V transmission, is decomposed into a channel multiplexing allocation sub-problem and a power control sub-problem based on the block coordinate descent framework.

[0009] The NOMA sensing switching matching algorithm is used to optimize the multiplexing relationship between V2V links and V2I subbands under fixed power conditions, so as to obtain a stable multiplexing topology that meets the interference and reliability constraints.

[0010] A heterogeneous multi-agent model is constructed, treating V2I link pairs and V2V links as independent agents. The continuous power joint optimization is completed using the heterogeneous multi-agent multi-agent proximal policy optimization (MAPPO) framework, which features centralized training and distributed execution.

[0011] Based on a block coordinate descent alternating iterative mechanism, stable reuse matching and MAPPO power control are repeatedly executed until system performance converges, outputting the final reuse topology and power strategy. The beneficial effects of this invention are:

[0012] This invention addresses the challenge of strong coupling between spectrum reuse and power control in NOMA-V2X scenarios by employing a two-stage joint optimization method combining exchange matching and heterogeneous MAPPO, offering significant technical advantages. By decoupling the mixed-integer non-convex problem through a block coordinate descent framework, the solution complexity is effectively reduced, enabling rapid resource allocation in highly dynamic vehicular environments. The NOMA-aware exchange matching algorithm, guided by a global situation function, accurately suppresses aggregation interference, ensuring stable base station SIC decoding and significantly improving the overall V2I link rate. Heterogeneous MAPPO enables centralized training and distributed execution, with the agent relying solely on local observations for decision-making, drastically reducing signaling overhead and meeting online real-time control requirements. This invention simultaneously achieves high-speed V2I transmission and highly reliable V2V communication, maintaining a low interruption probability even in high-density link scenarios, effectively improving system spectral and energy efficiency. Compared to traditional OMA and centralized resource allocation schemes, this invention comprehensively improves system throughput, reliability, robustness, and deployment practicality, providing stable support for efficient communication in large-scale vehicle networks. Attached Figure Description

[0013] Figure 1 This is a schematic diagram of a NOMA-based V2X system model in the vehicle-to-everything (V2X) communication of the present invention;

[0014] Figure 2 This is a schematic diagram illustrating the interaction between various execution entities in the Internet of Vehicles (IoV) based on the present invention.

[0015] Figure 3 This is a flowchart of a heterogeneous DRL framework for NOMA-V2X communication matching and combination in a vehicle network according to the present invention;

[0016] Figure 4 This is a schematic diagram of the exchange and matching process based on NOMA perception in the Internet of Vehicles according to the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] In the embodiments provided by this invention, a NOMA-based V2X communication network is considered, consisting of a single base station and several vehicle terminals on urban roads, such as... Figure 1 As shown, where:

[0019] NOMA: A power-domain non-orthogonal multiple access technology that enables concurrent transmission of multiple users within the same time-frequency resource block through power superposition coding. The receiver uses serial interference cancellation (SIC) decoding, which can significantly improve spectrum efficiency and user access capacity, and is suitable for high-density vehicle-to-everything (V2X) communication scenarios.

[0020] V2X: Vehicle-to-everything communication is the core communication technology of the Internet of Vehicles. It includes communication modes such as vehicle-to-vehicle (V2V) and vehicle-to-infrastructure (V2I) and is used to transmit safety information and business data to ensure reliable interconnection of autonomous driving and intelligent transportation.

[0021] like Figure 2 As shown, this embodiment of the invention provides a multi-agent spectrum sharing method based on NOMA-V2X. The heterogeneous multi-agent spectrum sharing system constructed in this method involves a base station, V2I link pairs, and V2V links. The V2I link pairs consist of two uplinks with different channel quality and are used for high-speed data transmission between vehicles and base stations. The V2V links are used for reliable interaction of secure information between vehicles. The V2I link employs NOMA power domain multiplexing transmission, and the base station uses serial interference cancellation (SIC) decoding. The V2V link shares the V2I licensed spectrum through underlying multiplexing. Under given channel model and power constraints, the mixed-integer non-convex resource allocation problem is decomposed into two sub-problems: multiplexing structure optimization and power control, based on a block coordinate descent framework. A NOMA-aware switching matching algorithm is used, guided by the global situation function, to perform improved switching operations and obtain a stable multiplexing topology that satisfies interference and reliability constraints. A heterogeneous multi-agent model is constructed, treating the V2I and V2V links as independent agents. A MAPPO framework with centralized training and distributed execution is used to complete continuous power joint optimization. Matching and power control are cyclically executed based on an alternating iterative mechanism until system performance converges, outputting the optimal multiplexing topology and power strategy to achieve coordinated optimization of V2I throughput maximization and V2V high-reliability transmission.

[0022] like Figure 2 As shown, this embodiment of the invention provides a multi-agent spectrum sharing method based on NOMA-V2X, specifically including the following steps:

[0023] S1: Construct a NOMA-V2X communication system with a single base station and multiple vehicle terminals, divide the V2I uplink pairs and V2V links, establish channel gain, signal-to-interference-plus-noise ratio and transmission rate models, and determine spectrum reuse constraints and power constraints.

[0024] S2: The mixed integer non-convex optimization problem with the goal of maximizing the total V2I rate and taking into account the reliability of V2V transmission is decomposed into a channel multiplexing allocation sub-problem and a power control sub-problem based on the block coordinate descent framework.

[0025] S3: The NOMA sensing switching matching algorithm is used to optimize the multiplexing relationship between V2V links and V2I subbands under fixed power conditions to obtain a stable multiplexing topology that meets interference and reliability constraints.

[0026] S4: Construct a heterogeneous multi-agent model, treating V2I link pairs and V2V links as independent agents, and use the MAPPO framework of centralized training and distributed execution to complete the joint optimization of continuous power.

[0027] S5: Based on the block coordinate descent alternating iteration mechanism, it cyclically executes stable reuse matching and MAPPO power control until the system performance converges, and outputs the final reuse topology and power strategy.

[0028] In this embodiment of the invention, step S1 requires system initialization to construct a NOMA-V2X communication system with a single base station and multiple vehicle terminals. The specific steps for system initialization of the NOMA-V2X communication system with a single base station and multiple vehicle terminals include:

[0029] S101: Vehicle links are divided into two categories based on service characteristics: V2I link sets. ,common Yes, it is responsible for high-speed data transmission between vehicles and infrastructure.

[0030] S102: V2V Link Set It enables reliable interaction of safety-critical information between vehicles.

[0031] S103: All V2I links are grouped as A V2I link pair is formed by two V2I links multiplexed on the same subband via NOMA. Let the first V2I link pair be... The two V2I links in each link pair are respectively and And satisfy as well as in Indicates the first and the Channel gain of the base station of the V2I link Indicates the first and the The transmit power of the vehicles on each V2I link is such that each V2I link pair is allocated an independent subband with a bandwidth of [missing information]. The total available bandwidth is divided into Each one contains resource blocks.

[0032] In this embodiment of the invention, a NOMA-V2X communication system with a single base station and multiple vehicle terminals is constructed. After dividing the V2I uplink pairs and V2V links, it is also necessary to establish channel gain, signal-to-interference-plus-noise ratio and transmission rate models, and determine spectrum reuse constraints and power constraints.

[0033] In this embodiment of the invention, the process of establishing channel gain, signal-to-interference-plus-noise ratio, and transmission rate models, and determining spectrum reuse constraints and power constraints, may include:

[0034] S111: For V2I link pairs, the first... and the The channel gains of the V2I links to the base station are respectively and It is composed of both large-scale fading and small-scale fading. in , These are small-scale, rapidly fading power components that follow an exponential distribution with a mean of 1 and are independent of each other across different resource blocks and links. , For large-scale fading components associated with log-normal shadowing fading, , This refers to the distance from the V2I link transmitter to the base station. This is the path loss index.

[0035] S112: The The channel gain of the V2V link is in For power components, For log-normal shadowing fading, The distance from the transmitter to the receiver. This is the path loss index.

[0036] S113: Order Assign indicator variables to the spectrum. Indicates the first The V2V link in the first Each V2I link performs underlying spectrum sharing on the allocated channels; otherwise... .

[0037] S114: Assume that each V2V link is multiplexed with at most one V2I link pair, and each V2I link pair can be multiplexed with at most one V2I link pair. V2V link multiplexing , At the base station, the first The uplink received signal-to-interference-plus-noise ratio of each V2I link pair is in For the first The transmit power of the V2V link, , The first V2V link transmitter to the first Article and No. Interference channel gain at the receiver of a V2I link. This represents the power of additive white Gaussian noise.

[0038] S115: The The V2V link is subject to two types of interference, and the cross-layer interference caused by the V2I link. and co-layer interference between V2V links using the same frequency . No. The height with the first The signal-to-interference-plus-noise ratio (SINR) of the V2V link is: in , For cross-layer interference channel gain, For V2V inter-layer interference channel gain, Let be the channel gain of this link. According to Shannon's theorem, the transmission rate of each link is... , , in The expected value is represented by the statistical mean of the variables that follow. For the first The average rate of the V2I links To and The first pairing The average rate of each V2I link For the first The V2V link in the first Average transmission rate per subband.

[0039] S116: The V2V link is responsible for the reliable transmission of security-critical information. Its transmission reliability is ensured by acceptable latency. The size of the successful internal transfer is The probability of bit packets is used as a metric, with the following constraints: in The probability operator represents the probability of the event within the parentheses occurring. For the first SINR of a V2V link The minimum SINR threshold for V2V transmission. This is the upper bound of the transmission interruption probability. For a generalized fast fading distribution with a mean of 1, the minimum SINR threshold required for V2V transmission is equivalent to a deterministic constraint. in This is the equivalent SINR threshold.

[0040] In this embodiment of the invention, the NOMA-V2X resource allocation stage in step S2 also needs to consider the dual objective constraints of maximizing the total V2I rate and ensuring the reliability of V2V transmission, and generate corresponding spectrum reuse decision variables and power control variables, such as... Figure 3 As shown. Therefore, taking the decomposition and solution process of a mixed-integer nonconvex optimization problem as an example, this resource allocation process for NOMA-V2X can include:

[0041] S201: The resource allocation optimization problem of V2X-NOMA networks can be formulated as follows: Constraints C1-C3 characterize the spectrum reuse relationship between V2V links and V2I subbands, C4 ensures the transmission reliability of allocated V2V links, and C5 and C6 limit the transmit power of each link to an acceptable range.

[0042] S202: Decision-making regarding discrete multiplexing With continuous power control The essence of strong coupling is to decompose the optimization problem into two alternating subproblems using a block coordinate descent framework. In the exchange matching stage, a stable reused structure is obtained through rigorous improvement of the potential function, while in the power control stage, an approximate optimal strategy is learned through MAPPO.

[0043] S203: Subproblem A is the channel allocation subproblem, under fixed power... Optimize the reuse matrix The corresponding constraints are C1-C3.

[0044] S204: Problem B is a power control subproblem, in a fixed multiplexing structure. The following joint optimization is performed on the continuous power variables, corresponding to constraints C4-C6.

[0045] In this embodiment of the invention, the multiplexing allocation stage in step S3 also needs to consider the dual constraints of NOMA SiC decoding stability and V2V transmission reliability, and generate a stable multiplexing topology that satisfies externality awareness. Therefore, taking the NOMA-aware switching matching solution subproblem A under fixed power as an example, as follows... Figure 4 As shown, the multiplexing optimization process for V2V links and V2I subbands can include:

[0046] S301: In a fixed transmit power set Under these conditions, the channel allocation problem can be expressed as ,st C1,C2,C3.

[0047] S302: If the matching mechanism only considers interference isolation for a single link, it is difficult to suppress the superposition effect of interference from multiple links within the same subband, which will reduce the decoding stability of SIC and may lead to... and Simultaneously, the problem worsens. Based on this, an exchange matching mechanism incorporating externality awareness is designed to find locally optimal stable matches within the feasible region through heuristic exchange operations.

[0048] S303: Note The current matching state, which is related to the assignment indicator variable. Equivalent. In a given match Define subbands under fixed power conditions The expected value and rate of the upper V2I link pair are the utility functions of that subband. in, , By introducing Externalities caused by aggregation interference are explicitly incorporated into the evaluation metrics of matching decisions.

[0049] S304: The matching phase adopts the deterministic equivalent threshold given by equation (C4). The feasibility requirement for matching applies to all conditions. The V2V link whose SINR satisfies in, Under the current match, by Calculation. Meanwhile... Still through and It indirectly reflects changes in the reception quality at the base station.

[0050] S305: To measure the system-level benefits of this switching operation, the total rate of the global V2I link pairs is defined as the system potential function. The algorithm accepts candidate matches The following three conditions must be met simultaneously: First, Cardinality constraints C2 and C3 must be satisfied; secondly, in Under these conditions, all active V2V links must meet reliability requirements.

[0051] S306: Under constraints C2 and C3, the matching space is finite. Simultaneously, the power is constrained by C5 and C6, and the channel gain is also bounded. Therefore... There exists a finite upper bound. Therefore, the exchange matching process can terminate within a finite number of steps and output a stable exchange matching, which is the NOMA-aware exchange matching algorithm process with externalities.

[0052] In this embodiment of the invention, the power control stage in step S4 also needs to consider multiple constraints such as the feasibility of NOMA serial interference elimination (SIC), V2I rate maximization, and V2V transmission reliability, and generate a continuous power allocation strategy that satisfies the power boundary. Therefore, taking the process of solving sub-problem B of heterogeneous MAPPO power optimization under a fixed multiplexing topology as an example, the continuous power joint optimization process of V2I links and V2V links may include:

[0053] S401: Fixed multiplexing structure is a switched stable multiplexing structure. Subsequently, the power control subproblem is formulated. The decisions of each link are coupled with each other through subband interference and base station reception process. The power control subproblem can be regarded as a multi-agent cooperative optimization problem under local observation.

[0054] S402: Local observation of V2I agents Including uplink channel information, same-subband V2V interference and Local observation of V2V agents This includes the status of direct links, interference from V2I, mutual interference from other V2V links, and service status.

[0055] S403: V2I link to intelligent agent With V2V link agents In time slots The normalized actions output by the policy network are scaled and mapped to the actual transmit power. It is a two-dimensional action vector. , For time slots At that time, link to link and links The transmission power, To ensure that the exploration process strictly meets the hard constraints of the transmit power limit and the allocation of higher power to users in NOMA weak channels, the mapped actions must satisfy the feasible region truncation condition. .

[0056] S404: During the intensive training phase, the centralized value network (critic) takes the global state as input. Output the state value in The vector concatenation operation combines the three parts of information into a single global state vector. For all V2I link pairs of agents In the time slot Local observation set, For all V2V link pairs of agents In the time slot The local set of observations. During the execution phase, each agent relies only on its own observations. Independent decision-making is required to meet the needs of distributed online control.

[0057] S405: To guide the two types of agents to form cooperative power control, a system-level shared reward is adopted, and the optimization objective is written as a weighted combination of V2I throughput and V2V transmission completion. in, and Indicates the first A V2I link connects two users in a time slot. The instantaneous achievable speed, Used to depict the first A V2V link in time slot Transmission completion rate and These are the weighting coefficients. In this way, policy updates will not only pursue the local gains of a single link, but will incorporate both V2I throughput and V2V task completion into the learning objective.

[0058] S406: During training, this invention employs the MAPPO centralized training framework to mitigate non-stationarity in multi-agent learning. The temporal difference (TD) error is defined as... The advantage function is estimated using Generalized Advantage Estimation (GAE). in, As a discount factor, These are the GAE coefficients. For any agent... Importance sampling ratio is defined as in For the first Parameters of the agent policy network, For the new strategy, i.e., in the current parameters Below, the observation is Take action at the time The probability, Under the old strategy, i.e., before the parameter update, the observation is Take action at the time The probability of.

[0059] S407: The actor network is updated by maximizing the PPO pruning target. in For the first PPO pruning objective function for each agent For generalized advantage estimation, For cutting The upper and lower limits are limited to the range In the middle, the critic is updated by minimizing the mean squared error loss. in This is the cutting factor. To achieve the desired reward, the critic utilizes the global state. Guide all actors to improve synchronously to reduce the non-stationarity caused by simultaneous policy updates of multiple agents.

[0060] In this embodiment of the invention, the joint optimization stage in step S5 also needs to consider the iterative convergence condition, the outer loop termination threshold, and the matching-power bidirectional feedback update, and output the globally optimal reuse topology and power strategy. Therefore, taking the complete joint allocation process of alternating block coordinate descent iteration as an example, the system joint resource allocation process may include:

[0061] S501: Let This represents the current matching status. This serves as a power reference for matching utility assessment. In the... In the next outer iteration, first at a fixed power reference... The following calls the NOMA-aware exchange matching algorithm with externalities and outputs exchange-stable matches. Then fix By interacting with the environment through MAPPO to sample the trajectory and update the policy parameters, a power policy is obtained. Finally, a new power reference is constructed from the current policy output. This is used for externality evaluation in the next round of matching. When the system objective improvement is less than a threshold... Or the number of outer iterations reaches the upper limit. The algorithm stops when the time is right.

[0062] S502: Complexity of the Exchange Matching Phase: The exchange matching phase mainly includes two processes: initial matching and iterative matching. During initial matching, each V2V link is traversed at most once. There are 1 reusable sub-band, therefore the complexity of this part is O(n). During the exchange process, a maximum of [number] checks are required per round. There are 10 candidate swap pairs, therefore the complexity of a single round of swap search is O(n). Suppose the algorithm executes a total of [number] times before convergence. With each valid swap, the total complexity of the swap matching phase is... .

[0063] S503: Power control stage complexity based on MAPPO: Assume the total number of agents in the system is... The sampling length for each round is The number of rounds for updating parameters in Proximal Policy Optimization (PPO) is... Under the condition of a fixed neural network structure, the complexity of trajectory sampling is O(n). The complexity of parameter updates is Therefore, the complexity of the single-round power control stage can be expressed as: In summary, if the outer layer performs alternating optimizations and executes concurrently... If the rounds are repeated, the overall complexity upper bound of the proposed joint algorithm is: in, This indicates the number of valid swaps during the swap matching phase. This indicates the number of iterations for the outer layer's alternating optimization.

[0064] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A multi-agent spectrum sharing method for vehicle-to-everything (NOMA-V2X) communication based on non-orthogonal multiple access, characterized in that, The method includes: S1: Construct a NOMA-V2X communication system with a single base station and multiple vehicle terminals, divide the vehicle-to-infrastructure (V2I) uplink pair and the vehicle-to-vehicle (V2V) link, establish channel gain, signal-to-interference-plus-noise ratio and transmission rate models, and determine spectrum reuse constraints and power constraints. S2: The mixed integer non-convex optimization problem with the goal of maximizing the total V2I rate and taking into account the reliability of V2V transmission is decomposed into a channel multiplexing allocation sub-problem and a power control sub-problem based on the block coordinate descent framework. S3: The NOMA sensing switching matching algorithm is used to optimize the multiplexing relationship between V2V links and V2I subbands under fixed power conditions to obtain a stable multiplexing topology that meets interference and reliability constraints. S4: Construct a heterogeneous multi-agent model, treating V2I link pairs and V2V links as independent agents, and use the Multi-Agent Proximal Policy Optimization (MAPPO) framework with centralized training and distributed execution to complete the joint optimization of continuous power. S5: Based on the block coordinate descent alternating iteration mechanism, it cyclically executes stable reuse matching and MAPPO power control until the system performance converges, and outputs the final reuse topology and power strategy.

2. The multi-agent spectrum sharing method based on NOMA-V2X according to claim 1, characterized in that, In the NOMA-V2X communication system, the V2I link transmits uplink in NOMA mode, the base station uses serial interference cancellation (SIC) decoding, and the V2V link shares the V2I licensed spectrum in a low-level multiplexing manner.

3. The multi-agent spectrum sharing method based on NOMA-V2X according to claim 1, characterized in that, The channel gain is composed of large-scale fading, log-normal shadowing fading, and small-scale fast fading. During the power control phase, the agent relies only on locally observed channel information, while during the matching phase, global periodic channel state information with estimation error is used.

4. The multi-agent spectrum sharing method based on NOMA-V2X according to claim 1, characterized in that, Using the total rate of V2I link pairs as the system potential function, an aggregation interference externality evaluation mechanism is introduced; the matching relationship is iteratively optimized through heuristic switching operations to satisfy V2V reliability constraints and SIC decoding stability constraints, and finally outputs a switching-stable multiplexed topology.

5. The multi-agent spectrum sharing method based on NOMA-V2X according to claim 1, characterized in that, The heterogeneous multi-agent model includes: setting V2I link pairs and V2V links as different types of agents, with each agent making independent decisions based solely on local observations; during the centralized training phase, the global state is input by a centralized Critic network, and during the execution phase, each agent outputs power actions in a distributed manner, reducing signaling overhead and latency.

6. The multi-agent spectrum sharing method based on NOMA-V2X according to claim 1, characterized in that, The MAPPO power optimization includes: adopting a system-level shared reward, weighting and combining V2I throughput and V2V transmission completion as the optimization objective; using Generalized Advantage Estimation (GAE) to estimate the advantage function, and pruning the objective to update the policy network through Proximal Policy Optimization (PPO) to minimize the mean square error loss of the value network and alleviate the non-stationarity of multi-agent learning.

7. The multi-agent spectrum sharing method based on NOMA-V2X according to claim 1, characterized in that, The block coordinate descent alternating iteration includes: performing exchange matching with a fixed power to obtain a reused topology, then performing MAPPO power control with a fixed reused topology; iterating cyclically until the overall system rate increase is less than the convergence threshold, and outputting the globally optimal spectrum reuse and power allocation joint strategy.

8. The multi-agent spectrum sharing method based on NOMA-V2X according to claim 1, characterized in that, The method is applicable to high-density, high-dynamic vehicle-to-everything (V2X) scenarios, and can simultaneously improve system spectrum efficiency, energy efficiency and V2V transmission reliability, while reducing aggregation interference and signaling overhead.