User association and resource allocation method based on fairness perception in full decoupling network

By constructing a fairness-aware user association and resource allocation method in 6G wireless communication networks, and utilizing a distributed two-layer multi-armed slot machine algorithm and congestion pricing mechanism, the unfairness problem of user association and resource allocation in fully decoupled networks is solved, thereby maximizing system utility and improving user experience quality.

CN121397615BActive Publication Date: 2026-03-27NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In 6G wireless communication networks, unfairness exists in user association and resource allocation in fully decoupled wireless access networks. Existing methods cannot effectively guarantee fairness across time and space, leading to a decline in service quality and damage to supplier revenue.

Method used

We construct a fairness-aware user association and resource allocation method. By building network topology, resource allocation and fairness quantification models, we use a distributed two-layer multi-armed slot machine algorithm for solution. Combined with congestion pricing and Jain fairness index, we achieve optimization of user association and resource allocation.

Benefits of technology

It achieves dual protection of long-term and short-term fairness in large-scale heterogeneous networks, improves system utility and user experience quality, and reduces spatial unfairness in resource contention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121397615B_ABST
    Figure CN121397615B_ABST
Patent Text Reader

Abstract

The application discloses a user association and resource allocation method based on fairness perception in a full decoupling network, aiming at the asymmetric interference, resource coupling and user fairness loss caused by the independent operation of uplink and downlink in a 6G full decoupling wireless access network, a double-layer mechanism integrating short-term and long-term fairness is designed: short-term spatial fairness is realized through a congestion pricing mechanism, and long-term time fairness is realized by embedding a Jain fairness index; a joint optimization problem is converted into a distributed two-layer multi-arm bandit model, a distributed algorithm based on mean index exploration, adaptive utilization interval and interruption driving coordination is proposed, while ensuring logarithmic regret and balanced convergence, efficient learning and fairness convergence are realized. The application can make the Jain fairness index stable at a high level, reduce the interruption probability, and is significantly superior to existing schemes in fairness and efficiency, and is suitable for heterogeneous deployment scenarios of the 6G full decoupling wireless access network.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of 6G wireless communication networks, and particularly relates to a user association and resource allocation method based on fairness perception in a fully decoupled network. BACKGROUND

[0002] With the evolution of ultra-5G and 6G wireless networks towards ultra-dense heterogeneous deployment, large-scale connectivity and diversified service requirements, traditional coupled wireless access networks, in which each base station simultaneously handles uplink and downlink traffic, often fail to fully utilize spectrum resources under asymmetric load and heterogeneous infrastructure. Therefore, a fully decoupled wireless access network is proposed to separate uplink and downlink operations, and each user equipment can independently associate with different base stations for uplink and downlink transmission, thereby improving spectrum efficiency and coverage flexibility. However, this architecture decoupling introduces new challenges in user association and resource allocation. Due to different interference patterns and heterogeneous base station capacities experienced by uplink and downlink, distributed users compete for limited resources in the absence of centralized coordination, and these dynamics often lead to spatial unfairness. Researchers have studied these issues and designed various methods, but existing research mainly focuses on improving spectrum efficiency in fully decoupled wireless access networks, while ignoring fairness stability across time and space. Methods based on matching and reinforcement learning usually maximize aggregate utility or total rate, achieving transient load balancing but lacking sustained fairness guarantees. Optimization-based methods have high computational complexity, require perfect information assumption, and lack fairness mechanisms for fully decoupled systems, limiting their scalability in large-scale networks. The multi-armed bandit framework faces the following limitations: no fairness guarantee, slow convergence speed, etc. Previous research has mainly focused on maximizing user coverage, minimizing service delay, or optimizing user-perceived quality of service, but has failed to consider service fairness. Therefore, service providers may not be able to provide sustained fair service to their users, thereby reducing service quality and ultimately harming the service provider's revenue. SUMMARY

[0003] To solve the above problems, the application discloses a user association and resource allocation method based on fairness perception in a fully decoupled network, which considers the heterogeneous service latency and reliability requirements of users, the heterogeneous limitations of base station resources, and the characteristics of independent decision-making of uplink and downlink, optimizes the user association and resource allocation decisions for different types of services, thereby maximizing the overall system revenue and reducing the impact of potential unfairness on service availability.

[0004] To achieve the above purpose, the technical scheme of the application is as follows:

[0005] The user association and resource allocation method based on fairness perception in a fully decoupled network comprises:

[0006] Step S1: Construct a fully decoupled wireless access network system model, including a network topology model, a user association model, a resource allocation model, a congestion pricing model, and a fairness quantification model.

[0007] Step S2: In view of the heterogeneous service delay and reliability requirements of users, the heterogeneous limitation of base station resources, and the characteristics of uplink and downlink independent decision-making, overall fairness benefits and system utility are considered to construct a user association and resource allocation benefit maximization problem model based on fairness protection.

[0008] Step S3: For the characteristics of the problem, the optimization problem based on fairness protection is decomposed into two independent sub-problems in the uplink and downlink directions, and is further transformed into a distributed two-layer multi-armed bandit model, which is transformed into a non-monotonic submodular maximization problem.

[0009] Step S4: A distributed two-layer fairness-driven multi-armed bandit algorithm is used to solve the problem, and through the mean index exploration, adaptive utilization interval, and interruption-driven coordination mechanism, a user association and resource allocation scheme with theoretical performance guarantee is obtained.

[0010] Further, in step S1, a fully decoupled wireless access network system is constructed, which is composed of user equipment (UE), uplink base stations, downlink base stations, uplink orthogonal channels, downlink orthogonal channels, and discrete power levels. Let denote the user set, denote the base station set (superscript denote the uplink or downlink), denote the channel set.

[0011] Network topology model:

[0012] Each user each time slot is associated with at most one base station and occupies one channel in each direction. Let denote the association indicator variable, where denotes that user is associated with base station in direction at time slot , and 0 otherwise. Let denote the channel allocation indicator variable, where denotes that user uses channel in direction at time slot , otherwise 0. These variables satisfy the following constraints:

[0013] ,

[0014] base station and user on channel is denoted by , which contains large-scale path loss and small-scale fading. The large-scale path loss follows the standard path loss model , where is the distance between base station and user , and is the path loss exponent. The small-scale fading is independent Rayleigh fading. Based on the above network topology, further define the physical transmission characteristics. For uplink transmission, the interference suffered by user n at base station m on channel k is:

[0015] ,

[0016] where is the transmit power of user at base station on channel . For downlink transmission, the interference suffered by user n at base station m on channel is:

[0017] ,

[0018] where is the transmit power of base station to user on channel . According to the calculation result of the interference, given the bandwidth and noise power spectral density , the signal-to-interference-and-noise ratio (SINR) of user at base station on channel is:

[0019] ,

[0020] Based on the Shannon capacity formula, the achievable rate is:

[0021] ,

[0022] The total transmission rate of user in direction is:

[0023] ,

[0024] Congestion pricing model:

[0025] In fully decoupled radio access networks, independent uplink and downlink operations can easily lead to temporary congestion of certain base stations or channels. To capture and regulate this dynamics, two instantaneous indicators are introduced: channel occupancy and power consumption.

[0026] Base station In the channel direction , the channel occupancy is:

[0027] ,

[0028] The total transmit power consumption of the base station is:

[0029] ,

[0030] To quantify the degree of overload, the normalized violation of channel and power constraints is defined as:

[0031] ,

[0032] ,

[0033] where is the maximum allowed occupancy, is the maximum power limit.

[0034] These violation quantities are translated into dynamic congestion prices as a penalty for overusing resources. Specifically, the channel-level and base station-level prices are iteratively updated according to the following rules:

[0035] ,

[0036] ,

[0037] where is a decreasing step size, controls the decay of historical prices, defines a dead-zone threshold for insignificant violations, denotes the projection to .

[0038] Fairness quantification model:

[0039] The Jain fairness index is adopted as a normalized indicator to evaluate and maintain the balance of throughput over time:

[0040] ,

[0041] The user throughput is almost uniform, while Reflecting strong performance disparity.

[0042] A unified fairness-aware utility model:

[0043] To couple short-term and long-term fairness goals, a unified fairness-aware utility for each user is defined as:

[0044] ,

[0045] where denotes the total transmission rate; and denote the congestion-dependent power and channel costs; imposes a penalty when the user rate is below its QoE threshold; as a fairness reinforcement term, rewards behavior that contributes to global fairness.

[0046] Further, in step S2, in view of the heterogeneous service delay and reliability requirements of users, the heterogeneous finiteness of base station resources, and the characteristics of uplink and downlink independent decision-making, the fairness guarantee-based user association and resource allocation benefit maximization problem model is constructed by taking into account the fairness benefit and system utility, and the specific construction is as follows:

[0047] In each time slot , the overall system utility aggregates the uplink and downlink utilities of all users, weighted by a factor that reflects their relative importance:

[0048] ,

[0049] where and denote the uplink and downlink utilities of user in time slot .

[0050] The optimization goal is to maximize the long-term average system utility while maintaining short-term and long-term fairness:

[0051] ,

[0052] where , and denote the association, channel, and power allocation decisions. The constraints collectively ensure feasible user-base station association, congestion-aware resource usage, and fairness regulation through adaptive pricing and long-term balancing.

[0053] Further, in step S3, the optimization problem based on fairness guarantee is decomposed into two sub-problems for uplink and downlink directions independently, and further transformed into a distributed two-layer multi-armed bandit model. Specifically as follows:

[0054] The system-level optimization problem is decomposed into two directionally separated sub-problems, corresponding to uplink (UL) and downlink (DL) respectively:

[0055]

[0056]

[0057] where and denote the uplink and downlink utility of user at time slot . This decomposition is consistent with the FD-RAN architecture, where uplink and downlink run on orthogonal frequency spectrum with different sets of base stations.

[0058] To ensure the optimality of the decomposition, each-user instantaneous utility is rewritten under congestion pricing:

[0059]

[0060] where and are the uplink / downlink actions of user , and price evolves according to a price update rule. Term depends on the aggregate uplink+downlink rate of all users, and is thus a broadcast, user-independent component.

[0061] Since uplink and downlink run on orthogonal frequency spectrum, their feasible sets are independent. The fairness index is an additive broadcast term that does not change the ordering of uplink or downlink actions. Therefore, maximizing the sum of the two additively separable directional utilities simplifies to maximizing each on its respective feasible set.

[0062] With this decomposition, the resource allocation task is further modeled as a two-layer multi-agent multi-armed bandit (MAMAB) framework, where each user acts as an autonomous learning agent that adaptively selects its uplink and downlink associated policies based on local observations, and the dynamic price signal provides light-weight coordination to align individual decisions with the global fairness goal.

[0063] The decision framework is defined as:

[0064]

[0065] where​​​​ and represent the uplink and downlink slot processes. Each user n interacts with both layers simultaneously by choosing an action:

[0066] ,

[0067] and receives the corresponding stochastic reward and . The user joint decision at time slot is:

[0068] ,

[0069] The overall expected utility is expressed as:

[0070] ,

[0071] where balances the uplink and downlink priorities.

[0072] For each transmission direction , the action space for user is:

[0073] ,

[0074] where denotes the set of candidate base stations, denotes the available channels, denotes the discrete power levels. At each decision time slot , user chooses an action:

[0075] ,

[0076] which corresponds to a base station association, channel selection, and power control in the given direction.

[0077] Given an action , the instantaneous reward is defined as a fairness-aware utility:

[0078] ,

[0079] where integrates the throughput, congestion penalty and , QoE penalty, and long-term fairness index . Thus, the reward reflects both local efficiency and system-level fairness.

[0080] To maintain consistency between the two layers, both uplink and downlink rewards are adjusted by the shared fairness signal:

[0081] ,

[0082] in It is a direction-specific fairness-perceived reward function that shares pricing and fairness terms to implicitly coordinate users' independent decisions, guiding them toward overall network efficiency and balanced performance.

[0083] Furthermore, in step S4, a distributed two-layer fairness-driven multi-armed slot machine algorithm is used to solve the problem, resulting in a user association and resource allocation scheme with theoretical performance guarantees, as detailed below.

[0084] In each refresh time slot, the user By selecting the index with the highest mean from the neighborhood To construct a candidate set, actions are combined with random actions. (size Each candidate action Executed Next update:

[0085] ,

[0086] in It's a cumulative reward. It refers to the number of executions. The optimal action is... .

[0087] The algorithm uses exponentially increasing intervals. Phase randomization During the interval During this period, the user utilizes the current best action. Instead of exploring, this adaptive scheduling balances exploration overhead and convergence speed.

[0088] When the user changes their action selection, it broadcasts an interrupt signal to the neighbors, who then set a flag. This triggers an immediate refresh, enabling rapid adaptation and exponentially increasing intervals. Prevent cascading oscillations.

[0089] Base stations update prices according to price update rules. and These prices appear in the rewards China. Global Fairness Index Excitation equilibrium distribution. Projection operator. and Ensure bounded price growth, while the dead zone threshold Prevent oscillation.

[0090] The specific algorithm flow is as follows:

[0091] Step S4-1: Initialization. The algorithm initializes the statistics of all actions from the beginning and . The last action and the interrupt flag are initialized. The phase is drawn, the power index is set, the counter is initialized, and the interval is set.

[0092] Step S4-2: Refreshing. In each time slot , the counter is updated. The condition is checked. If the refreshing condition is met, the exploration phase is entered; otherwise, the exploitation phase is entered.

[0093] Step S4-3: Exploration phase. The uplink candidate set (mean sorted neighbors + random) is constructed, . For each , the following is performed times: the action is performed, the reward from the reward function is observed, and the is updated. The optimal uplink action is selected. Similarly, the downlink candidate set is constructed, and the optimal downlink action is selected.

[0094] Step S4-4: Action updating and interrupt broadcasting. If or , the interrupt signal is broadcast to the neighbors, and the and are updated. The interrupt flag is reset, the power index is updated, the counter is reset, and the interval is updated.

[0095] Step S4-5: Exploitation phase. If refreshing is not needed, the last action is performed once, and the is updated; the last action is performed once, and the is updated.

[0096] Step S4-6: Joint action formation and price updating. The joint action is formed. The base station updates the price using the price updating rule (sparse broadcasting). The Jain fairness index is updated using the Jain fairness index formula.

[0097] Step S4-7: Repeat steps S4-2 to S4-6 until the time range is reached .

[0098] Step S4-8: Return the final user association and resource allocation scheme.

[0099] The beneficial effects of the present application are:

[0100] Double-layer fairness guarantee mechanism: The present application fuses short-term and long-term double-layer fairness mechanisms, uses a congestion pricing mechanism to punish overloaded base stations and channels in real time, effectively alleviates short-term spatial resource contention; at the same time, by embedding the Jain fairness index as a long-term incentive, the user throughput is ensured to be balanced in the time dimension. Simulation results show that in a large-scale heterogeneous network, the present application can keep the Jain fairness index stable at above 0.96, which is significantly better than traditional methods.

[0101] Efficient distributed learning architecture: A distributed two-layer multi-armed bandit model is designed for the uplink and downlink decoupling characteristics of FD-RAN. The model uses direction separability to decompose the complex joint optimization problem into independent sub-problems, combined with the mean index exploration strategy, which greatly reduces the search complexity of the action space and realizes the rapid convergence of logarithmic regret.

[0102] Low-overhead dynamic adaptability: The interrupt-driven coordination and phase randomization mechanism is introduced, which enables users to quickly respond to network state changes and avoid group synchronous oscillation. In addition, the base station uses a sparse broadcast strategy to update the price only when it is overloaded, which significantly reduces the signaling overhead, making the method highly scalable and suitable for 6G large-scale dense deployment scenarios.

[0103] Dual protection of QoS and QoE: By introducing a QoE penalty term in the utility function, the minimum rate requirement of users is strictly guaranteed in resource allocation, effectively reducing the interruption probability and improving the perceived service quality of users. BRIEF DESCRIPTION OF DRAWINGS

[0104] Figure 1 is a schematic diagram of a fully decoupled wireless access network in the present application example;

[0105] Figure 2 is a specific flowchart in the present application example;

[0106] Figure 3 is an approximate algorithm flowchart in the present application example. DETAILED DESCRIPTION

[0107] The present application will be further illustrated below in conjunction with the drawings and specific embodiments, and it should be understood that the following specific embodiments are only used to illustrate the present application and not to limit the scope of the present application.

[0108] As Figure 2 shown, the user association and resource allocation method based on fairness perception in the full decoupling network of the application comprises:

[0109] Step S1: Construct a full decoupling wireless access network system model, including network topology model, user association model, resource allocation model, congestion pricing model and fairness quantification model.

[0110] As Figure 1 shown, a full decoupling wireless access network system is constructed, which is composed of user equipment (UE), uplink base stations, downlink base stations, uplink orthogonal channels, downlink orthogonal channels and discrete power level groups. Let denote the user set, denote the base station set (superscript denotes applicable to uplink or downlink), denote the channel set.

[0111] For each user and each time slot , define the association indicator variable , where denotes that user associates base station in direction in time slot , otherwise 0. Define the channel allocation indicator variable , where denotes that user uses channel in direction in time slot , otherwise 0. These variables satisfy the following constraints:

[0112] ,

[0113] The channel gain of base station and user on channel is denoted as , which includes large-scale path loss and small-scale fading. The large-scale path loss follows the standard path loss model , where is the distance between base station and user , is the path loss exponent. The small scale fading is independent Rayleigh fading.

[0114] For uplink transmission, the base station receives the signal from the user on the channel .

[0115] ,

[0116] where is the transmit power of the user on the channel .

[0117] For downlink transmission, the user receives the signal from the base station on the channel .

[0118] ,

[0119] where is the transmit power of the base station to the user on the channel .

[0120] Given the bandwidth and the noise power spectral density , the signal-to-interference-and-noise ratio (SINR) of the user on the channel is:

[0121] ,

[0122] Based on the Shannon capacity formula, the achievable rate is:

[0123] ,

[0124] The total transmission rate of the user in the direction is:

[0125] .

[0126] In a fully decoupled radio access network, independent uplink and downlink operations can easily lead to temporary congestion of certain base stations or channels. To capture and regulate this dynamics, two instantaneous indicators are introduced: channel occupancy and power consumption.

[0127] The channel occupancy of the base station in the direction of the channel is:

[0128] ,

[0129] The total transmit power consumption of base station m is:

[0130] ,

[0131] To quantify the degree of overload, the normalized violation quantities of channel and power constraints are defined as:

[0132] ,

[0133] ,

[0134] where , is the maximum allowed occupancy, is the maximum power limit.

[0135] These violation quantities are translated into dynamic congestion prices as a penalty for overusing resources. Specifically, the channel-level and base station-level prices are iteratively updated according to the following rules:

[0136] ,

[0137] ,

[0138] where is a decreasing step size, controls the decay of historical prices, defines a dead zone threshold for insignificant violations, denotes the projection to .

[0139] The Jain fairness index is adopted as a normalized indicator to evaluate and maintain the balance of throughput over time:

[0140] ,

[0141] indicates that user throughputs are almost uniform, while reflects strong performance disparity.

[0142] To couple short-term and long-term fairness objectives, a unified fairness-aware utility for each user is defined as:

[0143] ;

[0144] where denotes the instantaneous throughput reward for user ; and denote the congestion-related power and channel costs; Penalize users when their rate is below their QoE threshold; As a fairness reinforcement, reward behaviors that help global fairness.

[0145] Step S2: Considering the heterogeneous service latency and reliability requirements of users, the heterogeneous limitation of base station resources, and the characteristics of uplink and downlink independent decision-making, the problem of maximizing the benefits of user association and resource allocation based on fairness protection is constructed.

[0146] In each time slot , the overall system utility aggregates the uplink and downlink utilities of all users, weighted by a factor that reflects their relative importance:

[0147] ,

[0148] The optimization goal is to maximize the long-term average system utility while maintaining short-term and long-term fairness:

[0149] ,

[0150] where , and represent association, channel, and power allocation decisions, respectively. The constraints collectively ensure feasible user-base station association, congestion-aware resource usage, and fairness regulation through adaptive pricing and long-term balancing.

[0151] Step S3: Based on the characteristics of the problem, the optimization problem based on fairness protection is decomposed into two independent sub-problems for uplink and downlink directions, and further transformed into a distributed two-layer multi-armed bandit model.

[0152] The system-level optimization problem is decomposed into two directionally separated sub-problems, corresponding to the uplink (UL) and downlink (DL):

[0153] ,

[0154] ,

[0155] where and represent the uplink and downlink utilities of user in time slot . This decomposition is consistent with the FD-RAN architecture, where uplink and downlink run on orthogonal frequency spectrums with different sets of base stations.

[0156] To ensure that the decomposition remains optimal, rewrite the instantaneous utility of each user under congestion pricing:

[0157] ,

[0158] where and are the user's uplink / downlink actions, and the price evolves according to a price update rule. The term depends on the aggregate uplink+downlink rates of all users, and is thus a broadcast, user-agnostic component.

[0159] Since uplink and downlink operate on orthogonal spectrums, their feasible sets are independent. The fairness index is an additive broadcast term that does not change the ordering of uplink or downlink actions. Thus, based on the independence of the uplink and downlink decision spaces, the joint utility maximization problem is decomposed into utility maximization sub-problems in each direction independently.

[0160] With this decomposition, the resource allocation task is further modeled as a two-layered multi-agent multi-armed bandit (MAMAB) framework, where each user acts as an autonomous learning agent that adaptively selects its uplink and downlink associated policies based on local observations, and the dynamic price signal provides light coordination to align individual decisions with the global fairness objective.

[0161] The decision framework is defined as:

[0162] ,

[0163] where and denote the uplink and downlink bandit processes. Each user n interacts with both layers simultaneously by selecting actions:

[0164] ,

[0165] and receives the corresponding stochastic rewards and . The user's joint decision at time slot is:

[0166] ,

[0167] The overall expected utility is expressed as:

[0168] ,

[0169] where balances the uplink and downlink priorities.

[0170] For each transmission direction , the user​ The action space is:

[0171] ,

[0172] in Represents the set of candidate base stations. Indicates available channels. This represents discrete power levels. In each decision time slot... ,user Choose one action:

[0173] ,

[0174] This corresponds to base station association, channel selection, and power control in a given direction.

[0175] Given action Instantaneous reward is defined as perceived fairness utility:

[0176] ,

[0177] in Integrated throughput, congestion penalty and QoE penalties and long-term fairness index Therefore, rewards reflect both local efficiency and system-wide fairness.

[0178] To maintain consistency between the two layers, both uplink and downlink rewards are adjusted by sharing a fairness signal:

[0179] ,

[0180] in It is a direction-specific fairness-perceived reward function. Shared pricing and fairness terms implicitly coordinate users' independent decisions, guiding them toward overall network efficiency and balanced performance.

[0181] Step S4: A distributed, two-layer fairness-driven multi-armed slot machine algorithm is used to solve the problem, obtaining a user association and resource allocation scheme with theoretical performance guarantees. The specific algorithm flow is as follows: Figure 3 As shown.

[0182] In each refresh time slot, the user By selecting the index with the highest mean from the neighborhood To construct a candidate set, actions are combined with random actions. (size Each candidate action a is executed Z times to update:

[0183] ,

[0184] in It's a cumulative reward. It refers to the number of executions. The optimal action is... .

[0185] The algorithm uses exponentially increasing intervals. Phase randomization During the interval During this period, the user utilizes the current best action. Instead of exploring, this adaptive scheduling balances exploration overhead and convergence speed.

[0186] When the user changes their action selection, it broadcasts an interrupt signal to the neighbors, who then set a flag. This triggers an immediate refresh. This enables rapid adaptation, while the intervals grow exponentially. Prevent cascading oscillations.

[0187] Base stations update prices according to price update rules. and These prices appear in the rewards In China, global fairness refers to... Numerical excitation equilibrium distribution. Projection operator. and Ensure bounded price growth, while the dead zone threshold Prevent oscillation.

[0188] The algorithm flow is as follows:

[0189] Step S4-1: Starting from initialization, the algorithm initializes the statistical information of all actions. and Initialize the last action. and interruption flag Plot the phase Set power index ,counter ,interval .

[0190] Step S4-2: Determine if a refresh is needed. In each time slot... Update the counter Judgment conditions If the refresh conditions are met, proceed to the exploration phase; otherwise, proceed to the utilization phase.

[0191] Step S4-3: Exploration Phase. Constructing the Upstream Candidate Set. (Neighbors sorted by mean + random) For each ,implement Next: Execution of action , observe the reward function from , update . Select the optimal uplink action . Similarly, construct the downlink candidate set and select the optimal downlink action .

[0192] Step S4-4: Action update and interrupt broadcast. If or , broadcast an interrupt signal to the neighbors, update and . Reset the interrupt flag , update the power index , reset the counter , update the interval .

[0193] Step S4-5: Utilization phase. If no refresh is needed, execute the last action once, update ; execute the last action once, update .

[0194] Step S4-6: Form joint action and price update. Form the joint action . The base station updates the price (sparse broadcast) using the price update rule. Update the fairness index using the Jain fairness index formula .

[0195] Step S4-7: Repeat steps S4-2 to S4-6 until the time range is reached.

[0196] Step S4-8: Return the final user association and resource allocation scheme.

[0197] It should be noted that the above merely illustrates the technical idea of the present application and cannot limit the protection scope of the present application. For those of ordinary skill in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, which fall within the protection scope of the claims of the present application.

Claims

1. A method for user association and resource allocation based on fairness perception in fully decoupled network, characterized in that: The method comprises the following steps: Step S1: According to the heterogeneous service delay and reliability requirements of users, the heterogeneous limitation of base station resources, and the characteristics of uplink and downlink independent decision, the fairness benefit and system utility are overall planned, a completely decoupled wireless access network system model is constructed, including a network topology model, a congestion pricing model, a fairness quantification model, and a user association and resource allocation benefit maximization problem model based on fairness guarantee; Based on the fairness of the user associated with the protection and the construction of the maximum profit of resource allocation model, as follows: in each time slot , the overall system utility aggregates all the uplink and downlink utility of users, weighted by the factor reflecting its relative importance, that is, , wherein and denote the uplink and downlink utility of a user at time slot respectively; The optimization goal is to maximize the long-term average system utility while maintaining short-term and long-term fairness, that is , wherein , and denote association, channel and power allocation decisions respectively, the constraints collectively ensure feasible user-base station association, congestion-aware resource usage and fairness regulation through adaptive pricing and long-term balancing, the fairness-aware utility function exhibits non-convexity due to the interdependence among throughput, congestion price and global fairness index in view of problem characteristics, the recursive price update and time-varying fairness index further introduce time coupling, making the problem dynamic and non-stationary; Step S3: According to the problem characteristics, the optimization problem based on fairness guarantee is decomposed into two independent sub-problems in uplink and downlink directions, and is further transformed into a distributed two-layer multi-armed bandit model, and the problem is transformed into a non-monotonic submodular maximization problem; Specifically as follows: The system-level optimization problem is decomposed into two direction-separated sub-problems, corresponding to uplink and downlink respectively, that is And , where and denote the uplink and downlink utility of user at time slot , this decomposition is consistent with the FD-RAN architecture where uplink and downlink run on orthogonal frequency spectrum with different sets of base stations, in order to ensure that the decomposition remains optimal, rewrite each user's instantaneous utility under congestion pricing as , wherein and are the uplink / downlink actions of users , the price evolves according to a price update rule, the item depends on the aggregate uplink+downlink rates of all users, based on the independence of the uplink and downlink decision spaces, the joint utility maximization problem is decomposed into two independent directions of utility maximization sub-problems, using this decomposition, the resource allocation task is modeled as a two-level multi-agent multi-armed bandit framework, where each user acts as an autonomous learning agent, adaptively selecting its uplink and downlink associated policies based on local observations, the dynamic price signal provides lightweight coordination to align individual decisions with the global fairness goal, the decision framework is defined as , where and denote the uplink and downlink slot processes, respectively, for each user interacts with both layers simultaneously by choosing an action, i.e., and and receives the corresponding stochastic reward and The joint decision of the users at time slot is The overall expected utility is represented as , wherein Balancing uplink and downlink priorities, For each transmission direction ,user The action space is , where denotes the set of candidate base stations, denotes the available channels, denotes the discrete power levels, at each decision slot , the user selects an action , which corresponds to base station association, channel selection and power control in a given direction, given action , the instantaneous reward is defined as fairness-aware utility where integrates throughput, congestion penalty and , QoE penalty and long-term fairness index Thus, the reward reflects both local efficiency and system-level fairness, in order to maintain consistency between the two layers, both uplink and downlink rewards are adjusted by the shared fairness signal, i.e. where is the direction-specific fairness-aware reward function, the shared pricing and fairness terms implicitly coordinate the users' independent decisions, guiding them towards overall network efficiency and balanced performance; Step S4: A distributed two-layer fairness-driven multi-armed bandit algorithm is used for solving, through mean index exploration, adaptive utilization interval and interruption-driven coordination mechanism, a user association and resource allocation scheme with theoretical performance guarantee is obtained; At each refresh slot, the user The candidate set is constructed by selecting from the neighborhood the actions with the highest mean index plus a random action Each candidate action is executed times to update , where is the cumulative reward, is the number of executions, the optimal action is , the algorithm adopts exponentially increasing intervals , where the phase is randomized , denotes the ceiling function, during the interval , the user exploits the current best action without exploration, this adaptive schedule balances the exploration overhead and convergence speed, when the user changes the action selection, it broadcasts an interrupt signal to the neighbors, the neighbors set a flag to trigger an immediate refresh, this enables fast adaptation, while the exponentially increasing intervals prevent cascading oscillations, the base station updates the prices according to the price update rule and , these prices appear in the reward , the global fairness index encourages balanced distribution, the projection operator and ensure bounded price growth, while the dead zone threshold prevents oscillations; The specific process of the algorithm in step S4 is as follows: Step S4-1 : From the initialization, the algorithm initializes the statistics of all actions and , initializes the last action and the interrupt flag , draws the phase , sets the power index , the counter , the interval ; Step S4-2: judging whether refreshing is needed, in each time slot , updating the counter , judging condition , if the refreshing condition is met, entering the exploring phase, otherwise entering the exploiting phase; Step S4-3: Exploration phase, build the uplink candidate set , , for each , perform times: perform action , observe from the reward function, update , select the optimal uplink action , then build the downlink candidate set and select the optimal downlink action ; Step S4-4: Action update and interruption broadcast, if or , broadcast interruption signal to neighbors, update and , reset interruption flag , update power index , reset counter , update interval ; Step S4-5: Utilization phase, if no refresh is needed, perform last action Once, update , perform last action Once, update ; Step S4-6: Forming joint action and price update, forming joint action wherein representing the optimal joint action of the users the base station updates the prices using the price update rule and updates the fairness index using the Jain's fairness index formula ; Step S4-7: Steps S4-2 to S4-6 are repeated until a time range is reached ; Step S4-8: Return the final user association and resource allocation scheme.

2. The method of claim 1, wherein the method is implemented in a fully decoupled network. In step S1, the network topology model is defined: A fully decoupled wireless access network system is constructed from a set of user equipments, a set of uplink base stations, a set of downlink base stations, a set of uplink orthogonal channels, a set of downlink orthogonal channels, and a set of discrete power levels, let denote the set of users, denote the set of base stations, denote the set of channels, each user is associated with at most one base station and occupies one channel in each time slot in each direction, let denote the association indicator variables, where denotes that user is associated with base station in direction in time slot , otherwise 0, let denote the channel allocation indicator variables, where denotes that user uses channel in direction in time slot , otherwise 0, these variables satisfy the following constraints: ; Base station With user On channel The channel gain on channel is denoted as , which contains large-scale path loss and small-scale fading, the large-scale path loss follows the standard path loss model where is the distance between base station and user , is the path loss exponent, and the small-scale fading is independent Rayleigh fading, based on the above network topology, further define the physical transmission characteristics, for uplink transmission, the interference suffered by base station on channel is: , wherein is a user at a base station channel power on the channel, for downlink transmission, a user experiences an interference on the channel ​ , wherein is a base station to a user transmit power on a channel , given a bandwidth and a noise power spectral density , a user signal-to-interference-and-noise ratio at a base station channel is , Based on the Shannon capacity formula, the achievable rate is , User The total transmission rate in the direction is 。 3.The method of claim 1, wherein: In step S1, the congestion pricing model is defined: In fully decoupled radio access networks, independent uplink and downlink operation can easily lead to temporary congestion of certain base stations or channels. In order to capture and regulate this dynamics, two instantaneous indicators are introduced: channel occupancy and power consumption, base station The channel occupancy The channel occupancy in the direction , Base station The total transmit power consumption is , In order to quantify the overload degree, the normalized violation quantity of channel and power constraint is defined as , wherein is the maximum allowed occupancy, is the maximum power limit, these violation quantities are translated into dynamic congestion prices as a punishment for overusing resources, specifically, the channel level and base station level prices are iteratively updated according to the following rules: , wherein is a decreasing step size, controls the decay of historical prices, defines a deadband threshold for insignificant violations, denotes the projection to .

4. The method of claim 1, wherein the method is characterized in that: In step S1, the fairness quantification model is defined: The Jain fairness index is used as a normalized index to evaluate and maintain the balance of throughput over time, that is , wherein represents a user throughput almost uniform, while reflects a strong performance disparity, In order to couple the short-term and long-term fairness goals, the unified fairness perception utility of each user is defined as , where denotes the users The total transmission rate, and denotes the congestion-dependent power and channel cost, imposes a penalty when the user rate is below its QoE threshold, as a fairness reinforcement term, rewards behavior that contributes to global fairness.​

Citation Information

Patent Citations

  • Uplink and downlink joint dynamic resource allocation method for 6G full decoupling network

    CN114554494A

  • Fair enhancement method for layered federated learning RSU scheduling in Internet of Vehicles scene

    CN118014305A