A method for inter-cell interference coordination based on multi-agent reinforcement learning

CN116489660BActive Publication Date: 2026-08-11SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-24
Publication Date
2026-08-11

AI Technical Summary

Benefits of technology

[0100]本发明的一种基于多智能体强化学习的小区间干扰协调方法,具有以下优点:利用多智能体强化学习算法,将系统内PBS划分为多个智能体,采用集中式训练、分布式执行的算法,使其根据智能体自身及其他智能体中的用户分布及所受干扰情况,单独调整小范围内或单个PBS的CRE偏置值,避免了偏置值空间尺寸膨胀、优化难度高的问题,综合考虑全局信息,有效提高了系统吞吐量并且改善系统内单个用户吞吐量。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116489660B_ABST
    Figure CN116489660B_ABST
Patent Text Reader

Abstract

This invention relates to a cell interference coordination method based on multi-agent reinforcement learning, belonging to the field of wireless mobile communication technology. To improve system throughput in heterogeneous network scenarios and address the problem of excessive action space when the cell range extension CRE bias value is adjusted individually by the picocell (PBS) within the system, this invention proposes a scheme for dynamically adjusting the CRE bias value using a multi-agent reinforcement learning algorithm. The scheme includes the following steps: constructing a downlink transmission link model of the heterogeneous network and proposing a problem to maximize system throughput; constructing an experience pool for the agents and a deep reinforcement learning neural network, initializing the CRE bias value of the PBS within the system; storing the experience pool through interaction between the PBS and the wireless communication system, and randomly sampling from the experience pool to train the neural network and update the neural network parameters; outputting the CRE bias value setting. This invention can effectively improve system throughput and alleviate interference experienced by users within the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wireless mobile communication technology, and in particular to a method for inter-cell interference coordination in heterogeneous networks. Background Technology

[0002] Heterogeneous networks refer to the deployment of low-power nodes such as relays and pico base stations (PBS) within the coverage area of ​​macro base stations (MBS). Due to their complex multi-layered and multi-access network structure, heterogeneous networks suffer from problems such as unbalanced network load and severe inter-cell interference.

[0003] To address the challenges in heterogeneous networks, Cell Range Expansion (CRE) technology has been introduced into the heterogeneous network architecture. The core principle of CRE is to add a positive offset, called the CRE offset value, to the Base Station Utility (PBS) that users typically use for access based on their maximum reference signal receiving power (Max-RSRP). The presence of a positive CRE offset value allows users to access PBSs that offer lower signal-to-interference-plus-noise ratio (SINR). In other words, by lowering the barrier to entry for users to access PBSs, more users choose PBSs, increasing the PBS's ability to transfer MBS load. Depending on how users are associated with base stations, users associated with PBSs can be divided into two groups: Pico Expansion User Equipment (PEUE) users associated with PBS base stations due to the positive CRE offset value, and Pico Center User Equipment (PCUE) users associated with PBS base stations even without the positive CRE offset value.

[0004] Users in areas where the coverage of base stations with different power levels overlap, especially those at the edge of pico cells, will suffer interference from high-power MBS. For CRE technology, on the one hand, if the offset value is set too small, the pico cell coverage area cannot be significantly expanded, and the CRE load balancing capability cannot be clearly demonstrated; on the other hand, if the offset value is set too large, the pico cell coverage area is unreasonably expanded, forcing users who should be better suited to access MBS to choose PBS, thus suffering strong interference from MBS. Therefore, it is necessary to choose an appropriate CRE offset value. Many technical personnel have conducted in-depth research on this issue. The paper "T.Jung,I.Song,S.Lee,S.Jung,S.Yoon and J.Kang,"CellRange Expansion with Geometric Information of Pico-Cell in HeterogeneousNetworks,"2018IEEE 87th Vehicular Technology Conference (VTC Spring),2018,pp.1-5,doi:10.1109 / VTCSpring.2018.8417616" considers the position of PBS in the system and proposes a CRE bias value adjustment scheme to maximize system performance and rate. The paper "GU Jing,DENG Yifei,ZHANGXin.Dynamic CRE bias selection algorithm based on heuristic reinforcement learning[J].Computer Engineering,2020,46(5):200-206" proposes an HSARSA algorithm for dynamic selection of CRE bias value based on heuristic function, in which the home base station independently learns the CRE bias value from experience.

[0005] Existing algorithms using deep reinforcement learning to solve the CRE bias optimization problem are mainly divided into centralized and distributed frameworks. In the centralized framework, a central controller acts as an agent, setting different CRE bias values ​​for all base stations in the system to improve system throughput. However, a drawback is that the bias value space expands rapidly with the number of base stations, increasing the optimization difficulty and complexity. In the distributed framework, multiple PBSs act as multiple agents, trained and executed independently. However, a drawback is insufficient global information interaction, ignoring the cooperative relationships between all base stations in the system; individual optimality does not represent global optimality. Summary of the Invention

[0006] This invention proposes a small-interval interference coordination technique based on a multi-agent reinforcement learning algorithm to dynamically adjust the CRE bias value in heterogeneous networks. This technique addresses the problems of uneven network load and severe small-interval interference while improving the total throughput of all users in the system.

[0007] To address the aforementioned technical problems, the present invention employs a method for inter-cell interference coordination based on multi-agent reinforcement learning, comprising the following steps:

[0008] Step 1, System Modeling Stage: Construct a Macro-Pico heterogeneous network interference scenario as a wireless communication system, considering the downlink transmission link of a two-layer heterogeneous network consisting of macro base stations (MBS) and pico base stations (PBS), and propose the problem of maximizing system throughput.

[0009] Step 1, the system modeling phase, includes the following specific steps:

[0010] Step 1.1: Consider a system with 1 MBS and N. P PBS, N U Users are randomly and uniformly distributed throughout the system, with each PBS containing N users within its coverage area. UP N are randomly distributed users. U This represents the total number of users within the system. When generating all users, it is required to generate N within the coverage area of ​​each PBS. UP The remaining N U -N UP ×N P Users are randomly distributed within the MBS coverage area.

[0011] Step 1.2: Propose a cell range expansion model, as follows:

[0012] In heterogeneous networks, users often use the Max-RSRP strategy to select associated base stations, meaning each user is associated with the base station providing the highest observed RSRP. To effectively broaden the service range of the PBS (Balanced Base Station) and encourage more users, especially those located at the boundary between macrocells and picocells, to select the PBS, the CRE (Continuous Enhancement Rejection) technique adds a positive bias value to each PBS. In this case, the user's associated base station is:

[0013]

[0014] Where i represents the i-th base station, u represents the u-th user, i = 0 refers to the macro base station MBS, and 1 ≤ i ≤ N. P The term refers to the picocell PBS; That is, the associated base station number of the u-th user when using CRE technology. σ i This represents the CRE bias value of base station i. Since MBS does not set a CRE bias value, σ0 = 0; RSRP iu The RSRP received by the u-th user from the i-th base station is calculated as follows:

[0015] RSRP iu =P i g iu (2)

[0016]

[0017] Among them, P i P represents the transmit power of the i-th base station. M P represents the transmit power of a macro base station (MBS). P g represents the transmit power of the picocell PBS. iu This represents the channel gain between the u-th user and the i-th base station;

[0018] Step 1.3: Propose the Almost Blank Subframe (ABS) model, as follows:

[0019] This invention combines the Absolute Block (ABS) technique from temporal interference coordination technology to mitigate the interference of CRE (Continuous Reaction) technology on edge users. The ABS technique is characterized by dividing subframes into two types: non-ABS subframes and ABS subframes. During ABS subframes, the MBS (Mobile Base Station) does not transmit any valid data, and the PBS (Balanced Partition Tool) schedules PEUEs (Peer-to-Peer Units) severely interfered with by the MBS. During non-ABS subframes, the MBS normally serves macro user equipment (MUE), and the PBS only schedules PCUEs (Peer-to-Peer Units). Users experience different degrees of interference during non-ABS and ABS periods; therefore, the SINR at the u-th user served by the PBS is represented as follows during different subframe periods:

[0020]

[0021]

[0022] in, This indicates the SINR of the u-th user associated with PBS during the ABS subframe. This indicates the SINR of the u-th user associated with PBS during a non-ABS subframe; This refers to the transmission power of the associated base station. This represents the channel gain between the u-th user and the associated base station; Indicates association with the first The co-channel interference experienced by the u-th user of a PBS base station; N0 is the noise power;

[0023] The SINR of the u-th user served by MBS is represented in different subframes as follows:

[0024]

[0025]

[0026] in, This indicates the SINR of the u-th user associated with MBS during the ABS subframe. This represents the SINR of the u-th user associated with the MBS during a non-ABS subframe period. Since the MBS does not transmit any valid data during the ABS subframe period, the SINR is represented as zero. P0 represents the transmit power of the 0th base station, i.e., the transmit power of the MBS, P0 = P M g 0u This represents the channel gain between the u-th user and the 0th base station, i.e., the channel gain between the u-th user and the MBS.

[0027] Define the number of users served by the i-th base station during the ABS subframe as: The number of users served during non-ABS subframes was The ABS ratio β is defined as the ratio of the number of ABS subframes to the total number of subframes in a frame, and β satisfies the following condition:

[0028] 0 < β < 1 (8)

[0029] Step 1.4: Propose a model to maximize system throughput, as follows:

[0030] This invention assumes that each user receives almost equal amounts of resources; therefore, the throughput of the u-th user can be calculated using Shannon's formula:

[0031]

[0032] Where W is the system bandwidth, and These represent the 1st, 2nd, and 3rd user-related terms associated with the uth user. The number of users served by each base station during ABS and non-ABS periods. This varies depending on the type of base station associated with the user, when the user is associated with a PBS. equal When a user associates MBS equal Similarly, when a user associates with PBS equal When a user associates MBS equal

[0033] The throughput of all users associated with the i-th base station can be calculated as follows:

[0034]

[0035]

[0036]

[0037] in, This represents the sum of the throughput of users within the i-th base station during the ABS period; This represents the sum of the throughput of users within the i-th base station during non-ABS periods.

[0038] System throughput is the total throughput of all cells within the system, and also the sum of the throughput of all users within the system. The objective of this invention is to maximize the sum of the throughput of all users within the system by jointly optimizing the CRE bias values ​​of all PBSs within the system.

[0039]

[0040] in, Indicates the 1st, 2nd... Nth P CRE bias values ​​for each base station; Indicates system throughput The maximum is the 1st, 2nd...Nth. P The CRE bias value of each base station.

[0041] Step 2, Initialization Phase: Construct the agent's experience pool and deep reinforcement learning neural network, including: target network and prediction neural network, and initialize the cell range expansion bias value of the pico base station PBS in the system.

[0042] The intelligent agent refers to dividing all PBSs within the system into K distinct regions, where all N in the k-th region... k Each PBS is considered as a single agent, sharing a single CRE bias value; where 1 ≤ k ≤ K; in this case, the agent is represented as follows:

[0043] {agent1,…,agent k} (14)

[0044]

[0045] It is obvious that when K=1, N1=N P When K = N, all PBSs within the system share the same dynamically adjusted CRE bias value;P N1 = ... = N K When =1, each PBS in the system has a unique, personalized CRE bias value.

[0046] Step 3, Interactive Training Phase: The PBS interacts with the wireless communication system to store the experience pool, and randomly samples from the experience pool to train the target network and the prediction neural network, updating the neural network parameters.

[0047] In step 3, the interaction between PBS and the wireless system during the interactive training phase refers to the environment generating information describing the system's state, i.e., the state, within the reinforcement learning system. The agent observes the state and uses this information to select an action. By executing the next action, the agent influences the environment, and the environment provides a reward based on the agent's action. When the cycle of "state → action → reward" is completed, an interaction between the agent and the environment is finished. Specifically, this invention includes the following steps:

[0048] Step 3.1: Set the number of rounds n episode =1.

[0049] Step 3.2: Set the time step n step =0, synchronous number n same_step =0, recent throughput C last =0, CRE bias value σ = σ0.

[0050] Step 3.3: State information observed by a single agent at time t

[0051] The number of PEUEs within the PBS coverage extension area and the interference status of the PEUEs are used as state information. Specifically, for the k-th agent, the state information is o. k It consists of two parts: local information and global information.

[0052] Part 1 Local Information: The number of PEUEs and the interference experienced by each PEUE within the PBS coverage extension area of ​​the k-th agent.

[0053]

[0054] in, This indicates that the existence of the CRE bias value is related to the number of PEUEs of the i-th PBS in the k-th agent; This represents the mean of the disturbance intensity experienced by the PEUE associated with the i-th PBS in the k-th agent.

[0055] Part Two: Global Information: The number of PEUEs within the coverage extension range of all PBSs in the range of the other K-1 agents, and the interference situation of the PEUEs.

[0056]

[0057] in, This indicates that the existence of the CRE bias value is related to the number of all PEUEs within the PBS extension range of the l-th agent; This represents the mean value of the disturbance intensity experienced by all PEUEs associated with the PBS in the l-th agent;

[0058] At this point, the state information of all agents in the reinforcement learning system constitutes the joint state information o:

[0059] o={o 1 ,…,o K} (18)

[0060] Step 3.4: Transfer the state information of a single agent. The agent ID is input into the prediction neural network to obtain the individual action value function value with different CRE bias values ​​in the action space, and an ε-greedy policy is used to select actions.

[0061] The agent's output is the CRE bias value of the PBS within the represented region. Therefore, the action a of the k-th agent... k for:

[0062]

[0063] Where, σ k This represents the CRE bias value; the CRE bias values ​​of all PBSs constitute a set. Let be the action space of the k-th agent, representing the set of all possible CRE bias values. The actions of all agents in a multi-agent system constitute the joint action 'a':

[0064] a={a 1 ,…,a K} (20)

[0065] Step 3.5: A single agent performs an action. That is, after setting a new CRE bias value, obtain the status information. and rewards for environmental feedback The reward is defined as a correlation function of the difference between the throughput of all users within the agent when the CRE bias is dynamically adjusted based on a multi-agent reinforcement learning algorithm and the system throughput without CRE. For the k-th agent, the reward is specifically defined as follows:

[0066]

[0067] Among them, C i,0 This represents the sum of throughput of users associated with the i-th base station when CRE technology is not used. This represents the difference between the total throughput of the PBS in the k-th agent when the CRE bias value is dynamically adjusted and when CRE technology is not used; C0-C 0,0 The input throughput of the MBS in a system with dynamically adjusted CRE bias is the difference between that of a system without CRE; μ is the reward adjustment factor; τ k This is the contribution factor of the k-th agent to the change in MBS throughput.

[0068] Step 3.6: After all agents have completed steps 3.1-3.5, calculate the global reward r. t And integrate to obtain joint action a t Joint state information t and the next joint state information o t+1 and with (o t ,a t ,r t ,o t+1 Stored in the experience pool in the form of ) The global reward r t The calculation is as follows:

[0069]

[0070] Step 3.7: If the number of records in the experience pool is less than the sampling quantity N batch Let n step =n step +1, proceed to step 3.3; if the number of records in the experience pool is greater than or equal to the sampling number N. batch Then proceed to step 3.8.

[0071] Step 3.8, from the experience pool A small batch of N is randomly selected from the data. batch Several empirical sample data points were used to train the predictive neural network.

[0072] Step 3.9: Calculate the objective function y t We construct a neural network loss function L(θ), and then update the parameters using the Mini-Batch Gradient Descent (MBGD) method.

[0073]

[0074]

[0075] Step 3.10, Update time step: n step =nstep +1 and calculate system throughput:

[0076] Step 3.11, if C t =C last Let n same_step =n same_step +1, otherwise n same_step =0, C last =C t .

[0077] Step 3.12, every N update Step-by-step update of target network parameters: θ - =θ.

[0078] Step 3.13, n episode =n episode +1.

[0079] Step 3.14, if n same_step ≤N same_step And n step ≤N step If yes, proceed to step 3.3; otherwise, proceed to step 3.15.

[0080] Step 3.15, if n episode ≤N episode If yes, proceed to step 3.2; otherwise, proceed to step 4.

[0081] Step 4, Output Stage: Set the CRE bias value of the picobase PBS in the output system.

[0082] Furthermore, step 2, the initialization phase, specifically includes: initializing the maximum number of training rounds N. episode Maximum number of training steps per round N step The maximum number of identical steps in each round, N same_step Learning rate r a Discount factor r d Greedy strategy ε, experience pool Predicting neural network weight parameters θ, target network weight parameters θ - Initialize the CRE bias value σ0 of the PBS within the system.

[0083] Furthermore, in step 3.3 and The calculation method is as follows:

[0084]

[0085]

[0086] in, This represents the co-frequency neighboring interference experienced by the u-th PEUE associated with the i-th PBS in the k-th agent; This represents the interference intensity experienced by the u-th PEUE associated with the i-th PBS in the l-th agent. This represents the number of PEUEs of the i-th PBS in the l-th agent, 1≤i≤N k

[0087] Furthermore, in step 3.5, τ k The calculation method is as follows:

[0088]

[0089] This invention provides a neural network training and update method based on the Value Decomposition Networks (VDN) algorithm. VDN is a classic method among value decomposition methods, employing a centralized training and distributed execution architecture. This algorithm considers the joint action value function Q... totoal The individual action value function Q of each agent m There exists a linear additive relationship between them, using Q m Q is represented by the sum of its components. totoal :

[0090]

[0091] Among them, Q totoal (o t ,a t ) represents the joint action value function of all intelligent agents. N represents the value function of a single action of a single agent. agent N represents the number of agents. In this invention, N agent =K, at this time the joint action value function Q totoal Represented as:

[0092]

[0093] In the VDN method, a single agent can be viewed as an approximate "Deep Q-Learning (DQN) network structure". In this case, the objective value TD Target and the loss function of the VDN algorithm are defined as follows:

[0094]

[0095]

[0096] Where, r t It is the global reward at time t; θ represents the joint action value function output by the target neural network. - q represents the parameters of all target neural networks associated with the joint action value function; totoal (o t ,a t ;θ) represents the joint action value function of the predictive neural network output, and θ represents all parameters of the predictive neural network associated with the joint action value function. With q totoal (o t ,a t ;θ) can be obtained by fusing the individual action value functions output by the target neural network and the prediction neural network respectively according to the formula:

[0097]

[0098]

[0099] in, This represents the target neural network parameters for the m-th agent. θ m Let θ = {θ1, ..., θ2} represent the parameters of the predictive neural network for the m-th agent. K}

[0100] The present invention provides a method for inter-small-scale interference coordination based on multi-agent reinforcement learning, which has the following advantages: By utilizing a multi-agent reinforcement learning algorithm, the PBS (Small Partitions) within the system are divided into multiple agents. An algorithm with centralized training and distributed execution is adopted, which allows the agents to adjust the CRE (Corrective Reaction) bias value of a small range or a single PBS individually according to the distribution of users in the agent and other agents and the interference they are subjected to. This avoids the problems of bias value space size expansion and high optimization difficulty. By comprehensively considering global information, the system throughput is effectively improved and the throughput of individual users within the system is also improved. Attached Figure Description

[0101] Figure 1 This is a flowchart of an embodiment of the present invention;

[0102] Figure 2 This is a model diagram of an embodiment of the present invention;

[0103] Figure 3 Comparison of system throughput under different CRE configuration methods;

[0104] Figure 4 The cumulative distribution function of user throughput within the system under different CRE settings. Detailed Implementation

[0105] To better understand the purpose, structure, and function of this invention, the following detailed description of an inter-cell interference coordination method based on multi-agent reinforcement learning is provided in conjunction with the accompanying drawings.

[0106] The following is the inter-cell interference coordination scheme based on multi-agent reinforcement learning proposed in this invention. The specific steps for the entire process are as follows:

[0107] Step 1: System Modeling Stage: Construct a Macro-Pico heterogeneous network interference scenario, consider the downlink transmission link of a two-layer heterogeneous network consisting of MBS and PBS, and propose the problem of maximizing system throughput.

[0108] The system modeling phase includes the following specific steps:

[0109] Step S101: As Figure 2 As shown, assume there is 1 MBS and N in the system. P =4 PBS, N U = 200 users are randomly and uniformly distributed throughout the system, where each PBS contains N users within its coverage area. UP = 35 randomly distributed users. The MBS cell radius is 289m, the PBS cell radius is 40m, the minimum distance between MBS and PBS is 75m, the minimum distance between PBS is 40m, the minimum distance between MBS and users is 35m, and the minimum distance between PBS and users is 10m.

[0110] Step S102: Propose a cell range expansion model, as follows:

[0111] Users use the Max-RSRP strategy to select associated base stations, meaning each user is associated with the base station providing the highest observed RSRP. To effectively broaden the service range of the PBS (Balanced Base Station) and encourage more users, especially those located at the boundary between macrocells and picocells, to select the PBS, the CRE (Continuous Rejection) technique adds a positive bias value to each PBS. In this case, the user's associated base station is:

[0112]

[0113] Where i and u represent the i-th base station and the u-th user, respectively, i = 0 refers to MBS, and 1 ≤ i ≤ N. P "Time" refers to PBS. That is, the associated base station number of the u-th user when using CRE technology. σ i This represents the CRE bias value of base station i. Since MBS does not set a CRE bias value, σ0 = 0. RSRP iu The RSRP received by the u-th user from the i-th base station is calculated as follows:

[0114] RSRP iu =P i g iu

[0115]

[0116] Among them, P i G represents the transmit power of the i-th base station. iu This represents the channel gain between the u-th user and the i-th base station. In this embodiment, a Rayleigh fading channel is used, and the distance-related path loss of MBS and PBS is calculated as follows:

[0117] L d,M =140.7 + 36.7 × log 10 (R)dB

[0118] L d,P =128.1 + 36.7 × log 10 (R)dB

[0119] Step S103: Propose an almost blank subframe model, as follows:

[0120] The SINR at the u-th user served by PBS is represented in different subframes as follows:

[0121]

[0122]

[0123] in, This indicates the SINR of the u-th user associated with PBS during the ABS subframe. This indicates the SINR of the u-th user associated with PBS during a non-ABS subframe. This represents the associated base station number of the u-th user. Indicates association with the first Interference from neighboring cells on the same frequency experienced by the u-th user of a PBS base station. N0 is the noise power, and the thermal noise density in this embodiment is -174dBm / Hz.

[0124] The SINR of the u-th user served by MBS is represented in different subframes as follows:

[0125]

[0126]

[0127] in, This indicates the SINR of the u-th user associated with MBS during the ABS subframe. This indicates the SINR of the u-th user associated with the MBS during a non-ABS subframe. Since the MBS does not transmit any valid data during ABS subframes, the SINR is represented as zero.

[0128] In this embodiment, the ABS ratio β is defined as 0.5.

[0129] Step S104: Propose a model to maximize system throughput, as follows:

[0130] This invention assumes that each user receives almost equal amounts of resources; therefore, the throughput of the u-th user can be calculated using Shannon's formula:

[0131]

[0132] Where W is the system bandwidth, W = 20MHz. and These represent the 1st, 2nd, and 3rd user-related terms associated with the uth user. The number of users served by each base station during ABS and non-ABS periods. This varies depending on the type of base station associated with the user, when the user is associated with a PBS. equal When a user associates MBS equal Similarly, when a user associates with PBS equal When a user associates MBS equal

[0133] The user throughput associated with the i-th base station can be calculated as follows:

[0134]

[0135]

[0136]

[0137] System throughput is the total throughput of all cells within the system, which is equal to the sum of the throughput of all users in the system. The objective of this invention is to maximize the sum of the throughput of all users in the system by jointly optimizing the CRE bias values ​​of all PBSs within the system.

[0138]

[0139] Step 2, Initialization Phase: Construct the agent's experience pool and deep reinforcement learning neural network, including: target network and prediction neural network, and initialize the CRE bias value of PBS in the system.

[0140] The initialization phase specifically includes the following steps:

[0141] Step S201: Initialize the maximum number of training rounds N episode =500, maximum number of training steps per round N step =100, Maximum number of identical steps per round N same_step =30, learning rate r a =0.0005, discount factor r d =0.88, Greedy Strategy ε=0.01, Experience Pool Capacity 10000, N batch =32, predicting neural network weight parameters θ, target network weight parameters θ - Initialize the CRE bias value of the PBS within the system.

[0142] Step 3, Interactive Training Phase: The PBS interacts with the wireless communication system to store the experience pool, and randomly samples from the experience pool to train the target network and the prediction neural network, updating the neural network parameters.

[0143] The interactive training phase includes the following steps:

[0144] Step S301: Set the number of rounds n episode =1.

[0145] Step S302: Set time step n step =0, synchronous number n same_step =0, recent throughput C last =0, CRE bias value σ = σ0.

[0146] Step S303: State information observed by a single agent k at time t Status information It consists of two parts: local information and global information.

[0147] In this embodiment, the intelligent agent refers to dividing all PBSs within the system into K distinct regions, where all N in the k-th (1≤k≤K) region... k Each PBS is considered a single agent, sharing the same CRE bias value. In this case, the agent is represented as follows:

[0148] {agent1,…,agent K}

[0149]

[0150] This embodiment considers one setting as VDN: K=2, N1=N2=2, that is, all PBS are divided into 2 regions, and 2 PBS in each region are regarded as 1 agent, and the same dynamic CRE bias value is set; another setting is considered as VDN: K=4, N1=…=N4=1, that is, all PBS are divided into 4 regions, and each PBS is regarded as 1 agent, and a separate dynamic CRE bias value is set.

[0151] In this embodiment, for the k-th agent, the local information refers to the number of PEUEs and the interference situation of each PEUE within the PBS coverage extension area of ​​the agent, specifically expressed as:

[0152]

[0153] in, This indicates that the existence of the CRE bias value is related to the number of PEUEs of the i-th PBS in the k-th agent; This represents the mean disturbance intensity experienced by the PEUE of the i-th PBS associated with the k-th agent. The specific calculation method is as follows:

[0154]

[0155] In this embodiment, for the k-th agent, the global information refers to the number of PEUEs within the coverage extension range of all PBSs and the interference situation of the PEUEs within the range of the other K-1 agents, specifically expressed as:

[0156]

[0157] in, This indicates that the existence of the CRE bias value is related to the number of all PEUEs within the PBS extension range of the l-th agent; This represents the average interference intensity experienced by all PEUEs associated with the PBS in the l-th agent. The specific calculation method is as follows:

[0158]

[0159] in, This represents the interference intensity experienced by the u-th PEUE associated with the i-th PBS in the l-th agent. This represents the number of PEUEs of the i-th PBS in the l-th agent.

[0160] Step S304: Transfer the state information of a single agent The agent ID is input into the prediction neural network to obtain the individual action value function value with different CRE bias values ​​in the action space, and an ε-greedy policy is used to select actions.

[0161] In this embodiment, for the k-th agent, the action a k Specifically, it is expressed as follows:

[0162]

[0163] Where, σ k This represents the CRE bias value; the CRE bias values ​​of all PBSs constitute a set. Let be the action space of the k-th agent, and let represent the set of all possible CRE bias values.

[0164] Step S305: A single agent performs an action. That is, after setting a new CRE bias value, obtain the status information. and rewards for environmental feedback

[0165] In this embodiment, for the k-th agent, the reward That is, the reward r obtained by the k-th agent at time t. k Specifically, it is expressed as:

[0166]

[0167] in, This represents the difference between the dynamically adjusted CRE bias value and the sum of the throughput of the PBS in the k-th agent without CRE technology; C0-C 0,0 The input throughput of the MBS in a system with dynamically adjusted CRE bias is the difference between that of a system without CRE; μ is the reward adjustment factor; τ k It is the contribution factor of the k-th agent to the change in MBS throughput.

[0168] Step S306: After all agents have completed the above process, calculate the global reward r. t And integrate to obtain joint action a t Joint state information t and the next joint state information o t+1 and with (o t ,a t ,r t ,o t+1 Stored in the experience pool in the form of )

[0169] In this embodiment, the global reward r t The calculation method is as follows:

[0170]

[0171] In this embodiment, the combined action a t Joint state information o t and the next joint state information o t+1 Each of these refers to a set consisting of the actions, states, and next states of all agents within the system, specifically represented as follows:

[0172]

[0173]

[0174]

[0175] Step S307: If the number of records in the experience pool is less than the sampling quantity N batch Let n step =n step +1, proceed to step S303; if the number of records in the experience pool is greater than or equal to the sampling number N. batch Then proceed to step S308.

[0176] Step S308: From the experience pool A small batch of N is randomly selected from the data. batch Several empirical sample data points were used to train the predictive neural network.

[0177] Step S309: Calculate the objective function y t We construct a neural network loss function L(θ), and then update the parameters using the MBGD method.

[0178] In this embodiment, the objective function y t The specific representations of the neural network loss function L(θ) and MBGD update parameters are as follows:

[0179]

[0180]

[0181]

[0182] Step S310: Update time step: n step =n step +1 and calculate system throughput:

[0183] Step S311: If C t =C last Let n same_step =n same_step +1, otherwise n same_step =0, C last =C t .

[0184] Step S312: After N update After rounds of iteration, the predictive neural network assigns weight parameters θ to the target network θ. - It is used to update the target network weight parameters.

[0185] Step S313: n episode =n episode +1.

[0186] Step S314: If n same_step ≤N same_step And n step ≤N step If the condition is met, proceed to step S303; otherwise, proceed to step S315.

[0187] Step S315: If n episode ≤N episode If the condition is met, proceed to step S302; otherwise, proceed to step 4.

[0188] Step 4, Output Stage: Set the CRE bias value of the PBS in the output system.

[0189] Experimental results:

[0190] Figure 3 This display compares system throughput under different CRE settings. The x-axis represents the number of users N included within the coverage area of ​​each PBS, while keeping the total number of users constant. UP The left y-axis corresponds to a bar chart, representing the difference in system throughput under different CRE settings compared to the system throughput without CRE. As shown in the figure, using CRE significantly increases system throughput. Furthermore, in terms of improving system throughput, the proposed method of dynamically adjusting the CRE bias value based on multi-agent reinforcement learning is better than the method in the 3GPP protocol that fixes the CRE bias value at 5dB. Moreover, dynamically adjusting the CRE bias value individually for each PBS (VDN, K=4) is no worse than dynamically adjusting the CRE bias value within a small range (VDN, K=2).

[0191] Figure 4 The display shows N UP When CRE = 35, the cumulative distribution function (CDF) of user throughput in the system under different CRE settings is shown in the figure. As can be seen from the figure, the user throughput distribution using CRE technology is significantly better than that without CRE technology, and the proportion of low-throughput users is significantly reduced; the user throughput distribution using dynamically adjusted CRE bias value is better than that using fixed CRE bias value.

[0192] It is understood that the present invention has been described through some embodiments, and those skilled in the art will recognize that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the invention. Furthermore, under the teachings of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the invention. Therefore, the present invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the protection scope of the present invention.

Claims

1. A method for inter-cell interference coordination based on multi-agent reinforcement learning, characterized in that, Includes the following steps: Step 1, System Modeling Stage: Construct a Macro-Pico heterogeneous network interference scenario as a wireless communication system, consider the downlink transmission link of a two-layer heterogeneous network consisting of macro base stations MBS and pico base stations PBS, and propose the problem of maximizing system throughput; Step 1.

1. Consider that there is 1 macro base station MBS and Step 1.

2. Consider that there is 1 macro base station MBS and Step 1.

3. Consider that there is 1 macro base station MBS and Step 1.

4. Consider that there is 1 macro base station MBS and Step 1.

5. Consider that there is 1 macro base station MBS and Step 1.2: Propose a cell range expansion model, as follows: In heterogeneous networks, users select their associated base station using the Maximum Reference Signal Received Power (Max-RSRP) strategy, meaning each user is associated with the base station providing the highest observed RSRP. Cell Range Extension (CRE) technology adds a positive bias value to each picocell's Basic Segmentation Base (PBS), at which point the user's associated base station is: (1) in, Indicates the first One base station, Indicates the first One user, The term "MBS" refers to a macro base station. The term refers to the picocell PBS; That is, when using CRE technology, the first The associated base station sequence number of each user. ; Indicates base station The CRE bias value, since MBS does not set a CRE bias value, therefore ; Representing the The user received from the first The RSRP of each base station is calculated as follows: (2) (3) in, Indicates the first The transmit power of each base station, This indicates the transmit power of the macro base station (MBS). This indicates the transmit power of the picocell PBS. Indicates the first The user and the first Channel gain between base stations; Step 1.3: Propose an almost blank subframe ABS model, as follows: Using the combined temporal interference coordination technique (ATC) ABS, subframes are divided into two types: non-ABS subframes and ABS subframes. During ABS subframes, the macro base station (MBS) does not transmit any valid data, and the pico base station (PBS) schedules extended pico user PEUEs that are severely interfered with by the macro base station (MBS). During non-ABS subframes, the MBS provides normal service to macro user MUEs, and the pico base station (PBS) only schedules pico center user PCUEs. Users experience different degrees of interference during non-ABS and ABS periods; therefore, the subframes served by the pico base station (PBS) are... The SINR at each user location during different subframes is represented as follows: (4) (5) in, Indicates the number associated with PBS SINR of each user during the ABS subframe. Indicates the number associated with PBS SINR of a user during a non-ABS subframe; This refers to the transmission power of the associated base station. Indicates the first Channel gain between a user and the associated base station; Indicates association with the first The first PBS base station Interference from neighboring cells on the same frequency suffered by individual users; It is noise power; The service provided by MBS The SINR at each user location during different subframes is represented as follows: (6) (7) in, Indicates the number associated with MBS SINR of each user during the ABS subframe. Indicates the number associated with MBS The SINR of a user during a non-ABS subframe is represented as zero because MBS does not transmit any valid data during ABS subframes. This represents the transmit power of the 0th base station, i.e., the transmit power of the MBS. ; Indicates the first The channel gain between the nth user and the 0th base station, i.e. the nth Channel gain between individual users and MBS; Definition of the first The number of users served by each base station during the ABS subframe is The number of users served during non-ABS subframes is Define ABS ratio It is the ratio of the number of ABS subframes to the total number of subframes within a single frame. The following conditions must be met: (8) Step 1.4: Propose a model to maximize system throughput, as follows: Each user is given an equal amount of resources, therefore the calculation is based on Shannon's formula. The throughput per user is: (9) in, It is system bandwidth. Indicates the first The first user associated with Number of users served by each base station during the ABS period Indicates the first The first user associated with The number of users served by each base station during non-ABS periods; depending on the type of base station associated with the user, when the user is associated with PBS. equal When a user associates with MBS equal Similarly, when a user associates with PBS... equal When a user associates with MBS equal ; Related to the first The throughput of all users at each base station is calculated as follows: (10) (11) (12) in, Indicates the period of ABS The sum of the throughput of users within each base station; Indicates the period outside of ABS. The sum of the throughput of users within each base station; System throughput is the total throughput of all cells within the system, and also the sum of the throughput of all users within the system. By jointly optimizing the CRE bias values ​​of all PBSs within the system, the sum of the throughput of all users within the system can be maximized. (13) in, Indicates the 1st, 2nd, ... 1st CRE bias values ​​for each base station; Indicates system throughput The largest is the 1st, 2nd... 1st The CRE bias value of each base station; Step 2, Initialization Phase: Construct the agent's experience pool and deep reinforcement learning neural network, including: target network and prediction neural network, and initialize the cell range extension CRE bias value of the pico base station PBS in the system; The intelligent agent refers to the agent that divides all PBSs within the system into... A different region, the first All within the region Each PBS is considered a single agent, sharing a single CRE bias value; where At this time, the intelligent agent It is expressed as follows: (14) (15) Step 3, Interactive Training Phase: The experience pool is stored through the interaction between the pico base station PBS and the wireless communication system, and random samples are taken from the experience pool to train the target network and the prediction neural network, and update the neural network parameters. Step 3.1: Set the number of rounds ; Step 3.2: Set the time step Phase number Recent throughput CRE bias value ; Step 3.3, a single agent in Observed state information at all times ; The status information It consists of two parts: local information and global information. The local information refers to the number of user PEUEs and the interference situation of each PEUE within the coverage extension area of ​​each PBS in the intelligent agent. (16) in, This indicates that due to the existence of the CRE bias value, where Related to the first The first intelligent agent Number of PEUEs per PBS; Indicates association with the first The first intelligent agent The mean of the interference intensity experienced by the PEUE of each PBS, where The global information refers to other... The number of PEUEs within the coverage extension range of all PBSs within the scope of each agent and the interference situation of PEUEs; (17) in, This indicates that the existence of the CRE bias value is related to the first... The number of all PEUEs within the PBS extension range of each agent; Indicates association with the first The mean of the disturbance intensity experienced by all PEUEs of PBS in each agent; Step 3.4: Transfer the state information of a single agent. Together with the agent ID, the prediction neural network is used to obtain the individual action value function value for different CRE bias values ​​within the action space, and then... Greedy strategy for selecting actions ; The action It refers to the first CRE bias values ​​for each agent: (18) in, This represents the CRE bias value; the CRE bias values ​​of all PBSs constitute a set. ; For the first The action space of an agent represents the set of all possible CRE bias values; Step 3.5: A single agent performs an action. That is, after setting a new CRE bias value, obtain the status information. and rewards for environmental feedback ; The reward This refers to the correlation function of the difference between the throughput of all users in the agent when the CRE bias value is dynamically adjusted based on the multi-agent reinforcement learning algorithm and the system throughput without using CRE technology; for the th The reward for each intelligent agent is specifically defined as follows: (19) in, Indicates the first The sum of the throughput of all users associated with each base station when CRE technology is not used. This indicates the difference between dynamically adjusting the CRE bias value and not using CRE technology. The difference in the total throughput of PBS among the agents; This represents the difference in throughput between the dynamically adjusted CRE bias value and the system's MBS throughput without using CRE technology. It is a reward adjustment factor; It is the first The contribution factor of each agent to the change in MBS throughput; Step 3.6: After all agents have completed steps 3.1-3.5, calculate the global reward. And integrate to achieve joint action Joint state information and the next joint state information and with Stored in the experience pool in the form of ; global rewards The calculation is as follows: (20) Step 3.7, if the experience pool The number of records in the sample is less than the number of samples. Then let Proceed to step 3.3; if the number of records in the experience pool is greater than or equal to the number of samples. Then proceed to step 3.8; Step 3.8, from the experience pool A small batch was randomly selected from the middle. A number of empirical sample data were used to train the predictive neural network. Step 3.9: Calculate the objective function Constructing the neural network loss function Then the parameters are updated using the mini-batch gradient descent (MBGD) method. Step 3.10, Update Time Step: And calculate the system throughput: ; Step 3.11, if Then let ,otherwise , ; Step 3.12, every Step-by-step update of target network parameters: ; Step 3.13 ; Step 3.14, if and If yes, proceed to step 3.3; otherwise, proceed to step 3.

15. Step 3.15, if If yes, proceed to step 3.2; otherwise, proceed to step 4. Step 4, Output Stage: Setting the CRE bias value of the PBS (PicoSat) within the output system; Step 2, the initialization phase, specifically includes: initializing the maximum number of training rounds. Maximum number of training steps per round Maximum number of identical steps per round Learning rate Discount Factor Greedy strategy Experience Pool Predicting neural network weight parameters Target network weight parameters Initialize the CRE bias value of the PBS within the system. ; In step 3.3 and The calculation method is as follows: (21) (22) in, Indicates association with the first The first intelligent agent The first PBS Interference from neighboring cells on the same frequency experienced by each PEUE; Indicates association with the first The first intelligent agent The first PBS The interference intensity experienced by each PEUE Indicates the first The first intelligent agent Number of PEUEs per PBS ; In step 3.5 The calculation method is as follows: 。