Joint Design of User Association and Hybrid Beamforming Method and System for sub-THz UDN using Multi-Agent Deep Reinforcment Learning

KR103002965B1Active Publication Date: 2026-08-11KOREA ADVANCED INST OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
KR1020220175665
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-12-15
Publication Date
2026-08-11
Estimated Expiration
2042-12-15

Smart Images

  • Figure 112022135051542-PAT00190_ABST
    Figure 112022135051542-PAT00190_ABST
Patent Text Reader

Abstract

An interference control and hybrid beamforming method and system applying multi-agent deep reinforcement learning for multiple users are presented. The interference control and hybrid beamforming method applying multi-agent deep reinforcement learning for multiple users proposed in the present invention includes the steps of: performing multi-agent reinforcement learning using expected interference and antenna gain information based on a gain table designed using Channel State Information (CSI) of all user terminals; searching for pairs of analog beamforming matrices corresponding to links that maximize antenna gain and minimize interference between user terminals through multi-agent reinforcement learning; applying a Signal to Leakage Plus Noise Ratio (SLNR) maximization technique that minimizes interference between user terminals based on the links; and optimizing transmission power for each link based on iterative waterfilling.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to an interference control and hybrid beamforming method and system applying multi-agent deep reinforcement learning for multiple users. Background Technology

[0002] Beamforming technology is a technology that uses multiple antennas to generate directional transmission and reception signals. In 5G and beyond 5G systems, high levels of path attenuation are experienced because ultra-high frequency channels are applied. However, since the channel frequency is very high, short-wavelength signals can be transmitted. Therefore, path attenuation can be overcome by designing a high-density array of antennas with very narrow spacing and applying beamforming technology to form a narrow beam.

[0003] When applying conventional digital antenna-based beamforming technology using multiple antennas, it requires as many RF (Radio Frequency) chains as there are antennas, resulting in high hardware complexity and power consumption. To overcome this, hybrid beamforming technology is gaining attention for adopting a narrow analog beam pattern utilizing multiple antennas while applying a much smaller number of RF chains compared to antennas.

[0004] Hybrid beamforming can obtain the spatial multiplexing gain of digital beamforming and the antenna beamforming gain of analog beamforming simultaneously. The analog beamforming block consists only of phase shifters to reduce circuit complexity.

[0005] To achieve high data transmission rates, a wide bandwidth must be secured. This can be achieved by utilizing ultra-high frequency channels, such as sub-THz (tera hertz), which are considered by beyond 5G mobile communication systems. However, high-frequency channels suffer from high path attenuation, which inevitably limits the transmission of signals to a narrow coverage area. Consequently, high data transmission rates can only be achieved by configuring a high-density network that is concentrated in a small space.

[0006] In contrast to conventional base station-centric networks, where users can only transmit and receive signals from a single base station, resulting in limited freedom in optimizing data transmission rates, user-centric networks allow each user to connect to multiple base stations, enabling connections to up to the number of base stations corresponding to the user's RF chain. This allows for the design of networks that achieve optimal transmission rates based on a high degree of freedom.

[0007] In the prior art [1], interference and power were controlled by applying a single-agent deep reinforcement learning technique that does not utilize channel information. However, because a single-agent-based deep reinforcement learning technique was applied, it became difficult to apply as the number of base stations and users in a high-density network increased, and spatial multiplexing gains could not be obtained because a single antenna system was applied to each user. In contrast, the present invention introduces a multi-agent deep reinforcement learning technique for multiple users and a hybrid beamformer technique utilizing it to increase power efficiency and achieve a high data transmission rate.

[0008] In the prior art [2], beamforming interference between users is controlled by applying a multi-agent deep reinforcement learning technique that incorporates channel information. However, as with the prior art [1], spatial multiplexing gain cannot be obtained because a single antenna system is applied to the user, and it is difficult to apply to multiple antenna systems because a hybrid beamforming system is not applied, and it is difficult to obtain a high data transmission rate and a high maximum transmission speed per unit area.

[0009] In the prior art [3], a technique was proposed to greedily select beam pairs with high channel coefficients based on a hybrid beamforming system that applies limited channel information. However, since beam pairs are greedily selected without applying artificial intelligence technology, it is difficult to optimally control residual interference between users, and consequently, it is difficult to obtain the optimal data transmission rate.

[0010] In conventional technology, beamforming was designed based on a single base station rather than multiple base stations, or for users based on a single antenna, in order to reduce the complexity of the problem. However, if only interference between users is considered for a single base station, interference caused by beams from other base stations cannot be controlled by this technology. Furthermore, if only users based on a single antenna are considered, only one RF chain can be used, making it impossible to obtain multiplexing gain per RF chain.

[0011] In the prior art [4], machine learning is used to find the optimal beamforming vector, and a vector quantization module is applied to quantize vector channels by dividing them into real and imaginary parts. Subsequently, the quantized vector channels are reassembled into the codeword of the beamforming vector. While this method may be convenient for obtaining a beamforming vector from a vector-shaped channel, if the dimension of the MIMO channel is a matrix of 2 or more, it is difficult to proceed with input after vectorization due to the sparsity and correlation of the channels. Furthermore, for channels of that form, if vectorization is performed and the real and imaginary parts are separated, the length of the vector input to the neural network becomes very large, increasing the size of the neural network and making it impossible to handle all parameters of the high-density network. Finally, since it assumes a user with a single antenna, multiplexing gain per RF chain cannot be obtained.

[0012] The prior art [5] proposed a user scheduling method capable of power allocation and transmitting signals to at least one user in a multiple antenna downstream system. Since the patent has multiple antennas, multiplexing gain per RF chain can be obtained. However, it assumes a single base station rather than multiple base stations, and since multiple base stations do not perform cooperative transmission, an improvement in data transmission rate with an increase in the number of base stations cannot be expected. Prior art literature

[0013] [1] FB Mismar, BL Evans and A. Alkhateeb, “Deep reinforcement learning for 5G networks: Joint beamforming power control and interference coordination”, IEEE Trans. Commun. , vol. 68, no. 3, pp. 1581-1592, Mar. 2020.

[0014] [2] J. Ge, Y.-C. Liang, J. Joung and S. Sun, "Deep reinforcement learning for distributed dynamic MISO downlink-beamforming coordination", IEEE Trans. Commun. , vol. 68, no. 10, pp. 6070-6085, Oct. 2020.

[0015] [3] G. Kwon and H. Park, “Joint user association and beamforming design for millimeter wave UDN with wireless backhaul”, IEEE J. Sel. Areas Commun. , vol. 37, no. 12, pp. 2653-2668, Dec. 2019.

[0016] [4] Korean Registered Patent No. 10-2168650 (October 15, 2020)

[0017] [5] Korean Registered Patent No. 10-1900607 (2018.09.13) The problem to be solved

[0018] The technical problem that the present invention aims to solve is to provide an interference control and hybrid beamforming method and system applying multi-agent deep reinforcement learning for multiple users to address the issue in mobile communication environments where multiple base stations and multiple users exist, where various types of interference, such as inter-base station interference and inter-beam interference, exist, making it impossible to numerically find the optimal solution, and where conventional machine learning is also unusable due to the lack of labels for training.

[0019] More specifically, the present invention proposes efficient interference control through multi-agent deep reinforcement learning operating in a user-centric network where multiple base stations of a more complex configuration can support a single user, by introducing a multi-agent deep reinforcement learning technique for multiple users and a hybrid beamformer technique utilizing the same to increase power efficiency and achieve a high data transmission rate. means of solving the problem

[0020] In one aspect, the interference control and hybrid beamforming method for multiple users proposed in the present invention, which applies multi-agent deep reinforcement learning, comprises the steps of: performing multi-agent reinforcement learning using expected interference and antenna gain information based on a gain table designed using Channel State Information (CSI) of all user terminals; searching for pairs of analog beamforming matrices corresponding to links that maximize antenna gain and minimize interference between user terminals through multi-agent reinforcement learning; applying a Signal to Leakage Plus Noise Ratio (SLNR) maximization technique that minimizes interference between user terminals based on the links; and optimizing transmission power for each link based on iterative waterfilling.

[0021] The step of performing multi-agent reinforcement learning using expected interference and antenna gain information based on a gain table designed using the CSI of the entire user terminal described above involves designing the Signal to Interference Plus Noise Ratio (SINR) for each SBS, predicted based on the expected interference and antenna gain during the multi-agent reinforcement learning process for a plurality of agents corresponding to each SBS for the multi-agent reinforcement learning, as the reward for each agent.

[0022] The step of performing multi-agent reinforcement learning using expected interference and antenna gain information based on a gain table designed using the CSI of all user terminals described above learns a link configuration that maximizes the data transmission rate through a trial-and-error method from a predetermined number of episodes and time intervals within the episodes for all agents.

[0023] The step of searching for analog beamforming matrix pairs corresponding to links that maximize antenna gain and minimize interference between user terminals through the above multi-agent reinforcement learning defines the link selected so far from the set of candidate links for each agent as the agent's state, defines the link selected in the current time interval as the agent's action, and defines the sum of the data transmission rates obtained from the links selected so far of the SBS associated with the agent in the corresponding time interval for each agent as the reward.

[0024] The step of searching for analog beamforming matrix pairs corresponding to links that maximize antenna gain and minimize interference between user terminals through the above multi-agent reinforcement learning involves constructing the state of the agent and the action of the agent into binary vectors to search for links between SBS and user terminals that minimize interference and maximize antenna gain for the selected links.

[0025] The step of applying an SLNR maximization technique to minimize interference between user terminals based on the above link involves applying a hybrid beamforming matrix to each SBS and applying an analog beamforming matrix without baseband beamforming to the user terminal, and when multiple SBSs simultaneously transmit signals to user terminals on the links connected to them, calculating the sum of the data transmission rates of the entire network based on the SINR of all links resulting from interference between links connected to user terminals, interference between user terminals caused by an SBS connected to a user terminal opening a link with another user terminal to transmit a stream, and interference received from other SBSs, and optimizing the beamforming matrix pairs and link configurations between SBSs and user terminals for maximizing the data transmission rate using the sum of the data transmission rates.

[0026] The step of applying an SLNR maximization technique to minimize interference between user terminals based on the above link transmits only a portion of the entire channel according to the channel gain between the SBS and the user terminal to prevent signal overhead for the MBS in an Ultra Massive Multiple Input Multiple Output (UM-MIMO) network, and reduces signal overhead through limited CSI acquisition by using the index of the above link as an index within the pre-input matrix of the analog beamformer.

[0027] The step of applying an SLNR maximization technique that minimizes interference between user terminals based on the above link involves multiple agents existing in the MBS performing learning simultaneously with other agents in a virtual network designed based on the gain table of the MBS, selecting multiple candidate links within each agent, and if there are links among the candidate links selected by each agent that match the row or column in the gain table, selecting only the link having the largest channel gain information among the candidate links selected by each agent within the gain table of the MBS and not selecting the remaining candidate links.

[0028] The above multi-agent reinforcement learning enables beam output in response to changes in the communication environment or channel inputs by immediately designing beams between multiple SBSs and multiple user terminals using online reinforcement learning.

[0029] The above multi-agent reinforcement learning assumes multiple SBSs that are not fixed in position in a UM-MIMO network, and increases the degrees of freedom and efficiency of the multi-agent reinforcement learning by matching multiple SBSs to multiple agents and learning beams simultaneously.

[0030] In one aspect, the interference control and hybrid beamforming system for multiple users proposed in the present invention, which applies multi-agent deep reinforcement learning, includes a learning unit that performs multi-agent reinforcement learning using expected interference and antenna gain information based on a gain table designed using Channel State Information (CSI) of all user terminals, and an interference control and beamforming execution unit that searches for pairs of analog beamforming matrices corresponding to links that maximize antenna gain and minimize interference between user terminals through multi-agent reinforcement learning, applies a Signal to Leakage Plus Noise Ratio (SLNR) maximization technique that minimizes interference between user terminals based on said links, and optimizes transmission power for each link based on iterative waterfilling. Effects of the invention

[0031] According to embodiments of the present invention, it can be applied to high-density network systems with a radius of tens of meters in the sub-THz band, which is a core technology and challenge of 5G and beyond 5G mobile communication systems. Sub-THz communication technology is currently receiving significant attention from academia and industry, along with ultra-massive MIMO technology. In particular, the present invention can achieve high maximum transmission speeds per unit area that meet the requirements of 5G and beyond 5G mobile communication systems, thus offering high market potential. Furthermore, by effectively controlling interference in high-density network systems through multi-agent deep reinforcement learning, it is possible to achieve high data transmission rates and maximum transmission speeds per unit area that were previously unattainable. Moreover, as AI-based communication systems are expected to represent the beyond 5G mobile communication technology market in the future, multi-agent deep reinforcement learning possesses advantages that differentiate it from existing AI technologies, allowing for the expectation of high technological superiority and market leadership. Brief explanation of the drawing

[0032] FIG. 1 is a diagram showing a user-centric network model using UM-MIMO according to an embodiment of the present invention. FIG. 2 is a diagram showing the configuration of an interference control and hybrid beamforming system applying multi-agent deep reinforcement learning for multiple users according to an embodiment of the present invention. FIG. 3 is a flowchart illustrating an interference control and hybrid beamforming method applying multi-agent deep reinforcement learning for multiple users according to an embodiment of the present invention. FIG. 4 is a diagram showing the interference control and hybrid beamforming process with multi-agent deep reinforcement learning applied according to one embodiment of the present invention. FIG. 5 is a diagram showing the simulation results according to one embodiment of the present invention. Specific details for implementing the invention

[0033] Hereinafter, embodiments of the present invention will be described in detail with reference to the attached drawings.

[0035] FIG. 1 is a diagram showing a user-centric network model using UM-MIMO according to an embodiment of the present invention.

[0036] Reinforcement learning has garnered attention as a highly efficient artificial intelligence technique for decision-making scenarios of a certain scale. However, in reality, numerous variables influence scenarios, and deep reinforcement learning has been proposed to handle large-scale scenarios that reflect these factors.

[0037] Deep reinforcement learning can efficiently find the optimal policy to maximize the reward of a scenario by utilizing deep neural networks and the interaction between the environment and objects called agents.

[0038] Previously, single-agent deep reinforcement learning has been extensively studied, but this method has difficulty controlling the explosively increasing number of possible states and actions due to growing variables. Furthermore, unlike what is assumed by single-agent deep reinforcement learning, in real-world environments, information sharing may be limited due to constraints between the agent and the environment, and multiple agents may be required.

[0039] To overcome this, multi-agent deep reinforcement learning has been proposed, which enables scalable scenario design by having multiple agents interact with the environment within the scenario.

[0040] As shown in FIG. 1, an interference control method is proposed that applies multi-agent deep reinforcement learning for multiple users (121, 122) in a user-centric network model based on a Macro Base Station (MBS) (110) using Ultra Massive Multiple Input Multiple Output (UM-MIMO).

[0041] In high-density networks operating in the ultra-high frequency band, eliminating interference between users is crucial for achieving high data transmission rates. In particular, when beamforming is applied by utilizing multiple antennas at both the base station and the user, interference can be controlled by enabling the simultaneous formation of links through more precise beams. However, since this significantly increases the complexity of interference control methods, existing optimization techniques struggle to efficiently eliminate interference.

[0042] To date, methods to reduce complexity by approximating existing problems and single-agent deep reinforcement learning, which exhibits high efficiency when the number of users is small, have been applied. However, as the number of users in the network increases, single-agent deep reinforcement learning becomes impossible to train as the size of the deep neural network increases significantly.

[0043] By introducing multi-agent deep reinforcement learning and training multiple agents corresponding to the number of small base stations (SBS) (131, 132, 133) responsible for link connections (140, 150) within the network simultaneously, the complexity can be significantly reduced compared to the single-agent deep reinforcement learning mentioned earlier, and a high data transmission rate can be obtained through interference removal more efficient than existing optimization techniques.

[0044] In existing technologies, beamforming for a single-antenna-based user or multi-user-based beamforming systems were designed based on semi-optimal interference control without incorporating AI technology. Alternatively, systems that transmit and receive signals only within a limited cell area to suppress interference have been proposed. However, the present invention enables the achievement of a high level of data transmission rate by efficiently eliminating interference in beamforming for multiple antenna-based users using limited channel information.

[0045] In addition, unlike existing single-agent deep reinforcement learning-based systems or systems designed for single-antenna-based users, the present invention enables the design of an AI-based hybrid beamforming system applicable even as the number of base stations and users increases, through a multi-agent deep reinforcement learning-based hybrid beamforming system for multiple antenna-based user systems exhibiting high complexity.

[0047] FIG. 2 is a diagram showing the configuration of an interference control and hybrid beamforming system applying multi-agent deep reinforcement learning for multiple users according to an embodiment of the present invention.

[0048] The interference control and hybrid beamforming system (200) according to the present embodiment may include a processor (210), a bus (220), a network interface (230), memory (240), and a database (250). The memory (240) may include an operating system (241) and an interference control and hybrid beamforming routine (242) that applies multi-agent deep reinforcement learning for multiple users. The processor (210) may include a learning unit (211) and an interference control and beamforming execution unit (212). In other embodiments, the interference control and hybrid beamforming system (200) may include more components than those of FIG. 2. However, it is not necessary to clearly illustrate most of the prior art components. For example, the interference control and hybrid beamforming system (200) may include other components such as a display or a transceiver.

[0049] Memory (240) is a computer-readable recording medium and may include a non-perishable permanent mass storage device such as RAM (random access memory), ROM (read only memory), and a disk drive. Additionally, program code for an operating system (241) and interference control and hybrid beamforming routines (242) applying multi-agent deep reinforcement learning for multiple users may be stored in memory (240). These software components may be loaded from a computer-readable recording medium separate from memory (240) using a drive mechanism (not shown). This separate computer-readable recording medium may include computer-readable recording media (not shown), such as a floppy drive, disk, tape, DVD / CD-ROM drive, or memory card. In another embodiment, software components may be loaded into memory (240) via a network interface (230) rather than a computer-readable recording medium.

[0050] The bus (220) can enable communication and data transmission between components of the interference control and hybrid beamforming system (200). The bus (220) can be configured using a high-speed serial bus, a parallel bus, a Storage Area Network (SAN), and / or other suitable communication technology.

[0051] The network interface (230) may be a computer hardware component for connecting the interference control and hybrid beamforming system (200) to a computer network. The network interface (230) may connect the interference control and hybrid beamforming system (200) to a computer network via a wireless or wired connection.

[0052] The database (250) can serve to store and maintain all information necessary for interference control and hybrid beamforming using multi-agent deep reinforcement learning for multiple users. Although FIG. 2 illustrates the database (250) being built and included inside the interference control and hybrid beamforming system (200), it is not limited thereto and may be omitted depending on the system implementation method or environment, or it is also possible for all or part of the database to exist as an external database built on a separate system.

[0053] The processor (210) may be configured to process instructions of a computer program by performing basic arithmetic, logic, and input / output operations of the interference control and hybrid beamforming system (200). Instructions may be provided to the processor (210) via memory (240) or a network interface (230) and via a bus (220). The processor (210) may be configured to execute program code for the learning unit (211) and the interference control and beamforming execution unit (212). Such program code may be stored in a recording device such as memory (240).

[0054] The learning unit (211) and the interference control and beamforming unit (212) may be configured to perform the steps (310–340) of FIG. 3.

[0055] The interference control and hybrid beamforming system (200) may include a learning unit (211) and an interference control and beamforming execution unit (212).

[0056] The learning unit (211) according to an embodiment of the present invention performs multi-agent reinforcement learning by utilizing expected interference and antenna gain information based on a gain table designed using the Channel State Information (CSI) of all user terminals.

[0057] The learning unit (211) according to an embodiment of the present invention designs the Signal to Interference Plus Noise Ratio (SINR) for each SBS (Small Base Station), which is predicted based on the interference and antenna gain expected during the multi-agent reinforcement learning process for a plurality of agents corresponding to each SBS for the multi-agent reinforcement learning, as the compensation for each agent.

[0058] A learning unit (211) according to an embodiment of the present invention learns a link configuration that maximizes the data transmission rate through a trial and error method from a predetermined number of episodes and time intervals within the episodes for all agents.

[0059] The interference control and beamforming execution unit (212) according to an embodiment of the present invention searches for pairs of analog beamforming matrices corresponding to links that maximize antenna gain and minimize interference between user terminals through multi-agent reinforcement learning. Based on the links, it applies a Signal to Leakage Plus Noise Ratio (SLNR) maximization technique that minimizes interference between user terminals and optimizes the transmission power for each link based on iterative waterfilling.

[0060] The interference control and beamforming execution unit (212) according to an embodiment of the present invention defines the link selected so far from the set of candidate links for each agent as the agent's state, defines the link selected in the current time interval as the agent's action, and defines the sum of the data transmission rates obtained from the links selected so far of the SBS associated with the agent in the corresponding time interval for each agent as a reward.

[0061] The interference control and beamforming execution unit (212) according to an embodiment of the present invention configures the state of the agent and the action of the agent into binary vectors to search for a link between the SBS and the user terminal that minimizes interference with the selected links and maximizes antenna gain.

[0062] The interference control and beamforming execution unit (212) according to an embodiment of the present invention applies a hybrid beamforming matrix to each SBS and applies an analog beamforming matrix to the user terminal without baseband beamforming. When a plurality of SBSs simultaneously transmit signals to user terminals of links connected to them, the sum of the data transmission rates of the entire network is calculated based on the SINR of all links due to interference between links connected to user terminals, interference between user terminals caused by an SBS connected to a user terminal establishing a link with another user terminal to transmit a stream, and interference received from other SBSs. Using the sum of the data transmission rates, the beamforming matrix pairs and the link configuration between the SBS and the user terminal are optimized to maximize the data transmission rate.

[0063] The interference control and beamforming unit (212) according to an embodiment of the present invention transmits only a portion of the entire channel according to the channel gain between the SBS and the user terminal to prevent signal overhead for the MBS in an Ultra Massive Multiple Input Multiple Output (UM-MIMO) network, and can reduce signal overhead through limited CSI acquisition by using the index of the link as an index within a pre-entered pre-matrix of the analog beamformer.

[0064] The interference control and beamforming execution unit (212) according to an embodiment of the present invention performs learning simultaneously with other agents in a virtual network designed based on the gain table of the MBS, with a plurality of agents existing in the MBS. Each agent selects a plurality of candidate links within the network, and if there are links among the selected candidate links that match a row or column in the gain table, only the link having the largest channel gain information among the candidate links is selected, and the remaining candidate links are not selected.

[0065] The interference control and beamforming unit (212) according to an embodiment of the present invention can immediately design beams between a plurality of SBSs and a plurality of user terminals using online reinforcement learning, thereby enabling beam output according to changes in the communication environment or changes in channel input. By assuming a plurality of SBSs that are not fixed in position in a UM-MIMO network and matching a plurality of SBSs to a plurality of agents to learn beams simultaneously, the degrees of freedom and efficiency of the multi-agent reinforcement learning can be increased.

[0067] FIG. 3 is a flowchart illustrating an interference control and hybrid beamforming method applying multi-agent deep reinforcement learning for multiple users according to an embodiment of the present invention.

[0068] The proposed interference control and hybrid beamforming method for multiple users using multi-agent deep reinforcement learning includes the step of performing multi-agent reinforcement learning (310) using expected interference and antenna gain information based on a gain table designed using Channel State Information (CSI) of all user terminals, the step of searching for pairs of analog beamforming matrices corresponding to links that maximize antenna gain and minimize interference between user terminals through multi-agent reinforcement learning (320), the step of applying a Signal to Leakage Plus Noise Ratio (SLNR) maximization technique that minimizes interference between user terminals based on the links (330), and the step of optimizing transmission power for each link based on iterative waterfilling.

[0069] In step (310), multi-agent reinforcement learning is performed using expected interference and antenna gain information based on a gain table designed using the CSI of all user terminals.

[0070] For the above multi-agent reinforcement learning, for multiple agents corresponding to each SBS (Small Base Station), the Signal to Interference Plus Noise Ratio (SINR) for each SBS, predicted based on the interference and antenna gain expected during the multi-agent reinforcement learning process, is designed as the reward for each agent.

[0071] For all agents, a link configuration that maximizes the data transmission rate is learned through a trial-and-error method using a predetermined number of episodes and time intervals within the episodes.

[0072] In step (310), a pair of analog beamforming matrices corresponding to a link that maximizes antenna gain and minimizes interference between user terminals is searched through multi-agent reinforcement learning.

[0073] For each agent, the link selected so far from the set of candidate links is defined as the agent's state, the link selected in the current time interval is defined as the agent's action, and the sum of the data transmission rates obtained from the links selected so far of the SBS associated with the agent in the corresponding time interval is defined as the reward.

[0074] The state of the agent and the action of the agent are configured as binary vectors to search for a link between the SBS and the user terminal that minimizes interference with the selected links and maximizes antenna gain.

[0075] In step (330), an SLNR maximization technique is applied to minimize interference between user terminals based on the link.

[0076] Each SBS applies a hybrid beamforming matrix, and the user terminal applies an analog beamforming matrix without baseband beamforming. When multiple SBSs simultaneously transmit signals to user terminals on the links connected to them, the sum of the data transmission rates of the entire network is calculated based on the SINR of all links resulting from interference between links connected to user terminals, interference between user terminals caused by an SBS connected to a user terminal establishing a link with another user terminal to transmit a stream, and interference received from other SBSs. Using the sum of the data transmission rates, beamforming matrix pairs and link configurations between SBSs and user terminals are optimized to maximize the data transmission rate.

[0077] In step (340), an SLNR maximization technique is applied to minimize interference between user terminals based on the link.

[0078] In order to prevent signal overhead for MBS in an Ultra Massive Multiple Input Multiple Output (UM-MIMO) network, only a portion of the entire channel is transmitted according to the channel gain between the SBS and the user terminal, and the index of the link is used as an index within the pre-entered matrix of the analog beamformer, thereby reducing signal overhead through limited CSI acquisition.

[0079] Multiple agents existing in the above MBS perform learning simultaneously with other agents in a virtual network designed based on the gain table of the above MBS. Each agent selects multiple candidate links within the network, and if there are links among the selected candidate links that match a row or column in the gain table, only the link with the largest channel gain information among those candidate links is selected, and the remaining candidate links are not selected.

[0080] Multi-agent reinforcement learning according to an embodiment of the present invention enables beam output in response to changes in the communication environment or channel inputs by immediately designing beams between a plurality of SBSs and a plurality of user terminals using online reinforcement learning.

[0081] Multi-agent reinforcement learning according to an embodiment of the present invention can increase the degrees of freedom and efficiency of the multi-agent reinforcement learning by assuming a plurality of SBSs that are not fixed in position in a UM-MIMO network and matching a plurality of SBSs to a plurality of agents to learn beams simultaneously.

[0082] As such, according to an embodiment of the present invention, a multi-agent deep reinforcement learning-based system for maximizing the transmission rate per unit area in a mobile communication environment where multiple base stations and multiple users exist is proposed.

[0083] According to an embodiment of the present invention, multi-agent deep reinforcement learning capable of optimal resource allocation and expansion without labels is performed in a mobile communication environment where multiple base stations and multiple users exist.

[0084] The design of beam patterns and interference control of beams between SBS and UEs, which are highly complex UM-MIMO-based, can be efficiently distributed through multi-agent deep reinforcement learning. Based on this, unlike existing single-agent deep reinforcement learning-based methods, it is possible to extend and apply this to high-density networks consisting of multiple base stations and multiple users.

[0085] According to an embodiment of the present invention, a state and behavior structure composed of binary vectors that reduces the training difficulty of deep reinforcement learning is applied, and by composing the state and behavior as binary vectors, even if the optimal state is not found with high probability after the learning process, a suboptimal state having a similar correlation can be found.

[0086] According to an embodiment of the present invention, complex channel and hybrid beamformer structures are simplified into an index and reflected in state and behavior.

[0087] While most machine learning techniques for designing existing channels and hybrid beamformers input elements consisting of complex numbers into deep neural networks by dividing them into real and imaginary parts, the present invention can reduce the complexity of states and behaviors by inputting a prior matrix in advance and defining states and behaviors based on indices within the said prior matrix.

[0088] According to an embodiment of the present invention, in a network composed of multiple base stations with non-fixed locations, efficient beams are simultaneously learned by associating base stations with agents using multi-agent deep reinforcement learning.

[0089] By using multi-agent deep reinforcement learning to train beams corresponding to each agent for base stations, it becomes possible to develop a deep reinforcement learning technique applicable even to the previously presented high-density networks. By assuming base stations that are not located in fixed positions, one can freely conceive of a network with a high degree of freedom—rather than a fixed network between base stations and users—and then simulate that network at any time within the environment of the presented reinforcement learning technique.

[0090] According to an embodiment of the present invention, beams between a plurality of base stations and a plurality of users are immediately designed using online reinforcement learning.

[0091] Conventional machine learning techniques have proposed methods capable of outputting suitable beams for various channels through extensive prior training. However, this method becomes unable to output a suitable beam for a channel if significant changes in the communication environment lead to the input of a channel not discovered during prior training. The online reinforcement learning according to the embodiment of the present invention prevents this by providing a structure that immediately learns a beam suitable for a new channel even when changes in the communication environment occur. Since it does not require extensive training processes like the aforementioned prior training, it can be considered more suitable.

[0092] According to an embodiment of the present invention, a multi-agent deep reinforcement learning technique with small information exchange overhead between agents is proposed based on a simple state and behavior structure.

[0093] Conventional multi-agent deep reinforcement learning techniques have suppressed policy divergence or instability through limited information exchange between agents. This limited information exchange is necessary because the state and behavior structures are complex, and the signal overhead required to exchange all such information is significant. In contrast, the proposed multi-agent deep reinforcement learning technique requires the actions of other agents to calculate the interference of the expected transmission rate—which serves as each agent's reward—for information exchange between agents. Since the state and behavior structures are already very simple, the overhead consumed is minimal, enabling more stable multi-agent reinforcement learning.

[0094] According to an embodiment of the present invention, candidate beams are selected to reduce the complexity of the behavior applied to deep reinforcement learning.

[0095] Analog beamforming based on multiple antenna base stations or user codebooks exhibits low efficiency and longer training periods if training is performed on every beam, as the beams are fine and the codebooks are large. To prevent this, beams are selected based on prior knowledge to avoid duplicate selection of highly correlated beam indices (beam pairs within the same row or column) expected to exhibit high interference; specifically, beams that share the same row or column in the gain table as previously selected beams are not chosen again. Based on beams selected in this manner, a condensed state and action space can be maintained. From the simulation results above, it can be confirmed that even as the number of users increases, the data transmission rate increases proportionally with the number of users by utilizing state and action spaces of a fixed size.

[0097] FIG. 4 is a diagram showing the interference control and hybrid beamforming process with multi-agent deep reinforcement learning applied according to one embodiment of the present invention.

[0098] First, each parameter for explaining the system model according to an embodiment of the present invention is described below:

[0099] : Number of small base stations

[0100] : Number of users

[0101] : Number of antennas in small base stations

[0102] Number of user antennas

[0103] : Array response vector of a small base station

[0104] : User's array response vector

[0105] : The th small base station and Channel matrix between the nth users

[0106] : The th small base station and LOS component of the channel matrix between the nth users

[0107] : The th small base station and non-LOS components of the channel matrix between the nth users

[0108] : Large-scale channel gain of the LOS component due to path loss and shadowing

[0109] : Large-scale channel gain of non-LOS components due to path loss and shadowing

[0110] : Number of RF chains in small base stations

[0111] : Number of user RF chains

[0112] : Transmission power of a small base station

[0113] : Transmission beamforming matrix of the nth small base station

[0114] : Transmit analog beamforming matrix of the nth small base station

[0115] : Transmit baseband beamforming matrix of the nth small base station

[0116] : Transmission power allocation matrix of the nth small base station

[0117] : Receive beamforming matrix of the th user

[0118] : i-th user's receiving analog beamforming matrix

[0119] : Receive baseband beamforming matrix of the i-th user

[0120] : Center frequency

[0121] : Objective function

[0122] Expected discount gain

[0123] : Loss function

[0124] : hour In State of the th agent

[0125] : hour In Actions of the th agent

[0126] : hour In Reward for the th agent

[0127] : Number of candidate links

[0128] : The th small base station and Link between the first users

[0129] : Bandwidth

[0130] : Noise power spectrum density

[0131] : conjugate transpose

[0132] Expectation operator

[0133] : The total number of links used by the nth small base station

[0134] : Total number of users supported by the nth small base station

[0135] : A set of all small base stations connected to the nth user

[0136] : Transmission power allocation matrix of the nth small base station

[0137] : Total number of dominant elements within the channel selected for restricted feedback

[0138] The proposed interference control and hybrid beamforming technology applying multi-agent deep reinforcement learning is a method for controlling interference between links in a user-centric high-density network in the sub-THz band based on deep reinforcement learning. To handle interference control with high complexity, multi-agent deep reinforcement learning is introduced to maximize the data transmission rate for all UEs (User Equipment). When applying deep reinforcement learning, multi-agent reinforcement learning is applied because using single-agent reinforcement learning would significantly increase complexity in controlling all beamforming within the entire high-density network. A two-stage hybrid beamforming design method is introduced; in the reinforcement learning execution stage, antenna gain is maximized through multi-agent reinforcement learning, and subsequently, in the interference control and beamforming execution stage, analog beamforming matrix pairs corresponding to links that mitigate user-to-user interference are searched, and a Signal to Leakage Plus Noise Ratio (SLNR) maximization technique is applied to the baseband beamforming matrix.

[0139] Specifically, in the reinforcement learning execution step (410), deep reinforcement learning is trained using expected interference and antenna gain information based on a gain table designed using the Channel State Information (CSI) of all users (411). The learning environment consists of I agents corresponding to each SBS. During the learning process, the Signal to Interference Plus Noise Ratio (SINR) for each SBS, predicted based on the expected interference and antenna gain, is designed as the reward for each agent (421). For each agent, the link selected so far from the set of candidate links is considered as the agent's state, and the link selected in the current time interval is considered as the agent's action (422). From this, an SBS-UE link that minimizes interference and maximizes antenna gain is searched (430).

[0140] In the interference control and beamforming step, an SLNR maximization technique is applied to mitigate interference between users based on the links discovered in the preceding reinforcement learning step (412). Additionally, the transmission power per link is optimized based on iterative water filling.

[0141] Below, regarding the channel model of a downlink multi-user environment, the Sub-THz channel is described.

[0142]

[0143]

[0144]

[0145] The characteristics of sub-THz band channels include a limited number of clusters and rays, and high path attenuation, and the parameters used in the above formula are as follows:

[0146] Bernoulli random variable

[0147] : MIMO channel gain for the LOS path

[0148] : MIMO channel gain for NLOS path

[0149] : Number of NLOS path components

[0150] : Complex Gaussian random variable

[0151] : Path loss index and log normal shadowing gain for LOS (NLOS) path

[0152] : The th SBS and Distance between the th UE

[0153] Is With a probability With a probability of 0, and Igo is, , When It is modeled as.

[0154] Next, regarding the channel model of a downlink multi-user environment, the signaling model will be explained.

[0155] The signal received by the i-th user and the signal passing through the receiver can be represented as follows:

[0156]

[0157]

[0158] Here, the user does not implement the baseband beamforming matrix among the hybrid beamforming matrices. Also, , , , , am. silver It is the transmission power supporting the stream of the nth link.

[0159] From the signal above and It consists only of RF phase shifters.

[0160] analog beamforming matrix and Orthogonal RF beamforming vectors are used when configuring it.

[0161] The baseband beamforming matrix within the SBS is configured differently for each connected user, and Igo silver It is a baseband beamforming column vector supporting the stream of the nth link.

[0162] Next, we will describe the design of a hybrid beamforming matrix for maximizing data transmission rates and the design of a problem optimizing the link configuration between a small base station and a user.

[0163] anticipant From the signal received by the nth user The th SBS and Each between the th UE The SINR for the links can be expressed as follows.

[0164]

[0165]

[0166]

[0167]

[0168] Here Is The th UE Within the analog beamforming matrix used for connection with the nth SBS It is the nth column vector.

[0169] gun SBSs simultaneously transmit signals to UEs on the links connected to them. Each SBS applies a hybrid beamforming matrix, and the UE applies an analog beamforming matrix without baseband beamforming. Here, interference between the links connected to the user Inter-user interference caused by a small base station connected to a user establishing a link with another user and transmitting a stream , interference received from other small base stations Defined as. is the user's noise power.

[0170] Based on the SINR of all links designed within the network as described above, the sum of the data transmission rates of the entire network can be calculated as follows.

[0171]

[0172] Based on this, the following data transmission rate maximization problem can be designed.

[0173]

[0174] The first and second constraints relate to the transmit power constraints of each SBS, the third constraint is due to analog beamforming constraints, and the last constraint is necessary because the number of streams that each SBS or UE can transmit over the entire link cannot exceed the number of RF chains.

[0175] A method for reducing signal overhead through limited CSI acquisition according to an embodiment of the present invention is described.

[0176] Assuming that all channel information between the small base station and the user is transmitted, the signal overhead concentrated on the MBS in ultra-dense networks becomes very large, so it is necessary to prevent this.

[0177] To reduce signal overhead toward the MBS, only some of the entire channel components are transmitted.

[0178]

[0179] Is It is a pre-designed angle such that the response vectors of the SBSs are orthogonal to each other, and It is also designed similarly. SBS's n The nth pre-response vector and UE's m This corresponds to the channel gain obtainable when using the i-th prior response vector. The most dominant of these channel gain components If you use the components, you can obtain the second equation.

[0180] In the preceding data transmission rate maximization problem, the analog beamformers of the SBS and UE are composed of a prior matrix of the beam directions of the designed link. For example, The th SBS and Between the first UE , If the link is designed in the direction, and each The th SBS and It is included as a column of analog beamformers to be applied to the th UE.

[0181] Since the analog beamformer for the corresponding beam direction is determined once the link is determined, the index of the link is assumed to be the index within the analog beamformer's dictionary matrix. For example When the nth prior response vector is selected as shown, the index of the corresponding link is assumed to be n.

[0182] A method for controlling interference between links using multi-agent deep reinforcement learning according to an embodiment of the present invention will be described.

[0183] In a multi-agent deep reinforcement learning system environment, MBS designs the channel gain table for all SBS-UE pairs as follows, based on limited channel information obtained from all UEs.

[0184]

[0185] represents the channel gain information between the first SBS and all UEs in the gain table above.

[0186] Within the MBS, design a pair of analog beamforming matrices for each SBS and the analog beamforming matrices of the UE to be connected to that SBS. I There are 2 agents.

[0187] These agents learn simultaneously with other agents within a virtual network designed based on the MBS gain table.

[0188] The ninth agent is Total inside Select candidate links. For example The candidate link selection method for the nth agent is as follows.

[0189] Select the largest element within.

[0190] After selection, to minimize interference, the corresponding element is within the matrix If an element is at a specific position, set all elements located at the nth row and mth column to 0.

[0191] Through the above process All candidate links are selected or all If my element is 0, terminate the above process.

[0192] The process of removing elements located in the same row or column as the selected element can prevent unnecessary learning processes in advance.

[0193] The Markov determination process, episode, and time interval structure according to an embodiment of the present invention will be described.

[0194] All agents learn link configurations that maximize the data transmission rate through a trial-and-error method using a predetermined number of episodes and the time intervals within them.

[0195] Each episode consists of time intervals equal to the number of RF chains. At the start of each episode, the state and behavior of each agent are initialized to a zero vector. Markov decision processes in reinforcement learning define the state, behavior, and reward type of each agent, and define the transition process from the current state to another state. Markov decision processes in reinforcement learning are defined as follows:

[0196]

[0197] Defines the currently selected link as the agent's state for each agent.

[0198] For each agent, define the link to be selected within the corresponding time interval as the agent's action.

[0199] For each agent, the reward is defined as the sum of the data transmission rates obtainable from the selected links from the corresponding time interval to the current time interval of the SBS associated with the agent:

[0200]

[0201] for example, In this case, the state vector can be defined as [0,0,0], [0,0,1], [0,1,0], [1,0,0], [0,1,1], [1,1,0], [1,0,1], and the action vector can be defined as [0,0,0], [0,0,1], [0,1,0], [1,0,0].

[0202] When states and actions are constructed using binary vectors as described above and one link is selected in a single time interval, similar states have a high correlation with each other. Consequently, reinforcement learning involving stochastic actors such as SAC is highly likely to converge to a state that is highly correlated with the optimal state.

[0203] Most deep reinforcement learning is designed based on deep neural networks. When using fully connected layers within a deep neural network, adjacent nodes are likely to have similar weights. In this case, it is advantageous to consider the correlations between adjacent nodes and configure the network so that states with similar correlations approximate the relationships between nodes located close to each other.

[0204] One of the SBS For example, some of my binary vectors [0,0,1,1,1,1,0] and [0,0,1,1,1,0,0] have a high correlation and are likely to have similar expected data transmission rates. In addition, since the input and output of a deep neural network are states and actions, the proposed structure takes into account the node-by-node correlations of the input and output neural layers.

[0205] The reward for each agent is defined as the sum of data transmission rates obtainable per SBS from the links selected up to the current time interval, provided that the number of links supporting the same UE exceeds the number of the UE's RF chains, or the beam on the UE side connecting the links If the beam on the UE side of the selected link in this preceding time interval is the same, the stream is not transmitted to the link with the lower channel gain for interference control.

[0206] The state transition process of the i-th agent is defined as follows, based on the definitions of state and behavior above:

[0207]

[0208] In this case, + corresponds to the following binary OR operation for each vector element.

[0209]

[0210] In the reinforcement learning technique SAC (Soft Actor-Critic) applied in the present invention, each agent searches for a state that creates an optimal reward based on a state in a discrete state space and an action in a discrete action space, based on the SAC technique, which is one of the policy-based learning methods.

[0211] In this case, the loss function is defined as follows:

[0212]

[0213] Is It is the state and behavior of the subsequent time interval.

[0214] , is a parameter of the goal critique, Updated policy for the next time interval of, is a parameter of the target actor.

[0215] In baseband beamforming and power allocation according to an embodiment of the present invention, the baseband beamforming is designed to maximize the SLNR maximization problem.

[0216]

[0217] Here, the total leakage power is given by the equation below.

[0218]

[0219] Unlike the existing SINR maximization problem, the problem of maximizing SLNR by link above has a closed solution as follows.

[0220]

[0221] The above constant is a normalization variable to satisfy the power transmission constraint of the data transmission rate maximization problem.

[0222] For power allocation, solve the following power allocation problems to explore the optimal power allocation ratio per link.

[0223]

[0224]

[0225] Since the above problem is local (non-convex), an iterative water filling method is introduced.

[0226] First, the above power allocation problem is expressed using a Lagrange function.

[0227]

[0228] Here, is the Lagrange multiplier, and am.

[0229] After that, the local optimal solution is found using the following two conditions among the KKT (Karush-Kuhn-Tucker) conditions.

[0230]

[0231]

[0232] The year is organized in the following form.

[0233]

[0234] Here It is as follows.

[0235]

[0236] Here , , .

[0237] Repeat until the two KKT conditions above are satisfied or the maximum number of iterations is reached.

[0239] FIG. 5 is a diagram showing the simulation results according to one embodiment of the present invention.

[0240] Based on the link configuration method obtained after the reinforcement learning process is completed and the corresponding hybrid beamformer, the data transmission rate of the user-centric network is maximized.

[0241] To verify the data transmission rate achievable by the links of the designed network considering the previously assumed overhead, the sum of the following overhead model and data transmission rates is assumed as a performance metric.

[0242] In the case of the overhead model, the overhead that occurs when each UE predicts the channel corresponding to each link connected to a different SBS ( ), Overhead incurred when transmitting 2 dominant channel elements to MBS ( ), overhead incurred when transmitting the index of the corresponding link to each SBS and UE after the reinforcement learning process is completed ( It consists of ), and the sum of achievable data transfer rates, considering overhead, is as follows:

[0243]

[0244] Here is the conversion coefficient between bits and symbols, and is assumed to be 1 since binary phase shift modulation (BPSK) is assumed, and Q is the number of bits required to convert a scalar value to a bit.

[0245] Each SBS and UE of the user-centric network was deployed within a sector with a radius of 50m and a range of 120 degrees using the Poisson process.

[0246] Other parameters required for network configuration were set as follows.

[0247]

[0248] The environment parameters within the multi-agent deep reinforcement learning were set as follows.

[0249]

[0250] From the simulation results of Figure 5, it can be seen that the sum of the data transmission rates obtained from the proposed link configuration method using multi-agent deep reinforcement learning is higher than that of the conventional link configuration method [3]. In other words, the proposed technology controls interference between links more efficiently.

[0251] In addition, it can be seen that even if the number of I SBSs and K UEs increases to 30, the multi-agent reinforcement learning-based data transmission rate is higher than the configuration method of the prior art [3].

[0252] In addition, the length of the state space and action vector , Despite being limited to [a certain amount], an increase in the data transmission rate with increasing user numbers can be observed. This can be analyzed as a result of the increased number of high-gain beams available for selection and the increased degree of freedom in beam selection methods that can avoid interference as the number of users increases.

[0254] The device described above may be implemented as a hardware component, a software component, and / or a combination of a hardware component and a software component. For example, the device and components described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and one or more software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include a plurality of processing elements and / or a plurality of types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. Additionally, other processing configurations, such as parallel processors, are also possible.

[0255] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or instruct the processing unit independently or collectively. Software and / or data may be embodied in any type of machine, component, physical device, virtual equipment, computer storage medium, or device so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and may be stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media.

[0256] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either alone or in combination. The program instructions recorded on the medium may be those specifically designed and configured for the embodiment, or they may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc.

[0257] Although the embodiments have been described above with reference to limited examples and drawings, those skilled in the art can make various modifications and variations from the description above. For example, suitable results can be achieved even if the described techniques are performed in a different order than described, and / or the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.

[0258] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims set forth below.

Claims

Claim 1 An interference control and hybrid beamforming method performed by a computer-implemented interference control and hybrid beamforming system at a Macro Base Station (MBS), wherein the learning unit of the interference control and hybrid beamforming system performs multi-agent reinforcement learning using expected interference and antenna gain information based on a gain table designed using Channel State Information (CSI) for downlinks between multiple user terminals located within the coverage of multiple Small Base Stations (SBS) and the multiple SBSs; the interference control and beamforming execution unit of the interference control and hybrid beamforming system searches for analog beamforming matrix pairs corresponding to the downlinks between the multiple SBSs and the multiple user terminals that maximize antenna gain and minimize interference between user terminals through multi-agent reinforcement learning; and the interference control and beamforming execution unit applies a technique to maximize the Signal to Leakage Plus Noise Ratio (SLNR) for each downlink of the downlinks, which is the ratio of the power of a signal transmitted through the corresponding downlink to the leakage interference and noise power exerted by the signal on other user terminals.The interference control and beamforming execution unit comprises a step of optimizing the transmission power of each of the downlinks based on iterative waterfilling, wherein the applying step comprises applying a hybrid beamforming matrix to each of the plurality of SBSs and applying an analog beamforming matrix without baseband beamforming to the user terminal, and when the plurality of SBSs simultaneously transmit signals to user terminals of the links connected to them, calculating the sum of the data transmission rates of the entire network based on the SINR of all links due to interference between links connected to the user terminals, interference between user terminals caused by an SBS connected to a user terminal opening a link with another user terminal and transmitting a stream, and interference received from other SBSs, and optimizing the beamforming matrix pairs and link configuration between the SBS and the user terminal to maximize the data transmission rate using the sum of the data transmission rates. Claim 2 In claim 1, the step of performing the multi-agent reinforcement learning comprises designing the Signal to Interference Plus Noise Ratio (SINR) for each SBS, predicted based on the interference and antenna gain expected during the multi-agent reinforcement learning process for each agent of the plurality of agents corresponding to each SBS for the multi-agent reinforcement learning, as the compensation for each agent, in an interference control and hybrid beamforming method. Claim 3 In paragraph 2, the step of performing the multi-agent reinforcement learning is an interference control and hybrid beamforming method that learns a link configuration that maximizes the data transmission rate through a trial and error method from a predetermined number of episodes and time intervals within the episodes for the plurality of agents. Claim 4 In paragraph 2, the searching step defines the link selected so far from a set of candidate links consisting of downlinks included in the gain table for each agent as the state of the agent, defines the link selected in the current time interval as the action of the agent, and defines the sum of the data transmission rates obtained from the links selected so far of the SBS associated with the agent in the corresponding time interval for each agent as the compensation, in an interference control and hybrid beamforming method. Claim 5 In claim 4, the searching step comprises an interference control and hybrid beamforming method for searching for links between an SBS and a user terminal that minimize interference and maximize antenna gain for the selected links by configuring the state of the agent and the behavior of the agent into binary vectors. Claim 6 delete Claim 7 An interference control and hybrid beamforming method performed by a computer-implemented interference control and hybrid beamforming system at a Macro Base Station (MBS), wherein the learning unit of the interference control and hybrid beamforming system performs multi-agent reinforcement learning using expected interference and antenna gain information based on a gain table designed using Channel State Information (CSI) for downlinks between multiple user terminals located within the coverage of multiple Small Base Stations (SBS) and the multiple SBSs; the interference control and beamforming execution unit of the interference control and hybrid beamforming system searches for analog beamforming matrix pairs corresponding to the downlinks between the multiple SBSs and the multiple user terminals that maximize antenna gain and minimize interference between user terminals through multi-agent reinforcement learning; and the interference control and beamforming execution unit applies a technique to maximize the Signal to Leakage Plus Noise Ratio (SLNR) for each downlink of the downlinks, which is the ratio of the power of a signal transmitted through the corresponding downlink to the leakage interference and noise power exerted by the signal on other user terminals. The interference control and beamforming performing unit comprises a step of optimizing the transmission power of each of the downlinks based on iterative waterfilling, and the applying step transmits only a portion of the entire channel according to the channel gain between the SBS and the user terminal to prevent signal overhead for the MBS in an Ultra Massive Multiple Input Multiple Output (UM-MIMO) network, and uses the indices of the downlinks as indices within a pre-entered pre-matrix of the analog beamformer to reduce signal overhead through limited CSI acquisition, thereby providing an interference control and hybrid beamforming method. Claim 8 In claim 7, the applying step comprises: a plurality of agents existing in the MBS performing learning simultaneously with other agents in a virtual network designed based on the gain table of the MBS; selecting a plurality of candidate links within each agent; and if there are links among the candidate links selected by each agent that match the row or column in the gain table, selecting only the link having the largest channel gain information among the candidate links selected by each agent within the gain table of the MBS and not selecting the remaining candidate links, in an interference control and hybrid beamforming method. Claim 9 In claim 1, the multi-agent reinforcement learning is an interference control and hybrid beamforming method capable of beam output according to changes in the communication environment or changes in channel input by immediately designing beams between multiple SBSs and multiple user terminals using online reinforcement learning. Claim 10 In claim 1, the interference control and hybrid beamforming method increases the degrees of freedom and efficiency of the multi-agent reinforcement learning by assuming a plurality of SBSs that are not fixed in position in a UM-MIMO network and matching a plurality of SBSs to a plurality of agents to simultaneously learn beams. Claim 11 In an interference control and hybrid beamforming system implemented by a computer at a Macro Base Station (MBS), a learning unit that performs multi-agent reinforcement learning using expected interference and antenna gain information based on a gain table designed using Channel State Information (CSI) for downlinks between multiple user terminals located within the coverage of multiple Small Base Stations (SBS) and said multiple SBSs; The interference control and beamforming execution unit includes a pair of analog beamforming matrices corresponding to the downlinks between the plurality of SBSs and the plurality of user terminals, which maximize antenna gain and minimize interference between user terminals through multi-agent reinforcement learning, and applies a technique to maximize the SLNR (Signal to Leakage Plus Noise Ratio), which is the ratio of the power of the signal transmitted through the corresponding downlink to the leakage interference and noise power exerted by the signal on other user terminals for each downlink of the downlinks, and optimizes the transmission power of each of the downlinks based on iterative waterfilling. The interference control and beamforming execution unit applies a hybrid beamforming matrix to each of the plurality of SBSs and applies an analog beamforming matrix without baseband beamforming to the user terminals, and when the plurality of SBSs simultaneously transmit signals to user terminals on the links connected to them, the data of the entire network based on the SINR of all links resulting from interference between links connected to user terminals, interference between user terminals caused by an SBS connected to a user terminal establishing a link with another user terminal and transmitting a stream, and interference received from other SBSs. An interference control and hybrid beamforming system that calculates the sum of transmission rates and optimizes beamforming matrix pairs for maximizing data transmission rates and link configurations between SBS and user terminals using the sum of the data transmission rates. Claim 12 In claim 11, the learning unit is an interference control and hybrid beamforming system that designs the Signal to Interference Plus Noise Ratio (SINR) for each SBS, predicted based on the interference and antenna gain expected during the multi-agent reinforcement learning process for each agent of the plurality of agents corresponding to each SBS of the plurality of SBSs for the multi-agent reinforcement learning, as the compensation for each agent. Claim 13 In claim 12, the learning unit is an interference control and hybrid beamforming system that learns a link configuration that maximizes the data transmission rate through a trial and error method from a predetermined number of episodes and time intervals within the episodes for the plurality of agents. Claim 14 In claim 12, the interference control and beamforming performing unit defines the link selected so far from a set of candidate links consisting of downlinks included in the gain table for each agent as the agent's state, defines the link selected in the current time interval as the agent's action, and defines the sum of the data transmission rates obtained from the links selected so far of the SBS associated with the agent in the corresponding time interval for each agent as a reward, in an interference control and hybrid beamforming system. Claim 15 In claim 14, the interference control and beamforming unit comprises an interference control and hybrid beamforming system that searches for links between an SBS and a user terminal to minimize interference on the selected links and maximize antenna gain by configuring the state of the agent and the action of the agent into binary vectors. Claim 16 delete Claim 17 In an interference control and hybrid beamforming system implemented by a computer at a Macro Base Station (MBS), a learning unit that performs multi-agent reinforcement learning using expected interference and antenna gain information based on a gain table designed using Channel State Information (CSI) for downlinks between multiple user terminals located within the coverage of multiple Small Base Stations (SBS) and said multiple SBSs; An interference control and beamforming system comprising an interference control and beamforming unit that searches for analog beamforming matrix pairs corresponding to the downlinks between the plurality of SBSs and the plurality of user terminals, which maximize antenna gain and minimize interference between user terminals through multi-agent reinforcement learning, applies a technique to maximize the SLNR (Signal to Leakage Plus Noise Ratio), which is the ratio of the power of the signal transmitted through the corresponding downlink to the leakage interference and noise power that the signal exerts on other user terminals, for each downlink of the downlinks, and optimizes the transmission power of each of the downlinks based on iterative waterfilling, wherein the interference control and beamforming unit transmits only a portion of the entire channel according to the channel gain between the SBS and the user terminal to prevent signal overhead for the MBS in an Ultra Massive Multiple Input Multiple Output (UM-MIMO) network, and reduces signal overhead through limited CSI acquisition by using the indices of the downlinks as indices within a pre-input matrix of the analog beamformer. Claim 18 In claim 17, the interference control and beamforming unit performs learning simultaneously with other agents in a virtual network designed based on the gain table of the MBS, selects multiple candidate links within each agent, and if there are links among the candidate links selected by each agent that match the row or column in the gain table, selects only the link having the largest channel gain information among the candidate links selected by each agent within the gain table of the MBS and does not select the remaining candidate links, thereby constituting an interference control and hybrid beamforming system. Claim 19 In claim 11, the interference control and beamforming unit enables beam output according to changes in the communication environment or changes in channel input by immediately designing beams between a plurality of SBSs and a plurality of user terminals using online reinforcement learning, and assumes a plurality of SBSs that are not fixed in position in a UM-MIMO network, and increases the degrees of freedom and efficiency of the multi-agent reinforcement learning by matching a plurality of SBSs to a plurality of agents and learning beams simultaneously. Claim 20 A program stored in a computer-readable storage medium for executing an interference control and hybrid beamforming method performed by a computer-implemented interference control and hybrid beamforming system at a Macro Base Station (MBS), wherein the interference control and hybrid beamforming method comprises: a step in which a learning unit of the interference control and hybrid beamforming system performs multi-agent reinforcement learning using expected interference and antenna gain information based on a gain table designed using Channel State Information (CSI) for downlinks between multiple user terminals located within the coverage of multiple Small Base Stations (SBS) and the multiple SBSs; a step in which an interference control and beamforming execution unit of the interference control and hybrid beamforming system searches for analog beamforming matrix pairs corresponding to the downlinks between the multiple SBSs and the multiple user terminals that maximize antenna gain and minimize interference between user terminals through multi-agent reinforcement learning; and a step in which the interference control and beamforming execution unit applies a Signal to Leakage Plus Noise Ratio (SLNR) maximization technique that minimizes interference between user terminals based on the downlinks.A program stored on a computer-readable storage medium, wherein the interference control and beamforming execution unit of the interference control and hybrid beamforming system comprises a step of optimizing the transmission power of each of the downlinks based on iterative waterfilling, wherein each of the plurality of SBSs applies a hybrid beamforming matrix and the user terminal applies an analog beamforming matrix without baseband beamforming, and when the plurality of SBSs simultaneously transmit signals to user terminals of the links connected to them, the sum of the data transmission rates of the entire network is calculated based on the SINR of all links due to interference between links connected to user terminals, interference between user terminals caused by an SBS connected to a user terminal establishing a link with another user terminal and transmitting a stream, and interference received from other SBSs, and the beamforming matrix pairs and link configurations between SBSs and user terminals are optimized for maximizing the data transmission rate using the sum of the data transmission rates.

Citation Information

Patent Citations

  • Communication devices and methods

    US20220200678A1

  • Multi-agent policy machine learning

    WO2022159008A1