Method for performing random access procedure of small cell base station in ultra-dense network and apparatus therefor

The joint optimization of power allocation and access control through multi-agent reinforcement learning enhances SIC decryption success rates and reduces access delay in ultra-dense networks by addressing inter-cell interference in machine-type communication devices.

WO2025143554A1PCT designated stage expired Publication Date: 2025-07-03IND UNIV COOP FOUND HANYANG UNIV ERICA CAMPUS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/018478
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-28
Filing Date
2024-11-21
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing unlicensed non-orthogonal random access technologies in ultra-dense networks suffer from high inter-cell interference, leading to SIC decoding failures and reduced device access performance in machine-type communication devices.

Method used

A joint optimization scheme for power allocation and access control using a multi-agent reinforcement learning method, employing a partially observable distributed Markov decision process model (Dec-POMDP) and QMIX reinforcement learning, to minimize SIC decoding failures by optimizing power levels and access control elements.

Benefits of technology

The proposed method increases SIC decryption success rates by up to 15% and reduces access delay, effectively managing inter-cell interference in ultra-dense networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024018478_03072025_PF_FP_ABST
    Figure KR2024018478_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a method and an apparatus for performing a random access procedure of a small cell base station in an ultra-dense network. Specifically, the method comprises the steps of: transmitting, to a machine-type communication device (MTCD), an SIB-2 message including the number of power level pools for power allocation and an access control element for access control; receiving, from the MTCD, a composite signal based on the SIB-2 message; and performing successive interference cancellation (SIC) decoding on the SIB-2 message.
Need to check novelty before this filing date? Find Prior Art

Description

Method for performing random access procedure of small cell base station in ultra-dense network and device therefor

[0001] Embodiments disclosed herein relate to a method for performing a random access procedure of a small cell base station and a device therefor.

[0002] Machine-Type Communication (MTC) is a promising new technology that can realize a ubiquitous computing environment, a concept similar to the Internet of Things (IoT). These MTC devices require little human interaction and can communicate captured data via wireless networks. Potential MTC-based applications and services include smart metering, healthcare monitoring, remote security surveillance, intelligent transportation monitoring systems, and supply chain monitoring. Some forecasts suggest that the number of deployed MTC devices could exceed one billion in the near future. However, the network resources of existing mobile broadband networks can be burdened by the numerous data transmissions by numerous MTC devices, in addition to the numerous data transmissions by human-operated mobile communication devices such as smartphones and tablet personal computers. Furthermore, power consumption and computing resources can be taxed as human-operated mobile communication devices, and MTC devices compete for a limited amount of bandwidth for data communication.

[0003] Accordingly, unlicensed non-orthogonal random access (UAR) is receiving significant attention as a key technology to support the ultra-high-capacity Machine-Type Communication Devices (MTCDs) that will be discussed in 6G. UUAR improves device connectivity based on grant-free random access (grant-free random access) and non-orthogonal multiple access (NOMA) technologies. However, in the next-generation ultra-dense networks anticipated for next-generation networks, existing UUAR technologies have the problem of causing high inter-cell interference, significantly degrading device access performance.

[0004] Meanwhile, the background technology described above is technical information that the inventor possessed for the purpose of deriving the present invention or acquired during the process of deriving the present invention, and cannot necessarily be said to be publicly known technology disclosed to the general public prior to the application for the present invention.

[0005] The present invention proposes a joint optimization method for power allocation and access control to address interference problems in ultra-dense networks.

[0006] The present invention proposes applying a multi-agent reinforcement learning method because the problem of inter-cell interference control is a combinatorial optimization that is difficult to compute within polynomial time.

[0007] The present invention proposes a partially observable distributed Markov decision process model (Dec-POMDP) ​​for machine learning. Furthermore, it proposes a joint optimization algorithm for power allocation and access control based on a multi-agent reinforcement learning approach.

[0008] As a technical means for achieving the above-described technical task, according to one embodiment of the present invention, a method for performing a random access procedure of a small cell base station in an ultra-dense network may include the steps of: transmitting an SIB-2 message including the number of power level pools for power allocation and an access control element for access control to a machine-type communication device (MTCD); receiving a composite signal based on the SIB-2 message from the MTCD; and performing SIC decoding (Successive Interference Cancellation Decoding) on ​​the SIB-2 message.

[0009] Furthermore, the number of power level pools and the access control elements may be determined based on observation values ​​obtained through local observation of the small base station, based on a mixed network based on QMIX reinforcement learning based on a Dec-POMDP (Decentralized Partially Observable Markov Decision Process) model.

[0010] Furthermore, the hybrid network may be characterized in that it is trained to maximize the benefit of the hybrid network based on the behavior of all small cell base stations within the ultra-dense network including the small cell base station.

[0011] Furthermore, the mixed network may be characterized in that it is trained using sampling data of a replay buffer in which experience from the mixed network is stored.

[0012] Furthermore, the above experience may be characterized as consisting of a previous state of the ultra-dense network, a previous set of actions of the ultra-dense network, a reward function, a state of the ultra-dense network, and a set of actions of the ultra-dense network.

[0013] Furthermore, the observation value by local observation of the small base station may be characterized as being composed of the number of power level pools of the previous time period, the access control element of the previous time period, and the compensation function of the previous time period, which are observation values ​​that the small base station can measure.

[0014] Furthermore, the machine-type communication device may be characterized in that it is set to perform access control based on ACB-BO (Access Class Barring-Backoff) using the access control element.

[0015] Furthermore, the machine-type communication device may be characterized in that it determines transmission power based on an extended exponential path attenuation model.

[0016] Furthermore, the random access procedure may be characterized as a random access procedure based on unauthorized non-orthogonal random access.

[0017] As a technical means for achieving the above-described technical task, according to another embodiment of the present invention, a small cell base station performing a random access procedure based on unlicensed non-orthogonal random access in an ultra-dense network, the small cell base station comprising: a radio frequency unit; a memory; and a processor controlling the radio frequency unit and the memory, wherein the processor transmits an SIB-2 message including the number of power level pools for power allocation and an access control element for access control to a machine-type communication device (MTCD), receives a composite signal based on the SIB-2 message from the MTCD, and performs SIC decoding (Successive Interference Cancellation Decoding) on ​​the SIB-2 message.

[0018] As a technical means for achieving the above-described technical task, according to another embodiment of the present invention, a method for performing a random access procedure of a Machine-Type Communication Device (MTCD) for unlicensed non-orthogonal random access in an ultra-dense network may be characterized by comprising the steps of: receiving, from a small cell base station, a SIB-2 message including the number of power level pools for power allocation and an access control element for access control; and transmitting, to the small cell base station, a composite signal based on the SIB-2 message.

[0019] Any one of the aforementioned problem solving methods can achieve up to 15% higher data decryption rate and lower connection delay in an umMTC environment.

[0020] The effects that can be obtained from the present invention are not limited to the effects mentioned above, and other effects not mentioned can be clearly understood by a person having ordinary skill in the art to which the present invention belongs from the description below.

[0021] FIG. 1 is a reference diagram showing an ultra dense network (UDN) environment to which one embodiment of the present invention can be applied.

[0022] Figure 2 illustrates an unlicensed non-orthogonal random access procedure to which the present invention is applied.

[0023] FIG. 3 is a reference diagram for explaining an embodiment of applying Qmix in UDN (Ultra-Dense Networks) using the Dec-POMDP model proposed in the present invention.

[0024] FIG. 4 is a reference diagram for explaining an embodiment of the Qmix-based power allocation and access control joint optimization technique proposed in the present invention.

[0025] Figure 5 is a graph comparing the SIC decryption success rates of four technologies while changing the number of user devices.

[0026] Figure 6 is a graph comparing the average access delay times of four technologies while changing the number of user devices.

[0027] Figure 7 is a graph comparing the SIC decoding success rates of four technologies while changing the SINR threshold value.

[0028] Figure 8 is a graph comparing the average access delay times of four technologies while changing the SINR threshold value.

[0029] Below, various embodiments are described in detail with reference to the attached drawings. The embodiments described below may be modified and implemented in various different forms. To more clearly explain the features of the embodiments, detailed descriptions of matters widely known to those skilled in the art to which the embodiments pertain below have been omitted. In addition, parts of the drawings that are not related to the description of the embodiments have been omitted, and similar parts have been designated with similar drawing reference numerals throughout the specification.

[0030] Throughout the specification, when a component is said to be "connected" to another component, this includes not only the "direct connection" but also the "connection with other components in between." Furthermore, when a component is said to "include" another component, this does not exclude other components, but rather implies that other components may be included, unless otherwise specifically stated.

[0031] The embodiments will be described in detail with reference to the attached drawings below.

[0032]

[0033] Figure 1 is a reference diagram illustrating an ultra-dense network (UDN) environment to which one embodiment of the present invention can be applied. As shown in Figure 1, the present invention considers an ultra-dense network environment for supporting ultra-high-capacity machine-type communication in next-generation wireless networks.

[0034] Unlike traditional cellular networks, ultra-large-scale network environments consist of numerous small cells, which are small-power base stations designed to increase the coverage and capacity of a wireless network.

[0035] Small cells are primarily used in urban areas or densely populated areas. They cover a much smaller area than conventional large-cell base stations, and are effective in improving data transmission speeds, reducing network congestion, and enhancing signal strength. Therefore, operating multiple, densely populated small cells effectively achieves massive connectivity by distributing the massive number of machine-to-machine communication device connections across multiple small cell base stations (SBSs).

[0036] Machine-to-machine communication devices (MTCDs) perform device connection to adjacent small cell base stations using an unlicensed, non-orthogonal random access method.

[0037] Unlicensed, non-orthogonal random access reduces access latency and improves device connectivity by omitting the access authorization process and transmitting data packets simultaneously with the preamble resource. At the same time, non-orthogonal multiple access improves device connectivity by allowing multiple devices to transmit data using a common spectrum resource. Consequently, data transmitted using the same spectrum resource are separated in the power domain.

[0038] To distinguish transmitted data, base stations utilize a Successive Interference Cancellation (SIC) receiver. A SIC receiver can efficiently separate and receive signals from multiple users simultaneously. For composite signals transmitted over the same spectrum resource, it first detects the strongest signal (i.e., the signal with the highest signal quality), then removes it and extracts the next strongest signal from the remaining signals.

[0039] Accordingly, SIC decoding is successful when the signal-to-interference-plus-noise-ratio (SINR) for the target signal is above a threshold.

[0040] However, factors that can cause SIC decoding failure include power conflicts and high inter-cell interference. Power conflicts occur when multiple machine-to-machine communication devices select the same subchannel and power level, resulting in SIC decoding failure. High inter-cell interference occurs when signals transmitted on the same subchannel from adjacent small cells interfere with each other, preventing the quality difference between signals at the SINR threshold from being guaranteed.

[0041] As shown in Fig. 1, the SIC (Successive Interference Cancellation) decoding process is repeated from the signal with the strongest signal strength to the signal with the weakest signal strength, and if a signal quality below the SINR threshold occurs, SIC decoding of signals with weak reception strength, including that signal, fails.

[0042] Therefore, to increase the SIC decoding success rate, machine-to-machine communication devices must exclusively select subchannels and power levels while maintaining low inter-cell interference so that the quality difference between signals at the SINR threshold is maintained.

[0043] To maintain this signal quality difference, the base station broadcasts a power level pool in advance to machine-type communication devices via the system information block-2 (SIB-2). Accordingly, the machine-type communication devices select a power level randomly from the power level pool to transmit the preamble and data.

[0044] In the present invention, it is necessary to efficiently perform strategic allocation for a machine-type communication device, and for this purpose, it is necessary to perform optimal power allocation. Hereinafter, designing an optimal power level pool is defined as power allocation for a machine-type communication device.

[0045] Through this, the present invention can minimize the SIC decryption failure rate in a network model in which a large number of machine-type communication devices exist.

[0046] That is, the present invention proposes a method for minimizing the SIC decoding failure rate due to interference through joint optimization for optimal power allocation and access control, and considers the trade-off relationship between power collisions and inter-cell interference to calculate optimal power allocation.

[0047] For example, the number of power level pools within a transmit power range that a machine-type communication device can transmit may vary depending on the SINR threshold. If the SINR threshold is low, the base station can increase the number of power level pools to increase the probability that the machine-type communication devices will exclusively select a power level. However, this reduces the difference between power levels, which increases the probability that the target signal will not achieve a value above the SINR threshold due to inter-cell interference. Conversely, increasing the difference between power levels to reduce the impact of interference reduces the probability that the machine-type communication devices will exclusively select a power level. At the same time, the power collision rate and inter-cell interference increase with the number of machine-type communication devices, requiring appropriate access control.

[0048] Therefore, since computing inter-cell interference is a combinatorial optimization problem that is difficult to compute within polynomial time, the present invention proposes an effective method for achieving joint optimization of power allocation and access control below.

[0049]

[0050] Figure 2 illustrates an unlicensed, non-orthogonal random access procedure to which the present invention applies. The present invention utilizes the unlicensed, non-orthogonal random access procedure proposed in the 3GPP standard, but differs from the SIB-2 message in Step 1 of the conventional random access procedure.

[0051] In the present invention, the SIB-2 message includes the number of power level pools for power allocation as an additional component different from the conventional component. and access control elements for access control is added.

[0052] That is, small cell base stations have their own coverage area. Divide the area into equal parts and calculate the distance appropriate for each power level.

[0053] Accordingly, the power level pool calculated by the small cell base station and the distance for each power level are broadcasted through SIB-2 messages, and machine-type communication devices select a power level appropriate for their distance.

[0054] The present invention also considers the Access Class Barring-Backoff (ACB-BO) scheme proposed by 3GPP for access control for machine-type communication devices. ACB-BO (Access Class Barring-Backoff) is a traffic management and control technique primarily used in mobile communication networks, and is used to manage overload situations in wireless networks and promote efficient resource utilization.

[0055] Specifically, ACB-BO has two timers and one access control element. Operates access control elements represents the access barring rate for all access attempts, and each device receives a random value between 0 and 1. and compare the access control element and value by creating In this case, access to small cell base stations is prohibited.

[0056] Machine-type communication devices can determine transmission power based on the distance from the small cell base station. Each machine-type communication device can determine the distance to the small cell base station through continuous communication with the small cell base station. Accordingly, machine-type communication devices measure path loss based on the reception strength and distance of the small cell base station's broadcast signal.

[0057] In ultra-dense networks where the present invention can be applied, stretched exponential path loss is considered as a representative path loss model. Based on the path loss model, each machine-type communication device determines its transmission power as shown in Mathematical Equation 1 below.

[0058]

[0059] Silver base station at It means the second target receiving power intensity, and is the tuning parameter of the path attenuation model. refers to the maximum transmission power of machine-type communication devices.

[0060]

[0061]

[0062] A small cell base station performs SIC decoding on the composite signal transmitted to it from a machine-type communication device. Each signal is first sorted by signal strength, and SINR measurement is performed starting with the signal with the highest received power, as shown in Equation 2 below.

[0063]

[0064] In mathematical expression 2 represents small-scale fading. is a small cell base station It means interference of signals about, is another small cell base station It means interference of signals for small cell base stations. If so, one signal can be successfully extracted. Here refers to the SINR threshold for signal extraction.

[0065]

[0066] As described in Fig. 1, the SIC decoding process is repeated until the signal with the weakest signal strength is reached, and if a signal quality below the SINR threshold occurs, SIC decoding of signals with weak reception strength, including the signal, fails. That is, since failure of the SIC decoding process is directly related to random access failure, the present invention broadcasts optimal power allocation and access control elements through a SIB-2 message, allowing each machine-type communication device to select the optimal radio resource.

[0067]

[0068] Figures 3 and 4 are reference diagrams illustrating embodiments of the power allocation and access control joint optimization technique proposed in the present invention. Referring to Figures 3 and 4, two specific methods of the present invention are proposed. For convenience of explanation, the agent is the same as the small cell base station described above, and can be used interchangeably.

[0069] The first of the methods proposed in the present invention is the design of a Dec-POMDP (Decentralized Partially Observable Markov Decision Process) model for applying multi-agent reinforcement learning, and the second is a power allocation and access control joint optimization technique based on multi-agent reinforcement learning.

[0070] The Dec-POMDP model is a distributed game model for applying multi-agent reinforcement learning. Differentiating itself from conventional Dec-POMDP models, the proposed Dec-POMDP model is designed to allow small cell base stations in Ultra-Dense Networks (UDNs) to act as agents, implementing optimal power allocation and access control factors.

[0071] That is, the Dec-POMDP model proposed in the present invention includes the design of state / action / reward functions to ensure learning convergence of agents in a large-scale network scale and the design of episodes that take into account sporadic data transmission of random access.

[0072] Accordingly, the present invention proposes a power allocation and access control joint optimization technique based on QMIX, a representative multi-agent reinforcement learning model shown in FIG. 3.

[0073] The QMIX network can address the non-stationarity problem, a key issue in multi-agent environments, and each agent achieves optimal power allocation and access control based on historical random access performance results through distributed reinforcement learning.

[0074] This can be expressed as a power allocation and access control joint optimization technique, as illustrated in Fig. 4, and when the optimal power allocation and access control method proposed in the present invention is applied to an existing unlicensed non-orthogonal random access procedure, the random access success rate is improved as the SIC decryption success rate increases.

[0075]

[0076] More specifically, the present invention explores various elements through multi-agent reinforcement learning. The Dec-POMDP model for performing multi-agent reinforcement learning in the present invention can be composed of the seven elements represented in Mathematical Formula 3 below.

[0077]

[0078] represents a set of agents in the Dec-POMDP model, and in the present invention, the agents mean small cell base stations.

[0079] refers to the state information of the environment, and in the present invention, it consists of shared environment information between agents. Therefore, the state includes global network information that cannot be observed by individual agents.

[0080] In the present invention, the environment is a shared condition in which a small base station (SBS, i.e., an agent) operates for an uplink transmission scenario of a machine-to-machine communication device (MTCD) within an ultra-dense network (UDN), and represents the complex interaction of communication, interference, and learning strategies in a dense network setting.

[0081] represents the collective action set of all agents. In this invention, considering the transmission power limitations of machine-type communication devices, a power level pool of 2 to 4 is considered. In addition, access control elements are discretized into 0.1 units, defining a total of 3*10=30 action areas.

[0082] represents the state transition probability, which represents the probability of transition from one state to another based on the joint actions of the agents.

[0083] stands for the reward function. This represents the reward value assigned to the agents' joint action A. Since each agent makes individual decisions, "selfish" behavior can occur, where agents prioritize their own reward over the interests of the entire network. To mitigate this, the system is designed to ensure that all agents receive the same reward from the shared environment for a given joint action, thereby encouraging cooperation among agents.

[0084] Therefore, the purpose of the present invention is to maximize the SIC decryption success rate, so the reward function is designed as in mathematical formula 4.

[0085]

[0086] reward function The value is calculated across the entire network and transmitted to all agents, and the agents have the same reward function. Because they share values, each agent performs collaborative reinforcement learning to maximize overall network performance.

[0087] is a discount factor for future rewards, and is a reward function The value is a discount factor is deducted over time by

[0088] are the observation values ​​of the agents, and the actual agent selects an action based on the observation values / states that it can measure (i.e., locally observe) rather than the state information of the entire environment. Here, the observable / measurable state of an individual agent can be expressed as in mathematical equation 5.

[0089]

[0090] In mathematical equation 5 is the number of power level pools (i.e., power design) in the t-1th time slot. means the access control element in the t-1th time zone. means the reward function value at the t-1th time zone.

[0091] That is, the agents' observations The value is the previous state information of the environment is determined based on the observed values, and individual agents According to, action a i Select (t).

[0092]

[0093]

[0094] Accordingly, a collective action set With, It can be composed of a number of power level pools and a set of access control elements, and can be configured as in mathematical expression 6.

[0095]

[0096] Each agent selects the action corresponding to the maximum Q value for a given observation, based on Q-learning, a typical reinforcement learning method. The Q value is calculated as shown in Equation 7.

[0097]

[0098] In mathematical expression 7, represents the policy set of agents excluding agent i. That is, is the policy of other agents from the perspective of agent i. Considering the Q-value, which represents the limit of the expected reward when taking the observation value zi and the action ai, each agent uses a common reward r to calculate the Q-value based on its own observations and actions. At time t, agent i observes value z i and action a i This is the reward you receive when you take it. is the policy of agent i for the observation value zi and action a i is the probability. is the observation value of agent i at the initial time. is the action of agent i at the initial time.

[0099] That is, with reference to FIG. 3, further explanation is made that each individual agent (i.e., small cell base station) has its own agent network (e.g., Agent Network 1, Agent Network 2, …, Agent Network s), which is composed of one GRU (Gated Recurrent Unit) layer and two fully connected layers.

[0100] Here, individual agents take current actions (i.e., a1(t), a2(t)… a) based on their observations and previous actions. s (t)) is selected and performed.

[0101] Accordingly, the joint action set of all agents is entered into the environment and is in the state (t) and compensation (t) is calculated, and this information is stored as an experience as in Equation 8.

[0102]

[0103] The saved experience is stored in the replay buffer. As much as is sampled, it is used to train the mixture network.

[0104] That is, based on the actions of all agents, the reward function is calculated, and the previous state (C(t-1)), the previous action set (A(t-1)) and the reward function (r(t)), and the state (C(t)) and action set (A(t)) of the next agents are stored in a replay buffer as a series of experiences (E(t)). The replay buffer is used to train the mixture network in batch size ( ) is sampled.

[0105]

[0106] The mixed network generates a total Q-value (Q) based on the individual Q-values ​​derived from each agent network. tot ) and the mixed network considers the degree to which individual Q-values ​​contribute to the total Q-value, and allows each agent to learn in the direction of maximizing the total Q-value.

[0107] Additionally, the mixed network updates the weight parameters of each agent network and provides feedback. Agents then use the provided feedback to estimate their local observations. Choose the appropriate action based on:

[0108]

[0109] Fig. 4 is a reference diagram of a decision code for explaining a QMIX-based power allocation and access control joint optimization technique, and a detailed description of the decision code of Fig. 4 is as follows.

[0110] Each agent has an observation Observe - Choose actions according to greedy methods. Early A high value of leads to sufficient exploration of all devices and as the epoch increases, The value decreases, allowing each agent to choose the optimal learned action, which is given by Equation 9.

[0111]

[0112] Here, a i (t) is the action at a specific time t, and z i (t) is the agent's local observation, a i (t-1) is the action at time t-1, are network parameters for agent i (Lines 4-11).

[0113] The environment trains a hybrid network by utilizing the experiences of batch units. In the QMIX of the present invention, a double deep Q-network is used, and the hybrid network is trained based on the temporal difference error (Lines 17-19).

[0114] Here is a periodically mixed network It refers to the target network parameters that are updated in . Finally, the loss value is calculated and the mixture network is trained using batch gradient descent (Line 21, 22).

[0115]

[0116] According to the above-described Dec-POMDP model and QMIX-based power allocation and access control joint optimization technique, each small cell base station searches for optimal power allocation and access control elements in real time through multi-agent reinforcement learning.

[0117] The generated element values ​​are transmitted to machine-type communication devices via SIB-2 messages. These devices perform transmission power settings and access control procedures as specified in the SIB-2 messages, ultimately improving random access performance.

[0118]

[0119] Below, the simulation results according to the present invention are described.

[0120] Conventional techniques fail to achieve optimal power allocation while simultaneously coordinating inter-cell interference due to the lack of joint optimization for power allocation and access control. In ultra-dense networks, SIC decoding failures due to inter-cell interference directly lead to random access failures, severely reducing random access success rates. In contrast, the present invention coordinates inter-cell interference through access control and improves the exclusive wireless resource occupancy of machine-to-machine communication devices through optimal power allocation.

[0121] At the same time, the method proposed in the present invention is based on low-complexity multi-agent reinforcement learning, so it can replace network congestion in real time by performing online learning, and network congestion at a specific point in time can be appropriately distributed over the time domain through access control, which can have the effect of increasing overall random access.

[0122] The performance of the present invention was evaluated through simulation. In the simulation environment, the number of machine-type communication devices was set to 10 according to the 6G Key Performance Indicator (KPI). 7 Device / km 2 This is considered. Specifically, 50m 2 Up to 25,000 machine-type communication devices are considered within a square radius. Furthermore, a distribution environment of 10,000 to 25,000 devices is considered to measure performance based on the number of machine-type communication devices.

[0123] To account for the sporadic communication traffic of ultra-high-capacity machine-type communication, machine-type communication devices are activated within 10 seconds according to a Beta distribution (α=3. β=4). Upon activation, machine-type communication devices transmit a preamble and a certain size of data according to an unlicensed, non-orthogonal random access procedure. Each machine-type communication device is T DIt has a delay time constraint of 10 seconds, so the total random access execution period is 10+T. D It becomes a snail.

[0124] Small cell areas are evenly divided into power level pools based on distance, with lower target transmit powers set based on distance. This means that devices closer to the small cell base station select relatively higher power levels, while devices farther away select lower power levels.

[0125] To compare the performance of the technique according to the present invention, three benchmark techniques are considered.

[0126] The first benchmark technology is Access Class Barring and Back-off (ACB-BO), a representative access control scheme proposed by 3GPP. Because ACB-BO does not consider power allocation, two benchmark technologies were utilized, each adopting a maximum (power level pool number: 4) and minimum (power level pool number: 2) power level pools.

[0127] The second technology is the benchmark technology PPPD (Prototype Power Pool Design), which considers optimal power allocation without access control.

[0128] The third benchmark technology is Extended Cluster-based Power Selection (ECPS), which allows each machine-to-machine communication device to select the optimal power level and subchannel.

[0129] Because existing methods consider a small, fixed number of machine-type communication devices, we extended the existing technology to avoid radio resource conflicts instead of increasing the number of devices. All methods adhere to the same unlicensed, non-orthogonal random access procedure.

[0130] Figure 5 is a graph comparing the SIC decoding success rates of four technologies while varying the number of user devices. Here, the criterion for SIC decoding success is determined by whether the SINR for the target signal exceeds the SINR threshold. The SIC decoding rate is the ratio of the number of successful SIC decodings to the total number of SIC operations performed. The number of user devices increases from 10,000 to 5,000, with a total of 25,000 machine-type communication devices being considered.

[0131] Figure 6 is a graph comparing the average access delay of four technologies while changing the number of user devices. Access delay refers to the delay time until the actual transmission data is decrypted at the base station, including random access retransmission of machine-type communication devices. Average access delay refers to the average access delay time for all transmitting devices, and in the case of a device that fails random access within the delay time constraint, the delay time constraint T D It is calculated as

[0132] Figure 7 is a graph comparing the SIC decoding success rates of four technologies while varying the SINR threshold value. The SINR threshold is set in advance for all small cell base stations, and a higher SINR threshold indicates a higher sensitivity to inter-cell interference during the SIC decoding process. All schemes showed a decrease in the SIC decoding success rate as the SINR threshold increased, but the SIC decoding success rate according to the present invention achieved the highest performance with a success rate of over 60%, even with an SINR threshold of 4 dB.

[0133] Figure 8 is a graph comparing the average access delay of four technologies while changing the SINR threshold value. It shows that the technology according to the present invention achieved the lowest access delay for various SINR thresholds. As the SINR threshold increased, the average access delay according to the present invention approached the access delay of ECPS. This is because the influence of inter-cell interference increases as the required SINR interval increases, and the advantage of inter-cell interference control according to the present invention decreases.

[0134] Although not limited thereto, the various descriptions, functions, procedures, proposals, methods and / or operational flowcharts of the present invention disclosed in this document may be applied to various fields requiring wireless communication / connection (e.g., 5G) between devices.

[0135] Hereinafter, more specific examples will be provided with reference to the drawings. In the drawings / descriptions below, the same drawing reference numerals may represent identical or corresponding hardware blocks, software blocks, or functional blocks, unless otherwise described.

[0136] Figure 9 illustrates a communication system applied to the present invention.

[0137] Referring to FIG. 9, a communication system applied to the present invention includes a wireless device, a base station, and a network. Here, the wireless device refers to a device that performs communication using a wireless access technology (e.g., 6G, 5G NR (New RAT), LTE (Long Term Evolution)) and may be referred to as a communication / wireless / 5G / 6G device. Although not limited thereto, the wireless device may include a robot (100a), a vehicle (100b-1, 100b-2), an XR (eXtended Reality) device (100c), a hand-held device (100d), a home appliance (100e), an IoT (Internet of Things) device (100f), and an AI device / server (400). For example, the vehicle may include a vehicle equipped with a wireless communication function, an autonomous vehicle, a vehicle capable of performing vehicle-to-vehicle communication, etc. Here, the vehicle may include an Unmanned Aerial Vehicle (UAV) (e.g., a drone). XR devices include AR (Augmented Reality) / VR (Virtual Reality) / MR (Mixed Reality) devices, and can be implemented in the form of HMD (Head-Mounted Device), HUD (Head-Up Display) installed in a vehicle, television, smartphone, computer, wearable device, home appliance, digital signage, vehicle, robot, etc. Mobile devices can include smartphone, smart pad, wearable device (e.g. smartwatch, smartglass), computer (e.g. laptop, etc.), etc. Home appliances can include TV, refrigerator, washing machine, etc.

[0138] IoT devices may include sensors, smart meters, etc. For example, base stations and networks may also be implemented as wireless devices, and a specific wireless device (200a) may act as a base station / network node to other wireless devices.

[0139] Wireless devices (100a to 100f) can be connected to a network (300) via a base station (200). Artificial Intelligence (AI) technology can be applied to the wireless devices (100a to 100f), and the wireless devices (100a to 100f) can be connected to an AI server (400) via the network (300). The network (300) can be configured using a 3G network, a 4G (e.g., LTE) network, a 5G (e.g., NR) network, etc. The wireless devices (100a to 100f) can communicate with each other via the base station (200) / network (300), but can also communicate directly (e.g., sidelink communication) without going through the base station / network.

[0140] For example, vehicles (100b-1, 100b-2) can communicate directly (e.g., V2V (Vehicle to Vehicle) / V2X (Vehicle to Everything) communication). In addition, IoT devices (e.g., sensors) can communicate directly with other IoT devices (e.g., sensors) or other wireless devices (100a to 100f).

[0141] Wireless communication / connection (150a, 150b, 150c) can be established between wireless devices (100a~100f) / base stations (200), and base stations (200) / base stations (200). Here, wireless communication / connection can be achieved through various wireless access technologies (e.g., 5G NR) such as uplink / downlink communication (150a), sidelink communication (150b) (or, D2D communication), and communication between base stations (150c) (e.g., relay, IAB (Integrated Access Backhaul). Through wireless communication / connection (150a, 150b, 150c), wireless devices and base stations / wireless devices, and base stations and base stations can transmit / receive wireless signals to each other. For example, wireless communication / connection (150a, 150b, 150c) can transmit / receive signals through various physical channels. To this end, at least some of various configuration information setting processes for transmitting / receiving wireless signals, various signal processing processes (e.g., channel encoding / decoding, modulation / demodulation, resource mapping / demapping, etc.), and resource allocation processes can be performed based on various proposals of the present invention.

[0142]

[0143] Figure 10 illustrates a wireless device applicable to the present invention.

[0144] Referring to FIG. 10, the first wireless device (100) and the second wireless device (200) can transmit and receive wireless signals through various wireless access technologies (e.g., LTE, NR). Here, {the first wireless device (100), the second wireless device (200)} can correspond to {the wireless device (100x), the base station (200)} and / or {the wireless device (100x), the wireless device (100x)} of FIG. 13.

[0145] A first wireless device (100) includes one or more processors (102) and one or more memories (104), and may further include one or more transceivers (106) and / or one or more antennas (108). The processor (102) controls the memories (104) and / or the transceivers (106), and may be configured to implement the descriptions, functions, procedures, proposals, methods, and / or operational flowcharts disclosed in this document. For example, the processor (102) may process information in the memory (104) to generate first information / signal, and then transmit a wireless signal including the first information / signal via the transceiver (106). In addition, the processor (102) may receive a wireless signal including second information / signal via the transceiver (106), and then store information obtained from signal processing of the second information / signal in the memory (104). The memory (104) may be connected to the processor (102) and may store various information related to the operation of the processor (102). For example, the memory (104) may perform some or all of the processes controlled by the processor (102), or may store software code including commands for performing the descriptions, functions, procedures, proposals, methods, and / or operation flowcharts disclosed in this document. Here, the processor (102) and the memory (104) may be part of a communication modem / circuit / chip designed to implement a wireless communication technology (e.g., LTE, NR). The transceiver (106) may be connected to the processor (102) and may transmit and / or receive a wireless signal via one or more antennas (108).

[0146] The transceiver (106) may include a transmitter and / or a receiver. The transceiver (106) may be used interchangeably with an RF (Radio Frequency) unit. In the present invention, a wireless device may also refer to a communication modem / circuit / chip.

[0147] The second wireless device (200) includes one or more processors (202), one or more memories (204), and may further include one or more transceivers (206) and / or one or more antennas (208). The processor (202) controls the memories (204) and / or the transceivers (206), and may be configured to implement the descriptions, functions, procedures, proposals, methods, and / or operational flowcharts disclosed in this document. For example, the processor (202) may process information in the memory (204) to generate third information / signals, and then transmit a wireless signal including the third information / signals via the transceivers (206). Furthermore, the processor (202) may receive a wireless signal including fourth information / signals via the transceivers (206), and then store information obtained from signal processing of the fourth information / signals in the memory (204). The memory (204) may be connected to the processor (202) and may store various information related to the operation of the processor (202). For example, the memory (204) may perform some or all of the processes controlled by the processor (202), or may store software code including commands for performing the descriptions, functions, procedures, proposals, methods, and / or operation flowcharts disclosed in this document. Here, the processor (202) and the memory (204) may be part of a communication modem / circuit / chip designed to implement a wireless communication technology (e.g., LTE, NR). The transceiver (206) may be connected to the processor (202) and may transmit and / or receive a wireless signal via one or more antennas (208).

[0148] The transceiver (206) may include a transmitter and / or a receiver. The transceiver (206) may be used interchangeably with an RF unit. In the present invention, a wireless device may also mean a communication modem / circuit / chip.

[0149] Hereinafter, the hardware elements of the wireless device (100, 200) will be described in more detail. Although not limited thereto, one or more protocol layers may be implemented by one or more processors (102, 202). For example, one or more processors (102, 202) may implement one or more layers (e.g., functional layers such as PHY, MAC, RLC, PDCP, RRC, SDAP). One or more processors (102, 202) may generate one or more Protocol Data Units (PDUs) and / or one or more Service Data Units (SDUs) according to the descriptions, functions, procedures, proposals, methods, and / or operation flowcharts disclosed in this document. One or more processors (102, 202) may generate messages, control information, data, or information according to the descriptions, functions, procedures, proposals, methods, and / or operation flowcharts disclosed in this document. One or more processors (102, 202) can generate signals (e.g., baseband signals) including PDUs, SDUs, messages, control information, data or information according to the functions, procedures, proposals and / or methods disclosed herein, and provide the signals to one or more transceivers (106, 206). One or more processors (102, 202) can receive signals (e.g., baseband signals) from one or more transceivers (106, 206) and obtain PDUs, SDUs, messages, control information, data or information according to the descriptions, functions, procedures, proposals, methods and / or operational flowcharts disclosed herein.

[0150] One or more processors (102, 202) may be referred to as a controller, a microcontroller, a microprocessor, or a microcomputer. One or more processors (102, 202) may be implemented by hardware, firmware, software, or a combination thereof. For example, one or more Application Specific Integrated Circuits (ASICs), one or more Digital Signal Processors (DSPs), one or more Digital Signal Processing Devices (DSPDs), one or more Programmable Logic Devices (PLDs), or one or more Field Programmable Gate Arrays (FPGAs) may be included in one or more processors (102, 202). The descriptions, functions, procedures, proposals, methods, and / or operational flowcharts disclosed in this document may be implemented using firmware or software, and the firmware or software may be implemented to include modules, procedures, functions, etc. The descriptions, functions, procedures, suggestions, methods and / or operation flowcharts disclosed in this document may be implemented using firmware or software configured to perform one or more processors (102, 202) or stored in one or more memories (104, 204) and executed by one or more processors (102, 202). The descriptions, functions, procedures, suggestions, methods and / or operation flowcharts disclosed in this document may be implemented using firmware or software in the form of codes, instructions and / or sets of instructions.

[0151] One or more memories (104, 204) may be coupled to one or more processors (102, 202) and may store various forms of data, signals, messages, information, programs, codes, instructions, and / or commands. The one or more memories (104, 204) may be configured as ROM, RAM, EPROM, flash memory, hard drives, registers, cache memory, computer-readable storage media, and / or combinations thereof. The one or more memories (104, 204) may be located internally and / or externally to the one or more processors (102, 202). Additionally, the one or more memories (104, 204) may be coupled to the one or more processors (102, 202) via various technologies, such as wired or wireless connections.

[0152] One or more transceivers (106, 206) can transmit user data, control information, wireless signals / channels, etc., as mentioned in the methods and / or flowcharts of this document, to one or more other devices. One or more transceivers (106, 206) can receive user data, control information, wireless signals / channels, etc., as mentioned in the descriptions, functions, procedures, proposals, methods and / or flowcharts of this document, from one or more other devices. For example, one or more transceivers (106, 206) can be connected to one or more processors (102, 202) and can transmit and receive wireless signals. For example, one or more processors (102, 202) can control one or more transceivers (106, 206) to transmit user data, control information, or wireless signals to one or more other devices. Additionally, one or more processors (102, 202) may control one or more transceivers (106, 206) to receive user data, control information, or wireless signals from one or more other devices. Additionally, one or more transceivers (106, 206) may be coupled to one or more antennas (108, 208), and one or more transceivers (106, 206) may be configured to transmit and receive user data, control information, wireless signals / channels, or the like, as referred to in the descriptions, functions, procedures, proposals, methods, and / or operational flowcharts disclosed herein, via one or more antennas (108, 208). In this document, one or more antennas may be multiple physical antennas or multiple logical antennas (e.g., antenna ports). One or more transceivers (106, 206) can convert received user data, control information, wireless signals / channels, etc. from RF band signals to baseband signals in order to process the received user data, control information, wireless signals / channels, etc. using one or more processors (102, 202).One or more transceivers (106, 206) may convert user data, control information, wireless signals / channels, etc. processed by one or more processors (102, 202) from baseband signals to RF band signals. For this purpose, one or more transceivers (106, 206) may include an (analog) oscillator and / or filter.

[0153] It is clear that the examples of the proposed methods described above can also be considered as a type of proposed methods, as they can be included as one of the implementation methods of the present disclosure. Furthermore, the proposed methods described above can be implemented independently, but can also be implemented in the form of a combination (or merge) of some of the proposed methods. A rule can be defined so that the base station notifies the terminal of whether the proposed methods are applicable (or information about the rules of the proposed methods) through a predefined signal (e.g., a physical layer signal or a higher layer signal).

[0154] The present disclosure may be embodied in other specific forms without departing from the technical ideas and essential features described herein. Therefore, the above detailed description should not be construed as limiting in all respects but rather as illustrative. The scope of the present disclosure should be determined by a reasonable interpretation of the appended claims, and all modifications within the equivalent scope of the present disclosure are intended to be included within the scope of the present disclosure. Furthermore, claims that do not explicitly cite each other in the claims may be combined to form embodiments or incorporated into new claims through post-filing amendments.

Claims

1. A method for performing a random access procedure of a small cell base station in an ultra-dense network, A step of transmitting a SIB-2 message including a number of power level pools for power allocation and an access control element for access control to a Machine-Type Communication Device (MTCD); A step of receiving a synthetic signal based on the SIB-2 message from the machine-type communication device; and Comprising a step of performing SIC decoding (Successive Interference Cancellation Decoding) on ​​the above SIB-2 message. How to perform random access procedures.

2. In paragraph 1, The number of power level pools and the access control elements are, Based on a mixed network based on QMIX reinforcement learning based on the Dec-POMDP (Decentralized Partially Observable Markov Decision Process) model, characterized in that it is determined according to the observation value by local observation of the small base station. How to perform random access procedures.

3. In paragraph 2, The above mixed network is, characterized in that the hybrid network is trained based on the behavior of all small cell base stations in the ultra-dense network including the small cell base station, so as to maximize the benefit of the hybrid network. How to perform random access procedures.

4. In paragraph 3, The above mixed network is, characterized in that the training is performed using sampled data of a replay buffer in which experience from the above mixed network is stored. How to perform random access procedures.

5. In paragraph 4, The above experience is, The previous state of the above ultra-dense network. The previous action set of the above ultra-dense network, the reward function, the state of the above ultra-dense network. The action set of the above ultra-dense network is characterized by being composed of, How to perform random access procedures.

6. In paragraph 5, The above reward function is, As expressed by the following mathematical formula 1, [Mathematical Formula 1] (Here, R(t) is a reward function, n and N are natural numbers, t is a specific time, SINR n (t) is the signal-to-interference-plus-noise ratio of the nth machine-type communication device, R th is the SINR threshold for signal extraction) How to perform random access procedures.

7. In paragraph 2, The observation values ​​by local observation of the above small base station are, The above small base station is characterized in that the measurable observation values ​​are composed of the number of power level pools of the previous time period, the access control elements of the previous time period, and the reward function of the previous time period. How to perform random access procedures.

8. In paragraph 2, The behavior of the above small base station is characterized in that it is determined according to the following mathematical expression 2. [Mathematical formula 2] (Here, a i (t) is the action at a specific time t, z i (t) is the observation value by local observation of agent i, a i (t-1) is the action at time t-1, are network parameters for agent i) How to perform random access procedures.

9. In paragraph 2, The behavior of the above small base station is characterized in that it is determined according to the maximum value of Q according to the following mathematical expression 3. [Mathematical Formula 3] (Here, is the policy set of agents excluding agent i, is the policy of other agents from the perspective of agent i. Q-value, which represents the limit of the expected reward when taking the observation value zi and action ai, considering . At time t, agent i observes value z i and action a i The reward you receive when you take it, is the policy of agent i for the observation value z. i and action a i The probability of is the observation value of agent i at the initial time, is the action of agent i at initial time) How to perform random access procedures.

10. In paragraph 1, The above mechanical communication device, It is characterized in that it is set to perform access control based on ACB-BO (Access Class Barring-Backoff) using the above access control element. How to perform random access procedures.

11. In paragraph 1, The above mechanical communication device, characterized in that the transmission power is determined based on an extended exponential path attenuation model. How to perform random access procedures.

12. In paragraph 11, The above transmission power is characterized by being determined by the following mathematical formula 4. [Mathematical formula 4] (Here, Silver base station at th target receiving power intensity, and is the tuning parameter of the path attenuation model, is the maximum transmission power of mechanical communication devices) How to perform random access procedures.

13. In paragraph 1, The above random access procedure is, Characterized by a random access procedure based on non-permissioned non-orthogonal random access, How to perform random access procedures.

14. In a small cell base station performing a random access procedure based on unlicensed non-orthogonal random access in an ultra-dense network, radio frequency unit; memory; and A processor comprising: a radio frequency unit and a memory; The above processor, Transmitting a SIB-2 message containing the number of power level pools for power allocation and an access control element for access control to a Machine-Type Communication Device (MTCD), From the above machine-type communication device, a synthetic signal based on the SIB-2 message is received, Performing SIC decoding (Successive Interference Cancellation Decoding) on ​​the above SIB-2 message. Small cell base station.

15. A method for performing a random access procedure of a machine-type communication device (MTCD) for unauthorized non-orthogonal random access in an ultra-dense network, A step of receiving, from a small cell base station, a SIB-2 message including a number of power level pools for power allocation and an access control element for access control; A step of transmitting a composite signal based on the SIB-2 message to the small cell base station, How to perform random access procedures.