Method and device for scheduling in wireless communication system
An AI model trained through reinforcement learning optimizes data transmission in high-frequency wireless communication systems by determining optimal scheduling policies, addressing power consumption and latency issues in 6G networks.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2026-01-21
- Publication Date
- 2026-07-30
AI Technical Summary
Existing wireless communication systems face challenges in efficiently managing data transmission in high-frequency bands, particularly in 6G mobile communication technology, where explosive device connectivity and increased bandwidth requirements lead to high power consumption and latency, necessitating improved scheduling and resource allocation.
Implementing an artificial intelligence model trained through reinforcement learning to determine optimal data transmission policies, considering minimum delay and resource block size, to minimize power consumption and latency while ensuring quality of service.
The AI model optimizes data scheduling, reducing power consumption and latency, thereby enhancing the efficiency and performance of wireless communication systems.
Smart Images

Figure KR2026001265_30072026_PF_FP_ABST
Abstract
Description
Method and device for scheduling in a wireless communication system
[0001] The present disclosure relates to a wireless communication system, and more specifically to a method and apparatus for scheduling in a wireless communication system.
[0002] 5G mobile communication technology defines a wide frequency band to enable fast transmission speeds and new services, and can be implemented not only in frequency bands below 6 GHz ('Sub 6 GHz'), such as 3.5 gigahertz (3.5 GHz), but also in ultra-high frequency bands called millimeter waves (mmWave), such as 28 GHz and 39 GHz ('Above 6 GHz'). In addition, for 6G mobile communication technology, which is referred to as a system beyond 5G, implementation in the terahertz (THX) band (e.g., the 3 terahertz band at 95 GHz) is being considered to achieve transmission speeds 50 times faster and ultra-low latency reduced to one-tenth compared to 5G mobile communication technology.
[0003] In the early stages of 5G mobile communication technology, aiming to satisfy service support and performance requirements for enhanced Mobile BroadBand (eMBB), Ultra-Reliable Low-Latency Communications (URLLC), and massive Machine-Type Communications (mMTC), technologies such as beamforming and Massive MIMO to mitigate path loss and increase transmission distance in ultra-high frequency bands, support for various numerologies (such as the operation of multiple subcarrier spacings) and dynamic operation of slot formats for the efficient utilization of ultra-high frequency resources, initial access techniques to support multi-beam transmission and broadband, definition and operation of Band-Width Parts (BWP), Low Density Parity Check (LDPC) codes for high-volume data transmission, new channel coding methods such as Polar Codes for the reliable transmission of control information, and L2 pre-processing (L2 Standardization has been carried out for pre-processing, network slicing which provides a dedicated network specialized for specific services, and other methods.
[0004] When 5G mobile communication systems are commercialized, connected devices, which are increasing explosively, will be connected to communication networks. Accordingly, it is expected that there will be a need to enhance the functionality and performance of 5G mobile communication systems and to integrate the operation of connected devices. To this end, new research is planned to be conducted on 5G performance improvement and complexity reduction, support for AI services, support for metaverse services, and drone communication using eXtended Reality (XR), Artificial Intelligence (AI), and Machine Learning (ML) to efficiently support Augmented Reality (AR), Virtual Reality (VR), and Mixed Reality (MR).
[0005] Furthermore, the advancement of these 5G mobile communication systems encompasses multi-antenna transmission technologies such as new waveforms to guarantee coverage in the terahertz band of 6G mobile communication technology, Full Dimensional MIMO (FD-MIMO), array antennas, and large-scale antennas; metamaterial-based lenses and antennas to improve terahertz band signal coverage; high-dimensional spatial multiplexing technology using OAM (Orbital Angular Momentum); and Reconfigurable Intelligent Surface (RIS) technology; as well as Full Duplex technology for enhancing frequency efficiency and system networks in 6G mobile communication technology; AI-based communication technologies that realize system optimization by utilizing satellites and AI from the design stage and internalizing end-to-end AI support functions; and the realization of services of complexity exceeding the limits of terminal computing capabilities by utilizing ultra-high-performance communication and computing resources. It could serve as a foundation for the development of next-generation distributed computing technologies.
[0006] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art in relation to the present disclosure.
[0007] In a wireless communication system according to the present disclosure, a base station may include a transceiver, a memory for storing instructions, and at least one processor. The instructions are executed individually or collectively by the at least one processor so that the base station identifies, in a time slot, a minimum delay required time for each data packet requiring transmission and the size of a resource block required to transmit each data packet, and for the time slot, applies the identified minimum delay required time and the size of the required resource block to an artificial intelligence model, wherein the artificial intelligence model is trained to receive input data related to the minimum delay required time for each data packet requiring transmission and data related to the size of the resource block required to transmit each data packet requiring transmission, and outputs a value related to whether to transmit each data packet in the time slot, and determines whether to transmit each data packet through the transceiver in the time slot based on the value output from the artificial intelligence model.
[0008] In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components.
[0009] FIG. 1 illustrates a wireless communication system according to various embodiments of the present disclosure.
[0010] FIG. 2 illustrates an example of a fronthaul structure according to the function split of a base station according to various embodiments of the present disclosure.
[0011] FIG. 3 illustrates the configuration of a distributed unit (DU) according to various embodiments of the present disclosure.
[0012] FIG. 4 illustrates the configuration of a radio unit (RU) according to various embodiments of the present disclosure.
[0013] FIG. 5 is a flowchart illustrating a method for an electronic device according to one embodiment to perform data scheduling.
[0014] FIG. 6 is a block diagram illustrating data received by an artificial intelligence model according to one embodiment of the present disclosure and a value output from the artificial intelligence model accordingly.
[0015] FIG. 7 is a flowchart illustrating the process of training an artificial intelligence model according to one embodiment.
[0016] FIG. 8 is a flowchart illustrating a specific method for performing data scheduling using an artificial intelligence model in one embodiment.
[0017] FIG. 9 is an example diagram illustrating the delay time and number of transmissions obtained by adjusting the weights for each policy for training an artificial intelligence model according to one embodiment.
[0018] FIG. 10a is an example diagram illustrating the effect of performing data scheduling using an artificial intelligence model trained according to one embodiment.
[0019] FIG. 10b is an example diagram illustrating the effect of performing data scheduling using an artificial intelligence model trained according to one embodiment.
[0020] Embodiments of the present disclosure are described below in detail with reference to the attached drawings so that those skilled in the art can easily implement them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein. Furthermore, in order to clearly explain the present disclosure in the drawings, parts unrelated to the explanation have been omitted, and similar parts throughout the specification are denoted by similar reference numerals.
[0021] The terms used in this disclosure are described in their current, general form considering the functions mentioned herein; however, they may refer to various other terms depending on the intent of those skilled in the art, case law, or the emergence of new technologies. Accordingly, the terms used in this disclosure should not be interpreted solely by their names, but should be interpreted based on the meaning of the terms and the overall content of this disclosure.
[0022] Additionally, terms such as the first, second, third, ..., Nth may be used to describe various components, but the components should not be limited by these terms. These terms are used for the purpose of distinguishing one component from another.
[0023] Throughout the specification, when a part is described as being "connected" to another part, this includes not only cases where they are "directly connected," but also cases where they are "electrically connected" with other components interposed between them. Furthermore, when a part is described as "including" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components.
[0024] Phrases such as "in one embodiment" appearing in various places in this disclosure do not necessarily refer to the same embodiment.
[0025] One embodiment of the present disclosure may be represented by functional block configurations and various processing steps. Some or all of these functional blocks may be implemented by various numbers of hardware and / or software configurations that execute specific functions. For example, the functional blocks of the present disclosure may be implemented by one or more microprocessors or by circuit configurations for a specific function. Additionally, for example, the functional blocks of the present disclosure may be implemented in various programming or scripting languages. The functional blocks may be implemented as algorithms executed on one or more processors. Furthermore, the present disclosure may employ prior art for electronic configuration, signal processing, and / or data processing. Terms such as “mechanism,” “element,” “means,” and “configuration” may be used broadly and are not limited to mechanical and physical configurations.
[0026] Furthermore, the connecting lines or connecting members between the components depicted in the drawings are merely illustrative of functional connections and / or physical or circuit connections. In the actual device, connections between components may be represented by various alternative or added functional connections, physical connections, or circuit connections.
[0027] Artificial intelligence technology consists of machine learning (e.g., deep learning) and elemental technologies utilizing machine learning.
[0028] Machine learning is an algorithmic technology that classifies and learns the features of input data on its own, and the elemental technology is a technology that mimics functions such as cognition and judgment of the human brain by utilizing deep learning machine learning algorithms, and consists of the fields of linguistic understanding, visual understanding, reasoning / prediction, knowledge representation, and motion control.
[0029] The various fields where artificial intelligence technology is applied are as follows. Linguistic understanding is a technology that recognizes, applies, and processes human language and text, and includes natural language processing, machine translation, dialogue systems, question answering, and speech recognition / synthesis. Visual understanding is a technology that perceives and processes objects like human vision, and includes object recognition, object tracking, image search, person recognition, scene understanding, spatial understanding, and image enhancement. Inference and prediction is a technology that judges information to logically infer and predict, and includes knowledge / probability-based inference, optimization prediction, preference-based planning, and recommendation. Knowledge representation is a technology that automatically processes human experiential information into knowledge data, and includes knowledge construction (data generation / classification) and knowledge management (data utilization). Motion control is a technology that controls the autonomous driving of vehicles and the movement of robots, and includes motion control (navigation, collision, driving) and manipulation control (behavior control).
[0030] Reinforcement learning is a machine learning technique that enables an agent to learn an optimal policy of action through interaction in a given environment. The agent observes state information from the environment, selects and performs specific actions accordingly, and receives rewards or penalties as a result. Based on these rewards, the agent can explore and learn a policy of action that maximizes long-term cumulative rewards. Reinforcement learning is characterized by its ability to learn by accumulating experience in simulations or real-world environments without data or prior knowledge, thereby enabling the resolution of complex decision-making problems. In the present disclosure, the artificial intelligence model may include an artificial intelligence model trained through reinforcement learning.
[0031] Reinforcement learning can be composed of three elements: state, action, and reward. The state represents the current situation of the environment, and the agent can decide on its next action based on this. An action is one of various strategies the agent can choose, allowing it to perform an appropriate action in a specific situation. The reward is an evaluation criterion that expresses the result of the selected action numerically; through this, the agent can learn the optimal policy based on experience.
[0032] The present disclosure proposes an energy-saving scheduler for data management and power optimization of a base station, which can operate based on reinforcement learning. The elements of reinforcement learning, namely state, action, and reward, can be applied to the system of the present disclosure as follows.
[0033] In the 'state observation' stage, the state(s) can be verified by measuring the number of data requests from users and packets waiting in the queue (unit: resource block, m) and the minimum delay requirement (d) for each time slot. Accordingly, the state information in the present invention includes the total sum (m) of currently waiting resource blocks and the minimum delay requirement (d), and this may be information reflecting the urgency of transmission related to the amount of data and processing time that the system must process.
[0034] 'Action' can refer to a movement that determines whether or not to transmit data. Transmission ( ) and non-transmission( It may be composed of two parts. An electronic device or method of operation according to the present disclosure may provide an electronic device or method that determines whether to transmit data by selecting the most appropriate action in the current state.
[0035] 'Compensation' can be designed based on multi-objective reinforcement learning techniques with the goal of saving power to the base station and minimizing delays in data transmission. In this disclosure, the compensation structure is set in a vector form rather than a scalar form, and data packet transmission ( ) and waiting( A compensation of -1 can be granted for each according to ). Through this, the electronic device or method according to the present disclosure can make an optimal decision to minimize data delay time while saving the power of the base station as much as possible.
[0036] In the present disclosure, state information is an element representing the situation that an agent can observe in the current environment and may include data necessary for the agent to determine an optimal action. State information may quantify or quantitatively express the current state of the environment so that the agent can select a course of action based on it during the learning and decision-making process.
[0037] In the present disclosure, the first policy may be an AS (always send) policy. In one embodiment, the first policy may include a policy of always transmitting data packets stored in a data buffer regardless of conditions. An electronic device according to one embodiment may continuously transmit data packets during a time slot when it decides to schedule data according to the first policy in a time slot.
[0038] In the present disclosure, the second policy may be an always-wait (AW) policy. In one embodiment, the second policy may include a policy of not always transmitting data packets. An electronic device according to one embodiment may not transmit data packets during a time slot if it decides to perform data scheduling according to the second policy in a time slot. For example, if it decides to perform data scheduling according to the second policy in a time slot, it may not transmit any data packets during the time slot. For example, if it decides to perform data scheduling according to the second policy in a time slot, it may transmit data packets only when the minimum delay required time of the data packets stored in the data buffer becomes zero during the time slot.
[0039] The method according to the present disclosure may be performed by an electronic device. In the present disclosure, the electronic device may include a base station or a device of a base station. The device of a base station may include at least one of a centralized unit (CU), a distributed unit (DU), or a radio unit (RU) of a base station. However, it is not limited thereto.
[0040] The method according to the present disclosure may be performed by an external electronic device of a base station. The external electronic device of the base station may include an electronic device that supports a server or network connection. However, it is not limited thereto.
[0041] The present disclosure will be described in detail below with reference to the attached drawings.
[0042] FIG. 1 is a drawing illustrating a wireless communication system according to various embodiments of the present disclosure.
[0043] FIG. 1 illustrates a base station (110), a terminal (120), and a terminal (130) as part of nodes utilizing a wireless channel in a wireless communication system. FIG. 1 illustrates only one base station, but other base stations identical or similar to the base station (110) may be included.
[0044] A base station (110) is a network infrastructure that provides wireless access to terminals (120, 130). The base station (110) has coverage defined as a certain geographical area based on the distance at which it can transmit signals. In addition to being a base station, the base station (110) may be referred to as an 'access point (AP)', 'eNodeB (eNB)', '5G node (5th generation node)', 'next generation nodeB (gNB)', 'wireless point', 'transmission / reception point (TRP)', or other terms having an equivalent technical meaning.
[0045] Each of the terminal (120) and terminal (130) is a device used by a user and communicates with the base station (110) via a wireless channel. The link from the base station (110) toward the terminal (120) or terminal (130) is referred to as a downlink (DL), and the link from the terminal (120) or terminal (130) toward the base station (110) is referred to as an uplink (UL). Additionally, the terminal (120) and terminal (130) can communicate with each other via a wireless channel. At this time, the link between the terminal (120) and terminal (130) (device-to-device link, D2D) is referred to as a sidelink, and the sidelink may be used interchangeably with the PC5 interface. In some cases, at least one of the terminal (120) and terminal (130) may be operated without user intervention. That is, at least one of the terminal (120) and the terminal (130) is a device that performs machine type communication (MTC) and may not be carried by a user. Each of the terminal (120) and the terminal (130) may be referred to as 'user equipment (UE)', 'mobile station', 'subscriber station', 'remote terminal', 'wireless terminal', or 'user device', or other terms having an equivalent technical meaning, in addition to 'terminal'.
[0046] The base station (110), terminal (120), and terminal (130) can perform beamforming. The base station and the terminal can transmit and receive wireless signals in a relatively low frequency band (e.g., FR1 (frequency range 1) of NR). Additionally, the base station and the terminal can transmit and receive wireless signals in a relatively high frequency band (e.g., FR2 of NR, millimeter wave (mmWave) band (e.g., 28 GHz, 30 GHz, 38 GHz, 60 GHz)). In some embodiments, the base station (110) can communicate with the terminal (120) within a frequency range corresponding to FR1. In some embodiments, the base station can communicate with the terminal (120) within a frequency range corresponding to FR2. At this time, to improve channel gain, the base station (110), terminal (120), and terminal (130) can perform beamforming. Here, beamforming may include transmit beamforming and receive beamforming. That is, the base station (110), terminal (120), and terminal (130) can give directivity to the transmitted signal or the received signal. To this end, the base station (110) and terminals (120, 130) can select serving beams (112, 113, 121, 131) through a beam search or beam management procedure. After the serving beams (112, 113, 121, 131) are selected, subsequent communication can be performed through a resource that is in a quasi-co-located (QCL) relationship with the resource that transmitted the serving beams (112, 113, 121, 131).
[0047] If large-scale characteristics of the channel transmitting the symbol on the first antenna port can be inferred from the channel transmitting the symbol on the second antenna port, the first antenna port and the second antenna port may be evaluated to have a QCL relationship. For example, the large-scale characteristics may include at least one of a delay spread, a Doppler spread, a Doppler shift, an average gain, an average delay, and a spatial receiver parameter.
[0048] Although FIG. 1a illustrates that both the base station and the terminal perform beamforming, various embodiments of the present disclosure are not necessarily limited thereto. In some embodiments, the terminal may or may not perform beamforming. Additionally, the base station may or may not perform beamforming. That is, either the base station or the terminal may perform beamforming, or neither the base station nor the terminal may perform beamforming.
[0049] In the present disclosure, a beam refers to a spatial flow of a signal in a wireless channel, formed by one or more antennas (or antenna elements), and this formation process may be referred to as beamforming. Beamforming may include analog beamforming and digital beamforming (e.g., precoding). Reference signals transmitted based on beamforming may include, for example, DM-RS (demodulation-reference signal), CSI-RS (channel state information-reference signal), SS / PBCH (synchronization signal / physical broadcast channel), and SRS (sounding reference signal). Additionally, as a configuration for each reference signal, an IE such as a CSI-RS resource or an SRS-resource may be used, and such a configuration may include information associated with the beam. Information associated with a beam may refer to whether the relevant configuration (e.g., CSI-RS resource) uses the same spatial domain filter as other configurations (e.g., other CSI-RS resources within the same CSI-RS resource set) or a different spatial domain filter, or whether it is quasi-colocated (QCL) with a reference signal, and if so, what type (e.g., QCL type A, B, C, D).
[0050] Conventionally, in communication systems with a relatively large cell radius of base stations, each base station was installed to include the functions of a digital processing unit (or DU (digital unit)) and a radio frequency (RF) processing unit (or RU (radio unit)). However, in 4G (4th generation) and / or later communication systems, as high frequency bands are used and the cell radius of base stations decreases, the number of base stations required to cover a specific area has increased, and the burden of installation costs for operators to install these increased base stations has increased. To minimize base station installation costs, a structure has been proposed in which the DU and RU of a base station are separated, with one or more RUs connected to a single DU via a wired network, and one or more geographically distributed RUs deployed to cover a specific area. Below, base station deployment structures and extension examples according to various embodiments of the present disclosure are described with reference to FIG. 1b.
[0051] FIG. 2 is a drawing illustrating an example of a fronthaul structure according to the function split of a base station according to various embodiments of the present disclosure.
[0052] Fronthaul refers to the space between entities between a wireless LAN and a base station, unlike backhaul between a base station and a core network. Although FIG. 1b discloses a DU and an RU, it is not limited thereto. In various embodiments of the present disclosure, DU and RU may respectively refer to O-DU and O-RU, which are terms in the O-RAN specification, and the aforementioned terms may be used interchangeably. FIG. 2 illustrates an example of a fronthaul structure between a DU (160) and a RU (180), but this is merely for convenience of explanation and the present disclosure is not limited thereto. In other words, embodiments of the present disclosure may also be applied to a fronthaul structure between a single O-DU and a plurality of O-RUs. For example, embodiments of the present disclosure may be applied to a fronthaul structure between a single O-DU and two O-RUs. Additionally, embodiments of the present disclosure may also be applied to a fronthaul structure between a single O-DU and three O-RUs.
[0053] Referring to FIG. 2, the base station (110) may include a DU (160) and an RU (180). The fronthole (170) between the DU (160) and the RU (180) may be operated via an Fx interface. Various fronthole interfaces defined in the standard (e.g., eCPRI (enhanced common public radio interface), ROE (radio over ethernet)) may be used for the operation of the fronthole (170).
[0054] As communication technology develops, mobile data traffic increases, and consequently, the bandwidth requirements for the fronthaul between the digital unit and the wireless unit have increased significantly. In deployments such as a centralized / cloud radio access network (C-RAN), the DU performs functions for the packet data convergence protocol (PDCP), radio link control (RLC), media access control (MAC), and physical (PHY), while the RU can be implemented to perform additional functions for the PHY layer in addition to radio frequency (RF) functions.
[0055] The DU (160) can perform upper-layer functions of the wireless network. For example, the DU (160) can perform functions of the MAC layer and parts of the PHY layer. Here, parts of the PHY layer are functions of the PHY layer that are performed at a higher level, and may include, for example, channel encoding (or channel decoding), scrambling (or descrambling), modulation (or demodulation), and layer mapping (or layer demapping). According to one embodiment, if the DU (160) conforms to the O-RAN standard, it may be referred to as an O-DU (O-RAN DU).
[0056] The RU (180) can perform lower-layer functions of the wireless network. For example, the RU (180) can perform RF functions, which are part of the PHY layer. Here, part of the PHY layer refers to functions of the PHY layer that are performed at a level relatively lower than that of the DU (160), and may include, for example, inverse fast Fourier transform (IFFT) transformation (or fast Fourier transform) transformation, cyclic prefix (CP) insertion (CP removal), and digital beamforming. Examples of such specific functional separation are described in detail in FIG. 4. The RU (180) may be referred to as an 'access unit (AU)', 'access point (AP)', 'transmission / reception point (TRP)', 'remote radio head (RRH)', 'radio unit (RU)', or other terms having an equivalent technical meaning. According to one embodiment, if the RU (180) conforms to the O-RAN standard, it may be referred to as an O-RU (O-RAN RU).
[0057] Although FIG. 1b describes a base station comprising a DU and an RU, various embodiments of the present disclosure are not limited thereto. In some embodiments, the base station may be implemented in a distributed deployment according to a centralized unit (CU) configured to perform the functions of the upper layers of the access network (e.g., packet data convergence protocol, RRC (PDCP)) and a distributed unit (DU) configured to perform the functions of the lower layers. In this case, the distributed unit (DU) may include the digital unit (DU) and the radio unit (RU) of FIG. 1b. Between a core network (e.g., 5G core or next generation core (NGC)) and a radio network (RAN), the base station may be implemented in a structure arranged in the order of CU, DU, and RU. The interface between the CU and the distributed unit (DU) may be referred to as the F1 interface.
[0058] A centralized unit (CU) is connected to one or more DUs and can perform functions at a higher layer than the DUs. For example, the CU may perform functions at the radio resource control (RRC) and packet data convergence protocol (PDCP) layers, while the DU and RU may perform functions at lower layers. The DU may perform radio link control (RLC), media access control (MAC), and some functions of the physical (PHY) layer (high PHY), while the RU may perform the remaining functions of the PHY layer (low PHY). Additionally, as an example, a digital unit (DU) may be included in a distributed unit (DU) depending on the implementation of a distributed deployment of base stations. Hereinafter, for convenience of explanation, the digital unit (DU) may be understood as having the same meaning as a distributed unit (DU) that does not include an RU. In addition, DU (digital unit) (or DU (distributed unit)) can be understood as having the same meaning as O-DU (O-RAN digital unit) or O-DU (O-RAN distributed unit).
[0059] FIG. 3 is a diagram illustrating the configuration of a distributed unit (DU) according to various embodiments of the present disclosure.
[0060] The configuration exemplified in FIG. 3 can be understood as the configuration of the DU (160) of FIG. 1b as part of a base station. Terms such as '...part', '...unit' used below refer to a unit that processes at least one function or operation, and this can be implemented in hardware or software, or a combination of hardware and software.
[0061] Referring to FIG. 3, the DU (160) includes a communication unit (210), a storage unit (220), and a control unit (230).
[0062] The communication unit (210) can perform functions for transmitting and receiving signals in a wired communication environment. The communication unit (210) may include a wired interface for controlling a direct connection between devices through a transmission medium (e.g., copper wire, optical fiber). For example, the communication unit (210) can transmit an electrical signal to another device through a copper wire or perform conversion between an electrical signal and an optical signal. The communication unit (210) can be connected to a radio unit (RU). The communication unit (210) can be connected to a core network or to a distributed CU.
[0063] The communication unit (210) may perform functions for transmitting and receiving signals in a wireless communication environment. For example, the communication unit (210) may perform a conversion function between a baseband signal and a bit sequence according to the physical layer specifications of the system. For example, when transmitting data, the communication unit (210) generates complex symbols by encoding and modulating the transmitted bit sequence. Also, when receiving data, the communication unit (210) restores the received bit sequence by demodulating and decoding the baseband signal. Additionally, the communication unit (210) may include a plurality of transmission and reception paths. Also, according to one embodiment, the communication unit (210) may be connected to a core network or to other nodes (e.g., an integrated access backhaul).
[0064] The communication unit (210) can transmit and receive signals. To this end, the communication unit (210) may include at least one transceiver. For example, the communication unit (210) can transmit a synchronization signal, a reference signal, system information, a message, a control message, a stream, control information, or data.
[0065] The communication unit (210) transmits and receives signals as described above. Accordingly, all or part of the communication unit (210) may be referred to as a 'transmitter', a 'receiver', or a 'transmitter / receiver'. Furthermore, in the following description, transmission and reception performed via a wireless channel are used to mean that processing as described above is performed by the communication unit (210).
[0066] Although not illustrated in FIG. 3, the communication unit (210) may further include a backhaul communication unit for connecting to a core network or another base station. The backhaul communication unit provides an interface for performing communication with other nodes within the network. That is, the backhaul communication unit converts a bit sequence transmitted from a base station to another node, e.g., another connection node, another base station, an upper node, a core network, etc., into a physical signal, and converts a physical signal received from another node into a bit sequence.
[0067] The storage unit (220) stores data such as basic programs, application programs, and configuration information for the operation of the DU (160). The storage unit (220) may include memory. The storage unit (220) may be composed of volatile memory, non-volatile memory, or a combination of volatile memory and non-volatile memory. Additionally, the storage unit (220) provides the stored data upon the request of the control unit (230).
[0068] The control unit (230) controls the overall operations of the DU (160). For example, the control unit (230) transmits and receives signals through the communication unit (210) (or through the backhaul communication unit). In addition, the control unit (230) writes and reads data to and from the storage unit (220). Furthermore, the control unit (230) can perform the functions of the protocol stack required by the communication standard. To this end, the control unit (230) may include at least one processor.
[0069] The configuration of the DU (160) shown in FIG. 3 is merely an example, and the examples of DUs performing various embodiments of the present disclosure are not limited to the configuration shown in FIG. 2. Depending on the various embodiments, some or identical configurations may be added, deleted, or changed.
[0070] FIG. 4 is a drawing illustrating the configuration of a radio unit (RU) according to various embodiments of the present disclosure.
[0071] The configuration exemplified in FIG. 4 can be understood as the configuration of the RU (180) of FIG. 1b as part of a base station. Terms such as '...part', '...unit' used below refer to a unit that processes at least one function or operation, and this can be implemented in hardware or software, or a combination of hardware and software.
[0072] Referring to FIG. 4, the RU (180) includes a communication unit (310), a storage unit (320), and a control unit (330).
[0073] The communication unit (310) performs functions for transmitting and receiving signals through a wireless channel. For example, the communication unit (310) upconverts a baseband signal into an RF band signal and transmits it through an antenna, and downconverts the RF band signal received through the antenna into a baseband signal. For example, the communication unit (310) may include a transmission filter, a reception filter, an amplifier, a mixer, an oscillator, a DAC, an ADC, etc.
[0074] Additionally, the communication unit (310) may include a plurality of transmission and reception paths. Furthermore, the communication unit (310) may include an antenna unit. The communication unit (310) may include at least one antenna array composed of a plurality of antenna elements. In terms of hardware, the communication unit (310) may be composed of a digital circuit and an analog circuit (e.g., a radio frequency integrated circuit (RFIC)). Here, the digital circuit and the analog circuit may be implemented as a single package. Additionally, the communication unit (310) may include a plurality of RF chains. The communication unit (310) may perform beamforming. The communication unit (310) may apply beamforming weights to a signal to give directionality according to the settings of the control unit (330) to the signal to be transmitted or received. According to one embodiment, the communication unit (310) may include an RF (radio frequency) block (or RF unit).
[0075] Additionally, the communication unit (310) can transmit and receive signals. To this end, the communication unit (310) may include at least one transceiver. The communication unit (310) can transmit a downlink signal. The downlink signal may include a synchronization signal (SS), a reference signal (RS) (e.g., CRS (cell-specific reference signal), DM (demodulation)-RS), system information (e.g., MIB, SIB, RMSI (remaining system information), OSI (other system information)), a configuration message, control information, or downlink data. Additionally, the communication unit (310) can receive an uplink signal. Uplink signals may include random access-related signals (e.g., random access preamble (RAP) (or Msg1 (message 1)), Msg3 (message 3)), reference signals (e.g., sounding reference signal (SRS), DM-RS), or power headroom report (PHR), etc.
[0076] The communication unit (310) transmits and receives signals as described above. Accordingly, all or part of the communication unit (310) may be referred to as a 'transmitter', a 'receiver', or a 'transmitter / receiver'. Additionally, in the following description, transmission and reception performed via a wireless channel are used to mean that processing as described above is performed by the communication unit (310).
[0077] The storage unit (320) stores data such as basic programs, application programs, and configuration information for the operation of the RU (180). The storage unit (320) may be composed of volatile memory, non-volatile memory, or a combination of volatile memory and non-volatile memory. The storage unit (320) provides the stored data upon request from the control unit (330). According to various embodiments of the present disclosure described below, the storage unit (320) may include memory for performing operations for processing for beamforming weights or for IFFT transformation.
[0078] The control unit (330) (processor or controller) controls the overall operations of the RU (180). For example, the control unit (330) transmits and receives signals through the communication unit (310). Additionally, the control unit (330) writes and reads data to and from the storage unit (320). Furthermore, the control unit (330) can perform the functions of the protocol stack required by the communication standard. To this end, the control unit (330) may include at least one processor. In some embodiments, the control unit (330) may be configured to transmit an SRS to the DU (160) based on the antenna number. Additionally, in some embodiments, the control unit (330) may be configured to transmit an SRS to the DU (160) after an uplink transmission. The control unit (330) may be a set of commands or codes stored in the storage unit (320) according to the SRS transmission method, or a storage space storing commands / codes or commands / codes that are temporarily resided in the control unit (330), or may be part of the circuitry constituting the control unit (330). Additionally, the control unit (330) may include various modules for performing communication. According to various embodiments, the control unit (330) may control the RU (180) to perform operations according to various embodiments described below. For example, the control unit (330) may control the beamforming block, IFFT block, and specific components included in each block included in the RU to operate according to various embodiments of the present disclosure described below.
[0079] According to various embodiments of the present disclosure, the aforementioned unit is merely an example and is not limited thereto, and the RU (180) may include various blocks or units according to various embodiments described below. For example, the RU may further include a block for precoding data or applying beamforming weights, a block for IFFT transformation, or a block for performing IFFT and inserting CP. Below, examples of more specific RU structures including various blocks are described, and the operations performed by the RU are described for convenience and may refer to the operations performed by each specific component of the RU.
[0080] FIG. 5 is a flowchart illustrating a method for an electronic device according to one embodiment to perform data scheduling.
[0081] Referring to identification number 510, an electronic device according to one embodiment can identify, in one time slot, the minimum delay required time for each data packet that needs to be transmitted and the size of the resource block required to transmit each data packet.
[0082] In one embodiment, the electronic device can identify the minimum delay required time for each data packet that needs to be transmitted in a single time slot. The electronic device can determine the shortest of the delay required times for each data packet stored in the base station's data buffer as the minimum delay required time.
[0083] In one embodiment, the electronic device can identify the size of a resource block required to transmit each data packet. The electronic device can identify the size of a resource block required to transmit each data packet stored in the base station's data buffer. The electronic device can determine the sum of the sizes of each identified resource block as the size of a resource block required to transmit a data packet stored in the device.
[0084] In one embodiment, the electronic device can identify the data buffer of the electronic device in a time slot. For example, the electronic device can identify the amount of available data buffer of the electronic device in the current time slot.
[0085] Referring to identification number 520, an electronic device according to one embodiment can apply the identified minimum delay required time and the size of the required resource block to an artificial intelligence model.
[0086] In one embodiment, the electronic device may apply data related to the identified minimum delay required time, the size of the required resource block, and the amount of the data buffer of the electronic device to an artificial intelligence model.
[0087] An artificial intelligence model according to one embodiment may be trained to receive input data related to the minimum delay required time for each data packet requiring transmission and data related to the size of a resource block required to transmit each data packet requiring transmission, and to output a value related to whether to transmit each data packet in the time slot. An artificial intelligence model according to one embodiment may receive input data related to the minimum delay required time for each data packet requiring transmission and data related to the size of a resource block required to transmit each data packet requiring transmission. Based on the input data, the artificial intelligence model may output one of a value instructing to transmit each data packet in the time slot or a value instructing not to transmit each data packet in the time slot. For example, the artificial intelligence model may output 1 as a value instructing to transmit each data packet in the time slot. For example, the artificial intelligence model may output 0 as a value instructing not to transmit each data packet in the time slot.
[0088] Referring to identification number 530, an electronic device according to one embodiment can transmit each data packet through a transceiver in a time slot based on a value output from an artificial intelligence model.
[0089] As described above, the artificial intelligence model may output a value instructing the transmission of each data packet in the time slot. The electronic device may decide to transmit data packets in the time slot based on the value output from the artificial intelligence model instructing the transmission of data packets. For example, if 1 is output from the artificial intelligence model, the electronic device may decide to transmit each data packet in the time slot. For example, if 0 is output from the artificial intelligence model, the electronic device may decide not to transmit data packets in the time slot.
[0090] In one embodiment, the electronic device may continuously transmit the data packets requiring transmission during the time slot based on a value output from the artificial intelligence model instructing the transmission of each data packet. The electronic device may sequentially transmit the data packets requiring transmission during the time slot, starting with the data packet received first, based on a value output from the artificial intelligence model instructing the transmission of each data packet.
[0091] In one embodiment, the artificial intelligence model may output a value instructing not to transmit a data packet during the time slot. The electronic device may determine not to transmit a data packet during the time slot based on the value output from the artificial intelligence model instructing not to transmit a data packet. The electronic device may not transmit any data packet during the time slot based on the value output from the artificial intelligence model instructing not to transmit a data packet.
[0092] An electronic device according to one embodiment may perform the operations of identification numbers 510 to 530 for a time slot and sequentially perform the operations of identification numbers 510 to 530 for the next time slot.
[0093] Accordingly, the electronic device according to the present disclosure can perform data scheduling according to an optimized scheduling policy by reducing the power consumption of the base station and increasing the user's QoS (quality of service).
[0094] FIG. 6 is a block diagram illustrating data received by an artificial intelligence model according to one embodiment of the present disclosure and a value output from the artificial intelligence model accordingly.
[0095] An artificial intelligence model (610) according to one embodiment may receive data related to the minimum delay required time of a data packet that needs to be transmitted. For example, an electronic device may determine the shortest of the delay required times of each data packet stored in the base station's data buffer as the minimum delay required time and apply the determined minimum delay required time to the artificial intelligence model (610). An artificial intelligence model (610) according to one embodiment may receive data related to the size of a resource block required to transmit each data packet that needs to be transmitted. An electronic device may determine the sum of the sizes of each identified resource block as the size of the resource block required to transmit the data packet stored in the electronic device and apply the determined resource block size to the artificial intelligence model (610). An artificial intelligence model (610) according to one embodiment may receive data related to the size of a resource block required to transmit each data packet that needs to be transmitted.
[0096] An artificial intelligence model (610) according to one embodiment may output a value that instructs to transmit a data packet in a time slot or a value that instructs not to transmit a data packet in a time slot. For example, the artificial intelligence model may output 1 as a value that instructs to transmit each of the data packets in the time slot. For example, the artificial intelligence model may output 0 as a value that instructs not to transmit each of the data packets in the time slot.
[0097] FIG. 7 is a flowchart illustrating the process of training an artificial intelligence model according to one embodiment.
[0098] The artificial intelligence model according to the present disclosure may include an algorithm or artificial intelligence model trained through reinforcement learning. In one embodiment, the artificial intelligence model may include an artificial intelligence model trained by the Q-learning method. Below, the process of training the artificial intelligence model of the present disclosure for data scheduling of a base station will be described in sequence.
[0099] Referring to identification number 710, an electronic device according to one embodiment can determine state information for a time slot. The state information may include data related to the minimum delay required time for each data packet required to be transmitted in the time slot and data related to the size of the resource block required to be transmitted. In one embodiment, the electronic device can identify, for each time slot, the number of data packets requested by the base station and data packets received by the base station and waiting, and the minimum delay required time. Accordingly, the state information is the total sum of the currently waiting resource blocks ( ) and minimum delay required time( It may include data related to ). That is, the state information may include information related to the amount of data that an electronic device or system must process and the minimum time that can be used to process that data. In one embodiment, the state information may take the form of a pair consisting of a value related to the minimum delay required time for each data packet that needs to be transmitted in a time slot, and a value related to the total amount of data of the resource block that needs to be transmitted. For example, state information can be determined according to the following mathematical formula 1. of mathematical formula 1 can be determined by the following mathematical formula 2, and of mathematical formula 1 It can be determined by the following mathematical formula 3.
[0100] [Mathematical Formula 1]
[0101]
[0102] In the following mathematical formula 2, represents the total amount of data of packets stored in the data buffer of the electronic device in a time slot. represents the amount of data of each packet stored in the data buffer.
[0103] [Mathematical Formula 2]
[0104]
[0105] In the following mathematical formula 3, represents the minimum delay required time for each packet stored in the data buffer in the time slot. represents the delay required time for each packet stored in the data buffer in the time slot.
[0106] [Mathematical Formula 3]
[0107]
[0108] In one embodiment, the electronic device can determine state information using an artificial intelligence model.
[0109] In identification number 720, the artificial intelligence model can determine the Q values that can be obtained according to the first policy and the second policy.
[0110] In one embodiment, it may be determined whether to transmit data in a time slot. For example, if it is determined to transmit data in a time slot, the artificial intelligence model
[0111] In one embodiment, the first policy may be an AS (always send) policy. In one embodiment, the first policy may include a policy of always transmitting data packets stored in a data buffer regardless of conditions. An electronic device according to one embodiment may continuously transmit data packets during a time slot when it decides to schedule data according to the first policy in a time slot.
[0112] In one embodiment, the second policy may be an always-wait (AW) policy. Additionally, the second policy may include a policy of not always transmitting data packets. An electronic device according to one embodiment may not transmit data packets during a time slot if it decides to perform data scheduling according to the second policy in a time slot. For example, if it decides to perform data scheduling according to the second policy in a time slot, it may not transmit any data packets during the time slot. If it decides to perform data scheduling according to the second policy in a time slot, it may transmit data packets only when the minimum delay requirement time of the data packets stored in the data buffer becomes zero during the time slot.
[0113] In identification number 730, the artificial intelligence model can perform policy evaluation for the first policy and the second policy.
[0114] In one embodiment, the artificial intelligence model, according to the following mathematical formula 4, Q value according to the first policy ( ) can be obtained. According to the mathematical formula 5 below, the AI model obtains the Q value according to the second policy ( ) can be obtained. The artificial intelligence model can perform policy evaluation on Q values derived based on the following mathematical formula 4 or mathematical formula 5.
[0115] [Mathematical Formula 4]
[0116]
[0117] [Mathematical Formula 5]
[0118]
[0119] In identification number 740, This can be obtained.
[0120] In one embodiment, is obtained based on a first policy that selects transmission in all states and a second policy that selects waiting in all states. and It can be obtained based on. and It can be used to optimize the decision of transmission or waiting. In determining whether to transmit data and The weight that it possesses is the weight and It can be adjusted by. Weight and It can be adjusted to reflect the dynamic requirements of the traffic model. In one embodiment, It can be obtained according to the following mathematical formula 6. can be a value representing the internal division point between the first policy and the second policy.
[0121] [Mathematical Formula 6]
[0122]
[0123] In mathematical formula 6, and represents the weights for the first policy and the second policy, respectively. and Is, In order to coordinate the influence of the first and second policies on [it], it may be adjusted. is the Q value according to the first policy, may refer to the ratio reflecting the Q value according to the second policy. Weights for the first and second policies and It can be adjusted according to the system's dynamic requirements to obtain the optimal policy. Through this calculation, target Q values for all states are derived, and the derived target Q values It is set to.
[0124] In identification number 750, through the following policy evaluation The Q value closest to can be obtained.
[0125] In one embodiment, the artificial intelligence model It can be trained to select the action with the Q value closest to it. Through the policy evaluation process for the first and second policies, the respective Q values are and can be obtained. Based on each obtained Q value, weights class applied This can be derived.
[0126] According to one embodiment, the Q value can be updated through iterative policy evaluation. The Q value can be updated through policy evaluation according to the following Equation 7. At this time, A policy can be selected to derive the Q value closest to it.
[0127] [Mathematical Formula 7]
[0128]
[0129] According to one embodiment, for each state information, Q target To select the action closest to, the decision function ...can be used. Referring to Equations 8 and 9 below, for each pair of state information and action, a decision function can be obtained. Based on each state information, among the possible actions Q target The action with the Q value closest to can be selected. Referring to Equation 9, It is based on the p-norm between two vectors, with weight μ assigned to each element of the vector. i It can be obtained by applying and summing and taking the p-th root.
[0130] [Mathematical Formula 8]
[0131]
[0132] [Mathematical Formula 9]
[0133]
[0134] In identification number 760, the artificial intelligence model, based on the evaluated Q value, You can obtain.
[0135] Once training for the AI model is complete, Q values for all state information-action pairs are calculated. It can be stored in. In order to determine the optimal action corresponding to each state information, The decision function again By applying, The action that minimizes the value can be selected. That is, after training is complete, Decision function for each state information By applying [this], the action that yields the smallest value can be selected as the optimal action for the corresponding state information. Accordingly, an optimal policy for each state information can be obtained.
[0136] In one embodiment, Each state Output value corresponding to May include. Output value It can contain a value of 0 or 1.
[0137] For example, according to the first policy It can be configured as shown in Table 1 below. The first policy may include a policy of transmitting data packets in all states. Therefore, regarding the first policy silver, It does not include the Q value when α is 0.
[0138]
[0139] For example, according to the second policy It can be configured as shown in Table 2 below. The second policy may include a policy of waiting without transmitting data packets in all states. Therefore, regarding the second policy silver, It does not include the Q value when α is 1.
[0140]
[0141] FIG. 8 is a flowchart illustrating a specific method for performing data scheduling using an artificial intelligence model in one embodiment.
[0142] In identification number 810, the electronic device can observe the state of the time slot. In one embodiment, the electronic device can determine state information for the current time slot. For example, if the current time slot is the t-th time slot, the set of requested data packets is It can be defined as. represents the total number of data packets requested in the corresponding time slot. It is the set of data packets remaining in the base station's data buffer that were not transmitted in the previous time slot. In that case, the data buffer before determining whether to transmit in the current time slot is It can be determined as. Data buffer When the state is measured through the state observation device, the state information of the corresponding time slot It could be. represents the total amount of data of data packets stored in the data buffer, and represents the minimum delay required time of the above data packets stored in the data buffer.
[0143] In identification number 820, the electronic device can calculate the target Q value of the time slot. The electronic device, through a policy evaluation process for the first policy and the second policy, calculates each Q value and Calculate and the weights for these values class The target Q value by applying It is possible to obtain the target Q value. A detailed description of the process for obtaining the target Q value corresponds to the description of identification number 740 in Fig. 7, so it will be omitted here.
[0144] In identification number 830, the electronic device selects the Q value closest to the target value of the time slot through an artificial intelligence model trained via reinforcement learning (deep-learning) (e.g., the artificial intelligence model (610) of FIG. 6), and based on the selected Q value You can obtain. A detailed explanation regarding the process of obtaining [it] corresponds to the explanation regarding identification number 760 in Fig. 7, so it will be omitted here.
[0145] FIG. 9 is an example diagram illustrating the delay time and number of transmissions obtained by adjusting the weights for each policy for training an artificial intelligence model according to one embodiment.
[0146] In FIG. 9, the graph connecting the coordinates representing the delay time and number of transmissions in the case following the first policy and the coordinates representing the delay time and number of transmissions in the case following the second policy is a graph representing the number of transmissions relative to the delay time in an ideal situation.
[0147] The method according to the present disclosure provides a method for determining an optimal policy for each state information. An artificial intelligence model can determine the state-action pair closest to a graph representing the number of transmissions relative to a delay time in an ideal situation by performing reinforcement learning to select an optimal policy for each state information. For example, the coordinates of identification number 920 may represent the number of transmissions according to a policy that obtains the coordinates closest to the graph representing the number of transmissions relative to a delay time in an ideal situation when the delay time is 38. In this case, the weight is 0.4, and the weight can be 0.4. For example, the coordinates of identification number 930 may represent the number of transmissions according to a policy that obtains the coordinates closest to a graph representing the number of transmissions relative to the delay time in an ideal situation, in the case where the delay time is 42. In this case, the weight is 0.5, and the weight It can be 0.1.
[0148] FIG. 10a is an example diagram illustrating the effect of performing data scheduling using an artificial intelligence model trained according to one embodiment. FIG. 10b is an example diagram illustrating the effect of performing data scheduling using an artificial intelligence model trained according to one embodiment.
[0149] The method according to the present disclosure can provide an optimized policy that takes into account delay time and power consumption by adjusting weights that follow a first policy or a second policy.
[0150] FIGS. 10a and 10b illustrate a comparison of delay times in cases following policies that adjust whether to transmit data packets considering delay times within an electronic device, and in cases following the method of the present disclosure. For example, in order to minimize the delay time of data packets while maintaining the number of data transmissions of a base station at a second policy level Fix the value to 0.01 and The values can be adjusted. In this case, The lower the value, the more the latency can be reduced without approaching the first policy.
[0151] In a wireless communication system according to the present disclosure, a base station may include a transceiver, a memory for storing instructions, and at least one processor. The instructions are executed individually or collectively by the at least one processor so that the base station identifies, in a time slot, a minimum delay required time for each data packet requiring transmission and the size of a resource block required to transmit each data packet, and for the time slot, applies the identified minimum delay required time and the size of the required resource block to an artificial intelligence model, wherein the artificial intelligence model is trained to receive input data related to the minimum delay required time for each data packet requiring transmission and data related to the size of the resource block required to transmit each data packet requiring transmission, and outputs a value related to whether to transmit each data packet in the time slot, and determines whether to transmit each data packet through the transceiver in the time slot based on the value output from the artificial intelligence model.
[0152] The artificial intelligence model may receive input data related to the minimum delay required time for each data packet requiring transmission and data related to the size of the resource block required to transmit each data packet requiring transmission. The artificial intelligence model may include a model trained through reinforcement learning to output one of a value instructing to transmit each data packet in the time slot or a value instructing not to transmit each data packet in the time slot, based on the input data.
[0153] The artificial intelligence model may include an artificial intelligence model trained to determine state information for the time slot based on data related to the minimum delay required time for each data packet requiring transmission and the size of the resource block requiring transmission, determine the value to be output in correspondence with the determined state information, store the determined value in correspondence with the determined state information, and output the stored value in correspondence with the determined state information.
[0154] The above artificial intelligence model may include an artificial intelligence model trained to determine the output value based on the data buffer of the base station, the identified minimum delay required time, and the amount of power consumed by the base station.
[0155] The artificial intelligence model may include an artificial intelligence model trained to perform a policy evaluation for a first policy of immediately transmitting data stored in the data buffer of the base station and a second policy of not transmitting data based on the determined state information, and to determine the value to be output corresponding to each of the state information of the determined time slot according to the result of the policy evaluation.
[0156] The artificial intelligence model may include an artificial intelligence model trained to assign a first weight to the first policy, assign a second weight to the second policy, and, based on the determined state information, adjust the first weight and the second weight to determine a target cumulative reward expectation for the determined state information, and, according to the determined target cumulative reward expectation, determine the value to be output corresponding to the determined state information.
[0157] The value to be output in response to the above-determined state information may be a value that obtains the cumulative reward expectation closest to the above-determined target cumulative reward expectation.
[0158] A method performed by a base station in a wireless communication system may include, in a time slot, an operation of identifying a minimum delay required time for each data packet requiring transmission and the size of a resource block required to transmit each data packet, and for the time slot, an operation of applying the identified minimum delay required time and the size of the required resource block to an artificial intelligence model, wherein the artificial intelligence model is trained to receive input data related to the minimum delay required time for each data packet requiring transmission and data related to the size of the resource block required to transmit each data packet requiring transmission, and to output a value related to whether to transmit each data packet in the time slot, and may include an operation of transmitting each data packet through a transceiver in the time slot based on the value output from the artificial intelligence model.
[0159] The artificial intelligence model may receive input data related to the minimum delay required time for each data packet requiring transmission and data related to the size of the resource block required to transmit each data packet requiring transmission. The artificial intelligence model may include a model trained through reinforcement learning to output one of a value instructing to transmit each data packet in the time slot or a value instructing not to transmit each data packet in the time slot, based on the input data.
[0160] The artificial intelligence model may include an artificial intelligence model trained to determine state information for the time slot based on data related to the minimum delay required time for each data packet requiring transmission and the size of the resource block requiring transmission, determine the value to be output in correspondence with the determined state information, store the determined value in correspondence with the determined state information, and output the stored value in correspondence with the determined state information.
[0161] The above artificial intelligence model may include an artificial intelligence model trained to determine the output value based on the data buffer of the base station, the identified minimum delay required time, and the amount of power consumed by the base station.
[0162] The artificial intelligence model may include an artificial intelligence model trained to perform a policy evaluation for a first policy of immediately transmitting data stored in the data buffer of the base station and a second policy of not transmitting data based on the determined state information, and to determine the value to be output corresponding to each of the state information of the determined time slot according to the result of the policy evaluation.
[0163] The artificial intelligence model may include an artificial intelligence model trained to assign a first weight to the first policy, assign a second weight to the second policy, and, based on the determined state information, adjust the first weight and the second weight to determine a target cumulative reward expectation for the determined state information, and, according to the determined target cumulative reward expectation, determine the value to be output corresponding to the determined state information.
[0164] The value to be output in response to the above-determined state information may be a value that obtains the cumulative reward expectation closest to the above-determined target cumulative reward expectation.
[0165] A recording medium according to the present disclosure may include a computer-readable recording medium storing a program for executing a method comprising: an operation of identifying, in a time slot, a minimum delay required time for each data packet requiring transmission and the size of a resource block required to transmit each data packet; an operation of applying the identified minimum delay required time and the size of the required resource block to an artificial intelligence model for the time slot; wherein the artificial intelligence model is trained to receive input data related to the minimum delay required time for each data packet requiring transmission and data related to the size of the resource block required to transmit each data packet requiring transmission, and output a value related to whether to transmit each data packet in the time slot; and an operation of transmitting each data packet through a transceiver in the time slot based on the value output from the artificial intelligence model. The artificial intelligence model may receive input data related to the minimum delay required time for each data packet requiring transmission and data related to the size of the resource block required to transmit each data packet requiring transmission. The artificial intelligence model may include one trained through reinforcement learning to output one of a value instructing to transmit each data packet in the time slot or a value instructing not to transmit each data packet in the time slot, based on the input data.
[0166] The artificial intelligence model may include an artificial intelligence model trained to determine state information for the time slot based on data related to the minimum delay required time for each data packet requiring transmission and the size of the resource block requiring transmission, determine the value to be output in correspondence with the determined state information, store the determined value in correspondence with the determined state information, and output the stored value in correspondence with the determined state information.
[0167] The above artificial intelligence model may include an artificial intelligence model trained to determine the output value based on the data buffer of the base station, the identified minimum delay required time, and the amount of power consumed by the base station.
[0168] The artificial intelligence model may include an artificial intelligence model trained to perform a policy evaluation for a first policy of immediately transmitting data stored in the data buffer of the base station and a second policy of not transmitting data based on the determined state information, and to determine the value to be output corresponding to each of the state information of the determined time slot according to the result of the policy evaluation.
[0169] The artificial intelligence model may include an artificial intelligence model trained to assign a first weight to the first policy, assign a second weight to the second policy, and, based on the determined state information, adjust the first weight and the second weight to determine a target cumulative reward expectation for the determined state information, and, according to the determined target cumulative reward expectation, determine the value to be output corresponding to the determined state information.
[0170] The value to be output in response to the above-determined state information may be a value that obtains the cumulative reward expectation closest to the above-determined target cumulative reward expectation.
[0171] Various embodiments of this document may be implemented as software comprising one or more instructions stored on a storage medium readable by a machine. For example, the processor of the machine may call at least one of the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, "non-transitory" simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily in the storage medium.
[0172] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or an application store (e.g., Play Store). TMIt can be distributed online (e.g., downloaded or uploaded) through ) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0173] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In a base station of a wireless communication system, Transmitter / receiver; Memory for storing instructions; and It includes at least one processor, The above instructions are executed individually or collectively by the at least one processor, so that the base station: In a time slot, identify the minimum delay required time for each data packet requiring transmission and the size of the resource block required to transmit each data packet, and For the above time slot, the identified minimum delay required time and the size of the required resource block are applied to an artificial intelligence model, wherein the artificial intelligence model is trained to receive input data related to the minimum delay required time for each data packet requiring transmission and data related to the size of the resource block required to transmit each data packet requiring transmission, and to output a value related to whether to transmit each data packet in the above time slot. Determining whether to transmit each of the data packets through the transceiver in the time slot based on the value output from the artificial intelligence model, Base station.
2. In Claim 1, The above artificial intelligence model is: Data related to the minimum delay required time for each data packet requiring transmission and data related to the size of the resource block required to transmit each data packet requiring transmission are received as input, A model of artificial intelligence trained through reinforcement learning, which outputs one of a value instructing to transmit each of the data packets in the time slot or a value instructing not to transmit each of the data packets in the time slot, based on the input data received above. Base station.
3. In Claim 2, The above artificial intelligence model is: Based on data related to the minimum delay required time for each data packet requiring transmission and the size of the resource block requiring transmission, state information for the time slot is determined, and Determine the value to be output in correspondence with the above-determined state information, and Storing the determined value to correspond to the determined state information, and including an artificial intelligence model trained to output a stored value corresponding to the above-determined state information, Base station.
4. In Claim 3, The above artificial intelligence model includes an artificial intelligence model trained to determine the output value based on the data buffer of the base station, the identified minimum delay required time, and the amount of power consumed by the base station. Base station.
5. In Claim 3, The above artificial intelligence model is: Based on the above-determined state information, a policy evaluation is performed for a first policy of immediately transmitting data stored in the data buffer of the base station and a second policy of not transmitting data, and A model of artificial intelligence trained to determine the value to be output corresponding to each of the state information of the determined time slot according to the result of the above policy evaluation, comprising Base station.
6. In Claim 5, The above artificial intelligence model is: A first weight is assigned to the above first policy, and a second weight is assigned to the above second policy, and Based on the above-determined state information, by adjusting the first weight and the second weight, the target cumulative reward expectation for the above-determined state information is determined, and A model of artificial intelligence trained to determine the value to be output in response to the determined state information according to the above-determined target cumulative reward expectation, Base station.
7. In Claim 6, The value to be output in response to the above-determined state information is a value that obtains the cumulative reward expectation closest to the above-determined target cumulative reward expectation, Base station.
8. A method performed by a base station in a wireless communication system, An operation to identify, in a time slot, the minimum delay required time for each data packet requiring transmission and the size of the resource block required to transmit each data packet; With respect to the above time slot, the operation of applying the identified minimum delay required time and the size of the required resource block to an artificial intelligence model, wherein the artificial intelligence model is trained to receive input data related to the minimum delay required time for each data packet requiring transmission and data related to the size of the resource block required to transmit each data packet requiring transmission, and to output a value related to whether to transmit each data packet in the above time slot; and A method comprising the operation of transmitting each of the data packets through a transmitting and receiving unit in the time slot based on a value output from the artificial intelligence model.
9. In Claim 8, The above artificial intelligence model is: Data related to the minimum delay required time for each data packet requiring transmission and data related to the size of the resource block required to transmit each data packet requiring transmission are received as input, A method comprising an artificial intelligence model trained through reinforcement learning to output one of a value instructing each data packet to be transmitted in the time slot or a value instructing not to transmit each data packet in the time slot, based on the input data received above.
10. Regarding the recording media of the base station, An operation to identify, in a time slot, the minimum delay required time for each data packet requiring transmission and the size of the resource block required to transmit each data packet; With respect to the above time slot, the operation of applying the identified minimum delay required time and the size of the required resource block to an artificial intelligence model, wherein the artificial intelligence model is trained to receive input data related to the minimum delay required time for each data packet requiring transmission and data related to the size of the resource block required to transmit each data packet requiring transmission, and to output a value related to whether to transmit each data packet in the above time slot; and A computer-readable recording medium having a program for executing a method including the operation of transmitting each of the data packets through a transmitting and receiving unit in the time slot based on a value output from the artificial intelligence model.
11. In Claim 10, The above artificial intelligence model is: Data related to the minimum delay required time for each data packet requiring transmission and data related to the size of the resource block required to transmit each data packet requiring transmission are received as input, A recording medium comprising an artificial intelligence model trained through reinforcement learning to output one of a value instructing to transmit each of the data packets in the time slot or a value instructing not to transmit each of the data packets in the time slot, based on the input data received above.
12. In Claim 11, The above artificial intelligence model is: Based on data related to the minimum delay required time for each data packet requiring transmission and the size of the required resource block, state information for the time slot is determined, and Determine the value to be output in correspondence with the above-determined state information, and Storing the determined value to correspond to the determined state information, and A recording medium comprising an artificial intelligence model trained to output a value stored corresponding to the above-determined state information.
13. In Claim 12, A recording medium comprising an artificial intelligence model trained to determine the output value based on the data buffer of the base station, the identified minimum delay required time, and the amount of power consumed by the base station.
14. In Claim 12, The above artificial intelligence model is: Based on the above-determined state information, a policy evaluation is performed for a first policy of immediately transmitting data stored in the data buffer of the base station and a second policy of not transmitting data, and A model of artificial intelligence trained to determine the value to be output corresponding to each of the state information of the determined time slot according to the result of the above policy evaluation, comprising Recording media.
15. In Claim 14, The above artificial intelligence model is: A first weight is assigned to the above first policy, and a second weight is assigned to the above second policy, and Based on the above-determined state information, by adjusting the first weight and the second weight, the target cumulative reward expectation for the above-determined state information is determined, and A model of artificial intelligence trained to determine the value to be output in response to the determined state information according to the above-determined target cumulative reward expectation, Recording media.