Electronic device and method for allocating computing resources of plurality of cells on basis of deep reinforcement learning algorithm
The method uses a deep reinforcement learning algorithm to optimize computing resource allocation in C-RAN and O-RAN architectures, addressing network inefficiencies and latency by reallocating resources based on Q-values for improved throughput.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2025-12-10
- Publication Date
- 2026-07-23
AI Technical Summary
Existing methods for allocating computing resources in C-RAN and O-RAN architectures fail to optimize the resources required by each cell in real time, leading to network inefficiencies and potential latency issues.
An electronic device and method utilize a deep reinforcement learning algorithm to acquire network state information, generate control messages for reallocating computing resources, and optimize throughput by applying Q-values to improve network performance.
The method enhances network performance by dynamically reallocating computing resources, reducing waste and improving transmission speeds and stability in C-RAN and O-RAN environments.
Smart Images

Figure KR2025021207_23072026_PF_FP_ABST
Abstract
Description
Electronic device and method for allocating computing resources of multiple cells based on a deep reinforcement learning algorithm
[0001] The present disclosure relates to an electronic device and method for allocating computing resources of a plurality of cells based on a deep reinforcement learning algorithm.
[0002] C-RAN and O-RAN are next-generation network architectures introduced to enhance the efficiency and flexibility of wireless communication systems. Existing Radio Access Networks (RANs) had the disadvantage of low resource utilization and high maintenance costs due to their distributed structure across each base station. To address this, C-RAN proposed a method that separates base station functions, centrally manages Baseband Units (BBUs), and deploys only Remote Radio Heads (RRHs) in the field. O-RAN further strengthens flexibility and interoperability by adding open interfaces and software-based network functions.
[0003] In this structure, the situation of connecting with multiple cells is emerging as a significant challenge in terms of network performance and resource management. The BBU plays a central role in collecting and processing data from various cells, and as the number of cells increases, the burden on computing resources and bandwidth grows. In particular, in environments where traffic surges, problems such as latency or slowdowns in data processing speed may occur due to the performance limitations of the BBU. To address this, high-performance processors and efficient resource management algorithms are required.
[0004] Furthermore, C-RAN and O-RAN architectures can provide enhanced services by combining with network slicing and Multi-access Edge Computing (MEC) technologies. For example, wasting network resources can be reduced by analyzing and optimizing the resources required by each cell in real time. Additionally, thanks to centralized management, cooperation between cells becomes possible, which can significantly improve transmission speeds and network stability. These technological advancements are establishing themselves as key elements of the 5G and 6G eras.
[0005] As such, with the advancement of communication technology and the proliferation of C-RAN and O-RAN architectures, methods and devices are needed to reduce the waste of network resources by analyzing and optimizing the resources required by each cell in real time.
[0006] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art in relation to the present disclosure.
[0007] The present disclosure provides an electronic device and method for allocating computing resources of a plurality of cells based on a deep reinforcement learning algorithm.
[0008] A method performed by a first electronic device according to one embodiment of the present disclosure may include: acquiring network state information of a plurality of cells connected to a second electronic device; acquiring a plurality of Q values by applying the network state information to a deep reinforcement learning algorithm for improving the network performance of the plurality of cells; generating a control message instructing the re-allocation of computing resources allocated to the plurality of cells based on the plurality of Q values; transmitting the control message to the second electronic device; and acquiring information related to the throughput of the plurality of cells from the second electronic device.
[0009] A first electronic device of a wireless communication system according to one embodiment of the present disclosure comprises: a transceiver; a memory for storing instructions; and one or more processors, wherein the instructions are executed by the one or more processors so that the first electronic device: obtains network state information of a plurality of cells connected to a second electronic device, applies the network state information to a deep reinforcement learning algorithm for improving the network performance of the plurality of cells to obtain a plurality of Q values, generates a control message instructing the re-allocation of computing resources allocated to the plurality of cells based on the plurality of Q values, transmits the control message to the second electronic device, and obtains information related to the throughput of the plurality of cells from the second electronic device.
[0010] FIG. 1 illustrates a wireless communication system according to various embodiments of the present disclosure.
[0011] FIG. 2 illustrates an example of a fronthaul structure according to the functional separation of base stations according to various embodiments of the present disclosure.
[0012] FIG. 3 illustrates an example of a centralized-RAN (C-RAN) structure in a wireless communication system according to one embodiment of the present disclosure.
[0013] FIG. 4 illustrates the configuration of an electronic device according to one embodiment of the present disclosure.
[0014] FIG. 5 is a diagram illustrating an example of reallocating computing resources allocated to a plurality of cells according to the traffic distribution of a plurality of cells, according to one embodiment of the present disclosure.
[0015] FIG. 6 is a diagram illustrating an operation to reallocate computing resources allocated to a plurality of cells based on a deep reinforcement learning algorithm in a first electronic device and a second electronic device according to an embodiment of the present disclosure.
[0016] FIG. 7 is a flowchart illustrating the operation of reallocating computing resources allocated to a plurality of cells based on a deep reinforcement learning algorithm in a first electronic device according to one embodiment of the present disclosure.
[0017] FIG. 8 is a drawing for explaining an example of a plurality of Q values obtained by a first electronic device according to one embodiment of the present disclosure.
[0018] FIG. 9 is a diagram illustrating a method for a first electronic device to reallocate computing resources according to a Q value, according to one embodiment of the present disclosure.
[0019] FIG. 10 is a drawing for illustrating an example in which a first electronic device determines a compensation value for the reallocation of computing resources based on information related to throughput, according to one embodiment of the present disclosure.
[0020] FIG. 11 is a flowchart illustrating the operation of reallocating computing resources allocated to a plurality of cells based on a deep reinforcement learning algorithm in a first electronic device according to one embodiment of the present disclosure, and the operation of training a deep reinforcement learning algorithm in the first electronic device.
[0021] FIG. 12 is a flowchart illustrating the operation of a first electronic device training a deep reinforcement learning algorithm according to one embodiment of the present disclosure.
[0022] In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components.
[0023] The technical problems to be solved in this document are not limited to those mentioned above, and other unmentioned technical problems will be clearly understood by those skilled in the art to which this invention belongs from the description below.
[0024] Hereinafter, embodiments are described in detail with reference to the attached drawings so that those skilled in the art can easily implement the contents of the present disclosure. However, the disclosed embodiments may be implemented in various different forms and are not limited to the embodiments described herein. Furthermore, in order to clearly explain the present disclosure in the drawings, parts unrelated to the explanation have been omitted, and similar parts throughout the specification have been given similar reference numerals.
[0025] The terms used in this disclosure are described in their current, general form considering the functions mentioned herein; however, they may refer to various other terms depending on the intent of those skilled in the art, case law, the emergence of new technologies, etc. Accordingly, the terms used in this disclosure should not be interpreted solely by their names, but should be interpreted based on the meaning of the terms and the overall content of this disclosure.
[0026] Additionally, terms such as first, second, etc., may be used to describe various components, but the components are not limited by these terms. These terms are used for the purpose of distinguishing one component from another.
[0027] In the present disclosure, when a part is described as being "connected" to another part, this includes not only cases where they are "directly connected," but also cases where they are "electrically connected" or "operationally connected" with other elements interposed between them. Furthermore, when a part is described as "comprising" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components.
[0028] Phrases such as "in one embodiment" appearing in various places in this disclosure do not necessarily refer to the same embodiment.
[0029] Furthermore, the connecting lines or connecting members between the components depicted in the drawings are merely illustrative of functional connections and / or physical or circuit connections. In the actual device, connections between components may be represented by various alternative or added functional connections, physical connections, or circuit connections.
[0030] A method performed by a first electronic device of a wireless communication system according to various embodiments of the present disclosure may include: acquiring network state information of a plurality of cells connected to a second electronic device; acquiring a plurality of Q values by applying the network state information to a deep reinforcement learning algorithm for improving the network performance of the plurality of cells; generating a control message instructing the re-allocation of computing resources allocated to the plurality of cells based on the plurality of Q values; transmitting the control message to the second electronic device; and acquiring first information related to the throughput of the plurality of cells from the second electronic device.
[0031] In one embodiment, the operation of generating the control message may include selecting a maximum value and a minimum value among the plurality of Q values, and generating the control message that instructs to reallocate the computing resources of the cell corresponding to the minimum value to the cell corresponding to the maximum value.
[0032] In one embodiment, the computing resources required for packet processing allocated to one slot are defined as one computing resource unit, and the control message may be a message instructing to reallocate one computing resource unit of the cell corresponding to the minimum value to the cell corresponding to the maximum value.
[0033] In one embodiment, the method may further include: identifying an increase in throughput resulting from the reassignment based on the first information; and determining a reward resulting from the reassignment based on the increase.
[0034] In one embodiment, the operation of storing the network state information, the reallocation, and the reward value in a buffer for learning the deep reinforcement learning algorithm may be further included.
[0035] In one embodiment, the method further includes an operation of training the deep reinforcement learning algorithm by applying a mini-batch of information sampled from the buffer at each first period, wherein the mini-batch may include network state information at a first time point, a reallocation at the first time point, a reward value based on the reallocation at the first time point, and network state information at a second time point, which is the time point following the first time point.
[0036] In one embodiment, the deep reinforcement learning algorithm is a double deep Q network (DDQN) algorithm, and the deep reinforcement learning algorithm includes a main deep learning neural network and a target deep learning neural network, and the target deep learning neural network can be updated according to the main deep learning neural network.
[0037] In one embodiment, the operation of training the deep reinforcement learning algorithm may include applying network state information at the first time point and a reallocation at the first time point among the network state information to the main deep learning neural network to obtain a first Q value, applying network state information at the second time point among the network state information and a reward value according to the reallocation at the first time point to the target deep learning neural network to obtain a target Q value, updating the first weight of the main deep learning neural network based on a DDQN loss value representing the difference between the first Q value and the target Q value, and updating the second weight of the target deep learning neural network with the first weight of the main deep learning neural network every second period.
[0038] In one embodiment, the network status information may include information regarding the number of terminals activated in the plurality of cells, the utilization rate of radio resources of the plurality of cells, and the state in which computing resources are allocated to the plurality of cells.
[0039] In one embodiment, the operation of obtaining the plurality of Q values by applying the network state information to the deep reinforcement learning algorithm may include: an operation of setting an arbitrary value; an operation of generating the control message instructing to reallocate the computing resources arbitrarily when the arbitrary value is smaller than the epsilon value; and an operation of obtaining the plurality of Q values by applying the network state information to the deep reinforcement learning algorithm and generating the control message instructing to reallocate the computing resources based on the plurality of Q values when the arbitrary value is greater than or equal to the epsilon value.
[0040] A first electronic device of a wireless communication system according to various embodiments of the present disclosure may include a transceiver, a memory for storing instructions, and one or more processors. The instructions may be executed by the one or more processors so that the first electronic device: obtains network state information of a plurality of cells connected to a second electronic device; obtains a plurality of Q values by applying the network state information to a deep reinforcement learning algorithm for improving the network performance of the plurality of cells; generates a control message instructing the re-allocation of computing resources allocated to the plurality of cells based on the plurality of Q values; transmits the control message to the second electronic device; and obtains first information related to the throughput of the plurality of cells from the second electronic device.
[0041] In one embodiment, the operation of the first electronic device generating the control message may include the operation of generating the control message that selects a maximum value and a minimum value among the plurality of Q values and instructs the computing resources of the cell corresponding to the minimum value to be reallocated to the cell corresponding to the maximum value.
[0042] In one embodiment, in a first electronic device, the computing resources required for packet processing allocated to one slot are defined as one computing resource unit, and the control message may be a message instructing to reallocate one computing resource unit of the cell corresponding to the minimum value to the cell corresponding to the maximum value.
[0043] In one embodiment, the instructions may be executed by the one or more processors to enable the first electronic device to: identify an increase in throughput resulting from the reallocation based on the first information, and determine a reward resulting from the reallocation based on the increase.
[0044] In one embodiment, the instructions may be executed by the one or more processors to cause the first electronic device to store the network state information, the reallocation and the reward value in a buffer for learning the deep reinforcement learning algorithm.
[0045] In one embodiment, the instructions are executed by the one or more processors to enable the first electronic device to train the deep reinforcement learning algorithm by applying a mini-batch of information sampled from the buffer at each first period. The mini-batch may include network state information at a first time point, a reallocation at the first time point, a reward value for the reallocation at the first time point, and network state information at a second time point, which is the time point following the first time point.
[0046] In one embodiment, in a first electronic device, the deep reinforcement learning algorithm is a double deep Q network (DDQN) algorithm, and the deep reinforcement learning algorithm includes a main deep learning neural network and a target deep learning neural network, and the target deep learning neural network may be a neural network that is updated according to the main deep learning neural network.
[0047] In one embodiment, the instructions are executed by the one or more processors, and the operation of the first electronic device to train the deep reinforcement learning algorithm may include applying the network state information at the first time point and the reallocation at the first time point among the network state information to the main deep learning neural network to obtain a first Q value, applying the network state information at the second time point among the network state information and the reward value according to the reallocation at the first time point to the target deep learning neural network to obtain a target Q value, updating the first weight of the main deep learning neural network based on a DDQN loss value representing the difference between the first Q value and the target Q value, and updating the second weight of the target deep learning neural network with the first weight of the main deep learning neural network every second period.
[0048] In one embodiment, the network status information may include information regarding the number of terminals activated in the plurality of cells, the utilization rate of radio resources of the plurality of cells, and the state in which computing resources are allocated to the plurality of cells.
[0049] In one embodiment, the instructions are executed by the one or more processors, and the operation of the first electronic device applying the network state information to the deep reinforcement learning algorithm to obtain the plurality of Q values may include: an operation of setting an arbitrary value; an operation of generating the control message instructing to reallocate the computing resources arbitrarily when the arbitrary value is less than the epsilon value; and an operation of applying the network state information to the deep reinforcement learning algorithm to obtain the plurality of Q values and generating the control message instructing to reallocate the computing resources based on the plurality of Q values when the arbitrary value is greater than or equal to the epsilon value.
[0050] The present disclosure will be described in detail below with reference to the attached drawings.
[0051] FIG. 1 illustrates a wireless communication system according to various embodiments of the present disclosure.
[0052] FIG. 1 illustrates a base station (110), a terminal (120), and a terminal (130) as part of nodes utilizing a wireless channel in a wireless communication system. FIG. 1 illustrates only one base station, but other base stations identical or similar to the base station (110) may be included.
[0053] A base station (110) is a network infrastructure that provides wireless access to terminals (120, 130). The base station (110) has coverage defined as a certain geographical area based on the distance at which it can transmit signals. In addition to being a base station, the base station (110) may be referred to as an 'access point (AP)', 'eNodeB (eNB)', '5G node (5th generation node)', 'next generation nodeB (gNB)', 'wireless point', 'transmission / reception point (TRP)', or other terms having an equivalent technical meaning.
[0054] Each of the terminal (120) and terminal (130) is a device used by a user and communicates with the base station (110) via a wireless channel. The link from the base station (110) toward the terminal (120) or terminal (130) is referred to as a downlink (DL), and the link from the terminal (120) or terminal (130) toward the base station (110) is referred to as an uplink (UL). Additionally, the terminal (120) and terminal (130) can communicate with each other via a wireless channel. In this case, the link between the terminal (120) and terminal (130) (device-to-device link; D2D) is referred to as a sidelink, and the sidelink may be used interchangeably with the PC5 interface. In some cases, at least one of the terminal (120) and terminal (130) may be operated without user intervention. That is, at least one of the terminal (120) and the terminal (130) is a device that performs machine type communication (MTC) and may not be carried by a user. Each of the terminal (120) and the terminal (130) may be referred to as 'user equipment (UE)', 'customer premises equipment (CPE)', 'mobile station', 'subscriber station', 'remote terminal', 'wireless terminal', 'electronic device', or 'user device' or other terms having an equivalent technical meaning.
[0055] The base station (110), terminal (120), and terminal (130) can perform beamforming. The base station (110), terminal (120), and terminal (130) can transmit and receive wireless signals in relatively low frequency bands (e.g., FR1 (frequency range 1) of NR) as well as in high frequency bands (e.g., FR2 of NR, millimeter wave (mmWave) bands (e.g., 28 GHz, 30 GHz, 38 GHz, 60 GHz)). In some embodiments, the base station can communicate with the terminal within a frequency range corresponding to FR1. In some embodiments, the base station can communicate with the terminal within a frequency range corresponding to FR2. At this time, to improve channel gain, the base station (110), terminal (120), and terminal (130) can perform beamforming. Here, beamforming may include transmit beamforming and receive beamforming. That is, the base station (110), terminal (120), and terminal (130) can impart directivity to the transmit signal or the receive signal. This For this purpose, the base station (110) and terminals (120, 130) can select serving beams (112, 113, 121, 131) through a beam search or beam management procedure. After the serving beams (112, 113, 121, 131) are selected, subsequent communication can be performed through a resource that is in a quasi-co-located (QCL) relationship with the resource that transmitted the serving beams (112, 113, 121, 131).
[0056] If large-scale characteristics of the channel transmitting the symbol on the first antenna port can be inferred from the channel transmitting the symbol on the second antenna port, the first antenna port and the second antenna port may be evaluated to have a QCL relationship. For example, the large-scale characteristics may include at least one of a delay spread, a Doppler spread, a Doppler shift, an average gain, an average delay, and a spatial receiver parameter.
[0057] Although FIG. 1 illustrates that both the base station and the terminal perform beamforming, various embodiments of the present disclosure are not necessarily limited thereto. In some embodiments, the terminal may or may not perform beamforming. Additionally, the base station may or may not perform beamforming. That is, either the base station or the terminal may perform beamforming, or neither the base station nor the terminal may perform beamforming.
[0058] In the present disclosure, a beam refers to a spatial flow of a signal in a wireless channel, formed by one or more antennas (or antenna elements), and this formation process may be referred to as beamforming. Beamforming may include analog beamforming and digital beamforming (e.g., precoding). Reference signals transmitted based on beamforming may include, for example, DM-RS (demodulation-reference signal), CSI-RS (channel state information-reference signal), SS / PBCH (synchronization signal / physical broadcast channel), and SRS (sounding reference signal). Additionally, as a configuration for each reference signal, an IE such as a CSI-RS resource or an SRS-resource may be used, and such a configuration may include information associated with the beam. Information associated with a beam may refer to whether the configuration (e.g., CSI-RS resource) uses the same spatial domain filter as other configurations (e.g., other CSI-RS resources within the same CSI-RS resource set) or a different spatial domain filter, or which reference signal it is quasi-colocated with, and if so, what type (e.g., QCL type A, B, C, D).
[0059] Although FIG. 1 illustrates that both the base station and the terminal perform beamforming, various embodiments of the present disclosure are not necessarily limited thereto. In some embodiments, the terminal may or may not perform beamforming. Additionally, the base station may or may not perform beamforming. That is, either the base station or the terminal may perform beamforming, or neither the base station nor the terminal may perform beamforming.
[0060] Conventionally, in communication systems with a relatively large cell radius of base stations, each base station was installed to include the functions of a digital processing unit (or DU (digital unit)) and a radio frequency (RF) processing unit (or RU (radio unit)). However, in 4G (4th generation) and / or later communication systems, as high frequency bands are used and the cell radius of base stations decreases, the number of base stations required to cover a specific area has increased, and the burden of installation costs for operators to install these increased base stations has increased. To minimize base station installation costs, a structure has been proposed in which the DU and RU of a base station are separated, with one or more RUs connected to a single DU via a wired network, and one or more geographically distributed RUs deployed to cover a specific area. Below, base station deployment structures and extension examples according to various embodiments of the present disclosure are described through FIG. 2.
[0061] FIG. 2 illustrates an example of a fronthaul structure according to the functional separation of base stations according to various embodiments of the present disclosure. Unlike backhaul between a base station and a core network, fronthaul refers to the space between entities between a wireless LAN and a base station.
[0062] Referring to FIG. 2, the base station (110) may include a DU (160) and an RU (180). The front hole (170) between the DU (160) and the RU (180) is F x It can be operated through an interface. For the operation of the fronthole (170), an interface such as eCPRI (enhanced common public radio interface) or ROE (radio over ethernet) may be used.
[0063] As communication technology develops, mobile data traffic increases, and consequently, the bandwidth requirements for the fronthaul between the digital unit and the wireless unit have increased significantly. In deployments such as a centralized / cloud radio access network (C-RAN), the DU performs functions for the packet data convergence protocol (PDCP), radio link control (RLC), media access control (MAC), and physical (PHY), while the RU can be implemented to perform additional functions for the PHY layer in addition to radio frequency (RF) functions.
[0064] The DU (160) may perform upper-layer functions of the wireless network. For example, the DU (160) may perform functions of the MAC layer and parts of the PHY layer. Here, parts of the PHY layer are functions of the PHY layer performed at a higher level, and may include, for example, channel encoding (or channel decoding), scrambling (or descrambling), modulation (or demodulation), and layer mapping (or layer demapping). According to one embodiment, if the DU (160) conforms to the O-RAN standard, it may be referred to as an O-DU (O-RAN DU). The DU (160) may be represented as a first network entity for a base station (e.g., gNB) in the embodiments of the present disclosure as needed.
[0065] The RU (180) can perform lower-layer functions of the wireless network. For example, the RU (180) can perform RF functions, which are part of the PHY layer. Here, part of the PHY layer refers to functions of the PHY layer that are performed at a level relatively lower than that of the DU (160), and may include, for example, IFFT transformation (or FFT transformation), CP insertion (CP removal), and digital beamforming. Examples of such specific functional separation are described in detail in FIG. 4. The RU (180) may be referred to as an 'access unit (AU)', 'access point (AP)', 'transmission / reception point (TRP)', 'remote radio head (RRH)', 'radio unit (RU)', or other terms having an equivalent technical meaning. According to one embodiment, if the RU (180) conforms to the O-RAN standard, it may be referred to as an O-RU (O-RAN RU). The DU (180) may be replaced with a second network entity for a base station (e.g., gNB) in the embodiments of the present disclosure as needed.
[0066] Although FIG. 2 describes a base station comprising a DU and an RU, various embodiments of the present disclosure are not limited thereto. In some embodiments, the base station may be implemented in a distributed deployment according to a centralized unit (CU) configured to perform the functions of the upper layers of the access network (e.g., packet data convergence protocol, RRC (PDCP)) and a distributed unit (DU) configured to perform the functions of the lower layers. In this case, the distributed unit (DU) may include the digital unit (DU) and radio unit (RU) of FIG. 1. Between a core network (e.g., 5G core or next generation core (NGC)) and a radio network (RAN), the base station may be implemented in a structure arranged in the order of CU, DU, and RU. The interface between the CU and the distributed unit (DU) may be referred to as the F1 interface.
[0067] A centralized unit (CU) is connected to one or more DUs and can perform functions at a higher layer than the DUs. For example, the CU may perform functions at the radio resource control (RRC) and packet data convergence protocol (PDCP) layers, while the DU and RU may perform functions at lower layers. The DU may perform functions at the radio link control (RLC), media access control (MAC), and some functions of the physical (PHY) layer (high PHY), while the RU may perform the remaining functions of the PHY layer (low PHY). Additionally, as an example, a digital unit (DU) may be included in a distributed unit (DU) depending on the implementation of a distributed base station deployment. Unless otherwise defined, the operations of the digital unit (DU) and the RU are described below; however, various embodiments of the present disclosure may be applied to both base station deployments including a CU and deployments where the DU is directly connected to the core network without a CU (i.e., implemented by integrating the CU and DU into a single entity).
[0068] FIG. 3 illustrates an example of a centralized-RAN (C-RAN) structure in a wireless communication system according to one embodiment of the present disclosure.
[0069] In a C-RAN structure, baseband units (BBUs) can be installed in a clustered state at a hub site. In a C-RAN structure, a base station may include only an RU and an antenna. According to one embodiment, a vRAN can be constructed based on the C-RAN structure even when virtualization is applied.
[0070] Referring to Figure 3, the Centralized Radio Access Network (C-RAN) structure can improve network management efficiency and performance by centralizing BBUs and deploying RRHs in the field. Through the centralized structure, it is possible to reduce maintenance costs and optimize resources, and increase network flexibility. In the C-RAN structure, since multiple base stations can be connected through a single BBU (or DU), there are economic advantages in terms of BBU hardware or rack space usage, but maximum capacity cell resources may be required to support each cell.
[0071] According to FIG. 3, a wireless communication system according to one embodiment may include components such as a Baseband Unit (BBU), a Remote Radio Head (RRH), and a server.
[0072] The Baseband Unit (BBU) serves as the central device of a base station and is responsible for data processing and control. The BBU performs signal processing at the physical and upper layers and can be connected to multiple RRHs to support remote control and communication. Through a centralized processing structure, the BBU can contribute to optimizing network performance and enhancing management efficiency.
[0073] A Remote Radio Head (RRH) is a device that transmits and receives radio signals from a base station and is typically installed alongside the base station antenna. It can be connected to a BBU via fiber optics and designed to amplify and process signals even at remote locations. This improves communication quality and allows for more flexible base station installation.
[0074] The server can play a role in supporting network management, data storage, and the execution of various applications. In particular, in a centralized-RAN (C-RAN) or open-RAN (O-RAN) architecture according to the example illustrated in Fig. 3, the server can operate with the BBU as a central management unit, thereby increasing the flexibility and scalability of the network. Additionally, it can play an important role in enhancing the user experience through real-time data analysis and network optimization.
[0075] According to Fig. 3, in a C-RAN structure, a BBU can be directly connected to a server. The BBU can communicate with the server via wireless or wired connection, and the BBU and the server can transmit and receive signals and / or messages via wireless or wired connection.
[0076] The server may be connected indirectly to the BBU. Although Figure 3 illustrates the BBU and the server being directly connected, this is not limited to this, and the server and the BBU may be connected through other network entities.
[0077] Although the server and BBU are shown as being configured separately in FIG. 3, according to one embodiment, the server and BBU may be included in a single electronic device and are not limited to the illustrated drawings.
[0078] Although BBU is illustrated in FIG. 3, according to one embodiment, BBU may correspond to DU, and is not limited to the illustrated figure.
[0079] Although a C-RAN structure is illustrated in FIG. 3, the present disclosure is not limited to a C-RAN structure and can be applied to other RAN structures including an O-RAN (open-RAN) structure. In other RAN structures including an O-RAN structure, a connection to a server can be made in the same way as in FIG. 3.
[0080] At least some of the network entities described in FIG. 3 may be virtualized. For example, at least some of the network entities may be virtualized on a cloud platform (e.g., open chassis and blade specification edge cloud) and configured on a device (e.g., a server). This virtualization may support services in dense urban areas due to latency low enough to meet latency requirements and rich fronthaul capacity that allows for baseband unit (BBU) functions pooled at a central location. Since there is no need to attempt near-real-time centralization beyond the limits, the cloud platform may be optimized for the RAN deployment scenario of the present disclosure.
[0081] FIG. 4 illustrates the configuration of an electronic device according to one embodiment of the present disclosure.
[0082] The configuration exemplified in FIG. 4 can be understood as the configuration of the base station (110) of FIG. 1. The configuration exemplified in FIG. 4 can also be understood as the configuration of the DU (160) of FIG. 2. The configuration exemplified in FIG. 4 can also be understood as the configuration of the BBU of FIG. 3. The configuration exemplified in FIG. 4 can also be understood as the configuration of an electronic device including a server of FIG. 3. That is, the electronic device (400) of FIG. 4 may include at least one of a base station, a BBU, a server, or a DU. Terms such as '... unit', '... device' used below refer to a unit that processes at least one function or operation, and this may be implemented as hardware or software, or a combination of hardware and software.
[0083] Referring to FIG. 4, the electronic device (400) may include a transceiver (410), a storage unit (420), and a processor (430).
[0084] The transceiver (410) can perform functions for transmitting and receiving signals in a wired communication environment. The transceiver (410) may include an interface for wired communication to control a direct connection between devices through a transmission medium (e.g., copper wire, optical fiber). For example, the transceiver (410) can transmit an electrical signal to another device through a copper wire or perform conversion between an electrical signal and an optical signal. The transceiver (410) may be connected to a radio unit (RU). The transceiver (410) may be connected to a core network or to a distributed CU. The transceiver (410) may be connected to a server.
[0085] The transceiver (410) may perform functions for transmitting and receiving signals in a wireless communication environment. For example, the transceiver (410) may perform a conversion function between a baseband signal and a bit sequence according to the physical layer specifications of the system. For example, when transmitting data, the transceiver (410) may generate complex symbols by encoding and modulating the transmitted bit sequence. Also, when receiving data, the transceiver (410) may restore the received bit sequence by demodulating and decoding the baseband signal. Additionally, the transceiver (410) may include a plurality of transmission and reception paths. Also, according to one embodiment, the transceiver (410) may be connected to a core network or to other nodes (e.g., an integrated access backhaul).
[0086] The transceiver (410) can transmit and receive signals. To this end, the transceiver (410) may include at least one transceiver. For example, the transceiver (410) can transmit a synchronization signal, a reference signal, system information, a message, a control message, a stream, control information, or data. Additionally, the transceiver (410) can perform beamforming.
[0087] The transmitting and receiving unit (410) can transmit and receive signals as described above. Accordingly, all or part of the transmitting and receiving unit (410) may be referred to as a 'transmitting unit', a 'receiving unit', or a 'transmitting and receiving unit'. Furthermore, in the following description, transmission and reception performed via a wireless channel may be used to mean that processing as described above is performed by the transmitting and receiving unit (410).
[0088] Although not illustrated in FIG. 4, the transceiver (410) may further include a backhaul communication unit for connecting to a core network or another base station. The backhaul communication unit provides an interface for performing communication with other nodes within the network. That is, the backhaul communication unit can convert a bit sequence transmitted from a base station to another node, e.g., another connection node, another base station, an upper node, a core network, etc., into a physical signal, and can convert a physical signal received from another node into a bit sequence.
[0089] The memory (420) can store data such as basic programs, applications, and configuration information for the operation of a base station, DU, BBU, and / or server. The memory (420) may include a buffer. The memory (420) may be composed of volatile memory, non-volatile memory, or a combination of volatile and non-volatile memory. Additionally, the memory (420) can provide stored data upon request from the processor (430).
[0090] The processor (430) controls the overall operations of the electronic device (400). For example, the processor (430) can transmit and receive signals through the transceiver (410) (or through the backhaul communication unit). Additionally, the processor (430) can write and read data to and from memory (420). Furthermore, the processor (430) can perform the functions of a protocol stack required by the communication standard. To this end, the processor (430) may include at least one processor. In some embodiments, the processor (430) may include a control message generator that generates a control plane message having an extended field containing a regularization factor, and a management message generator that generates a management message to disable the regularization factor field of a message containing an existing regularization factor (e.g., a control plane message of Section Type 6 of O-RAN). The control message generation unit and the management message generation unit may be a set of instructions or code stored in the processor (430), at least temporarily resided in the processor (430), or a storage space storing instructions / code, or a part of the circuitry constituting the processor (430). According to various embodiments, the processor (430) may control the electronic device (400) to perform operations according to various embodiments described below.
[0091] The configuration of the electronic device illustrated in FIG. 4 is merely an example, and the examples of electronic devices for performing various embodiments of the present disclosure are not limited to the configuration illustrated in FIG. 4. That is, it is obvious that some configurations may be added, deleted, or changed depending on various embodiments.
[0092] FIG. 5 is a diagram illustrating an example of reallocating computing resources allocated to a plurality of cells based on a deep reinforcement learning algorithm according to one embodiment of the present disclosure.
[0093] Referring to FIG. 5, according to one embodiment, a plurality of cells may be connected to an electronic device. The plurality of cells may be connected to a baseband unit (BBU) or a DU via wired or wireless communication. A cell in the present disclosure may include an area serviced by a base station. A cell in the present disclosure may include a service coverage area of a network. For example, referring to 510 in FIG. 5, N cells may be connected to a BBU. However, the electronic device connected to the plurality of cells is not limited to a BBU.
[0094] According to one embodiment, each of the plurality of cells may have a different traffic density. The traffic density in the present disclosure may indicate the degree to which data transmission demand is concentrated for each of the plurality of cells. That is, the traffic density may indicate the degree of data transmission demand of user terminals in each of the plurality of cells. Accordingly, the traffic density may indicate the concentration of radio resources allocated to process data for each of the plurality of cells. That is, the traffic density may indicate the degree of capacity of radio resources allocated to each of the plurality of cells. For example, a high traffic density may indicate that a large amount of radio resources may be allocated to the corresponding cell, and a low traffic density may indicate that a small amount of radio resources may be allocated to the corresponding cell. For example, referring to 510 in FIG. 5, it can be seen that each of the N cells has a different traffic density. For example, in 510 of FIG. 5, cell 5 may have a relatively large amount of radio resources allocated and may have a high traffic density because there is a large amount of data to process. Accordingly, cell 5 may be in a network overload state. Also, in 510 of FIG. 5, cell 1 may have relatively less wireless resources allocated because it has less data to process, and may have a low traffic density. Accordingly, cell 1 may be processing relatively less traffic compared to cell 5. Referring to FIG. 5, among the multiple cells, the cell with high traffic density may be shown in bold, and among the multiple cells, the cell with low traffic density may be shown in light.
[0095] According to one embodiment, the traffic density of a plurality of cells may vary. The traffic density of a plurality of cells may vary in real time as the wireless resources being transmitted and received change. That is, the traffic density of a plurality of cells may not be static, and the traffic density of a plurality of cells may vary dynamically. For example, in 510 of FIG. 5, the traffic density of cell 5 is shown as high, but as the data transmission demand of user terminals in each of the plurality of cells changes dynamically, the traffic density of other cells may become higher than the traffic density of cell 5. Therefore, in order to improve network performance by efficiently processing wireless resources in response to dynamically changing traffic density, it is necessary to process wireless resources by allocating computing resources in response to the dynamically changing traffic density.
[0096] According to one embodiment, wireless resources can be processed according to computing resources allocated to a plurality of cells. In the present disclosure, computing resources may be hardware (and / or software) resources required to process and manage data at a base station. Computing resources may be computational resources required to perform tasks for controlling, scheduling, traffic processing, and / or increasing transmission efficiency of wireless resources. Computing resources in the present disclosure may include computing resources of a BBU and / or computing resources of a DU. For example, a plurality of cells may process the wireless resources of the cells according to computing resources allocated to the plurality of cells.
[0097] According to one embodiment, a computing resource unit may be a computing resource required for processing a packet allocated to a slot. A computing resource may be a set of computing resource units. For example, it may be assumed that in a cell allocated with one computing resource unit, one packet can be processed in one slot. In this case, in a cell allocated with three computing resource units, three packets can be processed in one slot. However, computing resource units are not limited to the examples described above.
[0098] According to one embodiment, if computing resources allocated to multiple cells are not dynamically allocated, there may be computing resources wasted depending on wireless resources. That is, if computing resources allocated to multiple cells remain fixed and do not change in real time, computing resources may be wasted as the wireless resources used by the multiple cells change. For example, referring to 520 in FIG. 5, two computing resources may be allocated to each of the N cells, from cell 1 to cell N. In this case, although computing resources are allocated equally in the same number of two to the multiple cells, since the traffic density differs for each cell, there may be a shortage of computing resources in cells 2, 4, 5, and 6, resulting in a large number of waiting packets. On the other hand, computing resources may be wasted in cells 1, 3, 7, and N because the traffic density is low. In 520 of FIG. 5, overall network performance may deteriorate because there is a large amount of wasted computing resources.
[0099] According to one embodiment, the first electronic device may dynamically allocate computing resources allocated to a plurality of cells. That is, the first electronic device may reallocate wasted computing resources to other cells. For example, computing resources may be allocated according to 530 of FIG. 5 by reallocating wasted computing resources from cell 1, cell 3, cell 7, and cell N to other cells according to 520 of FIG. 5. Referring to 530 of FIG. 5, as the first electronic device reallocates computing resources, network performance can be improved by efficiently allocating limited computing resources. Referring to 530 of FIG. 5, data packets can be processed efficiently by the first electronic device dynamically allocating more computing resources to cells with high data processing demand.
[0100] According to one embodiment, the first electronic device may reallocate computing resources allocated to a plurality of cells according to a fixed period. For example, the first electronic device may reallocate wasted computing resources to another cell every 5 seconds. However, the period during which the first electronic device reallocates computing resources is not limited to the example described above. Although an embodiment has been described in which the first electronic device reallocates wasted computing resources to another cell, it is not limited thereto, and even in cases where computing resources are not wasted, the first electronic device may reallocate computing resources allocated to a plurality of cells according to a fixed period in order to improve network performance.
[0101] According to one embodiment, the first electronic device may reallocate computing resources allocated to a plurality of cells based on a deep reinforcement learning (DRL) algorithm. The first electronic device may reallocate computing resources allocated to a plurality of cells based on a deep Q network (DQN) algorithm. The first electronic device may reallocate computing resources allocated to a plurality of cells based on a double DQN (DDQN) algorithm. The specific operation of the deep reinforcement learning algorithm will be described later.
[0102] According to one embodiment, the first electronic device may apply network state information of a plurality of cells to deep reinforcement learning to reallocate computing resources allocated to a plurality of cells. For example, the first electronic device may apply information regarding at least one of the number of terminals activated in a plurality of cells, the utilization rate of wireless resources of a plurality of cells, or the state in which computer resources are allocated to a plurality of cells to deep reinforcement learning to reallocate computing resources allocated to a plurality of cells.
[0103] In one embodiment, the network performance of a plurality of cells can be improved by causing the first electronic device to reallocate the computing resources of the second electronic device. For example, the network performance of a plurality of cells connected to the BBU can be improved by causing the first electronic device to reallocate the computing resources of the BBU. That is, by reducing the computing resources allocated to cells where computing resources are being wasted and allocating more computing resources to cells with high traffic density, the effect of efficiently handling traffic load and efficiently utilizing limited computing resources to improve network performance can be achieved.
[0104] FIG. 6 is a diagram illustrating an operation to reallocate computing resources allocated to a plurality of cells based on a deep reinforcement learning algorithm in a first electronic device and a second electronic device according to an embodiment of the present disclosure.
[0105] Referring to FIG. 6, according to one embodiment, a first electronic device can obtain network status information of a plurality of cells connected to a second electronic device. The network status information is information indicating the network status of a plurality of cells connected to the second electronic device, and may include signal information of the plurality of cells, traffic load being processed in the plurality of cells, delay time of data packets transmitted in the plurality of cells, activation status of the plurality of cells, transmission bandwidth of the plurality of cells, and status of terminals connected to the plurality of cells. For example, the network status information may include information regarding the number of terminals activated in the plurality of cells, the utilization rate of radio resources of the plurality of cells, and the state in which computing resources are allocated to the plurality of cells. For example, referring to FIG. 6, the first electronic device can obtain network status information of a plurality of cells connected to the second electronic device in the "Observe State" stage.
[0106] According to one embodiment, the first electronic device can obtain a plurality of Q values by applying network state information of a plurality of cells to a deep reinforcement learning algorithm. The first electronic device can obtain Q values equal to the number of cells connected to the second electronic device through the deep reinforcement learning algorithm. A Q value is a value representing the expected cumulative reward that can be obtained when a specific action is performed in a specific state in reinforcement learning, and in the present disclosure, it may refer to a value representing the cumulative degree of improvement in network performance when the computing resources of the second electronic device are reallocated. The Q value in the present disclosure may be a value calculated cumulatively as the degree of increase in the throughput of a cell when additional computing resources are allocated to the cell. For example, referring to FIG. 6, the first electronic device can perform a "Take Action" step using the obtained plurality of Q values.
[0107] According to one embodiment, the first electronic device may instruct the second electronic device to reallocate computing resources based on a plurality of Q values. The first electronic device may generate a control message instructing the second electronic device to reallocate computing resources using a plurality of Q values, and may reallocate computing resources by transmitting the control message to the second electronic device. For example, referring to FIG. 6, the first electronic device may generate a control message and instruct the second electronic device to reallocate computing resources in the "Take Action" step.
[0108] According to one embodiment, the second electronic device can reallocate computing resources based on a control message of the first electronic device. According to one embodiment, an operation of reallocating computing resources based on a Q value will be described later.
[0109] According to one embodiment, the first electronic device can obtain information related to the throughput of a plurality of cells. The first electronic device can obtain information related to the throughput of a plurality of cells from the second electronic device. The information related to the throughput of a plurality of cells may include information regarding the throughput of a plurality of cells according to the reallocated computing resources after reallocating computing resources based on a control message. Based on the information related to the throughput of a plurality of cells, the first electronic device can determine a reward for the reallocation. For example, referring to FIG. 6, the first electronic device can obtain information related to the throughput of a plurality of cells from the second electronic device and determine a reward in the "Get Reward" step. The operation of determining the reward will be described later.
[0110] According to one embodiment, the first electronic device may store information in a buffer for training a deep reinforcement learning algorithm. In one embodiment, the first electronic device may store network state information received from the second electronic device in the buffer. For example, the first electronic device may receive network state information of a plurality of cells from the MAC layer of the BBU and store it in the buffer. In one embodiment, the first electronic device may store information related to the reallocation of a plurality of cells in the buffer. For example, the first electronic device may store information regarding cells that it has decided to reduce computing resources based on a Q value, and cells that it has decided to add computing resources to, in the buffer. In one embodiment, the first electronic device may store a reward value in the buffer. For example, the first electronic device may store a reward value determined based on information related to throughput in the buffer.
[0111] According to one embodiment, the first electronic device can train a deep reinforcement learning algorithm by applying a mini-batch of information sampled from a buffer. The mini-batch may include network state information at a first time point, a reallocation at the first time point, a reward value at the first time point, and network state information at a second time point, which is the time point following the first time point. Detailed operations for training a deep reinforcement learning algorithm by applying a mini-batch will be described later.
[0112] According to one embodiment, the first electronic device may train a deep reinforcement learning algorithm through offline learning. For example, the first electronic device may train the deep reinforcement learning algorithm by applying a mini-batch, which is a pre-collected data set stored in a buffer. However, the method of training the deep reinforcement learning algorithm according to the present disclosure is not limited to offline learning and may be trained through online learning in real time.
[0113] FIG. 7 is a flowchart (700) illustrating an operation of reallocating computing resources allocated to a plurality of cells based on a deep reinforcement learning algorithm in a first electronic device according to one embodiment of the present disclosure.
[0114] Referring to FIG. 7, in operation 710, the first electronic device can obtain network status information of a plurality of cells.
[0115] According to one embodiment, the second electronic device can collect network status information of a plurality of cells connected to the second electronic device. For example, the BBU can collect at least one of terminal-specific scheduling information at the MAC layer, traffic information of a plurality of cells, information regarding a packet queue when packets are waiting at a plurality of cells, or network performance information including the data transmission speed of a plurality of cells. However, the second electronic device is not limited to collecting network status information from the BBU and may also collect network status information through other entities.
[0116] According to one embodiment, the first electronic device can obtain network state information from the second electronic device. The first electronic device can receive network state information collected by the second electronic device. The first electronic device can select necessary information from the network state information received from the second electronic device. For example, the first electronic device can select information to be applied to a deep reinforcement learning algorithm from the received network state information.
[0117] According to one embodiment, the first electronic device can acquire network status information at regular statistical intervals. For example, the first electronic device can acquire network status information at intervals of 5 seconds. However, the regular statistical interval is not limited to 5 seconds, and the first electronic device can acquire network status information at shorter or longer intervals. Additionally, in one embodiment, the first electronic device may acquire network status information at irregular intervals.
[0118] According to one embodiment, network status information may include the number of terminals activated in a plurality of cells. The activated terminals may include terminals connected to the cell that transmit or receive data or maintain a connection state. However, the information included in the network status information is not limited to the number of terminals activated in a plurality of cells, and the network status information may include other information indicating the cell load status or other information indicating the demand for computing resources.
[0119] According to one embodiment, network status information may include the utilization rate of wireless resources of a plurality of cells. The utilization rate of wireless resources may be information indicating how efficiently wireless resources allocated at the cell level are being used. However, the information included in the network status information is not limited to the utilization rate of wireless resources and may include other information indicating opportunities for packet processing, such as the number of waiting packets and the traffic density of the cell.
[0120] According to one embodiment, network state information may include a state in which computing resources are allocated to a plurality of cells. For example, network state information may include information on how many computing resource units are allocated to each of the plurality of cells. The first electronic device may adjust the minimum number of computing resource units or the maximum number of computing resource units allocated to the plurality of cells by utilizing the state in which computing resources are allocated to the plurality of cells.
[0121] In operation 720, the first electronic device can obtain multiple Q values by applying network state information to a deep reinforcement learning algorithm.
[0122] According to one embodiment, the first electronic device can apply network state information to a deep reinforcement learning algorithm. The first electronic device can input network state information to a main deep learning neural network trained by the deep reinforcement learning algorithm.
[0123] According to one embodiment, the first electronic device can obtain a plurality of Q values equal to the number of cells. For example, if N cells are connected to the second electronic device, the first electronic device can obtain N Q values from a deep reinforcement learning algorithm. Each of the N Q values can correspond to each of the N cells. An example of the plurality of Q values will be described later in FIG. 8.
[0124] According to one embodiment, the Q value may be an expected cumulative reward value when computing resources are reallocated based on network state information applied to a deep reinforcement learning algorithm. For example, the Q value may increase if network performance improves as computing resources are reallocated. For example, for cells with a large amount of wasted computing resources, network performance may not improve even if computing resources are reallocated, so the Q value for cells with a large amount of wasted computing resources may be relatively small. For example, for cells with high traffic density, network performance may improve when computing resources are reallocated, so the Q value for cells with high traffic density may be relatively large. However, the Q value is not limited to the examples described above.
[0125] According to one embodiment, the first electronic device can acquire a Q value by utilizing the epsilon greedy technique and collect various state information, behavior information, and reward information. The first electronic device can acquire network state information of a plurality of cells by utilizing the epsilon greedy technique, generate a control message instructing the reallocation of computing resources, and acquire information related to the throughput of a plurality of cells. The operation of acquiring a Q value and collecting various information by utilizing the epsilon greedy technique will be described later in FIG. 12.
[0126] According to one embodiment, the first electronic device may store network state information applied to a deep reinforcement learning algorithm in a buffer. The network state information stored in the buffer may be used to train the deep reinforcement learning algorithm. An operation to train the deep reinforcement learning algorithm using the buffer in which the network state information is stored will be described later.
[0127] In operation 730, the first electronic device may generate a control message instructing the reallocation of computing resources based on a plurality of Q values.
[0128] According to one embodiment, the first electronic device may generate a control message instructing to additionally allocate computing resources to cells with a high Q value and to reduce computing resources to cells with a low Q value. For example, the first electronic device may generate a control message instructing to reallocate computing resources from cells with a low Q value among N cells to computing resources from cells with a high Q value.
[0129] According to one embodiment, the first electronic device may generate a control message instructing the reallocation of computing resources by selecting a maximum value (max) and a minimum value (min) among a plurality of Q values. For example, the first electronic device may select a maximum value and a minimum value among a plurality of Q values and generate a control message instructing the reallocation of computing resources of a cell corresponding to the minimum value to a cell corresponding to the maximum value. For example, the first electronic device may generate a control message instructing the reallocation of one computing resource unit of a cell corresponding to the minimum value to a cell corresponding to the maximum value. For example, the first electronic device may generate a control message instructing the reallocation of one computing resource unit of a cell corresponding to the value smaller than the minimum value to a cell corresponding to the value larger than the maximum value. However, the operation of generating a control message instructing the reallocation of computing resources based on Q values is not limited to the examples described above, and a control message instructing the reallocation of computing resources may also be generated based on other methods.
[0130] According to one embodiment, a control message instructing to reallocate computing resources may be a message instructing to reallocate a computing resource unit of one cell to another cell. For example, the control message may include a cell number to reduce the computing resource unit and a cell number to reallocate the computing resource unit. However, the above-described example is not limited, and the control message may include instructions to reallocate multiple computing resource units.
[0131] In operation 740, the first electronic device can transmit a control message to the second electronic device.
[0132] According to one embodiment, the first electronic device may transmit a control message to the second electronic device instructing it to re-allocate computing resources allocated to a plurality of cells. For example, the first electronic device may transmit a control message to the second electronic device instructing it to re-allocate the computing resource unit of the cell corresponding to the minimum value among the Q values to the cell corresponding to the maximum value among the Q values.
[0133] According to one embodiment, the second electronic device may reallocate computing resources in response to receiving a control message from the first electronic device. For example, the BBU may receive a control message and reallocate computing resources based on the control message. For example, the DU may receive a control message and reallocate computing resources based on the control message.
[0134] According to one embodiment, the first electronic device may store reassignment information in a buffer. The reassignment information stored in the buffer may be used to train a deep reinforcement learning algorithm. For example, the reassignment information stored in the buffer may include re-allocation. For example, the reassignment information stored in the buffer may include the number of a cell to be reassigned. An operation to train a deep reinforcement learning algorithm using a buffer in which reassignment information is stored will be described later.
[0135] According to one embodiment, the first electronic device and the second electronic device according to the present disclosure may be located within a single device. For example, the first electronic device and the second electronic device may be located within a single electronic device, and within the single electronic device, the first electronic device may transmit a control message to the second electronic device. However, the present disclosure is not limited thereto, and the first electronic device and the second electronic device may be located within a single electronic device or may be separate electronic devices.
[0136] In operation 750, the first electronic device can obtain information related to the throughput of multiple cells from the second electronic device.
[0137] According to one embodiment, information regarding the throughput of a plurality of cells may be information regarding the amount of data transmitted and received from the cells during a specific period. For example, a cell that has been reallocated computing resources may have its throughput increased.
[0138] According to one embodiment, the first electronic device may determine a reward for reallocation based on information related to throughput. For example, the first electronic device may identify an increase in throughput for a cell that has been reallocated computing resources and determine a reward based on the increase in throughput. For example, the first electronic device may determine that the reward is large if the increase in throughput of the cell that has been reallocated computing resources is large. However, the example in which the first electronic device determines the reward is not limited to the example described above, and the reward may be determined according to other examples in which the reward is determined based on the degree to which network performance is improved by reallocation. The operation of the first electronic device determining the reward based on information related to throughput will be described later in FIG. 11.
[0139] According to one embodiment, the first electronic device can determine a reward for reallocation based on information related to network status information, reallocation, and throughput. For example, the first electronic device can determine the reward by using not only information related to throughput, but also network status information obtained in operation 710 and reallocation-related information included in the control message generated in operation 730.
[0140] According to one embodiment, the first electronic device may store a reward value resulting from reallocation in a buffer. The reward value stored in the buffer may be used to train a deep reinforcement learning algorithm. An operation to train a deep reinforcement learning algorithm using a buffer in which a reward value is stored will be described later.
[0141] According to one embodiment, the first electronic device may perform operation 710 after performing operation 750. For example, the first electronic device may perform an operation to reallocate computing resources by performing operation 710 at regular intervals. For example, the first electronic device may identify network information of a plurality of cells by performing operation 710 every 5 seconds.
[0142] According to one embodiment, a mini-batch may be a data bundle sampled from information stored in a buffer. The buffer may include network status information at a current time point, reallocation, compensation value, and network status information at a next time point. The mini-batch may include network status information at a first time point, reallocation at the first time point, compensation value at the first time point, and network status information at a second time point, which is the time point next to the first time point, as a single data bundle. The first time point and the second time point may be distinguished based on the time point at which the network status information is acquired. For example, the buffer may store network status information at time point x, information reallocated based on the network status information at time point x, compensation value resulting from the reallocation based on the network status information at time point x, and network status information at time point x+1, which is the time point next to time point x, as a single data bundle, a single mini-batch. The network status information at time point x+1 may include network status information acquired as the first electronic device performs operation 710 after acquiring the network status information at time point x.
[0143] According to one embodiment, the first electronic device can train a deep reinforcement learning algorithm by applying a mini-batch of sampled information stored in a buffer at each first period. For example, the first period may be a specific time unit. For example, the first period may be determined according to the capacity of the information stored in the buffer. For example, the first period may be set when the capacity of the data stored in the buffer reaches 80%. For example, the first period may be set when the number of data bundles stored in the buffer reaches 50,000. However, it is not limited to the examples described above.
[0144] FIG. 8 is a drawing for explaining an example of a plurality of Q values obtained by a first electronic device according to one embodiment of the present disclosure.
[0145] Referring to FIG. 8, the first electronic device can obtain multiple Q values by applying network state information to a deep reinforcement learning algorithm.
[0146] According to one embodiment, a plurality of Q values obtained by the first electronic device may correspond to the number of cells connected to the second electronic device. For example, if there are N cells connected to the second electronic device, referring to FIG. 8, the first electronic device can obtain N Q values by applying network state information regarding the N cells to a deep reinforcement learning algorithm.
[0147] According to one embodiment, the Q value of each cell may be a value representing the cumulative reward when computing resources are reallocated to each cell. For example, referring to FIG. 8, the cumulative reward of cell 1, which is the Q value of cell 1, may be a value representing the cumulative reward when computing resources are reallocated to cell 1. The cumulative reward of cell 2, which is the Q value of cell 2, may be a value representing the cumulative reward when computing resources are reallocated to cell 2. The cumulative reward of cell N, which is the Q value of cell N, may be a value representing the cumulative reward when computing resources are reallocated to cell N.
[0148] According to one embodiment, the first electronic device can obtain a plurality of Q values by inputting network state information into a main deep learning neural network included in a deep reinforcement learning algorithm. The deep reinforcement learning algorithm may include a main deep learning neural network and a target deep learning neural network. The first electronic device can obtain Q values by inputting network state information into the main deep learning neural network other than the target deep learning neural network. However, it is not limited to the examples described above.
[0149] FIG. 9 is a diagram illustrating a method for a first electronic device to reallocate computing resources according to a Q value, according to one embodiment of the present disclosure.
[0150] Referring to FIG. 9, according to one embodiment, a first electronic device may select a maximum value and a minimum value among a plurality of Q values and generate a control message instructing to reallocate the computing resources of the cell corresponding to the minimum value to the cell corresponding to the maximum value. For example, referring to 910 in FIG. 9, if the Q value of cell 5 is the largest and the Q value of cell 3 is the smallest among N Q values, the first electronic device may generate a control message instructing to reallocate the computing resources of cell 3 to cell 5. More specifically, the first electronic device may generate a control message instructing to reallocate the computing resource units of cell 3 to cell 5. For example, referring to 910 in FIG. 9, the first electronic device may generate a control message instructing to reallocate the wasted computing resource units (911) among the computing resources of cell 3 to cell 5.
[0151] According to one embodiment, the second electronic device may reallocate computing resources in accordance with a control message of the first electronic device. For example, referring to 920 in FIG. 9, the second electronic device may reallocate one computing resource unit (921) to cell 5 by reallocating a computing resource unit of cell 3 to cell 5. Referring to 920 in FIG. 9, in accordance with the reallocation of computing resources, one computing resource unit may be allocated to cell 3 and four computing resource units may be allocated to cell 5.
[0152] In one embodiment, an example is presented in which a first electronic device generates a control message and a second electronic device reallocates computing resources according to the control message, but this is not limited thereto. The first electronic device and the second electronic device may be included in a single electronic device, and the aforementioned control message may be transmitted and received within the single electronic device, or the first electronic device may be made to reallocate computing resources through signals other than the control message.
[0153] FIG. 10 is a drawing for illustrating an example in which a first electronic device determines a reward for reallocating computing resources based on information related to throughput, according to one embodiment of the present disclosure.
[0154] According to one embodiment, the first electronic device may transmit a control message to the second electronic device instructing it to reallocate computing resources allocated to a plurality of cells. The second electronic device may reallocate computing resources allocated to a plurality of cells in accordance with the control message. For example, referring to 1010 in FIG. 10, the second electronic device may receive a control message instructing it to reallocate one computing resource unit of cell 3 to cell 5. Referring to FIG. 10, the second electronic device may reallocate computing resources (1011). 1020 in FIG. 10 may be a state in which the second electronic device has reallocated computing resources.
[0155] According to one embodiment, the first electronic device may reallocate computing resources and obtain information related to throughput from the second electronic device. The first electronic device may obtain information related to the throughput of a plurality of cells in a reallocated state according to 1020 of FIG. 10 from the second electronic device.
[0156] According to one embodiment, the first electronic device can identify an increase in throughput resulting from reassignment based on information related to throughput. For example, referring to FIG. 10, the first electronic device can obtain information related to throughput resulting from reassignment (1011) based on information related to throughput of a plurality of cells obtained from the second electronic device. For example, referring to 1021 in FIG. 10, the first electronic device can identify an increase in throughput resulting from reassignment based on information related to throughput at 1010 in FIG. 10 before reassignment (1011) and information related to throughput at 1020 in FIG. 10 after reassignment (1011). For example, the first electronic device can obtain information regarding the throughput of a plurality of cells at 1010 included in the network status information before reassignment (1011) in FIG. 10. The first electronic device can obtain information regarding throughput after the reallocation (1011) of FIG. 10 and identify the increase in throughput.
[0157] According to one embodiment, information related to throughput may include the throughput of a cell having reallocated computing resources. For example, referring to FIG. 10, computing resources of cell 5 may be reallocated according to reallocation (1011), and the first electronic device may obtain the throughput of cell 5. However, information related to throughput is not limited to the throughput of a cell having reallocated computing resources and may include the throughput of a plurality of cells.
[0158] According to one embodiment, the first electronic device may determine a reward value for reallocation based on an increase in throughput. For example, referring to 1022 in FIG. 10, the first electronic device may determine a reward value based on an increase in throughput resulting from reallocation (1011). The reward value may be determined to have a larger value as the increase in throughput is relatively larger. The reward value may be determined to have a smaller value as the increase in throughput is relatively smaller. For example, if the increase in throughput resulting from reallocation for cell 2 is 50 Mbps and the increase in throughput resulting from reallocation for cell 4 is 100 Mbps, the reward value for reallocation for cell 4 may be determined to have a larger value than the reward value for reallocation for cell 2.
[0159] FIG. 11 is a flowchart (1100) illustrating the operation of reallocating computing resources allocated to a plurality of cells based on a deep reinforcement learning algorithm in a first electronic device according to one embodiment of the present disclosure, and the operation of training a deep reinforcement learning algorithm in the first electronic device. FIG. 11 may be described with reference to FIG. 7.
[0160] Referring to FIG. 11, in operation 710, the first electronic device can identify network status information of a plurality of cells. Operation 710 of FIG. 11 may be an operation corresponding to operation 710 of FIG. 7.
[0161] In operation 1110, the first electronic device can determine whether any value is lower than the Epsilon value.
[0162] According to one embodiment, the first electronic device can obtain multiple Q values and perform various reassignment actions by applying network state information to a deep reinforcement learning algorithm using an Epsilon greedy technique. The Epsilon greedy technique may be a technique that compares a random value between 0 and 1 with an Epsilon value, performs an exploration phase when the Epsilon value is greater than the random value, and performs an exploitation phase when the Epsilon value is smaller than the random value. In the Epsilon greedy technique, the degree of the exploration phase is controlled by Epsilon, and the deep reinforcement learning algorithm can be trained efficiently by gradually decreasing the Epsilon value.
[0163] According to one embodiment, the random value may be a random value between 0 and 1 that is set whenever the first electronic device identifies network status information of a plurality of cells according to operation 710.
[0164] According to one embodiment, the Epsilon value may be a value that is gradually reduced according to the Epsilon greedy technique. For example, the first electronic device may gradually reduce the Epsilon value by subtracting a constant value (e.g., 0.001) from the Epsilon value whenever operation 710 is performed. For example, the first electronic device may gradually reduce the Epsilon value by reducing the Epsilon value when the average reward value exceeds a threshold value. For example, the first electronic device may gradually reduce the Epsilon value based on the number of times operation 710 is performed. However, the method by which the first electronic device reduces the Epsilon value is not limited to the examples described above and may vary depending on the learning environment and learning objective of the first electronic device.
[0165] According to one embodiment, if a random value is lower than the Epsilon value, the first electronic device can perform operation 1120. According to one embodiment, if a random value is higher than the Epsilon value, the first electronic device can perform operation 720.
[0166] In operation 1120, the first electronic device may generate a control message instructing to reallocate computing resources at will. Operation 1120 of FIG. 11 may correspond to the exploration phase of the Epsilon greedy technique.
[0167] According to one embodiment, the first electronic device may generate a control message instructing the reallocation of computing resources randomly without relying on a plurality of Q values. For example, the first electronic device may generate a control message instructing the reallocation of computing resources completely randomly without obtaining a plurality of Q values corresponding to the output of a deep reinforcement learning algorithm. For example, the first electronic device may generate a control message instructing the reallocation of computing resources by selecting two cells among a plurality of cells without obtaining a plurality of Q values. In one embodiment, the first electronic device may instruct the reallocation of computing resources randomly to a plurality of cells. For example, the first electronic device may generate a control message instructing the reallocation of computing resources completely randomly without utilizing a plurality of Q values obtained by applying acquired network state information. However, the control message instructing the reallocation of computing resources randomly is not limited to the examples described above.
[0168] According to one embodiment, the first electronic device can obtain a plurality of Q values by inducing an exploration action based on a hint. For example, the first electronic device can generate a control message instructing to reallocate computing resources randomly using a Boltzmann exploration technique or a Softmax exploration technique.
[0169] However, the method by which the first electronic device generates a control message according to operation 1120 is not limited to the example described above, and the first electronic device may generate a control message instructing the reallocation of computing resources by utilizing other methods used in general Epsilon greedy techniques.
[0170] In operation 720, the first electronic device can obtain multiple Q values by applying network state information to a deep reinforcement learning algorithm. Operation 720 of FIG. 11 may be an operation corresponding to operation 720 of FIG. 7. Operation 720 of FIG. 11 may correspond to the exploitation phase of the Epsilon greedy technique.
[0171] In operation 730, the first electronic device may generate a control message instructing the reallocation of computing resources based on a plurality of Q values. Operation 730 of FIG. 11 may be an operation corresponding to operation 730 of FIG. 7.
[0172] In operation 740, the first electronic device can transmit a control message to the second electronic device.
[0173] According to one embodiment, in operation 740 performed after operation 1120, the first electronic device may transmit a control message to the second electronic device instructing it to reallocate a computing resource generated by the first electronic device.
[0174] According to one embodiment, in operation 740 performed after operation 730, the first electronic device may transmit a control message to the second electronic device instructing the first electronic device to reallocate computing resources generated based on a plurality of Q values.
[0175] In operation 1130, the first electronic device can store network status information, re-allocation and compensation values in a buffer.
[0176] According to one embodiment, the first electronic device may store network status information in a buffer. The network status information may include information regarding the number of terminals activated in a plurality of cells, the utilization rate of wireless resources in a plurality of cells, and the state in which computing resources are allocated to a plurality of cells.
[0177] According to one embodiment, the first electronic device may store the re-allocation of computing resources in a buffer. The re-allocation of computing resources may include information regarding the re-allocation of computing resources allocated to a plurality of cells based on a plurality of Q values. For example, the re-allocation may include information regarding a cell among the plurality of cells in which a change in allocated computing resources has occurred. For example, the re-allocation may include information regarding a cell among the plurality of cells in which computing resources have decreased due to the re-allocation and information regarding a cell in which computing resources have increased due to the re-allocation. However, the re-allocation of computing resources is not limited to the examples described above and may include a control message instructing the re-allocation of computing resources allocated to a plurality of cells and the included information.
[0178] According to one embodiment, the first electronic device may store a compensation value resulting from the reallocation of computing resources in a buffer. For example, the first electronic device may determine an increase in throughput resulting from the reallocation based on information related to throughput, and may determine (or extract) a compensation value based on the increase in throughput. The first electronic device may store the determined (or extracted) compensation value in a buffer.
[0179] According to one embodiment, the buffer may be volatile memory. The buffer may be used to train a deep reinforcement learning algorithm, and since the buffer is configured as volatile memory, the training speed of the deep reinforcement learning algorithm using the buffer can be increased. The first electronic device can reduce the data access time through the buffer, which is volatile memory. However, the buffer is not limited to volatile memory, and the buffer may be configured as non-volatile memory depending on the implementation purpose or implementation environment of the first electronic device.
[0180] According to one embodiment, a mini-batch may be a bundle of data sampled to use information stored in a buffer for training.
[0181] According to one embodiment, a mini-batch may include network status information at a first time point, reallocation at the first time point, a compensation value based on the reallocation at the first time point, and network status information at a second time point, which is the time point following the first time point. For example, at the first time point, a first electronic device may acquire network status information of a plurality of cells. Based on the network status information at the first time point, the first electronic device may be made to reallocate computing resources. The first electronic device may determine a compensation value based on the reallocation based on the network status information at the first time point. The first electronic device may acquire network status information of a plurality of cells at a second time point, which is the time point following the first time point in which the network status information of the plurality of cells was acquired. The network status information at the second time point may be network status information based on computing resources reallocated according to the reallocation based on the network status information at the first time point.
[0182] In operation 1140, the first electronic device can determine whether the data stored in the buffer is greater than or equal to a threshold.
[0183] According to one embodiment, the first electronic device can determine whether the capacity of data stored in the buffer is greater than or equal to a threshold capacity. For example, the first electronic device can determine whether the capacity of data stored in the buffer is greater than or equal to a threshold capacity of 700 MB. However, the threshold capacity is not limited to the example described above. In one embodiment, if the first electronic device determines that the capacity of data stored in the buffer is greater than or equal to a threshold capacity, it can perform operation 1150.
[0184] According to one embodiment, the first electronic device can determine whether the number of mini-batches, which are bundles of data stored in a buffer, is greater than or equal to a threshold number. For example, the first electronic device can determine whether the number of mini-batches, which are bundles of data including network status information at a first time point, reallocation at the first time point, a compensation value due to reallocation at the first time point, and network status information at a second time point, which is the time point following the first time point, is greater than or equal to a threshold number of 50,000. However, the threshold number is not limited to the example described above. In one embodiment, if the first electronic device determines that the number of mini-batches, which are bundles of data stored in a buffer, is greater than or equal to a threshold number, it can perform operation 1150.
[0185] According to one embodiment, the first electronic device can determine whether the number of times data is stored in the buffer is greater than or equal to a threshold number. For example, the first electronic device can determine whether the number of times a reward value is stored in the buffer is greater than or equal to a threshold number of 50,000 times. However, the threshold number is not limited to the example described above. In one embodiment, if the first electronic device determines that the number of times data is stored in the buffer is greater than or equal to a threshold number, it can perform operation 1150.
[0186] According to one embodiment, if the first electronic device determines YES in operation 1140, the first electronic device can reset the buffer.
[0187] According to one embodiment, if the first electronic device determines NO in operation 1140, the first electronic device may perform operation 710 without resetting the buffer. For example, the first electronic device may perform an operation to reallocate computing resources by performing operation 710 at regular intervals. For example, the first electronic device may identify network information of multiple cells by performing operation 710 every 5 seconds.
[0188] In operation 1150, the first electronic device can train a deep reinforcement learning algorithm.
[0189] In one embodiment, the deep reinforcement learning algorithm may be a DQN (deep Q network) algorithm or a DDQN (double deep Q network) algorithm. A deep reinforcement learning algorithm that is a DDQN algorithm may include a main deep learning neural network and a target deep learning neural network.
[0190] According to one embodiment, the operation of a first electronic device training a deep reinforcement learning algorithm is examined in FIG. 12.
[0191] FIG. 12 is a flowchart (1150) showing the operation of a first electronic device training a deep reinforcement learning algorithm according to one embodiment of the present disclosure.
[0192] Referring to FIG. 12, the flowchart (1150) of FIG. 12 may be a flowchart illustrating the operation of a first electronic device training a deep reinforcement learning algorithm according to the operation 1150 of FIG. 11.
[0193] According to one embodiment, if the first electronic device determines that the data stored in the buffer is greater than or equal to a threshold according to operation 1140, it may perform operation 1210.
[0194] In operation 1210, the first electronic device can extract a mini-batch from the buffer.
[0195] According to one embodiment, the first electronic device can extract a mini-batch from a buffer, which is a bundle of data including network state information at a first time point, a reallocation at the first time point, a compensation value based on the reallocation at the first time point, and network state information at a second time point. For example, the first electronic device can randomly extract a mini-batch from the data stored in the buffer.
[0196] In operation 1220, the first electronic device can obtain a first Q value by inputting network state information at a first time point and a reallocation at a first time point into the main deep learning neural network.
[0197] According to one embodiment, the first electronic device may input network state information at a first time point and reallocation at a first time point to a main deep learning neural network included in a deep reinforcement learning algorithm. The deep reinforcement learning algorithm may include a main deep learning neural network and a target deep learning neural network. For example, the main deep learning neural network may be a deep learning neural network for selecting an action to reallocate computing resources, and the target deep learning neural network may be a deep learning neural network for stable training separated from the main deep learning neural network.
[0198] According to one embodiment, the first electronic device can obtain a first Q value, which is the sum of reward values that can be cumulatively obtained by inputting network state information at a first time point and a reallocation at a first time point into a main deep learning neural network.
[0199] In operation 1230, the first electronic device can obtain a target Q value by inputting network state information at a second time point and a reward value based on reallocation at a first time point into the target deep learning neural network.
[0200] According to one embodiment, the first electronic device can apply network information at a second time point to a target deep learning neural network and obtain a plurality of Q values for a plurality of cells.
[0201] According to one embodiment, the first electronic device can obtain a target Q value by using the highest Q value among a plurality of Q values obtained by applying network information to a target deep learning neural network at a second time point and a reward value according to reallocation at a first time point.
[0202] In operation 1240, the first electronic device can update the first weight of the main deep learning neural network based on the difference between the target Q value and the first Q value.
[0203] According to one embodiment, the first electronic device can calculate a DDQN loss (double DQN loss) value by calculating the difference between the target Q value and the first Q value.
[0204] According to one embodiment, the first electronic device can update the first weight of the main deep learning neural network based on the DDQN loss value. The first electronic device can update the first weight of the main deep learning neural network in a direction in which the DDQN loss value is minimized.
[0205] The weights of a deep learning neural network in the present disclosure may be values that determine how much an input value influences the result in a deep learning model comprising a main deep learning neural network and a target deep learning neural network. A deep learning model receives numerous inputs to predict an output, and in this process, it may learn which inputs are more important by assigning weights to each input. The deep learning model can learn to make accurate predictions by adjusting the weights.
[0206] In operation 1250, the first electronic device can determine whether the weight update cycle has been reached.
[0207] According to one embodiment, the weight update cycle may be a cycle for updating the second weight of the target deep learning neural network. For example, the weight update cycle may be determined based on the number of times the first weight of the main deep learning neural network has been updated. For example, if the first electronic device updates the first weight of the main deep learning neural network 100 times, the first electronic device may be determined to have reached the weight update cycle. However, the weight update cycle is not limited to the examples described above.
[0208] In operation 1260, the first electronic device can update the second weight of the target deep learning neural network to the first weight of the main deep learning neural network.
[0209] According to one embodiment, the weight update cycle may be a cycle for updating the weights of a target deep learning neural network. For example, the first electronic device may update the first weight of the main deep learning neural network according to operation 1240, but may not update the second weight of the target deep learning neural network. When the first electronic device reaches the weight update cycle according to operation 1250, it may update the second weight of the target deep learning neural network.
[0210] According to one embodiment, when the first electronic device reaches the weight update cycle, it can update the second weight of the target deep learning neural network to the first weight of the main deep learning neural network. As the first electronic device reaches the weight update cycle and updates the second weight of the target deep learning neural network to the first weight of the main deep learning neural network, the effect of increasing the stability of learning for the deep reinforcement learning algorithm and reducing excessive variability can be achieved.
[0211] In FIG. 12 according to one embodiment, the operation of the first electronic device training a deep reinforcement learning algorithm can be understood as the operation of training a DDQN algorithm. Content not described in the above example can be supplemented with an explanation based on the DDQN algorithm and can be understood based on a general DDQN algorithm.
[0212] Methods according to the claims or embodiments described in the specification of the present disclosure may be implemented in the form of hardware, software, or a combination of hardware and software.
[0213] When implemented in software, a computer-readable storage medium may be provided for storing one or more programs (software modules). One or more programs stored in the computer-readable storage medium are configured for execution by one or more processors within an electronic device. One or more programs include instructions that cause the electronic device to execute methods according to the claims or embodiments described in the specification of this disclosure.
[0214] In the present disclosure, the function or operation performed by an electronic device may be performed by one or more processors executing one or more instructions stored in memory. The function or operation of the electronic device mentioned in the present disclosure may be performed by a single processor executing one or more instructions, or by a combination of multiple processors executing one or more instructions. A processor mentioned in the present disclosure is understood to include a circuit for performing operations or controlling other components of the electronic device. For example, the one or more processors may include a central processing unit (CPU), a micro-processor unit (MPU), an application processor (AP), a communication processor (CP), a neural processing unit (NPU), a system on chip (SoC), or an integrated circuit (IC) configured to execute one or more instructions. The one or more processors may be configured to perform the operation of the electronic device described above.
[0215] In the present disclosure, a program (software module, software) may be stored in a random access memory, a non-volatile memory including flash memory, a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic disc storage device, a compact disc-ROM (CD-ROM), digital versatile discs (DVDs), or other forms of optical storage devices, or a magnetic cassette. Alternatively, it may be stored in a memory composed of some or all of these. The memory may be composed of a single storage medium or a combination of multiple storage media. The one or more instructions may be stored in a single storage medium or distributed across multiple storage media.
[0216] Additionally, the above program may be stored on an attachable storage device that can be accessed via a communication network such as the Internet, Intranet, LAN (local area network), WLAN (wide LAN), or SAN (storage area network), or a combination thereof. Such a storage device may be connected to a device performing an embodiment of the present disclosure through an external port. Additionally, a separate storage device on a communication network may be connected to a device performing an embodiment of the present disclosure.
[0217] In the specific embodiments of the present disclosure described above, the components included in the disclosure are expressed in a singular or plural form according to the specific embodiments presented. However, the singular or plural expression is selected to suit the situation presented for convenience of explanation, and the present disclosure is not limited to singular or plural components; even if a component is expressed in the plural form, it may be composed of a singular form, and even if a component is expressed in the singular form, it may be composed of a plural form.
[0218] Additionally, in the present disclosure, terms such as "part," "module," etc. may be hardware components such as a processor or circuit, and / or software components executed by hardware components such as a processor.
[0219] "Parts" and "modules" may be implemented by a program that is stored on an addressable storage medium and can be executed by a processor. For example, "parts" and "modules" may be implemented by components such as software components, object-oriented software components, class components, and task components, as well as by processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables.
[0220] The specific embodiments described in this disclosure are merely examples and do not limit the scope of this disclosure in any way. For the sake of brevity, descriptions of prior electronic configurations, control systems, software, and other functional aspects of said systems may be omitted.
[0221] Additionally, in the present disclosure, "comprising at least one of a, b, or c" may mean "comprising only a, comprising only b, comprising only c, or comprising a combination of two or more (comprising a and b, comprising b and c, comprising a and c, or comprising all of a, b, and c)."
[0222] Meanwhile, although specific embodiments have been described in the detailed description of the present disclosure, it is understood that various modifications are possible within the scope of the present disclosure. Therefore, the scope of the present disclosure should not be limited to the described embodiments, but should be defined by the claims set forth below as well as equivalents thereof.
Claims
1. A method performed by a first electronic device of a wireless communication system, An operation to obtain network status information of a plurality of cells connected to a second electronic device; An operation to obtain multiple Q values by applying the network state information to a deep reinforcement learning algorithm to improve the network performance of the plurality of cells; An operation to generate a control message instructing the re-allocation of computing resources allocated to the plurality of cells based on the plurality of Q values; The operation of transmitting the control message to the second electronic device; and A method comprising the operation of obtaining first information related to the throughput of the plurality of cells from the second electronic device.
2. In Paragraph 1, The operation of generating the above control message is, Select the maximum and minimum values from the above plurality of Q values, and A method comprising the operation of generating the control message that instructs the computing resources of the cell corresponding to the minimum value to be reallocated to the cell corresponding to the maximum value.
3. In Paragraph 2, The computing resources required for packet processing allocated to a single slot are defined as a single unit of computing resources, and A method in which the above control message is a message instructing to reallocate one computing resource unit of the cell corresponding to the minimum value to the cell corresponding to the maximum value.
4. In Paragraph 1, An operation to identify an increase in throughput resulting from the reassignment based on the first information above; and A method further comprising the operation of determining a reward value according to the reassignment based on the above increase amount.
5. In Paragraph 4, A method further comprising the operation of storing the network state information, the reallocation and the reward value in a buffer for learning the deep reinforcement learning algorithm.
6. In Paragraph 5, The operation further includes training the deep reinforcement learning algorithm by applying a mini-batch of information sampled from the buffer at each first cycle. A method comprising: the above mini-batch including network state information at a first time point, a reallocation at the first time point, a compensation value according to the reallocation at the first time point, and network state information at a second time point, which is the time point following the first time point.
7. In Paragraph 6, The above deep reinforcement learning algorithm is the DDQN (double deep Q network) algorithm, and The above deep reinforcement learning algorithm includes a main deep learning neural network and a target deep learning neural network, and A method in which the target deep learning neural network is updated according to the main deep learning neural network.
8. In Paragraph 7, The operation of training the above deep reinforcement learning algorithm is, A first Q value is obtained by applying the network state information at the first time point and the reallocation at the first time point among the network state information to the main deep learning neural network, and A target Q value is obtained by applying the network state information at the second time point and the compensation value resulting from the reallocation at the first time point among the network state information to the target deep learning neural network, and Update the first weight of the main deep learning neural network based on the DDQN loss value representing the difference between the first Q value and the target Q value, and A method comprising the operation of updating the second weight of the target deep learning neural network with the first weight of the main deep learning neural network every second period.
9. In Paragraph 1, A method comprising the above network status information including the number of terminals activated in the plurality of cells, the utilization rate of radio resources of the plurality of cells, and information regarding the state in which computing resources are allocated to the plurality of cells.
10. In Paragraph 1, The operation of obtaining the plurality of Q values by applying the network state information to the deep reinforcement learning algorithm is: Action of setting an arbitrary value; The operation of generating the control message that instructs to reallocate the computing resource arbitrarily when the above arbitrary value is smaller than the epsilon value; and A method comprising the operation of, when the arbitrary value is greater than or equal to the epsilon value, applying the network state information to the deep reinforcement learning algorithm to obtain the plurality of Q values, and generating the control message that instructs to reallocate the computing resources based on the plurality of Q values.
11. In the first electronic device of a wireless communication system, Transmitter / receiver; Memory for storing instructions; and Includes one or more processors, The above instructions are executed by the above one or more processors, and the first electronic device: Obtaining network status information of multiple cells connected to the second electronic device, and A plurality of Q values are obtained by applying the network state information to a deep reinforcement learning algorithm to improve the network performance of the plurality of cells, and Based on the plurality of Q values above, a control message is generated to instruct the re-allocation of computing resources allocated to the plurality of cells, and Transmit the control message to the second electronic device, and A first electronic device that obtains first information related to the throughput of the plurality of cells from the second electronic device.
12. In Paragraph 11, The above instructions are executed by the above one or more processors, and the first electronic device: The operation of generating the above control message is, Select the maximum and minimum values from the above plurality of Q values, and A first electronic device comprising the operation of generating the control message that instructs the computing resources of the cell corresponding to the minimum value to be reallocated to the cell corresponding to the maximum value.
13. In Paragraph 12, The computing resources required for packet processing allocated to a single slot are defined as a single computing resource unit, and A first electronic device, wherein the above control message is a message instructing to reallocate one computing resource unit of the cell corresponding to the minimum value to the cell corresponding to the maximum value.
14. In Paragraph 11, The above instructions are executed by the above one or more processors, and the first electronic device: Based on the first information above, identify the increase in throughput resulting from the reallocation, and A first electronic device that determines a reward value according to the reallocation based on the above increase amount.
15. In Paragraph 14, The above instructions are executed by the above one or more processors, and the first electronic device: A first electronic device that stores the network state information, the reallocation, and the reward value in a buffer for learning the deep reinforcement learning algorithm.