Large model-based policy lightweighting method for real-time network scheduling
The two-stage model inference structure addresses the limitations of large-scale models by using a parent model to learn optimal policies and update a child model in real-time, enhancing network scheduling efficiency and adaptability in dynamic environments.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
- Filing Date
- 2026-01-14
- Publication Date
- 2026-07-23
AI Technical Summary
Existing network scheduling and resource allocation technologies face challenges in adapting to diverse environments due to high computational costs and long inference times of large-scale models, limiting their applicability in real-time changing conditions.
A two-stage model inference structure is employed, comprising a large-scale parent model and a smaller child model, where the parent model learns optimal policies for various environments and updates the child model when necessary, enabling real-time inference and adaptation to environmental changes.
The two-stage model inference structure reduces computational resources and inference time, allowing for efficient real-time network scheduling and resource allocation even in resource-constrained environments, adapting to rapid changes and ensuring high performance.
Smart Images

Figure KR2026000828_23072026_PF_FP_ABST
Abstract
Description
Large Model-Based Policy Lightweighting Method for Real-Time Network Scheduling
[0001] The present invention relates to a large-scale model-based policy lightweighting method for real-time network scheduling, and more specifically, to an optimal behavior inference method and apparatus based on a two-stage model inference structure as a large-scale model-based policy lightweighting method for real-time network scheduling. The present invention is a research result resulting from project support under Project No. RS-2024-00405128 and Project Unique Number 2710007976 of the Institute of Information & Communications Technology Planning & Evaluation (IITP) under the Ministry of Science and ICT.
[0002] URLLC in 5G and 6G communications includes new services that will transform industries through ultra-reliable / low-latency links, such as remote control of critical infrastructure, self-driving vehicles, and end-to-end ultra-precision networking. Levels of reliability and latency are essential for 6G end-to-end ultra-precision networking, smart grid control, industrial automation, robotics, and drone control and coordination. For this reason, many resource management methods are being researched in next-generation communication networks, such as 6G, to guarantee performance or required latency for ultra-reliable, low-latency, and citizen-sense data.
[0003] Existing network scheduling and resource allocation technologies have primarily relied on rule-based methods or reinforcement learning-based algorithms trained in a single environment. Rule-based methods, such as MaxWeight, are designed to satisfy specific objectives, often failing to find the optimal policy based on those objectives. While these algorithms operate efficiently in complex situations, they cannot flexibly respond to requirements that deviate from specific objectives. Reinforcement learning-based algorithms can learn optimal policies in a specific environment, but the learned policies cannot generalize to other environments, making it difficult to handle diverse situations. In particular, approaches utilizing large-scale models like Transformers face the problem of limited applicability in real-time changing environments due to high computational costs and long inference times.
[0004] The technical problem to be solved by the present invention is to provide an optimal behavior inference method based on a two-stage model inference structure.
[0005] Another technical objective of the present invention is to provide an optimal behavior inference device based on a two-stage model inference structure.
[0006] Another technical objective of the present invention is to provide a computer-readable recording medium that records a program for executing an optimal behavior inference method based on a two-stage model inference structure on a computer.
[0007] The technical problems to be solved by the present invention are not limited to the above technical problems, and other technical problems not mentioned will be clearly understood by those skilled in the art to which the present invention belongs from the description below.
[0008] A method for inferring an optimal action based on a two-stage model inference structure according to the present invention, for achieving the above technical problem, may include: a step of inputting context information for a specific environment into a first artificial intelligence model that has learned an optimal policy for each environment in advance; a step of inputting information about a current state related to the specific environment into a second artificial intelligence model that approximates an optimal policy for the context information for the specific environment; and a step of outputting an optimal action in the second artificial intelligence model based on model parameters related to the second artificial intelligence model that correspond to the input information about the current state and the optimized policy for the specific environment.
[0009] The above model parameters may include one or more parameters connecting each layer in the second artificial intelligence model. The optimal behavior inference method based on the two-stage model inference structure may further include a step of deciding to update the model parameters when the inference performance for the specific environment on a pre-prepared simulator is better than a predetermined amount than the inference performance for the optimal behavior output from the second artificial intelligence model.
[0010] The above-described optimal behavior inference method based on a two-stage model inference structure may further include the step of applying context information of the changed environment to the first artificial intelligence model and outputting model parameters of the second artificial intelligence model corresponding to the context information of the changed environment when some environment changes or when some environment changes more than a predetermined amount in each environment.
[0011] The optimal action output above may include network scheduling information or user scheduling information requiring real-time control. The first artificial intelligence model is a meta-reinforcement learning-based transformer model, and the second artificial intelligence model may be a model that is smaller in size than the first artificial intelligence model or has an inference time shorter than the first artificial intelligence model. The optimal action inference method based on the two-stage model inference structure may further include a step in which the first artificial intelligence model learns one or more parameters connecting each context information for each environment and each layer in the second artificial intelligence model through supervised learning.
[0012] The above-described optimal behavior inference method based on a two-stage model inference structure may further include the step of updating the model parameters of the inferred second artificial intelligence model corresponding to the context information of the changed environment by applying them to the second artificial intelligence model.
[0013] An optimal behavior inference device based on a two-stage model inference structure according to the present invention, for achieving other technical objectives described above, may include: a memory; and a processor that learns an optimal policy for each environment in advance and inputs context information for a specific environment into a first artificial intelligence model stored in the memory, controls the input of information about a current state related to the specific environment into a second artificial intelligence model stored in the memory as a model that approximates the optimal policy for the context information for the specific environment, and controls the output of an optimal behavior in the second artificial intelligence model based on model parameters related to the second artificial intelligence model that correspond to the input information about the current state and the optimized policy for the specific environment.
[0014] The processor may decide to update the model parameters when it is determined that the inference performance for the specific environment on a pre-prepared simulator is better than the inference performance for the optimal action output from the second artificial intelligence model.
[0015] The processor can output model parameters of the second artificial intelligence model corresponding to the context information of the changed environment by applying context information of the changed environment to the first artificial intelligence model when a part of the environment changes or a part of the environment changes by more than a predetermined amount in each of the above environments.
[0016] The optimal behavior inference technique based on a two-stage model inference structure according to the present invention can solve the problem that traditional DRL-based approaches have difficulty changing the model dynamically in response to changes in the driving environment.
[0017] In addition, the optimal behavior inference technique based on a two-stage model inference structure according to the present invention can solve problems such as the significant computational resources, time required, and power consumption of existing large models in a single inference.
[0018] The present invention provides a resource-efficient real-time inference system capable of adapting to various goals and environments by generating a small-scale child model based on an optimal policy learned by a large-scale parent model for each environment. Through this, the system can infer an optimal policy in real time even within a short control cycle, rapidly adapt to environmental changes, and demonstrate high efficiency even in resource-constrained environments.
[0019] In addition, the present invention can overcome problems such as the conventional lack of goal adaptability, limitations of real-time processing, and inefficiency in resource-constrained environments, and can provide excellent performance in network scheduling, wireless resource allocation, and other real-time resource management systems.
[0020] In addition, solving these problems enables accurate real-time control, such as when the scheduling control cycle is very short, the channel state changes rapidly, or scheduling must be performed within 1ms to guarantee performance.
[0021] The effects obtainable from the present invention are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art from the description below.
[0022] The accompanying drawings, which are included as part of the detailed description to aid in understanding the present invention, provide embodiments of the present invention and explain the technical concept of the present invention together with the detailed description.
[0023] Figure 1 is a diagram illustrating the overall system structure of 5G NR.
[0024] Figure 2 is a diagram illustrating the logical architecture in an Open RAN (O-RAN) system.
[0025] Figure 3 is a diagram illustrating a protocol stack-user plane for 5G, 6G, etc.
[0026] Figure 4 is a diagram illustrating the RIC architecture in O-RAN.
[0027] Figure 5 is a diagram illustrating the arrangement of O-RNA and applications (Apps).
[0028] Figure 6 is a diagram illustrating an example of network scheduling.
[0029] Figure 7 is a diagram illustrating context-based meta-reinforcement learning.
[0030] Figure 8 is a table showing an example of a new architecture that can reflect the scheduler's decision every 1ms in an O-RAN architecture.
[0031] FIG. 9 is a diagram illustrating a comparison between a conventional learning-based inference structure (20) and a two-stage model inference structure (10) for real-time control according to the present invention.
[0032] FIG. 10 is a diagram illustrating network scheduling as an example of a two-stage model inference structure according to the present invention.
[0033] FIG. 11 is a drawing for explaining the details of learning a parent model (11) according to the present invention.
[0034] FIG. 12 is a block diagram illustrating the function of an optimal behavior inference device (100) based on a two-stage model inference structure according to the present invention.
[0035] Hereinafter, preferred embodiments according to the present invention will be described in detail with reference to the accompanying drawings. The detailed description disclosed below in conjunction with the accompanying drawings is intended to describe exemplary embodiments of the present invention and is not intended to represent the only embodiment in which the present invention may be practiced. The following detailed description includes specific details to provide a complete understanding of the present invention. However, those skilled in the art know that the present invention may be practiced without such specific details. For example, for convenience of explanation, the following detailed description is described specifically assuming cases such as a mobile communication system 3GPP 5G NR, a next-generation communication system 6G, and an Open-Radio Access Network (O-RAN) system; however, it is applicable to any other mobile communication system except for specific details specific to 3GPP 5G NR, a next-generation communication system 6G, and Open RAN (O-RAN).
[0036] In some cases, to avoid obscuring the concept of the present invention, known structures and devices may be omitted or illustrated in the form of block diagrams focusing on the core functions of each structure and device. Additionally, throughout this specification, the same components are described using the same reference numerals.
[0037] The present invention is capable of various modifications and may have various embodiments, and specific embodiments are illustrated in the drawings and described in detail. However, this is not intended to limit the invention to specific embodiments, and it should be understood that the invention includes all modifications, equivalents, and substitutions that fall within the spirit and scope of the invention.
[0038] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. On the other hand, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.
[0039] Terms such as "first," "second," etc., may be used to describe various components, but said components should not be limited by said terms. These terms are used solely for the purpose of distinguishing one component from another.
[0040] The terms used herein are merely for describing specific embodiments and are not intended to limit the invention. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this specification, terms such as “comprising” or “having” are intended to indicate the presence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0041] Furthermore, the components of the embodiments described with reference to each drawing are not limited to the respective embodiments and may be implemented to be included in other embodiments within the scope of maintaining the technical spirit of the present invention. It is also obvious that multiple embodiments may be re-implemented as a single embodiment that integrates multiple embodiments, even if a separate description is omitted.
[0042] Additionally, terms such as “…part,” “…unit,” “…module,” and “…device” described in the specification refer to a unit that processes at least one function or operation, and this may be implemented in hardware, software, or a combination of hardware and software.
[0043] In addition, for the following description, the terminal can receive information from the base station / network via the downlink, and the terminal can also transmit information to the base station / network via the uplink. The information transmitted or received by the terminal includes data and various control information, and various physical channels exist depending on the type and purpose of the information transmitted or received by the terminal.
[0044] Embodiments of the present invention may be supported by disclosed standard documents for wireless access systems such as 5G NR, 6G, and Open RAN. That is, steps or parts among the embodiments of the present invention that are not described in order to clearly reveal the technical concept of the present invention may be supported by said documents. In addition, all terms disclosed in this document may be explained by said standard documents.
[0045] The present invention can be applied not only to 6G but also to 3GPP 5G and Open RAN, which are currently undergoing standardization. First, we will briefly describe the details regarding 5G NR and Open RAN systems to which the method proposed in this specification can be applied, and then describe the present invention.
[0046] Figure 1 is a diagram illustrating the overall system structure of 5G NR.
[0047] Referring to FIG. 1, the NG-RAN consists of base stations (gNBs) that provide control plane (RRC) protocol terminations for the NG-RA user plane (new AS sublayer / PDCP / RLC / MAC / PHY) and User Equipment (UE). The gNBs are interconnected via the Xn interface. The gNBs are also connected to the NGC via the NG interface. The gNBs are connected to the Access and Mobility Management Function (AMF) via the N2 interface and to the User Plane Function (UPF) via the N3 interface.
[0048] NG-C represents the Control Plane interface used for the NG2 reference point between the new RAN and the NGC, and NG-U represents the User Plane interface used for the NG3 reference point between the new RAN and the NGC. Non-standalone NR has a deployment configuration where the gNB requires an LTE eNB as an anchor for the Control Plane connection to the EPC, or requires an eLTE eNB as an anchor for the Control Plane connection to the NGC. Non-standalone E-UTRA has a deployment configuration where the eLTE eNB requires a gNB as an anchor for the Control Plane connection to the NGC. The User Plane Gateway is the endpoint of the NG-U interface.
[0049] In a 5G system, the Radio Unit (RU) is a device that handles the Digital Front-End (DFE), part of the PHY layer, and digital beamforming functions. While 5G RU designs must be inherently intelligent, the primary considerations for RU design are size, weight, and power consumption. The Distributed Unit (DU) is a distributed processing unit located near the RU that executes parts of the RLC, MAC, and PHY layers. Depending on the function partitioning option, this logical node includes a subset of eNB / gNB functions, and its operation is controlled by the Centralized Unit (CU). The DU can be distributed unit software deployed on-site at Commercial Off-the-Shelf (COTS) servers, and the DU software is typically deployed near the RU on-site and can execute parts of the RLC, MAC, and PHY layers.
[0050] The CU is a centralized unit that executes the RRC and PDCP layers. A gNB consists of a CU connected to the CU via Fs-C and Fs-U interfaces for the CP and UP, respectively, and one DU. A CU with multiple DUs supports multiple gNBs. Through a partitioned architecture, 5G networks can utilize different deployments of protocol stacks between the CU and DU depending on medium availability and network design. It is a logical node that includes gNB functions such as user data transport, mobility control, RAN sharing (MORAN), positioning, and session management, with the exception of functions assigned exclusively to the DU. The CU controls the operation of multiple DUs through the midhaul interface. As the central unit, the CU primarily includes the RRC, SDAP, and PDCP protocol layers and is mainly responsible for non-real-time RRC and PDCP protocol stack functions. The CU can be deployed to the cloud or servers to support the integrated deployment of core network UPF sinking and edge computing. The CU and DU can be connected via the F1 interface. One CU can manage one or more DUs. In short, DUs are responsible for real-time Layer 1 (L1, physical layer) and lower Layer 2 (L2), which include data link layer and scheduling functions, while CUs are responsible for non-real-time upper L2 and L3 (network layer) functions.
[0051] The following is a brief description of the definitions of each term in an Open RAN (O-RAN) system to which the present invention can be applied.
[0052] A Near-RT RIC is an O-RAN Near-Real-Time RAN intelligent controller. It is a logical function that enables near-real-time control and optimization of RAN elements and resources through granular data collection and processing via the E2 interface. This may include AI / ML (Artificial Intelligence / Machine Learning) workflows, including model training, inference, and updates. An xApp is an application designed to run on a near-RT RIC. Such applications are likely to consist of one or more microservices and identify the data consumed and provided at the time of onboarding. Applications are independent of the near-RT RIC and can be provided by any third party. Refer to Table 1 for details.
[0053] The Non-RT RIC is an O-RAN non-real-time RAN intelligent controller and a logical function within the SMO that drives content delivered through the A1 interface. The Non-RT RIC framework and functions consist of the Non-RT RIC applications (rApps) defined below.
[0054] Non-RT RIC Applications (rApps): These are modular applications that provide value-added services related to RAN operations, such as A1 interface operation, by leveraging functions exposed through the R1 interface of the Non-RT RIC framework. They recommend values and actions that can be subsequently applied through the O1 / O2 interfaces and generate "enrichment information" regarding the use of other rApps. The functions of rApps within the Non-RT RIC enable non-real-time control and optimization of RAN elements and resources, as well as policy-based guidance for applications / functions of the Near-RT RIC.
[0055] The Non-RT RIC framework is an internal SMO function that logically terminates the A1 interface to the Near-RT RIC and exposes the set of internal SMO services required for runtime processing to the rApp via the R1 interface. The Non-RT RIC framework functions within the Non-RT RIC provide AI / ML workflows, including model training, inference, and updates, required by the rApp.
[0056] NMS is a network management system for O-RUs designated to support legacy Open Fronthaul M-Plane deployments. O-Cloud is a cloud computing platform consisting of a collection of physical infrastructure nodes that meet O-RAN requirements for hosting software components (e.g., operating systems, virtual machine monitors, container runtimes, etc.) that support related O-RAN functions (Near-RT RIC, O-CU-CP, O-CU-UP, and O-DU, etc.).
[0057] O-CU-CP (O-RAN Central Unit-Control Plane) is a logical node that hosts the RRC and control plane portions of the PDCP protocol. O-CU-UP (O-RAN Central Unit-User Plane) is a logical node that hosts the user plane portions of the PDCP and SDAP protocols.
[0058] O-DU (O-RAN Distributed Unit) is a logical node that hosts the RLC / MAC / High-PHY layer based on lower-layer function partitioning. O-eNB is an eNB or ng-eNB that supports the E2 interface. O-RU (O-RAN Radio Unit) is a logical node that hosts the Low-PHY layer and RF processing based on lower-layer function partitioning. This is similar to 3GPP's "TRP" or "RRH," but is more specific in that it includes the Low-PHY layer (FFT / iFFT, PRACH extraction).
[0059] Figure 2 is a diagram illustrating the logical architecture in an Open RAN (O-RAN) system.
[0060] Referring to Fig. 2, within the logical architecture of the O-RAN, the wireless side includes Near-RT RIC, O-CU-CP, O-CU-UP, O-DU, and O-RU functions. The E2 interface connects the O-eNB to the Near-RT RIC. Although not shown in Fig. 2, the O-eNB supports O-DU and O-RU functions through the Open Fronthaul interface.
[0061] On the management side, an SMO framework including Non-RT-RIC capabilities is included. On the other hand, O-Cloud is a cloud computing platform composed of a collection of physical infrastructure nodes that meet O-RAN requirements. It includes relevant O-RAN capabilities (Near-RT RIC, O-CU-CP, O-CU-UP, and O-DU, etc.), supporting software components (e.g., operating system, virtual machine monitor, container runtime, etc.), and appropriate management and orchestration functions. The virtualization of O-RU will be investigated further in the future. As shown in Fig. 2, O-RU terminates the Open Fronthaul M-Plane interface for O-DU and SMO.
[0062] The following Table 1 is a table defining the concepts of the O-RAN architecture components described in the present invention.
[0063] 본 발명에서 설명하는 O-RAN 아키텍처 구성요소 정의Near-RT RIC: O-RAN Near-Real-Time RAN Intelligent Controller: A logical function that enables near-real-time control and optimization of RAN elements and resources via fine-grained data collection and actions over E2 interface. It may include AI / ML (Artificial Intelligence / Machine Learning) workflow including model training, inference and updates. Please refer to
[0022] for more informationNon-RT RIC: O-RAN Non-Real-Time RAN Intelligent Controller: A logical function within SMO that drives the content carried across the A1 interface. It is comprised of the Non-RT RIC Framework and the Non-RT RIC Applications (rApps) whose functions are defined below. Please refer to
[0020] for more information.Non-RT RIC Applications (rApps): Modular applications that leverage the functionality exposed via the Non-RT RIC Framework’s R1 interface to provide added value services relative to RAN operation, such as driving the A1 interface, recommending values and actions that may be subsequently applied over the O1 / O2 interface and generating "enrichment information" for the use of other rApps. The rApp functionality within the Non-RT RIC enables non-real-time control and optimization of RAN elements and resources and policy-based guidance to the applications / features in Near-RT RIC. Please refer to
[0021] for more information.Non-RT RIC Framework: That functionality internal to the SMO that logically terminates the A1 interface to the Near-RT RIC and exposes to rApps, via its R1 interface, the set of internal SMO services needed for their runtime processing.The Non-RT RIC Framework functionality within the Non-RT RIC provides AI / ML workflow including model training, inference and updates needed for rApps. Please refer to
[0021] for more information.NMS: A Network Management System for the O-RU as specified in
[0027] to support legacy Open Fronthaul M-Plane deployments (prior to version 5 of
[0027] ).O-Cloud: O-Cloud is a cloud computing platform comprising a collection of physical infrastructure nodes that meet O-RAN requirements to host the relevant O-RAN functions (such as Near-RT RIC, O-CU-CP, O-CU-UP, and O-DU), the supporting software components (such as Operating System, Virtual Machine Monitor, Container Runtime, etc.) and the appropriate management and orchestration functions. Please refer to
[0028] for more information.O-CU-CP: O-RAN Central Unit - Control Plane: a logical node hosting the RRC and the control plane part of the PDCP protocol. Please refer to Section 4.3.3 for more information.O-CU-UP: O-RAN Central Unit - User Plane: a logical node hosting the user plane part of the PDCP protocol and the SDAP protocol. Please refer to Section 4.3.4 for more information.O-DU: O-RAN Distributed Unit: a logical node hosting RLC / MAC / High-PHY layers based on a lower layer functional split. Please refer to Section 4.3.5 for more information.O-eNB: An eNB [6] or ng-eNB [8] that supports E2 interface. Please refer to Section 4.3.7 for more information.O-RU: O-RAN Radio Unit: a logical node hosting Low-PHY layer and RF processing based on a lower layer functional split. This is similar to 3GPP's "TRP" or "RRH" but more specific in including the Low-PHY layer (FFT / iFFT, PRACH extraction). Please refer to Section 4.3.6 for more information.O1: Interface between SMO framework as specified in Section 4.3.1 and O-RAN managed elements, for operation and management, by which FCAPS management, PNF (Physical Network Function) software management, File management shall be achieved.O2: Interface between SMO framework as specified in Section 4.3.1 and the O-Cloud for supporting O-RAN virtual network functions. Please refer to
[0028] for more information.Open FH M-Plane: Management interface controlling the O-RU, generally driven from the O-DU but in the case of the hybrid topology also driven from the SMO. Please refer to
[0027] for more details.SMO: A Service Management and Orchestration system as described in Section 4.3.1.xApp: An application designed to run on the near-RT RIC. Such an application is likely to consist of one or more microservices and at the point of on-boarding will identify which data it consumes and which data it provides. The application is independent of the near-RT RIC and may be provided by any third party. The E2 enables a direct association between the xApp and the RAN functionality
[0022] .dApps: Distributed applications for real-time inference and control in O-RAN.
[0064] 도 3은 5G, 6G 등의 프로토콜 스택-사용자 평면을 예시한 도면이다.
[0065] Referring to Figure 3, ultra-low latency in lower layers is required to achieve the goals of URLLC. That is, to provide URLLC services, ultra-low latency is required in lower layers (layers 2 / 3) such as IP / SDAP / PDCP / RLC / MAC layers (i.e., at packet-level low latency). As such, services where latency is critical in 5G systems, etc., must be delivered within the time desired by the ADU.
[0066] Figure 4 is a diagram illustrating the RIC architecture in O-RAN.
[0067] It is necessary to satisfy the requirement that data be transmitted within the total latency for each application service. Therefore, it is necessary to manage both the network domain and the computing domain simultaneously. However, as illustrated in Figure 4, the RIC architecture in O-RAN focuses on network control, so separate definitions are required for controlling edge / commercial servers. When servers or computing nodes are used as Virtual Network Functions (VNFs), details regarding how computing nodes (units) are deployed and how interactions are performed also need to be defined. An architecture capable of simultaneously managing and controlling both the network domain and the computing domain is also required. O-RAN components can be deployed to COTS clouds using Docker. Depending on latency requirements, SMOs can be deployed in regional clouds / regional data centers, O-CUs in regional or edge devices (servers / cloud / edge data centers), and O-DUs in cell sites.
[0068] Figure 5 is a diagram illustrating the arrangement of O-RNA and applications (Apps).
[0069] Referring to Figure 5, RICs in O-RAN require much stricter latency constraints. Non-Real Time RICs require units in seconds, Near-RT RICs require units in the 10 to 1000ms range, and Real-Time RICs require units within 1ms. The distributed architecture of O-RAN forms a multi-layered computing constraint. A 1:N mapping between upper and lower layers can be established from SMO-CU-DU-RU. Large Language Models (LMs) and Transformers can be employed in RAN control. In RAN control, Transformer-based neural networks are used to solve coupling problems (Based on Branch & Bound, Based on Graph analysis), while LLMs can be used to optimize network performance (e.g., Inter-based networking, Time-series data, Sequence data).
[0070] Figure 6 is a diagram illustrating an example of network scheduling.
[0071] Referring to Figure 6, while scheduling is crucial for network resource management, finding optimal solutions for most of the goals required by various network services is NP-hard (referring to problems that cannot be solved in polynomial time, but can be solved using exponential equations instead of polynomial equations). Examples of NP-hard problems include achieving maximum throughput / minimum queue delay and achieving minimum deadline violations. Traditional rule-based scheduling algorithms (e.g., maxweight) rely mostly on heuristics and fail to achieve optimal performance. Deep Reinforcement Learning (DRL) is suitable for solving these NP-hard problems and is gaining popularity in various network control applications; however, traditional DRL-based approaches have the problem of being difficult to adapt to dynamic network environments because a new model is required whenever the operating environment changes. To address these issues with traditional DRL-based approaches, reinforcement learning-based approaches can also be considered.
[0072] Reinforcement learning is a type of machine learning in which an agent learns actions to maximize rewards while interacting with an environment. The agent refers to a machine learning model as the entity performing the learning. The agent selects an action from a specific state within the environment. A state is information representing the current situation of the environment, and the agent determines its next action based on this state. An action is a behavior that the agent can choose based on the state, and each action changes the state of the environment. The environment is the world in which the agent interacts; it changes its state according to the agent's actions and provides rewards in return. A policy is a strategy that determines which action an agent will take in a given state. Policies can be defined probabilistically or deterministically. A reward is the feedback obtained as a result of the agent performing a specific action.
[0073] The Q-function is a function that derives the value of a given action from a given state. It provides an estimate of the total reward expected when a specific action is selected in a specific state. The Q-function is used by an agent to compare the value of each action. The reinforcement learning process consists of: 1) an initialization process in which the agent observes the state within the environment and recognizes the initial state; 2) an action selection process in which the agent selects an action based on the current state and the agent's policy; 3) an environment response process in which the environment transitions to a new state and provides a reward based on the agent's action; 4) a learning process in which the agent evaluates how good the chosen action was based on the received reward and updates the policy or value function; and 5) an iterative process in which the agent repeats this process to learn better actions and form a policy that obtains the maximum reward in the long term.
[0074] Figure 7 is a diagram illustrating context-based meta-reinforcement learning.
[0075] Referring to Fig. 7, the goal of meta-reinforcement learning is to learn initialization parameters or policies that can rapidly adapt to new environments or tasks. An example of reinforcement learning is context-based meta-reinforcement learning (context-based meta-RL), which can represent optimal policies for various environments as a single model and thus serve as a solution to traditional DRL-based approaches. Context-based meta-reinforcement learning has a structure that takes information about the current state as input and outputs the optimal action for a single environment. For example, context-based meta-reinforcement learning can be a solution for deriving optimal scheduling policies in dynamic network environments. However, since this context-based meta-reinforcement learning uses a general feedforward network, a problem arises where the structure of the artificial intelligence model must change because the state and action dimensions of reinforcement learning change when the number of queues in the network fluctuates. In other words, the input and output sizes of the DNN must change dynamically depending on the number of queues, but traditional DNNs have limitations in dynamically changing these dimensions.
[0076] The Transformer is a deep learning model architecture designed to solve Natural Language Processing (NLP) and other data sequence problems, with its core idea being the Self-Attention mechanism. Large models generally refer to deep learning models with billions to hundreds of billions of parameters. These models are characterized by being trained on massive datasets, possessing high performance across various tasks, being pre-trained on large data and fine-tuned for specific tasks, and requiring massive computational resources during the training and inference processes. Most large models are designed based on the Transformer architecture. In particular, the successful development of large models in the field of Natural Language Processing (NLP) relies heavily on the performance and scalability of the Transformer. Because the Transformer (base) model takes a sequence of arbitrary length as input and learns the attention levels between tokens, it is widely used for problems such as language models where the length of the input sequence varies. Problems that traditional DRL techniques cannot solve can be addressed by training Transformer models based on meta-reinforcement learning. Transformers divide inputs into tokens to learn the relationships and importance levels between tokens. As one of the most expressive model architectures currently available, it is not constrained by input dimensions. As such, Transformer models can be used for real-time control (e.g., network control).
[0077] Figure 8 is a table showing an example of a new architecture that can reflect the scheduler's decision every 1ms in an O-RAN architecture.
[0078] Approximating knowledge of optimal policies for general environments with a single model using meta-reinforcement learning algorithms requires a very large model. For example, in network scheduling scenarios, scheduling control cycles are often very short, and particularly in wireless networks, channel states change rapidly; therefore, scheduling must be completed within 1ms to guarantee performance. However, such large-scale models require significant computational resources and time for even a single inference. Real-time control is a critical aspect of network control, but this necessitates larger AI models; consequently, this leads to longer computation times for a single action and high power consumption. Below, we propose a solution to address these issues regarding real-time control.
[0079] In the conventional O-RAN architecture of Fig. 2, the learning-based scheduler is deployed in the Near-Real Time RIC, whereas in the new O-RAN architecture illustrated in Fig. 8, a Realtime Edge RIC is located at the edge, μApps for policy execution for RAN control are deployed, and a real-time metric monitoring platform is deployed. Other details regarding the deployment in Fig. 8 are not explained in detail, and interpretation is based on the details illustrated in Fig. 8.
[0080] Below, we intend to describe a large-scale model-based policy lightweighting method for real-time control. In this invention, we propose a two-stage model inference structure, specifically an optimal action inference method and apparatus based on a two-stage model inference structure, as a large-scale model-based policy lightweighting method for real-time network scheduling. Examples of real-time control include network scheduling and network traffic control, but are not limited to these examples and can be applied to any environment requiring real-time control.
[0081] This invention is a technology that is easily applicable to various industries by utilizing a two-stage structure consisting of a large-scale model and a small-scale model. The lightweight structure of the small model is designed based on a standardized neural network architecture, allowing for rapid application in diverse fields such as network scheduling, wireless resource allocation, cloud computing, and edge computing. Since inference is performed using the small model, inference time can be significantly reduced compared to the large model, making it useful for real-time control. In particular, when this invention is applied to resource allocation problems in wireless communication environments, it can adapt in real-time to rapidly changing channel conditions and traffic patterns, providing higher efficiency and stability compared to existing rule-based resource allocation methods. The structure of this invention is designed with compatibility with existing network equipment and software in mind, and can be applied to various environments without additional modifications during initial implementation.
[0082] FIG. 9 is a diagram illustrating a comparison between a conventional learning-based inference structure (20) and a two-stage model inference structure (10) for real-time control according to the present invention.
[0083] Referring to the left side of FIG. 9, in a conventional learning-based model inference structure (20), a large model (Large mode) (e.g., Transformer) (21) is trained to receive context information and current state information as inputs and output an optimal action (e.g., optimal scheduling action). For example, when context information and the current state are input into a large model (21) that has been trained in advance, the large model (21) outputs a scheduling action, etc. However, when using a conventional large model (21), significant computational resources and time are required for a single inference, making it unsuitable for real-time control where the scheduling control cycle is very short, the channel state changes rapidly, or scheduling must be performed within 1ms to guarantee performance.
[0084] The right side of FIG. 9 illustrates a two-stage model inference structure (10) for real-time control according to the present invention. The two-stage model inference structure (10) is a two-stage inference structure including a parent model (11) and a child model (15). Here, the parent model (11) and the child model (15) are merely examples of names for the models (artificial intelligence models) and can be named in various ways. The parent model (11) contains a large model (12) that has general information about the environment. The large model (12) is, for example, a Transformer model. The child model (15) contains a small model (16), and the small model (16) can be, for example, a simple FFN (Feed-Forward Network) model, a CNN (Convolutional Neural Network) model having a number of parameters less than a predetermined number, a Transformer model, etc. In other words, any structure is possible as long as the small model (16) is smaller than the parent model (11) by a predetermined amount, that is, if the inference time is shorter by a predetermined amount. For convenience of explanation, the small model (16) is described below as a simple FFN (Feed-Forward Network) model. The large model (12) in the parent model (11) has a structure that receives context information and outputs a small model (16) optimized only for that environment. If the parent model (11) is a meta-reinforcement learning-based transformer model, the fact that the parent model (11) outputs a child model (15) may mean that it outputs parameters (e.g., weights or initialization parameters) to be applied to the algorithm of the child model (15) and applies them to the algorithm of the child model (15). As shown on the right side of FIG. 9, the output of the large model (12) is input into the small model (16) within the child model (15), and the input items are various parameters used for inference by the small model (16).The large model (12) performs inference only when necessary (e.g., when the environment changes or the environment changes beyond a certain standard), and the small model (16) within the child model (15) performs inference at every scheduling cycle (e.g., 1ms) to perform real-time scheduling according to the optimal policy in a power-efficient manner.
[0085] FIG. 10 is a diagram illustrating network scheduling as an example of a two-stage model inference structure according to the present invention, and FIG. 11 is a diagram illustrating matters regarding the learning of a parent model (11) according to the present invention.
[0086] Referring to FIG. 10, the process consists of data collection, model training, real-time inference, and update, centered around a two-stage model inference structure system composed of a parent model (11) and a child model (15). First, data for training the parent model (11) is collected as follows. A reinforcement learning-based algorithm (e.g., DQN) is utilized to learn the optimal policy for various environments. The optimal policy for each of multiple environments can be trained using a reinforcement learning (e.g., meta-reinforcement learning)-based algorithm (e.g., DQN (Deep Q Network)). The learned optimal policy is approximated by a small-sized model (e.g., Forward Neural Network (FNN)), and after training is completed, context information of the corresponding environment (e.g., time-series data of the state) and parameter information for each layer of the FNN model are stored in a buffer (40). The data collected in this way includes learning experiences in various environments and is used to train the parent model (11).
[0087] The model (30) is trained in an offline environment and collects and trains data in a given environment (e.g., e). The model (30) is trained based on input state(s) and behavior data (a). After training is complete, the model (30) can store context information of the environment (e.g., time series of the state) and information of the optimal policy model (e.g., parameters of each layer) in a buffer (40). Once data collection is complete, the parent model (11) is designed based on a transformer model and is trained through supervised learning to receive context information as input and output the parameters of the optimal policy model. During the training process, context information is used as input, and the parameters of the optimal policy model are used as labels. In this supervised learning, the input is context information, and the labels, which represent the correct answer or target value for the input context information, are model parameters (or artificial intelligence model parameters) (e.g., parameters of each layer). Data can be stored and updated using the buffer (40), thereby increasing training efficiency. The parent model (11) learns the ability to generalize environmental characteristics through data collected from various environments and to generate an optimal policy for a new environment.
[0088] The parent model (11) is learned using a supervised learning method, for example, the parameters being learned may be the input state in the corresponding environment and the corresponding policy. The parent model (11) learns parameters such as weights and bias values that connect each layer (or each layer unit) in the child model (15) which is optimized only for that environment. In this way, the parent model (11) learns optimal policies for various environments, and collects model information (e.g., parameters of each layer) and environment information regarding the optimal policies, so that the parent model (11) can be learned using a supervised learning method.
[0089] When executed, the child model (15) receives the current state as input and performs optimal scheduling in real time. The parent model (11) receives context information of the environment as input only when necessary (e.g., when a specific event occurs) to infer and update the parameters of the child model (15). The necessity of updating the child model (15) is determined by a critic. The critic implements a queue system identical to the current actual environment in the form of a simulator and evaluates the scheduling performance within the simulator using a basic scheduling algorithm (e.g., MaxWeight). If the performance on the simulator, which corresponds to the occurrence of the specific event mentioned above, is superior to the performance provided by the actual network by a certain level or more, the critic calls the parent model (11) and requests that the parameters of the child model (15) be updated.
[0090] Referring to FIG. 11, the Simple FFN model is a forward neural network, which is the most basic and simple artificial neural network model. The Simple FFN has an input layer that receives data, hidden layers that process data and extract features, and an output layer that outputs the final result. The nodes (neurons) of each layer are connected to all nodes (layers) of the previous layer, and weight and bias values from these connections are used. It is structured so that the input is fed into one or more hidden layers and then the output is produced. There is no recurrence or feedback, and data flows only in one direction (forward). At each node, data is processed using a linear transformation (the product of weights and input values) and an activation function (non-linear transformation). Examples of activation functions include ReLU. Due to the simplicity of the structure, there are no complex mechanisms (e.g., recurrence, convolution, attention, etc.). Furthermore, the Simple FFN operates effectively with a small model size using few parameters. The child model (15) according to the present invention can be implemented as a small model using the Simple FFN. Simple FFNs can be used as a standard when designing small models because they have few parameters and fast training. Simple FFN models provide interpretability and computational efficiency through a simple model structure.
[0091] Referring to FIGS. 10 and 11, as described above, the parent model (11) can update the child model (15) whenever necessary during execution. The parent model (11) serves to transmit optimized parameters (17, 18, 19) to the child model (15). The child model (15) operates in a real-time (Runtime) environment and performs decision-making based on state and action using the parameters set by the parent model (11). The child model (15) performs scheduling decisions based on the parameters (17, 18, 19) received from the parent model (11). The child model (15) can compare and optimize performance through Q-function calculation in a real-time environment. In this way, the child model (15) can perform optimal scheduling with the current state as input. The parent model (11) can infer parameters by receiving context information of the environment as input only when a specific event occurs or when necessary, update the parameters of the child model (15) with the inferred parameters, and pass the updated parameters (17, 18, 19) to the child model (15) to apply them.
[0092] Referring to FIG. 11, a system similar to the current real environment is implemented in the form of a simulator, and scheduling is performed within the simulator using a basic scheduling algorithm (e.g., non-learning-based maxweight). When the performance on the simulator is better than the performance on the actual network of the child model (15) by a certain amount, the child model (15) needs to be updated. The parent model (11) learns the relationships between input states by utilizing a transformer-based attention mechanism and generates an optimal policy suitable for the given environment. This overcomes the inefficiency of existing reinforcement learning-based technologies requiring separate learning in individual environments and enables the learning of generalized policies in various environments. The parent model (11) receives context information of the environment as input and infers the parameters of the child model (15) based on meta-information conditioned on the environment. Additionally, it evaluates the performance of the child model (15) as needed and continuously provides an optimal policy by updating it if the existing performance falls below a specific standard.
[0093] The child model (15) is designed with a lightweight structure based on meta-information generated by the parent model (11), enabling real-time inference within a short control cycle through a fixed input and output structure. This overcomes the disadvantage of existing large-scale models requiring high computational resources in resource-constrained environments and provides efficiency suitable for time-sensitive tasks such as real-time scheduling and wireless resource allocation. Additionally, by utilizing a critic to compare the performance of the child model (15) with the basic scheduling algorithm on the simulator, the child model (15) is enabled to respond quickly to environmental changes and maintain optimal performance. In other words, the child model (15) is updated according to environmental changes to efficiently reconfigure the child model (15) and continuously optimize network performance.
[0094] Referring to FIG. 11, the parent model (11) dynamically identifies the temporal relationship of input data through the attention mechanism of the Transformer and generates an optimal policy suitable for the given environment. The attention mechanism precisely analyzes the importance of each input state to derive the optimal parameters required for creating a small-scale model. This structure is designed based on the concept of meta-reinforcement learning, specifically context-based meta-reinforcement learning. Context-based meta-reinforcement learning uses context information that reflects the characteristics of the environment to effectively identify a given environment. In the present invention, time-series data of the state is utilized as context to model the dynamic characteristics of the environment, and the child model (15) can quickly and efficiently approximate an environment-optimized policy based on the context information learned by the parent model (11) in various environments. This context information captures temporal changes in the environment, helping the child model (15) to quickly adapt to changes in the environment and perform appropriate actions in real time.
[0095] The child model (15) has a fixed input and output structure and enables lightweight real-time inference through a simple structure such as a Feed-Forward Neural Network. When a specific environment changes or changes beyond a certain threshold, the parent model (11) regenerates the child model (15) based on new context information to ensure high adaptability and stability.
[0096] As illustrated in FIG. 11, the parent model (11) learns the parameter values of the model when a small model (16) of the target child model (15) is given. The parent model (11) is learned to infer weight values connecting the units of each layer of the neural network for a specific or corresponding environment.
[0097] The present invention combines the learning performance of a large model (12) and the inference efficiency of a small model (16) through a parent model-child model two-stage structure that did not exist in the past. This satisfies the requirements for real-time processing even in resource-constrained environments and overcomes the limitations of environmental adaptability and real-time inference of existing technologies. By combining the concepts of Transformer and context-based meta-reinforcement learning, the present invention can provide innovative performance and practicality in various application fields such as network scheduling and wireless resource allocation.
[0098] In this way, the present invention proposes a two-stage inference structure in which a large model (12) outputs a small model (16) that approximates an optimal policy conditioned to a given environment to perform fast inference. In the present invention, the parent model (11) learns to infer parameters of a different model other than the optimal action. The parent model (11) collects data such as context information and policy model parameters for learning.
[0099] In addition, a real-time network scheduling system can be implemented by utilizing a model based on a two-stage inference structure according to the present invention. A transformer-based network scheduler structure trained with meta-reinforcement learning capable of accommodating optimal policy information for a changing environment is implemented. An evaluator can determine the update time of the child model (15) by comparing the performance of the learning-based policy based on the current child model (15) with the policy based on the basic algorithm.
[0100] As described above, the present invention proposes a two-stage structure system that enables real-time inference by utilizing a parent model (11) to learn an optimal policy in various environments and generating a lightweight child model (15) based thereon. Existing technologies operate in a single environment or rely on a fixed model structure, making it difficult to adapt to dynamic environmental changes. However, the present invention solves this problem by having the parent model (11) perform meta-learning for each environment, and based on this, the child model (12) infers a policy optimized for the environment in real time. In particular, the present invention overcomes the high computational load and latency of transformer-based models, enabling fast operation even in resource-constrained environments. Furthermore, through various network simulations, high efficiency and adaptability have been experimentally verified even with a control cycle of less than 1ms.
[0101] This invention overcomes the limitations of environmental adaptability and real-time inference of existing technologies by combining the powerful learning performance of large-scale models with the real-time inference efficiency of small-scale models. Through this, it provides high performance and adaptability in various application fields such as network scheduling, wireless resource allocation, cloud, and edge computing, and ensures practical applicability even in resource-constrained environments.
[0102] FIG. 12 is a block diagram illustrating the function of an optimal behavior inference device (100) based on a two-stage model inference structure according to the present invention.
[0103] Referring to FIG. 12, the optimal behavior inference device (100) based on a two-stage model inference structure may include a memory (110) and a processor (120).
[0104] The algorithm itself of a parent model (11) including a large model (12) (e.g., a transformer model) and a child model (15) including a small model (16) (e.g., a simple FFN model) may be implemented as program code, and this program code is loaded into the memory (110) of a two-stage model inference structure-based optimal behavior inference device (100) for execution. The parent model (11) including the large model (12) and the child model (15) including the small model (16) (e.g., a simple FFN model) are stored in the memory (110) in an executable form.
[0105] The processor (120) can learn an optimal policy for each environment in advance and input context information for a specific environment into a large model (12) stored in the memory (110). The processor (120) can control the input of information regarding the current state related to the specific environment into a small model (16) stored in the memory (110) as an approximate model of the optimal policy for the context information regarding the specific environment. The processor (120) can control the output of an optimal action based on the information regarding the current state input into the small model (16) and model parameters corresponding to the optimized policy for the specific environment. Here, the model parameters may be model parameters related to the small model (16) and may be one or more parameters connecting each layer in the small model (16).
[0106] The processor (110) may decide to update the model parameters when it is determined that the inference performance for the specific environment on a pre-prepared simulator is better than the inference performance for the optimal action output from the small model (16). The processor (110) may control the output of the model parameters of the small model (16) corresponding to the context information of the changed environment by applying the context information of the changed environment to the large model (12) when some of the various environments change or when some of the environments change more than a predetermined amount.
[0107] The optimal behavior inference technique based on a two-stage model inference structure according to the present invention, as described above, solves the problem of difficulty in dynamically changing the model due to changes in the operating environment associated with traditional DRL-based approaches, as well as the problem of existing large models requiring significant computational resources, time, and power consumption for a single inference. This enables accurate real-time control, such as when the scheduling control cycle is very short, the channel state changes rapidly, or scheduling must be performed within 1ms to guarantee performance.
[0108] The present invention can be widely utilized in various industrial and technological fields. In particular, in the fields of network scheduling and resource management, it can be applied to next-generation communication systems such as 5G / 6G networks to efficiently perform real-time traffic management and resource allocation. In wireless communication environments, it is used for resource allocation considering frequency bands, transmission power, and channel conditions, and provides stable performance by adapting to rapidly changing network conditions. Furthermore, in edge computing and cloud environments, it maximizes data processing efficiency through resource optimization, and enables lightweight real-time inference of child models (15) even in edge environments where resources are limited, making it suitable for low-latency applications.
[0109] In particular, the present invention is suitable for applications requiring real-time resource management and decision-making in dynamically changing environments, and can be utilized in various industries such as network management, mobile communication, IoT platforms, and smart factories.
[0110] In IoT systems, the overall performance of the network can be optimized through real-time data processing and resource allocation in environments such as smart factories and smart cities. In the fields of autonomous driving and vehicle-to-vehicle networks (V2X), it can be utilized for path planning, communication resource management, and data transmission optimization for autonomous vehicles, and can respond quickly to changes in traffic conditions through the rapid inference capabilities of the child model (15). Furthermore, the present invention supports task scheduling and resource optimization in data centers and distributed computing environments to increase operational efficiency and enables real-time workload distribution in cloud infrastructure. Additionally, the present invention enables stable data transmission and real-time resource allocation in military and aviation communication networks, and provides high adaptability and efficiency in environments where response speed is critical. In conclusion, the present invention can be utilized in various industrial fields requiring real-time inference and adaptability, and provides high performance even in resource-constrained environments, thereby creating practical value in various industries such as networks, communications, IoT, and autonomous driving, and is very useful for industrial applications.
[0111] The device described above may be implemented as a hardware component, a software component, and / or a combination of a hardware component and a software component. For example, the device and components described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may perform an operating system (OS) and one or more software applications performed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. In addition, other processing configurations, such as parallel processors, are also possible.
[0112] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or command the processing unit independently or collectively. Software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computing devices and may be stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media.
[0113] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either alone or in combination. The program instructions recorded on the medium may be those specifically designed and configured for the embodiment, or they may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc. The hardware devices described above may be configured to operate as one or more software modules to perform the operation of the embodiment, and vice versa.
[0114] The embodiments described above are combinations of the components and features of the present invention in a specific form. Each component or feature should be considered optional unless otherwise explicitly stated. Each component or feature may be implemented in a form not combined with other components or features. Additionally, it is possible to construct embodiments of the present invention by combining some components and / or features. The order of operations described in the embodiments of the present invention may be changed. Some components or features of one embodiment may be included in another embodiment, or may be replaced with corresponding components or features of another embodiment. It is obvious that embodiments may be constructed by combining claims that do not have an explicit citation relationship in the claims, or that new claims may be included by amendment after filing.
[0115] It is obvious to those skilled in the art that the present invention may be embodied in other specific forms without departing from the essential features of the invention. Accordingly, the foregoing detailed description should not be interpreted restrictively in all respects but should be considered exemplary. The scope of the invention shall be determined by a reasonable interpretation of the appended claims, and all modifications within the equivalent scope of the invention are included within the scope of the invention.
[0116] Large model-based policy lightweighting methods for real-time network scheduling can be used industrially in 6G communication systems, etc.
Claims
1. A step of inputting context information for a specific environment into a first artificial intelligence model that has previously learned an optimal policy for each environment; A step of inputting information about the current state related to the specific environment into a second artificial intelligence model that approximates an optimal policy for context information regarding the specific environment; and A method for inferring an optimal action based on a two-stage model inference structure, comprising the step of outputting an optimal action based on information about the input current state in the second artificial intelligence model and model parameters related to the second artificial intelligence model corresponding to an optimized policy for the specific environment.
2. In Paragraph 1, The above model parameters include one or more parameters connecting each layer in the above second artificial intelligence model, a method for optimal behavior inference based on a two-stage model inference structure.
3. In Paragraph 1, A method for inferring an optimal action based on a two-stage model inference structure, further comprising the step of deciding to update the model parameters when the inference performance for the specific environment on a pre-prepared simulator is better than a predetermined amount than the inference performance for the optimal action output from the second artificial intelligence model.
4. In Paragraph 1, A method for optimal behavior inference based on a two-stage model inference structure, further comprising the step of applying context information of the changed environment to the first artificial intelligence model and outputting model parameters of the second artificial intelligence model corresponding to the context information of the changed environment when some environment changes or the said some environment changes by a predetermined amount in each of the above environments.
5. In Paragraph 1, The optimal action output above is an optimal action inference method based on a two-stage model inference structure, which includes network scheduling information or user scheduling information requiring real-time control.
6. In Paragraph 1, A method for optimal behavior inference based on a two-stage model inference structure, wherein the first artificial intelligence model is a meta-reinforcement learning-based transformer model, and the second artificial intelligence model is a model whose size is smaller than the first artificial intelligence model by a predetermined amount or whose inference time is shorter than the first artificial intelligence model by a predetermined amount.
7. In Paragraph 1, A method for optimal behavior inference based on a two-stage model inference structure, the first artificial intelligence model further includes the step of learning one or more parameters connecting each context information for each environment and each layer in the second artificial intelligence model through supervised learning.
8. In Paragraph 4, A method for optimal behavior inference based on a two-stage model inference structure, further comprising the step of updating the model parameters of the inferred second artificial intelligence model corresponding to the context information of the changed environment by applying them to the second artificial intelligence model.
9. Memory; and Context information for a specific environment is input into a first artificial intelligence model stored in the memory by learning an optimal policy for each environment in advance, and Controls inputting information about the current state related to the specific environment into a second artificial intelligence model stored in memory as a model that approximates the optimal policy for the context information regarding the specific environment, and An optimal action inference device based on a two-stage model inference structure, comprising a processor that controls the output of an optimal action based on information about the input current state in the second artificial intelligence model and model parameters related to the second artificial intelligence model corresponding to an optimized policy for the specific environment.
10. In Paragraph 9, An optimal behavior inference device based on a two-stage model inference structure, wherein the above model parameters include one or more parameters connecting each layer in the above second artificial intelligence model.
11. In Paragraph 9, The above processor is an optimal action inference device based on a two-stage model inference structure, which decides to update the model parameters when it is determined that the inference performance for the specific environment on a pre-prepared simulator is better than a predetermined amount than the inference performance for the optimal action output from the second artificial intelligence model.
12. In Paragraph 9, The above processor is, An optimal behavior inference device based on a two-stage model inference structure, which, when a part of the environment changes or a part of the environment changes by more than a predetermined amount in each of the above environments, applies context information of the changed environment to the first artificial intelligence model and outputs model parameters of the second artificial intelligence model corresponding to the context information of the changed environment.
13. A computer-readable recording medium storing a program for executing on a computer an optimal behavior inference method based on a two-stage model inference structure described in any one of claims 1 through 8.