Inference service in a network
By automating AI/ML model inference at the SMO/Non-RT RIC, the system addresses the complexity of O-RAN systems, enabling efficient and cost-effective deployment of high-sealability AI/ML applications.
Patent Information
- Application Number
- PCT/US2024/021275
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-05
- Filing Date
- 2024-03-25
- Publication Date
- 2025-07-10
AI Technical Summary
Existing O-RAN systems face complexity in performing AI/ML model inference due to the need for rApps to manage deployment, scaling, networking, and cloud orchestration, complicating application implementation and increasing development costs.
The system performs AI/ML model inference directly at the Service Management and Orchestration (SMO)/Non-RT RIC, automating the deployment and scaling of AI/ML models without relying on rApps or xApps, simplifying the process by handling workload distribution, cloud orchestration, and networking.
This approach simplifies the development and deployment of high-sealability AI/ML applications, reducing entry barriers and lowering costs while broadening the supply chain of advanced AI/ML solutions for O-RAN systems.
Smart Images

Figure US2024021275_10072025_PF_FP_ABST
Abstract
Description
INFERENCE SERVICE IN A NETWORK
[0001] This application claims priority from US Provisional Patent Application No. 63 / 617,880, filed with the United States Patent and Trademark Office on January 5, 2024 and entitled “O-RAN SMO / NON-RT RIC AI / ML INFERENCE SERVICE,” the disclosure of which is incorporated herein by reference in its entirety.FIELD
[0002] Apparatuses and methods consistent with example embodiments of the present disclosure relate to inference service in a telecommunication network.BACKGROUND
[0003] A radio access network (RAN) is an important component in a telecommunications system, as it connects end-user devices (or user equipment) to other parts of the network. The RAN includes a combination of various network elements (NEs) that connect end-users to a core network. Traditionally, hardware and / or software of a particular RAN is vendor specific.
[0004] Open RAN (O-RAN) technology has emerged to enable multiple vendors to provide hardware and / or software to a telecommunications system. Since different vendors are involved, the type of hardware and / or software provided may also be different. That is, different types of NEs may be provided by different vendors, and depending on the specific service, the NE could be virtualized in software form (e.g., virtual machine (VM)-based), or could be in physical hardware form (e.g., non-VM based).
[0005] To this end, O-RAN disaggregates the RAN functions into a centralized unit (CU), a distributed unit (DU), and a radio unit (RU). The CU may be a logical node for hosting RadioResource Control (RRC), Service Data Adaptation Protocol (SDAP), and / or Packet Data Convergence Protocol (PDCP) sublayers of the RAN. The DU may be a logical node hosting Radio Link Control (RLC), Media Access Control (MAC), and Physical (PHY) sublayers of the RAN. The RU may be a physical node that converts radio signals from antennas to digital signals that can be transmitted over the Front Haul to a DU. Because these entities have open protocols and interfaces between them, they can be developed by different vendors.
[0006] FIG. 1 illustrates an 0-RAN architecture in the related art. RAN functions in the O-RAN architecture may be controlled and optimized by a RAN Intelligent Controller (RIC). The RIC may be a software-defined component that implements modular applications to facilitate the multivendor operability required in the 0-RAN system, as well as to automate and optimize RAN operations. As shown in FIG. 1, the RIC may be divided into two types: a non-real-time RIC (Non- RT RIC) 120 and a near-real-time RIC (Near-RT RIC) 130.
[0007] The Non-RT RIC 120 may be the control point of a non-real-time control loop and may operate on a timescale greater than 1 second within a Service Management and Orchestration (SMO) framework 110. Its functionalities may be implemented through modular applications called rApps, and may include: providing policy based guidance and enrichment across the Al interface, which is the interface that enables communication between the Non-RT RIC and the Near-RT RIC; performing data analytics; Artificial Intelligence / Machine Learning (AI / ML) training and inference for RAN optimization; and / or recommending configuration management actions over the 01 interface, which may be the interface that connects the SMO to RAN managed elements (e.g., Near-RT RIC 130, 0-RAN Centralized Unit (O-CU) 140,150, 0-RAN Distributed Unit (O-DU) 170, etc.).
[0008] The Near-RT RIC 130 may operate on a timescale between 10 milliseconds and 1 second and may be coupled with the O-DU 170, the O-CU (disaggregated into the O-CU control plane (O-CU-CP) 140 and the O-CU user plane (O-CU-UP) 150), and an open evolved NodeB (O- eNB) 160 via the E2 interface. The Near-RT RIC 130 may use the E2 interface to control the underlying RAN elements (E2 nodes / network functions (NFs)) over a near-real-time control loop. The Near-RT RIC 130 may monitor, suspend / stop, override, and control the E2 nodes (O-CU 140,150, O-DU 170, and O-eNB 160) via policies. For example, the Near-RT RIC 130 may set policy parameters on activated functions of the E2 nodes. Further, the Near-RT RIC 130 may host xApps to implement functions such as quality of service (QoS) optimization, mobility optimization, slicing optimization, interference mitigation, load balancing, security, etc.
[0009] Here, the O-CU-CP 140 and the O-CU-UP 150 may be coupled to each other via the El interface, and may be coupled to the O-DU 170 via the Fl-c interface and Fl-u interface, respectively. Further, the O-RU 180 may be coupled to the O-DU 170 via the Open Fronthaul (OF) Control (C), User (U), Synchronization (S), and Management (M) Planes, and may be coupled to the SMO 110 via the OF M-Plane.
[0010] The two types of RICs work together to optimize the O-RAN. For example, the Non-RT RIC 120 may provide the policies, data, and Artificial Intelligence / Machine Learning (AI / ML) models enforced and used by the Near-RT RIC 130 for RAN optimization, and the Near- RT RIC 130 may return policy feedback (i.e., how the policy set by the Non-RT RIC 120 works).
[0011] As mentioned above, the Non-RT RIC 120 may be located within the SMO framework 110, which manages and orchestrates RAN elements. Specifically, the SMO 110 may manage and orchestrate what is referred to as the O-RAN Cloud (O-Cloud) 190. The O-Cloud 190may be a collection of physical RAN nodes that host the RICs, O-CUs, and O-DUs, the supporting software components (e.g., the operating systems and runtime environments), and the SMO 110 itself. In other words, the SMO 110 may manage the O-Cloud 190 from within. The 02 interface may be the interface between the SMO 110 and the O-Cloud 190 it resides in. Through the 02 interface, the SMO 110 may provide infrastructure management services (IMS) and deployment management services (DMS).SUMMARY
[0012] Example embodiments of the present disclosure automatically perform AI / ML model inference. As such, example embodiments of the present disclosure provide key architecture solution in SMO / Non-RT RIC and Near-RT RIC to perform AI / ML model inference without relying on the rApps and xApps.
[0013] According to embodiments, an apparatus is provided. The apparatus may be configured to: receive an inference request; obtain one or more ALML models based on the inference request; and perform an inference at a Service Management and Orchestration (SMO) / Non-RT RIC or Near-RT RIC using the one or more AI / ML models.
[0014] According to embodiments, a method is provided. The method may include: receiving an inference request; obtaining one or more ALML models based on the inference request; and performing an inference at a Service Management and Orchestration (SMO) / Non-RT RIC or Near-RT RIC using the one or more AI / ML models.
[0015] According to embodiments, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium may have recorded thereon instructions executable by an apparatus to cause the apparatus to perform a method including:receiving an inference request; obtaining one or more AI / ML models based on the inference request; and performing an inference at a Service Management and Orchestration (SMO) using one or more AI / ML models.
[0016] Additional aspects will be set forth in part in the description that follows and, in part, will be apparent from the description, or may be realized by practice of the presented embodiments of the disclosure.BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Features, advantages, and significance of exemplary embodiments of the disclosure will be described below with reference to the accompanying drawings, in which like signs denote like elements, and wherein:
[0018] FIG. 1 illustrates an O-RAN architecture in the related art;
[0019] FIG. 2 illustrates an exemplary embodiment of a AI / ML model inference (MLI) system, according to one or more embodiments;
[0020] FIG. 3 illustrates an example flow of data in a system for performing AI / ML model inference in a network, according to one or more embodiments;
[0021] FIG. 4 illustrates a block diagram of an example of deploying one or more AI / ML models on a node cluster, according to one or more embodiments;
[0022] FIG. 5 illustrates an example flow of data in a system for performing AI / ML model inference in a network, according to one or more embodiments;
[0023] FIG. 6 illustrates a flow diagram of an example method for performing machine learning inference, according to one or more embodiments;
[0024] FIG. 7 illustrates a flow diagram of an example method for performing an inference at the SMO, according to one or more embodiments;
[0025] FIG. 8 illustrates a diagram of an example environment in which systems and / or methods, described herein, may be implemented;DETAILED DESCRIPTION
[0026] The following detailed description of example embodiments refers to the accompanying drawings. The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit the implementations to the precise form disclosed. Modifications and variations are possible in light of the above disclosure or may be acquired from practice of the implementations. Further, one or more features or components of one embodiment may be incorporated into or combined with another embodiment (or one or more features of another embodiment). Additionally, in the flowcharts and descriptions of operations provided below, it is understood that one or more operations may be omitted, one or more operations may be added, one or more operations may be performed simultaneously (at least in part), and the order of one or more operations may be switched, as long as these modifications may not affect the resulting scope of the invention.
[0027] It will be apparent that systems and / or methods, described herein, may be implemented in different forms of hardware, software, or a combination of hardware and software. The actual specialized control hardware or software code used to implement these systems and / or methods is not limiting of the implementations. Thus, the operation and behavior of the systems and / or methods were described herein without reference to specific software code. It is understoodthat software and hardware may be designed to implement the systems and / or methods based on the description herein.
[0028] Even though particular combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features may be combined in ways not specifically recited in the claims and / or disclosed in the specification. Although each dependent claim listed below may directly depend on only one claim, the disclosure of possible implementations includes each dependent claim in combination with every other claim in the claim set.
[0029] No element, act, or instruction used herein should be construed as critical or essential unless explicitly described as such. Also, as used herein, the articles “a” and “an” are intended to include one or more items, and may be used interchangeably with “one or more.” Where only one item is intended, the term “one” or similar language is used. Also, as used herein, the terms “has,” “have,” “having,” “include,” “including,” or the like are intended to be open- ended terms. Further, the phrase “based on” is intended to mean “based, at least in part, on” unless explicitly stated otherwise. Furthermore, expressions such as “at least one of [A] and [B]”, “[A] and / or [B]”, or “at least one of [A] or [B]” are to be understood as including only A, only B, or both A and B.
[0030] The foregoing disclosure provides illustration and description but is not intended to be exhaustive or to limit the implementations to the precise form disclosed. Modifications and variations are possible in light of the above disclosure or may be acquired from practice of the implementations.
[0031] It shall be noted that, descriptions of example embodiments of the present disclosure may include terms and names defined in one or more standard organizations, such as the 3rd Generation Partnership Project (3GPP) standard organization, the European Telecommunications Standards Institute (ETSI) standard organization, the Open RAN (O-RAN) Alliance, and the like. For instance, the terms [5GC, RRC, SR, etc ], and the like, as well as the associated features and operations, are to be interpreted as consistent with those specified in one or more [3GPP, O-RAN, etc.] specifications, unless being described otherwise.
[0032] As described above, the Non-RT RIC included in the SMO framework may include functionalities related to Artificial Intelligence / Machine Learning (AI / ML) training and inference for RAN optimization. In particular, the Non-RT RIC may include AI / ML workflow services provided to the rApps, which may include training of a AI / ML model, registration of a AI / ML model, discovery of a AI / ML model, subscription to a change of a AI / ML model, registration of capability to train a AI / ML model, storage of a AI / ML model, monitoring the performance of a deployed AI / ML model, retrieving location details of a AI / ML model, requesting deployment of a AI / ML model, and the like.
[0033] Accordingly, AI / ML applications may benefit from a platform level inference service, including distributed computing, high scalability, high portability running on either CPU, GPU, or TPU, standardized interface protocol for pre / post processing, model inference, and monitoring, auto networking among AI / ML model instances, and simple and pluggable production serving.
[0034] In the related art, in order to perform the AI / ML inference, an rApp may first request the Non-RT RIC to collect data and train a AI / ML model. Once the AI / ML model is trained,the AI / ML model may be transmitted back to the rApp, where the rApp itself may perform the inference using the AI / ML model.
[0035] In this regard, the above approach for performing the inference in the related art may have the following shortcomings. Since the rApp itself performs the inference, the rApp may have to manage all relevant processes related to the deployment for the inference, including management of the scalability and distribution based on the usage (demand) of the AI / ML service, acceleration layer abstraction for the inference between the application layer and the cloud orchestration functions (e.g., O-Cloud DMS, K8S, and the like), networking, platform capability discovery, acceleration hardware access, and the like. Accordingly, the above approach complicates the application implementation which needs to take care of the distribution, deployment, scaling, networking, Accelerator Adaptation Layer (AAL), and the like, as well as complicates the cloud orchestration framework within the SMO for the AI / ML model instance deployment and scaling.
[0036] Further, while the Non-RT RIC include functionalities to support AI / ML inference, there are currently no solutions for how to perform such AI / ML inference.
[0037] Accordingly, system, methods, devices, and the like, provided in the example embodiments of the present disclosure automatically perform AI / ML model inference.
[0038] According to embodiments, in order to perform a AI / ML model inference, the system may obtain the data regarding deployment of AI / ML models used for inference and perform the inference at a Service Management and Orchestration (SMO) or Non-RT RIC using the one or more AI / ML models.
[0039] Ultimately, example embodiments of the present disclosure automatically perform AI / ML model inference, which provide key architecture solution in SMO / Non-RT RIC and Near- RT RIC to perform AI / ML model inference without relying on the rApps and xApps, as shown inFIG. 3 and FIG. 5.
[0040] It is contemplated that features, advantages, and significances of example embodiments described hereinabove are merely a portion of the present disclosure, and are not intended to be exhaustive or to limit the scope of the present disclosure.
[0041] Further descriptions of the features, components, configuration, operations, and implementations of the threshold tuning system of the present disclosure, according to one or more embodiments, are provided in the following.Example System Architecture
[0042] FIG. 2 illustrates an exemplary embodiment of a AI / ML model inference (MLI) system 200, according to one or more embodiments. As shown in FIG. 2, the MLI system 200 may include a processor 210, a memory 220, a storage component 230, an input component 240, an output component 250, a communication interface 260, and a bus 270.
[0043] According to embodiments, the MLI system 200 may include an apparatus, a system, a platform, a module, or the like, which may be configured to perform one or more operations or actions for performing AI / ML model inference. According to embodiments, the MLI system 200 may include a Service Management and Orchestration (SMO), or a non-real time RAN Intelligent Controller (Non-RT RIC) included in the SMO. According to embodiments, the MLI system 200 may include a near-real-time RAN Intelligent Controller (Near-RT RIC).
[0044] The processor 210, as used herein, means any type of computational circuit that may comprise hardware elements and software elements. The processor 210 may be embodied as a multi-core processor, a single core processor, or a combination of one or more multi-core processors and one or more single core processors, a distributed processing system, or the like. The processor 210 may be a Central Processing Unit (CPU) a graphics processing unit (GPU), an accelerated processing unit (APU), an application-specific integrated circuit (ASIC), or another type of processing component.
[0045] Memory 220 may include a random-access memory (RAM), a read only memory (ROM), and / or another type of dynamic or static storage device (e.g., a flash memory, a magnetic memory, and / or an optical memory) that stores information and / or instructions for use by processor 210. The memory 220 may include machine-readable instructions which are executable by the processor 210. These machine-readable instructions when executed by the processor 210 causes the processor 210 to perform method steps of an exemplary embodiment described herein.
[0046] Storage component 230 may store information and / or software related to the operation and use of the MLI system 200. For example, storage component 230 may include a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optic disk, and / or a solid-state disk), a compact disc (CD), a digital versatile disc (DVD), a floppy disk, a cartridge, a magnetic tape, and / or another type of non-transitory computer-readable medium, along with a corresponding drive.
[0047] Input component 240 may be configured to receive information, such as via user input. For example, the input component 240 may include, but not be limited to, a touch screen display, a keyboard, a keypad, a mouse, a button, a switch, and / or a microphone. Additionally, oralternatively, the input component 240 may include a sensor for sensing information (e.g., a global positioning system (GPS), an accelerometer, a gyroscope, and / or an actuator).
[0048] Output component 250 may be configured to provide output information from the MLI system 200. For example, the output component 250 may be, but not limited to, a display, a speaker, and / or one or more light-emitting diodes (LEDs).
[0049] Communication interface 260 may be an interface that provides a communication connection to other devices. The connection by the communication interface 260 may be a wired connection, a wireless connection, or a combination of wired and wireless connections, and may be a direct connection or an indirect connection via a communication network that exists between other devices. In other words, the standard of the communication interface 260 is not limited.
[0050] The bus 270 may act as an interconnect between the processor 210, the memory 220, the storage component 230, the input component 240, the output component 250, and the communication interface 260 of the MLI system 200.
[0051] The number and arrangement of components shown in FIG. 2 are provided as an example. In practice, MLI system 200 may include additional components, fewer components, different components, or differently arranged components than those shown in FIG. 2. Additionally, or alternatively, a set of components (e.g., one or more components) of MLI system 200 may perform one or more functions described as being performed by another set of components of MLI system 200.
[0052] Descriptions of several example operations which may be performed by the processor 210 are provided below with reference to FIG. 6 to FIG. 7.Example of AI / ML model inference in the Present Disclosure
[0053] FIG. 3 illustrates an example flow of data in a system 300 for performing AI / ML model inference in a network, according to one or more embodiments.
[0054] As illustrated in FIG. 3, system 300 may include an rApp / Other Service Management and Orchestration Functions (SMOF) 310, an SMO / Non-RT RIC 320, and an SMO cloud 330.
[0055] According to embodiments, the rApp or Other Service Management and Orchestration Functions (SMOF) 310 are communicatively coupled to the SMO / Non-RT RIC 320 which provide an AI / ML workflow service (e.g., include an AI / ML workflow service provider 322 in the SMO / Non-RT RIC 320). According to embodiments, the SMO / Non-RT RIC 320 may correspond to a Service Management and Orchestration (SMO) itself, or a non-real time RAN Intelligent Controller (Non-RT RIC) included in the SMO. According to embodiments, the SMO cloud 330 may correspond to a cloud computing environment associated with the deployment of SMO / Non-RT RIC 320, rApps and other SMOFs. The SMO cloud 330 may correspond to a cloud computing environment separate from an O-Cloud in the O-RAN architecture, and may be located in the same physical location and use the same hardware as the O-Cloud. Part of the SMO Cloud 330 can be O-Cloud which exposes standardized interfaces for deployment life cycle management of rApps, other SMOFs or AI / ML models.
[0056] Further, as shown in FIG. 3, the SMO / Non-RT RIC 320 may include an AI / ML workflow service provider 322. According to embodiments, the AI / ML workflow service provider 322 may refer to or include a group of functions that provide AI / ML workflow services for the SMO. In particular, the AI / ML workflow service provider 322 may include at least an AI / ML training function 324 and an AI / ML inference function 326, where the AI / ML training function324 may include a function that trains AI / ML models using data collected by the SMO / Non-RT RIC 320 and the AI / ML inference function 326 may include a function that perform inference using AI / ML models.
[0057] Furthermore, as shown in FIG. 3, the SMO cloud 330 may include one or more node clusters 332 and resource pools 334. For example, the resource pool 334 may include one or more central processing units (CPUs), one or more graphical processing units (GPUs), one or more field programmable gate arrays (FPGAs), one or more application-specific integrated circuit (ASICs), one or more memory, one or more cloud tensor processing units (TPUs), and the like.
[0058] In order to perform AI / ML model inference, the rApp / Other SMOF 310 may first transmit an inference request to the SMO / Non-RT RIC 320. The inference request may include data regarding one or more AI / ML models which the inference is to be performed based on, as well as any other relevant artifacts and metadata related to the inference.
[0059] It may be understood that, prior to step 1 (i.e., prior to the process to perform AI / ML model inference), the rApp / Other SMOF 310 may transmit a AI / ML model to the SMO / Non-RT RIC 320 for training. The AI / ML model may be generated from any type of open-source tool, such as Tensorflow, PyTorch, Scikit-learn, and the like. The training of AI / ML models may be performed by the AI / ML training function 324 of the AI / ML workflow service provider 322 in the SMO / Non-RT RIC 320.
[0060] Once the AI / ML model is trained, the trained AI / ML model may be stored in a storage at the SMO / Non-RT RIC 320. In this regard, if the rApp / Other SMOF 310 would like to utilize (deploy) the trained AI / ML model that is stored in the storage at the SMO / Non-RT RIC 320 for the inference, the rApp / Other SMOF 310 may indicate a storage location at the SMO / Non-RTRIC 320 of such stored AI / ML model in the inference request. Alternatively, if the rApp / Other SMOF 310 would like to utilize a new AI / ML model that is not stored in the storage at the SMO / Non-RT RIC 320, the rApp / Other SMOF 310 may provide such new AI / ML model in the inference request. According to embodiments, the rApp / Other SMOF 310 may also indicate a storage location at the rApp / Other SMOF 310 itself of a stored AI / ML model in the inference request. Further, if the rApp / Other SMOF 310 would like to utilize more than one AI / ML models for the inference, where one of the AI / ML models is stored in the storage at the SMO / Non-RT RIC 320 and the other one of the AI / ML model is new, the rApp / Other SMOF 310 may provide both the storage location of the stored AI / ML model as well as the new AI / ML model in the inference request.
[0061] According to embodiments, the inference request may further include a scaling policy. The scaling policy may define policies for how to scale the deployment of a AI / ML model, where the SMO / Non-RT RIC 320 may automatically scale the deployment of the AI / ML model based on such scaling policy. According to embodiments, the deployment of a AI / ML model may be scaled based on a performance of the SMO cloud 330 that is deploying the AI / ML models. For example, the scaling policy may specify a number of instances of a AI / ML model that should be deployed in a Kubernetes node cluster, a number of pods and nodes in the Kubernetes node cluster that should be created to host the instances of the AI / ML model, and an assignment of the instances of the AI / ML model to the pods / nodes in the Kubernetes node cluster for different levels of performance of the SMO cloud 330. According to embodiments, the performance of the SMO cloud may include a speed for performing an inference, an accuracy of the results of the inference, and the like, and may be affected by resource required to deploy the AI / ML models (i.e., asspecified in the resource descriptor described below), available resource in the SMO cloud 330, workload demand from the rApp or other SMOF 310, and the like.
[0062] For example, as inference is being performed with the one or more AI / ML models deployed at the SMO cloud, the demand for such inference may increase (e.g., more users are utilizing services that requires the results of the inference) which may lead to decrease in speed for performing the inference. Accordingly, the deployment of the one or more AI / ML models may be scaled up in order to accommodate for the decrease in speed, where more instances of the one or more AI / ML models may be deployed in a Kubernetes node cluster to accommodate for the increase in demand. According to embodiments, the scaling policy may be specific for and defined by different rApps / SMOFs.
[0063] According to embodiments, the inference request may further include an inference graph. The inference graph may define interactions among the one or more AI / ML models for performing inference, where the SMO / Non-RT R1C 320 may automatically deploy the one or more AI / ML models having interactions based on such inference graph as shown in FIG. 4. It may be understood that, for a complex AI / ML application involving an inference with a plurality of AI / ML models, such plurality of AI / ML models may need to coordinate and interact in a certain manner in a pipeline (sequence). Accordingly, the inference graph may define how such plurality of AI / ML models connect and interact with each other, as well as how they are executed in a sequence. For example, the inference graph may define pipeline configuration, model input / output data routing targets and routing conditions, and the like.
[0064] According to embodiments, the inference request may further include a resource descriptor. The resource descriptor may define minimum requirements on computational resources(e.g., CPU, GPU, TPU resources and the like) to be allocated for the deployment of the one or more AI / ML models as well as any preference on a specific resources to be utilized for the deployment of the one or more AI / ML models (e.g., a GPU is to be utilized for acceleration), where the SMO / Non-RT RIC 320 may automatically allocate resources for deploying the AI / ML models according to such resource descriptor.
[0065] According to embodiments, the inference request may further include a measurement configuration defining measurement metrics that need to be reported, where the SMO / Non-RT RIC 320 may automatically report such measurement metrics according to such measurement configuration.
[0066] According to embodiments, the inference request may need to be authorized by the SMO / Non-RT RIC following the inference request being received by the SMO / Non-RT RIC.
[0067] According to embodiments, one or more data included in the inference request described above may instead be provided to the SMO / Non-RT RIC 320 before and separate from the inference request. For example, data related to the infrastructure and platform level configuration of the SMO / Non-RT RIC 320, such as the scaling policy and resource descriptor may be provided to the SMO / Non-RT RIC 320 by a user (e.g., platform developer) before and separate from the inference request, where the rApp / Other SMOF 310 simply needs to provide data related to the one or more AI / ML models, such as the storage location and the inference graph (which is specific to the one or more AI / ML models) to the SMO / Non-RT RIC 320 without having to prepare the scaling policy and resource descriptor for the inference.
[0068] Once the inference request is received by the SMO / Non-RT RIC 320, theSMO / Non-RT RIC 320 may obtain the one or more AI / ML models indicated in the inferencerequest. For example, if the inference request includes a storage location of a stored AI / ML model, the SMO / Non-RT RIC 320 may retrieve the one or more AI / ML models from the storage location. In another example, if the inference request includes the one or more AI / ML models, the SMO / Non-RT RIC 320 may simply receive the one or more AI / ML models from the inference request. Accordingly, the binary data of the one or more AI / ML models may be stored and loaded to the SMO / Non-RT RIC 320 following the inference request.
[0069] Once the one or more AI / ML models are obtained, the SMO / Non-RT RIC 320 may perform the inference using the one or more AI / ML models. In particular, the AI / ML inference function 326 of an AI / ML workflow service provider 322 in the SMO / Non-RT RIC 320 may perform the inference using the one or more AI / ML models.
[0070] According to embodiments, the AI / ML inference function 326 may perform the inference by first determining a cloud resource required and available for deploying the one or more AI / ML models. The cloud resource may be determined based on the resource descriptors. Once the cloud resource required and available for deploying the one or more AI / ML models is determined, the AI / ML inference function 326 may deploy the AI / ML models in the node cluster 332, where determined cloud resource may be allocated for the deployment of the deployed AI / ML models. Here, it may be understood that a configuration of the node cluster may be configured prior to performing the inference.
[0071] FIG. 4 illustrates a block diagram of an example of deploying one or more AI / ML models on a node cluster 400 which is automatically handled by the AI / ML inference function, according to one or more embodiments.
[0072] In the example shown in FIG. 4, the inference request may include AI / ML model X, Y, and Z for performing an inference, along with an inference graph defining the interactions among the AI / ML model X, Y, and Z and a scaling policy defining the scaling rules of the deployment of the AI / ML model X, Y, and Z.
[0073] In particular, a pipeline for performing an inference involving the AI / ML model X, Y, and Z may be as follows: the AI / ML model X may perform an initial step of the inference (e.g., image pre-processing for image classification) and provide a result to the AI / ML model Y; the AI / ML model Y may then receive the result from the AI / ML model X, perform a subsequent step of the inference (e.g., classify an image for image classification), and provide a result to the AI / ML model Z; and the AI / ML model Z may then receive the result from the AI / ML model Y, perform a final step of the inference (e.g., decision making for image classification), and provide a final result of the inference. Accordingly, the inference graph may define that an output from AI / ML model X is provided to an input of AI / ML model Y, and an output from AI / ML model Y is provided to an input of AI / ML model Z. Further, the scaling policy may define that, at the current performance of the SMO cloud, two instances of AI / ML model X, three instances of AI / ML model Y, and one instance of AI / ML model Z should be deployed.
[0074] Accordingly, as shown in FIG. 4, a node cluster 400 may include one or more server nodes, e.g. server node 400A and 400B. The AI / ML inference function may automatically deploy the AI / ML model X, Y, and Z at the node cluster 400 by creating K8S pods 410A, 420A, 410B, and 420B on the server nodes 400A and 400B to host the AI / ML model inference workloads, and then hosting instances of the AI / ML model X, Y, and Z at one or more of the K8S pods 410A, 420A, 410B, and 420B. For example, as shown in FIG. 4, pod 410A of server node 400A may hostan instance of AI / ML model X 411 A and an instance of AI / ML model Y 412A, pod 420 A of server node 400A may host an instance of AI / ML model X 421 A and an instance of AI / ML model Y 422A, pod 410B of server node 400B may host an instance of AI / ML model Y 41 IB, and pod 420B of server node 400B may host an instance of AI / ML model Z 42 IB.
[0075] According to embodiments, the number of the instances of the AI / ML models, the number of the pods, the number of the server nodes, as well as an assignment of the AI / ML models to a specific pod of a specific node may be determined automatically by the AI / ML inference function based on the scaling policy. Further, the above may also be changed in run-time according to the scaling policy as described above in relation to the scaling policy.
[0076] Further, once the instances of the AI / ML model X, Y, and Z are hosted at one or more of the K8S pods 410A, 420A, 410B, and 420B, data input and output of the instances of the AI / ML model X, Y, and Z may be routed to each other. For example, as shown in FIG. 4, the outputs from AI / ML model X 411 A and 421 A are provided (routed) to the inputs of AI / ML model Y 421 A, 422A, and 41 IB, and the outputs from AI / ML model Y 421 A, 422A, and 41 IB are provided to an input of AI / ML model Z 421B. The data input and output routing from one AI / ML model instance to another AI / ML model instance is automatically handled by the AI / ML inference function based on the inference graph received in the AI / ML inference request from the rApp or other SMO function.
[0077] According to embodiments, the AI / ML inference function 326 may further monitor the performance of the SMO cloud 330 (which is deploying the one or more AI / ML models), and then scale a deployment of the one or more AI / ML models based on the performance and a scaling policy. The deployment of the one or more AI / ML models may be scaled by, for example,increasing or decreasing a number of instances of the AI / ML model, the number of the pods hosting the instances of the AI / ML model, and / or server nodes of the node cluster 322. According to embodiments, the AI / ML inference function 326 may also monitor a resource utilization of the SMO cloud 330. Further, the AI / ML inference function 326 may also report the performance of the SMO cloud 330 to a user (e.g., platform developer, vendor, operator, and the like). For example, Key Performance Indicators (KPIs) and healthy data may be reported to the user for monitoring purposes.
[0078] For example, returning to FIG. 4, the scaling policy may specify that, at a certain performance level, three instances of AI / ML model X should be deployed. Accordingly, the AI / ML inference function 326 may monitor the performance of the SMO cloud 330, and once the performance of the SMO cloud 330 reaches the certain performance level, the AI / ML inference function 326 may deploy an additional instance of AI / ML model X at pod 410B in server node 400B.
[0079] It may be understood that, while the above descriptions and examples are provided in relation to SMO / Non-RT RIC, similar processes may be applicable to Near-RT RIC. For example, referring to FIG. 5, which illustrates an example flow of data in a system 500 for performing AI / ML model inference in a Near-RT RIC, according to one or more embodiments.
[0080] As illustrated in FIG. 5, the configuration of system 500 may be similar to the configuration of system 300 in FIG. 3, where the rApp / Other SMOF 310 is replaced with xApp 510, the SMO / Non-RT RIC 320 is replaced with Near-RT RIC 520, and an SMO cloud 330 is replaced with O-Cloud 530.
[0081] According to embodiments, the xApp 510 may correspond to an xApp that is communicatively coupled to the Near-RT RIC 520 via a Near-RT RIC API. Further, according to embodiments, the Near-RT RIC 520 may be communicatively coupled to the O-Cloud 530 via an 02 interface orE2 interface. In particular, the Near-RT RIC 520 may be communicatively coupled to a Deployment Management Service (DMS) in the O-Cloud 530 via the 02 or E2 interfaces, where the DMS may manage the deployment of the one or more AI / ML models in the node cluster 532.
[0082] Accordingly, the above process for performing AI / ML model inference provide key architecture solution in SMO / Non-RT RIC and Near-RT RIC framework to perform AI / ML model inference for rApps, other SMO functions / services and xApps.
[0083] Further, the above process may provide key architecture solution in SMO or Non- RT RIC for AI / ML model inference and identify the service to be exposed to the application layer including rApps and other SMO functions provided by the Al / AIL model inference service, service exposure via a SMOS or the R1 service interfaces, integration of AI / ML model inference service with other open-source tool and Kubernetes and AI / M1 model inference workflow in the SMO and RIC platform, and the like.
[0084] Furthermore, the above process may also allow application vendors to only focus on the AI / ML and business logics implementation in the rApps, without worrying about the complicated platform tasks such as workload distribution, cloud orchestration, networking, interface adaption, acceleration hardware access, element configurations (container, pods, K8S, helm charts, autoscaling rules, deployment artifacts, CPU / GPU / TPU, networking among the AI / ML model instances, and the like), and the like. Subsequently, development and onboardingof high seal ability, high performance AI / ML applications on SMO / Non-RT RIC platform may be simplified for vendors and operators, which may greatly lower the entry barriers, broadening the diversified supply chain of advanced AI / ML applications, lowering the development cost, and speeding up the introduction of new AI / ML solutions for O-RAN.Example Operations for Performing AI / ML Model Inference in the Present Disclosure
[0085] In the following, several example operations performable by the MLI system of the present disclosure are described with reference to FIG. 6 to FIG. 7.
[0086] FIG. 6 illustrates a flow diagram of an example method 600 for performing AI / ML model inference, according to one or more embodiments. One or more operations in method 600 may be performed by at least one processor (e.g., processor 210) of the MLI system.
[0087] As illustrated in FIG. 6, at operation S610, the at least one processor may be configured to receive an inference request. The inference may include one or more storage locations of one or more AI / ML models and / or the one or more AI / ML models themselves. According to embodiments, the inference may further include one or more of a scaling policy, an inference graph, a resource requirement, and a measurement configuration. According to embodiments, the scaling policy may be provided to the MLI system by a user (e.g., platform developer) before and separate from the inference request.
[0088] According to embodiments, the inference request may be received from an rApp via an R1 interface. According to embodiments, the inference request may be received from a Service Management and Orchestration (SMO) function via an Artificial Intelligence / Machine Learning (AI / ML) workflow service Application Programming Interface (API). According toembodiments, the inference request may be received from an xApp via a Near-RT RIC API. The method then proceeds to operation S620.
[0089] At operation S620, the at least one processor may be configured to obtain one or more AI / ML models based on the inference request. According to embodiments, the one or more AI / ML models may be obtained by retrieving the one or more AI / ML models from the one or more storage locations indicated in the inference request, and / or receiving the one or more AI / ML models via the inference request. The method then proceeds to operation S630.
[0090] At operation S630, the at least one processor may be configured to perform an inference. According to embodiments, the inference may be performed at a Service Management and Orchestration (SMO) / Non-RT RIC using the one or more AI / ML models. According to embodiments, the inference may be performed at a Near-RT RIC using the one or more AI / ML models. Examples of operations for performing an inference at the SMO are described below with reference to FIG. 7.
[0091] Upon performing operation S630, the method 600 may be ended or be terminated. Alternatively, method 600 may return to operation S610, such that the at least one processor may be configured to repeatedly perform, for at least a predetermined amount of time, the receiving the inference request (at operation S610), the obtaining the one or more AI / ML models (at operation S620), and the performing the inference (at operation S630). For instance, the at least one processor may continuously (or periodically) receive the request to perform inference from the rApp / xApp, and then restart the receiving the inference request (at operation S610), the obtaining the one or more AI / ML models (at operation S620), and the performing the inference (at operation S630).
[0092] To this end, the system of the present disclosure may automatically deploy one or more AI / ML models at the SMO / Non-RT RIC and Near-RT RIC in order to perform AI / ML model inference.Example Operations for Performing an Inference at the SMO in the Present Disclosure
[0093] FIG. 7 illustrates a flow diagram of an example method 700 for performing an inference at the SMO, according to one or more embodiments. One or more operations of method 700 may be part of operation S630 in method 600, and may be performed by at least one processor (e.g., processor 210) of the MLI system.
[0094] As illustrated in FIG. 7, at operation S710, the at least one processor may be configured to determine cloud resource required and available for deploying the one or more AI / ML models. The method then proceeds to operation S720.
[0095] At operation S720, the at least one processor may be configured to deploy the one or more AI / ML models at the SMO cloud. According to embodiments, the determined cloud resource (i.e., cloud resource required and available for deploying the one or more AI / ML models determined during operation S710) may be allocated for deployment of the one or more AI / ML models.
[0096] According to embodiments, the at least one processor may be configured to deploy the one or more AI / ML models at the SMO cloud by: creating one or more nodes comprising one or more pods in a Kubemetes node cluster; hosting one or more instances of the one or more AI / ML models at the one or more pods; and routing data input and data output of the one or more instances of the one or more AI / ML models.
[0097] According to embodiments, the one or more nodes and the one or more pods may be created based on a scaling policy. According to embodiments, the one or more instances of the one or more AI / ML models may be hosted at the one or more pods based on the scaling policy. According to embodiments, the data input and data output of the one or more instances of the one or more AI / ML models may be routed based on an inference graph. The method then proceeds to operation S730.
[0098] At operation S730, the at least one processor may be configured to monitor a performance of the SMO cloud. According to embodiments, the monitored performance of the SMO cloud may be reported to a user (e.g., platform developer, vendor, operator, and the like). The method then proceeds to operation S740.
[0099] At operation S740, the at least one processor may be configured to scale a deployment of the one or more AI / ML models based on the performance and a scaling policy. According to embodiments, the deployment of the one or more AI / ML models may be scaled by increasing or decreasing the number of instances of the one or more AI / ML models, pods, and / or server nodes of the Kubemetes node cluster.
[0100] It may be understood that, while the above examples operations in FIG. 7 are described in relation to SMO / Non-RT RIC, similar processes may be applicable to Near-RT RIC as explained above in relation to FIG. 5.Example Implementation Environment
[0101] FIG. 8 illustrates a diagram of an example environment 800 in which systems and / or methods, described herein, may be implemented. As shown in FIG. 8, environment 800 may include a device 810, a platform 820, and a network 830. Devices of environment 800 mayinterconnect via wired connections, wireless connections, or a combination of wired and wireless connections. In some embodiments, any of the functions and operations described with reference to FIG. 2 to FIG. 7 above may be performed by any combination of elements illustrated in FIG. 8.
[0102] Device 810 may include one or more devices capable of receiving, generating, storing, processing, and / or providing information associated with platform 820. For example, device 810 may include a computing device (e.g., a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc.), a mobile phone (e.g., a smart phone, a radiotelephone, etc.), a wearable device (e.g., a pair of smart glasses or a smart watch), or a similar device. In some implementations, device 810 may receive information from and / or transmit information to platform 820.
[0103] Platform 820 includes one or more devices capable of receiving, generating, storing, processing, and / or providing information. In some implementations, platform 820 may include a cloud server or a group of cloud servers. In some implementations, platform 820 may be designed to be modular such that certain software components may be swapped in or out depending on a particular need. As such, platform 820 may be easily and / or quickly reconfigured for different uses.
[0104] In some implementations, as shown, platform 820 may be hosted in cloud computing environment 822. Notably, while implementations described herein describe platform 820 as being hosted in cloud computing environment 822, in some implementations, platform 820 may not be cloud-based (i.e., may be implemented outside of a cloud computing environment) or may be partially cloud-based.
[0105] Cloud computing environment 822 includes an environment that hosts platform 820.Cloud computing environment 822 may provide computation, software, data access, storage, etc.services that do not require end-user (e.g., user device 810) knowledge of a physical location and configuration of system(s) and / or device(s) that hosts platform 820. As shown, cloud computing environment 822 may include a group of computing resources 824 (referred to collectively as “computing resources 824” and individually as “computing resource 824”).
[0106] Computing resource 824 includes one or more personal computers, a cluster of computing devices, workstation computers, server devices, or other types of computation and / or communication devices. In some implementations, computing resource 824 may host platform 820. The cloud resources may include compute instances executing in computing resource 824, storage devices provided in computing resource 824, data transfer devices provided by computing resource 824, etc. In some implementations, computing resource 824 may communicate with other computing resources 824 via wired connections, wireless connections, or a combination of wired and wireless connections.
[0107] As further shown in FIG. 8, computing resource 824 includes a group of cloud resources, such as one or more applications (“APPs”) 824-1, one or more virtual machines (“VMs”) 824-2, virtualized storage (“VSs”) 824-3, one or more hypervisors (“HYPs”) 824-4, or the like. While the current example embodiment is with reference to virtualized network functions, it is understood that one or more other embodiments are not limited to a particular type of cloud computing environment, and may be implemented in at least one of containers, cloud-native services, one or more container platforms, etc. For example, in one or more other example embodiments, any of the above-described components may be a software-based component deployed or hosted in, for example, a server cluster such as a hybrid cloud server, data center servers, and the like. The software-based component may be containerized and may be deployedand controlled by one or more machines, called “nodes”, that run or execute the containerized network elements and are addressable. In this regard, a server cluster may contain at least one master node and a plurality of worker nodes, wherein the master node(s) controls and manages a set of associated worker nodes.
[0108] According to embodiments, example embodiments described herein may be implemented or be deployed in the server platform described above, in the form of virtualized network function (VNF). In this regard, it is contemplated that the terms “virtual”, “virtualized”, or the like, described hereinabove are merely intended to specify the nature of the machine (and the elements and resources associated therewith) being provided in virtual or software form. In this regard, the “virtual machine”, “virtualized storage”, and the like, described hereinabove should not be limited to any specific type of virtual machine or virtual element. Accordingly, it can be understood that the (or operations associated therewith) may be defined or presented in the form of a containerized network function, of which the functions may be provided in the form of containers.
[0109] Application 824- 1 includes one or more software applications that may be provided to or accessed by user device 810. Application 824-1 may eliminate a need to install and execute the software applications on user device 810. For example, application 824-1 may include software associated with platform 820 and / or any other software capable of being provided via cloud computing environment 822. In some implementations, one application 824-1 may send / receive information to / from one or more other applications 824-1, via virtual machine 824-2.
[0110] Virtual machine 824-2 includes a software implementation of a machine (e.g., a computer) that executes programs like a physical machine. Virtual machine 824-2 may be either asystem virtual machine or a process virtual machine, depending upon use and degree of correspondence to any real machine by virtual machine 824-2. A system virtual machine may provide a complete system platform that supports execution of a complete operating system (“OS”). A process virtual machine may execute a single program, and may support a single process. In some implementations, virtual machine 824-2 may execute on behalf of a user (e g., user device 810), and may manage infrastructure of cloud computing environment 822, such as data management, synchronization, or long-duration data transfers.
[0111] Virtualized storage 824-3 includes one or more storage systems and / or one or more devices that use virtualization techniques within the storage systems or devices of computing resource 824. In some implementations, within the context of a storage system, types of virtualizations may include block virtualization and file virtualization. Block virtualization may refer to abstraction (or separation) of logical storage from physical storage so that the storage system may be accessed without regard to physical storage or heterogeneous structure. The separation may permit administrators of the storage system flexibility in how the administrators manage storage for end users. File virtualization may eliminate dependencies between data accessed at a file level and a location where files are physically stored. This may enable optimization of storage use, server consolidation, and / or performance of non-disruptive file migrations.
[0112] Hypervisor 824-4 may provide hardware virtualization techniques that allow multiple operating systems (e g., “guest operating systems”) to execute concurrently on a host computer, such as computing resource 824. Hypervisor 824-4 may present a virtual operating platform to the guest operating systems, and may manage the execution of the guest operatingsystems. Multiple instances of a variety of operating systems may share virtualized hardware resources.
[0113] Network 830 may include one or more wired and / or wireless networks. For example, network 830 may include a cellular network (e.g., a fifth generation (5G) network, a long-term evolution (LTE) network, a third generation (3G) network, a code division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a telephone network (e.g., the Public Switched Telephone Network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber optic-based network, or the like, and / or a combination of these or other types of networks.
[0114] The number and arrangement of devices and networks shown in FIG. 8 is provided as an example. In practice, there may be additional devices and / or networks, fewer devices and / or networks, different devices and / or networks, or differently arranged devices and / or networks than those shown in FIG. 8. Furthermore, two or more devices shown in FIG. 8 may be implemented within a single device, or a single device shown in FIG. 8 may be implemented as multiple, distributed devices. Additionally, or alternatively, a set of devices (e.g., one or more devices) of environment 800 may perform one or more functions described as being performed by another set of devices of environment 800.Various Aspects of Embodiments
[0115] According to example embodiments, a AI / ML model inference is performed at aService Management and Orchestration (SMO) / Non-RT RIC or Near-RT RIC using one or moreAI / ML models provided by an rApp / Other SMO functions or an xApp. In this regard, since theinference is performed at the SMO / Non-RT RIC or Near-RT RIC instead of the rApp / Other SMO functions or the xApp, the SMO / Non-RT RIC and Near-RT RIC may be provided with key architecture solution to perform AI / ML model inference without relying on the rApps and xApps. Subsequently, the application vendors are allowed to only focus on the AI / ML and business logics implementation in the rApps and xApps, without worrying about the complicated platform tasks such as workload distribution, cloud orchestration, networking, interface adaption, acceleration hardware access, and the like.
[0116] The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit the implementations to the precise form disclosed. Modifications and variations are possible in light of the above disclosure or may be acquired from practice of the implementations.
[0117] Some embodiments may relate to a system, a method, and / or a computer readable medium at any possible technical detail level of integration. Further, one or more of the above components described above may be implemented as instructions stored on a computer readable medium and executable by at least one processor (and / or may include at least one processor). The computer readable medium may include a computer-readable non-transitory storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out operations.
[0118] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storagedevice, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.
[0119] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.
[0120] Computer readable program code / instructions for carrying out operations may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions,machine dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuitry, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++, or the like, and procedural programming languages, such as the “C” programming language or similar programming languages. The computer readable program instructions may execute entirely on the user’ s computer, partly on the user’s computer, as a standalone software package, partly on the user’s computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user’s computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects or operations.
[0121] These computer readable program instructions may be provided to a processor of a general-purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function ina particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram block or blocks.
[0122] The computer readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0123] The flowchart and block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer readable media according to various embodiments. In this regard, each block in the flowchart or block diagrams may represent a microservice(s) module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s). The method, computer system, and computer readable medium may include additional blocks, fewer blocks, different blocks, or differently arranged blocks than those depicted in the Figures. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the Figures. For example, two blocks shown in succession may, in fact, be executed concurrently or substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems thatperform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions.
[0124] It will be apparent that systems and / or methods, described herein, may be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual specialized control hardware or software code used to implement these systems and / or methods is not limiting of the implementations. Thus, the operation and behavior of the systems and / or methods were described herein without reference to specific software code-it being understood that software and hardware may be designed to implement the systems and / or methods based on the description herein.
[0125] Various further respective aspects and features of embodiments of the present disclosure may be defined by the following items:Item [1]: An apparatus that may be configured to: receive an inference request; obtain one or more AI / ML models based on the inference request; and perform an inference at a Service Management and Orchestration (SMO) / Non-RT RIC or Near-RT RIC using the one or more AI / ML models.Item [2]: The apparatus according to item [1], wherein the apparatus may be configured to perform the inference by: determining cloud resource required and available for deploying the one or more AI / ML models; and deploying the one or more AI / ML models at an SMO cloud or an O-Cloud, wherein the determined cloud resource may be allocated for deployment of the one or more AI / ML models.Item [3]: The apparatus according to item [2], wherein the apparatus may be configured to perform the inference by: monitoring a performance of the SMO cloud or theO-Cloud; and scaling a deployment of the one or more AI / ML models based on the performance and a scaling policy.Item [4]: The apparatus according to one of items [2]-[3], wherein the Near-RT RIC may be communicatively coupled to a Deployment Management Service (DMS) in the O- Cloud via an 02 interface or an E2 interface.Item [5]: The apparatus according to one of items [ l]-[4], wherein the inference request may include one or more storage locations of the one or more AI / ML models, and wherein the apparatus may be configured to obtain the one or more AI / ML models by retrieving the one or more AI / ML models from the one or more storage locations.Item [6]: The apparatus according to one of items [l]-[5], wherein the inference request may include the one or more AI / ML models, and wherein the apparatus may be configured to obtain the one or more AI / ML models by receiving the one or more AI / ML models via the inference request.Item [7]: The apparatus according to one of items [l]-[6], wherein the inference request may include one or more of a model location, a resource descriptor, a scaling policy, an inference graph, and a measurement configuration.Item [8]: The apparatus according to one of items [l]-[7], wherein the inference request may be received from an rApp via an R1 interface, from a Service Management and Orchestration (SMO) function via an Artificial Intelligence / Machine Learning (AI / ML) workflow service Application Programming Interface (API), or from an xApp via a Near-RT RIC API.Item [9]: A method that may include: receiving an inference request; obtaining one or more AI / ML models based on the inference request; and performing an inference at a Service Management and Orchestration (SMO) / Non-RT RIC or Near-RT RIC using the one or more AI / ML models.Item
[0010] : The method according to item [9], wherein the performing the inference may include: determining cloud resource required and available for deploying the one or more AI / ML models; and deploying the one or more AI / ML models at an SMO cloud or an O-Cloud, wherein the determined cloud resource may be allocated for deployment of the one or more AI / ML models.Item
[0011] : The method according to item
[0010] , wherein the performing the inference may include: monitoring a performance of the SMO cloud or the O-Cloud; and scaling a deployment of the one or more AI / ML models based on the performance and a scaling policy.Item
[0012] : The method according to one of items
[0010] -[l 1], wherein the Near-RT RIC may be communicatively coupled to a Deployment Management Service (DMS) in the O-Cloud via an 02 interface or an E2 interface.Item
[0013] : The method according to one of items [9]-
[0012] , wherein the inference request may include one or more storage locations of the one or more AI / ML models, and wherein the obtaining the one or more AI / ML models may include retrieving the one or more AI / ML models from the one or more storage locations.Item
[0014] : The method according to one of items [9]-
[0013] , wherein the inference request may include the one or more AI / ML models, and wherein the obtaining the one ormore AI / ML models may include receiving the one or more AI / ML models via the inference request.Item
[0015] : The method according to one of items [9]-
[0014] , wherein the inference request may include one or more of a model location, a resource descriptor, a scaling policy, an inference graph, and a measurement configuration.Item
[0016] : The method according to one of items [9]-[l 5], wherein the inference request may be received from an rApp via an R1 interface, from a Service Management and Orchestration (SMO) function via an Artificial Intelligence / Machine Learning (AI / ML) workflow service Application Programming Interface (API), or from an xApp via a Near- RT RIC API.Item
[0017] : A non-transitory computer-readable recording medium that may have recorded thereon instructions executable by an apparatus to cause the apparatus to perform a method including: receiving an inference request; obtaining one or more AI / ML models based on the inference request; and performing an inference at a Service Management and Orchestration (SMO) / Non-RT RIC or Near-RT RIC using the one or more AI / ML models.Item
[0018] : The non-transitory computer-readable recording medium according to item
[0017] , wherein the performing the inference may include: determining cloud resource required and available for deploying the one or more AI / ML models; and deploying the one or more AI / ML models at an SMO cloud or an O-Cloud, wherein the determined cloud resource may be allocated for deployment of the one or more AI / ML models.Item
[0019] : The non-transitory computer-readable recording medium according to item
[0018] , wherein the performing the inference may include: monitoring a performance ofthe SMO cloud or the O-Cloud; and scaling a deployment of the one or more AI / ML models based on the performance and a scaling policy.Item
[0020] : The non-transitory computer-readable recording medium according to item
[0018] , wherein the Near-RT RIC may be communicatively coupled to a Deployment Management Service (DMS) in the O-Cloud via an 02 interface or an E2 interface.
[0126] It can be understood that numerous modifications and variations of the present disclosure are possible in light of the above teachings. It will be apparent that within the scope of the appended clauses, the present disclosures may be practiced otherwise than as specifically described herein.
Claims
What is claimed is:
1. An apparatus configured to: receive an inference request; obtain one or more AI / ML models based on the inference request; and perform an inference at a Service Management and Orchestration (SMO) / Non-RT RIC or Near-RT RIC using the one or more AI / ML models.
2. The apparatus according to claim 1, wherein the apparatus is configured to perform the inference by: determining cloud resource required and available for deploying the one or more AI / ML models; and deploying the one or more AI / ML models at an SMO cloud or an O-Cloud, wherein the determined cloud resource is allocated for deployment of the one or more AI / ML models.
3. The apparatus according to claim 2, wherein the apparatus is configured to perform the inference by: monitoring a performance of the SMO cloud or the O-Cloud; and scaling a deployment of the one or more AI / ML models based on the performance and a scaling policy.
4. The apparatus according to claim 2, wherein the Near-RT RIC is communicatively coupled to a Deployment Management Service (DMS) in the O-Cloud via an 02 interface or an E2 interface.
5. The apparatus according to claim 1, wherein the inference request comprises one or more storage locations of the one or more AI / ML models, and wherein the apparatus is configured to obtain the one or more AI / ML models by retrieving the one or more AI / ML models from the one or more storage locations.
6. The apparatus according to claim 1, wherein the inference request comprises the one or more AI / ML models, and wherein the apparatus is configured to obtain the one or more AI / ML models by receiving the one or more AI / ML models via the inference request.
7. The apparatus according to claim 1, wherein the inference request comprises one or more of a model location, a resource descriptor, a scaling policy, an inference graph, and a measurement configuration.
8. The apparatus according to claim 1, wherein the inference request is received from an rApp via an R1 interface, from a Service Management and Orchestration (SMO) function via an Artificial Intelligence / Machine Learning (AI / ML) workflow service Application Programming Interface (API), or from an xApp via a Near-RT RIC API.
9. A method comprising: receiving an inference request; obtaining one or more AI / ML models based on the inference request; and performing an inference at a Service Management and Orchestration (SMO) / Non-RT RIC or Near-RT RIC using the one or more AI / ML models.
10. The method according to claim 9, wherein the performing the inference comprises: determining cloud resource required and available for deploying the one or more AI / ML models; and deploying the one or more AI / ML models at an SMO cloud or an O-Cloud, wherein the determined cloud resource is allocated for deployment of the one or more AI / ML models.
11. The method according to claim 10, wherein the performing the inference comprises: monitoring a performance of the SMO cloud or the O-Cloud; and scaling a deployment of the one or more AI / ML models based on the performance and a scaling policy.
12. The method according to claim 10, wherein the Near-RT RIC is communicatively coupled to a Deployment Management Service (DMS) in the O-Cloud via an 02 interface or an E2 interface.
13. The method according to claim 9, wherein the inference request comprises one or more storage locations of the one or more AI / ML models, and wherein the obtaining the one or more AI / ML models comprises retrieving the one or more AI / ML models from the one or more storage locations.
14. The method according to claim 9, wherein the inference request comprises the one or more AI / ML models, and wherein the obtaining the one or more AI / ML models comprises receiving the one or more AI / ML models via the inference request.
15. The method according to claim 9, wherein the inference request comprises one or more of a model location, a resource descriptor, a scaling policy, an inference graph, and a measurement configuration.
16. The method according to claim 9, wherein the inference request is received from an rApp via an R1 interface, from a Service Management and Orchestration (SMO) function via an Artificial Intelligence / Machine Learning (AI / ML) workflow service Application Programming Interface (API), or from an xApp via a Near-RT RIC API.
17. A non-transitory computer-readable recording medium having recorded thereon instructions executable by an apparatus to cause the apparatus to perform a method comprising: receiving an inference request;obtaining one or more AI / ML models based on the inference request; and performing an inference at a Service Management and Orchestration (SMO) / Non-RTRIC or Near-RT RIC using the one or more AI / ML models.
18. The non-transitory computer-readable recording medium according to claim 17, wherein the performing the inference comprises: determining cloud resource required and available for deploying the one or more AI / ML models; and deploying the one or more AI / ML models at an SMO cloud or an O-Cloud, wherein the determined cloud resource is allocated for deployment of the one or more AI / ML models.
19. The non-transitory computer-readable recording medium according to claim 18, wherein the performing the inference comprises: monitoring a performance of the SMO cloud or the O-Cloud; and scaling a deployment of the one or more AI / ML models based on the performance and a scaling policy.
20. The non-transitory computer-readable recording medium according to claim 18, wherein the Near-RT RIC is communicatively coupled to a Deployment Management Service (DMS) in the O-Cloud via an 02 interface or an E2 interface.
Citation Information
Patent Citations
Data processing method and device and communication equipment
CN116419209A
Optimized diagnostics plan for an information handling system
US20220342738A1
Radio access network intelligent application manager
WO2023091664A1
Network aware compute resource management use case for o-ran non-RT ric
WO2023177419A1
Centralized machine learning model configurations
WO2023184310A1