Machine learning model remote management in advanced communication networks
By publishing machine learning model capability data in O-RAN and remotely managing it, the performance degradation and cost increase caused by multi-vendor deployment is solved, and the accuracy of autonomous RAN optimization and real-time decision-making is achieved.
Patent Information
- Application Number
- CN202380087232.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-20
- Filing Date
- 2023-10-26
- Publication Date
- 2025-07-25
AI Technical Summary
In Open Radio Access Networks (O-RAN), the management of machine learning models results in hardware and software differences due to multi-vendor deployment, resulting in performance degradation and increased operating costs. The existing approach relies on hardware and software matching, resulting in inefficiency.
Through the release and remote management of machine learning model capabilities data, model training and selection is used using the interface defined by O-RAN to eliminate hardware and software dependencies, realize autonomous RAN optimization, and adapt to changes in network conditions.
Improve the inference accuracy and efficiency of machine learning models, reduce operational costs, and provide greater supplier flexibility and real-time decision-making accuracy.
Smart Images

Figure CN120380486A_ABST
Abstract
Description
[0001] Related Applications
[0002] This application claims priority to U.S. Non - Provisional Patent Application No. 18 / 068,842, filed on December 20, 2022, entitled "MACHINE LEARNING MODEL REMOTE MANAGEMENT IN ADVANCED COMMUNICATION NETWORKS", the entire content of which is incorporated herein by reference. BACKGROUND OF THE INVENTION
[0003] Different radio access network (RAN) optimization use cases can apply different machine learning models for real - time prediction. This includes supervised learning models, unsupervised learning models, and reinforcement learning models to balance training complexity and use - case - specific performance. Such models need to be managed throughout the RAN lifecycle to address various new RAN functions and business models emerging in the network.
[0004] In an Open Radio Access Network (O - RAN or sometimes also referred to as Open RAN), machine learning models can be deployed at edge nodes (e.g., near - real - time RAN intelligent controller or near RT RIC), distributed units (DUs), or centralized units (CUs) for fast optimization actions. These machine learning models are remotely managed by an operator (e.g., via Service Management and Orchestration or SMO component) to change the selected machine learning models or adjust some of their hyperparameters (after external training on a powerful machine).
[0005] In an O - RAN deployment, the model operator (coupled via SMO, for example) and the model host (e.g., near RT RIC) can be from different vendors, thus adopting different software images, platforms, and hardware models. In a multi - vendor deployment, such differences make machine learning model management very challenging when considering dynamic network conditions (e.g., time - varying traffic types and loads). As a result, the network will likely use machine learning models that are not optimally trained, leading to performance degradation. Alternatively, the operator will need to perform manual model updates based on humans (e.g., remotely logging into the inference host), which increases the overall operational cost of the O - RAN deployment.
[0006] Existing O-RAN methods for machine learning model deployment include image-based options and software-based options. In the image-based option, the operator sends a container with the machine model software and dependencies (e.g., libraries). One of the drawbacks of the image-based option is the dependence on the model host hardware capabilities; for example, the model host must support the same container runtime environment. Additionally, a model host with lower hardware capabilities (compared to the model operator hardware) has lower runtime efficiency. In the software-based option, the operator sends the model file; this approach results in significant drawbacks as it requires both the model operator and the host to support the same machine learning software library. Consider an example where the model operator runs on a powerful machine (e.g., higher central processing unit (CPU) capability = 100) and supports a machine learning model as a python file, while the model host runs on a more modest machine (e.g., CPU capability = 10) and only supports machine learning models in the form of C++ (cpp files). Consequently, engineers must perform software upgrades to match the software of the two entities (high cost), and model training and inference at the host will take a long time to provide network recommendations, resulting in sub-optimal network performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] The techniques described herein are illustrated by way of example and not limitation in the figures, in which like reference numerals indicate similar elements and in which:
[0008] Figure 1 and Figure 2 A block diagram representation of a system / architecture including example radio access network (RAN) components according to various aspects and embodiments of the present disclosure, wherein a machine learning model can be remotely managed by the example RAN components.
[0009] Figure 3 An example component, data flow, and sequence diagram showing an overview of operations related to the remote management of a machine learning model according to various aspects and embodiments of the present disclosure.
[0010] Figure 4 An example component and signaling diagram showing an example data flow sequence related to an embodiment of machine learning model metadata and configuration exchange according to various aspects and embodiments of the present disclosure.
[0011] Figure 5 An example component and signaling diagram showing an example data flow sequence related to an embodiment of remote machine learning model selection according to various aspects and embodiments of the present disclosure.
[0012] Figure 6is a flowchart showing example operations related to receiving, at a model host, a machine learning model and associated data in response to a model host publishing machine learning model capabilities data, in accordance with various aspects and embodiments disclosed herein.
[0013] Figure 7 is a flowchart showing example operations related to receiving, at a model host, a set of machine learning models and associated data for selection in response to a model host publishing machine learning model capabilities data, in accordance with various aspects and embodiments disclosed herein.
[0014] Figure 8 is a flowchart showing example operations related to receiving, at a model host, a machine learning model and associated data and training the machine learning model at the model host, in accordance with various aspects and embodiments disclosed herein.
[0015] Figure 9 is a block diagram representing an example computing environment in which aspects of the subject matter described herein may be incorporated.
[0016] Figure 10 depicts an example schematic block diagram of a computing environment in accordance with various aspects and embodiments disclosed herein, with which the disclosed subject matter may interact at least in part / be at least partially implemented. Detailed Description
[0017] Aspects of the techniques described herein generally relate to remotely managing machine learning models, including in O-RAN-based embodiments. A network function publishes its machine learning capabilities for post-processing on a remote machine (e.g., via a Service Management and Orchestration (SMO) component). Example network functions that can advertise their capabilities data can include, but are not limited to, a Near Real-Time RAN Intelligent Controller (Near RT RIC), a Distributed Unit (DU), a Centralized Unit (CU), a Radio Unit (RU), etc. In this way, each network function hosting at least one machine learning model exposes its machine learning model capabilities data, such as including but not limited to the supported model name, input Key Performance Indicators (KPIs), features, output actions, etc. As a result, software images, platforms, and hardware models can be matched to the published machine learning capabilities of the model host.
[0018] As will be appreciated, machine learning model post-processing may include training, model selection, configuration, and / or activation / deactivation of the model via interfaces (e.g., A1 and O1) defined by connecting an operator to a network function that will host or is currently hosting the model (e.g., O-RAN). The operator trains the model via the SMO orchestrator, where parameters and coefficients are adjusted to improve inference performance. Also as described herein, one or more input features may be added (to improve the performance of the model) or removed (to reduce the complexity of the model), which promotes adaptability. As a result of such adaptability, any mismatch between the (e.g., powerful) hardware of the model operator and the (e.g., less powerful) hardware of the model host is eliminated.
[0019] In Figure 4 One example embodiment generally represented therein, the machine learning model is communicated and deployed using the O-RAN defined configuration O1 interface. To this end, the O1 interface (e.g., yang file) is updated to include machine learning model parameters such as model name, model specific parameters, and coefficients determined during training. For example, the operator may specify a neural network model such as including coefficients, layers, input key performance indicators (KPIs), and model outputs (either KPIs or network decisions such as configuration parameter values). The criterion-defined network configuration (NETCONF) protocol may be used for initial model configuration as well as model updates during runtime.
[0020] In Figure 5 Another example embodiment generally represented therein, the host is configured for remote machine learning model selection. In this embodiment, the model host supports multiple machine learning models (e.g., communicated and deployed as above); however, the host does not have sufficient resources to train all models and instead selects the most likely optimal model. The host may evaluate the performance of the model against (multiple) reselection criteria, where the (multiple) reselection criteria allow the local host to perform local reselection between different models during runtime when performance is insufficient. The host may also request the operator to re-evaluate and re-select or re-deploy a new model (or group of models) via the O1 or A1 O-RAN interface. The selection made by the operator may be based on the model performance indicated by the performance monitoring module (at the SMO) and may also indicate performance-based reselection criteria, where the performance-based reselection criteria allow the local host to perform local reselection between different models during runtime.
[0021] It should be understood that any example herein is non-limiting. As an example, the technology has been described in an O-RAN environment. However, this is merely an example and the technology can be implemented in similar environments (including those not yet implemented). Thus, any embodiment, aspect, concept, structure, function, or example described herein is non-limiting, and the technology can be used in various ways that generally provide benefits and advantages in the areas of communication and computing. It should also be noted that terms such as "optimize" or "optimal" used herein (e.g., "maximize", "minimize", etc.) only indicate a goal of moving towards a better state and do not necessarily result in an ideal outcome.
[0022] References to "one embodiment", "an embodiment", "one implementation", "an implementation", etc. throughout this specification mean that the particular features, structures, or characteristics described in connection with that embodiment / implementation are included in at least one embodiment / implementation. Thus, phrases such as "in one embodiment" or "in an implementation" that appear throughout this specification do not necessarily all refer to the same embodiment / implementation. Further, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments / implementations.
[0023] Aspects of the present subject disclosure will now be described more fully hereinafter with reference to the accompanying drawings, in which example components, diagrams, and / or operations are shown. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the various embodiments. However, the present subject disclosure may be embodied in many different forms and should not be construed as limited to the examples set forth herein.
[0024] Figure 1 and Figure 2 FIG. 100 shows a system / architecture of a RAN agent, including radio units 102(1)-102(n) coupled to a distributed unit (DU) 104 (with a high-speed (PCIe) accelerator) via an enhanced common public radio interface (eCPRI). In turn, the DU 104 is coupled to a centralized unit control plane (CU-CP) 106 (via an F1-C fronthaul control plane interface) and a centralized unit user plane (CU-UP) 108 (via an F1-U fronthaul user plane interface) through defined interfaces, and is coupled to a router 110 via an O1 / E2 interface. The router 110 is also communicatively coupled to the CU-CP 106.
[0025] The router 110 facilitates communication between these components and a near real-time RIC 112, a new radio core 114, and a service management and orchestration component 116. The service management and orchestration component 116 includes an application 118 and a non-real-time RIC 120.
[0026] Figure 1 The dashed arrows depicted in communicate data of the system agent via the interface are coupled / stored to the data hub 222( Figure 2 ), whereby the communication via the interface can be Figure 2 accessed by other components depicted in. In Figure 2 , the thick line 224 indicates that the data-related metadata in the yang model for machine learning is coupled to the model pipeline 226. The model parameters and registry metadata in the yang model for machine learning are represented by the thick line 228 and can be accessed by the model parameter registry 230, the performance metric registries 232 and 234, where the performance metric registries 232 and 234 can be used for performance verification (box) 236. The solid line 240 indicates the output decision / feature metadata in the yang model for machine learning, where the decision trigger component 242 is coupled to Figure 1 the non-real-time RIC 120 and the near-real-time RIC 112 of
[0027] Figure 3 shows an overview of the sequence and data flow between the operator 330, the orchestration component (SMO) 316, and the model host 312. As shown by arrow one (1), the model host 312 presents its machine learning model capabilities, such as the supported model name, input KPIs (features), and output actions.
[0028] As shown by box two (2), the operator trains the model via the SMO orchestrator 316 (2a), where the parameters and coefficients are adjusted to improve the inference performance. Some input features can be added (to improve performance) or removed (to reduce complexity). The operator 330 can also select (2b) a machine learning model (or model group) for the model host 312.
[0029] As shown by arrow three (3), the operator 330 / orchestration component 316 sends the machine learning model metadata and parameters to the model host (the network function that publishes its capability data), whereby the machine learning model host 312 applies the associated parameters to the existing local model, i.e., applies them in the host's environment and software, and sends an acknowledgement message (not explicitly shown). As the network function of the model host 312 uses the indicated parameters for local decision-making and model operation, including for real-time inference using local / actual observation data to trigger model retraining, thereby reducing features when appropriate to reduce complexity.
[0030] Generally speaking, application metadata can be used and pipelined through the yang model. The components of the RAN can be identified using their specific data and a related model whose structure consists of "affiliation" and "agent", where the affiliation defines the components associated with the data / model. For example, data can be obtained from a radio device and used to build a model whose output is adopted by the L1 CPU component; in such an example, the radio device and the L1 CPU are affiliated with each other.
[0031] Each affiliation can be associated with an agent for data generation. For example, in the above example, the radio device can be the agent for data generation. In another example, the fronthaul user plane or physical resource block (PRB) data can be affiliated with the DU, CU, and RU performance, thus acting as an agent for the fronthaul traffic management model effect. Generally speaking, the affiliation (DU / CU / RU) can refer to a model dedicated to the DU / CU / RU components, and the example affinity group (PRB) can refer to the overall model affinity group (PRB, cell site, cell, etc.). The agent (associated agent) can refer to any associated agent; there can be agent ID, agent communication API, agent function, agent database, and agent security classification data.
[0032] Machine learning model metadata allows any agent (e.g., RIC, CU, DU, etc.) to recreate and train the model locally and then use the model for inference. The model types can include but are not limited to regression / classification / clustering / time series / neural network (NN) / federated learning (FL) / reinforcement learning (RL), etc. Data metadata represents the model key data used to trigger the ingestion of data into ML training and can include but are not limited to describing the input feature pipeline, output statistics, data fields, data tables, data timestamps, and data pipeline API names. Training parameter metadata can include but are not limited to: describing the split ratio indicating the ratio of training data to validation data (e.g., 70 / 30, 60 / 40, 80 / 20); test k-fold (e.g., 2, 4, 6, 8, 10, 12): used for local model hyperparameter tuning; and testing the feedback loop parameters for the lag differential in the time series analysis model, such as autoregressive integrated moving average (ARIMA). Tuning metadata can include tuning hyperparameters, such as the number of trees, pdq (where: p is the number of autoregressive terms, d is the number of non-seasonal differences required for stationarity, q is the number of lagged prediction errors in the prediction equation), and the number of layers. Debugging metadata can include model debugging metadata with acute results, including the decisive threshold for RAN components. Validation metadata can include feature importance (FI) and the impact on the output (metadata) using plots.
[0033] Generally, the configuration (e.g., via a yang file) includes a complete pipeline for a specific application. Application metadata model configuration, status data, and administrative actions can be manipulated via the Network Configuration (NETCONF) protocol.
[0034] As described herein, in a first embodiment, in Figure 4 the techniques described herein are represented, Figure 4 an example representation sequence and data flow diagram of an embodiment is shown, starting from arrow one (1), where the model host 412 (in this example, the host network function is a near-real-time RIC) publishes its machine learning model capability data (via the SMO 416), and the machine learning model capability data shows host-specific information such as the supported model name, input KPIs (features), and output actions. As shown via box two (2) and arrow three (3), the operator 440 selects a machine learning model based on the model host's capability data and transmits the selected model for deployment by the model host.
[0035] The host 412 performs local training (e.g., periodically or otherwise triggered), as shown in box four (4). The local training uses actual local observation data and the training percentage indicated by the operator in the metadata via the O1 interface.
[0036] The host 412 evaluates the locally trained model based on the training percentage indicated in the metadata and compares the results with the error metric (e.g., mean squared error) also indicated by the operator 440 in the metadata and / or one or more quality of service (QoS) KPIs (e.g., handover success rate) specified by the operator.
[0037] If the model does not meet the value(s) specified by the operator in the metadata, remote training is required. This is generally represented in Figure 4 box 444 of. At arrow five (5), the host 412 sends its model capabilities to the operator 440 via the SMO 416 to indicate that remote training is needed.
[0038] As shown via box six (6) and arrow seven (7), the operator 440 trains the model and sends the updated hyperparameters back to the host 412 for further inference by the host 412. As part of the inference process, as shown in box eight (8), the host 412 calculates / estimates complexity data such as the time and resources required to collect each input feature (storage time and resources as well as processing time and resources) and its impact on the inference latency. As shown in box nine (9), in the case of computational resource congestion, the host 412 removes one or more less important features, where such less important features are indicated by the operator 440 in the metadata.
[0039] Similarly, as described herein, in another embodiment, in Figure 5 illustrates the techniques described herein. In this embodiment, multiple machine learning models are deployed and hosted at the network function / RAN agent (node) 512 for use during runtime and continuously / regularly evaluated based on RAN KPIs. In this example, the model host network function / RAN agent 512 can be a central unit. Model reselection occurs in cases of data anomalies or low RAN KPIs.
[0040] Thus, using Figure 5 the arrow labeled one (1) in Figure 3 which generally corresponds to arrow three (3) in Figure 5 the operator configures multiple machine learning models via the SMO 516 and deploys them to the RAN node 512, and can indicate the state of each machine learning model as active or deactivated. Only one machine learning model is activated at a time. Note that as shown via block 512 in
[0041] if not preselected for the RAN node 512, the RAN node 512 can select its own machine learning model (such as based on metadata sent with each model), e.g., by matching training parameter data with local observation data.
[0041] The operator 550 / SMO 516 can provide reselection criteria associated with each machine learning model to the RAN node 512. Examples include providing target RAN KPIs as criteria data for evaluation and / or processing requirements; (e.g., in the case of high CPU utilization, the RAN node should select a simpler model).
[0042] During runtime, the RAN node 512 evaluates the performance of the currently active machine learning model ( Figure 5 block 3 in
[0043] and triggers reselection (block 4) when appropriate. In one implementation, this occurs based on meeting (or not meeting) conditional criteria, e.g., including when the RAN KPI has dropped below a target value (such as below the handover success rate or user throughput metric).
[0044] When one or more reselection criteria are met, the RAN node may reselect another model (if available locally), or request a new model (or model group) from the SMO 516 / Operator 550, as Figure 5 shown by the arrow labeled seven (7) in
[0045] Such a request occurs when all available models (block 544) have been evaluated, such as based on a significant training data / observation data mismatch, or by a previously selected and tried model being determined to have low performance, or by a model being considered likely to fail (e.g., relatively high certainty). After receiving the request from the RAN node 512, the Operator 550 provides at least one new model for deployment via the SMO 516 and may remove one or more existing models. Note that the RAN node 512 may supply information explaining / suggesting the reason for the request in association with the request, such as a mismatch between observed values (e.g., including the actual observed value range) and training values (e.g., the range used in training). If information about the observed values is provided by the RAN node 512, the Operator 550 may provide the RAN node 512 with a model trained with a more appropriate value range.
[0046] One or more aspects may be embodied in a network device as represented in example operations such as Figure 6 and may include, for example, a memory storing computer-executable components and / or operations, and a processor executing the computer-executable components and / or operations stored in the memory. Example operations may include operation 602, which represents the network model host of the network device publishing machine learning model capability data of the network model host. Example operation 604 represents receiving a machine learning model and machine learning model data associated with the machine learning model in response to the publication, where the machine learning model data may include machine learning model metadata and machine learning parameter data. Example operation 606 represents deploying the machine learning model for network communication operations.
[0047] Further operations may include: training the machine learning model using local data to obtain an inference result; evaluating the inference result with respect to the information in the machine learning model metadata; and outputting a request for remote training in response to the evaluation of the inference result not meeting the information in the machine learning model metadata.
[0048] Further operations may include: receiving updated hyperparameter data, updated machine learning model metadata, and updated machine learning parameter data in response to the request; and retraining the machine learning model based on the updated hyperparameter data.
[0049] Further operations may include: estimating retraining complexity data prior to retraining; and in response to the complexity data exceeding complexity criterion data, removing at least one input feature from a set of input features related to retraining for collection based on feature importance data in the updated machine learning model metadata.
[0050] The network model host may include a radio access network intelligent controller of a network service management and orchestration device.
[0051] The network device may include a radio access network node.
[0052] The machine learning model metadata may include at least one of the following: model type data, training-related data, training parameter data, training parameter data, tuning metadata, debugging metadata, or validation metadata.
[0053] Receiving a machine learning model may include receiving the machine learning model as part of a set of received machine learning models, and further operations may include selecting the machine learning model for deploying the machine learning model.
[0054] The machine learning model may be a first machine learning model, and further operations may include: evaluating operation result data obtained after deploying the machine learning model relative to reselection criterion data obtained in connection with the set of received machine learning models; selecting a second machine learning model from the set of received machine learning models based on the operation result data and the reselection criterion data; and deploying the second machine learning model for network communication operations.
[0055] Further operations may include: determining that no machine learning model in the set of received machine learning models can satisfy the operation result data and the reselection criterion data; and based on the determination, requesting a different machine learning model not in the set of received machine learning models.
[0056] Further operations may include: generating feedback information associated with the request; the feedback information may include information to help obtain a model in the set of received machine learning models that can satisfy the operation result data and the reselection criterion data.
[0057] At Figure 7One or more example aspects are shown, such as example aspects corresponding to example operations of a method. Example operation 702 represents a network system including a processor publishing machine learning model capability data via a communication network. Example operation 704 represents the network system receiving a set of machine learning models and associated model reselection criterion data. Example operation 706 represents the network system operating a machine learning model in the set of machine learning models as an active machine learning model for network-related operations. Example operation 708 represents the network system evaluating the active machine learning model relative to the reselection criterion data to determine whether operating with the active machine learning model results in model reselection.
[0058] The active machine learning model can include a first machine learning model in the set, and further includes determining that the active machine learning model triggers a reselection operation based on the evaluation, and further includes: the network system performing a reselection operation to select a second machine learning model in the set; and the network system operating the second machine learning model in the set as the active machine learning model for network-related operations.
[0059] Further operations can include: the network system determining that each machine learning model in the set triggers a model reselection operation; and the network system requesting a different machine learning model that is not part of the set from the communication network to be used as the active machine learning model.
[0060] Further operations can include: the network system providing explanatory data associated with the request to help obtain a different machine learning model that is not part of the set.
[0061] Determining that each machine learning model in the set triggers a model reselection operation can include: operating each machine learning model in the set of machine learning models as an active machine learning model instance; and determining that each machine learning model instance triggers the associated model reselection criterion data.
[0062] Figure 8 Various example operations are outlined, for example, example operations corresponding to a machine-readable medium including executable instructions, where the executable instructions, when executed by a processor of a network model host, facilitate the execution of the operations. Example operation 802 represents publishing machine learning model capability data of a network model host. Example operation 804 represents receiving a machine learning model associated with machine learning model metadata and machine learning parameter data based on the publishing. Example operation 806 represents training the machine learning model using local data to obtain an inference result, where the local data is local to the network model host. Example operation 808 represents evaluating the inference result relative to criterion data in the machine learning model metadata.
[0063] Further operations may include: outputting a request for remote training in response to an evaluation of an inference result being determined to not meet criterion data.
[0064] The inference result may be a first inference result, the criterion data may include first criterion data, and further operations may include: receiving updated hyperparameter data, updated machine learning model metadata, and updated machine learning parameter data in response to the request; retraining the machine learning model based on the updated hyperparameter data to obtain a second inference result; and re-evaluating the second inference result against second criterion in the machine learning model metadata.
[0065] The updated machine learning model metadata may include feature importance data, and further operations may include: determining complexity data representative of the complexity of the retraining prior to retraining; and in response to the complexity data being determined to exceed complexity criterion data, removing input features from a set of retraining-related features based on the feature importance data.
[0066] It can be seen that the techniques described herein facilitate remote management of machine learning models. The inference model host can self-optimize model selection based on factors such as processing overhead and RAN KPIs. As a result of the model selection and format based on the inference host capabilities, the techniques described herein facilitate interoperability. The models can be deployed similarly to the current configuration management policies for CU RAN parameters and DU RAN parameters.
[0067] As a result of the techniques described herein, the existing hardware dependencies and software dependencies between both the model host and the operator are eliminated, thus providing the operator with greater flexibility in selecting RAN vendors and SMO vendors. The techniques described herein achieve autonomous RAN optimization based on actively training machine learning models, which improves the accuracy of real-time decision-making; the techniques adapt the selected machine learning model and its configuration according to the use cases, deployment types, and time-varying processing capabilities of network functions.
[0068] Figure 9 FIG. is a schematic block diagram of a computing environment 900 that can interact with the disclosed subject matter. System 900 includes one or more remote components 910. The (multiple) remote components 910 can be hardware and / or software (e.g., threads, processes, computing devices). In some embodiments, the (multiple) remote components 910 can be a distributed computer system connected via a communication framework 940 to local auto-scaling components and / or programs that utilize distributed computer system resources. The communication framework 940 can include wired network devices, wireless network devices, mobile devices, wearable devices, radio access network devices, gateway devices, femtocell devices, servers, etc.
[0069] System 900 also includes one or more local components 920. The local component(s) 920 can be hardware and / or software (e.g., threads, processes, computing devices). In some embodiments, the local component(s) 920 can include auto-scaling components and / or programs that transmit / use remote resources 910, etc., that are connected to a remotely located distributed computing system via a communication framework 940.
[0070] One possible communication between the remote component(s) 910 and the local component(s) 920 can be in the form of data packets adapted to be transmitted between two or more computer processes. Another possible communication between the remote component(s) 910 and the local component(s) 920 can be in the form of circuit-switched data adapted to be transmitted between two or more computer processes in a radio time slot. System 900 includes a communication framework 940, where the communication framework 940 can be used to facilitate communication between the remote component(s) 910 and the local component(s) 920, and can include an air interface, e.g., the Uu interface of a UMTS network via a Long-Term Evolution (LTE) network, etc. The remote component(s) 910 can be operatively connected to one or more remote data stores 950 that can be used to store information on the remote component(s) 910 side of the communication framework 940, such as hard disk drives, solid-state drives, SIM cards, device memories, etc. Similarly, the local component(s) 920 can be operatively connected to one or more local data stores 930 that can be used to store information on the local component(s) 920 side of the communication framework 940.
[0071] To provide additional context for the various embodiments described herein, Figure 10 and the following discussion is intended to provide a brief general description of a suitable computing environment 1000 in which the various embodiments described herein can be implemented. Although these embodiments have been described above in the general context of computer-executable instructions that can run on one or more computers, those skilled in the art will recognize that these embodiments can also be implemented in combination with other program modules and / or as a combination of hardware and software.
[0072] Generally, program modules include routines, programs, components, data structures, etc., that perform particular tasks or implement particular abstract data types. Moreover, those skilled in the art will understand that these methods can be practiced using other computer system configurations, including single-processor or multi-processor computer systems, minicomputers, mainframe computers, Internet of Things (IoT) devices, distributed computing systems, and personal computers, handheld computing devices, microprocessor-based or programmable consumer electronics, etc., each of which can be operatively coupled to one or more associated devices.
[0073] The illustrated embodiments in the examples of this document can also be practiced in a distributed computing environment where some tasks are performed by remote processing devices linked through a communication network. In a distributed computing environment, program modules can be located in local and remote memory storage devices.
[0074] Computing devices generally include various media, where the various media can include computer-readable storage media, machine-readable storage media, and / or communication media, and the two terms machine-readable storage media and / or communication media are used differently from each other in this document as follows. Computer-readable storage media or machine-readable storage media can be any available storage media accessible by a computer, and include both volatile and non-volatile media, as well as removable and non-removable media. By way of example and not limitation, computer-readable storage media or machine-readable storage media can be implemented in conjunction with any method or technology for storing information such as computer-readable or machine-readable instructions, program modules, structured data, or unstructured data, etc.
[0075] Computer-readable storage media can include, but are not limited to, random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), Blu-ray disc (BD) or other optical disc storage, magnetic tape cartridges, tapes, magnetic disk storage or other magnetic storage devices, solid state drives or other solid state storage devices, or other tangible and / or non-transitory media that can be used to store the required information. In this regard, when applied to a storage device, memory, or computer-readable medium in this document, the term "tangible" or "non-transitory" should be understood to exclude only propagating transitory signals themselves as a modifier, without relinquishing the right to storage devices, memories, or computer-readable media that meet all criteria other than just propagating transitory signals themselves.
[0076] Computer-readable storage media can be accessed by one or more local or remote computing devices (e.g., via an access request, query, or other data retrieval protocol) to perform various operations with respect to the information stored by the media.
[0077] A communication medium typically embodies computer-readable instructions, data structures, program modules, or other structured or unstructured data in a data signal, such as a modulated data signal, for example, a carrier wave or other transmission mechanism, and includes any information delivery or transmission medium. The term "(multiple) modulated data signal" refers to a signal whose one or more characteristics are set or changed in a manner that encodes information in one or more signals. By way of example and not limitation, communication media include wired media, such as a wired network or a direct wired connection, and wireless media, such as acoustic, RF (radio frequency), infrared, and other wireless media.
[0078] Referring again to Figure 10 , an example environment 1000 for implementing various embodiments of the aspects described herein includes a computer 1002, where the computer 1002 includes a processing unit 1004, a system memory 1006, and a system bus 1008. The system bus 1008 couples system components, including but not limited to the system memory 1006, to the processing unit 1004. The processing unit 1004 can be any of a variety of commercially available processors. Dual microprocessors and other multiprocessor architectures can also be used as the processing unit 1004.
[0079] The system bus 1008 can be any of several types of bus structures, which can further interconnect to a memory bus (with or without a memory controller), a peripheral bus, and a local bus using any of a variety of commercially available bus architectures. The system memory 1006 includes a ROM 1010 and a RAM 1012. The basic input / output system (BIOS) can be stored in a non-volatile memory such as ROM, erasable programmable read-only memory (EPROM), EEPROM, where the BIOS contains basic routines (such as during startup) that help to deliver information between elements within the computer 1002. The RAM 1012 can also include high-speed RAM, such as static RAM for caching data.
[0080] The computer 1002 also includes an internal hard disk drive (HDD) 1014 (e.g., EIDE, SATA), and can include one or more external storage devices 1016 (e.g., a magnetic floppy disk drive (FDD) 1016, a memory stick or flash drive reader, a memory card reader, etc.). Although the internal HDD 1014 is shown as being within the computer 1002, the internal HDD 1014 can also be configured for external use in a suitable chassis (not shown). Additionally, although not shown in the environment 1000, a solid-state drive (SSD) can be used in addition to, or in place of, the HDD 1014.
[0081] Other internal or external storage devices may include at least one other storage device 1020 having a storage medium 1022 (e.g., a solid state storage device, a non-volatile memory device, and / or an optical disk drive that can read from or write to removable media such as CD-ROM disks, DVDs, BDs, etc.). The external storage device 1016 may be facilitated by a network virtual machine. The HDD 1014, the external storage device(s) 1016, and the storage device (e.g., drive) 1020 may be connected to the system bus 1008 via an HDD interface 1024, an external storage interface 1026, and a drive interface 1028, respectively.
[0082] The drives and their associated computer-readable storage media provide non-volatile storage of data, data structures, computer-executable instructions, etc. For the computer 1002, the drives and storage media accommodate the storage of any data in a suitable digital format. Although the above description of computer-readable storage media refers to the corresponding types of storage devices, those skilled in the art should understand that other types of storage media that are computer-readable, whether currently available or developed in the future, may also be used in the exemplary operating environment, and further, any such storage media may contain computer-executable instructions for performing the methods described herein.
[0083] Many program modules may be stored in the drives and the RAM 1012, including an operating system 1030, one or more application programs 1032, other program modules 1034, and program data 1036. All or part of the operating system, applications, modules, and / or data may also be cached in the RAM 1012. The systems and methods described herein may be implemented using various commercially available operating systems or combinations of operating systems.
[0084] The computer 1002 may optionally include emulation technology. For example, a hypervisor (not shown) or other intermediary may emulate the hardware environment for the operating system 1030, and the emulated hardware may optionally be different from Figure 10 the hardware shown. In such an embodiment, the operating system 1030 may include one VM among multiple virtual machines (VMs) hosted on the computer 1002. Further, the operating system 1030 may provide a runtime environment for the application 1032, such as a Java runtime environment or a.NET framework. A runtime environment is a consistent execution environment that allows the application 1032 to run on any operating system that includes the runtime environment. Similarly, the operating system 1030 may support containers, and the application 1032 may be in the form of a container, i.e., a lightweight, independent executable software package that includes, for example, code for the application, runtime, system tools, system libraries, and settings.
[0085] Further, computer 1002 can be implemented to have a security module, such as a Trusted Platform Module (TPM). For example, in the case of having a TPM, the boot component hashes the next boot component over time before loading the next boot component and waits for the result to match a security value. Such a process can occur at any layer in the code execution stack of computer 1002. For example, an application is performed at the application execution level or the operating system (OS) kernel level, thereby implementing security at any code execution level.
[0086] The user can input commands and information into computer 1002 through one or more wired / wireless input devices (e.g., keyboard 1038, touch screen 1040, and a pointing device such as mouse 1042). Other input devices (not shown) can include a microphone, an infrared (IR) remote control, a radio frequency (RF) remote control or other remote controls, a joystick, a virtual reality controller and / or a virtual reality headset, a game pad, a stylus, an image input device (e.g., (a) camera), a gesture sensor input device, a vision motion sensor input device, an emotion or face detection device, a biometric input device (e.g., a fingerprint or iris scanner), etc. These and other input devices are often connected to the processing unit 1004 through an input device interface 1044, where the input device interface 1044 can be coupled to the system bus 1008, but can also be connected through other interfaces (such as a parallel port, an IEEE 1094 serial port, a game port, a USB port, an IR interface, interface, etc.).
[0087] A monitor 1046 or other type of display device can also be connected to the system bus 1008 via an interface such as a video adapter 1048. In addition to the monitor 1046, a computer generally also includes other peripheral output devices (not shown), such as speakers, printers, etc.
[0088] The computer 1002 can operate in a networked environment using logical connections via wired and / or wireless communication to one or more remote computers, such as the remote computer(s) 1050. The remote computer(s) 1050 can be a workstation, server computer, router, personal computer, portable computer, microprocessor-based entertainment appliance, peer device, or other common network node, and typically includes many or all of the elements described relative to the computer 1002, although only the memory / storage device 1052 is shown for simplicity. The depicted logical connections include wired / wireless connections to a local area network (LAN) 1054 and / or a larger network, such as a wide area network (WAN) 1056. Such LAN networking environments and WAN networking environments are common in offices and companies and facilitate enterprise-wide computer networks, such as intranets, all of which can be connected to a global communications network, such as the Internet.
[0089] When used in a LAN networking environment, the computer 1002 can be connected to the local area network 1054 via a wired and / or wireless communication network interface or adapter 1058. The adapter 1058 can facilitate wired or wireless communication to the LAN 1054, which can also include a wireless access point (AP) deployed thereon for communicating with the adapter 1058 in wireless mode.
[0090] When used in a WAN networking environment, the computer 1002 can include a modem 1060, or can be connected to a communication server on the WAN 1056 via other means for establishing communication on the WAN 1056 (such as via the Internet). The modem 1060 can be an internal or external wired or wireless device that can be connected to the system bus 1008 via the input device interface 1044. In a networked environment, program modules depicted relative to the computer 1002 or portions thereof can be stored in the remote memory / storage device 1052. It will be understood that the network connections shown are examples, and other means of establishing a communication link between computers can be used.
[0091] When used in a LAN networking environment or a WAN networking environment, in addition to, or instead of, the external storage device 1016 described above, the computer 1002 may also access a cloud storage system or other network-based storage systems. Generally, the connection between the computer 1002 and the cloud storage system may be established, for example, by an adapter 1058 or a modem 1060 through the LAN 1054 or the WAN 1056, respectively. After connecting the computer 1002 to an associated cloud storage system, the external storage interface 1026 may manage the storage devices provided by the cloud storage system with the assistance of the adapter 1058 and / or the modem 1060, just as it manages other types of external storage devices. For example, the external storage interface 1026 may be configured to provide access to cloud storage sources as if these sources were physically connected to the computer 1002.
[0092] The computer 1002 may be operable to communicate with any wireless device or entity operably disposed in a wireless communication (e.g., a printer, a scanner, a desktop and / or portable computer, a portable data assistant, a communication satellite, any device or location associated with a wirelessly detectable tag (e.g., a phone booth, a newsstand, a store shelf, etc.), and a telephone). This may include Wi-Fi and technologies. Thus, the communication may be a predefined structure like a traditional network, or merely an ad-hoc communication between at least two devices.
[0093] The foregoing description of the embodiments of the present disclosure that includes what is described in the abstract is not intended to be exhaustive or to limit the disclosed embodiments to the precise forms disclosed. While specific embodiments and examples have been described herein for purposes of illustration, those skilled in the relevant art will recognize that various modifications are possible within the scope of such embodiments and examples.
[0094] In this regard, while the disclosed subject matter has been described in connection with various embodiments and the corresponding drawings, it should be understood that, where applicable, other similar embodiments may be used, or modifications and additions may be made to the described embodiments so as to perform the same, similar, alternative, or substitute functions of the disclosed subject matter without departing therefrom. Accordingly, the disclosed subject matter should not be limited to any single embodiment described herein, but should be construed in accordance with the breadth and scope of the appended claims.
[0095] As used in the subject specification, the term "processor" can refer to substantially any computing processing unit or device, including but not limited to: a single-core processor; a single processor with software multithreading execution capabilities; a multi-core processor; a multi-core processor with software multithreading execution capabilities; a multi-core processor utilizing hardware multithreading technology; a parallel platform; and a parallel platform with distributed shared memory. Additionally, a processor can refer to an integrated circuit, an application specific integrated circuit, a digital signal processor, a field programmable gate array, a programmable logic controller, a complex programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A processor can utilize nanoscale architectures (such as but not limited to molecule and quantum dot-based transistors, switches, and gates) in order to optimize space usage or enhance the performance of a user device. A processor can also be implemented as a combination of computing processing units.
[0096] As used in this application, the terms "component", "system", "platform", "layer", "selector", "interface", etc. are intended to refer to a computer-related entity or an entity related to an operating device having one or more specific functions, where the entity can be hardware, a combination of hardware and software, software, or software in execution. By way of example, a component can be but is not limited to a process running on a processor, a processor, an object, an executable program, an execution thread, a program, and / or a computer. By way of illustration and not limitation, an application running on a server and the server can both be components. One or more components can reside within a process and / or an execution thread, and a component can be located on one computer and / or distributed between two or more computers. Additionally, these components can execute from various computer-readable media on which various data structures are stored. These components can communicate, such as via a signal having one or more data packets, through local and / or remote processes (e.g., data from one component interacts with another component in a local system, a distributed system, and / or via the signal interacts with other systems through a network such as the Internet). As another example, a component can be a device having a specific function provided by a mechanical component operated by an electrical or electronic circuit and operated by a software or firmware application executed by a processor, where the processor can be internal or external to the device and executes at least a portion of the software or firmware application. As yet another example, a component can be a device that provides a specific function through an electronic component without mechanical components, where the electronic component can include a processor therein to execute at least a portion of the software or firmware that imparts the function to the electronic component.
[0097] In addition, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified or clear from the context, "X employs A or B" is intended to mean any natural inclusive permutation. That is, if X employs A; X employs B; or X employs both A and B, then "X employs A or B" is satisfied in any of the foregoing instances.
[0098] Although these embodiments are susceptible to various modifications and alternative constructions, certain illustrated embodiments thereof are shown in the drawings and have been described in detail above. However, it should be understood that the various embodiments are not intended to be limited to the specific forms disclosed, but rather, the invention is intended to cover all modifications, alternative constructions, and equivalents falling within the spirit and scope.
[0099] In addition to the various embodiments described herein, it should be understood that other similar embodiments may be used, or modifications and additions may be made to the described embodiments in order to perform the same or equivalent functions of the corresponding embodiments without departing therefrom. Further still, the performance of one or more of the functions described herein may be shared by multiple processing chips or multiple devices, and similarly, storage may be implemented among multiple devices. Accordingly, the various embodiments are not limited to any single embodiment, but rather are construed in accordance with the breadth, spirit, and scope of the appended claims.
Claims
1. A system, comprising: a processor; and a memory that stores executable instructions which, when executed by the processor, facilitate performance of operations, the operations including: publishing, by a network model host of a network device, machine learning model capability data of the network model host; receiving, in response to the publishing, a machine learning model and machine learning model data associated with the machine learning model, the machine learning model data including machine learning model metadata and machine learning parameter data; and deploying the machine learning model for network communication operations.
2. The system according to claim 1, wherein the operation further comprises: Training the machine learning model using local data to obtain an inference result; Evaluating the inference result relative to information in the machine learning model metadata; and outputting a request for remote training in response to the evaluation of the inference result not satisfying the information in the machine learning model metadata.
3. The system according to claim 2, wherein the operation further comprises: Receiving updated hyperparameter data, updated machine learning model metadata, and updated machine learning parameter data in response to the request; and retraining the machine learning model based on the updated hyperparameter data.
4. The system according to claim 3, wherein the operation further comprises: Before the retraining, estimating complexity data of the retraining; and in response to the complexity data exceeding complexity criterion data, removing at least one input feature from a set of input features related to the retraining for collection based on feature importance data in the updated machine learning model metadata.
5. The system according to claim 1, wherein the network model host includes a radio access network intelligent controller of a network service management and orchestration device.
6. The system according to claim 1, wherein the network device includes a radio access network node.
7. The system according to claim 1, wherein the machine learning model metadata includes at least one of the following: model type data, training related data, training parameter data, training parameter data, tuning metadata, debugging metadata, or validation metadata.
8. The system according to claim 1, wherein the receiving of the machine learning model includes receiving the machine learning model as part of a set of received machine learning models, and wherein the operations further include selecting the machine learning model for the deployment of the machine learning model.
9. The system according to claim 8, wherein the machine learning model is a first machine learning model, and wherein the operation further comprises: Evaluating operation result data obtained after the deployment of the machine learning model relative to reselection criterion data obtained in connection with the set of received machine learning models; selecting a second machine learning model from the set of received machine learning models based on the operation result data and the reselection criterion data; and deploying the second machine learning model for the network communication operations.
10. The system according to claim 9, wherein the operation further comprises: Determining that no machine learning model in the set of received machine learning models can satisfy the operation result data and the reselection criterion data; and requesting a different machine learning model not in the set of received machine learning models based on the determination.
11. The system according to claim 10, wherein the operation further includes generating feedback information associated with the request, the feedback information including information for assisting in obtaining a model from the set of received machine learning models that can satisfy the operation result data and the reselection criterion data.
12. A method includes: publishing, by a network system including a processor, machine learning model capability data via a communication network; receiving, by the network system, a set of machine learning models and associated model reselection criterion data; operating, by the network system, a machine learning model from the set of machine learning models as an active machine learning model for an operation related to the network; and evaluating, by the network system, the active machine learning model against the reselection criterion data to determine whether operating with the active machine learning model results in model reselection.
13. The method according to claim 12, wherein the active machine learning model comprises a first machine learning model in the group, and the method further comprises determining, based on the evaluation, that the active machine learning model triggers a reselection operation, and further comprises: performing, by the network system, the reselection operation to select a second machine learning model from the set; and operating, by the network system, the second machine learning model from the set as the active machine learning model for an operation related to the network.
14. The method according to claim 13, further comprising: determining, by the network system, that each machine learning model from the set triggers the model reselection operation; and requesting, by the network system, from the communication network a different machine learning model that is not part of the set to be used as the active machine learning model.
15. The method according to claim 14 further comprises: providing, by the network system, explanatory data associated with the request to assist in obtaining the different machine learning model that is not part of the set.
16. The method according to claim 14, wherein determining that each machine learning model in the group triggers the model reselection operation comprises: operating each machine learning model from the set of machine learning models as an active machine learning model instance; and determining that each machine learning model instance triggers the associated model reselection criterion data.
17. A non - transitory machine - readable medium includes executable instructions that, when executed by a processor of a network model host, facilitate performance of operations including: publishing the machine learning model capability data of the network model host; receiving, based on the publishing, a machine learning model associated with machine learning model metadata and machine learning parameter data; training the machine learning model using local data to obtain an inference result, wherein the local data is local to the network model host; and evaluating the inference result against criterion data in the machine learning model metadata.
18. The non-transitory machine-readable medium according to claim 17, wherein the operation further comprises: outputting a request for remote training in response to determining that the evaluation of the inference result does not satisfy the criterion data.
19. The non - transitory machine - readable medium according to claim 18, wherein the inference result is a first inference result, wherein the criterion data includes first criterion data, and wherein the operation further includes: receiving updated hyperparameter data, updated machine learning model metadata, and updated machine learning parameter data in response to the request; retraining the machine learning model based on the updated hyperparameter data to obtain a second inference result, and Re-evaluate the second inference result relative to a second criterion in the machine learning model metadata.
20. The non-transitory machine-readable medium according to claim 19, wherein the updated machine learning model metadata includes feature importance data, and wherein the operation further comprises: Before the retraining, determine complexity data indicative of the complexity of the retraining; and in response to the complexity data being determined to exceed complexity criterion data, remove input features from a set of retraining-related features based on the feature importance data.