Methods and apparatus for ai / ml model monitoring
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2024-02-14
- Publication Date
- 2026-05-06
AI Technical Summary
Current AI/ML model monitoring systems lack effective mechanisms for selecting appropriate metrics and providing timely recommendations to address performance degradation and reliability issues in dynamic telecommunications networks, particularly in 5G and beyond systems.
A novel monitoring mechanism that categorizes metrics into stability, performance, and operational categories, allowing for the selection of relevant metrics based on AI/ML model type, use case, and environmental context, and provides recommendations for updating, retraining, or switching models to maintain reliability and performance.
This approach enables continuous monitoring and improvement of AI/ML models, ensuring optimal performance and reliability by detecting anomalies and providing tailored recommendations for maintaining or updating models, thereby enhancing the overall success of AI/ML-powered networks.
Smart Images

Figure KR2024002092_22082024_PF_FP
Abstract
Description
METHODS AND APPARATUS FOR AI / ML MODEL MONITORING
[0001] Certain examples of the present disclosure relate to methods, apparatus and / or systems for monitoring metrics and / or KPIs for AI / ML models. Further, certain examples of the present disclosure relate to methods and apparatus for selecting metrics to monitor based on the AI / ML model, or information thereon, and for recommending or identifying solutions to address an issue identified in the AI / ML model through monitoring of the selected metrics (e.g., by identifying abnormal behaviour of a selected metric). Furthermore, in certain examples of the present disclosure, the monitored metrics are each one of a stability metric, a performance metric and an operational metric; and the metrics include at least one metric from each of these categories.
[0002] At the beginning of the development of 5G mobile communication technologies, in order to support services and to satisfy performance requirements in connection with enhanced Mobile BroadBand (eMBB), Ultra Reliable Low Latency Communications (URLLC), and massive Machine-Type Communications (mMTC), there has been ongoing standardization regarding beamforming and massive MIMO for mitigating radio-wave path loss and increasing radio-wave transmission distances in mmWave, supporting numerologies (for example, operating multiple subcarrier spacings) for efficiently utilizing mmWave resources and dynamic operation of slot formats, initial access technologies for supporting multi-beam transmission and broadbands, definition and operation of BWP (BandWidth Part), new channel coding methods such as a LDPC (Low Density Parity Check) code for large amount of data transmission and a polar code for highly reliable transmission of control information, L2 pre-processing, and network slicing for providing a dedicated network specialized to a specific service.
[0003] Currently, there are ongoing discussions regarding improvement and performance enhancement of initial 5G mobile communication technologies in view of services to be supported by 5G mobile communication technologies, and there has been physical layer standardization regarding technologies such as V2X (Vehicle-to-everything) for aiding driving determination by autonomous vehicles based on information regarding positions and states of vehicles transmitted by the vehicles and for enhancing user convenience, NR-U (New Radio Unlicensed) aimed at system operations conforming to various regulation-related requirements in unlicensed bands, NR UE Power Saving, Non-Terrestrial Network (NTN) which is UE-satellite direct communication for providing coverage in an area in which communication with terrestrial networks is unavailable, and positioning.
[0004] Moreover, there has been ongoing standardization in air interface architecture / protocol regarding technologies such as Industrial Internet of Things (IIoT) for supporting new services through interworking and convergence with other industries, IAB (Integrated Access and Backhaul) for providing a node for network service area expansion by supporting a wireless backhaul link and an access link in an integrated manner, mobility enhancement including conditional handover and DAPS (Dual Active Protocol Stack) handover, and two-step random access for simplifying random access procedures (2-step RACH for NR). There also has been ongoing standardization in system architecture / service regarding a 5G baseline architecture (for example, service based architecture or service based interface) for combining Network Functions Virtualization (NFV) and Software-Defined Networking (SDN) technologies, and Mobile Edge Computing (MEC) for receiving services based on UE positions.
[0005] As 5G mobile communication systems are commercialized, connected devices that have been exponentially increasing will be connected to communication networks, and it is accordingly expected that enhanced functions and performances of 5G mobile communication systems and integrated operations of connected devices will be necessary. To this end, new research is scheduled in connection with eXtended Reality (XR) for efficiently supporting AR (Augmented Reality), VR (Virtual Reality), MR (Mixed Reality) and the like, 5G performance improvement and complexity reduction by utilizing Artificial Intelligence (AI) and Machine Learning (ML), AI service support, metaverse service support, and drone communication.
[0006] Furthermore, such development of 5G mobile communication systems will serve as a basis for developing not only new waveforms for providing coverage in terahertz bands of 6G mobile communication technologies, multi-antenna transmission technologies such as Full Dimensional MIMO (FD-MIMO), array antennas and large-scale antennas, metamaterial-based lenses and antennas for improving coverage of terahertz band signals, high-dimensional space multiplexing technology using OAM (Orbital Angular Momentum), and RIS (Reconfigurable Intelligent Surface), but also full-duplex technology for increasing frequency efficiency of 6G mobile communication technologies and improving system networks, AI-based communication technology for implementing system optimization by utilizing satellites and AI (Artificial Intelligence) from the design stage and internalizing end-to-end AI support functions, and next-generation distributed computing technology for implementing services at levels of complexity exceeding the limit of UE operation capability by utilizing ultra-high-performance communication and computing resources.
[0007] The content of the following documents is referred to below and / or their content provides background information and context that the following disclosure should be considered in view of:
[0008] [1] 3GPP TS 22.261;
[0009] 3rdGeneration Partnership Project; Technical Specification Group Services and System Aspects; Service requirements for the 5G system; SA1;
[0010] Release 18 (e.g., V18.8.0) / Release 19 (e.g., V19.1.0);
[0011] [2] RP-213599, 3GPP TSG RAN Meeting #94e;
[0012] Study on Artificial Intelligence (AI) / Machine Learning (ML) for NR Air Interface.
[0013] [3] 3GPP TS 38.413;
[0014] 3rd Generation Partnership Project; Technical Specification Group Radio Access Network; NG-RAN; NG Application Protocol (NGAP);
[0015] Release 17 (e.g., V17.0.0).
[0016] [4] 3GPP TS 38.423;
[0017] 3rd Generation Partnership Project; Technical Specification Group Radio Access Network; NG-RAN; Xn application protocol (XnAP);
[0018] Release 17 (e.g., V17.3.0).
[0019] [5] 3GPP TS 38.331;
[0020] 3rd Generation Partnership Project; Technical Specification Group Radio Access Network; NR; Radio Resource Control (RRC) protocol specification
[0021] Release 17 (e.g., V17.3.0).
[0022] [6] Sharifi, Sepehr, et al. "Identifying the Hazard Boundary of ML-enabled Autonomous Systems Using Cooperative Co-Evolutionary Search." arXiv preprint arXiv:2301.13807 (2023).
[0023] (Note: the example versions shown for each TS are non-limiting, other versions of the TS may be considered also)
[0024] The present disclosure relates to wireless communication systems and, more specifically, the invention relates to methods and apparatus for AI / ML model monitoring.
[0025] It is an aim of certain examples of the present disclosure to address, solve and / or mitigate, at least partly, at least one of the problems and / or disadvantages associated with the related art, for example at least one of the problems and / or disadvantages described herein. It is an aim of certain examples of the present disclosure to provide at least one advantage over the related art, for example at least one of the advantages described herein.
[0026] According to an aspect of the present disclosure, there is provided a first entity for monitoring an artificial intelligence / machine learning (AI / ML) model deployable in a network, the first entity comprising: a transmitter; a receiver; and at least one processor configured to: obtain information relating to the AI / ML model; and based on the information relating to the AI / ML model, identify one or more metric for monitoring the AI / ML model from among a plurality of metrics categorized based on one or more of the AI / ML model, a use case for the AI / ML model, functionality of the AI / ML model, interdependence between metrics, or data obtained for monitoring; wherein each metric of the plurality of metrics is included in one or more of the plurality of categories.
[0027] According to various examples, identifying the one or more metric comprises: based on the information relating to the AI / ML model, identifying at least one first metric included in a first category of the plurality of categories and at least one second metric included in a second category of the plurality of categories.
[0028] According to various examples, identifying the one or more metric comprises: based on the information relating to the AI / ML model, identifying at least one metric for each of the plurality of categories.
[0029] According to various examples, the plurality of categories comprises one or more of: a category including metrics relating to the ability of the AI / ML model to maintain consistent results over time; a category including metrics relating to accuracy and effectiveness of outputs of the AI / ML model; and a category including metrics relating to operational behaviour of a system in which the AI / ML model is deployed.
[0030] According to various examples, one or more of: the category including metrics relating to the ability of the AI / ML model to maintain consistent results over time comprises at least one of: a metric for identifying a prior probability shift, Population Stability Index (PSI), Divergence Index, Maximum Mean Discrepancy (MMD), Kullback-Leibler (KL) Divergence, Energy Distance (ED), Cramer-von Mises (CvM) Test, Anderson-Darling Test, a metric for identifying a covariate shift, Characteristic Stability Index (CSI), or Wasserstein's Distance; the category including metrics relating to accuracy and effectiveness of outputs of the AI / ML model comprises one or more of: Root Mean Squared Error (RMSE), R-Square, F1-Score, Mean Squared Error (MSE), Area Under the Receiver Operating Characteristic curve (AUC-ROC), Cross-Entropy Loss, Kullback-Leibler Divergence Loss, Hit Rate, Diversity, Novelty, or a metric for identifying concept shift; and the category including metrics relating to operational behaviour of a system in which the AI / ML model is deployed comprises one or more of: ML throughput, ML Latency, Resource Usage, Power Consumption, Scalability, Availability, a metric for identifying system health issues, a metric for identifying endpoint latency, or a metric for identifying input / output (I / O), memory, and / or CPU issues.
[0031] According to various examples, the at least one processor is further configured to: transmit first information on the identified one or more metric to a second entity; and receive second information on at least one selected metric, wherein the at least one selected metric is included among the one or more metric.
[0032] According to various examples, the at least one processor is further configured to: monitor the at least one selected metric; and upon detecting at least one monitored metric to exhibit unusual behaviour, provide a recommendation or carry out the recommendation.
[0033] According to various examples, the at least one monitored metric is detected to exhibit unusual behaviour when it is detected, by the at least one processor, that: measurements associated with the at least one monitored metric fall outside of expected ranges, or are above / below specified threshold(s), the at least one monitored metric is underperforming, or the at least one monitored metric diverges from expected results.
[0034] According to various examples, the recommendation comprises one of: updating the AI / ML model, retraining the AI / ML model, switching to another AI / ML model, maintaining deployment of the AI / ML model in the network, deactivating the AI / ML model, or implementing a fallback solution.
[0035] According to various examples, one of: if the at least one monitored metric exhibiting unusual behaviour includes only metric(s) relates to accuracy and effectiveness of outputs of the AI / ML model, the at least one processor is configured to provide a recommendation to retrain the AI / ML model to the second entity; or if the at least one monitored metric exhibiting unusual behaviour includes a metric relating to accuracy and effectiveness of outputs of the AI / ML model and a metric relating to the ability of the AI / ML model to maintain consistent results over time, the at least one processor is configured to provide a recommendation to implement a fallback solution for ensuring safety and reliability of a system in which the AI / ML model is deployed and update, or a recommendation to update the AI / ML model.
[0036] According to various examples, the recommendation is provided to the second entity.
[0037] According to various examples, the information relating to the AI / ML model is obtained from the second entity.
[0038] According to various examples, the at least one processor is configured to: identify a proposed monitoring frequency for each of the one or more metric based on third information; and include the proposed monitoring frequency in the first information; and wherein the second information comprises a monitoring frequency for each of the at least one selected metric. For example, the third information is information configured in the first entity or information signalled to the first entity by another entity.
[0039] According to various examples, the information relating to the AI / ML model indicates one or more key performance indicator (KPI) for the AI / ML model; and / or wherein the AI / ML model is deployed at, or to be deployed at, the second entity and / or a third entity in the network.
[0040] According to various examples, the at least one processor is configured to: identify the one or more metric based on a type of the AI / ML model and / or an application being used in associated with the AI / ML model as indicated by the information relating to the AI / ML model.
[0041] According to various examples, the one or more identified metric comprises at least one recommended monitoring metric identified, by the at least one processor, based on one or more of the AI / ML model, expected performance of the AI / ML model or environmental context.
[0042] According to another aspect of the present disclosure, there is provided a second entity for managing an artificial intelligence / machine learning (AI / ML) model deployable in a network, the second entity comprising: a transmitter; a receiver; and at least one processor configured to: receive, from the first entity, first information on one or more metric identified by the first entity; select at least one metric to be monitored from among the one or more metric; and transmit, to the first entity, second information on the at least one selected metric; wherein each of the one or more metric is included in one or more of a plurality of categories, having been categorised based on one or more of the AI / ML model, a use case for the AI / ML model, functionality of the monitored AI / ML model, interdependence between metrics, or data obtained for monitoring.
[0043] According to various examples, the plurality of categories comprises one or more of: a category including metrics relating to the ability of the AI / ML model to maintain consistent results over time; a category including metrics relating to accuracy and effectiveness of outputs of the AI / ML model; and a category including metrics relating to operational behaviour of a system in which the AI / ML model is deployed.
[0044] According to various examples, one or more of: the category including metrics relating to the ability of the AI / ML model to maintain consistent results over time comprises at least one of: a metric for identifying a prior probability shift, Population Stability Index (PSI), Divergence Index, Maximum Mean Discrepancy (MMD), Kullback-Leibler (KL) Divergence, Energy Distance (ED), Cramer-von Mises (CvM) Test, Anderson-Darling Test, a metric for identifying a covariate shift, Characteristic Stability Index (CSI), or Wasserstein's Distance; the category including metrics relating to accuracy and effectiveness of outputs of the AI / ML model comprises one or more of: Root Mean Squared Error (RMSE), R-Square, F1-Score, Mean Squared Error (MSE), Area Under the Receiver Operating Characteristic curve (AUC-ROC), Cross-Entropy Loss, Kullback-Leibler Divergence Loss, Hit Rate, Diversity, Novelty, or a metric for identifying concept shift; and the category including metrics relating to operational behaviour of a system in which the AI / ML model is deployed comprises one or more of: ML throughput, ML Latency, Resource Usage, Power Consumption, Scalability, Availability, a metric for identifying system health issues, a metric for identifying endpoint latency, or a metric for identifying input / output (I / O), memory, and / or CPU issues.
[0045] According to various examples, the at least one processor is configured to: receive, from the first entity, a recommendation relating to unusual behaviour exhibited by at least one metric, among the at least one selected metric, being monitored at the first entity; and carry out the recommendation.
[0046] According to various examples, the recommendation comprises one of: updating the AI / ML model, retraining the AI / ML model, switching to another AI / ML model, maintaining deployment of the AI / ML model in the network, deactivating the AI / ML model, or implementing a fallback solution.
[0047] According to various examples, the first information comprises a proposed monitoring frequency for each of the one or more metric; and wherein the at least one processor is configured to: select a monitoring frequency for each of the at least one selected metric; and include the selected monitoring frequency for each of the at least one selected metric in the second information.
[0048] According to various examples, the at least one processor is configured to transmit information relating to the AI / ML model to the first entity, prior to receiving the first information; and / or wherein the information relating to the AI / ML model indicates one or more key performance indicator (KPI) for the AI / ML model; and / or wherein the AI / ML model is deployed at, or to be deployed at, the second entity and / or a third entity in the network.
[0049] According to another aspect, there is provided a method of a first entity for monitoring an artificial intelligence / machine learning (AI / ML) model deployable in a network, the method comprising: obtaining information relating to the AI / ML model; and based on the information relating to the AI / ML model, identifying one or more metric for monitoring the AI / ML model from among a plurality of metrics categorized based on one or more of the AI / ML model, a use case for the AI / ML model, functionality of the AI / ML model, interdependence between metrics, or data obtained for monitoring; wherein each metric of the plurality of metrics is included in one or more of the plurality of categories.
[0050] According to another aspect of the present disclosure, there is provided a method of a second entity for managing an artificial intelligence / machine learning (AI / ML) model deployable in a network, the method comprising: receiving, from the first entity, first information on one or more metric identified by the first entity; selecting at least one metric to be monitored from among the one or more metric; and transmitting, to the first entity, second information on the at least one selected metric; wherein each of the one or more metric is included in one or more of a plurality of categories, having been categorised based on one or more of the AI / ML model, a use case for the AI / ML model, functionality of the monitored AI / ML model, interdependence between metrics, or data obtained for monitoring.
[0051] According to other aspects of the present disclosure, there is provided a computer-readable data carrier having stored thereon a computer program which, when executed by at least one processor of an electronics device, causes the electronic device to perform a method according to one of those described above or elsewhere herein.
[0052] According to other aspects of the present disclosure, there is provided a network comprising a first entity according to any of the examples or aspects disclosed and / or a second entity according to any of the examples or aspects disclosed above.
[0053] It will be appreciated that the present disclosure envisages and includes all possible combinations of the above aspects and examples.
[0054] Other aspects, advantages, and salient features of the invention will become apparent to those skilled in the art from the following detailed description taken in conjunction with the accompanying drawings.
[0055] In one embodiment, a first entity for monitoring an artificial intelligence / machine learning (AI / ML) model deployable in a network, the first entity comprising: a transceiver; and a processor configured to: obtain information relating to the AI / ML model, and identify one or more metric for monitoring the AI / ML model from among a plurality of metrics categorized based on one or more of the AI / ML model, a use case for the AI / ML model, functionality of the AI / ML model, interdependence between metrics, or data obtained for monitoring, based on the information relating to the AI / ML model, wherein each metric of the plurality of metrics is included in one or more of the plurality of categories.
[0056] In another embodiment, a second entity for managing an artificial intelligence / machine learning (AI / ML) model deployable in a network, the second entity comprising: a transceiver; and the processor configured to: receive, from the first entity, first information on one or more metric identified by the first entity, select at least one metric to be monitored from among the one or more metric, and transmit, to the first entity, second information on the at least one selected metric; wherein each of the one or more metric is included in one or more of a plurality of categories, having been categorised based on one or more of the AI / ML model, a use case for the AI / ML model, functionality of the monitored AI / ML model, interdependence between metrics, or data obtained for monitoring.
[0057] In yet another embodiment, a method of first entity for monitoring an artificial intelligence / machine learning (AI / ML) model deployable in a network, the method comprising: obtaining information relating to the AI / ML model, and identifying one or more metric for monitoring the AI / ML model from among a plurality of metrics categorized based on one or more of the AI / ML model, a use case for the AI / ML model, functionality of the AI / ML model, interdependence between metrics, or data obtained for monitoring, based on the information relating to the AI / ML model, wherein each metric of the plurality of metrics is included in one or more of the plurality of categories.
[0058] In yet another embodiment, a method of second entity for managing an artificial intelligence / machine learning (AI / ML) model deployable in a network, the method comprising: receiving, from the first entity, first information on one or more metric identified by the first entity, selecting at least one metric to be monitored from among the one or more metric, and transmitting, to the first entity, second information on the at least one selected metric; wherein each of the one or more metric is included in one or more of a plurality of categories, having been categorised based on one or more of the AI / ML model, a use case for the AI / ML model, functionality of the monitored AI / ML model, interdependence between metrics, or data obtained for monitoring.
[0059] Advantages, and salient features of the invention will become apparent to those skilled in the art from the following detailed description, which, taken in conjunction with the annexed drawings, discloses exemplary embodiments of the invention.
[0060] According to various embodiments of the present disclosure, method and apparatus for AI / ML data collection and model monitoring in communications networks.
[0061] Embodiments / examples of the present disclosure are further described hereinafter with reference to the accompanying drawings, in which:
[0062] Figure 1 is an illustrated of three categories of metric described in the present disclosure.
[0063] Figure 2 is a flow diagram illustrating a method (e.g., a monitoring method) in accordance with various examples of the present disclosure.
[0064] Figure 3 is a flow diagram illustrating a method (e.g., an action recommendation method) in accordance with various examples of the present disclosure.
[0065] Figure 4 is a block diagram illustrating an example structure of a network entity in accordance with certain examples of the present disclosure.
[0066] Figure 5 is a representation of continuous delivery for ML end-to-end process.
[0067] Wireless or mobile (cellular) communications networks in which a mobile terminal (e.g., user equipment (UE), such as a mobile handset) communicates via a radio link with a network of base stations, or other wireless access points or nodes, have undergone rapid development through a number of generations. The 3rdGeneration Partnership Project (3GPP) design, specify and standardise technologies for mobile wireless communication networks. Fourth Generation (4G) and Fifth Generation (5G) systems are now widely deployed.
[0068] 3GPP standards for 4G systems include an Evolved Packet Core (EPC) and an Enhanced-UTRAN (E-UTRAN: an Enhanced Universal Terrestrial Radio Access Network). The E-UTRAN uses Long Term Evolution (LTE) radio technology. LTE is commonly used to refer to the whole system including both the EPC and the E-UTRAN, and LTE is used in this sense in the remainder of this document. LTE should also be taken to include LTE enhancements such as LTE Advanced and LTE Pro, which offer enhanced data rates compared to LTE.
[0069] In 5G systems a new air interface has been developed, which may be referred to as 5G New Radio (5G NR) or simply NR. NR is designed to support the wide variety of services and use case scenarios envisaged for 5G networks, though builds upon established LTE technologies. New frameworks and architectures are also being developed as part of 5G networks in order to increase the range of functionality and use cases available through 5G networks.
[0070] In recent years, autonomous systems that incorporate machine learning (ML) solutions have become more prevalent in telecommunication networks, performing tasks such as prediction, planning, control, etc. These systems are fundamentally different from conventional software components and present novel challenges and safety risks that are not manageable through traditional software engineering practices. The main reason for this difference is that the logic of these ML solutions is not defined by source code or specifications, but rather it is determined by the training process, the training data, and the input data at inference time.
[0071] A new framework developed as part of 5G networks (and beyond) is the use of artificial intelligence / machine learning (AI / ML), which may be used for the optimisation of the operation of 5G networks. In AI / ML operation, AI / ML models and / or data might be transferred across the AI / ML applications (e.g., application functions (AFs)), 5GC (5G core), UEs (user equipments) etc.). Without limitation, the AI / ML works could be divided into two main phases: model training and inference. During model training and inference, multiple rounds of interaction may be required.
[0072] In Section 6.40 ('AI / ML model transfer in 5GS') in TS 22.261 [1], three types of AI / ML operations to be supported in Release 18 and in Release 19 are described as follows:
[0073] a)AI / ML operation splitting between AI / ML endpoints
[0074] The AI / ML operation / model is split into multiple parts according to the current task and environment. The intention is to offload the computation-intensive, energy-intensive parts to network endpoints, whereas leave the privacy-sensitive and delay-sensitive parts at the end device. The device executes the operation / model up to a specific part / layer and then sends the intermediate data to the network endpoint. The network endpoint executes the remaining parts / layers and feeds the inference results back to the device.
[0075] b)AI / ML model / data distribution and sharing over 5G system
[0076] Multi-functional mobile terminals might need to switch the AI / ML model in response to task and environment variations. The condition of adaptive model selection is that the models to be selected are available for the mobile device. However, given the fact that the AI / ML models are becoming increasingly diverse, and with the limited storage resource in a UE, it can be determined to not pre-load all candidate AI / ML models on-board. Online model distribution (i.e., new model downloading) is needed, in which an AI / ML model can be distributed from a NW (network) endpoint to the devices when they need it to adapt to the changed AI / ML tasks and environments. For this purpose, the model performance at the UE needs to be monitored constantly.
[0077] c)Distributed / Federated Learning over 5G system
[0078] The cloud server trains a global model by aggregating local models partially-trained by each end devices. Within each training iteration, a UE performs the training based on the model downloaded from the AI server using the local training data. Then the UE reports the interim training results to the cloud server via 5G UL channels. The server aggregates the interim training results from the UEs and updates the global model. The updated global model is then distributed back to the UEs and the UEs can perform the training for the next iteration.
[0079] Model monitoring refers to the control and evaluation of the performance of a ML model to determine whether or not it is operating properly and efficiently. When the AI / ML model experiences some performance decay, appropriate maintenance actions should be taken to restore performance. AI / ML analytics and AI / ML model outputs are expected to drive most strategic decisions in Beyond 5G (B5G) and 6thGeneration (6G) application functions and network orchestration entities. However, the performance of AI / ML models degrades over time. This can lead to non-optimal decisions, which may lead to performance degradation and reduction in quality of experience (QoE), and thus profit or revenue may decline.
[0080] To prevent such negative effects, network operators should closely track AI / ML model performance and metrics on their networks (e.g., refer to [6]). Consequently, a monitoring mechanism should be carefully crafted in order to select the right metrics and key performance indicators (KPIs) to report and track. The primary motivation for a model monitoring method or apparatus is to create an effective feedback loop post-deployment back to the model building phase (an example of this is illustrated in Figure 5). This helps to improve the deployed AI / ML model by deciding whether to update the model, continue with the existing model, implement a fallback solution, etc. To enable this decision, the monitoring mechanism should track and report various metrics belonging to different domains or categories. In addition, acceptable performance, model risks, and costs of errors vary across use cases.
[0081] In Figure 5, the numerals define the following:
[0082] · 501 - Training code;
[0083] · 503 - Training data;
[0084] · 505 - Candidate models;
[0085] · 511 - Test data;
[0086] · 513 - Metrics;
[0087] · 515 - Chosen model;
[0088] · 521 - Productionised model;
[0089] · 531 - Test data;
[0090] · 533 - Test code;
[0091] · 535 - Model;
[0092] · 541 - Application code;
[0093] · 543 - Code and model in production;
[0094] · 551 - Production data.
[0095] A monitoring procedure should therefore understand the type of AI / ML solution and application in order to select the appropriate metrics accordingly.
[0096] The following is a selection of some agreements in 3GPP RAN1 that are related to the present disclosure:
[0097] RP-213599 [1]:
[0098] 3GPP agreed "Study on Artificial Intelligence (AI) / Machine Learning (ML) for NR Air Interface".
[0099] RAN1#110bis-e Agreements
[0100] Study AI / ML model monitoring for at least the following purposes: model activation, deactivation, selection, switching, fallback, and update (including re-training).
[0101] Study at least the following metrics / methods for AI / ML model monitoring in lifecycle management per use case:
[0102] . Monitoring based on inference accuracy, including metrics related to intermediate KPIs
[0103] i. Monitoring based on system performance, including metrics related to system peformance KPIs
[0104] ii. Other monitoring solutions, at least following 2 options.
[0105] · Monitoring based on data distribution
[0106] a) Input-based: e.g., Monitoring the validity of the AI / ML input, e.g., out-of-distribution detection, drift detection of input data, or something simple like checking SNR, delay spread, etc.
[0107] b) Output-based: e.g., drift detection of output data
[0108] · Monitoring based on applicable condition
[0109] Note: Model monitoring metric calculation may be done at NW or UE
[0110] Study performance monitoring approaches, considering the following model monitoring KPIs as general guidance
[0111] iii. Accuracy and relevance (i.e., how well does the given monitoring metric / methods reflect the model and system performance)
[0112] iv. Overhead (e.g., signaling overhead associated with model monitoring)
[0113] v. Complexity (e.g., computation and memory cost for model monitoring)
[0114] vi. Latency (i.e., timeliness of monitoring result, from model failure to action, given the purpose of model monitoring)
[0115] vii. FFS: Power consumption
[0116] viii. Other KPIs are not precluded.
[0117] Note: Relevant KPIs may vary across different model monitoring approaches.
[0118] FFS: Discussion of KPIs for other LCM procedures
[0119] In CSI compression using two-sided model use case, study potential specification impact for performance monitoring including:
[0120] ·NW-side performance monitoring: NW monitors the performance and make decisions of model activation / deactivation / updating / switching
[0121] ·UE-side performance monitoring: UE monitors the performance and reports to Network, NW makes decisions of model activation / deactivation / updating / switching
[0122] In CSI compression using two-sided model use case, further study potential specification impact related to assistance signaling and procedure for model performance monitoring.
[0123] In CSI compression using two-sided model use case, further study at least the following options for performance monitoring metrics / methods:
[0124] ·Intermediate KPIs as monitoring metrics (e.g., SGCS)
[0125] ·Eventual KPIs (e.g., Throughput, hypothetical BLER, BLER, NACK / ACK).
[0126] ·Legacy CSI based monitoring: schemes using additional legacy CSI reporting
[0127] ·Other monitoring solutions, at least including the following option:
[0128] oInput or Output data based monitoring: such as data drift between training dataset and observed dataset and out-of-distribution detection
[0129] For BM-Case1 and BM-Case2 with a UE-side AI / ML model, study the following alternatives for model monitoring with potential down-selection:
[0130] ·Atl1. UE-side Model monitoring
[0131] oUE monitors the performance metric(s)
[0132] oUE makes decision(s) of model selection / activation / deactivation / switching / fallback operation
[0133] ·Atl2. NW-side Model monitoring
[0134] oNW monitors the performance metric(s)
[0135] oNW makes decision(s) of model selection / activation / deactivation / switching / fallback operation
[0136] ·Alt3. Hybrid model monitoring
[0137] oUE monitors the performance metric(s)
[0138] oNW makes decision(s) of model selection / activation / deactivation / switching / fallback operation
[0139] For BM-Case1 and BM-Case2 with a network-side AI / ML model, study the NW-side model monitoring:
[0140] ·NW monitors the performance metric(s) and makes decision(s) of model selection / activation / deactivation / switching / fallback operation
[0141] Regarding NW-side model monitoring for a network-side AI / ML model of BM-Case1 and BM-Case2, study the potential specification impacts from the following aspects
[0142] · Beam measurement and report for model monitoring
[0143] Note: This may or may not have specification impact.
[0144] RAN1#111 Agreements:
[0145] Regarding AI / ML model monitoring for AI / ML based positioning, to study and provide inputs on feasibility, potential benefits (if any) and potential specification impact at least for the following aspects
[0146] · At least the following are identified for further study as potential data for calculating monitoring metric
[0147] o If monitoring based on model output
[0148] § E.g. , estimated UE location corresponding to model output for direct AI / ML positioning, estimated intermediate parameter(s) corresponding to model output for AI / ML assisted positioning, ground truth label corresponding to model inference output for both direct and AI / ML assisted positioning
[0149] o If monitoring based on model input
[0150] § E.g., measurement corresponding to model inference input
[0151] o Note1: other type of potential data for model monitoring is not precluded
[0152] o Note2: combination of one or more type of potential data for monitoring is not precluded
[0153] · If a given type of data is necessary for calculating monitoring metric, study whether and if so
[0154] o How an entity can be used to provide the given type of data for calculating monitoring metric
[0155] § Companies are requested to report their assumption of the entity (or entities) used to provide the given type of data for calculating monitoring metric for each case
[0156] o Potential signalling for provisioning of the given type of data for calculating associated monitoring metric
[0157] o Potential assistance signaling and procedure to facilitate an entity providing data for calculating monitoring metric
[0158] o Potential UE-network interaction
[0159] § E.g., model monitoring decision indication between UE and network
[0160] The following description of examples of the present disclosure, with reference to the accompanying drawings, is provided to assist in a comprehensive understanding of certain examples of the present invention. The description includes various specific details to assist in that understanding but these are to be regarded as merely exemplary. Accordingly, those of ordinary skill in the art will recognize that various changes and modifications of the examples described herein can be made without departing from the scope of the invention or disclosure.
[0161] The same or similar components may be designated by the same or similar reference numerals, although they may be illustrated in different drawings.
[0162] Detailed descriptions of techniques, structures, constructions, functions or processes known in the art may be omitted for clarity and conciseness, and to avoid obscuring the subject matter of the present disclosure.
[0163] The terms and words used herein are not limited to the bibliographical or standard meanings, but are merely used to enable a clear and consistent understanding of the invention.
[0164] Throughout the description of this specification, the words "comprise", "include" and "contain" and variations of the words, for example "comprising" and "comprises", means "including but not limited to", and is not intended to (and does not) exclude other features, elements, components, integers, steps, processes, operations, functions, characteristics, properties and / or groups thereof.Also, throughout the description of this specification, the word "Cramer-von Mises" means " ".
[0165] Throughout the description of this specification, the singular form, for example "a", "an" and "the", encompasses the plural unless the context otherwise requires. For example, reference to "an object" includes reference to one or more of such objects.
[0166] Throughout the description, the expression "at least one of A, B and / or C" (or the like) and the expression "one or more of A, B and / or C" (or the like) should be seen to separately include all possible combinations, for example: A, B, C, A and B, A and C, A and B and C.
[0167] Throughout the description of this specification, language in the general form of "X for Y" (where Y is some action, process, operation, function, activity or step and X is some means for carrying out that action, process, operation, function, activity or step) encompasses means X adapted, configured or arranged specifically, but not necessarily exclusively, to do Y.
[0168] Features, elements, components, integers, steps, processes, operations, functions, characteristics, properties and / or groups thereof described or disclosed in conjunction with a particular aspect, embodiment or example are to be understood to be applicable to any other aspect, embodiment or example described herein unless incompatible therewith..
[0169] Certain examples of the present disclosure relate to methods, apparatus and / or systems etc. for monitoring metrics and / or KPIs for AI / ML models. Further, certain examples of the present disclosure relate to methods and apparatus for selecting metrics to monitor based on the AI / ML model, or information thereon, and for recommending or identifying solutions to address an issue identified in the AI / ML model through monitoring of the selected metrics (e.g., by identifying abnormal behaviour of a selected metric). Furthermore, in certain examples of the present disclosure, the monitored metrics are each one of a stability metric, a performance metric and an operational metric; and the metrics include at least one metric from each of these categories
[0170] The following examples are applicable to, and use terminology associated with, 3GPP 4G (e.g., LTE) and / or 5G (e.g., NR). However, the skilled person will appreciate that the techniques disclosed herein are not limited to these examples or to 3GPP 4G (e.g., LTE) and / or 5G (e.g., NR), and may be applied in any suitable system or standard, for example one or more existing and / or future generation wireless communication systems or standards (e.g., B5G, 6G etc.). The skilled person will appreciate that the techniques disclosed herein may be applied in any existing or future releases of 3GPP 4G (e.g., LTE) and / or 5G (e.g., NR) or any other relevant standard. For example, the functionality of the various network entities and other features disclosed herein may be applied to corresponding or equivalent entities or features in other communication systems or standards. Corresponding or equivalent entities or features may be regarded as entities or features that perform the same or similar role, function, operation or purpose within the network.
[0171] A particular network entity may be implemented as a network element on a dedicated hardware, as a software instance running on a dedicated hardware, and / or as a virtualised function instantiated on an appropriate platform, e.g. on a cloud infrastructure.
[0172] The skilled person will appreciate that the present invention is not limited to the specific examples disclosed herein. For example:
[0173] · The techniques disclosed herein are not limited to 3GPP 4G or 5G.
[0174] · One or more entities in the examples disclosed herein may be replaced with one or more alternative entities performing equivalent or corresponding functions, processes or operations.
[0175] · One or more of the messages in the examples disclosed herein may be replaced with one or more alternative messages, signals or other type of information carriers that communicate equivalent or corresponding information.
[0176] · One or more further elements, entities and / or messages may be added to the examples disclosed herein.
[0177] · One or more non-essential elements, entities and / or messages may be omitted in certain examples.
[0178] · The functions, processes or operations of a particular entity in one example may be divided between two or more separate entities in an alternative example.
[0179] · The functions, processes or operations of two or more separate entities in one example may be performed by a single entity in an alternative example.
[0180] · Information carried by a particular message in one example may be carried by two or more separate messages in an alternative example.
[0181] · Information carried by two or more separate messages in one example may be carried by a single message in an alternative example.
[0182] · The order in which operations are performed may be modified, if possible, in alternative examples.
[0183] · The transmission of information between network entities is not limited to the specific form, type and / or order of messages described in relation to the examples disclosed herein.
[0184] Certain examples of the present disclosure may be provided in the form of an apparatus / device / network entity configured to perform one or more defined network functions and / or a method therefor. Such an apparatus / device / network entity may comprise one or more elements, for example one or more of receivers, transmitters, transceivers, processors, controllers, modules, units, and the like, each element configured to perform one or more corresponding processes, operations and / or method steps for implementing the techniques described herein. For example, an operation / function of X may be performed by a module configured to perform X (or an X-module). Certain examples of the present disclosure may be provided in the form of a system (e.g., a network) comprising one or more such apparatuses / devices / network entities, and / or a method therefor.
[0185] It will be appreciated that examples of the present disclosure may be realized in the form of hardware, software or a combination of hardware and software. Certain examples of the present disclosure may provide a computer program comprising instructions or code which, when executed, implement a method, system and / or apparatus in accordance with any aspect, example and / or embodiment disclosed herein. Certain embodiments of the present disclosure provide a machine-readable storage storing such a program.
[0186] As discussed above, one new framework developed as part of 5G networks (and beyond, e.g., B5G, 6G) is the use of artificial intelligence / machine learning (AI / ML), which may be used for the optimisation of the operation of 5G networks (and beyond).
[0187] The dynamic nature of deployment environments for ML models means that AI solutions can become outdated and lose their effectiveness. This underscores the importance of ongoing monitoring and evaluation of ML-powered autonomous systems to maintain their safety and reliability in real-world conditions. An objective is to develop ML-powered autonomous systems that can operate securely and efficiently in changing environments, without posing any undue risk to business operations.
[0188] Model monitoring plays a critical role in the machine learning lifecycle management, as it helps to sustain the performance of models in changing conditions. Real-world data is constantly evolving, and models trained on old data may no longer produce accurate results. Through monitoring, data scientists can monitor the performance of models over time, detect any decline and take corrective measures, such as retraining the model or adjusting its parameters, to ensure optimal performance.
[0189] 3GPP has recently begun to explore the field of model monitoring for AI / ML solutions used in telecommunications networks. This is due to the increasing integration of AI / ML components in various applications, such as management and orchestration, resource allocation, etc. With this integration, the importance of ensuring the reliability and safety of these components has become a major concern for the industry.
[0190] Various examples, aspects, embodiments etc. of present disclosure propose a novel mechanism for AI / ML monitoring of models deployed in a telecommunications network and / or UEs based on model monitoring metrics and / or KPIs defined herein.
[0191] In particular, certain examples of the present disclosure introduce solutions, e.g., methods, apparatus, systems etc., for a novel monitoring mechanism of an AI / ML model, using newly defined performance monitoring metrics (and / or KPIs) that are categorised based on their dependence on monitoring data (i.e., data collected for monitoring purpose), monitored AI / ML model(s), and / or monitored model functionality or use case. Certain examples provide methods, apparatus, systems etc., for recommending or identifying solutions / actions for addressing an issue with an AI / ML model which is identified through the monitoring. Further, certain examples provide methods, apparatus, systems etc. combining such methods or any other such methods disclosed herein.
[0192] In certain embodiments, a disclosed monitoring mechanism may reside in a newly defined network entity / network function (NF) or in an existing network entity / network function), in a given UE (or a group of UEs), in a server, and / or in an application, etc.
[0193] For each particular AI / ML model (and model information, data, parameters, etc.), a monitoring method according to various examples of the present disclosure may include selecting one or more distinct monitoring metrics to track the model performance, model stability, and / or other model operation processes. For example, a monitoring metrics stack may be defined or configured in an exemplary method. A monitoring metrics stack may include one or more (and, in some examples, ideally all) of three broad categories of monitoring metrics, which are stability metrics, performance metrics, and operational metrics. These categories are based on the interdependence between the specific metrics, the monitoring data, monitored model (or model use case or functionality) and / or model operational behaviour. These three monitoring metrics are summarized as follows:
[0194] · Stability metrics that focus on the ability of the AI / ML model to maintain consistent results over a given time period. These metrics are dependent on the model (or the model functionality or use case) and model data. Stability metrics ensure that the model is not becoming over-fit or under-fit to the data, or that it is not exhibiting any other behavior that would indicate that the model is becoming unstable.
[0195] ·Performance metricsthat focus on the accuracy and effectiveness of the model's outputs. These metrics are dependent on the data, the ML model, and the operational behavior of the system. Performance metrics help to evaluate the overall quality of the model's results and to identify areas where the model may be underperforming.
[0196] ·Operational metricsthat focus on the operational behavior of the system. These metrics are dependent on the operational behavior of the system and they help to ensure that the system is functioning as intended and that the overall performance of the system is not being negatively impacted by operational issues such as hardware or software errors.
[0197] In various examples, the monitoring metrics stack includes all three categories of monitoring metrics, thereby including at least one metric from each category.
[0198] The distinct metrics defined in the above categories may provide a comprehensive approach to monitoring and evaluating the performance of ML models, for distinct types of applications and situations.
[0199] The AI / ML or application life-cycle manager (this may be an NF) may share, with the monitoring entity (e.g., a new NF or existing NF etc., as indicated above), information and / or KPIs about the ML model (or AI / ML model) that is currently deployed or is about to be deployed.
[0200] From the aforementioned three categories of metric, the monitoring mechanism according to the present disclosure recommends (e.g., determines, identifies, indicates etc.) a set or subset of metrics to which the AI / ML life-cycle manager should subscribe. It should be noted that different AI / ML solutions and applications may have varying acceptable performance levels, model risks, and costs of errors. Therefore, monitoring methods disclosed in examples of the present disclosure may take into account the specific type of AI / ML solution or model and / or the application being used, and may choose the appropriate metrics accordingly. That is, for example, a monitoring method may determine one or more metrics (for the life cycle manager to subscribe to) based on the type of AI / ML solution / model and / or the application being used. This may allow for a customized and effective monitoring process that promotes the success of AI / ML models deployed in networks. The aim of this mechanism is to create a feedback loop between the deployment phase and the model building phase, allowing not only to detect performance degradation but for continuous improvement of the AI / ML model.
[0201] Following this, the AI / ML or application life-cycle manager may select a set of measurements to be monitored, and send the selected set back to the monitoring function / mechanism (e.g., indicate the selected set of measurements to the monitoring mechanism). A monitoring instance or NF is deployed to track the selected measurements / metrics and provide regular updates on their values to the ML or application lifecycle manager.
[0202] In the event that any of the metrics exhibit unusual behavior, the monitoring mechanism may provide recommendations to the life-cycle manager. These recommendations can range from updating the existing model, switching to a more suitable alternative, implementing a fallback solution, or maintaining the current deployment. The aim of various methods disclosed herein is to ensure the reliability and performance of the AI / ML-powered network, contributing to its overall success. Examples relating to this are described in relation to Figure 3 below.
[0203] In various examples, if the monitoring function detects that the model is underperforming in any metric of the class performance, but not in any other KPI from a distinct category, the monitoring function might recommend retraining the deployed model to improve its performance.
[0204] In other examples, should the metrics where the abnormality is detected belong to both the performance and stability classes, the monitoring function may recommend the implementation of a fallback solution for ensuring the safety and reliability of the system, and may update, or recommend updating, the current model with the most recent data.
[0205] Furthermore, in various embodiments the monitoring function may provide insights into the changing operational environment and recommend switching to a more suitable model that aligns with the current conditions, when (and, in some examples, only when) a measurement of class operation falls below a predefined threshold. For example, if the RAM utilization of the ML's container increases, the monitoring entity may suggest switching to a model with lower RAM consumption, while still meeting the required KPIs.
[0206] The monitoring mechanism(s) according to various embodiments of the present disclosure may provide tailored recommendations based on the reported abnormal measurements, to ensure the reliable and efficient operation of the ML-powered network. The disclosed monitoring mechanisms may be designed to support the decision-making process of the AI / ML life-cycle manager by providing valuable insights and recommendations based on the performance metrics of the model. This framework may categorize the monitoring metrics into three groups, which allows for a more comprehensive and targeted evaluation of the model's behavior. The recommendations may aid the AI / ML life-cycle manager in making informed decisions regarding the deployment and maintenance of the ML module, thereby leading to the success of the network infrastructure.
[0207] An embodiment of the present disclosure provides a method for monitoring an AI / ML model, the method performed by a first network entity (e.g., a monitoring entity) and comprising one or more of the following operations: receiving, from a second network entity (e.g., an AI / ML or application life cycle manager) information and / or KPIs about an ML model that is being deployed or will be deployed (e.g., at the second network entity or at a third network entity in communication with the second network entity); identifying or determining one or more metrics (e.g., this may be based on the AI / ML model type and / or an application being used in association with the AI / ML model), and indicating the one or more metrics to the second network entity (e.g., indicating the second network entity to subscribe to the one or more metrics); receiving, from the second network entity, a selection from the one or more metrics (e.g., at least one of the one or more metrics) to monitor; tracking or monitoring the selected metrics and / or associated measurements; providing updates on the values of the selected metrics (e.g., providing tracking / monitoring results or information) to the second network entity; and if one or more of the selected metrics are identified or determined to exhibit unusual behaviour (e.g., measurements fall outside of expected ranges, or are above / below specified threshold(s), or are underperforming, or diverge from expected results etc.), providing one or more recommendations, actions or instructions to the second network entity (e.g., the one or more recommendations or instructions may include one or more of updating the existing model, switching to a more suitable alternative, implementing a fallback solution, or maintaining the current deployment).
[0208] According to another embodiment of the present disclosure, there is provided a method for monitoring an AI / ML model, the method performed by a second network entity (e.g., an AI / ML or application life cycle manager) and comprising one or more of the following operations: transmitting, to a first network entity (e.g., a monitoring entity) information and / or KPIs about an ML model that is being deployed or will be deployed (e.g., at the second network entity or at a third network entity in communication with the second network entity); receiving one or more metrics from the first network entity (e.g., receiving an indication of the one or more metrics, so as to subscribe to the one or more metrics); transmitting, to the first network entity, a selection from the one or more metrics to monitor (e.g., at least one of the one or more metrics); receiving updates on the values of the selected metrics being monitored at the first network entity (e.g., receiving tracking / monitoring results or information); receiving one or more recommendations, actions or instructions to the second network entity (e.g., the one or more recommendations or instructions may include one or more of updating the existing model, switching to a more suitable alternative, implementing a fallback solution, or maintaining the current deployment) - this may be in the event that the first network entity identifies or determines that one or more of the selected metrics exhibit unusual behaviour (e.g., measurements fall outside of expected ranges, or are above / below thresholds, or are underperforming, or diverge from expected results etc.); and implementing one of the received recommendations, actions or instructions.
[0209] A feature of various embodiments of the present disclosure is to provide a model monitoring method and apparatus to gather distinct monitoring metrics and / or KPIs for various types of AI / ML applications available in the network side (e.g. RAN, core network, and / or any network entity and / or network function) and / or the UE side, or two-sided (i.e. UE and Network -side), or server, cloud, or other.
[0210] More specifically, various embodiments define three broad types of monitoring metrics - stability, performance, and operational metrics - based on the dependency of the distinct metrics on the data, the AI / ML model, and / or the operational behaviour. An example on the monitoring metrics is shown in Figure 1.
[0211] Stability metrics
[0212] The stability metrics play a role in detecting a distribution shift on either data or features. A distribution shift refers to a change in the distribution of input data used for training a machine learning model and the distribution of data that the model will encounter in real-world scenarios or its features. This shift may result in a poor performance of the model due to the invalidation of underlying assumptions about the data. There can be various reasons for distribution shifts, including alterations in data collection procedures, changes in the environment where the model is applied, or modifications in the target population served by the model. It is beneficial, and at times even essential, to detect and address distribution shifts as they can negatively impact model accuracy and lead to biased or incorrect predictions. The solution presented in this disclosure addresses two types of distribution shifts: prior probability shift and covariate shifts.
[0213] A prior probability shift occurs when the distribution of the input data in the training phase of the model is different from the distribution of the input data in the production phase. One or more of the following metrics may be used to identify a prior probability shift (among others, i.e., this list is not limiting):
[0214] ·Population Stability Index (PSI). PSI is used to measure the difference between the distribution of the input data in the training phase and the distribution of the input data in the production phase. PSI is calculated by dividing the data into equal-sized bins and comparing the cumulative distribution functions (CDF) of the training data and production data for each bin (see annex). A large PSI value indicates a significant difference between the two distributions and, therefore, a potential shift in prior probabilities.
[0215] ·The Divergence Indexis a metric used to measure the difference between two probability distributions. It provides a numerical value that quantifies the dissimilarity between two distributions. The value ranges from 0, indicating that the two distributions are identical, to 1, indicating that the two distributions are completely different. The Divergence Index can be calculated using various methods, such as, Jensen-Shannon divergence, or Total Variation Distance.
[0216] ·Maximum Mean Discrepancy (MMD):The MMD is a non-parametric metric that compares the mean of two distributions. A large MMD value indicates a significant difference between the two distributions, which could indicate a prior probability shift.
[0217] ·Kullback-Leibler (KL) Divergence:KL divergence is a measure of the difference between two distributions. A large KL divergence value indicates a significant difference between the two distributions, which could indicate a prior probability shift.
[0218] ·Energy Distance (ED):The ED is a non-parametric measure of the difference between two distributions. ED is a statistical distance measure between two probability distributions. It measures the amount of work required to transform one distribution into the other. A large ED value indicates a significant difference between the two distributions, which could indicate a prior probability shift.
[0219] ·Cramer-von Mises (CvM) Test: The CvM test measures the difference between the empirical distribution function (EDF) of the test data and the CDF of the training data. A large difference between the two functions indicates a prior probability shift.
[0220] ·Anderson-Darling Test: This statistical test can be used to determine if two probability distributions come from the same population or if there is a shift in the distribution.
[0221] A covariate shift occurs when the distribution of features in the training data and the distribution of features in the deployment data are different. To detect a covariate shift one or more of the following metrics could be used (among others, i.e., this list is not limiting):
[0222] ·Characteristic Stability Index (CSI)is a measure of the difference between two datasets, with a higher CSI indicating a greater difference between the two datasets. It is calculated by comparing the distribution of the feature values in the training data to the distribution of the feature values in the deployment data, and it can be used to detect covariate shift by measuring the change in the distribution of the features over time.
[0223] ·Wasserstein's Distance, also known as Earth Mover's Distance, is a distance metric that measures the amount of effort needed to transform one distribution into another. It is calculated as the minimum amount of "work" required to transform the cumulative distribution function of the training data to match the cumulative distribution function of the deployment data. It can be used to detect covariate shift by measuring the amount of difference between the two distributions.
[0224] Further discussion / details of some of the above metrics may be found in the Annex later in this document.
[0225] These are some of the metrics that may be used to detect a prior probability shift. However, no single metric is perfect; the choice of metric may depend on the particular problem at hand and the assumptions about the underlying data. These metrics can provide insights into the source of the shift, which can then inform decisions about how to adjust the model or collect more representative training data.
[0226] According to various examples, stability metrics may be defined as being dependent on the input data and the output predictions. Stability metrics may: identify data probability shifts (e.g., PSI, divergence index etc), and / or identify covariate shifts (e.g., novelty index etc.).
[0227] According to various examples of the present disclosure, an example of a metric in the category of stability metrics is one of any of the metrics indicated above. For example, one or more stability metrics may include any combination of one or more of the above listed metrics.
[0228] Performance metrics
[0229] The performance metrics are designed to detect a shift in the underlying relationship between the independent and dependent variables, also known as a conceptual shift. These metrics measure the performance of a deployed machine learning model compared to when it was trained or compared to previous outcomes obtained post-deployment. The focus of these metrics is to evaluate the quality of the model and to determine if it is still performing as expected. Some commonly used performance metrics for different types of machine learning models include, but are not limited to:
[0230] ·Root Mean Squared Error (RMSE): RMSE is a popular performance metric for regression problems. It measures the average magnitude of the error in predictions made by the model, by taking the square root of the mean of the squared differences between the predicted and actual values.
[0231] ·R-Square: is a measure of goodness-of-fit for regression models. It measures the proportion of variability in the dependent variable that the independent variables can explain in the model. R-Square can be calculated by dividing the explained variance by the total variance.
[0232] ·F1-Score:is a measure of a model's performance that considers both precision and recall, taking their harmonic mean. This means that it evaluates the model's ability to identify positive cases (recall) as well as the accuracy of the cases it does identify (precision).What makes the F1 score a robust measure of model performance is its approach of considering both attributes, in contrast to accuracy which only measures the number of correct predictions. By accounting for precision and recall, the F1 score provides a balanced evaluation of a model's overall performance.
[0233] ·Mean Squared Error (MSE):is a performance metric for regression problems that measures the average of the squared differences between the predicted and actual values.
[0234] ·AUC-ROC:is a performance metric for binary classification problems that measures the ability of a classifier to distinguish between positive and negative classes. AUC-ROC is the area under the receiver operating characteristic (ROC) curve.
[0235] ·Cross-Entropy Loss: Cross-Entropy Loss is a performance metric for multi-class classification problems that measures the difference between the predicted probabilities and the actual class labels. The definition of Cross-Entropy Loss is the negative log-likelihood of the predicted probabilities.
[0236] ·Kullback-Leibler Divergence Loss:is a performance metric for classification problems that measures the difference between the predicted and actual class distributions. It is obtained by computing the sum of the differences between the predicted and actual class distributions.
[0237] ·Hit Rate:is a performance metric for recommendation systems that measures the proportion of recommendations that result in user engagement. The definition of Hit Rate is the number of engaged recommendations divided by the number of total recommendations.
[0238] ·Diversity:this is a performance metric for recommendation systems that measures the variety of recommendations.
[0239] ·Novelty:Novelty is a performance metric for recommendation systems that measures the degree of uniqueness of recommendations. Hence, it is the proportion of items in the recommendation list that are not present in the user's history. Implementation: Novelty can be calculated by dividing the number of unique items in the recommendation list by the total number of items.
[0240] Further discussion / details of some of the above metrics may be found in the Annex later in this document.
[0241] These metrics may provide a way to assess the health of a machine learning model and / or detect shifts in the data over time, which can help ensure that the model continues to produce accurate results.
[0242] According to various examples, performance metrics may be defined as being dependent on the ML mode. Performance metrics may identify concept shift, and / or may use ML metrics (e.g., accuracy, F1, RMSE, MSE, ROC etc.)
[0243] According to various examples of the present disclosure, an example of a metric in the category of performance metrics is one of any of the metrics indicated above. For example, one or more performance metrics may include any combination of one or more of the above listed metrics.
[0244] Operational metrics
[0245] Operational metrics are key performance indicators that evaluate the practical implementation of a deployed model from a usage perspective. Unlike performance metrics, which may focus on the quality of the model itself, operational metrics may focus on the physical infrastructure and operational behavior of the model. These metrics provide valuable insights into the real-world performance and efficiency of the model and can be used to optimize its deployment for maximum efficiency and cost-effectiveness. Some common operational metrics include, but are not limited to:
[0246] ·ML Throughput: This metric measures the number of requests that a model can handle per unit of time, and provides a quantifiable measure of the model's ability to handle large volumes of requests.
[0247] ·ML Latency: This metric evaluates the average response time of the model, which is an important factor in determining the user experience of the model.
[0248] ·Resource Usage: This metric evaluates the average consumption of resources such as memory, CPU, disk usage, and IO when making predictions. It provides valuable information about the computational requirements of the model and can help inform decisions about hardware and infrastructure.
[0249] ·Power Consumption: This metric measures the average power consumption of the model, which is important in cloud-based deployment scenarios where cost is a major concern. By monitoring power consumption, data scientists can identify areas for optimization to minimize costs and improve efficiency.
[0250] ·Scalability: The ability of the system to handle increasing loads or demands, such as number of requests or concurrent users.
[0251] ·Availability: The percentage of time that the system is operational and accessible to users.
[0252] These operational metrics may provide a comprehensive view of the practical implementation of the model and may help inform ongoing efforts to improve its performance and efficiency in real-world deployment scenarios.
[0253] According to various examples, operational metrics may be defined as being independent of the ML model and / or the data. Operational metrics may: identify ML system health issues, identify endpoint latency, and / or identify IO / memory / CPU usage etc.
[0254] According to various examples of the present disclosure, an example of a metric in the category of operational metrics is one of any of the metrics indicated above. For example, one or more operational metrics may include any combination of one or more of the above listed metrics.
[0255] An example in accordance with the present disclosure will now be described with reference to Figure 2. Figure 2 shows an example flow chart of possible setups according to various examples of the present disclosure in relation to the creation / execution of a monitoring workload.
[0256] Figure 2 illustrates message exchange between the network entity in charge of AI / ML life-cycle manager (in one example, the application function (AF)) and / or the monitoring entity in relation to the monitoring workload subscription request and response procedure.
[0257] The text shown in Fig. 2 should not be seen as limiting but is provided merely to indicate options according to an example of the present disclosure. Below, several examples in accordance with the general method of Fig. 2 (as indicated by numerals 210, 220, 230, 240 and 250, which may be taken in any combination) are disclosed. Furthermore, it will be appreciated that one or more of the operations shown in Fig. 2 may be omitted, reordered or modified, if desired - the present disclosure should be seen to include all combinations of 210, 220, 230, 240 and 250, as well as each operation individually.
[0258] A description of one example according to Fig. 2 is as follows:
[0259] When a new AI / ML model is to be deployed in the network, the entity in charge of its life cycle manager (e.g., UE, network entity, and / or network function (NF), etc.) may trigger, for example, a MonitoringRequest workload. To do so, the entity 10 (e.g., ML NF 10) sends a request message to the Monitoring Network Function 20 (MNF 20) with information related to the AI / ML application or NF.
[0260] For example, in operation 210 of Figure 2, the ML NF 10 (or network entity 10) may send, to the MNF 20, a MonitoringRequest message (or, more generally, a message) that may comprise, one, some, or all of (i.e., one or more of) the following information:
[0261] · Model type: Different learning types of AI / ML models, such as supervised, unsupervised, or reinforcement learning, may require different monitoring metrics. For example, accuracy and precision are important metrics for supervised models, while stability and reliability are important, even critical, for reinforcement learning models.
[0262] · Model architecture: The architecture of the AI / ML model, such as deep neural networks or decision trees, will also impact the monitoring metrics. For example, monitoring the activations of hidden nodes in a neural network can provide insights into the model's behaviour.
[0263] · Data characteristics: The monitoring metrics will also be impacted by the data used to train and validate the model. For example, monitoring the distribution of input data and its changes over time can help detect data drift, which can affect the model's accuracy.
[0264] · Operational environment: The environment in which the model should be deployed, such as a UE, cloud or edge device, and the available computational resources, will also influence the monitoring metrics.
[0265] · Performance objectives: The desired performance objectives of the model, such as accuracy, speed, or stability, will also impact the monitoring metrics. For example, monitoring the latency and resource utilization can help ensure that the model meets performance requirements.
[0266] Upon receiving the MonitoringRequest message, in operation 220 the MNF 20 may analyse the monitoring request and may select a list of one or more measurements available based on the information received from the ML NF 10 (e.g., the AI / ML life-cycle manager). For example, the MNF 10 may filter a list of possible metrics / measurements based on the information in the message, or may otherwise determine a list of possible metrics / measurements based on the message. For example, the one or more selected metrics may be any of those disclosed herein, such as in relation to the three categories defined above.
[0267] Furthermore, in operation 220 the MNF 10 may also include or configure a list of recommended monitoring metrics taking into account one or more of the proposed AI / ML model, expected performance and environment context. The recommended list of measurements may include a list of metrics from all the three categories described above operational, performance, stability), or a list of metrics from one or more of these three categories. For example, the recommended metrics may be any of those disclosed herein, such as in relation to the three categories defined above.
[0268] In addition, the monitoring entity 20 and ML NF 10 (or network entity 10) may agree on the frequency of the monitoring, such as real-time monitoring or periodic monitoring (5 min, 10 min, threshold based, etc). Thus, for each recommended metric, a proposed monitoring frequency is included in the MNF 20 (e.g., this may be preconfigured in the MNF 20, or signalled to the MNF 20 by some other network entity). Operation 220 may include the MNF 220 identifying a frequency corresponding to each filtered / determined metric.
[0269] Through a MonitoringResponse, this information (e.g., a list of available metrics based on the features and characteristics of the model, the recommended monitoring metrics, and / or the frequency of each metric in the list etc.) may be shared with the AI / ML NF 10 (e.g., transmitting to the AI / ML NF 10 by the MNF 20). That is, in operation 230, the MNF 20 may transmit the information determined in operation 220 to the ML NF 10.
[0270] After receiving the list of measurements, in operation 240 the ML NF 10 may select the desired metric(s) to monitor as well as the monitoring frequency. The AI / ML NF 10 (or network entity 10) may send the selected desired metric(s) (or an indication thereof) to the MNF 10 through a MonitoringDecision message (or, more generally, a message), optionally including the monitoring frequency of each selected metric (or an indication thereof).
[0271] Optionally, upon or after sending the MonitoringDecision, the ML NF 10 AI / ML life-cycle manager may start a timer, e.g., a count-down timer such as a monitoring response timer.
[0272] Optionally, the MonitoringDecision is acknowledged by the MNF 20 with an ACK message to the ML NF 10 in operation 250. If no response is received by the AI / ML NF 10 (or network entity10 ) by the time the timer expires, a new MNF instance may be queried.
[0273] Another example according to Fig. 2 is as follows.
[0274] In operation 210, a first network entity 10 (e.g., ML NF 10) may transmit a request including one or more of the above-mentioned types of information (i.e., those previously described in relation to operation 210 for the preceding example), to a second network entity 20 (e.g., MNF 20).
[0275] In operation 220, the second network entity 20 may identify at least one of: one or more metrics or measurements based at least part on the information in the request, one or more recommended metrics or measurements based on at least part of the information in the request; and a monitoring frequency for each metric in either list. For example, the metrics referred to here may be any of those disclosed herein (e.g., in relation to one or more of the three different categories described above) which are appropriate for the information in the request (i.e., which are identified based on the information).
[0276] In operation 230, the second network entity 20 may transmit the identified data from operation 220 to the first network entity 10; e.g., may transmit the list of recommended metrics and corresponding monitoring frequency information (or an indication thereof).
[0277] In operation 240, the first network entity 10 may select or identify one or more metrics from the list and, optionally, the corresponding monitoring frequency for each selected / identified metric, and may transmit the selected / identified metrics (or an indication thereof) to the second network entity 20 (and, optionally, transmit the corresponding monitoring frequencies (or an indication thereof) to the second network entity 20.
[0278] In operation 250, the second network entity 20 may transmit an acknowledgement message to the first network entity 10 upon receiving the transmission from operation 240. The first network entity 10, if not acknowledgement is received, may query a third network entity (e.g., a new MNF instance). In certain examples, no acknowledgement is required.
[0279] Note that, in certain embodiments, the first network entity 10 and the second network entity 20 correspond to the same entity, network node, apparatus etc. For example, both may be implemented in a UE, if the model and monitoring are performed at the UE (a discussion of deployment cases is provided below). As such, reference to communications (e.g., transmissions, receptions) between the first and second network entities includes local (e.g., in-device) communications, such as from one component to another.
[0280] An AI / ML model may be deployed at three different points in the network RAN, including the edge (e.g., a point in the network close to end users), UEs, or both (i.e. two-sided models):
[0281] ·Network edge: When deployed at the network edge, the AI / ML model may operate at the nearest point to the end-users, thereby reducing the latency and providing a faster response time. This deployment scenario is beneficial for use cases that require real-time data processing, high computation capabilities and low latency.
[0282] ·UE: At the user equipment, the model is deployed directly on the end-user device, providing a personalized experience and improving privacy as the data never leaves the device. This deployment scenario is ideal for use cases that require personalization.
[0283] ·Both: Joint deployment is a combination of both edge and UE deployment, where a part of the model is deployed at the edge or UE and another part in the opposite element. This deployment scenario is beneficial for use cases that require both high computational data processing capabilities and personalization.
[0284] Regardless of the deployment scenario, monitoring the performance of deployed AI / ML model(s) is crucial to guarantee their safety and reliability in production environments. The deployment scenario has a significant impact on the type of monitoring that can be performed and the challenges that need to be addressed. Various examples relating to different monitoring scenarios for each deployment type are now described:
[0285] a.Monitoring at the Edge: In various examples, when the model is deployed solely at the edge, an edge server may collect and store relevant metrics. Monitoring may involve counting measurements from one or more of three distinct types of metrics: operational, performance, and stability (e.g., as described above).
[0286] In various examples, when the model is deployed at the UE, the UE may (e.g., must) transmit relevant metrics to the edge for monitoring. The UE may periodically send the collected metrics to the edge server, which then may process and store the metrics for analysis.
[0287] In various examples, for a joint Edge and UE model, the monitoring network function may only need to receive information regarding the part of the model that executes on the UE side. For example, if the initial part of the model is at the UE then only operational metrics will be transmitted to the UE; however, if it is the final part of the AI / ML model residing at the UE, then all three types of metrics (operational, performance and stability) will be collected and reported by the UE.
[0288] b.Monitoring at the UE: In various examples, when the model is deployed at the UE, monitoring may be performed from within the UE, i.e., monitoring the performance of the model on the device. This type of monitoring may impact resource-constrained UEs. Monitoring metrics such as power consumption and resource utilization can be used to understand the impact of the model on the device. In addition, they can be used to assess the feasibility of deploying the model and the MNF on the UE. This type of monitoring may be particularly useful for models deployed jointly at the edge server and UE, where the final part of the inference process occurs on the UE. With all relevant metrics available on the device, there is no need to transmit data over the network, reducing latency and preserving privacy.
[0289] c.Monitoring Jointly at the Edge Server and UE: In various examples, monitoring is split between the edge server and the UE, with one part deployed at the edge server and the other part deployed at the UE. This type of monitoring is only suitable for models that are deployed jointly at the edge server and UE; and may provide insights into the interplay between the network and the UE, and help identify any challenges related to network-UE interactions.
[0290] Having a comprehensive understanding of the different monitoring scenarios for each deployment type is crucial for data scientists and network operators to make informed decisions about deploying and operating ML models in real-world environments. The various monitoring methods disclosed herein may be implemented according to any of the deployments types mentioned above, as appropriate.
[0291] In an ongoing monitoring process according to various embodiments of the present disclosure, the monitoring entity / function (e.g., MNF) may send measurement readings of the metrics selected by the AI / ML life-cycle manager (e.g., ML NF) through a MonitoringDecision message (or, more generally, a message); in certain examples, these messages may be sent regularly, e.g., at a predetermined frequency.
[0292] These readings may be both delivered to the AI / ML life-cycle manager and analyzed by the monitoring function to understand the current state of the AI / ML system. Based on the category or categories of the metrics showing abnormal behavior (assuming there are such), the monitoring function may provide specific recommendations on what actions to take in order to revert the situation. Further, the AI / ML life-cycle manager may act on the recommendations.
[0293] Various embodiments therefore include one or more distinct recommendations, and the determining and / or providing thereof, that can be made based on the category or categories to which the metrics exhibiting abnormal behavior belong. A logical graph illustrating an example of the present disclosure in relation to making recommendations is presented in Figure 3.
[0294] It will be appreciated that the various states and operations of Fig. 3 may be re-ordered and / or one or more states or operations may be omitted or modified, if desired. It will also be appreciated that the method of Fig. 3 may be implemented by one entity (e.g., an entity executing the AI / ML model and performing the monitoring, such as an entity comprising the ML NF and the MNF), or by a plurality of entities (e.g., a ML NF deployed separately to a MNF).
[0295] In 310 of Fig. 3, one or more metrics may be monitored (to give a non-limiting example, by a MNF following a method such as illustrated in Fig. 2). The one or more metrics may be one or more of the metrics described herein, for example.
[0296] In 320, if there is no detected decline in any of the monitored metrics (e.g., no abnormal behaviour is identified), the MNF may continue to monitor the one or more metrics. If there is a detected decline in a monitored metric, the method proceeds to 330.
[0297] In 330, the MNF determines the class (or type) of the declining / declined metric; e.g., stability, performance or operation.
[0298] If the determined class is stability (as in 332) or performance (as in 334), the method proceeds to 360. If the determined class is operation (as in 336), the method proceeds to 340.
[0299] In 350, in the event that any of the metrics that fall under the category of stability or performance report abnormal behaviour, or fall below a predefined threshold, a pure ML model problem may have arisen or may arise. Action may therefore be needed to prevent potential negative effects on the model's performance in the network.
[0300] According to various examples, the method may include one or more of the following operations (e.g., as may be performed by the MNF or performed by an ML NF acting on a recommendation from the MNF) in the event of an issue with stability and / or performance metrics:
[0301] 1.Identifying the cause of the shift: E.g., understanding the root cause of the metric drop can help determine a course of action, such as an optimal course. The drop may be due to a variety of reasons, poor algorithm design, a change in the distribution of the input data, a change in the domain of the problem, or other factors. This may be considered part of 350 shown in Figure 3, or a separate operation.
[0302] 2.Evaluating the impact on performance(indicated in 350): E.g., comparing the model's performance before and after the problem was detected. This may help determine the extent to which the new scenario has affected the accuracy and reliability of the model's predictions. Assessing the impact on performance may help in order to determine the steps to rectify the situation. If the degradation is small the following steps can be considered:
[0303] I.Retraining the model(indicated in 360): E.g., retrain the same model available in production with updated training data that is representative of the new distribution. The model may then learn to better handle the new distribution, and improve its performance metrics. One step may be to collect updated training data that accurately represents the new distribution. This data should be large enough to provide a robust representation of the new distribution, and it should be labelled correctly to ensure that the model can learn from it effectively. In some cases, the new data may not be sufficient on its own to retrain the model effectively. In these cases, it may be necessary to combine the new data with historical (training) data, and assign higher weights to features that have drifted significantly from each other. This can help to ensure that the updated model is able to effectively learn from the new distribution. The next step is to retrain the model using the updated training data.
[0304] II.Using a different model (i.e., model switch, indicate in 370): If retraining the model is not feasible or if it does not produce the desired results (following a check in 365, which results in a 'No' outcome), it may be necessary to consider using a different model. E.g., this could involve simple procedures such as updating the model architecture, adjusting the hyperparameters, or using different training algorithms to ensure that the model is able to effectively learn from the new data. Alternatively, switching model may include more complex procedures; e.g., selecting a model that has a more robust prior probability distribution, or using an ensemble of models to capture the changing distribution better. This may be done by testing several different models and selecting the one that best fits the new distribution.
[0305] III. After retraining the model or switching to a different model, it's important to evaluate its performance to determine if it has improved (indicated in 365 and / or 375, as appropriate). For example, this may be done using shadow testing or A / B testing approaches, where the new model (i.e., retrained or switched) is compared to the existing (e.g., champion) model in production. The results of these tests should be used to determine if the retrained model is a better fit for the new distribution, and if it should be deployed in production.
[0306] IV. If the retrained or different model performs well and is deemed to be a better fit for the new distribution (indicated in the 'Yes' outcome of 365, 375), the retrained or different model may be deployed in production (as indicated in 390). It is important to monitor the performance of this updated model closely to ensure that it continues to perform well and does not experience any unexpected issues.
[0307] 3. If the degradation is so significant that it is not possible to obtain satisfactory results using the above methods (indicated in the 'No' outcome of 375 or, if the option of switching is not implemented in a method, the 'No' outcome of 365), it may be necessary to implement a fallback solution. For example, a fallback solution may involve switching to a rule-based system, which is a system that operates based on a set of predefined rules and does not rely on machine learning algorithms. This may help to ensure the continued operation of the system, even if the ML model is not performing as expected. In another example, a fallback solution may be to switch to a manual process, where tasks that were previously automated are now performed by humans. This may also help to ensure the continued operation of the system, even if the ML model is not performing as expected. It is important to note that fallback solutions may not be able to fully compensate for the reduced performance of the ML model. Therefore, it is important to carefully evaluate the trade-off between the benefits of using an ML model and the risks associated with relying on fallback solutions. Additionally, fallback solutions should be designed and tested in advance to ensure they are effective and can be deployed quickly in the event of a failure of the ML model.
[0308] Referring to 336, operational metrics are a crucial aspect of any machine learning ML system. E.g., operational metrics may determine the practical efficiency and effectiveness of the system in real-world scenarios. In particular, CPU utilization, ML throughput, ML latency, power consumption, scalability, etc. are key performance indicators of the operational efficiency of the ML system. These metrics provide critical information about the resource utilization and performance characteristics of the ML system and allow us to monitor the performance of the system over time and make informed decisions about how to improve its operation.
[0309] In 350, in the event that any of the metrics that fall under the category of operational falls below the expected level of performance (e.g., show abnormal behaviour), it is may be necessary to, or is recommended to, take prompt action to address the issue. According to various examples, the method may include one or more of the following operations (e.g., as may be performed by the MNF or performed by an ML NF acting on a recommendation from the MNF) in the event of an issue with operational metrics:
[0310] 1.Identifying impact on performance(indicated in 340): E.g., similar to 350, the impact on performance may be evaluated such as by comparing the model's performance before and after the problem was detected.
[0311] 2.Root cause identification(indicated in 342): E.g., after gathering data (e.g., data relating to impact on performance or data relating to the abnormally behaving operational metric), the method may comprise performing a root cause analysis to determine the underlying reason for the decline in the operational metric. E.g., the reasons for the decline may include insufficient resources, poor resource utilization, flawed algorithm design, or any other factor that could impact the system's performance.
[0312] 3.Formulation of solutions(not illustrated in Fig. 3): E.g., based on the root cause analysis, formulating a set of solutions aimed at resolving the issue and restoring the operational metric to its desired level. In some examples such solutions may involve increasing available resources, optimizing the algorithm design, deploying a lighter-weight model, or other relevant actions.
[0313] 4.Evaluation(not illustrated in Fig. 3): E.g., before implementation of a solution(s), evaluating each proposed solution to determine its impact and determine the optimal (e.g., best) course of action. In some examples, this may can be done using simulation, testing, experimentation, or a combination of methods to accurately assess the solution's impact on the operational metric.
[0314] 5.Implementation(indicated in 344): E.g., once the optimal (e.g., most effective) solution has been identified, may should be implemented in the production environment. This may involve modifying resources such as updating the ML model, adjusting system configurations, or making any necessary changes to the system.
[0315] 6.Monitoring and Evaluation(indicated in 390): E.g., after the solution has been implemented, it is important to monitor the system's performance and evaluate the effectiveness of the solution in restoring the operational metric to its desired level.
[0316] According to various embodiments such as described above, when an operational metric falls below the expected level of performance, action may be taken to address the issue. This may involve one or more of data collection and analysis, root cause identification, solution formulation, solution evaluation, solution implementation, and monitoring and evaluation.
[0317] In an example of the present disclosure, a method of a monitoring function (or other network entity) may comprise: identifying (320) a metric used for monitoring an AI / ML model to be behaving abnormally (such as by one of the methods described herein); identifying (330) a type of the metric; and based on the type of the metric, determining (340-344, 350-380) a solution to address an issue in the AI / ML model related to (e.g., resulting from) the abnormal behaviour of the metric. For example, determining the solution may include updating the AI / ML model; for example, by retraining (360) the AI / ML model, changing (370) to a different AI / ML model, modifying (344) system resources of an entity on which the AI / ML model is executed, or implementing (380) a fallback solution. Optionally, the updated AI / ML model may be monitored, once in production (390), to identify any further issues (as may be indicated though a monitored metric showing abnormal behaviour) or persistence of the original issue, or to check the issue is resolved, or to check for improvement of the model.
[0318] Figure 4 is a block diagram illustrating an exemplary network entity 400 (or electronic device, or network node etc.) that may be used in examples of the present disclosure. For example, a ML NF, MNF, UE, eNB, device, network entity, network node, network function, network etc. as described in any of the embodiments / examples disclosed above may be implemented by or comprise network entity 400 (or be in combination with network entity 400). For example, an AI / ML or application life-cycle manager, ML NF or MNF in accordance with any of the examples / embodiments / aspects etc. described above may be implemented by or in combination with, or comprise, network entity 400.
[0319] The network entity 400 comprises a controller 405 (or at least one processor) and at least one of a transmitter 401, a receiver 403, or a transceiver (not shown). It will be appreciated that network entity may comprise an antenna also.
[0320] For example: controller 405 may be arranged to control the network entity 400 to perform any of the one or more features, operations or functions disclosed in relation to a network entity above; transmitter 401 may be arranged to transmit any one or more of the information, signals, data etc. mentioned above; and receiver 403 may be arranged to receive any one or more of the information, signals, data etc. mentioned above. The person skilled in the art would understand how such a network entity 400 in accordance with anyone or more example / embodiment disclosed herein may be provided.
[0321] For all of the examples / aspects / embodiments etc. described above / herein, it should be considered that the corresponding features / operations apply in any order or combination, and that furthermore there exists the possibility to omit one or more features / operations.
[0322] Moreover, for all of the examples, embodiments, aspects etc. above, these apply to at least LTE, NR, NR NTN or IoT NTN (note this list is merely to give some examples and should not be seen as limiting), including any related signalling / messages on any of the inferences X2, Xn, NG, S1, F1, etc (again, this list is merely to give some examples and should not be seen as limiting).It will be appreciated that, in each example / embodiment / aspect etc. described above, one or more features or operations may be omitted, modified or moved (e.g., to change the order of the features or the operations), if desired and appropriate.
[0323] Additionally, where the figures illustrating example method flows include text in relation to a specific step / operation, it will be appreciated that this text is simply an example of the corresponding step / operation, where a more general definition (such as may be found in the description of the corresponding step) may apply for the step / operation.
[0324] Additionally, regarding all of the above, one or more features or operations etc. from any example / embodiment may be combined with features or operations from any other example / embodiment. That is, the present disclosure should be considered to include all combinations of examples / embodiments disclosed herein, as appropriate, as well as combinations of individual features within and between each example / embodiment, as appropriate.
[0325] The techniques described herein may be implemented using any suitably configured apparatus and / or system. Such an apparatus and / or system may be configured to perform a method according to any aspect, embodiment or example disclosed herein. Such an apparatus may comprise one or more elements, for example one or more of receivers, transmitters, transceivers, processors, controllers, modules, units, and the like, each element configured to perform one or more corresponding processes, operations and / or method steps for implementing the techniques described herein. For example, an operation / function of X may be performed by a module configured to perform X (or an X-module). The one or more elements may be implemented in the form of hardware, software, or any combination of hardware and software.
[0326] It will be appreciated that examples of the present disclosure may be implemented in the form of hardware, software or any combination of hardware and software. Any such software may be stored in the form of volatile or non-volatile storage, for example a storage device like a ROM, whether erasable or rewritable or not, or in the form of memory such as, for example, RAM, memory chips, device or integrated circuits or on an optically or magnetically readable medium such as, for example, a CD, DVD, magnetic disk or magnetic tape or the like.
[0327] It will be appreciated that the storage devices and storage media are embodiments of machine-readable storage that are suitable for storing a program or programs comprising instructions that, when executed, implement certain examples of the present disclosure. Accordingly, certain examples provide a program comprising code for implementing a method, apparatus or system according to any example, embodiment and / or aspect disclosed herein, and / or a machine-readable storage storing such a program. Still further, such programs may be conveyed electronically via any medium, for example a communication signal carried over a wired or wireless connection.
[0328] While the invention has been shown and described with reference to certain examples, it will be understood by those skilled in the art that various changes in form and detail may be made therein without departing from the scope of the invention.
[0329] The reader's attention is directed to all papers and documents which are filed concurrently with or previous to this specification in connection with this application and which are open to public inspection with this specification, and the contents of all such papers and documents are incorporated herein by reference.
[0330] Annex
[0331] There now follows additional details relating to some of the metrics and measurements described above:
[0332]
[0333]
[0334]
[0335]
[0336]
[0337]
[0338]
[0339]
[0340]
[0341]
[0342]
[0343]
[0344]
[0345]
[0346] Acronyms and Definitions
[0347] 3GPP 3rdGeneration Partnership Project
[0348] 5G 5thGeneration
[0349] 5GC 5G Core
[0350] 5QI 5G QoS Identifier
[0351] 5GS 5G System
[0352] 5GSM 5G System Session Management
[0353] 5GMM 5G System Mobility Management
[0354] AF Application Function
[0355] AI Artificial Intelligence
[0356] AM Acknowledged Mode
[0357] AMF Access and Mobility Management Function
[0358] AS Application Server
[0359] ASP Application Service Provider
[0360] AUSF Authentication Server Function
[0361] CDN Content Delivery Network
[0362] DCAF Data Collection Application Function
[0363] DNAI Data Network Access Identifier
[0364] DNN Data Network Name
[0365] DNS Domain Name Server
[0366] DRB Data Radio Bearer
[0367] eNB Evolved Node B
[0368] EPC Evolved Packet Core
[0369] FEC Forward Error Correction
[0370] FQDN Fully Qualified Domain Name
[0371] GBR Guaranteed Bit Rate
[0372] gNB Next generation Node B
[0373] GPSI Generic Public Subscription Identifier
[0374] HSS Home Subscriber Service
[0375] IAB Integrated Access and Backhaul
[0376] ID Identity / Identifier
[0377] IIoT Industrial Internet of Things
[0378] IMEI International Mobile Equipment Identities
[0379] IP Internet Protocol
[0380] I-SMF Intermediate SMF
[0381] LADN Local Area Data Network
[0382] LL SSM Lower Layer SSM
[0383] MBMS Multimedia Broadcast / Multicast Service
[0384] MBS Multicast / Broadcast Service
[0385] MBSF Multicast / Broadcast Service Function
[0386] MBSTF Multicast / Broadcast Service Transport Function
[0387] MB-SMF Multicast / Broadcast Session Management Function
[0388] MB-UPF Multicast / Broadcast User Plane Function
[0389] ML Machine Learning
[0390] MME Mobility Management Entity
[0391] MN Master Node
[0392] MNF Monitoring Network Function
[0393] MNO Mobile Network Operator
[0394] MT Mobile Termination
[0395] NAS Non-Access Stratum
[0396] NEF Network Exposure Function
[0397] NRF Network Repository Function
[0398] NG-RAN Next Generation Radio Access Network
[0399] NG-eNB Next Generation eNB
[0400] NSA Non-Standalone
[0401] NSSF Network Slice Selection Function
[0402] NTN Non-Terrestrial Networks
[0403] NW Network
[0404] NWDAF Network Data Analytics Function
[0405] OS Operating System
[0406] OSAPP OS Application
[0407] PCF Policy Control Function
[0408] PCO Protocol Configuration Options
[0409] PDR Packet Detection Rule
[0410] PDU Protocol Data Unit
[0411] PTM Point To Multipoint
[0412] PTP Point to Point
[0413] QFI QoS Flow Identifier (ID)
[0414] QoS Quality of Service
[0415] RACH Random Access Channel
[0416] RAN Radio Access Network
[0417] RRC Radio Resource Control
[0418] RSD Route Selection Descriptor
[0419] SA Standalone
[0420] SDAP Service Data Adaptation Protocol
[0421] SDU Service Data Unit
[0422] SGW Serving Gateway
[0423] SIM Subscriber Identity Module
[0424] SLA Service Level Agreement
[0425] SM Session Management
[0426] SMF Session Management Function
[0427] SN Secondary Node
[0428] S-NSSAI Single Network Slice Selection Assistance Information
[0429] SSB Synchronization Signal Block
[0430] SSM Source Specific IP Multicast address
[0431] SSC Session and Service Continuity
[0432] SRB Signaling Radio Bearer
[0433] SUPI Subscription Permanent Identifier
[0434] TA Tracking Area
[0435] TAI Tracking Area Identity
[0436] TE Terminal Equipment
[0437] TM Transparent Mode
[0438] TMGI Temporary Mobile Group Identity
[0439] TS Technical Specification
[0440] UDM Unified Data Manager
[0441] UDR Unified Data Repository
[0442] UE User Equipment
[0443] UL Uplink
[0444] UM Unacknowledged Mode
[0445] UP User Plane
[0446] UPF User Plane Function
[0447] URLLC Ultra-Reliable and Low-Latency Communication
[0448] URSP UE Route Selection Policy
Claims
1.A first entity for monitoring an artificial intelligence / machine learning (AI / ML) model deployable in a network, the first entity comprising:a transceiver; anda processor configured to:obtain information relating to the AI / ML model, andidentify one or more metric for monitoring the AI / ML model from among a plurality of metrics categorized based on one or more of the AI / ML model, a use case for the AI / ML model, functionality of the AI / ML model, interdependence between metrics, or data obtained for monitoring, based on the information relating to the AI / ML model,wherein each metric of the plurality of metrics is included in one or more of the plurality of categories.2.The first entity of claim 1,wherein the processor is further configured to:identify at least one first metric included in a first category of the plurality of categories and at least one second metric included in a second category of the plurality of categories, based on the information relating to the AI / ML model, andidentify at least one metric for each of the plurality of categories, based on the information relating to the AI / ML model, andwherein the plurality of categories includes one or more of:a category including metrics relating to the ability of the AI / ML model to maintain consistent results over time,a category including metrics relating to accuracy and effectiveness of outputs of the AI / ML model, anda category including metrics relating to operational behaviour of a system in which the AI / ML model is deployed.3.The first entity of claim 2, wherein one or more of:the category including metrics relating to the ability of the AI / ML model to maintain consistent results over time comprises at least one of:a metric for identifying a prior probability shift,Population Stability Index (PSI),Divergence Index,Maximum Mean Discrepancy (MMD),Kullback-Leibler (KL) Divergence,Energy Distance (ED),Cramer-von Mises (CvM) Test,Anderson-Darling Test,a metric for identifying a covariate shift,Characteristic Stability Index (CSI), orWasserstein's Distance;the category including metrics relating to accuracy and effectiveness of outputs of the AI / ML model comprises one or more of:Root Mean Squared Error (RMSE),R-Square,F1-Score,Mean Squared Error (MSE),Area Under the Receiver Operating Characteristic curve (AUC-ROC),Cross-Entropy Loss,Kullback-Leibler Divergence Loss,Hit Rate,Diversity,Novelty, ora metric for identifying concept shift; andthe category including metrics relating to operational behaviour of a system in which the AI / ML model is deployed comprises one or more of:ML throughput,ML Latency,Resource Usage,Power Consumption,Scalability,Availability,a metric for identifying system health issues,a metric for identifying endpoint latency, ora metric for identifying input / output (I / O), memory, and / or CPU issues.4.The first entity of claim 1, wherein the processor is further configured to:transmit first information on the identified one or more metric to a second entity,receive second information on at least one selected metric, wherein the at least one selected metric is included among the one or more metric,monitor the at least one selected metric, andprovide a recommendation or carry out the recommendation, upon detecting at least one monitored metric to exhibit unusual behaviour,wherein the at least one monitored metric is detected to exhibit unusual behaviour when it is detected, by the processor, that:measurements associated with the at least one monitored metric fall outside of expected ranges, or are above / below specified threshold,the at least one monitored metric is underperforming, orthe at least one monitored metric diverges from expected results, andwherein the recommendation comprises one of:updating the AI / ML model,retraining the AI / ML model,switching to another AI / ML model,maintaining deployment of the AI / ML model in the network,deactivating the AI / ML model, orimplementing a fallback solution, andwherein in case that the at least one monitored metric exhibiting unusual behaviour includes only metric(s) relates to accuracy and effectiveness of outputs of the AI / ML model, the at least one processor is configured to provide a recommendation to retrain the AI / ML model to the second entity,wherein in case that the at least one monitored metric exhibiting unusual behaviour includes a metric relating to accuracy and effectiveness of outputs of the AI / ML model and a metric relating to the ability of the AI / ML model to maintain consistent results over time, the at least one processor is configured to provide a recommendation to implement a fallback solution for ensuring safety and reliability of a system in which the AI / ML model is deployed and update, or a recommendation to update the AI / ML model,wherein the recommendation is provided to the second entity, orwherein the information relating to the AI / ML model is obtained from the second entity.5.The first entity of claim 4, wherein the processor is further configured to:identify a proposed monitoring frequency for each of the one or more metric based on third information,include the proposed monitoring frequency in the first information, andidentify the one or more metric based on a type of the AI / ML model and / or an application being used in associated with the AI / ML model as indicated by the information relating to the AI / ML model, andwherein the second information comprises a monitoring frequency for each of the at least one selected metric,wherein the AI / ML model is deployed at, or to be deployed at, the second entity and / or a third entity in the network, orwherein the one or more identified metric comprises at least one recommended monitoring metric identified, by the at least one processor, based on one or more of the AI / ML model, expected performance of the AI / ML model or environmental context.6.A second entity for managing an artificial intelligence / machine learning (AI / ML) model deployable in a network, the second entity comprising:a transceiver; anda processor configured to:receive, from the first entity, first information on one or more metric identified by the first entity,select at least one metric to be monitored from among the one or more metric, andtransmit, to the first entity, second information on the at least one selected metric;wherein each of the one or more metric is included in one or more of a plurality of categories, having been categorised based on one or more of the AI / ML model, a use case for the AI / ML model, functionality of the monitored AI / ML model, interdependence between metrics, or data obtained for monitoring,wherein the plurality of categories comprises one or more of:a category including metrics relating to the ability of the AI / ML model to maintain consistent results over time,a category including metrics relating to accuracy and effectiveness of outputs of the AI / ML model, ora category including metrics relating to operational behaviour of a system in which the AI / ML model is deployed, orwherein one or more of:the category including metrics relating to the ability of the AI / ML model to maintain consistent results over time comprises at least one of:a metric for identifying a prior probability shift,Population Stability Index (PSI),Divergence Index,Maximum Mean Discrepancy (MMD),Kullback-Leibler (KL) Divergence,Energy Distance (ED),Cramer-von Mises (CvM) Test,Anderson-Darling Test,a metric for identifying a covariate shift,Characteristic Stability Index (CSI), orWasserstein's Distance;the category including metrics relating to accuracy and effectiveness of outputs of the AI / ML model comprises one or more of:Root Mean Squared Error (RMSE),R-Square,F1-Score,Mean Squared Error (MSE),Area Under the Receiver Operating Characteristic curve (AUC-ROC),Cross-Entropy Loss,Kullback-Leibler Divergence Loss,Hit Rate,Diversity,Novelty, ora metric for identifying concept shift; andthe category including metrics relating to operational behaviour of a system in which the AI / ML model is deployed comprises one or more of:ML throughput,ML Latency,Resource Usage,Power Consumption,Scalability,Availability,a metric for identifying system health issues,a metric for identifying endpoint latency, ora metric for identifying input / output (I / O), memory, and / or CPU issues.7.The second entity of claim 6, wherein the processor is further configured to:receive, from the first entity, a recommendation relating to unusual behaviour exhibited by at least one metric, among the at least one selected metric, being monitored at the first entity, andcarry out the recommendation, andwherein the recommendation comprises one of:updating the AI / ML model,retraining the AI / ML model,switching to another AI / ML model,maintaining deployment of the AI / ML model in the network,deactivating the AI / ML model, orimplementing a fallback solution.8.The second entity of claim 6,wherein the processor is further configured to:select a monitoring frequency for each of the at least one selected metric, andinclude the selected monitoring frequency for each of the at least one selected metric in the second information, andwherein the first information comprises a proposed monitoring frequency for each of the one or more metric,wherein the processor is configured to transmit information relating to the AI / ML model to the first entity, prior to receiving the first information,wherein the information relating to the AI / ML model indicates one or more key performance indicator (KPI) for the AI / ML model, orwherein the AI / ML model is deployed at, or to be deployed at, the second entity and / or a third entity in the network.9.A method of first entity for monitoring an artificial intelligence / machine learning (AI / ML) model deployable in a network, the method comprising:obtaining information relating to the AI / ML model, andidentifying one or more metric for monitoring the AI / ML model from among a plurality of metrics categorized based on one or more of the AI / ML model, a use case for the AI / ML model, functionality of the AI / ML model, interdependence between metrics, or data obtained for monitoring, based on the information relating to the AI / ML model,wherein each metric of the plurality of metrics is included in one or more of the plurality of categories.10.The method of claim 9, the method further comprising:identifying at least one first metric included in a first category of the plurality of categories and at least one second metric included in a second category of the plurality of categories, based on the information relating to the AI / ML model, andidentifying at least one metric for each of the plurality of categories, based on the information relating to the AI / ML model, andwherein the plurality of categories includes one or more of:a category including metrics relating to the ability of the AI / ML model to maintain consistent results over time,a category including metrics relating to accuracy and effectiveness of outputs of the AI / ML model, ora category including metrics relating to operational behaviour of a system in which the AI / ML model is deployed.11.The method of claim 10, wherein one or more of:the category including metrics relating to the ability of the AI / ML model to maintain consistent results over time comprises at least one of:a metric for identifying a prior probability shift,Population Stability Index (PSI),Divergence Index,Maximum Mean Discrepancy (MMD),Kullback-Leibler (KL) Divergence,Energy Distance (ED),Cramer-von Mises (CvM) Test,Anderson-Darling Test,a metric for identifying a covariate shift,Characteristic Stability Index (CSI), orWasserstein's Distance;the category including metrics relating to accuracy and effectiveness of outputs of the AI / ML model comprises one or more of:Root Mean Squared Error (RMSE),R-Square,F1-Score,Mean Squared Error (MSE),Area Under the Receiver Operating Characteristic curve (AUC-ROC),Cross-Entropy Loss,Kullback-Leibler Divergence Loss,Hit Rate,Diversity,Novelty, ora metric for identifying concept shift; andthe category including metrics relating to operational behaviour of a system in which the AI / ML model is deployed comprises one or more of:ML throughput,ML Latency,Resource Usage,Power Consumption,Scalability,Availability,a metric for identifying system health issues,a metric for identifying endpoint latency, ora metric for identifying input / output (I / O), memory, and / or CPU issues.12.The method of claim 10, wherein the method further comprising:transmitting first information on the identified one or more metric to a second entity,receiving second information on at least one selected metric, wherein the at least one selected metric is included among the one or more metric,monitoring the at least one selected metric, andproviding a recommendation or carry out the recommendation, upon detecting at least one monitored metric to exhibit unusual behaviour,wherein the at least one monitored metric is detected to exhibit unusual behaviour when it is detected, by the processor, that:measurements associated with the at least one monitored metric fall outside of expected ranges, or are above / below specified threshold,the at least one monitored metric is underperforming, orthe at least one monitored metric diverges from expected results,wherein the recommendation comprises one of:updating the AI / ML model,retraining the AI / ML model,switching to another AI / ML model,maintaining deployment of the AI / ML model in the network,deactivating the AI / ML model, orimplementing a fallback solution,wherein in case that the at least one monitored metric exhibiting unusual behaviour includes only metric(s) relates to accuracy and effectiveness of outputs of the AI / ML model, the at least one processor is configured to provide a recommendation to retrain the AI / ML model to the second entity,wherein in case that the at least one monitored metric exhibiting unusual behaviour includes a metric relating to accuracy and effectiveness of outputs of the AI / ML model and a metric relating to the ability of the AI / ML model to maintain consistent results over time, the at least one processor is configured to provide a recommendation to implement a fallback solution for ensuring safety and reliability of a system in which the AI / ML model is deployed and update, or a recommendation to update the AI / ML model,wherein the recommendation is provided to the second entity, orwherein the information relating to the AI / ML model is obtained from the second entity.13.The method of claim 12, wherein the method further comprising:identifying a proposed monitoring frequency for each of the one or more metric based on third information,including the proposed monitoring frequency in the first information,identifying the one or more metric based on a type of the AI / ML model and / or an application being used in associated with the AI / ML model as indicated by the information relating to the AI / ML model, andwherein the second information comprises a monitoring frequency for each of the at least one selected metric,wherein the AI / ML model is deployed at, or to be deployed at, the second entity and / or a third entity in the network, orwherein the one or more identified metric comprises at least one recommended monitoring metric identified, by the at least one processor, based on one or more of the AI / ML model, expected performance of the AI / ML model or environmental context.14.A method of second entity for managing an artificial intelligence / machine learning (AI / ML) model deployable in a network, the method comprising:receiving, from the first entity, first information on one or more metric identified by the first entity;selecting at least one metric to be monitored from among the one or more metric, andtransmitting, to the first entity, second information on the at least one selected metric;wherein each of the one or more metric is included in one or more of a plurality of categories, having been categorised based on one or more of the AI / ML model, a use case for the AI / ML model, functionality of the monitored AI / ML model, interdependence between metrics, or data obtained for monitoring,wherein the plurality of categories comprises one or more of:a category including metrics relating to the ability of the AI / ML model to maintain consistent results over time,a category including metrics relating to accuracy and effectiveness of outputs of the AI / ML model, ora category including metrics relating to operational behaviour of a system in which the AI / ML model is deployed, orwherein one or more of:the category including metrics relating to the ability of the AI / ML model to maintain consistent results over time comprises at least one of:a metric for identifying a prior probability shift,Population Stability Index (PSI),Divergence Index,Maximum Mean Discrepancy (MMD),Kullback-Leibler (KL) Divergence,Energy Distance (ED),Cramer-von Mises (CvM) Test,Anderson-Darling Test,a metric for identifying a covariate shift,Characteristic Stability Index (CSI), orWasserstein's Distance;the category including metrics relating to accuracy and effectiveness of outputs of the AI / ML model comprises one or more of:Root Mean Squared Error (RMSE),R-Square,F1-Score,Mean Squared Error (MSE),Area Under the Receiver Operating Characteristic curve (AUC-ROC),Cross-Entropy Loss,Kullback-Leibler Divergence Loss,Hit Rate,Diversity,Novelty, ora metric for identifying concept shift; andthe category including metrics relating to operational behaviour of a system in which the AI / ML model is deployed comprises one or more of:ML throughput,ML Latency,Resource Usage,Power Consumption,Scalability,Availability,a metric for identifying system health issues,a metric for identifying endpoint latency, ora metric for identifying input / output (I / O), memory, and / or CPU issues.15.The method of claim 14, wherein the method further comprising:receiving, from the first entity, a recommendation relating to unusual behaviour exhibited by at least one metric, among the at least one selected metric, being monitored at the first entity, andcarrying out the recommendation, andwherein the recommendation comprises one of:updating the AI / ML model,retraining the AI / ML model,switching to another AI / ML model,maintaining deployment of the AI / ML model in the network,deactivating the AI / ML model, orimplementing a fallback solution.