Updating drift threshold values

By leveraging explainable machine learning to dynamically update drift threshold values based on feature importance, the solution addresses the limitations of existing drift detection techniques, enhancing the detection of data drifts and optimizing ML model retraining in dynamic environments.

WO2025122034A1PCT designated stage expired Publication Date: 2025-06-12TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/SE2023/051216
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Existing drift detection techniques in machine learning models, particularly in intent-based optimization for radio access networks (RAN), face challenges such as the use of static thresholds, lack of adaptability to changing feature importance, and limited ability to detect anomalies and drifts in dynamic environments.

Method used

The proposed solution involves using explainable machine learning (XAI) to dynamically update drift threshold values based on the importance of features, allowing for real-time adaptation to changes in data distributions and intent parameters.

Benefits of technology

This approach enhances the detection of data drifts, reduces false alarms, and optimizes retraining of ML models, ensuring continuous reliable performance in complex and dynamic environments like telecommunications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SE2023051216_12062025_PF_FP_ABST
    Figure SE2023051216_12062025_PF_FP_ABST
Patent Text Reader

Abstract

A method is provided. The method comprises obtaining a plurality of feature importance (FI) values associated with a plurality of features. Each FI value indicates an importance of one or more of the plurality of features with respect to an output of a machine learning (ML) model. The method further comprises obtaining one or more drift threshold values each of which is for detecting data drift of input data of the ML model. The input data of the ML model comprises values of the plurality of features. The method further comprises, based on the plurality of FI values, determining whether to update said one or more drift threshold values, and based on the determination, updating at least one of said one or more drift threshold values, thereby generating updated one or more drift threshold values.
Need to check novelty before this filing date? Find Prior Art

Description

UPDATING DRIFT THRESHOLD VALUESTECHNICAL FIELD

[0001] This disclosure relates to updating drift threshold values.BACKGROUND

[0002] Machine learning (ML) models are becoming commonplace in many product and service offerings. In many cases, the operating conditions of an ML model do not remain stationary. Instead, they drift over time. Two widely observed types of drift are data drift and concept drift.

[0003] Data drift occurs when the distribution of input data of the ML model evolves such that it no longer resembles the distribution of training input data that was used for training the ML model. Concept drift occurs when the relationship between input data of the ML model and output data of the ML model deviates from the relationship which was encoded by training data of the ML model. Data drift and concept drift, as well as other types of drift, can severely degrade the performance of the ML model.

[0004] Anomalies and data drifts can pose significant challenges to re-training and stable performance of ML models. Thus, in maintaining the model integrity of ML models, data monitoring, especially data drift detection, plays a crucial role.

[0005] The development of adaptive drift detection mechanisms for data in the field of telecommunication is rooted in the broader evolution of ML techniques targeting anomalies and drift detection across different domains. The complexity and dynamic nature of the telecom data requires advanced techniques not just to detect data drift but also to provide insights on the reasons and / or sources of the drift.

[0006] Intent-based optimization is an important driver for future radio access networks (RAN). With intents, an operator defines a set of high-level requirements for RAN behavior, which gets translated into an optimization objective in terms of system key performance indicators (KPIs). Intents can vary across deployments and across time for a given deployment. Dynamic optimization for a given intent, for example using reinforcement learning (RL), is a topic of active research as disclosed in reference [3], RL agent(s) solve for a high-level intentby taking sequential actions that guide a RAN system towards a desirable state. RL agent(s) may be trained for one or many intents prior to deployment. Since both the intent and the deployment environment can drift over time, RL agent(s) will need to be retrained periodically. However, so far, there has been a lack of efficient techniques that address drift detection in intent-based optimization in RAN.

[0007] Explainable artificial intelligence (XAI) is becoming a key component in academic research and industrial applications, emphasizing validation and comprehension of Al model predictions. Recently, there has been some work at the intersection of XAI and drift detection. One approach that stands out in recent literature is the methodology for domain-aware explainable anomaly and drift detection for multi-variate raw data using a constraint repository, as discussed in reference [1], This approach harnesses the power of a domain-indexed knowledge graph-based constraint repository. The system recognizes data anomalies and drift by comparing data against domain-specific constraints retrieved from this knowledge graph. The system also provides explanations for detected anomalies, referencing the violated constraints to offer clear insights into the root causes of anomalies and drift.

[0008] Another application of XAI in drift detection was demonstrated in the context of healthcare during the COVID-19 pandemic. For example, reference [2] discloses the dynamic nature of clinical settings, and how underlying data distributions can vary over time, leading to what is termed as data drift. More critically, the relationship between patient episode characteristics and clinical outcomes — concept drift — may also evolve. Using the COVID-19 pandemic as an example, the researchers highlighted the utility of explainable ML, particularly SHAP values (e.g., disclosed in reference [5]) to monitor such data drifts.

[0009] In data drift detection, the main goal is determining whether the distribution P(z) has drifted from a specific reference distribution P_ref(z). Depending on the drift problem, z can be input data of an ML model, true output data of the ML model, or a certain function of output data of the ML model. In real world problems, it is not realistic to expect samples from P(z) and P_ref(z) to be identical. In order to determine whether the difference between P(z) and P_ref(z) is due to some kind of drift or due to natural process noise, hypothesis testing techniques are widely used. Some examples of hypothesis testing techniques are Chi-Squared, Kolmogorov- Smirnov, Cramer-von Mises, Fisher’s exact test, Least-Squares Density Difference (LSDD) andMaximum Mean Discrepancy (MMD) (as disclosed in reference [6]). A basic process of a statistical two-sample hypothesis testing is shown in FIG. 6. Obtained p-value is compared to a threshold to decide between two hypotheses. Alpha represents the desired false positive rate.SUMMARY

[0010] Certain challenges presently exist in the existing drift detection techniques. Some of the challenges are as follows:

[0011] 1. Static Threshold: Some existing drift detection techniques often deploy fixed or static thresholds for detecting drifts. However, these fixed thresholds are not adjusted even when importance and / or relevance of specific features in data are changed. As a result, some data drifts may be undetected or may lead to an excessive number of false alarms and expensive retraining cycles, especially when the data drifts occur in critical features that have greater implications on system performance.

[0012] 2. Adaptive Threshold: Some existing drift detection techniques require a careful parameter tuning, and they may not always be well-suited for complex, multi-feature environments. They are largely statistical in nature and do not necessarily consider the structure of the data they are examining. Furthermore, they do not take into account the contextual or domain-specific importance of features, which is particularly crucial in applications like telecommunications where some features may have a greater impact on system performance and reliability than others. Finally, the complexity and computational overhead of these methods may also be a limiting factor, especially in real-time applications. As applications and technologies grow in complexity and the underlying data changes over time, it is vital that the algorithms can adapt to those changes and ensure continuous reliable performance.

[0013] 3. Limitations of Domain-Based Rules: These techniques are based on past and existing knowledge of an industry or domain. They don’t get automatically updated with new changes or emerging trends in that field. Because of this, there might be new situations or patterns that this system may not be able to recognize simply because they aren’t part of the “known facts” in its library. This means the system might miss out on detecting certain anomalies or drifts that arise from these new patterns.

[0014] 4. Lack of Adaptive Mechanisms: In the context of dynamic environments like telecommunication, performance of models needs to be constantly monitored and adjusted.Current technologies do recognize the need for retraining predictive models. However, there isn’t always an efficient, automated, and adaptive mechanism to flag when and how the retraining should occur.

[0015] 5. Intent-based RAN Optimization: The ML model performance degrades with changes in the intent as well as the RAN environment. Current approaches for intent-based optimization do not take into account the impact of intent on drift detection and retraining pipelines.

[0016] 6. Lack of Interpretability: Several drift detection methods do not provide insights into which features are contributing to the drift.

[0017] Accordingly, in one aspect of some embodiments of this disclosure, there is provided a method comprising obtaining a plurality of feature importance (FI) values associated with a plurality of features, wherein each FI value indicates an importance of one or more of the plurality of features with respect to an output of a machine learning (ML) model. The method further comprises obtaining one or more drift threshold values each of which is for detecting data drift of input data of the ML model, wherein the input data of the ML model comprises values of the plurality of features. The method further comprises, based on the plurality of FI values, determining whether to update said one or more drift threshold values; and based on the determination, updating at least one of said one or more drift threshold values, thereby generating updated one or more drift threshold values.

[0018] In another aspect, there is provided a computer program comprising instructions which when executed by processing circuitry cause the processing circuitry to perform the method of any one of the above embodiments.

[0019] In a different aspect, there is provided a carrier containing the computer program of the above embodiment, wherein the carrier is one of an electronic signal, an optical signal, a radio signal, and a computer readable storage medium.

[0020] In a different aspect, there is provided an apparatus being configured to obtain a plurality of feature importance (FI) values associated with a plurality of features, wherein each FI value indicates an importance of one or more of the plurality of features with respect to an output of a machine learning (ML) model. The apparatus is further configured to obtain one or more drift threshold values each of which is for detecting data drift of input data of the ML model, wherein the input data of the ML model comprises values of the plurality of features.The apparatus is further configured to, based on the plurality of FI values, determine whether to update said one or more drift threshold values; and based on the determination, update at least one of said one or more drift threshold values, thereby generating updated one or more drift threshold values.

[0021] In a different aspect, there is provided an apparatus comprising a processing circuitry and a memory, said memory containing instructions executable by said processing circuitry, whereby the apparatus is operative to perform the method of any one of the above embodiments.

[0022] Embodiments of this disclosure provide one or more of the following advantages:

[0023] 1. Dynamic Threshold Adaptation: Instead of using static thresholds, the system leverages explainable ML to adaptively set and modify thresholds for detecting drifts, based on real-time feature importance. This ensures that more critical features, which impact system performance significantly, are monitored with refined sensitivity.

[0024] 2. Computational Efficiency: While increasing the depth of monitoring, some embodiments of this disclosure ensure computational efficiency, enabling its application even in large-scale, real-time telecom environments.

[0025] 3. Optimized Model Retraining: Some embodiments of this disclosure allow performing retraining of ML models only when necessary.

[0026] 4. Usable in Intent-based Optimization: According to some embodiments, drift detection thresholds can be adapted with respect to intent parameters. This enables better overall system performance by optimizing the model retraining workflow.

[0027] 5. Adaptable to a Variety of ML algorithms: Some explainability techniques such as Kernel SHAP can explain any function that produces a prediction. With that in mind, some embodiments of this disclosure are adaptable to many types of ML algorithms including reinforcement learning (RL), as explained below.

[0028] 6. Ability to Dynamically Weigh the Importance of Different Features in Real-Time: Unlike traditional adaptive thresholding techniques, which often require manual tuning and lack domain-specific feature prioritization, the XALbased approach according to some embodiments self-adjusts without extensive tuning. This enables more accurate and more efficient detection of data drifts across a variety of domains. The approach is particularly beneficial in high-stakes, complex environments such as telecommunications whereunderstanding the relative importance of different data features is crucial.BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate various embodiments.

[0030] FIG. 1 shows a block diagram of an intent-based system according to some embodiments.

[0031] FIG. 2 shows a process according to some embodiments.

[0032] FIG. 3 shows feature importance values of input features.

[0033] FIG. 4 shows a process according to some embodiments.

[0034] FIG. 5 shows an apparatus according to some embodiments.

[0035] FIG. 6 shows a basic process of a statistical two-sample hypothesis testing.DETAILED DESCRIPTION

[0036] FIG. 1 A shows a block diagram of a simplified intent-based system 100. As shown in FIG. 1 A, the system 100 comprises an intent provider 102 and an ML model 104.

[0037] The intent provider 102 may be configured to provide, to the ML model 104, a high-level user intent related to an entity (e.g., a base station, a manufacturing process, etc.). In case the system 100 is used in the field of telecommunication, one specific example of the intent is “reducing energy consumption of a particular base station by 10%. ”

[0038] Based on the received intent, the ML model 104 may determine configurations of the entity, which may result in achieving the intent. For instance, in the above example, the ML model 104 may be configured to determine configurations of the base station, which would result in achieving the intent of reducing the energy consumption of the base station by 10%. One example of the ML 104 is a reinforcement learning (RL) agent.

[0039] FIG. IB shows an example of one aspect of the operation of the ML model 104 according to some embodiments. In FIG. IB, current input data 150 is provided to the ML model 104. The current input data 150 comprises values of five input features 112, 114, 116, 118, and 120. The ML model 104 is configured to generate, based on the values of the five input features, output data 160. For instance, in the above example, the ML model 104 may be configured toreceive values of five parameters of the base station (e.g., a number of user equipments (UEs) connected to the base station, a geographical region where the base station is located, signal processing overhead, power amplifier efficiency, etc.) and predict an amount of energy consumption of the base station based on the received values of the base station parameters.

[0040] As explained above, in some scenarios, data drift may occur in one or more of the values of the input features 112-120. For example, let’s assume that the ML model 104 was trained using training input data, and the value of the input feature 112 included in the training input data indicated that the number of UEs connected to the base station is 10. Let’s further assume that the value of the input feature 112 included in current input data 150 is 1000.

[0041] In this example, the significant difference between the value of the input feature 112 in the training input data and the value of the input feature 112 in the current input data constitutes data drift, and this data drift may result in the ML model to output bad output data (e.g., incorrectly predicting the amount of energy consumption of the base station).

[0042] More specifically, in the above example, the ML model was trained to predict the amount of energy consumption of the base station when the number of UEs connected to the base station is within a range between 0 and 100. Thus, the ML model may not accurately predict the amount of energy consumption of the base station when the number of UEs connected to the base station substantially deviates from this range. In this case, it may be desirable to retrain the ML model such that the ML model can accurately predict the amount of energy consumption of the base station when the number of UEs connected to the base station is within a range of, for example, 900 and 1100. Therefore, it is important to correctly detect data drift which triggers retraining of the ML model.

[0043] In detecting data drift of current input data, a drift threshold value may be used. More specifically, in case the difference between the value of an input feature in reference input data (e.g., the training input data discussed above) and the value of the input feature in the current input data is greater than or equal to the drift threshold value, it may be determined that data drift has occurred in the input feature of the current input data. On the other hand, in case the difference is less than the drift threshold value, it may be determined that data drift has not occurred in the input feature of the current input data.

[0044] For instance, in the above example, let’s assume that the drift threshold value for the input feature 112 (i.e., the number of UEs connected to the base station) is 500. In thisexample, since the difference between the value of the input feature 112 of the training input data — i.e., 10 — and the value of the input feature 112 of the current input data — i.e., 1000 — is 990 which is greater than 500, it may be determined that data drift has occurred in the input feature 112 of the current input data.

[0045] As discussed above, after detecting drift, the ML model may need to be retrained, and after retraining the ML model, the drift threshold value may need to be updated. For instance, in the above example, after retraining the ML model such that the ML model can accurately predict the amount of energy consumption of the base station when the number of UEs connected to the base station is within the range of 900 and 1100, no data drift should be detected when the value of the input feature 112 in the upcoming input data is 1000 again.

[0046] However, if the old drift threshold value, i.e., 500, is continuously used, since the difference between the value of the input feature 112 of the training input data — i.e., 10 — and the value of the input feature 112 of the upcoming input data — i.e., 1000 — is 990 which is greater than 500, it may be determined again that data drift has occurred in the input feature 112 of the upcoming input data, and thus unnecessary retraining of the ML model may be performed. In order to avoid performing this unnecessary retraining of the ML model, drift threshold values associated with input features may be updated. Also, in some embodiments, the reference input data used for detecting data drift may be updated. For instance, in the above example, the value of the input feature 112 of the training input data (i.e., the reference input data) was 10. In this example, the value of the input feature 112 for the reference input data may be updated to, for example, 500, to avoid performing the unnecessary retraining of the ML model.

[0047] In some cases, drift threshold values associated with different input features may need to be updated differently. This is because different input features may affect the output of the ML model differently. For example, in FIG. IB, 10% change in the value of the input feature 112 may result in 50% change in the value of the output data 160 of the ML model 104 while 10% change in the value of the input feature 120 may result in 1% change in the value of the output data 160. In this case, it may be desirable to determine that data drift has occurred even when the value of the input feature 112 changed a little bit while it may be desirable to determine that no data drift has occurred even when the value of the input feature 120 changed substantially.

[0048] In order to provide these different “sensitivities” for detecting drift for different input features, different threshold values may be provided for the different input features.Accordingly, in some embodiments of this disclosure, a process 200 shown in FIG. 2 is provided for updating and using different drift threshold values for different input features. Note that even though the embodiments of this disclosure are mainly explained using data drift as an example, the embodiments are equally applicable to any other type of drift (e.g., concept drift).

[0049] The process 200 may begin with step s202.

[0050] Step s202 - Deploying a Trained ML Model

[0051] The step s202 comprises deploying, in a target environment, a trained ML model which was trained using training input data (e.g., a pre-collected dataset). As mentioned above, one example of the ML model is an ML model for predicting power consumption of a base station. After performing the step s202, the process 200 may proceed to step s204.

[0052] Step s204 - Performing Ongoing Data Collection

[0053] The step s204 comprises performing ongoing data collection. The data collected may include current input data for the ML model, and the current input data may include values of input features. As mentioned above, in case the ML model is for predicting an amount of energy consumption of a base station, the current input data may comprise a value of the input feature 112 which indicates a number of UEs that are currently connected to the base station.

[0054] Even though FIG. 2 shows that the step s204 is performed after the step s202, in some embodiments, the step s204 may be performed before the step s202. Alternatively, the steps s202 and s204 may be performed simultaneously. After performing the step s204, the process 200 may proceed to step s206.

[0055] Step s206 - Determining Feature Importance (FI) Values of Input Features Using XAI

[0056] The step s206 comprises determining FI values of the input features of the ML model using XAI. Here, the determined FI value may be a global FI value of an input feature. The global FI values of the input features may indicate which input features have the most impact on the ML model’s predictions (i.e., the ML model’s output). Note that, in this disclosure, the term “global” FI values mean that the FI values are not associated with a single prediction but rather they correspond to a global representation of the ML model itself and / or an aggregate of many individual prediction feature rankings.

[0057] A global FI value may be calculated using many different explainability techniques. One example of such techniques is Shapley Additive Explanations (SHAP)technique. The SHAP technique is a perturbation-based explainability method that determines how changes to input data of the ML model impact the output of the ML model. In making such determination, the method calculates how each input feature value contributes to a given prediction. These importance contributions are known as “SHAP values” and can be positive or negative depending on whether the SHAP values push the ML model’s prediction to higher or lower values with respect to the ML model’s average output prediction.

[0058] By calculating SHAP values across a large number of samples that are representative of the input feature space and averaging their absolute values, a “global” importance for each input feature can be determined.

[0059] For an ML model having a single output node, this global FI can be interpreted as the average impact each input feature has on the ML model’s output magnitude. These FI values may be used for the drift detection thresholding which is described below.

[0060] For an ML model having multiple output nodes, each output node may have its own feature ranking since SHAP values are calculated with respect to each output node separately. In this case, an additional step may be taken to aggregate the FI values across the output nodes so that a single feature ranking can be used for the drift detection thresholding which is described below.

[0061] FIG. 3 shows a stacked bar plot of the FI values of ten input features obtained for SHAP calculations for an ML model having three output nodes. The figure shows that some input features are more, or less, important depending on the specific output node being analyzed. One way to calculate global FI values in multi-output models is to sum the individual FI values (e.g., mean SHAP values per input feature per output node) to obtain a single, summed importance value across output nodes. In FIG. 3, this would be a sum of the stacked bars for each y-axis input feature. Alternatively, one could take the average or median of each FI value across the output nodes. After calculating, the aggregated FI values may then be used as input to the drift detection thresholding described below. Note that the perturbation-based approach of calculating SHAP values make the technique adaptable to any type of ML model, which increases the flexibility of the embodiments to handle a variety of use-cases.

[0062] After performing the step s206, the process 200 may proceed to step s208.

[0063] Step s208 - Drift Detection Thresholding: Mapping FI Values to Thresholds

[0064] The step s208 comprises performing drift detection thresholding — i.e., convertingFI values of the input features into drift detection threshold values. The method used for such conversion may vary depending on the method used for drift detection. One example of the method used for such conversion is the Kolmogorov-Smirnov (KS) test.

[0065] The KS test is a commonly used drift detection method that uses the cumulative distribution functions of two samples to determine if they were drawn from the same distribution. To make the determination, a p-value threshold is selected prior to the KS test. The p-value dictates under what statistical significance the null hypothesis (i.e., that the two samples are drawn from the same distribution) is rejected. For a two-sample KS test, the null hypothesis is that the two samples are drawn from the same distribution. The null hypothesis is rejected if the p-value obtained from the KS test is less than a pre-defined p-value threshold.

[0066] In applying the KS test to the embodiments of this disclosure, the most important input features of the ML model may have the highest p-value thresholds (e.g., p < 0.1) when detecting drift. This condition would then trigger a drift detection when even the slightest deviations in the values of the most important input features are observed between the training input data and the current input data.

[0067] On the other hand, the lower importance input features of the ML model, which have less impact on the ML model’s predictions may have a lower p-value threshold (e.g., p < 0.001). Since those input features are less important to the ML model, it may be desirable that only highly significant KS detections are used to trigger drift. With this setup, drift detection (i.e., rejection of the null hypothesis) may occur more easily for high importance input features than for low importance features.

[0068] One way to determine the p-value thresholds from the FI values of the input features is to divide the input features into bins based on their ranking and then assign each bin a specific p-value. The p-values would be assigned in descending order according to the FI bins (i.e., the highest FI is assigned the highest p-value while the lowest FI bin is assigned the lowest p-value).

[0069] The number of bins to use for segmenting the input features may be adjustable by users as desired. Using FIG. 3 to illustrate this thresholding scheme when using only two importance bins assigned p-values [0.2, 0.05], the top five input features on the y-axis may be assigned a p-value of 0.2 and the bottom five input features may be assigned a p-value 0.05.

[0070] Alternatively, the p-values can be assigned based on a minimum / maximum scalingof the FI values, which maps the FI values to the desired minimum / maximum p-values. In this embodiment, the FI values of the input features may be scaled such that the most important feature is set to have the maximum p-value and the least important feature is set to have the minimum p-value considered. All feature importance values in-between may then be assigned with p-values based on a linear interpolation between the minimum / maximum p-values.

[0071] In other embodiments, p-values may be assigned only to the top N most important input features, where N is an adjustable parameter. In these embodiments, drift detection may only be tracked for the most important input features for the model’s predictions. All other input features may be considered inconsequential to the ML model’s predictions and may not be tracked for drift detection.

[0072] In other embodiments, an FI threshold may be used to determine a list of input features that are to be tracked for drift detection. For example, all input features that have global FI values greater than the FI threshold may be considered in the drift detection and all other input features may not be tracked for drift detection.

[0073] One advantage of these feature importance to p-value mapping schemes is that they are compatible with any drift detection method that uses p-value decision thresholds.

[0074] After performing the step s208, the process 200 may proceed to step s210.

[0075] Step s210 - Triggering Re-Training with Drift Detection

[0076] After determining the list of input features to use for drift detection and their corresponding drift detection thresholds, the step s210 may be performed. The step s210 comprises tracking incoming input data for drift and trigger retraining of the ML model in case drift is detected. Drift detection tests may be run on an hourly, daily, weekly, etc., basis depending on use-case requirements.

[0077] In some embodiments, the two samples compared in the drift detection may be the ML model’s training input data (i.e., training set feature distribution) and the incoming input data (i.e., the incoming dataset’s feature distribution). Drift can be determined individually for each input feature such that users will know which specific input features have drifted. If drift is detected in any input features, this may trigger the re-training of the ML model on a new training dataset that includes both the previous training set and the newly collected dataset that caused drift.

[0078] In some embodiments, the re-training of the ML model is not triggered if drift isdetected for one input feature. In these embodiments, the re-training of the ML model may be triggered only when a certain number (T) of input features have drifted. For instance, if T is set to 5, the ML model may be re-trained only when drift has been detected for 5 or more input features. T may be an adjustable parameter which may be determined based on how sensitive users want their re-training scheme to be.

[0079] Step s212 - Updating Drift Threshold Value

[0080] After obtaining the updated ML model (i.e., after retraining the ML model), the process 200 may proceed to step s212. The step s212 comprises updating the drift threshold value for the input feature according to the new FI value of the input feature.

[0081] The drift detection, model re-training, and drift thresholding processes may then repeated in a loop as new data is collected. In a summary:

[0082] At iteration 0, the initial drift detection metrics are inferred from the XAI method used.

[0083] At iteration > 0, the drift detection metrics are either continuously updated or remained the same depending on how significant the FI values of the input features change. The drift detection thresholds may be dynamically adjusted based on the magnitude of changes in FI values; specifically, updates are triggered when these changes exceed a predefined significance level.

[0084] Smart Drift Detection for Intent-based Radio Access Network (RAN) Optimization

[0085] With intent-based optimization, an RL model is trained on an objective function derived from a high-level RAN intent. A single RL model is typically trained for multiple intents. Subsequently, the RL model is monitored for good performance and retraining is triggered when degradation is observed or anticipated. Drift detection can be used for early detection of changes in operating conditions that can affect the performance of the RL model.

[0086] According to some embodiments, methods for determining drift detection parameters for intent-based optimization and triggering the retraining of the RL model are provided. The key insight is that each of the distribution of input features and the importance of input features (typically RAN KPIs) is a function of the intent. Since intent can change dynamically, the distribution and importance of input features may also vary based on the currently active intent. Drift detection needs to take this fact into account to avoid unnecessary retraining cycles and large performance degradation.

[0087] According to some embodiments, intent-conditioned drift detection process comprising the following sequential steps is provided.

[0088] The first step comprises, for an RL model trained to optimize for a set of intents, obtaining per-intent input FI values, for example, using SHAP analysis.

[0089] The second step comprises computing drift detection threshold values, as described above, for each intent.

[0090] The third step comprises, during online data collection from a model deployment, keeping track of the intent that generated the data, for example by including a unique intent identifier in addition to the (state, action, rewards) pairs.

[0091] The fourth step comprises applying drift detection on the recorded data, where per- intent detection thresholds are used in combination with per-intent online data.

[0092] The fifth step comprises, if drift is detected for one or more intents, triggering model retraining. This retraining can be configured in one of the following ways:

[0093] 1. Selecting data corresponding to the intents for which drift was detected and collected since the previous model retraining and using this data to retrain the existing model. This has the advantage of updating model parameters to address drift for specific intents. However, this can also cause the model to become biased to these intents, while degrading the performance for other intents that were not a part of this retraining cycle.

[0094] 2. Selecting data corresponding to all intents collected since the previous model retraining and using this data to retrain the existing model. This mitigates the risk of the model becoming biased towards a subset of intents over retraining cycles.

[0095] 3. Using data collected across multiple retraining cycles to retrain the current model, or to train a new intent-based model from scratch.

[0096] As explained above, intent-based frameworks are suitable use-case for smart drift detection. For these types of applications, the state space is very large, and re-training for any drift would be very computationally expensive. Adding the complexities of large action spaces, multi -objective reward functions and several intents, the space of possible drifts to be detected grows exponentially. Since not all drifts will be important enough to require action, a smart solution to prioritize drift may be essential for long term deployment of intent based frameworks.

[0097] FIG. 4 shows a process 400 according to some embodiments. The process 400 may begin with step s402. The step s402 comprises obtaining a plurality of feature importance, FI,values associated with a plurality of features, wherein each FI value indicates an importance of one or more of the plurality of features with respect to an output of a machine learning, ML, model. Step s404 comprises obtaining one or more drift threshold values each of which is for detecting data drift of input data of the ML model, wherein the input data of the ML model comprises values of the plurality of features. Step s406 comprises, based on the plurality of FI values, determining whether to update said one or more drift threshold values. Step s408 comprises, based on the determination, updating (s408) at least one of said one or more drift threshold values, thereby generating updated one or more drift threshold values.

[0098] In some embodiments, the plurality of FI values comprises a first FI value and a second FI value, the plurality of features comprises a first feature and a second feature, the first FI value indicates an importance of the first feature with respect to the output of the ML model, the second FI value indicates an importance of the second feature with respect to the output of the ML model, said one or more drift threshold values comprises a first drift threshold value and a second drift threshold value, the first drift threshold value is for detecting data drift of a value of the first feature, the second drift threshold value is for detecting data drift of a value of the second feature, the first drift threshold value is updated based on the first FI value and / or the second drift threshold value is updated based on the second FI value.

[0099] In some embodiments, the first and second drift threshold values are updated differently, or the first drift threshold value is updated while the second drift threshold value is not updated.

[0100] In some embodiments, whether to update the first drift threshold value is determined based on a difference between the first FI value and a previous FI value associated with the first feature, and the previous FI value indicates an importance of the first feature with respect to the output of the ML model, which was previously determined.

[0101] In some embodiments, whether to update the first drift threshold value is determined by: determining the difference between the first FI value and the previous FI value; comparing the difference to a FI threshold value; and determining to update the first drift threshold value in case the difference is greater than or equal to the FI threshold value or determining not to update the first threshold value in case the difference is less than the FI threshold value.

[0102] In some embodiments, an amount of updating the first drift threshold value is determined based on an amount of the difference.

[0103] In some embodiments, the ML model was trained using at least a training dataset of input data for the ML model, and the method comprises: obtaining a first dataset of input data for the ML model, which was not used for training the ML model; comparing the first dataset to the training dataset; and based on i) the comparison of the first dataset to the training dataset and ii) the updated drift threshold values, determining whether data drift has occurred in the first dataset.

[0104] In some embodiments, the training dataset of input data comprises a first value of the first feature, the first dataset of input data comprises a second value of the first feature, the process comprises comparing a difference between the first and second values of the first feature to the updated first drift threshold value, and whether data drift has occurred in the first dataset is determined based at least on the comparison of the difference to the updated first drift threshold value.

[0105] In some embodiments, the process further comprises determining that data drift has occurred in the first dataset; determining whether the occurred data drift satisfies a condition; and based on determining that the occurred data drift satisfies the condition, training the ML model.

[0106] In some embodiments, the ML model is retrained using at least the training dataset and the first dataset.

[0107] determining whether the occurred data drift satisfies the condition comprises determining whether data drift has occurred in values of at least N number of features included in the plurality of features, and N is a positive integer that is greater than or equal to 1.

[0108] In some embodiments, obtaining the plurality of FI values associated with the plurality of features comprises determining the plurality of FI values using an explainable artificial intelligence, XAI, technique.

[0109] In some embodiments, the XAI technique comprises shapley additive explanations, SHAP, technique.

[0110] In some embodiments, obtaining the plurality of FI values associated with the plurality of features comprises: obtaining a plurality of sample datasets; for each sample dataset,determining a contribution score of each feature, wherein the contribution score of each feature indicates how much a change in a value of the feature contributes to the output of the ML model when the sample dataset is used as input data of the ML model; and determining the FI value of each feature based on a combination of the contribution scores of the feature determined for the plurality of sample datasets.[OHl] In some embodiments, the output of the ML model comprises a first sub-output and a second sub-output, obtaining the first FI value comprises determining a combination of a third FI value indicating an importance of the first feature with respect to the first sub-output of the ML model and a fourth FI value indicating an importance of the first feature with respect to the second sub-output of the ML model, and obtaining the second FI value comprises determining a combination of a fifth FI value indicating an importance of the second feature with respect to the first sub-output of the ML model and a sixth FI value indicating an importance of the second feature with respect to the second sub-output of the ML model.

[0112] In some embodiments, the first FI value is determined based on a sum or an average of the third FI value and the fourth FI value, and the second FI value is determined based on a sum or an average of the fifth FI value and the sixth FI value.

[0113] In some embodiments, obtaining the plurality of FI values associated with the plurality of features comprises: dividing the plurality of features into a plurality of bins each of which is assigned with a ranking value; and for each of the plurality of features, determining the FI value based on the ranking value assigned to the bin to which the feature belongs.

[0114] In some embodiments, the plurality of FI values comprises a maximum FI value, a minimum FI value, and one or more middle FI values, each of said one or more middle FI values is determined based on an interpolation of the maximum FI value and the minimum FI value.

[0115] In some embodiments, the plurality of features includes a first set of features, the first set of features comprises features having TV highest FI values among the FI values of the plurality of features, where T is a positive integer, and each of said one or more drift threshold values is for detecting data drift in a value of each feature included in the first set of features.

[0116] In some embodiments, the process comprises comparing the FI value of each of the plurality of features to a FI threshold value, wherein obtaining said one or more drift thresholdvalues comprises obtaining a drift threshold value only for one or more features each of which has the FI value that is greater than or equal to the FI threshold value.

[0117] FIG. 5 is a block diagram of a node 500 capable of performing the process 200, according to some embodiments. As shown in FIG. 5, the node 500 may comprise: processing circuitry (PC) 502, which comprises one or more processors (P) 555 (e.g., one or more general purpose microprocessors and / or one or more other processors, such as an application specific integrated circuit (ASIC), field-programmable gate arrays (FPGAs), and the like), which processors may be co-located in a single housing or in a single data center or may be geographically distributed (e.g., the node 500 may be a distributed computing apparatus comprising two or more computers or a monolithic computing apparatus consisting of a single computer); at least one network interface 548 (e.g., a physical interface or air interface) comprising a transmitter (Tx) 545 and a receiver (Rx) 547 for enabling the node 500 to transmit data to and receive data from other nodes connected to network 110 (e.g., an Internet Protocol (IP) network) to which network interface 548 is connected (physically or wirelessly) (e.g., network interface 548 may be coupled to an antenna arrangement comprising one or more antennas for enabling the node 500 to wirelessly transmit / receive data); and a storage unit (a.k.a., “data storage system”) 508, which may include one or more non-volatile storage devices and / or one or more volatile storage devices. In embodiments where PC 502 includes a programmable processor, a computer readable storage medium (CRSM) 542 may be provided. CRSM 542 may store a computer program (CP) 543 comprising computer readable instructions (CRI) 544. CRSM 542 may be a non-transitory computer readable medium, such as, magnetic media (e.g., a hard disk), optical media, memory devices (e.g., random access memory, flash memory), and the like. In some embodiments, the CRI 544 of computer program 543 is configured such that when executed by PC 502, the CRI causes the node 500 to perform steps described herein (e.g., steps described herein with reference to the flow charts). In other embodiments, the node 500 may be configured to perform steps described herein without the need for code. That is, for example, PC 502 may consist merely of one or more ASICs. Hence, the features of the embodiments described herein may be implemented in hardware and / or software.

[0118] While various embodiments are described herein, it should be understood that they have been presented by way of example only, and not limitation. Thus, the breadth and scope ofthis disclosure should not be limited by any of the above-described exemplary embodiments. Moreover, any combination of the above-described elements in all possible variations thereof is encompassed by the disclosure unless otherwise indicated herein or otherwise clearly contradicted by context.

[0119] As used herein transmitting a message “to” or “toward” an intended recipient encompasses transmitting the message directly to the intended recipient or transmitting the message indirectly to the intended recipient (i.e., one or more other nodes are used to relay the message from the source node to the intended recipient). Likewise, as used herein receiving a message “from” a sender encompasses receiving the message directly from the sender or indirectly from the sender (i.e., one or more nodes are used to relay the message from the sender to the receiving node). Further, as used herein “a” means “at least one” or “one or more.”

[0120] Additionally, while the processes described above and illustrated in the drawings are shown as a sequence of steps, this was done solely for the sake of illustration. Accordingly, it is contemplated that some steps may be added, some steps may be omitted, the order of the steps may be re-arranged, and some steps may be performed in parallel.

[0121] Reference List

Claims

CLAIMS1. A method (400) comprising: obtaining (s402) a plurality of feature importance, FI, values associated with a plurality of features, wherein each FI value indicates an importance of one or more of the plurality of features with respect to an output of a machine learning, ML, model; obtaining (s404) one or more drift threshold values each of which is for detecting data drift of input data of the ML model, wherein the input data of the ML model comprises values of the plurality of features; based on the plurality of FI values, determining (s406) whether to update said one or more drift threshold values; and based on the determination, updating (s408) at least one of said one or more drift threshold values, thereby generating updated one or more drift threshold values.

2. The method of claim 1, wherein the plurality of FI values comprises a first FI value and a second FI value, the plurality of features comprises a first feature and a second feature, the first FI value indicates an importance of the first feature with respect to the output of the ML model, the second FI value indicates an importance of the second feature with respect to the output of the ML model, said one or more drift threshold values comprises a first drift threshold value and a second drift threshold value, the first drift threshold value is for detecting data drift of a value of the first feature, the second drift threshold value is for detecting data drift of a value of the second feature, the first drift threshold value is updated based on the first FI value and / or the second drift threshold value is updated based on the second FI value.

3. The method of claim 2, wherein the first and second drift threshold values are updated differently, orthe first drift threshold value is updated while the second drift threshold value is not updated.

4. The method of claim 2 or 3, wherein whether to update the first drift threshold value is determined based on a difference between the first FI value and a previous FI value associated with the first feature, and the previous FI value indicates an importance of the first feature with respect to the output of the ML model, which was previously determined.

5. The method of claim 4, wherein whether to update the first drift threshold value is determined by: determining the difference between the first FI value and the previous FI value; comparing the difference to a FI threshold value; and determining to update the first drift threshold value in case the difference is greater than or equal to the FI threshold value or determining not to update the first threshold value in case the difference is less than the FI threshold value.

6. The method of any one of claims 4 or 5, wherein an amount of updating the first drift threshold value is determined based on an amount of the difference.

7. The method of any one of claims 1-6, wherein the ML model was trained using at least a training dataset of input data for the ML model, and the method comprises: obtaining a first dataset of input data for the ML model, which was not used for training the ML model; comparing the first dataset to the training dataset; and based on i) the comparison of the first dataset to the training dataset and ii) the updated drift threshold values, determining whether data drift has occurred in the first dataset.

8. The method of claim 7 when claim 7 depends on claim 2, wherein the training dataset of input data comprises a first value of the first feature, the first dataset of input data comprises a second value of the first feature, the method comprises comparing a difference between the first and second values of the first feature to the updated first drift threshold value, and whether data drift has occurred in the first dataset is determined based at least on the comparison of the difference to the updated first drift threshold value.

9. The method of claim 7 or 8, the method further comprising: determining that data drift has occurred in the first dataset; determining whether the occurred data drift satisfies a condition; and based on determining that the occurred data drift satisfies the condition, training the ML model.

10. The method of claim 9, wherein the ML model is retrained using at least the training dataset and the first dataset.

11. The method of claim 9 or 10, wherein determining whether the occurred data drift satisfies the condition comprises determining whether data drift has occurred in values of at least N number of features included in the plurality of features, andN is a positive integer that is greater than or equal to 1.

12. The method of any one of claims 1-11, wherein obtaining the plurality of FI values associated with the plurality of features comprises determining the plurality of FI values using an explainable artificial intelligence, XAI, technique.

13. The method of any one of claims 1-12, wherein obtaining the plurality of FI values associated with the plurality of features comprises: obtaining a plurality of sample datasets;for each sample dataset, determining a contribution score of each feature, wherein the contribution score of each feature indicates how much a change in a value of the feature contributes to the output of the ML model when the sample dataset is used as input data of the ML model; and determining the FI value of each feature based on a combination of the contribution scores of the feature determined for the plurality of sample datasets.

14. The method of any one of claims 2-13, wherein the output of the ML model comprises a first sub-output and a second sub-output, obtaining the first FI value comprises determining a combination of a third FI value indicating an importance of the first feature with respect to the first sub-output of the ML model and a fourth FI value indicating an importance of the first feature with respect to the second suboutput of the ML model, and obtaining the second FI value comprises determining a combination of a fifth FI value indicating an importance of the second feature with respect to the first sub-output of the ML model and a sixth FI value indicating an importance of the second feature with respect to the second sub-output of the ML model.

15. An apparatus (500) being configured to: obtain (s402) a plurality of feature importance, FI, values associated with a plurality of features, wherein each FI value indicates an importance of one or more of the plurality of features with respect to an output of a machine learning, ML, model; obtain (s404) one or more drift threshold values each of which is for detecting data drift of input data of the ML model, wherein the input data of the ML model comprises values of the plurality of features; based on the plurality of FI values, determine (s406) whether to update said one or more drift threshold values; and based on the determination, update (s408) at least one of said one or more drift threshold values, thereby generating updated one or more drift threshold values.

16. The apparatus of claim 15, wherein the apparatus is further configured to perform the method of any one of claims 2-14.

17. An apparatus (500) comprising: a processing circuitry (502); and a memory (541), said memory containing instructions executable by said processing circuitry, whereby the apparatus is operative to perform the method of any one of claims 1-14.

Citation Information

Patent Citations

  • Source-free active adaptation to distributional shifts for machine learning

    US20230137905A1

  • Machine learning model change detection and versioning

    US20230144585A1

  • Data Drift Impact In A Machine Learning Model

    US20230186144A1

  • Adaptive retraining of an artificial intelligence model by detecting a data drift, a concept drift, and a model drift

    US20230376825A1