Systems and methods for dynamic activation of artificial intelligence / machine learning models based on inactive artificial intelligence / machine learning model assessment
Patent Information
- Application Number
- PCT/US2025/029994
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-27
- Filing Date
- 2025-05-19
- Publication Date
- 2026-10-01
Smart Images

Figure US2025029994_01102026_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS FOR DYNAMIC ACTIVATION OF ARTIFICIAL INTELLIGENCE / MACHINE LEARNING MODELS BASED ON INACTIVE ARTIFICIAL INTELLIGENCE / MACHINE LEARNING MODEL ASSESSMENTTECHNICAL FIELD
[0001] This application relates generally to wireless communication systems, including wireless communication systems that use artificial intelligence (AI) / machine learning (ML) models for one or more corresponding AI / ML model features.BACKGROUND
[0002] Wireless mobile communication technology uses various standards and protocols to transmit data between a base station and a wireless communication device. Wireless communication system standards and protocols can include, for example, 3rd Generation Partnership Project (3GPP) Long Term Evolution (LTE) (e.g., 4G), 3GPP New Radio (NR) (e.g., 5G), and Institute of Electrical and Electronics Engineers (IEEE) 802.11 standard for Wireless Local Area Networks (WLAN) (commonly known to industry groups as Wi-Fi®).
[0003] As contemplated by the 3GPP, different wireless communication systems' standards and protocols can use various radio access networks (RANs) for communicating between a base station of the RAN (which may also sometimes be referred to generally as a RAN node, a network node, or simply a node) and a wireless communication device known as a user equipment (UE). 3GPP RANs can include, for example. Global System for Mobile communications (GSM), Enhanced Data Rates for GSM Evolution (EDGE) RAN (GERAN), Universal Terrestrial Radio Access Network (UTRAN), Evolved Universal Terrestrial Radio Access Network (E-UTRAN), and / or Next- Generation Radio Access Network (NG-RAN).
[0004] Each RAN may use one or more radio access technologies (RATs) to perform communication between the base station and the UE. For example, the GERAN implements GSM and / or EDGE RAT, the UTRAN implements Universal Mobile Telecommunication System (UMTS) RAT or other 3GPP RAT, the E-UTRAN implements LTE RAT (sometimes simply referred to as LTE), and NG-RAN implements NR RAT (sometimes referred to herein as 5G RAT, 5G NR RAT, or simply NR). In certain deployments, the E-UTRAN may also implement NR RAT. In certain deployments, NG-RAN may also implement LTE RAT.
[0005] A base station used by a RAN may correspond to that RAN. One example of an E-UTRAN base station is an Evolved Universal Terrestrial Radio Access Network (E-14929-5814-3301'1 P71475WO1UTRAN) Node B (also commonly denoted as evolved Node B, enhanced Node B, eNodeB, or eNB). One example of an NG-RAN base station is a next generation Node B (also sometimes referred to as a g Node B or gNB).
[0006] A RAN provides its communication services with external entities through its connection to a core network (CN). For example, E-UTRAN may utilize an Evolved Packet Core (EPC) while NG-RAN may utilize a 5G Core Network (5GC).BRIEF DESCRIPTION OF THE SEVERAL VIEWS OF THE DRAWINGS
[0007] To easily identify the discussion of any particular element or act, the most significant digit or digits in a reference number refer to the figure number in which that element is first introduced.
[0008] FIG. 1 illustrates a diagram showing an example database of AI / ML models that is available at a UE.
[0009] FIG. 2 illustrates a first phase of an operational workflow between a UE and a base station.
[0010] FIG. 3 illustrates a second phase of an operational workflow between a UE and the base station.
[0011] FIG. 4 illustrates a method of a UE, according to embodiments discussed herein.
[0012] FIG. 5 illustrates a method of a base station, according to embodiments discussed herein.
[0013] FIG. 6 illustrates a method of a UE, according to embodiments discussed herein.
[0014] FIG. 7 illustrates a method of a UE, according to embodiments discussed herein.
[0015] FIG. 8 illustrates a method of a base station, according to embodiments discussed herein.
[0016] FIG. 9 illustrates a method of a UE, according to embodiments discussed herein.
[0017] FIG. 10 illustrates an example architecture of a wireless communication system, according to embodiments disclosed herein.
[0018] FIG. 11 illustrates a system for performing signaling between a wireless device and a network device, according to embodiments disclosed herein.
[0019] Various embodiments are described with regard to a UE. However, reference to a UE is merely provided for illustrative purposes. The example embodiments may be utilized with any electronic component that may establish a connection to a network and is configured with the hardware, software, and / or firmware to exchange information and data 24929-5814-3301'1 P71475WO1with the network. Therefore, the UE as described herein is used to represent any appropriate electronic component.
[0020] One aspect with respect to post-deployment handling of AI / ML models relates to the potential for frequent changes and / or updates to the Al / ML models. Viewed on way. an AI / ML model can be seen as a software component that can be substituted, upgraded, etc., and then executed on hardware in the device.
[0021] An Al / ML model may be understood to be used for a given AI / ML model feature. The AI / ML model feature, seen one way, represents a use case for the AI / ML model. For example, an AI / ML model may be understood to operate for a beam measurement and prediction AI / ML model feature, a channel state information (CSI) compression and / or decompression AI / ML model feature, and / or a positioning AI / ML model feature, etc. Other types of AI / ML model features may exist.
[0022] An AI / ML model may be understood to correspond to an AI / ML functionality. For example, an AI / ML model for beam measurement and prediction may be understood to correspond to an AI / ML model functionality' for the prediction of a particular set of 16 beams (a set A of beams) using beam measurements using a particular set of eight beams (a set B of beams). In another example, an AI / ML model for CSI compression and / or decompression may be understood to correspond to an AI / ML model functionality for the use of CSI compression using a particular encoder and / or CSI decompression using a particular decoder. One or more parameters used by an AI / ML model may be understood to represent and / or control the functionality' of the AI / ML model as used in this sense.
[0023] When considering AI / ML model functionality aspects, AI / ML model updating scenarios present considerations with respect to mechanisms for ensuring that a device that has passed conformance testing with one version of an AI / ML model for an AI / ML model functionality can also / still meet the minimum requirements for that AI / ML model functionality after the update to the AI / ML model.
[0024] In cases where one or more modifications for a given AI / ML model are made to a device over the device's lifetime (e.g., if there are post-deployment updates to an AI / ML model, and / or a new AI / ML model for the AI / ML functionality is developed at or provided to the device), various issues can arise. Firstly, it may be that the integration of a new or updated AI / ML model onto the device occurs without complete validation of the new / updated AI / ML model, resulting in inaccurate results. Secondly, a modification, updating, or fine-tuning of an existing AI / ML model could lead to degraded performance34929-5814-3301'1 P71475WO1under certain operational conditions, even if improvements occur under other operational conditions.
[0025] Embodiments discussed herein accordingly relate to mechanisms for ensuring the adaptability and flexibility of Al / ML models by providing for post-deployment validation / testing of AT / ML models. Note that a post-deployment phase for AI / ML model usage at the device can be considered within the broader context of an AI / ML model usage generalization framework.
[0026] Embodiments herein relate to the monitoring of one or more AI / ML models at the UE to test and / or validate those AI / ML models. Consider an AI / ML model that is in an active mode. The active mode may be understood to use inferencing result transmission, in which inferencing result(s) generated by the AI / ML model are reported out from the device and / or are substantively used by / within the system for purposes of facilitating the applied or “live” operation of the AI / ML model functionality corresponding to the AI / ML model.
[0027] If the AI / ML model in the active mode fails to test or validate, the AI / ML model may be switched from the active mode into an inactive mode. The inactive mode uses inferencing result suppression, in which any inference generated by the AI / ML model is not used by / within the system for the purposes of applied or “live” operation of the AI / ML model functionality corresponding to the AI / ML model. In some cases, it may be that the inferencing result suppression mode corresponds to a case in which an inferencing result of the AI / ML model is not reported out from the device. Note that the inferencing result suppression does not prevent the use of inferencing at the AI / ML model for purposes of facilitating monitoring as described herein, nor does it prevent transmission of performance metrics that are determined based on an inferencing result of the AI / ML model in applicable cases as is described herein.
[0028] In some cases where an active AI / ML model fails to test or validate, a device may then activate another AI / ML model for the corresponding AI / ML model feature. In other such cases, a device may “fall back” to the use of anon-AI / ML model-based behavior corresponding to the AI / ML model feature at the device instead of presently continuing to support the AI / ML model feature with any AI / ML model.
[0029] Various issues may be considered in the context of AI / ML model monitoring / potential AI / ML model deactivation / s witching. A first issue relates to the observation that if, as part of an AI / ML model monitoring procedure, reported key performance indicator(s) (KPI(s)) for an currently deployed / active AI / ML model fail to pass a minimum requirement, it may be that the currently deployed / active AI / ML model has 44929-5814-3301'1 P71475WO1drifted due to changing channel conditions / a changing propagation environment. The first issue corresponds to questions of how, for such cases, a UE or network device (as the case may be) selects a new AI / ML model to be deployed.
[0030] A second issue corresponds to questions of how it can be verified that any newly-selected / activated AL / ML model meets the requirements before being deployed.
[0031] Finally, a third issue relates to the observation that in current wireless communication systems, the existing framework for Al / ML model monitoring procedures is a reactive framework. That is, a performance degradation of the AI / ML model is first identified, and then the device (the UE or the network device, as the case may be) acts upon this degradation. This third issue corresponds to questions of how to arrange an AI / ML model monitoring procedure which is instead proactive, in that one or more actions to improve performance (e.g., switching AI / ML models) are taken before the performance of a currently active AI / ML model performance degrades (or, at least before it degrades overmuch). Such an approach would be capable of maintaining acceptable performance for the use of AI / ML model(s) without wasting resources for extra monitoring and / or for the latency associated with an intervening falling back to a non-AI / ML model-based behavior corresponding to cases that ultimately result in AI / ML model reselection.
[0032] In some cases, it may be that any verification(s) on updated AI / ML models are performed using data as collected in the field by a UE. This has the benefit of reflecting the exact UE hardware and implementation as part of such updates. Note that in some cases where AI / ML models have been trained at the network or at an over the air (OTA) server (even under identical conditions), that training may not reflect the exact UE hardware implementations. Therefore, verification of AI / ML models according to the actual UE hardware (radio frequency (RF) architecture, etc.) in this way remains useful.
[0033] Various methods to access / monitor the applicability and expected performance of an inactive AI / ML model are contemplated. Various examples with respect to the case of activation / selection / switching of UE-sided AI / ML models (or UE-parts of two-sided AI / ML models that have components at both the UE and the network) are now given. In a first example, assessment and / or monitoring for an AI / ML model may be based on any additional conditions associated with the AI / ML model. In a second example, assessment and / or monitoring for an AI / ML model may be based on an input and / or output data distribution for the AI / ML model. In a third example, assessment and / or monitoring for an AI / ML model may be based on past knowledge of the performance of that AI / ML model (e.g., based on the use of that AI / ML model at other UEs).54929-5814-3301'1 P71475WO1
[0034] In a fourth example, assessment and / or monitoring for AI / ML models may use a comparison of an active AI / ML model to inactive AI / ML model (s) for monitoring purposes (e.g., to measure the inference accuracy and / or other aspect(s) of the inactive AI / ML model(s) for comparison purposes). Various embodiments discussed herein relate to AI / ML model monitoring and / or assessment for this type of example.
[0035] In the context of switching between AI / ML models, various challenges can be considered. First, consider a case where a UE is configured to use a “fall back,'’ where the UE may switch from the use of an AI / ML model for an AI / ML model feature to the use of a non-AI / ML model-based behavior corresponding to the AI / ML model feature. For such cases, when the UE is operating based in the non-AI / ML-based manner, there may be questions with respect to when to (re)activate AI / ML-based operation, and correspondingly with respect to the identification of the AI / ML model that should be activated.Embodiments herein relate mechanisms for monitoring an inactive AI / ML model to assess whether a condition for activation of that AI / ML model is satisfied and / or which AI / ML model from a group of available AI / ML models is suitable for activation.
[0036] Second, consider a case where a UE is configured to be able to switch between different AI / ML models (e.g., switching to an updated AI / ML model for a given functionality7, and / or switching to an AI / ML model for a different functionality7within the applicable feature). In such cases, it may be determined that, corresponding to a case where the performance of an AI / ML model degrades, it is better to switch to the new AI / ML model in a proactive fashion rather than in a reactive fashion. Embodiments herein relate mechanisms for the determination of a target AI / ML model corresponding to such proactive uses. These embodiments involve the monitoring of inactive model(s) to determine / identify a target inactive AI / ML model to which to switch.Dynamic Management of Active / Inactive UE AI / ML Models
[0037] In some embodiments, a UE possesses multiple AI / ML models. Additionally, or alternatively, various AI / ML models are stored at OTA server and are downloadable at the UE fortraining and / or usage purposes. Additionally, or alternatively, various AI / ML models are stored at the network and are downloadable to the UE from the network for training and / or usage purposes.
[0038] FIG. 1 illustrates a diagram 100 showing an example database 102 of AI / ML models that is available (whether directly or via download, as just discussed) at a UE.64929-5814-3301'1 P71475WO1
[0039] The database 102 of AI / ML models includes the first AI / ML model 104. the second AI / ML model 106, the third AI / ML model 108, the fourth AI / ML model 110, and so on until the fifth AI / ML model 112, as illustrated.
[0040] As illustrated, each AI / ML model may correspond to various aspects associated or not associated with the UE capability or aspects of where a UE capability not specified. For example, FIG. 1 illustrates (in terms of the first AI / ML model 104) that each AI / ML model may include first aspects 114 associated with the UE capabilities and a set of inference parameters, second aspects 116 not associated with a UE capability, and / or third aspects 118 aspects that are not specified. Some conditions 120 as may be discussed herein may fall under either the first aspects 114 or the second aspects 116, while some additional conditions 122 as may be discussed herein may fall under either the second aspects 116 or the third aspects 118.
[0041] As illustrated, each of the AI / ML models of the database 102 may be identified using a functionality -based AI / ML model identifier (ID) 124. Each functionality -based AI / ML model ID 124 may be unique to its model to facilitate such identification. Each functionality-based AI / ML model ID 124 may be understood to correspond to a particular functionality for which the identified AI / ML model is to be used.
[0042] FIG. 1 may take a case where each of the first AI / ML model 104 and the second AI / ML model 106 is inactive, while another model (e.g., the first AI / ML model 104) is currently active.
[0043] Each of the AI / ML models in the database 102 may be associated with a set of inference parameters. FIG. 1 illustrates this in terms of the second AI / ML model 106 and the third AI / ML model 108. As can be seen, the second AI / ML model 106 corresponds to a first set of parameters 126. The “Associated ID” field may represent an identification of the functionality-based AI / ML model ID 124 that is assigned to the second AI / ML model 106.
[0044] The database 102 may be understood to be a grouping of AI / ML models for an AI / ML mode feature of beam measurement and prediction. Correspondingly, remaining fields of the AI / ML models of the database 102 include, as illustrated, a "set B pattern ID" field, a “size of setA / setB” field, an “order of setA / setB” field, a “classifier / regressor model” field, a “monitoring metric” field, and a “top-K report” field. Based on these remaining field types, it may be understood that the second AI / ML model 106 corresponds to an AI / ML model feature for predicting a set B of beams based on measurements of a set A of beams. Further, the settings for these fields may correspond to a particular functionality (e.g., a certain number / certain ones for the set A of beams, a certain74929-5814-3301'1 P71475WO1number / certain ones for the set B of beams, a particular classifier / regressor model) that is for the second AI / ML model 106 within this feature.
[0045] The third AI / ML model 108 uses a set of parameters 128. As can be seen, the field types are the same that were described in relation to the second AI / ML model 106.
[0046] Preliminarily, the "Associated ID” field may represent an identification of the functionality-based AI / ML model ID 124 that is assigned to the third AI / ML model 108 (and which is different from that assigned to the second AI / ML model 106).
[0047] Based on these remaining field ty pes, it is understood that the third AI / ML model 108, like the second AI / ML model 106, corresponds to an AI / ML model feature for predicting a set B of beams based on measurements of a set A of beams. However, settings for these fields for the third AI / ML model 108 may be different than those used in the second AI / ML model 106, such that the particular functionality7of the third AI / ML model 108 is different than the functionality' represented by the second AI / ML model 106. For example, the third AI / ML model 108 may use a different number / different ones for the set A of beams, a different number / different ones for the set B of beams, and / or a different classifier / regressor model than that which is used by the second AI / ML model 106.
[0048] Note that while this discussion of FIG. 1 assumes a case where the applicable AI / ML model feature is beam measurement and prediction (as may be concluded based on review of the particular parameters used in the AI / ML models thereol), this is given by way of example only. A database like the database 102 could just as well include AI / ML models for other types of AI / ML model features other than an AI / ML model feature of beam management and prediction (e.g., a CSI compression and decompression AI / ML model feature, a positioning AI / ML model feature, etc.).
[0049] A signaling workflow between a UE and a network (e g., a base station of a network) is now described.
[0050] In a first phase of the workflow, the UE performs monitoring of one or more of its inactive AI / ML model(s). Preliminarily, the monitoring may be configured. A base station signals, to a UE, which inactive AI / ML model(s) the UE is to monitor (e.g., via a radio resource control (RRC) reconfiguration message). This configuration message may use a flag (e.g., a “monitoringOnly flag” of an RRC message / an NR model configuration information element (IE)) to indicate to the UE that it should suppress AI / ML model outputs for the indicated AI / ML model(s). For example, in the context of a beam measurement and prediction AI / ML model feature, this may mean that no beam state updates generated84929-5814-3301'1 P71475WO1corresponding to the output of the indicated AI / ML models are delivered from the UE to the base station).
[0051] The messaging from the base station may further indicate a monitoring mode that the UE is to use to monitor the inactive Al / ML model(s). In some cases, a dataset-based method may be configured for use. The dataset-based monitoring method be considered an “offline” method. In this case, the UE operates the inactive monitored AI / ML model(s) with one or more datasets received from the base station and evaluates the results. In this case, the datasets may not directly represent an actual channel state between the UE and the base station, but may rather be datasets selected by the network to represent a channel state case that the network wishes the UE to test for. Note that in some cases, the network may select the dataset for this use after deeming that the dataset is at least reflective in one or more ways of the actual channel state between the UE and the base station.
[0052] In other cases, a shadow execution-based monitoring method may be configured for use. The shadow execution method may be considered an “online” method. In this case, the UE operates the inactive monitored AI / ML model(s) with data that is also be used as input for an active AI / ML model.
[0053] Once the desired operation type is configured-for, the UE proceeds to collects channel and / or performance data (as the case may be) for purposes of AI / ML model assessment of the indicated inactive AI / ML model(s).
[0054] The UE then evaluates the inactive AI / ML model(s) using delivered datasets (in the case of the dataset-based monitoring method) or “live” data (in the case of the shadowexecution based monitoring method.
[0055] FIG. 2 illustrates a first phase 202 of an operational workflow between a UE 204 and a base station 206. Note that FIG. 2 also assumes that the UE possesses and / or maintains a database 216 that is analogous to the one described previously in relation to FIG. 1. However, note that while FIG. 1 assumed an AI / ML model feature of beam measurement and prediction by way of example, the database 216 should be understood to maintain AI / ML models for any type of AI / ML model feature. In other words, it should be understood that the operational workflow under discussion is useable with a database 216 of AI / ML models for any AI / ML model feature (not only a database for an AI / ML model feature for beam measurement and prediction).
[0056] The operational workflow under discussion assumes that the UE 204 possesses or maintains the database 216. It may be that the first AI / ML model 218 is currently active for the corresponding AI / ML feature, while each of the second AI / ML model 220, the third 94929-5814-3301'1 P71475WO1AI / ML model 222, the fourth AI / ML model 224, and the fifth AI / ML model 226 are in an inactive mode.
[0057] First, the base station 206 signals 208 to the UE 204 which inactive AI / ML model(s) from the database 216 are to be monitored by the UE 204. This signaling may occur in, for example, an RRC reconfiguration message, as illustrated.
[0058] The base station 206 also configures 210 to the UE 204 a monitoring mode to use to monitor the indicated AI / ML model(s). (e.g., the base station 206 indicates to the UE 204 to use a dataset-based monitoring method, or indicates to the UE 204 to use a shadow execution-based monitoring method, as the case may be).
[0059] The base station 206 also configures 212 one or more resources for the UE to use to perform the monitoring. For example, in a dataset-based monitoring method, the base station 206 configures 212 the UE 204 to receive the dataset(s) used for monitoring. As another example, the in shadow execution-based monitoring method, the base station 206 configures 212 the UE 204 to use one or more sets of input data for the currently active AI / ML model for monitoring purposes with respect to the inactive AI / ML models a(s) well.
[0060] The first phase 202 then evaluates 214 the inactive AI / ML model(s) using the delivered datasets / according to the live shadow execution, as configured.
[0061] In some embodiments, the monitoring as configured occurs at the UE 204 when the UE 204 detects a given event. In some such embodiments, the base station 206 configures the event to the UE 204.
[0062] In some embodiments, the monitoring as configured occurs at the UE 204 according to a timer. In some such embodiments, the base station 206 configures the timer to the UE 204.
[0063] In some embodiments, the monitoring as configured occurs at the UE 204 according to a periodicity. In some such embodiments, the base station 206 configures the periodicity to the UE 204.
[0064] As part of performing the monitoring of the inactive AI / ML model(s), the UE computes performance metrics for the inactive AI / ML model(s). Examples of a performance metric may include, but are not limited to. a performance metric that represents the accuracy of the inactive AI / ML model, a latency (e.g., of inference) for the AI / ML model, which is or includes a pared-down inferencing result of an AI / ML model, etc.
[0065] Under the dataset-based monitoring method, the UE generates inferences at the inactive AI / ML model(s) by applying the network-provided datasets at the inactive AI / ML104929-5814-3301'1 P71475WO1model(s). The UE then determines the one or more performance metric(s) based on these generated inference(s).
[0066] Under the shadow execution-based monitoring method, the UE applies input data that is being used at an active Al / ML model at the inactive Al / ML model(s) as well, such that the inactive Al / ML model(s) generate inferences based on this ‘'live” data. The UE then determines the one or more performance metric(s) based on these generated inference(s).
[0067] Various options for enabling the UE to switch to a different Al / ML model based on performance metric(s) are contemplated. In a first option, the performance metric(s) so calculated are sent to a base station. In some examples, the performance metric(s) are sent using medium access control control element (MAC CE). In some examples, the performance metric(s) are sent to the base station in uplink control information (UCI).
[0068] Then, the base station evaluates the performance metric(s) and may determine, based on that evaluation, that a currently active Al / ML model at the UE should be switched with a selected inactive Al / ML model. Accordingly, the base station sends the UE a command to switch to the use of the selected inactive Al / ML model. In other words, the selected inactive Al / ML model is moved to an active state. Note additionally that in such cases, the currently active Al / ML model may also be moved to an inactive state.
[0069] In a second option, the network provides the UE with one or more thresholds for one or more of the performance metric(s) under use. Once the UE calculates the performance metric(s) are calculated at the UE, they are compared to the provided thresholds. If a threshold for a performance metric corresponding to an inactive Al / ML model is met, the UE may switch to the use of that Al / ML model. In other words, the selected inactive Al / ML model is moved to an active state. Note additionally that in such cases, the currently active Al / ML model may also be moved to an inactive state.
[0070] Corresponding to these cases, threshold values used by the base station or the UE (as the case may be) to evaluate performance metric(s) of the inactive Al / ML model(s) may be absolute threshold values (e.g., such that an inactive Al / ML model is activated once its performance metric(s) meet the associated thresholds). In some cases, the threshold values used may be relative threshold values (e.g., such that an inactive Al / ML model is activated once its performance metric(s) is / are a threshold amount better than corresponding performance metric(s) for a currently active Al / ML model).
[0071] Mechanisms for transitioning from the currently active Al / ML model to the selected inactive Al / ML model that is to replace the currently active Al / ML model are now discussed. The selected inactive Al / ML model may be loaded from UE storage. Then, the 114929-5814-3301'1 P71475WO1UE may execute one or more sanity checks on the selected inactive AI / ML model (e.g., to ensure input / output compatibility). Then, the UE applies the selected inactive AI / ML model as then new active AI / ML model (switches the inactive AI / ML model into the active mode). This is done such that service interruption for the AI / ML feature is minimized. The currently active AI / ML model may correspondingly be switched to the inactive mode going forward.
[0072] FIG. 3 illustrates a second phase 302 of an operational workflow between the UE 204 and the base station 206. Note that the second phase 302 may be understood to continue from the first phase 202 of the operational workflow as first introduced in FIG. 2 and related discussion. Note that FIG. 3 illustrates that the UE 204 possesses and / or maintains the database 216, as was previously discussed.
[0073] In a first option 306 for the second phase 302 of the operational workflow, the UE 204 sends 308 one or more performance metric(s) for one or more inactive AI / ML model(s) to the base station 206. As shown, MAC CE and / or UCI may be used for this purpose. Note that the performance metric(s) (e.g., KPI(s)) may be calculated by the UE 204 and / or sent to the base station 206 on a periodic basis. Performance metric(s) for a currently active AI / ML model may also be calculated and sent similarly.
[0074] The base station 206 receives the performance metric(s) for the AI / ML model(s) and makes decisions as to which (if any) inactive AI / ML models should be deployed due to the AI / ML model being an optimum available AI / ML model and / or better than a currently active AI / ML model.
[0075] Then, the network sends 310, based on its analysis of these one or more performance metric(s), an indication for the UE 204 to switch from a currently active AI / ML model to a selected inactive AI / ML model. Note that the label “AI / ML model X” as used here in FIG. 3 represents a notion that the inactive AI / ML model may be identified by its model ID (where A is the model ID of the selected inactive AI / ML model).
[0076] The UE correspondingly moves the selected inactive AI / ML model to an active state. Further, the currently active AI / ML model may be moved to an inactive state.
[0077] In a second option 304 for the second phase 302 of the operational workflow, the base station 206 signals 312 the UE 204 one or more thresholds for the UE 204 to use to evaluate calculated performance metric(s) for the AI / ML model(s).
[0078] The UE 204 then calculates one or more performance metric(s) for the inactive AI / ML models. The UE 204 then applies the calculated performance metric(s) with the provided threshold(s) and correspondingly selects an inactive AI / ML model that meets the 124929-5814-3301'1 P71475WO1applicable threshold(s). Note that the performance metric(s) (e.g., KPIs) may be calculated by the UE 204 and / or compared to the corresponding threshold(s) on a periodic basis.Performance metric(s) for a currently active AI / ML model may also be calculated similarly, attendant to this use.
[0079] The UE correspondingly autonomously switches 314 to the selected inactive AI / ML model. To facilitate the switch, the UE moves the selected inactive AI / ML model to an active state. Further, the currently active AI / ML model may be moved to an inactive state.
[0080] Various timing mechanisms for this monitoring are now discussed. In some cases, the monitoring may be event triggered. In some such cases, the base station 206 configures the UE 204 to perform the monitoring as just described in reaction to some trigger. In other such cases, the base station 206 configures the UE with a trigger for performing the monitoring, and the UE 204 activates the monitoring as configured in response to its detection of the trigger.
[0081] Various details of monitoring mechanisms for the inactive AI / ML models as discussed herein are now related. The system according to embodiments herein may use a lightweight "monitoring only” mode with respect to an inactive AI / ML model. In such circumstances, the UE may be configured to operate the inactive AI / ML model in such a monitoring-only mode, which is not considered a case of full activation for that AI / ML model.
[0082] In some cases, this monitoring-only mode corresponds to inferencing result suppression. In some cases, this means that a result is not transmitted to another entity in the wireless communication system. For example, in the case of a beam measurement and prediction AI / ML model feature, when operating an inactive AI / ML model in a monitoring-only mode, the UE may run the inactive AI / ML model to may a beam prediction, but suppress any reporting corresponding to that inactive AI / ML model. In other words, no beam state update information corresponding to the use of that inactive AI / ML model is sent to the network and / or used by the system at large directly for the beam measurement and prediction AI / ML model feature.
[0083] In some cases, the monitoring-only mode is implemented using a monitoring flag in a bitmap in a configuration information for the inactive AI / ML model (e g., in an RRC signaling message).
[0084] The use of the monitoring-only mode for inactive an AI / ML model may also configure the UE to limit computational resources used for performing inferencing using the134929-5814-3301'1 P71475WO1inactive AI / ML model (e.g., such that the UE uses a relatively lower inference frequency with the inactive AI / ML model, a reduced precision with the inactive AI / ML model, etc.).
[0085] Details with respect to a dataset-based monitoring method as discussed herein are now presented. When using the dataset-based monitoring method, a base station delivers representative datasets to a UE for the use by the UE in evaluating (in an offline fashion) one or more inactive AI / ML model(s). In some cases, a goal of dataset-based monitoring method is to assess the applicability and / or performance of the inactive AI / ML model(s) under something approaching current channel conditions. Accordingly, the datasets so provided may be configured to represent or estimate current channel conditions in one or more ways.
[0086] In some cases, synthetic datasets are used. These synthetic datasets may be understood as pre-configured datasets that the network sends or otherwise identifies to the UE for use.
[0087] In some cases, portions of field data (e.g., "‘field data snippets”) are used. These may be representative of (a portion) of currently applicable CSI, a (portion of) currently applicable signal to noise ratio (SNR) measurements, etc.
[0088] Note that a standard for the wireless communication system may define protocols for the delivery of such datasets from the network to the UE.
[0089] Details with respect to a shadow execution-based monitoring method as discussed herein are now presented. When using the shadow execution-based monitoring method, the UE (e.g., in response to a base station instruction / configuration) temporarily uses inactive AI / ML model(s) in parallel with any corresponding active AI / ML model for short durations. In some cases, this means that input data used for inferencing for an active AI / ML model is also applied for inferencing purposes at the inactive AI / ML model(s). In some cases, this facilitates the comparison of the performance of the inactive AI / ML model(s) to the performance of the active AI / ML model.
[0090] In one such example, the UE runs an inactive AI / ML model for the beam measurement and prediction feature alongside a currently active AI / ML model for the beam measurement and prediction feature. The results (or KPIs corresponding to the results) can then be compared for purposes of determining whether to switch to the inactive AI / ML model.
[0091] In some embodiments, the use of the shadow execution-based monitoring method is trigger based. In such cases, the base station configures the UE for event-based instances for144929-5814-3301'1 P71475WO1model monitoring (e.g.. when a performance degradation of a currently active AI / ML model is detected).
[0092] In some embodiments, the use of shadow execution-based monitoring method is periodic. In such cases, the base station configures the UE to perform the shadow executionbased monitoring mechanism according to a given periodicity.
[0093] Note that the use of the shadow execution-based monitoring method may be based on a UE capability (e.g., based on available processing capabilities of UE in terms of supporting the execution of multiple AI / ML models in parallel / at the same time / on the same input data).
[0094] Details of switching criteria as may be used in embodiments disclosed herein are now described. In some embodiments, a network provides criteria for switching to an inactive AI / ML model to the UE. In such cases, a base station signals performance metric thresholds (e.g., for accuracy, latency, robustness, etc.) for AI / ML model activation / switching to the UE. The UE then autonomously (e.g., without a direct switching instruction from the network) switches between its AI / ML models based on the network-provided criteria.
[0095] In one example corresponding to a beam measurement and prediction AI / ML model feature, it may be that the network configures the UE to switch from a currently active AI / ML model to an inactive AI / ML model in cases where the beam prediction accuracy of a currently active AI / ML model drops below' 80%.
[0096] In another example also corresponding to a beam measurement and prediction AI / ML model feature, it may be that a currently active AI / ML model's beam accuracy is 50%, while an inactive AI / ML model (e.g., of a set of such inactive AI / ML models) has a beam accurate of 85%. In such cases, the UE selects that inactive AI / ML model from the set of inactive models for deployment. In some such cases, pre-deployment verification may first be carried out with respect to the selected inactive AI / ML model.
[0097] Some such cases exhibit high levels of dynamic adaptation. For example, it may be that a UE uses local metrics (e.g., channel coherence time, interference levels) in addition to network-provided criteria to determine whether to switch to the use of an inactive AI / ML model.
[0098] In some embodiments, the UE and the network work collaboratively to determine when to switch to the use of an inactive AI / ML model. In such cases, the UE evaluates inactive AI / ML model(s) via a monitoring-only mode (in some cases, using a dataset-based assessment). The UE then reports the corresponding performance metric(s) (e.g., KPI 154929-5814-3301'1 P71475WO1scores) representing predicted performance of the inactive AI / ML model(s) to the network. The network analyzes these performance metrics and sends the UE an activation command / s witching command corresponding to a selected inactive AI / ML model.
[0099] In some cases, it is anticipated that a fall back to non-Al / ML model-based operation may be called for. In some cases, a switch to a baseline AI / ML model for a given AI / ML feature may be called for. Triggers for revert to non-AI / ML model-based operations, or to baseline AI / ML models, may be as follows.
[0100] In some cases, the fall back and / or reversion to a baseline AI / ML model may be UE-initiated. Such cases may exhibit relatively higher computational overhead and / or model instability.
[0101] In some cases, the fall back and / or reversion to a baseline AI / ML model may be network-initiated. Such cases may, upon detection of performance degradation of a currently active AI / ML model, cause the UE to fall back to non-AI / ML model-based operation in cases where all KPIs from all corresponding inactive AI / ML models are below minimum requirements.
[0102] Embodiments herein relate to the predictive activation of / switching to an inactive AI / ML model. In such cases, it may be that an (additional) AI / ML model may be used to predict when to switch to an inactive AI / ML model as discussed herein. Such a determination may be based on, for example UE mobility' patterns, a monitored history of KPI patterns associated with the set of available (active and / or inactive) AI / ML models, etc.
[0103] FIG. 4 illustrates a method 400 of a UE, according to embodiments discussed herein. The method 400 includes receiving 402, from a base station, configuration information configuring the UE to monitor a first AI / ML model using dataset-based monitoring while the first AI / ML model is in an inactive mode, wherein the inactive mode comprises inferencing result suppression. The method 400 further includes receiving 404, from the base station, a dataset corresponding to a use of the dataset-based monitoring. The method 400 further includes generating 406, while the first AI / ML model is in the inactive mode, a first inference using the first AI / ML model by applying the dataset at the first AI / ML model. The method 400 further includes determining 408, based on the first inference, a first performance metric for the first AI / ML model. The method 400 further includes sending 410, to the base station, the first performance metric for the first AI / ML model. The method 400 further includes receiving 412, in response to the first performance metric, an instruction to move the first AI / ML model to an active mode, wherein the active164929-5814-3301'1 P71475WO1mode comprises inferencing result transmission. The method 400 further includes moving 414, in response to the instruction, the first AI / ML model to the active mode.
[0104] In some embodiments, the method 400 further includes determining, based on a second inference of a second AI / ML model, a second performance metric for the second AI / ML model; and sending, to the base station, the second performance metric for the second AI / ML model; wherein the instruction to move the first AI / ML to the active mode is further received in response to the second performance metric.
[0105] In some embodiments of the method 400, the first performance metric indicates an accuracy of the first AI / ML model.
[0106] In some embodiments of the method 400, the first performance metric indicates a latency of the first AI / ML model.
[0107] In some embodiments of the method 400, the configuration information comprises a bitmap that identifies the first AI / ML model for use with the dataset-based monitoring.
[0108] In some embodiments of the method 400, the dataset comprises synthetic data.
[0109] In some embodiments of the method 400, the dataset comprises field data.
[0110] In some embodiments, the method 400 further includes moving a second AI / ML model from the active mode to the inactive mode in response to moving the first AI / ML model to the active mode.
[0111] In some embodiments of the method 400, the first AI / ML model is configured to use fewer computational resources when in the inactive mode as compared to when in the active mode.
[0112] FIG. 5 illustrates a method 500 of a base station, according to embodiments discussed herein. The method 500 includes sending 502, to a UE, configuration information configuring the UE to monitor a first AI / ML model using dataset-based monitoring while the first AI / ML model is in an inactive mode, wherein the inactive mode comprises inferencing result suppression. The method 500 further includes sending 504, to the UE, a dataset corresponding to a use of the dataset-based monitoring. The method 500 further includes receiving 506, from the UE. a first performance metric for the first AI / ML model that is based on the dataset. The method 500 further includes determining 508, based on the first performance metric, to send the UE an instruction to move the first AI / ML model to an active mode, wherein the active mode comprises inferencing result transmission. The method 500 further includes sending 510 the instruction to the UE.174929-5814-3301'1 P71475WO1
[0113] In some embodiments of the method 500, the determining, based on the first performance metric, to send the instruction to move the first AI / ML model to the active mode comprises determining that the first performance metric meets a threshold.
[0114] In some embodiments of the method 500, the determining, based on the first performance metric, to send the instruction to move the first AI / ML model to the active mode comprises determining that the first performance metric is better than a second performance metric for a second AI / ML model that is in the active mode. In some such embodiments, the method 500 further includes receiving the second performance metric from the UE.
[0115] In some embodiments of the method 500, the first performance metric indicates an accuracy of the first AI / ML model.
[0116] In some embodiments of the method 500, the first performance metric indicates a latency of the first AI / ML model.
[0117] In some embodiments of the method 500, the configuration information comprises a bitmap that identifies the first AI / ML model for use with the dataset-based monitoring.
[0118] In some embodiments of the method 500, the dataset comprises synthetic data.
[0119] In some embodiments of the method 500, the dataset comprises field data.
[0120] In some embodiments of the method 500, the first AI / ML model is configured to use fewer computational resources when in the inactive mode as compared to when in the active mode.
[0121] FIG. 6 illustrates a method 600 of a UE, according to embodiments discussed herein. The method 600 includes receiving 602, from a base station, configuration information configuring the UE to monitor a first AI / ML model using dataset-based monitoring while the first AI / ML model is in an inactive mode, wherein the inactive mode comprises inferencing result suppression. The method 600 further includes receiving 604. from the base station, a dataset corresponding to a use of the dataset-based monitoring. The method 600 further includes generating 606, while the first AI / ML model is in the inactive mode, a first inference using the first AI / ML model by applying the dataset at the first AI / ML model. The method 600 further includes determining 608 a first performance metric for the first AI / ML model corresponding to the first inference. The method 600 further includes determining 610, based on the first performance metric, to move the first AI / ML model to an active mode, wherein the active mode comprises inferencing result transmission. The method 600 further includes moving 612 the first AI / ML model to the active mode.184929-5814-3301'1 P71475WO1
[0122] In some embodiments of the method 600, the determining, based on the first performance metric, to move the first AI / ML model to the active mode comprises determining that the first performance metric meets a threshold. In some such embodiments, the method 600 further includes receiving, from the base station, the threshold.
[0123] In some embodiments, the method 600 further includes determining a second performance metric corresponding to a second inference of a second AI / ML model that is in the active mode; wherein the determining, based on the first performance metric, to move the first AI / ML model to the active mode comprises determining that the first performance metric is better than the second performance metric.
[0124] In some embodiments of the method 600, the determining to move the first AI / ML model to the active mode is further based on a channel condition between the UE and the base station.
[0125] In some embodiments of the method 600, the first performance metric indicates an accuracy of the first AI / ML model.
[0126] In some embodiments of the method 600, the first performance metric indicates a latency of the first AI / ML model.
[0127] In some embodiments of the method 600, the configuration information comprises a bitmap that identifies the first AI / ML model for use with the dataset-based monitoring.
[0128] In some embodiments of the method 600, the dataset comprises synthetic data.
[0129] In some embodiments of the method 600, the dataset comprises field data.
[0130] In some embodiments, the method 600 further includes moving a second AI / ML model from the active mode to the inactive mode in response to moving the first AI / ML model to the active mode.
[0131] In some embodiments of the method 600, the first AI / ML model is configured to use fewer computational resources when in the inactive mode as compared to when in the active mode.
[0132] FIG. 7 illustrates a method 700 of a UE, according to embodiments discussed herein. The method 700 includes receiving 702, from a base station, configuration information configuring the UE to monitor a first AI / ML model using shadow executionbased monitoring while the first AI / ML model is in an inactive mode, wherein the inactive mode comprises inferencing result suppression. The method 700 further includes generating 704, according to a use of the shadow execution-based monitoring, while the first AI / ML model is in the inactive mode, a first inference using the first AI / ML model by applying, at194929-5814-3301'1 P71475WO1the first AI / ML model, an input that is also applied at a second AI / ML model that is in an active mode to generate a second inference, wherein the active mode comprises inferencing result transmission. The method 700 further includes determining 706, based on the first inference, a first performance metric for the first AI / ML model. The method 700 further includes sending 708, to the base station, the first performance metric for the first AI / ML model. The method 700 further includes receiving 710, in response to the first performance metric, an instruction to move the first AI / ML to the active mode. The method 700 further includes moving 712, in response to the instruction, the first AI / ML model to the active mode.
[0133] In some embodiments, the method 700 further includes determining, based on the second inference, a second performance metric for the second AI / ML model; and sending, to the base station, the second performance metric for the second AI / ML model; wherein the instruction to move the first AI / ML to the active mode is further received in response to the second performance metric.
[0134] In some embodiments of the method 700, the first performance metric indicates an accuracy of the first AI / ML model.
[0135] In some embodiments of the method 700, the first performance metric indicates a latency of the first AI / ML model.
[0136] In some embodiments of the method 700, the configuration information comprises a bitmap that identifies the first AI / ML model for use with the shadow execution-based monitoring.
[0137] In some embodiments, the method 700 further includes moving the second AI / ML model from the active mode to the inactive mode in response to moving the first AI / ML model to the active mode.
[0138] In some embodiments of the method 700, the first AI / ML model is configured to use fewer computational resources when in the inactive mode as compared to when in the active mode.
[0139] FIG. 8 illustrates a method 800 of a base station, according to embodiments discussed herein. The method 800 includes sending 802, to a UE, configuration information configuring the UE to monitor a first AI / ML model using shadow execution-based monitoring while the first AI / ML model is in an inactive mode, wherein the inactive mode comprises inferencing result suppression. The method 800 further includes receiving 804, from the UE, a first performance metric for the first AI / ML model according to the shadow execution-based monitoring. The method 800 further includes determining 806, based on the 204929-5814-3301'1 P71475WO1first performance metric, to send the UE an instruction to move the first AI / ML model to an active mode, wherein the active mode comprises inferencing result transmission. The method 800 further includes sending 808 the instruction to the UE.
[0140] In some embodiments of the method 800, the determining, based on the first performance metric, to send the instruction to move the first AI / ML model to the active mode comprises determining that the first performance metric meets a threshold.
[0141] In some embodiments of the method 800, the determining, based on the first performance metric, to send the instruction to move the first AI / ML model to the active mode comprises determining that the first performance metric is better than a second performance metric for a second AI / ML model that is in the active mode. In some such embodiments, the method 800 further includes receiving the second performance metric from the UE.
[0142] In some embodiments of the method 800, the first performance metric indicates an accuracy of the first AI / ML model.
[0143] In some embodiments of the method 800, the first performance metric indicates a latency of the first AI / ML model.
[0144] In some embodiments of the method 800, the configuration information comprises a bitmap that identifies the first AI / ML model for use with the shadow execution-based monitoring.
[0145] In some embodiments of the method 800, the first AI / ML model is configured to use fewer computational resources when in the inactive mode as compared to when in the active mode.
[0146] FIG. 9 illustrates a method 900 of a UE, according to embodiments discussed herein. The method 900 includes receiving 902, from a base station, configuration information configuring the UE to monitor a first AI / ML model using shadow executionbased monitoring while the first AI / ML model is in an inactive mode, wherein the inactive mode comprises inferencing result suppression. The method 900 further includes generating 904, according to a use of the shadow execution-based monitoring, while the first AI / ML model is in the inactive mode, a first inference using the first AI / ML model by applying, at the first AI / ML model, an input that is also applied at a second AI / ML model that is in an active mode to generate a second inference, wherein the active mode comprises inferencing result transmission. The method 900 further includes determining 906, based on the first inference, a first performance metric for the first AI / ML model. The method 900 further includes determining 908, based on the first performance metric, to move the first AI / ML 214929-5814-3301'1 P71475WO1model to the active mode. The method 900 further includes moving 910 the first AI / ML model to the active mode.
[0147] In some embodiments of the method 900, the determining, based on the first performance metric, to move the first Al / ML model to the active mode comprises determining that the first performance metric meets a threshold. In some such embodiments, the method 900 further includes receiving, from the base station, the threshold.
[0148] In some embodiments, the method 900 further includes determining a second performance metric corresponding to the second inference of the second AI / ML model; wherein the determining, based on the first performance metric, to move the first AI / ML model to the active mode comprises determining that the first performance metric is better than the second performance metric.
[0149] In some embodiments of the method 900, the determining to move the first AI / ML model to the active mode is further based on a channel condition between the UE and the base station.
[0150] In some embodiments of the method 900, the first performance metric indicates an accuracy of the first AI / ML model.
[0151] In some embodiments of the method 900, the first performance metric indicates a latency of the first AI / ML model.
[0152] In some embodiments of the method 900, the configuration information comprises a bitmap that identifies the first AI / ML model for use with the shadow execution-based monitoring.
[0153] In some embodiments, the method 900 further includes moving the second Al / ML model from the active mode to the inactive mode in response to moving the first AI / ML model to the active mode.
[0154] FIG. 10 illustrates an example architecture of a wireless communication system 1000, according to embodiments disclosed herein. The following description is provided for an example wireless communication system 1000 that operates in conjunction with the LTE system standards and / or 5G or NR system standards as provided by 3GPP technical specifications.
[0155] As shown by FIG. 10, the wireless communication system 1000 includes UE 1002 and UE 1004 (although any number of UEs may be used). In this example, the UE 1002 and the UE 1004 are illustrated as smartphones (e.g., handheld touchscreen mobile computing devices connectable to one or more cellular networks), but may also comprise any mobile or non-mobile computing device configured for wireless communication.224929-5814-3301'1 P71475WO1
[0156] The UE 1002 and UE 1004 may be configured to communicatively couple with a RAN 1006. In embodiments, the RAN 1006 may be NG-RAN, E-UTRAN, etc. The UE 1002 and UE 1004 utilize connections (or channels) (shown as connection 1008 and connection 1010, respectively) with the RAN 1006, each of which comprises a physical communications interface. The RAN 1006 can include one or more base stations (such as base station 1012 and base station 1014) that enable the connection 1008 and connection 1010.
[0157] In this example, the connection 1008 and connection 1010 are air interfaces to enable such communicative coupling, and may be consistent with RAT(s) used by the RAN 1006, such as, for example, an LTE and / or NR.
[0158] In some embodiments, the UE 1002 and UE 1004 may also directly exchange communication data via a sidelink interface 1016. The UE 1004 is shown to be configured to access an access point (shown as AP 1018) via connection 1020. By way of example, the connection 1020 can comprise a local wireless connection, such as a connection consistent with any IEEE 802.11 protocol, wherein the AP 1018 may comprise a Wi-Fi® router. In this example, the AP 1018 may be connected to another network (for example, the Internet) without going through a CN 1024.
[0159] In embodiments, the UE 1002 and UE 1004 can be configured to communicate using orthogonal frequency division multiplexing (OFDM) communication signals with each other or with the base station 1012 and / or the base station 1014 over a multicarrier communication channel in accordance with various communication techniques, such as, but not limited to, an orthogonal frequency division multiple access (OFDMA) communication technique (e.g., for downlink communications) or a single carrier frequency division multiple access (SC-FDMA) communication technique (e g., for uplink and ProSe or sidelink communications), although the scope of the embodiments is not limited in this respect. The OFDM signals can comprise a plurality of orthogonal subcarriers.
[0160] In some embodiments, all or parts of the base station 1012 or base station 1014 may be implemented as one or more software entities running on server computers as part of a virtual network. In addition, or in other embodiments, the base station 1012 or base station 1014 may be configured to communicate with one another via interface 1022. In embodiments where the wireless communication system 1000 is an LTE system (e.g., when the CN 1024 is an EPC), the interface 1022 may be an X2 interface. The X2 interface may be defined between two or more base stations (e.g., two or more eNBs and the like) that connect to an EPC, and / or between two eNBs connecting to the EPC. In embodiments where 234929-5814-3301'1 P71475WO1the wireless communication system 1000 is an NR system (e.g., when CN 1024 is a 5GC). the interface 1022 may be an Xn interface. The Xn interface is defined between two or more base stations (e.g., two or more gNBs and the like) that connect to 5GC, between a base station 1012 (e.g., a gNB) connecting to 5GC and an eNB, and / or between two eNBs connecting to 5GC (e.g., CN 1024).
[0161] The RAN 1006 is shown to be communicatively coupled to the CN 1024. The CN 1024 may comprise one or more network elements 1026, which are configured to offer various data and telecommunications services to customers / subscribers (e.g., users of UE 1002 and UE 1004) who are connected to the CN 1024 via the RAN 1006. The components of the CN 1024 may be implemented in one physical device or separate physical devices including components to read and execute instructions from a machine-readable or computer-readable medium (e.g.. a non-transitory machine-readable storage medium).
[0162] In embodiments, the CN 1024 may be an EPC, and the RAN 1006 may be connected with the CN 1024 via an SI interface 1028. In embodiments, the SI interface 1028 may be split into two parts, an SI user plane (Sl-U) interface, which carries traffic data between the base station 1012 or base station 1014 and a serving gateway (S-GW), and the SI -MME interface, which is a signaling interface between the base station 1012 or base station 1014 and mobility management entities (MMEs).
[0163] In embodiments, the CN 1024 may be a 5GC, and the RAN 1006 may be connected with the CN 1024 via an NG interface 1028. In embodiments, the NG interface 1028 may be split into two parts, an NG user plane (NG-U) interface, which carries traffic data between the base station 1012 or base station 1014 and a user plane function (UPF), and the SI control plane (NG-C) interface, which is a signaling interface between the base station 1012 or base station 1014 and access and mobility management functions (AMFs).
[0164] Generally, an application server 1030 may be an element offering applications that use internet protocol (IP) bearer resources with the CN 1024 (e.g.. packet switched data services). The application server 1030 can also be configured to support one or more communication services (e.g., VoIP sessions, group communication sessions, etc.) for the UE 1002 and UE 1004 via the CN 1024. The application server 1030 may communicate with the CN 1024 through an IP communications interface 1032.
[0165] FIG. 11 illustrates a system 1100 for performing signaling 1134 between a wireless device 1102 and a network device 1118, according to embodiments disclosed herein. The system 1100 may be a portion of a wireless communications system as herein described. The wireless device 1102 may be, for example, a UE of a wireless communication system.244929-5814-3301'1 P71475WO1The network device 1118 may be, for example, a base station (e.g., an eNB or a gNB) of a wireless communication system.
[0166] The wireless device 1102 may include one or more processor(s) 1104. The processor(s) 1104 may execute instructions such that various operations of the wireless device 1102 are performed, as described herein. The processor(s) 1104 may include one or more baseband processors implemented using, for example, a central processing unit (CPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a controller, a field programmable gate array (FPGA) device, another hardware device, a firmware device, or any combination thereof configured to perform the operations described herein.
[0167] The wireless device 1102 may include a memory 1106. The memory 1106 may be a non-transitory computer-readable storage medium that stores instructions 1108 (which may include, for example, the instructions being executed by the processor(s) 1104). The instructions 1108 may also be referred to as program code or a computer program. The memory 1106 may also store data used by, and results computed by. the processor(s) 1104.
[0168] The wireless device 1102 may include one or more transceiver(s) 1110 that may include RF transmitter circuitry' and / or receiver circuitry that use the antenna(s) 1112 of the wireless device 1102 to facilitate signaling (e.g., the signaling 1134) to and / or from the wireless device 1102 with other devices (e.g., the network device 1118) according to corresponding RATs.
[0169] The wireless device 1102 may include one or more antenna(s) 1112 (e.g., one, two, four, or more). For embodiments with multiple antenna(s) 1112, the wireless device 1102 may leverage the spatial diversity of such multiple antenna(s) 1112 to send and / or receive multiple different data streams on the same time and frequency resources. This behavior may be referred to as, for example, multiple input multiple output (MIMO) behavior (referring to the multiple antennas used at each of a transmitting device and a receiving device that enable this aspect). MIMO transmissions by the wireless device 1102 may be accomplished according to precoding (or digital beamforming) that is applied at the wireless device 1102 that multiplexes the data streams across the antenna(s) 1112 according to known or assumed channel characteristics such that each data stream is received with an appropriate signal strength relative to other streams and at a desired location in the spatial domain (e.g., the location of a receiver associated with that data stream). Certain embodiments may use single user MIMO (SU-MIMO) methods (where the data streams are all directed to a single receiver) and / or multi user MIMO (MU-MIMO) methods (where 254929-5814-3301'1 P71475WO1individual data streams may be directed to individual (different) receivers in different locations in the spatial domain).
[0170] In certain embodiments having multiple antennas, the wireless device 1102 may implement analog beamforming techniques, whereby phases of the signals sent by the antenna(s) 1112 are relatively adjusted such that the (joint) transmission of the antenna(s) 1112 can be directed (this is sometimes referred to as beam steering).
[0171] The wireless device 1102 may include one or more interface(s) 1114. The interface(s) 1114 may be used to provide input to or output from the wireless device 1102. For example, a wireless device 1102 that is a UE may include interface(s) 1114 such as microphones, speakers, a touchscreen, buttons, and the like in order to allow for input and / or output to the UE by a user of the UE. Other interfaces of such a UE may be made up of transmitters, receivers, and other circuitry (e.g., other than the transceiver(s)1110 / antenna(s) 1112 already described) that allow for communication between the UE and other devices and may operate according to known protocols (e.g., Wi-Fi®, Bluetooth®, and the like).
[0172] The wireless device 1102 may include an AI / ML model activation module 1116. The AI / ML model activation module 1116 may be implemented via hardware, software, or combinations thereof. For example, the AI / ML model activation module 1116 may be implemented as a processor, circuit, and / or instructions 1108 stored in the memory 1106 and executed by the processor(s) 1104. In some examples, the AI / ML model activation module 1116 may be integrated within the processor(s) 1104 and / or the transceiver(s) 1110. For example, the AI / ML model activation module 1116 may be implemented by a combination of software components (e.g., executed by a DSP or a general processor) and hardware components (e.g., logic gates and circuitry) within the processor(s) 1104 or the transceiver(s) 1110.
[0173] The AI / ML model activation module 1116 may be used for various aspects of the present disclosure, for example, aspects of FIG. 4, FIG. 6, FIG. 7, and / or FIG. 9. The AI / ML model activation module 1116 may configure the wireless device 1102 to receive, from a base station, configuration information configuring the UE to monitor a first AI / ML model using dataset-based monitoring while the first AI / ML model is in an inactive mode, wherein the inactive mode comprises inferencing result suppression; receive, from the base station, a dataset corresponding to a use of the dataset-based monitoring; generate, while the first AI / ML model is in the inactive mode, a first inference using the first AI / ML model by applying the dataset at the first AI / ML model; determine, based on the first inference, a first 264929-5814-3301'1 P71475WO1performance metric for the first AI / ML model; send, to the base station, the first performance metric for the first AI / ML model; receive, in response to the first performance metric, an instruction to move the first AI / ML model to an active mode, wherein the active mode comprises inferencing result transmission; and move, in response to the instruction, the first AI / ML model to the active mode, in the manner discussed herein.
[0174] In some cases, the AI / ML model activation module 1132 may configure the network wireless device 1102 to receive, from a base station, configuration information configuring the UE to monitor a first AI / ML model using dataset-based monitoring while the first AI / ML model is in an inactive mode, wherein the inactive mode comprises inferencing result suppression; receive, from the base station, a dataset corresponding to a use of the dataset-based monitoring; generate, while the first AI / ML model is in the inactive mode, a first inference using the first AI / ML model by applying the dataset at the first AI / ML model; determine a first performance metric for the first AI / ML model corresponding to the first inference; determine, based on the first performance metric, to move the first AI / ML model to an active mode, wherein the active mode comprises inferencing result transmission; and move the first AI / ML model to the active mode, in the manner discussed herein.
[0175] In some cases, the AI / ML model activation module 1132 may configure the network wireless device 1102 to receive, from a base station, configuration information configuring the UE to monitor a first AI / ML model using shadow execution-based monitoring while the first AI / ML model is in an inactive mode, wherein the inactive mode comprises inferencing result suppression; generate, according to a use of the shadow execution-based monitoring, while the first AI / ML model is in the inactive mode, a first inference using the first AI / ML model by applying, at the first AI / ML model, an input that is also applied at a second AI / ML model that is in an active mode to generate a second inference, wherein the active mode comprises inferencing result transmission; determine, based on the first inference, a first performance metric for the first AI / ML model; send, to the base station, the first performance metric for the first AI / ML model; receive, in response to the first performance metric, an instruction to move the first AI / ML to the active mode; and move, in response to the instruction, the first AI / ML model to the active mode, in the manner discussed herein.
[0176] In some cases, the AI / ML model activation module 1132 may configure the network wireless device 1102 to receive, from a base station, configuration information configuring the UE to monitor a first AI / ML model using shadow execution-based274929-5814-3301'1 P71475WO1monitoring while the first AI / ML model is in an inactive mode, wherein the inactive mode comprises inferencing result suppression; generate, according to a use of the shadow execution-based monitoring, while the first AI / ML model is in the inactive mode, a first inference using the first AI / ML model by applying, at the first AI / ML model, an input that is also applied at a second AI / ML model that is in an active mode to generate a second inference, wherein the active mode comprises inferencing result transmission; determine, based on the first inference, a first performance metric for the first AI / ML model; determine, based on the first performance metric, to move the first AI / ML model to the active mode; and move the first AI / ML model to the active mode, in the manner discussed herein.
[0177] The network device 1118 may include one or more processor(s) 1120. The processor(s) 1120 may execute instructions such that various operations of the network device 1118 are performed, as described herein. The processor(s) 1120 may include one or more baseband processors implemented using, for example, a CPU, a DSP, an ASIC, a controller, an FPGA device, another hardw are device, a firmware device, or any combination thereof configured to perform the operations described herein.
[0178] The network device 1118 may include a memory 1122. The memory 1122 may be a non-transitory computer-readable storage medium that stores instructions 1124 (which may include, for example, the instructions being executed by the processor(s) 1120). The instructions 1124 may also be referred to as program code or a computer program. The memory 1122 may also store data used by, and results computed by, the processor(s) 1120.
[0179] The network device 1118 may include one or more transceiver(s) 1126 that may include RF transmitter circuitry and / or receiver circuitry that use the antenna(s) 1128 of the network device 1118 to facilitate signaling (e.g., the signaling 1134) to and / or from the network device 1118 with other devices (e.g., the wireless device 1102) according to corresponding RATs.
[0180] The network device 1118 may include one or more antenna(s) 1128 (e.g., one, two, four, or more). In embodiments having multiple antenna(s) 1128, the network device 1118 may perform MIMO, digital beamforming, analog beamforming, beam steering, etc., as has been described.
[0181] The netw ork device 1118 may include one or more interface(s) 1130. The interface(s) 1130 may be used to provide input to or output from the network device 1118. For example, a network device 1118 that is a base station may include interface(s) 1130 made up of transmitters, receivers, and other circuitry (e.g., other than the transceiver(s)284929-5814-3301'1 P71475WO11126 / antenna(s) 1128 already described) that enables the base station to communicate with other equipment in a core network, and / or that enables the base station to communicate with external networks, computers, databases, and the like for purposes of operations, administration, and maintenance of the base station or other equipment operably connected thereto.
[0182] The network device 1118 may include an AI / ML model activation module 1132. The AI / ML model activation module 1132 may be implemented via hardware, software, or combinations thereof. For example, the AI / ML model activation module 1132 may be implemented as a processor, circuit, and / or instructions 1124 stored in the memory 1122 and executed by the processor(s) 1120. In some examples, the AI / ML model activation module 1132 may be integrated within the processor(s) 1120 and / or the transceiver(s) 1126. For example, the AI / ML model activation module 1132 may be implemented by a combination of software components (e.g., executed by a DSP or a general processor) and hardware components (e.g., logic gates and circuitry) within the processor(s) 1120 or the transceiver(s) 1126.
[0183] The AI / ML model activation module 1132 may be used for various aspects of the present disclosure, for example, aspects of FIG. 5 and / or FIG. 8. In some cases, the AI / ML model activation module 1132 may configure the network device 1118 to send, to aUE, configuration information configuring the UE to monitor a first AI / ML model using datasetbased monitoring while the first AI / ML model is in an inactive mode, wherein the inactive mode comprises inferencing result suppression; send, to the UE, a dataset corresponding to a use of the dataset-based monitoring; receive, from the UE, a first performance metric for the first AI / ML model that is based on the dataset; determine, based on the first performance metric, to send the UE an instruction to move the first AI / ML model to an active mode, wherein the active mode comprises inferencing result transmission; and send the instruction to the UE, in the manner discussed herein.
[0184] In some cases, the AI / ML model activation module 1132 may configure the network device 1118 to send, to a UE, configuration information configuring the UE to monitor a first AI / ML model using shadow execution-based monitoring while the first AI / ML model is in an inactive mode, wherein the inactive mode comprises inferencing result suppression; receive, from the UE, a first performance metric for the first AI / ML model according to the shadow execution-based monitoring; determine, based on the first performance metric, to send the UE an instruction to move the first AI / ML model to an294929-5814-3301'1 P71475WO1active mode, wherein the active mode comprises inferencing result transmission; and send the instruction to the UE, in the manner discussed herein.
[0185] Embodiments contemplated herein include an apparatus comprising means to perform one or more elements of any one of the method 400. the method 600, the method 700, and / or the method 900. This apparatus may be, for example, an apparatus of a UE (such as a wireless device 1102 that is a UE, as described herein).
[0186] Embodiments contemplated herein include one or more non-transitory computer-readable media comprising instructions to cause an electronic device, upon execution of the instructions by one or more processors of the electronic device, to perform one or more elements of any one of the method 400, the method 600, the method 700, and / or the method 900. This non-transitory computer-readable media may be. for example, a memory of a UE (such as a memory 1106 of a wireless device 1102 that is a UE, as described herein).
[0187] Embodiments contemplated herein include an apparatus comprising logic, modules, or circuitry to perform one or more elements of any one of the method 400. the method 600. the method 700, and / or the method 900. This apparatus may be, for example, an apparatus of a UE (such as a wireless device 1102 that is a UE, as described herein).
[0188] Embodiments contemplated herein include an apparatus comprising: one or more processors and one or more computer-readable media comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform one or more elements of any one of the method 400, the method 600, the method 700, and / or the method 900. This apparatus may be, for example, an apparatus of a UE (such as a wireless device 1102 that is a UE, as described herein).
[0189] Embodiments contemplated herein include a signal as described in or related to one or more elements of any one of the method 400. the method 600, the method 700, and / or the method 900.
[0190] Embodiments contemplated herein include a computer program or computer program product comprising instructions, wherein execution of the program by a processor is to cause the processor to carry out one or more elements of any one of the method 400, the method 600, the method 700, and / or the method 900. The processor may be a processor of a UE (such as a processor(s) 1104 of a wireless device 1102 that is a UE, as described herein). These instructions may be. for example, located in the processor and / or on a memory of the UE (such as a memory 1106 of a wireless device 1102 that is a UE. as described herein).304929-5814-3301'1 P71475WO1
[0191] Embodiments contemplated herein include an apparatus comprising means to perform one or more elements of any one of the method 500 and / or the method 800. This apparatus may be, for example, an apparatus of a base station (such as a network device 1118 that is a base station, as described herein).
[0192] Embodiments contemplated herein include one or more non-transitory computer-readable media comprising instructions to cause an electronic device, upon execution of the instructions by one or more processors of the electronic device, to perform one or more elements of any one of the method 500 and / or the method 800. This non-transitory computer-readable media may be, for example, a memory of a base station (such as a memory 1122 of a network device 1118 that is a base station, as described herein).
[0193] Embodiments contemplated herein include an apparatus comprising logic, modules, or circuitry to perform one or more elements of any one of the method 500 and / or the method 800. This apparatus may be, for example, an apparatus of a base station (such as a network device 1118 that is a base station, as described herein).
[0194] Embodiments contemplated herein include an apparatus comprising: one or more processors and one or more computer-readable media comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform one or more elements of any one of the method 500 and / or the method 800. This apparatus may be. for example, an apparatus of a base station (such as a network device 1118 that is a base station, as described herein).
[0195] Embodiments contemplated herein include a signal as described in or related to one or more elements of any one of the method 500 and / or the method 800.
[0196] Embodiments contemplated herein include a computer program or computer program product comprising instructions, wherein execution of the program by a processing element is to cause the processing element to carry out one or more elements of any one of the method 500 and / or the method 800. The processor may be a processor of a base station (such as a processor(s) 1120 of a network device 1118 that is a base station, as described herein). These instructions may be. for example, located in the processor and / or on a memory of the base station (such as a memory 1122 of a network device 1118 that is a base station, as described herein).
[0197] For one or more embodiments, at least one of the components set forth in one or more of the preceding figures may be configured to perform one or more operations, techniques, processes, and / or methods as set forth herein. For example, a baseband processor as described herein in connection with one or more of the preceding figures may 314929-5814-3301'1 P71475WO1be configured to operate in accordance with one or more of the examples set forth herein. For another example, circuitry associated with a UE, base station, network element, etc. as described above in connection with one or more of the preceding figures may be configured to operate in accordance with one or more of the examples set forth herein.
[0198] Any of the above described embodiments may be combined with any other embodiment (or combination of embodiments), unless explicitly stated otherwise. The foregoing description of one or more implementations provides illustration and description, but is not intended to be exhaustive or to limit the scope of embodiments to the precise form disclosed. Modifications and variations are possible in light of the above teachings or may be acquired from practice of various embodiments.
[0199] Embodiments and implementations of the systems and methods described herein may include various operations, which may be embodied in machine-executable instructions to be executed by a computer system. A computer system may include one or more general-purpose or special-purpose computers (or other electronic devices). The computer system may include hardware components that include specific logic for performing the operations or may include a combination of hardware, software, and / or firmware.
[0200] It should be recognized that the systems described herein include descriptions of specific embodiments. These embodiments can be combined into single systems, partially combined into other systems, split into multiple systems or divided or combined in other ways. In addition, it is contemplated that parameters, attributes, aspects, etc. of one embodiment can be used in another embodiment. The parameters, attributes, aspects, etc. are merely described in one or more embodiments for clarity, and it is recognized that the parameters, attributes, aspects, etc. can be combined with or substituted for parameters, attributes, aspects, etc. of another embodiment unless specifically disclaimed herein.
[0201] It is well understood that the use of personally identifiable information should follow privacy policies and practices that are generally recognized as meeting or exceeding industry or governmental requirements for maintaining the privacy of users. In particular, personally identifiable information data should be managed and handled so as to minimize risks of unintentional or unauthorized access or use, and the nature of authorized use should be clearly indicated to users.
[0202] Although the foregoing has been described in some detail for purposes of clarity, it will be apparent that certain changes and modifications may be made without departing from the principles thereof. It should be noted that there are many alternative ways of implementing both the processes and apparatuses described herein. Accordingly, the present 324929-5814-3301'1 P71475WO1embodiments are to be considered illustrative and not restrictive, and the description is not to be limited to the details given herein, but may be modified within the scope and equivalents of the appended claims.334929-5814-3301'1 P71475WO1
Claims
CLAIMS1. A method of a user equipment (UE), comprising:receiving, from a base station, configuration information configuring the UE to monitor a first artificial intelligence (AI) / machine learning (ML) model using dataset-based monitoring while the first AT / ML model is in an inactive mode, wherein the inactive mode comprises inferencing result suppression;receiving, from the base station, a dataset corresponding to a use of the dataset-based monitoring;generating, while the first AI / ML model is in the inactive mode, a first inference using the first AI / ML model by applying the dataset at the first AI / ML model;determining, based on the first inference, a first performance metric for the first AI / ML model;sending, to the base station, the first performance metric for the first AI / ML model; receiving, in response to the first performance metric, an instruction to move the first AI / ML model to an active mode, wherein the active mode comprises inferencing result transmission; andmoving, in response to the instruction, the first AI / ML model to the active mode.
2. The method of claim 1, further comprising:determining, based on a second inference of a second AI / ML model, a second performance metric for the second AI / ML model; andsending, to the base station, the second performance metric for the second AI / ML model;wherein the instruction to move the first AI / ML to the active mode is further received in response to the second performance metric.
3. The method of claim 1, wherein the first performance metric indicates one of:an accuracy of the first AI / ML model; anda latency of the first AI / ML model.
4. The method of claim 1, wherein the configuration information comprises a bitmap that identifies the first AI / ML model for use with the dataset-based monitoring.
5. The method of claim 1. wherein the dataset comprises one of:synthetic data; andfield data.344929-5814-3301'1 P71475WO16. The method of claim 1, further comprising moving a second AI / ML model from the active mode to the inactive mode in response to moving the first AI / ML model to the active mode.
7. The method of claim 1. wherein the first AI / ML model is configured to use fewer computational resources when in the inactive mode as compared to when in the active mode.
8. A method of a base station, comprising:sending, to a user equipment (UE). configuration information configuring the UE to monitor a first artificial intelligence (AI) / machine learning (ML) model using dataset-based monitoring while the first AI / ML model is in an inactive mode, wherein the inactive mode comprises inferencing result suppression;sending, to the UE, a dataset corresponding to a use of the dataset-based monitoring; receiving, from the UE, a first performance metric for the first AI / ML model that is based on the dataset;determining, based on the first performance metric, to send the UE an instruction to move the first AI / ML model to an active mode, wherein the active mode comprises inferencing result transmission; andsending the instruction to the UE.
9. The method of claim 8, wherein the determining, based on the first performance metric, to send the instruction to move the first AI / ML model to the active mode comprises determining that the first performance metric meets a threshold.
10. The method of claim 8, wherein the determining, based on the first performance metric, to send the instruction to move the first AI / ML model to the active mode comprises determining that the first performance metric is better than a second performance metric for a second AI / ML model that is in the active mode.
11. The method of claim 10, further comprising receiving the second performance metric from the UE.
12. The method of claim 8, wherein the first performance metric indicates one of:an accuracy of the first AI / ML model; anda latency of the first AI / ML model.
13. The method of claim 8, wherein the configuration information comprises a bitmap that identifies the first AI / ML model for use with the dataset-based monitoring.354929-5814-3301'1 P71475WO114. The method of claim 8, wherein the dataset comprises one of:synthetic data; andfield data.
15. The method of claim 8, wherein the first AI / ML model is configured to use fewer computational resources when in the inactive mode as compared to when in the active mode.
16. A method of a user equipment (UE), comprising:receiving, from a base station, configuration information configuring the UE to monitor a first artificial intelligence (AI) / machine learning (ML) model using dataset-based monitoring while the first AI / ML model is in an inactive mode, wherein the inactive mode comprises inferencing result suppression;receiving, from the base station, a dataset corresponding to a use of the dataset-based monitoring;generating, while the first AI / ML model is in the inactive mode, a first inference using the first AI / ML model by applying the dataset at the first AI / ML model;determining a first performance metric for the first AI / ML model corresponding to the first inference;determining, based on the first performance metric, to move the first AI / ML model to an active mode, wherein the active mode comprises inferencing result transmission; and moving the first AI / ML model to the active mode.
17. The method of claim 16, wherein the determining, based on the first performance metric, to move the first AI / ML model to the active mode comprises determining that the first performance metric meets a threshold.
18. The method of claim 17, further comprising receiving, from the base station, the threshold.
19. The method of claim 16, further comprising determining a second performance metric corresponding to a second inference of a second AI / ML model that is in the active mode; wherein the determining, based on the first performance metric, to move the first AI / ML model to the active mode comprises determining that the first performance metric is better than the second performance metric.
20. The method of claim 16, wherein the determining to move the first AI / ML model to the active mode is further based on a channel condition between the UE and the base station.364929-5814-3301'1 P71475WO121. The method of claim 16, wherein the first performance metric indicates one of:an accuracy of the first AI / ML model; anda latency of the first AI / ML model.
22. The method of claim 16, wherein the configuration information comprises a bitmap that identifies the first AI / ML model for use with the dataset-based monitoring.
23. The method of claim 16, wherein the dataset comprises one of:synthetic data; andfield data.
24. The method of claim 16, further comprising moving a second AI / ML model from the active mode to the inactive mode in response to moving the first AI / ML model to the active mode.
25. The method of claim 16, wherein the first AI / ML model is configured to use fewer computational resources when in the inactive mode as compared to when in the active mode.
26. A method of a user equipment (UE), comprising:receiving, from a base station, configuration information configuring the UE to monitor a first artificial intelligence (AI) / machine learning (ML) model using shadow execution-based monitoring while the first AI / ML model is in an inactive mode, wherein the inactive mode comprises inferencing result suppression;generating, according to a use of the shadow execution-based monitoring, while the first AI / ML model is in the inactive mode, a first inference using the first AI / ML model by applying, at the first AI / ML model, an input that is also applied at a second AI / ML model that is in an active mode to generate a second inference, wherein the active mode comprises inferencing result transmission;determining, based on the first inference, a first performance metric for the first AI / ML model;sending, to the base station, the first performance metric for the first AI / ML model; receiving, in response to the first performance metric, an instruction to move the first AI / ML to the active mode; andmoving, in response to the instruction, the first AI / ML model to the active mode.
27. The method of claim 26, further comprising:374929-5814-3301'1 P71475WO1determining, based on the second inference, a second performance metric for the second AI / ML model; andsending, to the base station, the second performance metric for the second AI / ML model;wherein the instruction to move the first AI / ML to the active mode is further received in response to the second performance metric.
28. The method of claim 26, wherein the first performance metric indicates one of:an accuracy of the first AI / ML model; anda latency of the first AI / ML model.
29. The method of claim 26, wherein the configuration information comprises a bitmap that identifies the first AI / ML model for use with the shadow execution-based monitoring.
30. The method of claim 26, further comprising moving the second AI / ML model from the active mode to the inactive mode in response to moving the first AI / ML model to the active mode.
31. The method of claim 26, wherein the first AI / ML model is configured to use fewer computational resources when in the inactive mode as compared to when in the active mode.
32. A method of a base station, comprising:sending, to a user equipment (UE), configuration information configuring the UE to monitor a first artificial intelligence (AI) / machine learning (ML) model using shadow execution-based monitoring while the first AI / ML model is in an inactive mode, wherein the inactive mode comprises inferencing result suppression;receiving, from the UE, a first performance metric for the first AI / ML model according to the shadow execution-based monitoring;determining, based on the first performance metric, to send the UE an instruction to move the first AI / ML model to an active mode, wherein the active mode comprises inferencing result transmission; andsending the instruction to the UE.
33. The method of claim 32, wherein the determining, based on the first performance metric, to send the instruction to move the first AI / ML model to the active mode comprises determining that the first performance metric meets a threshold.384929-5814-3301'1 P71475WO134. The method of claim 32, wherein the determining, based on the first performance metric, to send the instruction to move the first AI / ML model to the active mode comprises determining that the first performance metric is better than a second performance metric for a second AI / ML model that is in the active mode.
35. The method of claim 34, further comprising receiving the second performance metric from the UE.
36. The method of claim 32, wherein the first performance metric indicates one of:an accuracy of the first AI / ML model; anda latency of the first AI / ML model.
37. The method of claim 32, wherein the configuration information comprises a bitmap that identifies the first AI / ML model for use with the shadow execution-based monitoring.
38. The method of claim 32, wherein the first AI / ML model is configured to use fewer computational resources when in the inactive mode as compared to when in the active mode.
39. A method of a user equipment (UE), comprising:receiving, from a base station, configuration information configuring the UE to monitor a first artificial intelligence (AI) / machine learning (ML) model using shadow execution-based monitoring while the first AI / ML model is in an inactive mode, wherein the inactive mode comprises inferencing result suppression;generating, according to a use of the shadow execution-based monitoring, while the first AI / ML model is in the inactive mode, a first inference using the first AI / ML model by applying, at the first AI / ML model, an input that is also applied at a second AI / ML model that is in an active mode to generate a second inference, wherein the active mode comprises inferencing result transmission;determining, based on the first inference, a first performance metric for the first AI / ML model;determining, based on the first performance metric, to move the first AI / ML model to the active mode; andmoving the first AI / ML model to the active mode.
40. The method of claim 39, wherein the determining, based on the first performance metric, to move the first AI / ML model to the active mode comprises determining that the first performance metric meets a threshold.394929-5814-3301'1 P71475WO141. The method of claim 40, further comprising receiving, from the base station, the threshold.
42. The method of claim 39, further comprising determining a second performance metric corresponding to the second inference of the second AI / ML model;wherein the determining, based on the first performance metric, to move the first AI / ML model to the active mode comprises determining that the first performance metric is better than the second performance metric.
43. The method of claim 39, wherein the determining to move the first AI / ML model to the active mode is further based on a channel condition between the UE and the base station.
44. The method of claim 39, wherein the first performance metric indicates one of:an accuracy of the first AI / ML model; anda latency of the first AI / ML model.
45. The method of claim 39, wherein the configuration information comprises a bitmap that identifies the first AI / ML model for use with the shadow execution-based monitoring.
46. The method of claim 39, further comprising moving the second AI / ML model from the active mode to the inactive mode in response to moving the first AI / ML model to the active mode.
47. An apparatus comprising means to perform the method of any of claim 1 to claim 46.
48. A computer-readable media comprising instructions to cause an electronic device, upon execution of the instructions by one or more processors of the electronic device, to perform the method of any of claim 1 to claim 46.
49. An apparatus comprising logic, modules, or circuitry to perform the method of any of claim 1 to claim 46.
50. A baseband processor for a user equipment (UE) that is configured to cause the UE to perform one or more elements of any one of claim 1 to claim 7, claim 16 to claim 31, and claim 39 to claim 46.
51. A baseband processor for a base station that is configured to cause the base station to perform one or more elements of any one of claim 8 to claim 15 and claim 32 to claim 38.404929-5814-3301'1 P71475WO1