Model scalability through radio propagation specific adaptation for artificial intelligence / machine learning models
By dynamically scaling and reconfiguring AI/ML models to match radio propagation conditions, the system achieves improved performance and reduced overhead, addressing inefficiencies in existing wireless communication systems.
Patent Information
- Application Number
- PCT/US2025/021562
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-02
- Filing Date
- 2025-03-26
- Publication Date
- 2025-10-09
AI Technical Summary
Existing wireless communication systems face challenges in achieving performance improvements through AI/ML models due to mismatches between radio propagation conditions and the configured size of the AI/ML models, leading to resource waste or system interruptions.
Adapting and scaling the operation of AI/ML models to match current radio propagation conditions by dynamically adjusting the size of beam sets and reconfiguring pre- and post-processing blocks, allowing the models to operate efficiently across varying conditions.
This approach reduces resource waste and interruptions by ensuring a balanced AI/ML model complexity that aligns with current radio conditions, enhancing system performance and reducing measurement overhead.
Smart Images

Figure US2025021562_09102025_PF_FP_ABST
Abstract
Description
MODEL SCALABILITY THROUGH RADIO PROPAGATION SPECIFICADAPTATION FOR ARTIFICIAL INTELLIGENCE / MACHINE LEARNINGMODELSTECHNICAL FIELD
[0001] This application relates generally to wireless communication systems, including wireless communication systems implementing mechanisms for scaling artificial intelligence (AI) / machine learning (ML) models.BACKGROUND
[0002] Wireless mobile communication technology uses various standards and protocols to transmit data between a base station and a wireless communication device. Wireless communication system standards and protocols can include, for example, 3rd Generation Partnership Project (3GPP) Long Term Evolution (LTE) (e.g., 4G), 3GPP New Radio (NR) (e.g., 5G), and Institute of Electrical and Electronics Engineers (IEEE) 802.11 standard for Wireless Local Area Networks (WLAN) (commonly known to industry groups as Wi-Fi®).
[0003] As contemplated by the 3GPP, different wireless communication systems' standards and protocols can use various radio access networks (RANs) for communicating between a base station of the RAN (which may also sometimes be referred to generally as a RAN node, a network node, or simply a node) and a wireless communication device known as a user equipment (UE). 3GPP RANs can include, for example. Global System for Mobile communications (GSM), Enhanced Data Rates for GSM Evolution (EDGE) RAN (GERAN), Universal Terrestrial Radio Access Network (UTRAN), Evolved Universal Terrestrial Radio Access Network (E-UTRAN), and / or Next-Generation Radio Access Network (NG-RAN).
[0004] Each RAN may use one or more radio access technologies (RATs) to perform communication between the base station and the UE. For example, the GERAN implements GSM and / or EDGE RAT, the UTRAN implements Universal Mobile Telecommunication System (UMTS) RAT or other 3GPP RAT, the E-UTRAN implements LTE RAT (sometimes simply referred to as LTE), and NG-RAN implements NR RAT (sometimes referred to herein as 5G RAT, 5G NR RAT, or simply NR). Incertain deployments, the E-UTRAN may also implement NR RAT. In certain deployments, NG-RAN may also implement LTE RAT.
[0005] A base station used by a RAN may correspond to that RAN. One example of an E-UTRAN base station is an Evolved Universal Terrestrial Radio Access Network (E- UTRAN) Node B (also commonly denoted as evolved Node B, enhanced Node B, eNodeB, or eNB). One example of an NG-RAN base station is a next generation Node B (also sometimes referred to as a g Node B or gNB).
[0006] A RAN provides its communication services with external entities through its connection to a core network (CN). For example, E-UTRAN may utilize an Evolved Packet Core (EPC) while NG-RAN may utilize a 5G Core Network (5GC).BRIEF DESCRIPTION OF THE SEVERAL VIEWS OF THE DRAWINGS
[0007] To easily identify the discussion of any particular element or act, the most significant digit or digits in a reference number refer to the figure number in which that element is first introduced.
[0008] FIG. 1 illustrates an example of grouping AI / ML models into various supporting conditions with an overarching AI / ML model functionality .
[0009] FIG. 2 illustrates an example of an AI / ML model that includes pre-processing blocks, an AI / ML core, and post-processing blocks, according to embodiments herein.
[0010] FIG. 3 illustrates a flow diagram for the scaling and / or configuring an AI / ML model at a UE, according to embodiments herein.
[0011] FIG. 4 illustrates a method of a UE operating an AI / ML model for network signal processing comprising an AI / ML core and first swappable processing blocks interfaced to the AI / ML core according to a first configuration of the AI / ML model, according to embodiments herein.
[0012] FIG. 5 illustrates a method of a UE operating an AI / ML model for network signal processing comprising an AI / ML core and first swappable processing blocks interfaced to the AI / ML core according to a first configuration of the AI / ML model, according to embodiments herein.
[0013] FIG. 6 illustrates a method of a base station in communication with a UE operating an AI / ML model for network signal processing comprising an AI / ML core and first swappable processing blocks interfaced to the AI / ML core according to a first configuration of the AI / ML model, according to embodiments herein.
[0014] FIG. 7 illustrates an example architecture of a wireless communication system, according to embodiments disclosed herein.
[0015] FIG. 8 illustrates a system for performing signaling between a wireless device and a network device, according to embodiments disclosed herein.DETAILED DESCRIPTION
[0016] Various embodiments are described with regard to a UE. However, reference to a UE is merely provided for illustrative purposes. The example embodiments may be utilized with any electronic component that may establish a connection to a network and is configured with the hardware, software, and / or firmware to exchange information and data with the network. Therefore, the UE as described herein is used to represent any appropriate electronic component.
[0017] In some wireless communication systems, the use of artificial intelligence (Al)Zmachine learning (ML) models may achieve performance improvements over non- AI / ML model operation in various circumstances. However, achieving performance improvements (e.g., achieving a reduction of measurement overhead, achieving reduced overhead through compression of channel state information (CSI), etc.) across all configurations and / or conditions may be challenging. Accordingly, it has been determined that adapting and / or scaling the operation of an AI / ML model to match the current radio propagation conditions may provide performance improvements. For example, for a beam management (BM) case and for a UE-sided AI / ML model, a downlink (DL) transmit (Tx) beam prediction is made for a first set (Set A) of one or more predicted DL Tx beams by an AI / ML model using actual measurement results of a (e.g.. different and / or smaller) second set (Set B) of one or more measured DL Tx beams. Note that there may be some size flexibility on the set B beams in such cases. The smaller the size of the set B beams, the smaller the overhead is for measurements and, correspondingly, the larger the size of the set B of beams, the larger the overhead is for measurements.
[0018] In some wireless communication mechanisms, the UE may be pre-configured to measure a set B of beams irrespective of the dynamics of radio conditions. Corresponding to such cases, during deployment there are two possible outcomes that correspond to a mismatch between the radio conditions and the configured 'size’ of the AI / ML model. In a first such outcome, the current radio conditions are favorable suchthat the achievement of the same throughput was otherwise possible using a smaller set of set B than what was actually used. Thus, there is a waste of resources (overdimensioned AI / ML model).
[0019] In a second such outcome, the radio conditions may be unfavorable, and the set B of beams is too small relative to the adverse radio conditions such that the spatial mapping / inference from set B to set A fails to preserve the system performance thus triggering interruptions (life cycle management (LCM) procedures of AI / ML model deactivation and / or fallback causing interruptions and / or unnecessary model transfer updating) that would not have occurred had a larger set B of beams been used (underdimensioned AI / ML model).
[0020] Embodiments discussed herein relate to the optimization of the tradeoff between AI / ML model complexity (e.g., a configured size of the AI / ML model) and any current radio propagation conditions (also referred to herein as “radio conditions”), so that a properly balanced / tuned AI / ML dimensionality for the radio conditions is achieved.
[0021] FIG. 1 illustrates an example of grouping AI / ML models into various supporting conditions with an overarching AI / ML model functionality.
[0022] In some wireless communication systems, AI / ML models may be grouped together into non-overlapping sets of supporting conditions as part of a UE capability report signaling.
[0023] For example, in the UE capability report signaling, with an overarching AI / ML model functionality 114 (e.g., a beam management functionality, a CSI compression / decompression functionality, and / or a positioning functionality, etc.), a first supporting condition(s) 102 may include, but is not limited to, various scenarios (e.g., urban macro (Uma) scenarios, urban micro (Umi) scenarios, indoor scenarios), signal to noise ratio (SNR) conditions, and / or Doppler effects.
[0024] Then, a number of first configurations for a set of AI / ML models may be grouped under each supporting condition. As shown, a first configuration 104 for a first model ("Model 1") may be configured for operation with first supporting condition(s) 102 may include various AI / ML model aspects and characteristics. For example, the first configuration 104 for the first model may operate according to a particular set of model input characteristics (e.g., input type, input size, pre-processing), a particular set of model output characteristics (e.g., output (inference) types, output (inference) size, post processing), a particular model complexity, particular performance metrics, particularmonitoring metrics, and / or additional conditions. In the case of an AI / ML model for beam management (where the AI / ML model functionality 114 is for beam management), it may be that particular additional conditions include set A and / or set B beam configuration(s), a particular pattern of set B beams, a particular shape of the Tx beams (e.g., 3 decibel (dB) bandwidth, pointing angles, beam shape), a particular wide set B beams (e.g., synchronization signal block (SSB)), and / or a particular narrow set B beams (e.g.. a channel state information reference signal (CSI-RS)).
[0025] Further, note that particular configurations for various additional models as may be compatible for use when the first supporting condition(s) 102 apply may also be available (e.g., a first configuration 106 for a second model ("Model 2"),.... a first configuration 108 for an Ath model ("Model N"). Note that these models may use same set or a different set of configuration items as have been described in relation to Model 1.
[0026] Then, a second configuration 116 for Model 1 may be configured for operation when second supporting condition(s) 110 apply. The second configuration 116 for Model 1 may use different values / choices for the various AI / ML model aspects and characteristics for Model 1 from those values / choices used by the first configuration 104 for Model 1. Similarly, a second configuration 120 for Model 2 may be configured for operation when the second supporting condition(s) 110 apply. The second configuration 120 for Model 2 may use different values / choices for the various AI / ML model aspects and characteristics for Model 2 from those values / choices used by the first configuration 106 for Model 2. Continuing on. a second configuration 124 for Model A may be configured for operation when the second supporting condition(s) 110 apply. The second configuration 124 for Model N may use values / choices for the various AI / ML model aspects and characteristics for Model N from those values / choices used by the first configuration 108 for Model N.
[0027] Subsequently, a Ath configuration 118 for Model 1 may be configured for operation when A th supporting condition(s) 112 apply. The Ath configuration 118 for Model 1 may use different values / choices for the various AI / ML model aspects and characteristics for Model 1 from those values / choices used by the first configuration 104 for Model 1 and / or the second configuration 116 for Model 1. Similarly, a Ath configuration 122 for Model 2 may be configured for operation when the Ath supporting condition(s) 112 apply. The Ath configuration 122 for Model 2 may use differentvalues / choices for the various AI / ML model aspects and characteristics for Model 2 from those values / choices used by the first configuration 106 for Model 2 and / or the second configuration 120 for Model 2. Continuing on, a th configuration 126 for Model A may be configured for operation when the Xth supporting condition(s) 112 apply. The 'th configuration 126 for Model N may use values / choices for the various AI / ML model aspects and characteristics for Model N from those values / choices used by the first configuration 108 for Model N and / or the second configuration 124 for Model N.
[0028] It is further contemplated that in at least some embodiments, a first configuration for an AI / ML model may substantively use more or fewer of the AI / ML model aspects and characteristics that are used when the (same) AI / ML model is in a second configuration.
[0029] In some cases, examples of supporting conditions and aspects of the AI / ML models may include, for example, the output size and / or the input size used for the AI / ML model. The output size of a particular configuration of a given AI / ML model may be controlled by the number of future measurement instances for beam management prediction, the number of set A beams, and / or the number K of top beams ("top- l" beams) according to the particular current configuration of the AI / ML model in question. The input size of a particular configuration of a given AI / ML model may be controlled by the number of history measurement instances for beam management prediction and / or the number of set B beams (e.g., used in spatial prediction).Matching AI / ML Model Functionalitv / Configuration with Radio Propagation Condition(s)
[0030] A device may be capable of determining that it cannot adequately generate a first inference with an AI / ML model in a first configuration by using input data for the AI / ML model that corresponds to (e.g., is derived from) signaling received on the channel between it and a transmitting device. This inability may be determined based on, for example, a radio condition incompatibility with the first configuration as detected / determined at the device, a deteriorated monitoring metric (e.g.. that is outside an acceptable range) as calculated at / identified to the device while using the AI / ML model according to the first configuration, and / or a determination at the device that it cannot / does not adequately receive signaling corresponding to input data for the AI / ML model (e.g., that is used to derive the input data).
[0031] In some such embodiments, to match an active AI / ML model functionality and / or configuration with active radio propagation conditions, the AI / ML model may be scaled to a second, different configuration for the AI / ML model. For example, in the case of an AI / ML model for beam management, the size of set B beams may be dynamically adjusted such that an optimized tradeoff between the reduction of measurement overhead and system performance may be achieved (e.g., a large size set B beams may achieve more robust performance, while a smaller set B beams may achieve lower overhead, both being useful in certain implementations).
[0032] As a result of this use of scaling, AI / ML model deactivation and / or fallback may take place less often. For example, instead of falling back when a mismatch between channel conditions and a current configuration for the in-use AI / ML model occurs (when it is therefore determined that a device cannot adequately generate a first inference with the AI / ML model by using input data corresponding to (e.g., derived from) signaling received on the channel), embodiments herein may consider a more gradual approach that involves utilizing a different configuration of the AI / ML model to adapt to the current propagation environment effectively (e.g., in terms of the example of FIG. 1, a different AI / ML model configuration for the AI / ML model residing within a different supporting condition grouping that is a better match for the current channel conditions).
[0033] Applicable radio conditions with respect to such operations may include, for example, SNR, Doppler effect, line of sight (LOS) versus non-line of sight (NLOS) (e.g., including delay spread and / or angular spread), and / or the presence of interference. If the radio conditions change and / or the change in radio conditions is detected by the UE or by the network, the AI / ML model may need to be scaled and / or reconfigured accordingly. For example, the AI / ML model may change its functionality to match the change in the radio conditions.
[0034] Additionally, the same principles (scaling / configuring the AI / ML model) may apply to temporal prediction in a beam management use case. For example, a downsampling of the beam management measurements takes place where the skipped measurements are replaced by the temporarily predicted outcomes from the AI / ML model. In static propagation environments (with a low Doppler effect) the downsampling factor may be large (thus skipping more measurements) and, inversely, in very high Doppler effect scenarios the downsampling factor may be small (more measurements may be needed to predict future measurements). As a result, changing configurationaspects for the AI / ML model such as the window of input measurements, and / or the number of future measurements (e.g., based on the Doppler effect condition) may change the AI / ML model functionality to match the radio propagation. The change in configuration aspects of the AI / ML model may correspond to a scaling of the output size and / or the input size of the AI / ML model, as discussed herein). For example, if the Doppler effect is very small and is not rotating, the AI / ML model may be scaled to be relatively smaller such that a waste of extra resources for performing measurements may be avoided.
[0035] FIG. 2 illustrates an example of an AI / ML model 200 that includes preprocessing blocks 202, an AI / ML core 204 and post-processing blocks 206, according to embodiments herein.
[0036] Embodiments herein discuss scaling and / or configuring the AI / ML model according to a change in propagation radio condition(s). Note that the scaling and / or configuring of an AI / ML model is different from the updating of an AI / ML model. For example, updating an AI / ML model may involve the acquisition and use of a completely different AI / ML model that involves model update transferring to the UE from an outside entity, thus causing latency and transfer overhead. Additionally, in such cases of updating the AI / ML model, the new model may need to be verified, whereas in cases of scaling / reconfiguring the AI / ML model, the scaled / reconfigured AI / ML model may not need to be verified again. As a result, scaling the AI / ML model is a simpler and more effective way to match radio conditions.
[0037] A structure for a scalable AI / ML model 200 is now discussed. The AI / ML model 200 consists of pre-processing blocks 202, an AI / ML core 204. and postprocessing blocks 206. Scaling may affect the pre-processing blocks 202 and / or the postprocessing blocks 206, leaving the AI / ML core 204 unmodified whereas updating an AI / ML model 200 affects the AI / ML core 204. Note that, in some embodiments for scaling the AI / ML functionality, reconfiguration of the AI / ML model 200 may entail swapping the pre-processing blocks 202 and / or the post-processing blocks 206 with different pre-processing blocks and / or post-processing blocks that correspond more closely to the changed radio conditions. Accordingly, each / both of the pre-processing blocks 202 and / or the post-processing blocks 206 may be understood as “swappable.”
[0038] It will accordingly be understood that scaling the AI / ML model 200 through the act of sw apping of one or more of the pre-processing blocks 202 and / or the post-processing blocks 206 with new blocks may cause the AI / ML model 200 to operate according to a different configuration for the AI / ML model 200 (see discussion of FIG. 1). For example, an input data size for the AI / ML model 200 may change, input data characteristics for the AI / ML model 200 may change, an inference size for the AI / ML model 200 may change, inference characteristics for the AI / ML model 200 may change, etc.
[0039] The replacement set of pre-processing blocks and / or post-processing blocks may be drawn from, for example, a memory of the device that is operating the AI / ML model 200.
[0040] In some embodiments, assistance information from the network may help scaling / reconfiguration the AI / ML model. In some cases, assistance information from the network may be in terms of the number of set B beams and / or Tx angles that will serve as an input to the AI / ML model, thus improving the performance of the model. In some other cases, Layer 1 (LI) signaling adaptation may be introduced (e.g., to report a larger number of top-A beams used for prediction and for finding the best beam for future transmission) if the SNR conditions change, or if model monitoring indicates performance degradation.
[0041] In some embodiments, a confidence level may be introduced as an output of an AI / ML model. For example, for a beam management case, for low SNR conditions, a beam prediction (e.g.. as corresponding to a beam ID) accuracy may deteriorate. In some cases, if the computed confidence per beam ID decreases, the AI / ML model may be scaled to signal a larger K value (through scaling the output of the AI / ML model), thus improving AI / ML model performance. As a result, further sweeping of top- A' beams for UE measurements may be necessary.
[0042] In some other cases, the AI / ML model for beam management may be scaled to use a greater number of set B beams (e.g., also due to dynamics of the radio propagation, due to some beams in the first set B being blocked, and / or due to an AI / ML model monitoring procedure experiencing degradation with an input of the first set B beams' reference signal received power (RSRP)), based on an indication from the UE or the network. For example, the UE or the network may indicate to change the set B beam pattern where beams 1, 3, 5, and 7 are blocked, thus beams 2, 4, 6, and 8 are the new beams corresponding to the set B beam pattern.
[0043] In yet some other cases, if the computed confidence per beam ID decreases, the AI / ML model for beam management may be scaled from using a narrow beam set B configuration to a wide beam set B configuration as indicated by the UE or by the network.
[0044] FIG. 3 illustrates a flow diagram 300 for the scaling and / or configuring of a UE- sided AI / ML model, according to embodiments herein.
[0045] The flow diagram 300 begins with an initial trigger for data collection 306 transmitted from the UE 304 to the network 302 (also referred to as “step 1”). It should be noted that LCM functionality resides at UE 304. Then, the network 302 transmits 308 the set B beams for use in generating an inference to the UE 304 (also referred to as “step 2”), where the UE 304, in some cases, may attempt to perform and / or actually perform inference based on the transmitted set B beams received from the network 302.
[0046] Accordingly, the UE 304 may predict, using an AI / ML model, the top- ? beam IDs, and, in at least some cases, monitoring metrics of the top-Al beams, such as their associated confidence levels, their associated RSRPs, and / or other monitoring metrics of the top-A' beams. It should be noted that the monitoring metrics may be key performance indicators (KPIs) (that may be measured and / or predicted) that indicate how well the AI / ML model performs / is currently performing. Subsequently the UE 304 may report 310 the predicted top-A? beams and, in at least some cases, any monitoring metrics such as the corresponding confidence levels, RSRPs. etc. to the network 302 (also referred to as “step 3”).
[0047] Then, the current radio conditions and / or a change in the radio conditions is detected 312 at the network 302 or is detected 312 at the UE 304 (e.g., a SNR change, a Doppler change, other radio condition change(s) are detected) (also referred to as “step 4”). Additionally, the model monitoring data is computed at least one of the UE 304 and the network 302 (e.g.. using the monitoring metrics). It is contemplated that the one of the UE 304 and the network 302 that computes the model monitoring data may share it with the other of the UE 304 and the network 302 in at least some cases.
[0048] In some cases, monitoring metrics from the AI / ML model monitoring may include the inputs to the AI / ML model and / or the outputs of the AI / ML model. In other words, model monitoring metrics are not to be limited to only aspects related to outputs of the AI / ML model. For example, an indication that one or more of the set B beams areblocked, and / or that the inputs of the AI / ML model are corrupted / some error has occurred / been detected in the AI / ML model.
[0049] In some examples, the UE 304 may send the LI measurements (as monitoring metrics) of the set B beams to the network 302.
[0050] In some cases, if the network 302 detects that some of a set of measurements at the UE are blocked or are of low quality, the network 302 may change the set B beam pattern or the number of set B beams. Note that the detected change in radio conditions may prevent the AI / ML model from performing an accurate inference or preventing the AI / ML model from performing inference altogether (in the event that scaling is not used).
[0051] Accordingly, based on the detected 312 radio condition change and / or model monitoring degradation, the AI / ML model may be reconfigured / scaled 314 (also referred to as “step 5”).
[0052] For example, for a beam management AI / ML model the number of set B beams and / or set A beams may be reconfigured, the set B beam pattern may be reconfigured, the top-A reporting number may be reconfigured.
[0053] Note that it is expressly contemplated that across the various cases, an AI / ML model functionality may be reconfigured / scaled autonomously at the UE 304 (in cases where the detection of the radio condition / model monitoring condition instead occurs at the UE 304). In such cases, the act of reconfiguring / scaling 314 may not require or use a network instruction.
[0054] Then, the UE 304 requests 316 data, from the network 302, for inference for the newly reconfigured AI / ML functionality (and, in cases of UE autonomous AI / ML model scaling, may also identifying the newly used configuration AI / ML model to the network) (also referred to as “step 6’?). The network 302. in response, transmits 318 data, to the UE 304 for inference corresponding to the new scaling of the AI / ML model (also referred to as “step 7”), thus effectively returning to step 2 of the flow diagram 300 discussed herein, where (new) set B beams are transmitted from the network 302 to the UE 304 for inference. Note that AI / ML model monitoring procedures may be maintained with respect to the AI / ML model after being scaled and / or reconfigured (such that any subsequent / additional scaling needs can be accordingly detected and carried out).
[0055] Corresponding scaling aspects for other type of AI / ML models that analogously use steps similar to those described in relation to FIG. 4 are also discussed. For example.a CSI compression / decompression AI / ML model may use steps analogous to step 1 and steps 3 through step 7. Note that for step 3, monitoring metrics for the CSI compression / decompression case may include the result of a comparison of (raw) CSI feedback at the UE 304 to decoded CSI feedback (e.g., a similarity metric) as calculated at either the network 302 or the UE 304. This result may be used at step 4 (by either the UE 304 or the network 302) to detect 312 the AI / ME model degradation, as indicated. Then, at step 5, the quantization scheme for the encoded CSI feedback may be reconfigured / scaled 314 (either by network instruction or autonomously at the UE, with follow up informational signaling to the network).
[0056] In various AI / ML model contexts, (e.g., beam management, CSI compression / decompression), it may be that the measurement window (i.e., at the input of the model) and / or the number of predictions may be reconfigured.
[0057] Note that particular examples of AI / ML model functionalities provided herein are not exhaustive; various other types of AI / ML models having various other types of functionalities may be reconfigured in view of a radio condition changes and / or monitoring metric degradation, consistent with the principles described herein.
[0058] FIG. 4 illustrates a method 400 of a UE operating an AI / ML model for network signal processing comprising an AI / ML core and first swappable processing blocks interfaced to the AI / ML core according to a first configuration of the AI / ML model, according to embodiments herein. The illustrated method 400 includes determining 402 that the UE cannot adequately generate a first inference by applying first input data corresponding to first network signaling received from a base station to the AI / ML model while the AI / ML model is in the first configuration. The method 400 further includes identifying 404, in a memory of the UE, second swappable processing blocks for a second configuration of the AI / ML model. The method 400 further includes scaling 406 the AI / ML model from the first configuration to the second configuration by swapping the second swappable processing blocks to interface with the AI / ML core in place of the first swappable processing blocks. The method 400 further includes generating 408 a second inference using the AI / ML model by applying second input data corresponding to second network signaling received from the base station to the AI / ML model while the AI / ML model is in the second configuration.
[0059] In some embodiments of the method 400, the first swappable processing blocks comprise first pre-processing blocks for the AI / ML model, and the second swappableprocessing blocks comprise second pre-processing blocks for the AI / ML model. In some such embodiments, the first pre-processing blocks are configured to operate according to a first input data size, and the second pre-processing blocks are configured to operate according to a second input data size that is different than the first data input size. In some other such embodiments, the first pre-processing blocks are configured to operate according to first input data characteristics, and the second pre-processing blocks are configured to operate according to second input data characteristics that are different than the first input data characteristics.
[0060] In some embodiments of the method 400, the first swappable processing blocks comprise first post-processing blocks for post-processing for the first inference, and the second swappable processing blocks comprise second post-processing blocks for postprocessing for the second inference. In some such embodiments, the first post-processing blocks are configured to operate according to a first inference size, and the second postprocessing blocks are configured to operate according to a second inference size that is different than the first inference size. In some other such embodiments, the first postprocessing blocks are configured to operate according to first inference characteristics, and the second post-processing blocks are configured to operate according to second inference characteristics that are different than the first inference characteristics.
[0061] In some embodiments of the method 400, the determining that the UE cannot adequately generate the first inference while the AI / ML model is in the first configuration comprises detecting, based on the first network signaling from the base station, that the first configuration of the AI / ML model is not compatible with a radio condition at the UE.
[0062] In some embodiments of the method 400, the determining that the UE cannot adequately generate the first inference while the AI / ML model is in the first configuration comprises determining that a model monitoring metric for the AI / ML model is outside of an acceptable range.
[0063] In some embodiments of the method 400, the determining that the UE cannot adequately generate the first inference using the first AI / ML model comprises determining that the UE cannot adequately receive the first network signaling corresponding to the first input data.
[0064] In some embodiments of the method 400, the AI / ML model comprises a beam measurement AI / ML model.
[0065] In some embodiments of the method 400, the AI / ML model comprises a CSI compression AI / ML model.
[0066] In some embodiments, the method 400 further comprises sending, to the base station, a request for the second network signaling.
[0067] FIG. 5 illustrates a method 500 of a UE operating an AI / ML model for network signal processing comprising an AI / ML core and first swappable processing blocks interfaced to the AI / ML core according to a first configuration of the AI / ML model, according to embodiments herein. The illustrated method 500 includes receiving 502, from a base station, an instruction to scale the AI / ML model from the first configuration to a second configuration of the AI / ML model. The method 500 further includes identifying 504, in a memory of the UE, second swappable processing blocks for the second configuration. The method 500 further includes scaling 506 the AI / ML model from the first configuration to the second configuration by swapping the second swappable processing blocks to interface with the AI / ML core in place of the first swappable processing blocks. The method 500 further includes generating 508 a first inference using the AI / ML model by applying input data corresponding to first network signaling received from the base station to the AI / ML model while the AI / ML model is in the second configuration.
[0068] In some embodiments of the method 500, the first swappable processing blocks comprise first pre-processing blocks for the AI / ML model, and the second swappable processing blocks comprise second pre-processing blocks for the AI / ML model. In some such embodiments, the first pre-processing blocks are configured to operate according to a first input data size, and the second pre-processing blocks are configured to operate according to a second input data size that is different than the first data input size. In some other such embodiments, the first pre-processing blocks are configured to operate according to first input data characteristics; and the second pre-processing blocks are configured to operate according to second input data characteristics that are different than the first input data characteristics.
[0069] In some embodiments of the method 500, the first swappable processing blocks comprise first post-processing blocks for post-processing for a second inference, and the second swappable processing blocks comprise second post-processing blocks for postprocessing for the first inference. In some such embodiments, the first post-processing blocks are configured to operate according to a first inference size, and the second post-processing blocks are configured to operate according to a second inference size that is different than the first inference size. In some other such embodiments, the first postprocessing blocks are configured to operate according to first inference characteristics, and the second post-processing blocks are configured to operate according to second inference characteristics that are different than the first inference characteristics.
[0070] In some embodiments, the method 500 further comprises detecting that the first configuration of the AI / ML model is not compatible with a radio condition at the UE, and sending, to the base station, an indication that the first configuration of the AI / ML model is not compatible with the radio condition at the UE.
[0071] In some embodiments, the method 500 further comprises determining that a model monitoring metric for the AI / ML model is outside of an acceptable range, and sending, to the base station, an indication that the model monitoring metric for the AI / ML model is outside of the acceptable range.
[0072] In some embodiments, the method 500 further comprises determining that the UE cannot adequately receive second network signaling from the base station for use with the AI / ML model according to the first configuration, and sending, to the base station, an indication that the UE cannot adequately receive the second network signaling.
[0073] In some embodiments, the method 500 further comprises sending, to the base station, a request for the first network signaling.
[0074] FIG. 6 illustrates a method 600 of a base station in communication with a UE operating an AI / ML model for network signal processing comprising an AI / ML core and first swappable processing blocks interfaced to the AI / ML core according to a first configuration of the AI / ML model, according to embodiments herein. The illustrated method 600 includes determining 602 that the UE cannot adequately generate a first inference using first network signaling sent by the base station to the AI / ML model while the AI / ML model is in the first configuration. The method 600 further includes sending 604, to the UE, an instruction to scale the AI / ML model from the first configuration to a second configuration by swapping second swappable processing blocks to interface with the AI / ML core in place of the first swappable processing blocks. The method 600 further includes sending 606, to the UE, second network signaling configured for use with the AI / ML model according to the second configuration.
[0075] In some embodiments of the method 600, the first swappable processing blocks comprise first pre-processing blocks for the AI / ML model, and the second swappable processing blocks comprise second pre-processing blocks for the AI / ML model. In some such embodiments, the first pre-processing blocks are configured to operate according to a first input data size, and the second pre-processing blocks are configured to operate according to a second input data size that is different than the first data input size. In some other such embodiments, the first pre-processing blocks are configured to operate according to first input data characteristics, and the second pre-processing blocks are configured to operate according to second input data characteristics that are different than the first input data characteristics.
[0076] In some embodiments of the method 600, the first swappable processing blocks comprise first post-processing blocks for post-processing for the first inference, and the second swappable processing blocks comprise second post-processing blocks for postprocessing for a second inference. In some such embodiments, the first post-processing blocks are configured to operate according to a first inference size, and the second postprocessing blocks are configured to operate according to a second inference size that is different than the first inference size. In some other such embodiments, the first postprocessing blocks are configured to operate according to first inference characteristics, and the second post-processing blocks are configured to operate according to second inference characteristics that are different than the first inference characteristics.
[0077] In some embodiments, the method 600 further comprises receiving, from the UE, feedback about the first network signaling, and wherein the determining that the UE cannot adequately generate the first inference while the AI / ML model is in the first configuration comprises detecting, based on the feedback from the UE about the first network signaling, that the first configuration of the AI / ML model is not compatible with a radio condition at the UE.
[0078] In some embodiments of the method 600, the determining that the UE cannot adequately generate the first inference while the AI / ML model is in the first configuration comprises determining that a model monitoring metric for the AI / ML model is outside of an acceptable range.
[0079] In some embodiments, the method 600 further comprises receiving, from the UE, an indication that the UE cannot adequately receive the first netw ork signaling from the base station, and wherein the determining that the UE cannot adequately generate thefirst inference while the AI / ML model is in the first configuration is based on the indication.
[0080] In some embodiments of the method 600, the AI / ML model comprises a beam measurement AI / ML model.
[0081] In some embodiments of the method 600, the AI / ML model comprises a CSI compression AI / ML model.
[0082] In some embodiments, the method 600 further comprises receiving, from the UE, a request for the second network signaling, wherein the sending the second network signaling occurs in response to the request.
[0083] FIG. 7 illustrates an example architecture of a wireless communication system 700, according to embodiments disclosed herein. The following description is provided for an example wireless communication system 700 that operates in conjunction with the LTE system standards and / or 5G or NR system standards as provided by 3GPP technical specifications.
[0084] As shown by FIG. 7, the wireless communication system 700 includes UE 702 and UE 704 (although any number of UEs may be used). In this example, the UE 702 and the UE 704 are illustrated as smartphones (e.g., handheld touchscreen mobile computing devices connectable to one or more cellular networks), but may also comprise any mobile or non-mobile computing device configured for wireless communication.
[0085] The UE 702 and UE 704 may be configured to communicatively couple with a RAN 706. In embodiments, the RAN 706 may be NG-RAN, E-UTRAN, etc. The UE 702 and UE 704 utilize connections (or channels) (shown as connection 708 and connection 710, respectively) with the RAN 706, each of which comprises a physical communications interface. The RAN 706 can include one or more base stations (such as base station 712 and base station 714) that enable the connection 708 and connection 710.
[0086] In this example, the connection 708 and connection 710 are air interfaces to enable such communicative coupling, and may be consistent with RAT(s) used by the RAN 706, such as, for example, an LTE and / or NR.
[0087] In some embodiments, the UE 702 and UE 704 may also directly exchange communication data via a sidelink interface 716. The UE 704 is shown to be configured to access an access point (shown as AP 718) via connection 720. By way of example, theconnection 720 can comprise a local wireless connection, such as a connection consistent with any IEEE 802.11 protocol, wherein the AP 718 may comprise a Wi-Fi® router. In this example, the AP 718 may be connected to another network (for example, the Internet) without going through a CN 724.
[0088] In embodiments, the UE 702 and UE 704 can be configured to communicate using orthogonal frequency division multiplexing (OFDM) communication signals with each other or with the base station 712 and / or the base station 714 over a multicarrier communication channel in accordance with various communication techniques, such as, but not limited to, an orthogonal frequency division multiple access (OFDMA) communication technique (e.g., for downlink communications) or a single carrier frequency division multiple access (SC-FDMA) communication technique (e.g., for uplink and ProSe or sidelink communications), although the scope of the embodiments is not limited in this respect. The OFDM signals can comprise a plurality of orthogonal subcarriers.
[0089] In some embodiments, all or parts of the base station 712 or base station 714 may be implemented as one or more software entities running on server computers as part of a virtual network. In addition, or in other embodiments, the base station 712 or base station 714 may be configured to communicate with one another via interface 722. In embodiments where the wireless communication system 700 is an LTE system (e.g., when the CN 724 is an EPC), the interface 722 may be an X2 interface. The X2 interface may be defined between two or more base stations (e.g., two or more eNBs and the like) that connect to an EPC. and / or between two eNBs connecting to the EPC. In embodiments where the wireless communication system 700 is an NR system (e.g., when CN 724 is a 5GC), the interface 722 may be an Xn interface. The Xn interface is defined between two or more base stations (e.g.. two or more gNBs and the like) that connect to 5GC. between a base station 712 (e.g.. a gNB) connecting to 5GC and an eNB. and / or between two eNBs connecting to 5GC (e g., CN 724).
[0090] The RAN 706 is shown to be communicatively coupled to the CN 724. The CN 724 may comprise one or more network elements 726, which are configured to offer various data and telecommunications services to customers / subscribers (e.g., users of UE 702 and UE 704) who are connected to the CN 724 via the RAN 706. The components of the CN 724 may be implemented in one physical device or separate physical devicesincluding components to read and execute instructions from a machine-readable or computer-readable medium (e.g., a non-transitory machine-readable storage medium).
[0091] In embodiments, the CN 724 may be an EPC, and the RAN 706 may be connected with the CN 724 via an S I interface 728. In embodiments, the S I interface 728 may be split into two parts, an SI user plane (Sl-U) interface, which carries traffic data between the base station 712 or base station 714 and a serving gateway (S-GW), and the SI -MME interface, which is a signaling interface between the base station 712 or base station 714 and mobility management entities (MMEs).
[0092] In embodiments, the CN 724 may be a 5GC, and the RAN 706 may be connected with the CN 724 via an NG interface 728. In embodiments, the NG interface 728 may be split into two parts, an NG user plane (NG-U) interface, which carries traffic data between the base station 712 or base station 714 and a user plane function (UPF). and the SI control plane (NG-C) interface, which is a signaling interface between the base station 712 or base station 714 and access and mobility management functions (AMFs).
[0093] Generally, an application server 730 may be an element offering applications that use internet protocol (IP) bearer resources with the CN 724 (e.g., packet switched data services). The application server 730 can also be configured to support one or more communication services (e.g., VoIP sessions, group communication sessions, etc.) for the UE 702 and UE 704 via the CN 724. The application server 730 may communicate with the CN 724 through an IP communications interface 732.
[0094] FIG. 8 illustrates a system 800 for performing signaling 834 between a wireless device 802 and a network device 818, according to embodiments disclosed herein. The system 800 may be a portion of a wireless communications system as herein described. The wireless device 802 may be, for example, a UE of a wireless communication system. The network device 818 may be, for example, a base station (e.g., an eNB or a gNB) of a wireless communication system.
[0095] The wireless device 802 may include one or more processor(s) 804. The processor(s) 804 may execute instructions such that various operations of the wireless device 802 are performed, as described herein. The processor(s) 804 may include one or more baseband processors implemented using, for example, a central processing unit (CPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a controller, a field programmable gate array (FPGA) device, another hardwaredevice, a firmware device, or any combination thereof configured to perform the operations described herein.
[0096] The wireless device 802 may include a memory 806. The memory7806 may be a non-transitory computer-readable storage medium that stores instructions 808 (which may include, for example, the instructions being executed by the processor(s) 804). The instructions 808 may also be referred to as program code or a computer program. The memory 806 may also store data used by, and results computed by, the processor(s) 804.
[0097] The wireless device 802 may include one or more transceiver(s) 810 that may include radio frequency (RF) transmitter circuitry and / or receiver circuitry that use the antenna(s) 812 of the wireless device 802 to facilitate signaling (e.g., the signaling 834) to and / or from the wireless device 802 with other devices (e.g., the network device 818) according to corresponding RATs.
[0098] The wireless device 802 may include one or more antenna(s) 812 (e.g., one, two, four, or more). For embodiments with multiple antenna(s) 812, the wireless device 802 may leverage the spatial diversity of such multiple antenna(s) 812 to send and / or receive multiple different data streams on the same time and frequency resources. This behavior may be referred to as, for example, multiple input multiple output (MIMO) behavior (referring to the multiple antennas used at each of a transmitting device and a receiving device that enable this aspect). MIMO transmissions by the wireless device 802 may be accomplished according to precoding (or digital beamforming) that is applied at the wireless device 802 that multiplexes the data streams across the antenna(s) 812 according to known or assumed channel characteristics such that each data stream is received with an appropriate signal strength relative to other streams and at a desired location in the spatial domain (e.g., the location of a receiver associated with that data stream). Certain embodiments may use single user MIMO (SU-MIMO) methods (where the data streams are all directed to a single receiver) and / or multi user MIMO (MU- MIMO) methods (where individual data streams may be directed to individual (different) receivers in different locations in the spatial domain).
[0099] In certain embodiments having multiple antennas, the wireless device 802 may implement analog beamforming techniques, whereby phases of the signals sent by the antenna(s) 812 are relatively adjusted such that the (joint) transmission of the antenna(s) 812 can be directed (this is sometimes referred to as beam steering).
[0100] The wireless device 802 may include one or more interface(s) 814. The interface(s) 814 may be used to provide input to or output from the wireless device 802. For example, a wireless device 802 that is a UE may include interface(s) 814 such as microphones, speakers, a touchscreen, buttons, and the like in order to allow for input and / or output to the UE by a user of the UE. Other interfaces of such a UE may be made up of transmitters, receivers, and other circuitry (e.g., other than the transceiver(s) 810 / antenna(s) 812 already described) that allow for communication between the UE and other devices and may operate according to known protocols (e.g., Wi-Fi®, Bluetooth®, and the like).
[0101] The wireless device 802 may include an AI / ML model scaling module 816. The AI / ML model scaling module 816 may be implemented via hardware, software, or combinations thereof. For example, the AI / ML model scaling module 816 may be implemented as a processor, circuit, and / or instructions 808 stored in the memory' 806 and executed by the processor(s) 804. In some examples, the AI / ML model scaling module 816 may be integrated within the processor(s) 804 and / or the transceiver(s) 810. For example, the AI / ML model scaling module 816 may be implemented by a combination of software components (e.g., executed by a DSP or a general processor) and hardware components (e.g., logic gates and circuitry) within the processor(s) 804 or the transceiver(s) 810.
[0102] The AI / ML model scaling module 816 may be used for various aspects of the present disclosure, for example, aspects of FIG. 1, FIG. 2, FIG. 3. FIG. 4, FIG. 5, FIG. 7, and FIG. 8. The AI / ML model scaling module 816 is configured to cause the wireless device 802 to determine that the wireless device 802 cannot adequately generate a first inference by applying first input data corresponding to first network signaling received from a base station to the AI / ML model while the AI / ML model is in the first configuration. The AI / ML model scaling module 816 is further configured to cause the wireless device 802 identify, in a memory of the UE, second swappable processing blocks for a second configuration of the AI / ML model and scale the AI / ML model from the first configuration to the second configuration by swapping the second swappable processing blocks to interface with the AI / ML core in place of the first swappable processing blocks. The AI / ML model scaling module 816 is further configured to cause the wireless device 802 to generate a second inference using the AI / ML model by applying second input data corresponding to second network signaling received from thebase station to the AI / ML model while the AI / ML model is in the second configuration. In some cases, the AI / ML model scaling module 816 is configured to cause the wireless device 802 to receive, from the base station, instruction to scale the AI / ML model from the first configuration to the second configuration. Additionally, in some cases, the AI / ML model scaling module 832 is configured to cause the wireless device 802 to determine that the UE cannot adequately generate the first inference by detecting, based on network signaling from the base station, that the first configuration of the AI / ML model is not compatible with a radio condition at the UE.
[0103] The network device 818 may include one or more processor(s) 820. The processor(s) 820 may execute instructions such that various operations of the network device 818 are performed, as described herein. The processor(s) 820 may include one or more baseband processors implemented using, for example, a CPU, a DSP, an ASIC, a controller, an FPGA device, another hardware device, a firmware device, or any combination thereof configured to perform the operations described herein.
[0104] The network device 818 may include a memory 822. The memory 822 may be a non-transitory computer-readable storage medium that stores instructions 824 (which may include, for example, the instructions being executed by the processor(s) 820). The instructions 824 may also be referred to as program code or a computer program. The memory 822 may also store data used by, and results computed by, the processor(s) 820.
[0105] The network device 818 may include one or more transceiver(s) 826 that may include RF transmitter circuitry and / or receiver circuitry that use the antenna(s) 828 of the network device 818 to facilitate signaling (e.g., the signaling 834) to and / or from the network device 818 with other devices (e.g., the wireless device 802) according to corresponding RATs.
[0106] The network device 818 may include one or more antenna(s) 828 (e.g., one, two, four, or more). In embodiments having multiple antenna(s) 828, the network device 818 may perform MIMO, digital beamforming, analog beamforming, beam steering, etc., as has been described.
[0107] The network device 818 may include one or more interface(s) 830. The interface(s) 830 may be used to provide input to or output from the network device 818. For example, a network device 818 that is a base station may include interface(s) 830 made up of transmitters, receivers, and other circuitry (e.g., other than the transceiver(s) 826 / antenna(s) 828 already described) that enables the base station to communicate withother equipment in a core network, and / or that enables the base station to communicate with external networks, computers, databases, and the like for purposes of operations, administration, and maintenance of the base station or other equipment operably connected thereto.
[0108] The network device 818 may include an AI / ML model scaling module 832. The AI / ML model scaling module 832 may be implemented via hardware, software, or combinations thereof. For example, the AI / ML model scaling module 832 may be implemented as a processor, circuit, and / or instructions 824 stored in the memory 822 and executed by the processor(s) 820. In some examples, the AI / ML model scaling module 832 may be integrated within the processor(s) 820 and / or the transceiver(s) 826. For example, the AI / ML model scaling module 832 may be implemented by a combination of software components (e.g., executed by a DSP or a general processor) and hardware components (e.g., logic gates and circuitry) within the processor(s) 820 or the transceiver(s) 826.
[0109] The AI / ML model scaling module 832 may be used for various aspects of the present disclosure, for example, aspects of FIG. 1, FIG. 2, FIG. 3, FIG. 6, FIG. 7, and FIG. 8. The AI / ML model scaling module 832 is configured to cause the network device 818 to determine that a UE cannot adequately generate a first inference using first network signaling sent by the network device 818 to the AI / ML model while the AI / ML model is in the first configuration. The AI / ML model scaling module 832 is further configured to cause the network device 818 to send, to the UE, an instruction to scale the AI / ML model from the first configuration to a second configuration by swapping second swappable processing blocks to interface with the AI / ML core in place of the first swappable processing blocks, and send, to the UE, second network signaling configured for use with the AI / ML model according to the second configuration. In some cases, the AI / ML model scaling module 832 is configured to cause the network device 818 to determine that the UE cannot adequately generate the first inference by detecting, based on feedback from the UE, that the first configuration of the AI / ML model is not compatible with a radio condition at the UE.
[0110] Embodiments contemplated herein include an apparatus comprising means to perform one or more elements of the method 400 and the method 500. This apparatus may be, for example, an apparatus of a UE (such as a wireless device 802 that is a UE, as described herein).[OHl] Embodiments contemplated herein include one or more non-transitory computer-readable media comprising instructions to cause an electronic device, upon execution of the instructions by one or more processors of the electronic device, to perform one or more elements of the method 400 and the method 500. This non- transitory' computer-readable media may be, for example, a memory of a UE (such as a memory 806 of a wireless device 802 that is a UE, as described herein).
[0112] Embodiments contemplated herein include an apparatus comprising logic, modules, or circuitry' to perform one or more elements of the method 400 and the method 500. This apparatus may be, for example, an apparatus of a UE (such as a wireless device 802 that is a UE, as described herein).
[0113] Embodiments contemplated herein include an apparatus comprising: one or more processors and one or more computer-readable media comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform one or more elements of the method 400 and the method 500. This apparatus may be, for example, an apparatus of a UE (such as a wireless device 802 that is a UE, as described herein).
[0114] Embodiments contemplated herein include a signal as described in or related to one or more elements of the method 400 and the method 500.
[0115] Embodiments contemplated herein include a computer program or computer program product comprising instructions, wherein execution of the program by a processor is to cause the processor to carry out one or more elements of the method 400 and the method 500. The processor may be a processor of a UE (such as a processor(s) 804 of a wireless device 802 that is a UE, as described herein). These instructions may be, for example, located in the processor and / or on a memory of the UE (such as a memory 806 of a wireless device 802 that is a UE, as described herein).
[0116] Embodiments contemplated herein include an apparatus comprising means to perform one or more elements of the method 600. This apparatus may be, for example, an apparatus of a base station (such as a network device 818 that is a base station, as described herein).
[0117] Embodiments contemplated herein include one or more non-transitory computer-readable media comprising instructions to cause an electronic device, upon execution of the instructions by one or more processors of the electronic device, to perform one or more elements of the method 600. This non-transitory computer-readablemedia may be. for example, a memory of a base station (such as a memory 822 of a network device 818 that is a base station, as described herein).
[0118] Embodiments contemplated herein include an apparatus comprising logic, modules, or circuitry to perform one or more elements of the method 600. This apparatus may be, for example, an apparatus of a base station (such as a network device 818 that is a base station, as described herein).
[0119] Embodiments contemplated herein include an apparatus comprising: one or more processors and one or more computer-readable media comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform one or more elements of the method 600. This apparatus may be, for example, an apparatus of a base station (such as a network device 818 that is a base station, as described herein).
[0120] Embodiments contemplated herein include a signal as described in or related to one or more elements of the method 600.
[0121] Embodiments contemplated herein include a computer program or computer program product comprising instructions, wherein execution of the program by a processing element is to cause the processing element to carry out one or more elements of the method 600. The processor may be a processor of a base station (such as a processor(s) 820 of a network device 818 that is a base station, as described herein). These instructions may be, for example, located in the processor and / or on a memory of the base station (such as a memory' 822 of a network device 818 that is a base station, as described herein).
[0122] For one or more embodiments, at least one of the components set forth in one or more of the preceding figures may be configured to perform one or more operations, techniques, processes, and / or methods as set forth herein. For example, a baseband processor as described herein in connection with one or more of the preceding figures may be configured to operate in accordance with one or more of the examples set forth herein. For another example, circuitry associated with a UE, base station, network element, etc. as described above in connection with one or more of the preceding figures may be configured to operate in accordance with one or more of the examples set forth herein.
[0123] Any of the above described embodiments may be combined with any other embodiment (or combination of embodiments), unless explicitly stated otherwise. Theforegoing description of one or more implementations provides illustration and description, but is not intended to be exhaustive or to limit the scope of embodiments to the precise form disclosed. Modifications and variations are possible in light of the above teachings or may be acquired from practice of various embodiments.
[0124] Embodiments and implementations of the systems and methods described herein may include various operations, which may be embodied in machine-executable instructions to be executed by a computer system. A computer system may include one or more general-purpose or special-purpose computers (or other electronic devices). The computer system may include hardware components that include specific logic for performing the operations or may include a combination of hardware, software, and / or firmware.
[0125] It should be recognized that the systems described herein include descriptions of specific embodiments. These embodiments can be combined into single systems, partially combined into other systems, split into multiple systems or divided or combined in other ways. In addition, it is contemplated that parameters, attributes, aspects, etc. of one embodiment can be used in another embodiment. The parameters, attributes, aspects, etc. are merely described in one or more embodiments for clarity, and it is recognized that the parameters, attributes, aspects, etc. can be combined with or substituted for parameters, attributes, aspects, etc. of another embodiment unless specifically disclaimed herein.
[0126] It is well understood that the use of personally identifiable information should follow privacy policies and practices that are generally recognized as meeting or exceeding industry or governmental requirements for maintaining the privacy of users. In particular, personally identifiable information data should be managed and handled so as to minimize risks of unintentional or unauthorized access or use, and the nature of authorized use should be clearly indicated to users.
[0127] Although the foregoing has been described in some detail for purposes of clarity, it will be apparent that certain changes and modifications may be made without departing from the principles thereof. It should be noted that there are many alternative ways of implementing both the processes and apparatuses described herein. Accordingly, the present embodiments are to be considered illustrative and not restrictive, and the description is not to be limited to the details given herein, but may be modified within the scope and equivalents of the appended claims.
Claims
CLAIMS1. A method of a user equipment (UE) operating an artificial intelligence (AI) / machine learning (ML) model for network signal processing comprising an AI / ML core and first swappable processing blocks interfaced to the AI / ML core according to a first configuration of the AI / ML model, the method comprising: determining that the UE cannot adequately generate a first inference by applying first input data corresponding to first network signaling received from a base station to the AI / ML model while the AI / ML model is in the first configuration; identifying, in a memory of the UE, second swappable processing blocks for a second configuration of the AI / ML model; scaling the AI / ML model from the first configuration to the second configuration by swapping the second swappable processing blocks to interface with the AI / ML core in place of the first swappable processing blocks; and generating a second inference using the AI / ML model by applying second input data corresponding to second network signaling received from the base station to the AI / ML model while the AI / ML model is in the second configuration.
2. The method of claim 1, wherein: the first swappable processing blocks comprise first pre-processing blocks for the AI / ML model; and the second swappable processing blocks comprise second pre-processing blocks for the AI / ML model.
3. The method of claim 2, wherein: the first pre-processing blocks are configured to operate according to a first input data size; and the second pre-processing blocks are configured to operate according to a second input data size that is different than the first data input size.
4. The method of claim 2, wherein: the first pre-processing blocks are configured to operate according to first input data characteristics; and the second pre-processing blocks are configured to operate according to second input data characteristics that are different than the first input data characteristics.
5. The method of claim 1, wherein: the first swappable processing blocks comprise first post-processing blocks for post-processing for the first inference; and the second swappable processing blocks comprise second post-processing blocks for post-processing for the second inference.
6. The method of claim 5, wherein: the first post-processing blocks are configured to operate according to a first inference size; and the second post-processing blocks are configured to operate according to a second inference size that is different than the first inference size.
7. The method of claim 5, wherein: the first post-processing blocks are configured to operate according to first inference characteristics; and the second post-processing blocks are configured to operate according to second inference characteristics that are different than the first inference characteristics.
8. The method of claim 1, wherein the determining that the UE cannot adequately generate the first inference while the AI / ML model is in the first configuration comprises detecting, based on the first network signaling from the base station, that the first configuration of the AI / ML model is not compatible with a radio condition at the UE.
9. The method of claim 1, wherein the determining that the UE cannot adequately generate the first inference while the AI / ML model is in the first configuration comprises determining that a model monitoring metric for the AI / ML model is outside of an acceptable range.
10. The method of claim 1, wherein the determining that the UE cannot adequately generate the first inference using the first AI / ML model comprises determining that the UE cannot adequately receive the first network signaling corresponding to the first input data.
11. The method of claim 1, wherein the AI / ML model comprises a beam measurement AI / ML model.
12. The method of claim 1, wherein the AI / ML model comprises a channel state information (CSI) compression AI / ML model.
13. The method of claim 1, further comprising sending, to the base station, a request for the second network signaling.
14. A method of a user equipment (UE) operating an artificial intelligence (Al)Zmachine learning (ML) model for network signal processing comprising an AI / ML core and first swappable processing blocks interfaced to the AI / ML core according to a first configuration of the AI / ML model, the method comprising: receiving, from a base station, an instruction to scale the AI / ML model from the first configuration to a second configuration of the AI / ML model; identifying, in a memory of the UE, second swappable processing blocks for the second configuration; scaling the AI / ML model from the first configuration to the second configuration by swapping the second swappable processing blocks to interface with the AI / ML core in place of the first swappable processing blocks; and generating a first inference using the AI / ML model by applying input data corresponding to first network signaling received from the base station to the AI / ML model while the AI / ML model is in the second configuration.
15. The method of claim 14, wherein: the first swappable processing blocks comprise first pre-processing blocks for the AI / ML model; and the second swappable processing blocks comprise second pre-processing blocks for the AI / ML model.
16. The method of claim 15, wherein: the first pre-processing blocks are configured to operate according to a first input data size; and the second pre-processing blocks are configured to operate according to a second input data size that is different than the first data input size.
17. The method of claim 15, wherein: the first pre-processing blocks are configured to operate according to first input data characteristics; andthe second pre-processing blocks are configured to operate according to second input data characteristics that are different than the first input data characteristics.
18. The method of claim 14, wherein: the first swappable processing blocks comprise first post-processing blocks for post-processing for a second inference; and the second swappable processing blocks comprise second post-processing blocks for post-processing for the first inference.
19. The method of claim 18, wherein: the first post-processing blocks are configured to operate according to a first inference size; and the second post-processing blocks are configured to operate according to a second inference size that is different than the first inference size.
20. The method of claim 18, wherein: the first post-processing blocks are configured to operate according to first inference characteristics; and the second post-processing blocks are configured to operate according to second inference characteristics that are different than the first inference characteristics.
21. The method of claim 14, further comprising: detecting that the first configuration of the AI / ML model is not compatible with a radio condition at the UE; and sending, to the base station, an indication that the first configuration of the AI / ML model is not compatible with the radio condition at the UE.
22. The method of claim 14, further comprising: determining that a model monitoring metric for the AI / ML model is outside of an acceptable range; and sending, to the base station, an indication that the model monitoring metric for the AI / ML model is outside of the acceptable range.
23. The method of claim 14, further comprising: determining that the UE cannot adequately receive second network signaling from the base station for use with the AI / ML model according to the first configuration; andsending, to the base station, an indication that the UE cannot adequately receive the second network signaling.
24. The method of claim 14, further comprising sending, to the base station, a request for the first network signaling.
25. A method of a base station in communication with a user equipment (UE) operating an artificial intelligence (AI) / machine learning (ML) model for network signal processing comprising an AI / ML core and first swappable processing blocks interfaced to the AI / ML core according to a first configuration of the AI / ML model, the method comprising: determining that the UE cannot adequately generate a first inference using first network signaling sent by the base station to the AI / ML model while the AI / ML model is in the first configuration; sending, to the UE, an instruction to scale the AI / ML model from the first configuration to a second configuration by swapping second swappable processing blocks to interface with the AI / ML core in place of the first swappable processing blocks: and sending, to the UE, second network signaling configured for use with the AI / ML model according to the second configuration.
26. The method of claim 25, wherein: the first swappable processing blocks comprise first pre-processing blocks for the AI / ML model; and the second swappable processing blocks comprise second pre-processing blocks for the AI / ML model.
27. The method of claim 26, wherein: the first pre-processing blocks are configured to operate according to a first input data size; and the second pre-processing blocks are configured to operate according to a second input data size that is different than the first data input size.
28. The method of claim 26, wherein: the first pre-processing blocks are configured to operate according to first input data characteristics; andthe second pre-processing blocks are configured to operate according to second input data characteristics that are different than the first input data characteristics.
29. The method of claim 25, wherein: the first swappable processing blocks comprise first post-processing blocks for post-processing for the first inference; and the second swappable processing blocks comprise second post-processing blocks for post-processing for a second inference.
30. The method of claim 29, wherein: the first post-processing blocks are configured to operate according to a first inference size; and the second post-processing blocks are configured to operate according to a second inference size that is different than the first inference size.
31. The method of claim 29, wherein: the first post-processing blocks are configured to operate according to first inference characteristics; and the second post-processing blocks are configured to operate according to second inference characteristics that are different than the first inference characteristics.
32. The method of claim 25, further comprising receiving, from the UE, feedback about the first network signaling, and wherein the determining that the UE cannot adequately generate the first inference while the AI / ML model is in the first configuration comprises detecting, based on the feedback from the UE about the first network signaling, that the first configuration of the AI / ML model is not compatible with a radio condition at the UE.
33. The method of claim 25, wherein the determining that the UE cannot adequately generate the first inference while the AI / ML model is in the first configuration comprises determining that a model monitoring metric for the AI / ML model is outside of an acceptable range.
34. The method of claim 25, further comprising receiving, from the UE, an indication that the UE cannot adequately receive the first network signaling from the base station,and wherein the determining that the UE cannot adequately generate the first inference while the AI / ML model is in the first configuration is based on the indication.
35. The method of claim 25, wherein the AI / ML model comprises a beam measurement AI / ML model.
36. The method of claim 25, wherein the AI / ML model comprises a channel state information (CSI) compression AI / ML model.
37. The method of claim 25, further comprising receiving, from the UE, a request for the second network signaling, wherein the sending the second network signaling occurs in response to the request.
38. An apparatus comprising means to perform the method of any of claim 1 to claim 37.
39. A computer-readable media comprising instructions to cause an electronic device, upon execution of the instructions by one or more processors of the electronic device, to perform the method of any of claim 1 to claim 37.
40. An apparatus comprising logic, modules, or circuitry to perform the method of any of claim 1 to claim 37.
41. A baseband processor for a user equipment (UE) that is configured to cause the UE to perform one or more elements of any one of claim 1 to claim 24.
42. A baseband processor for a base station that is configured to cause the base station to perform one or more elements of any one of claim 25 to claim 37.
Citation Information
Patent Citations
Method and apparatus for implementing ai-ML in a wireless network
WO2024039898A1