Delivery of ai / ML model parameters
By negotiating accuracy levels and using lookup tables, the mechanism addresses the challenge of delivering AI/ML model parameters over error-prone wireless channels, improving model performance and efficiency.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NOKIA TECHNOLOGIES OY
- Filing Date
- 2025-10-17
- Publication Date
- 2026-05-15
AI Technical Summary
The challenge of efficiently delivering AI/ML model parameters over error-prone wireless channels in 5G networks, particularly for localized models, is exacerbated by high sensitivity and volume, leading to potential distortions and performance degradation.
A mechanism for negotiating accuracy levels between base stations and UEs to establish optimal AI/ML model parameter representations using lookup tables, ensuring precise encoding and decoding based on fluctuating wireless channel conditions.
Minimizes errors during AI/ML model parameter transmission, enhancing model performance and efficiency by adapting to varying channel conditions and computational capabilities.
Smart Images

Figure IB2025060622_15052026_PF_FP_ABST
Abstract
Description
DELIVERY OF AI / ML MODEL PARAMETERSCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims priority from, and the benefit of, US Provisional Application No. 63 / 717699, filed November 7, 2024, which is hereby incorporated by reference in its entirety.FIELD
[0002] Various example embodiments of the present disclosure generally relate to the field of telecommunication and in particular, to methods, devices, apparatuses and computer readable storage medium for delivering AI / ML model parameters.BACKGROUND
[0003] 5G NR (New Radio) is the next-generation radio access technology developed by the 3rd Generation Partnership Project (3GPP) for the 5G mobile network. It is designed to be the global standard for the air interface of 5G networks, offering lower latency, higher system capacity, and massive device connectivity. Artificial intelligence (Al) and machine learning (ML) techniques are also introduced to the standard in which AI / ML is used in the field of channel state information (CSI) compression / decompression, CSI prediction, beam management and positioning. One potential approach for AI / ML model training is to perform it at base station or at 3rd party server based on massive data sharing. After the training, the model or at least the model parameters (weights) are then delivered to user equipment.SUMMARY
[0004] In a first aspect of the present disclosure, there is provided a first apparatus. The first apparatus comprises at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the first apparatus at least to: obtain, from a second apparatus, accuracy level information indicating parameter quantization for an artificial intelligence and machine learning (AI / ML) model, the accuracy level information indicating at least one accuracy level; obtain, from the second apparatus, indication information of at least one set of parameters of the AI / ML model corresponding to the at least one accuracy level; and generate, the at least one set of parameters based on the indication information and the at least one accuracy level.
[0005] In a second aspect of the present disclosure, there is provided a second apparatus. The second apparatus comprises at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the second apparatus at least to: transmit, to a first apparatus, accuracy level information indicating parameter quantization for an artificial intelligence and machine learning (AI / ML) model, the accuracy level information indicating at leastone accuracy level; and transmit, to the first apparatus, indication information of at least one set of parameters of the AI / ML model corresponding to the at least one accuracy level.
[0006] In a third aspect of the present disclosure, there is provided a method. The method comprises: obtaining, from a second apparatus, accuracy level information indicating parameter quantization for an artificial intelligence and machine learning (AI / ML) model, the accuracy level information indicating at least one accuracy level; obtaining, from the second apparatus, indication information of at least one set of parameters of the AI / ML model corresponding to the at least one accuracy level; and generating, the at least one set of parameters based on the indication information and the at least one accuracy level.
[0007] In a fourth aspect of the present disclosure, there is provided a method. The method comprises: transmitting, to a first apparatus, accuracy level information indicating parameter quantization for an artificial intelligence and machine learning (AI / ML) model, the accuracy level information indicating at least one accuracy level; and transmitting, to the first apparatus, indication information of at least one set of parameters of the AI / ML model corresponding to the at least one accuracy level.
[0008] In a fifth aspect of the present disclosure, there is provided a first apparatus. The first apparatus comprises means for obtaining, from a second apparatus, accuracy level information indicating parameter quantization for an artificial intelligence and machine learning (AI / ML) model, the accuracy level information indicating at least one accuracy level; means for obtaining, from the second apparatus, indication information of at least one set of parameters of the AI / ML model corresponding to the at least one accuracy level; and means for generating, the at least one set of parameters based on the indication information and the at least one accuracy level.
[0009] In a sixth aspect of the present disclosure, there is provided a second apparatus. The second apparatus comprises means for transmitting, to a first apparatus, accuracy level information indicating parameter quantization for an artificial intelligence and machine learning (AI / ML) model, the accuracy level information indicating at least one accuracy level; and means for transmitting, to the first apparatus, indication information of at least one set of parameters of the AI / ML model corresponding to the at least one accuracy level.
[0010] In a seventh aspect of the present disclosure, there is provided a computer readable medium. The computer readable medium comprises instructions stored thereon for causing an apparatus to perform at least the method according to the third aspect.
[0011] In an eighth aspect of the present disclosure, there is provided a computer readable medium. The computer readable medium comprises instructions stored thereon for causing an apparatus to perform at least the method according to the fourth aspect.
[0012] It is to be understood that the Summary section is not intended to identify key or essentialfeatures of embodiments of the present disclosure, nor is it intended to be used to limit the scope of the present disclosure. Other features of the present disclosure will become easily comprehensible through the following description.BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Some example embodiments will now be described with reference to the accompanying drawings, where:
[0014] FIG. 1 illustrates an example communication environment in which example embodiments of the present disclosure can be implemented;
[0015] FIG. 2 an example of AI / ML model parameters;
[0016] FIG. 3 illustrates an example of accuracy level agreement thresholds between the UE and the network;
[0017] FIGs. 4A-4C illustrate some examples of lookup tables in accordance with some embodiments in the disclosure;
[0018] FIGs. 5A-5B illustrate a diagram of an example message sequence in accordance with some embodiments;
[0019] FIG. 6 illustrates a flowchart of a method implemented at a first apparatus in accordance with some example embodiments of the present disclosure;
[0020] FIG. 7 illustrates a flowchart of a method implemented at a second apparatus in accordance with some example embodiments of the present disclosure;
[0021] FIG. 8 illustrates a simplified block diagram of a device that is suitable for implementing example embodiments of the present disclosure; and
[0022] FIG. 9 illustrates a block diagram of an example computer readable medium in accordance with some example embodiments of the present disclosure.
[0023] Throughout the drawings, the same or similar reference numerals represent the same or similar element.DETAILED DESCRIPTION
[0024] Principle of the present disclosure will now be described with reference to some example embodiments. It is to be understood that these embodiments are described only for the purpose of illustration and help those skilled in the art to understand and implement the present disclosure, without suggesting any limitation as to the scope of the disclosure. Embodiments described herein can be implemented in various manners other than the ones described below.
[0025] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.
[0026] References in the present disclosure to “one embodiment,” “an embodiment,” “an example embodiment,” and the like indicate that the embodiment described may include a particular feature, structure, or characteristic, but it is not necessary that every embodiment includes the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is submitted that it is within the knowledge of one skilled in the art to affect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.
[0027] It shall be understood that although the terms “first,” “second,”..., etc. in front of noun(s) and the like may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another and they do not limit the order of the noun(s). For example, a first element could be termed a second element, and similarly, a second element could be termed a first element, without departing from the scope of example embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.
[0028] As used herein, “at least one of the following: ” and “at least one of ” and similar wording, where the list of two or more elements are joined by “and” or “or”, mean at least any one of the elements, or at least any two or more of the elements, or at least all the elements.
[0029] As used herein, unless stated explicitly, performing a step “in response to A” does not indicate that the step is performed immediately after “A” occurs and one or more intervening steps may be included.
[0030] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments. As used herein, the singular forms “a”, “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises”, “comprising”, “has”, “having”, “includes” and / or “including”, when used herein, specify the presence of stated features, elements, and / or components etc., but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof.
[0031] As used in this application, the term “circuitry” may refer to one or more or all of the following:(a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry) and(b) combinations of hardware circuits and software, such as (as applicable):(i) a combination of analog and / or digital hardware ci rcuit(s) with software / firmware and(ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions) and(c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.
[0032] This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.
[0033] As used herein, the term “communication network” refers to a network following any suitable communication standards, such as New Radio (NR), Long Term Evolution (LTE), LTE-Advanced (LTE- A), Wideband Code Division Multiple Access (WCDMA), High-Speed Packet Access (HSPA), Narrow Band Internet of Things (NB-loT) and so on. Furthermore, the communications between a terminal device and a network device in the communication network may be performed according to any suitable generation communication protocols, including, but not limited to, the first generation (1 G), the second generation (2G), 2.5G, 2.75G, the third generation (3G), the fourth generation (4G), 4.5G, the fifth generation (5G), 5.5G, the sixth generation (6G) communication protocols, and / or any other protocols either currently known or to be developed in the future. Embodiments of the present disclosure may be applied in various communication systems. Given the rapid development in communications, there will of course also be future type communication technologies and systems with which the present disclosure may be embodied. It should not be seen as limiting the scope of the present disclosure to only the aforementioned system.
[0034] As used herein, the term “network device” refers to a node in a communication network via which a terminal device accesses the network and receives services therefrom. The network device may refer to a base station (BS) or an access point (AP), for example, a node B (NodeB or NB), an evolved NodeB (eNodeB or eNB), an NR NB (also referred to as a gNB), a Remote Radio Unit (RRU), a radio header (RH), a remote radio head (RRH), a relay, an Integrated Access and Backhaul (I AB) node, a low power node such as a femto, a pico, a non-terrestrial network (NTN) or non-ground network device such as a satellite network device, a low earth orbit (LEO) satellite and a geosynchronous earth orbit (GEO) satellite, an aircraft network device, and so forth, depending on the applied terminology and technology. In some example embodiments, radio access network (RAN) splitarchitecture comprises a Centralized Unit (CU) and a Distributed Unit (DU) at an IAB donor node. An IAB node comprises a Mobile Terminal (IAB-MT) part that behaves like a UE toward the parent node, and a DU part of an IAB node behaves like a base station toward the next-hop IAB node.
[0035] The term “terminal device” refers to any end device that may be capable of wireless communication. By way of example rather than limitation, a terminal device may also be referred to as a communication device, user equipment (UE), a Subscriber Station (SS), a Portable Subscriber Station, a Mobile Station (MS), or an Access Terminal (AT). The terminal device may include, but not limited to, a mobile phone, a cellular phone, a smart phone, voice over IP (VoIP) phones, wireless local loop phones, a tablet, a wearable terminal device, a personal digital assistant (PDA), portable computers, desktop computer, image capture terminal devices such as digital cameras, gaming terminal devices, music storage and playback appliances, vehicle-mounted wireless terminal devices, wireless endpoints, mobile stations, laptop-embedded equipment (LEE), laptop-mounted equipment (LME), USB dongles, smart devices, wireless customer-premises equipment (CPE), an Internet of Things (loT) device, a watch or other wearable, a head-mounted display (HMD), a vehicle, a drone, a medical device and applications (e.g., remote surgery), an industrial device and applications (e.g., a robot and / or other wireless devices operating in an industrial and / or an automated processing chain contexts), a consumer electronics device, a device operating on commercial and / or industrial wireless networks, and the like. The terminal device may also correspond to a Mobile Termination (MT) part of an IAB node (e.g., a relay node). In the following description, the terms “terminal device”, “communication device”, “terminal”, “user equipment” and “UE” may be used interchangeably.
[0036] As used herein, the term “resource,” “transmission resource,” “resource block,” “physical resource block” (PRB), “uplink resource,” or “downlink resource” may refer to any resource for performing a communication, for example, a communication between a terminal device and a network device, such as a resource in time domain, a resource in frequency domain, a resource in space domain, a resource in code domain, or any other combination of the time, frequency, space and / or code domain resource enabling a communication, and the like. In the following, unless explicitly stated, a resource in both frequency domain and time domain will be used as an example of a transmission resource for describing some example embodiments of the present disclosure. It is noted that example embodiments of the present disclosure are equally applicable to other resources in other domains.
[0037] To facilitate understanding of the terminologies, some definitions of the list of terminologies used for AI / ML are provided below.
[0038] AI / ML Model: A data driven algorithm that applies AI / ML techniques to generate a set of outputs based on a set of inputs.
[0039] AI / ML model parameters: Values within an AI / ML model that are learned during the model training process, representing the model’s adjustment variables to improve performance on a giventask. These parameters are crucial to the functionality of the model and directly influence its predictions and overall accuracy.
[0040] AI / ML model weights: Specific types of AI / ML model parameters used in neural networks, often referring to the connections between nodes in network layers. These weights are adjusted through the learning process to optimize the model’s ability to map input data to the desired output, thereby enhancing the model’s predictive accuracy.
[0041] AI / ML model delivery: A generic term referring to delivery of an AI / ML model from one entity to another entity in any manner. Note: An entity could mean a network node / function (e.g., gNB, location management function (LMF), etc.), UE, proprietary server, etc.
[0042] AI / ML model Inference: A process of using a trained AI / ML model to produce a set of outputs based on a set of inputs.
[0043] AI / ML model testing: A subprocess of training, to evaluate the performance of a final AI / ML model using a dataset different from one used for model training and validation. Differently from AI / ML model validation, testing does not assume subsequent tuning of the model.
[0044] AI / ML model training: A process to train an AI / ML Model [by learning the input / output relationship] in a data driven manner and obtain the trained AI / ML Model for inference.
[0045] AI / ML model transfer: Delivery of an AI / ML model over the air interface in a manner that is not transparent to 3GPP signaling, either parameters of a model structure known at the receiving end or a new model with parameters. Delivery may contain a full model or a partial model.
[0046] AI / ML model validation: A subprocess of training, to evaluate the quality of an AI / ML model using a dataset different from one used for model training, that helps selecting model parameters that generalize beyond the dataset used for model training.
[0047] Data collection: A process of collecting data by the network nodes, management entity, or UE for the purpose of AI / ML model training, data analytics and inference.
[0048] Federated learning / federated training: A machine learning technique that trains an AI / ML model across multiple decentralized edge nodes (e.g., UEs, gNEJs) each performing local model training using local data samples. The technique requires multiple interactions of the model, but no exchange of local data samples.
[0049] Functionality identification: A process / method of identifying an AI / ML functionality for the common understanding between the network and the UE. Note: Information regarding the AI / ML functionality may be shared during functionality identification. Where AI / ML functionality resides depends on the specific use cases and sub use cases.
[0050] Model activation: enable an AI / ML model for a specific function.
[0051] Model deactivation: disable an AI / ML model for a specific function.
[0052] Model download: Model transfer from the network to UE.
[0053] Model identification: A process / method of identifying an AI / ML model for the common understanding between the network (NW) and the UE. Note: The process / method of model identification may or may not be applicable. Note: Information regarding the AI / ML model may be shared during model identification.
[0054] Model monitoring: A procedure that monitors the inference performance of the AI / ML model.
[0055] Model parameter update: Process of updating the model parameters of a model.
[0056] Model selection: The process of selecting an AI / ML model for activation among multiple models for the same AI / ML enabled feature. Note: Model selection may or may not be carried out simultaneously with model activation.
[0057] Model switching: Deactivating a currently active AI / ML model and activating a different AI / ML model for a specific function.
[0058] Model update: Process of updating the model parameters and / or model structure of a model.
[0059] Model upload: Model transfer from UE to the network.
[0060] Network-side (AI / ML) model: An AI / ML Model whose inference is performed entirely at the network.
[0061] Offline field data: The data collected from field and used for offline training of the AI / ML model.
[0062] Offline training: An AI / ML training process where the model is trained based on collected dataset, and where the trained model is later used or delivered for inference. Note: This definition only serves as a guidance. There may be cases that may not exactly conform to this definition but could still be categorized as offline training by commonly accepted conventions.
[0063] Online field data: The data collected from field and used for online training of the AI / ML model.
[0064] Online training: An AI / ML training process where the model being used for inference) is (typically continuously) trained in (near) real-time with the arrival of new training samples. Note: the notion of (near) real-time vs. non real-time is context-dependent and is relative to the inference timescale. Note: This definition only serves as a guidance. There may be cases that may not exactly conform to this definition but could still be categorized as online training by commonly accepted conventions. Note: Fine-tuning / re-training may be done via online or offline training. (This note could be removed when we define the term fine-tuning.)
[0065] Reinforcement Learning (RL): A process of training an AI / ML model from input (a.k.a. state) and a feedback signal (a.k.a. reward) resulting from the model’s output (a.k.a. action) in an environment the model is interacting with.
[0066] Semi-supervised learning: A process of training a model with a mix of labelled data and unlabeled data.
[0067] Supervised learning: A process of training a model from input and its corresponding labels.
[0068] Two-sided (AI / ML) model: A paired AI / ML Model(s) over which joint inference is performed, where joint inference comprises AI / ML Inference whose inference is performed jointly across the UE and the network, i.e, the first part of inference is firstly performed by UE and then the remaining part is performed by the gNB, or vice versa.
[0069] UE-side (AI / ML) model: An AI / ML Model whose inference is performed entirely at the UE.
[0070] Unsupervised learning: A process of training a model without labelled data.
[0071] FIG. 1 illustrates an example communication environment 100 in which example embodiments of the present disclosure can be implemented. In the communication environment 100, a plurality of communication devices, including a terminal device 110 and a network device 120, can communicate with each other. In the example of FIG. 1 , the terminal device 110 may be a UE and the network device 120 may be a base station serving the UE. The serving area of the network device 120 may be called a cell.
[0072] It is to be understood that the number of devices and their connections shown in FIG. 1 are only for the purpose of illustration without suggesting any limitation. The communication environment 100 may include any suitable number of devices configured to implementing example embodiments of the present disclosure. Although not shown, it would be appreciated that one or more additional devices may be located in the cell, and one or more additional cells may be deployed in the communication environment 100. It is noted that although illustrated as a network device, the network device 120 may be another device than a network device. Although illustrated as a terminal device, the terminal device 110 may be another device than a terminal device.
[0073] In the following, for the purpose of illustration, some example embodiments are described with the terminal device 110 operating as a UE and the network device 120 operating as a base station. However, in some example embodiments, operations described in connection with a terminal device may be implemented at a network device or other device, and operations described in connection with a network device may be implemented at a terminal device or other device.
[0074] In some example embodiments, a transmission direction from the network device 120 to the terminal device 110 is referred to as a downlink (DL), while a transmission direction from the terminal device 110 to the network device 120 is referred to as an uplink (UL). In DL, the network device 120 is a transmitting (TX) device (or a transmitter) and the terminal device 110 is a receiving (RX) device (or a receiver). In UL, the terminal device 110 is a TX device (or a transmitter) and the network device 120 is a RX device (or a receiver).
[0075] Communications in the communication environment 100 may be implemented according to any proper communication protocol(s), comprising, but not limited to, cellular communication protocols, wireless local network communication protocols such as Institute for Electrical and Electronics Engineers (IEEE) 802.11 and the like, and / or any other protocols currently known or to be developedin the future. Moreover, the communication may utilize any proper wireless communication technology, comprising but not limited to: Code Division Multiple Access (CDMA), Frequency Division Multiple Access (FDMA), Time Division Multiple Access (TDMA), Frequency Division Duplex (FDD), Time Division Duplex (TDD), Multiple-Input Multiple-Output (MIMO), Orthogonal Frequency Division Multiple (OFDM), Discrete Fourier Transform spread OFDM (DFT-s-OFDM) and / or any other technologies currently known or to be developed in the future.
[0076] As discussed, as one of the potential approaches for AI / ML model training, the training of AI / ML models is performed at a base station or at a 3rd party server based on massive data sharing and the model or at least the model parameters (or weights, in the following, AI / ML model parameters and weights will be simply referred to as AI / ML model parameters for ease of description as weights are specific types of AI / ML model parameters) are then delivered to UE. Transmitting AI / ML parameters over the 5G air interface is challenging but essential for leveraging the capabilities of AI / ML.
[0077] There are a lot of ongoing discussions on the topic of AI / ML model / parameter transfer / delivery, e.g., how often the parameters need to be transferred, or when UE needs new parameters for a known model structure and when transferred parameters are ready for inference, etc. In certain deployments, frequent model transmission is increasingly necessary, particularly for AI / ML models trained in or adapted to specific areas or configurations. Such models, often referred to as “localized models,” are tailored to unique locations, propagation environments, or network (NW) configurations identified by specific IDs.
[0078] When an AI / ML model (or its parameters) is to be transmitted from a base station (e.g., gNB) or any third-party server to a UE, the transmission occurs over an inherently error-prone wireless channel, similar to other wireless signals. For instance, such a transfer may be necessary when delivering a cell-specific model or a newly localized model, such as an AI / ML-based CSI encoder, to the UE as it moves within the network. During this wireless transmission, the model parameters are susceptible to distortion, leading to potential errors in the model parameters. Such distortions may significantly degrade the performance of the AI / ML model, as any erroneous parameter may critically alter the functionality of the AI / ML model, including convolutional neural networks (CNNs), long shortterm memory networks (LSTMs), or other deep learning architectures.
[0079] Referring now to FIG. 2, which illustrates an example of AI / ML model parameters 200. As illustrated in FIG. 2, the model parameters 200, such as those for a CNN, appear as seemingly random numbers following the training process. In various AI / ML models used for CSI compression, decompression, or prediction, the total number of these parameters may reach or exceed 100,000. Due to the high volume and sensitivity of these parameters, transmitting them from the base station to the UE over the air interface may cause huge overload from signaling point of view.
[0080] The wireless channel exhibits dynamic characteristics that vary over time within individual cells, influenced by factors such as peak and non-peak traffic hours, seasonal variations, weekdays, weekends, and special periods, including festivals. The base station may identify and leverage these periodic patterns to trigger updates for AI / ML models as needed. However, the challenge arises in maintaining consistent model accuracy across updates; specifically, the precision of model representation, such as the accuracy of model parameters, and the overhead required for model transmission may fluctuate with each update. Consequently, it becomes essential to consider the feasibility of implementing different model representations tailored to the unique requirements of each update.
[0081] The primary concept disclosed herein is to provide a solution to minimize errors during the delivery or transmission of AI / ML model parameters to achieve desired performance levels over an inherently error-prone wireless air interface, where model parameters are subject to potential distortion, similar to other data. Given that the accuracy of AI / ML model parameters significantly impacts model performance, numerous studies have investigated parameter quantization methods, including reducing parameter resolution, if necessary, by rounding values from higher to lower precision.
[0082] The proposed mechanism enables negotiation between the base station and the UE to establish a desired performance level for an AI / ML model based on varying wireless channel conditions. This process involves determining the optimal accuracy level of a lookup table, which defines the accuracy of the model representation when encoding and decoding (i.e., reconstructing) AI / ML model parameters during transmission from the base station to the UE. Encoding parameters at a higher accuracy level correlates with quantizing the model parameters at higher precision, thus enhancing the accuracy of the model. Ultimately, using a more precise model representation results in improved performance of the AI / ML functionality provided by the model.Basic Scenario for Transferring Parameters of an AI / ML Model from base station to UE
[0083] Transmitter (NW side):
[0084] 1 ) The transmitter may assess the performance requirements for a specific AI / ML feature, functionality, or capability, and quantizes the AI / ML model parameters at selected accuracy levels to meet these performance criteria. The transmitter may then select an appropriately accurate lookup table aligned with the quantized parameters and encodes these parameters accordingly.
[0085] 2) The indexes from the lookup table may act as an additional overhead in a standard data transmission. This process may occur during initial setup, such as when the connection between the UE and the base station is first established, and the encoder model has not yet been activated. Alternatively, when the UE requires updated AI / ML model parameters as specified by the network, the lookup table indexes transmitted from the network may become part of the normal overheadassociated with CSI compression steps.
[0086] Receiver (UE side):
[0087] Upon receiving the parameters from the overhead data, the UE may decode the AI / ML model parameters according to the same or a specified lookup table. The UE may then operate the AI / ML model at the required accuracy level to achieve the targeted performance.
[0088] Channel Adaptation:
[0089] Notably, wireless channel conditions fluctuate over time due to various factors, such as peak and non-peak traffic periods, weekdays versus weekends, and special occasions like festivals. The base station or core network has insight into these periodic changes, which may occur hourly, daily, or according to other defined intervals. Based on the prevailing channel scenario, the base station may select the optimal accuracy level for the AI / ML model parameters and transmit or signal this selection to the UE. Consequently, the UE may apply this designated accuracy level to run the AI / ML model, achieving the expected performance under the current channel conditions.
[0090] Each base station possesses a distinct physical location, leading to unique channel conditions in each cell across the world. The wireless channel characteristics associated with the geographical positioning of the base station and the UE exhibit periodic changes, which may occur hourly, daily, weekly, and so forth, following predictable patterns. This pattern, based on collected CSI, may only be determined accurately by the base station. Prior to the transfer or delivery of the AI / ML model to the UE by the base station or network, the type and number of parameters of the model need to be fixed to facilitate the reliable transmission of model parameters.
[0091] The UE encoder is a predetermined model type (e.g., CNN, LSTM, MLP, transformer) with a baseline configuration that meets the minimum requirements defined in the 3GPP specifications. The model configuration includes a set number of layers and specified inputs and outputs for each layer. If the UE is required to support different levels of encoder complexity, the hyperparameters for the model may be negotiated with the base station in the radio resource control (RRC) connected mode. When only a partial model update is transmitted by the base station, the number of layers may vary. In such cases, the base station and the UE may re-align to agree on the layers applicable for model transfer, and this layer information may be provided alongside the model parameters during transmission.
[0092] Once the number of layers and inputs / outputs per layer are fixed, only the parameters within the model change dynamically in response to channel conditions.
[0093] The base station continuously fine-tunes both the encoder and decoder models at the base station side, gathering UE CSI reports and sounding reference signal (SRS) data as training inputs. After training is completed, the base station may update the model parameters on a periodic basis (e.g., every hour) to account for the evolving channel environment. By periodically adapting the modelparameters, the base station ensures that the UE receives the optimal model representation for current channel conditions, thereby improving AI / ML model performance and accuracy.
[0094] The types of AI / ML encoders supported by the UE, such as CNN, LSTM, MLP, and transformers, along with their associated complexities (e.g., models with 100k, 300k, 1 M, or 2M layers / parameters), may be incorporated into the standards. These model types and architectures are known and described in prior art. Based on the computational performance of the UE and the unique channel conditions, the UE may negotiate with the base station or network for an appropriate AI / ML model type to be utilized in line with channel conditions.
[0095] Alternatively, or additionally, the UE may indicate its supported accuracy levels for model parameter quantization. These accuracy levels specify the decimal resolution for the parameters, with examples including values such as exp(10A-3), exp(10A-4), exp(10A-6), exp(10A-8), and etc.. The specified accuracy affects the size of the lookup table used for parameter representation, which will be discussed in further detail later with reference to the examples of the lookup tables illustrated in FIGs. 4A-4C.
[0096] The UE may further indicate whether it may accept AI / ML model parameters in either a lookup table format or an approximation function format (or other formats that can be used for this purpose). This flexibility in model transfer formats enables optimized transmission based on the capabilities of the UE and network conditions, thereby enhancing model delivery efficiency.
[0097] The UE may also have the capability to reject AI / ML model configurations transmitted by the base station. This allows the UE to refuse configurations that exceed its processing or storage capacity, ensuring compatibility with its computational resources.
[0098] Alternatively, or additionally, the base station may transmit a plurality of AI / ML model configurations, including the corresponding parameters, to the UE simultaneously. Upon receiving the configurations, the UE may construct each AI / ML model based on the provided parameters. The UE may hold these the plurality of models in a standby mode, awaiting further instructions. Once the models are prepared, the based station may subsequently issue a command indicating which specific AI / ML model the UE should activate. This capability enables the network to dynamically select the most appropriate model for the current network conditions or application requirements without requiring an additional model / parameter transfer session. This approach optimizes the network efficiency by reducing latency associated with on-demand model transfer, as the UE may immediately switch to a pre-loaded AI / ML model based on the indication of the base station. Furthermore, it allows the network to respond flexibly to varying channel conditions and computational requirements, as the plurality of model configurations stored at the UE facilitate quick adaptation to these factors.
[0099] Alternatively, or additionally, the UE and the base station may negotiate the AI / ML model configuration and accuracy level. The UE may initiate communication with the base station byreporting its supported AI / ML model type and the maximum accuracy level it may handle, which is based on the calculation performance of the UE’s hardware. The base station may determine the AI / ML model type, and the minimum accuracy level required to achieve the desired model performance on the UE based on the reported capabilities and historical channel and performance data. The base station may assign corresponding identifiers (IDs) to inform the UE which encoder / decoder instances should be active on both the UE and base station sides for effective data exchange. The base station and UE engage in the negotiation process to finalize the accuracy level of the AI / ML model parameters that will be applied in the lookup table. The accuracy level is determined within a range bounded by the minimum accuracy level requested by the base station and the maximum accuracy level the UE’s hardware may support, as illustrated in FIG. 3.
[0100] Referring now to FIG. 3, which illustrates an example of accuracy level agreement 300 thresholds between the UE and the network, which may be applied to each AI / ML mode type. FIG. 3 illustrates the accuracy level agreement thresholds between UE and the network (e.g., the gNB), highlighting the range in which both the UE and network may negotiate an acceptable accuracy level for the transmission of AI / ML model parameters. The area labeled “Accuracy level agreement area” 320 represents the mutually agreed-upon range, where the lowest threshold 310 is set by the minimum accuracy requested by the gNB, and the highest threshold 320 is determined by the maximum accuracy level the UE may support based on its computational capabilities.
[0101] Using a higher accuracy level within this range, supported by the UE, may improve the performance of AI / ML models in tasks like CSI prediction and CSI compression. Although transmitting model parameters with higher accuracy might lead to an increased message delivery overhead, it may enhance the overall model performance. This improvement, in turn, boosts the efficiency of data transmission, potentially increasing throughput before the next required model update (e.g., on an hourly basis). FIG. 3 effectively shows the balancing act between computational demands on the UE and performance expectations from the network. Examples of how to generate the lookup tables based on different accuracy levels will be described in the following with reference to FIGs. 4A-4C, which illustrate some examples of lookup tables in accordance with some embodiments in the disclosure.
[0102] Referring now to FIGs. 4A-4B, which illustrate an approximation function 400A to generate the lookup table and a 2D lookup table generated based on the approximation function. In FIG. 4A, the use of a non-linear approximation function, specifically y = xA3 for values x between -1 and 1 , is illustrated as a method to generate a 2D lookup table 400B shown in FIG. 4B for AIML model parameter transmission. This function generates “y” values that increase gradually near zero but accelerate as they approach -1 or 1 , creating a higher resolution for small absolute parameters and allowing lower resolution for larger parameters. This approach optimizes accuracy where needed without unnecessary precision for larger values. For instance, when the parameter value is relativelylarge, such as 0.31415926, a high-resolution accuracy up to 0.1 x 1 Q-8may not be necessary. Conversely, for smaller parameter values, such as 0.00000011 , it is preferred to maintain a resolution accuracy up to 0.1 x 1 Q-8. This differential resolution allocation based on the magnitude of parameter values optimizes data transmission by allocating higher precision only where needed.
[0103] The lookup table based on y = xA3, shown in FIG. 4B, may be scaled to cover a broader range by extending both x and values (e.g., from -5 to 5). For example, if a parameter value of 0.31415926 is represented, only the row and column indexes (7,14) from the table need to be transmitted, requiring just 8 bits (4 bits for each index) per model parameter. This indexing method reduces the data size necessary for parameter transmission, thus minimizing the bandwidth required and improving the efficiency of model parameter delivery across the network.
[0104] In some embodiments, the accuracy level of the lookup table used for AI / ML model parameters may be defined based on the size of the table, allowing for flexible and scalable accuracy control in model configuration. It is specified that the increase of the table size (e.g., from 32x32 to 128x128) may result in higher accuracy levels, thus allowing finer resolution for model parameter representation. It is possible to define accuracy level from 1 to 3 or more, for example,
[0105] Level 1 (32x32 Table): This configuration uses a lookup table size of 32x32, requiring 5 bits per dimension (10 bits total) for each model weight or parameter,
[0106] Level 2 (64x64 T able): A lookup table size of 64x64 enables a higher accuracy level, requiring 6 bits per dimension (12 bits total) per parameter,
[0107] Level 3 (128x128 Table): For the highest accuracy, a 128x128 table is used, requiring 7 bits per dimension (14 bits total) for each model parameter.
[0108] To facilitate interoperability and flexibility, the configuration of the lookup table may be defined by the following parameters:
[0109] Equation Parameters: The equation governing the lookup table is pre-defined with configurable variables. For example, the equation format may be (x + a)b+ c, where a, b, and c are adjustable parameters. These parameters will be communicated between the UE and base station to ensure accurate model reconstruction and configuration,
[0110] Accuracy Levels: The accuracy levels (1 through 3 or more) are based on the size of the lookup table, and each level specifies the bit requirements and resolution for model parameters, as detailed above.
[0111] Alternatively, or additionally, the lookup table may be created in a linear way, for example, as shown in Table 1 below:Tab e 1
[0112] For the example illustrated in Table 1 , each instance of the lookup table has a defined interval of 0.0016 between consecutive values in each row. To ensure continuity, the starting value of each subsequent row is precisely 0.0016 greater than the last value of the preceding row. This arrangement allows for an accurate representation of values across the entire table range. To further enhance accuracy, an interpolation method may be applied between pairs of values within a row. The interpolation may be conducted using a linear equation, defined as follows: y = O.OOOlx, where x = 0 ... 15 (1 )
[0113] Alternatively, a polynomial interpolation equation may be employed to achieve the desired level of precision. The combination of this structured lookup table (in terms of row and column) and the chosen interpolation method enables an accuracy level of up to 0.0001. Through this approach, the lookup table (Table 1) may accurately represent values within the range of 0 to 0.1008.
[0114] Consider a model parameter with a value of 0.0916, represented to four decimal points. This value lies between two values in the lookup table: 0.0912 at coordinates (8, 2) and 0.0928 at coordinates (8, 3). The closer value, 0.0912 at (8, 2) and x = 5 in interpolation (i.e., 0.0004 = 0. 0001 x 4 (5thindex in the interpolation) may be selected. So, 0.0916 = 0.0912 (8, 2) + 0.0004. For the 8 rows and 8 columns, 3 binary number may be used to represent it. Also, 4 binary numbers may be used to represent the 16 intervals in the interpolation. In total, one parameter may be represented by 3+3+4 = 10 binary bits.
[0115] In order to create a lookup table for representing model parameters in the range between [- 1 , 1], the table may be structured with 20 rows and 32 columns. Each value in a given row increments by an interval of 0.0032. This configuration allows representation of parameters with greater precision within the specified range. To achieve an accuracy level of 0.0001 , a linear interpolation equation is employed, defined as: y = O.OOOlx, where % = 0 ... 31 (2)
[0116] In some embodiments, the accuracy level of the lookup table used for AI / ML model parameters may be determined by both the size of the lookup table and the interpolation precision applied within the table. It is possible to define accuracy level from 1 to 3 or more, for example,
[0117] Level 1 : Lookup table with a 32x32 structure,
[0118] Level 2: Lookup table with a 64x64 structure,
[0119] Level 3: Lookup table with a 128x128 structure,
[0120] and the interpolation precision based on the selected interpolation method, such as linear or polynomial interpolation, and further depends on the depth of interpolation.
[0121] The configuration of this lookup table may include the following elements:
[0122] Range of the Table: Defines the value range represented within the lookup table,
[0123] Accuracy Levels (1 to 3 or more): Specifies the table size based on row and column counts, correlating with the defined accuracy level,
[0124] Interpolation Equation: The interpolation function used, such as a linear or polynomial expression, for instance, x + a)b+ c, where a, b, and c are configurable parameters transmitted between the UE and the base station,
[0125] Interpolation Depth: Specifies the bit depth used for interpolation.
[0126] This configuration allows for enhanced precision and adaptability of model parameter representation within variable accuracy levels and interpolation depths, enabling efficient transmission and reconstruction of model parameters at the UE.
[0127] Referring now to FIG. 4C, which illustrates an example of complex-valued lookup table 400C. The method for constructing the lookup table, shown in FIG. 4G, involves applying two distinct interpolation equations to increase accuracy for complex values. For example, linear or polynomial interpolation equations may be applied independently to the real and imaginary components of each value. This dual interpolation method enhances precision by separately refining both the real and imaginary parts of the complex numbers in the lookup table.
[0128] The accuracy level for this lookup table may be determined similar to the approach as outlined in the above with reference to Table 1 , incorporating the following elements:
[0129] Accuracy Levels: Similar to the accuracy levels defined in the above with reference to Table 1 , this embodiment allows for various precision tiers based on the interpolation method and table size,
[0130] Separate Interpolation Equations: Two interpolation equations, one for the real part and one for the imaginary part, are applied independently to achieve higher accuracy,
[0131] Interpolation Equation Type: The interpolation method may utilize either linear or polynomial equations to maximize the precision of each component.
[0132] This approach ensures that the real and imaginary parts of complex parameter are accurately represented, thereby enhancing the overall fidelity of model parameters in the lookup table.
[0133] In the above, some examples of the lookup table instances were described with references to FIGs. 4A-4C. It is to be understood that the examples described are not intended to impose any limitations, other types of lookup table instances may be used.
[0134] Referring now to FIGs. 5A-5B, which illustrate a diagram of an example message sequence in accordance with some embodiments. The UE 110 may be the terminal device 110 in FIG. 1 and the base station 120 may be the network device 120 in FIG. 2.
[0135] At 501 , the base station 120 may train / update AI / ML models based on the collected history data for the environment in a cell the base station 120 serves. The history data may include variations in traffic density, user movement patterns, interference levels, and other relevant parameters. As discussed in the above, the base station 120 has the insight into the periodic changes to the wireless channel conditions. The history data collected may be used to establish patterns of environmental changes over time and generate a history dataset, which may be used to detect trends, seasonal changes, or other relevant characteristics of the environment based on which, the base station 120 may train / update related AI / ML models.
[0136] At 502 and 503, the UE 110 may transmit to the base station 120 capability information indicating its supported AI / ML model types and the maximum quantization accuracy level achievable based on the computational capacity of its hardware. The scope of the capability information may vary according to the current specifications and developments in 3GPP AI / ML model standards. For example, the capability information may include: 1 ) supported model complexities, e.g., the maximum model complexity; 2) supported accuracies for quantization, e.g., maximum quantization accuracy, 3) supported transfer formats, e.g., lookup table and / or approximation function, and the like. The capability information may further include the AI / ML capabilities / models supported or available and / or corresponding specifications, identifiers, and the like. The capability information may be transmitted: 1 ) in the standard payload of a message via Physical Downlink Shared Channel (PDSCH) while in RRC_CONNECTED mode; or 2) through Radio Resource Control (RRC) configuration signaling prior to the UE’s entry into RRC_CONNECTED mode.
[0137] At 504, based on the capability information provided by the UE 110, along with history data collected by the base station 120 at 501 , the base station 120 may select an appropriate AI / ML model type and establish a baseline quantization accuracy level as the network AI / ML configuration. This baseline accuracy level is set to meet minimum performance requirements for the AI / ML model to function effectively on the UE 110. The base station 120 may then assign specific IDs to the selected model configuration to inform the UE 110 of which encoder / decoder instances should be deployed on both the UE 110 and base station 120 sides for synchronized data exchange.
[0138] The final quantization accuracy level of the AI / ML model parameters may be determined through negotiation between the UE 110 and the base station 120, considering the performance limitations and computational capabilities of the UE’s hardware. The negotiation process ensures that the agreed-upon accuracy level falls between the minimum level requested by the base station 120 and the maximum level supported by the UE 110, as discussed in the above with reference to FIG. 3. And at 505, the UE 110 and the base station 120 may exchange the values of the AI / ML model parameters at the negotiated agreed-upon accuracy level.
[0139] The UE 110 may transmit negotiation messages to the base station 120 following the initialconfiguration at 504, proposing a preferred accuracy level for the AI / ML model parameters. If the accuracy level requested by the base station 120 is equal to or lower than the maximum accuracy level that the UE 110 can support, the UE 110 may send an acknowledgment to confirm acceptance of the configuration.
[0140] Upon receiving the UE’s negotiation messages, the base station 120 may update the accuracy level for the AI / ML model parameters to accommodate the UE’s selection. If the UE’s chosen accuracy level is at or exceeds the baseline requested by the base station 120, the base station 120 acknowledges the proposed configuration. As discussed in the above, using a higher accuracy level within the range illustrated in FIG. 3, supported by the UE 110, may improve the performance of AI / ML models in tasks like CSI prediction and CSI compression. Although transmitting model parameters with higher accuracy might lead to an increased message delivery overhead, it may enhance the overall model performance. This improvement, in turn, boosts the efficiency of data transmission, potentially increasing throughput before the next required model update (e.g., on an hourly basis).
[0141] The information exchange at 502 to 505 may occur over RRC signaling during connection (re)establishment or through RRC (re)configuration. Alternatively, the signaling exchange may also be conducted within RRC_CONNECTED mode.
[0142] The UE 110 may also have the capability to reject AI / ML model configurations transmitted by the base station 120 if the configuration exceeds its processing or storage capacity.
[0143] The exchange of the values of the AI / ML model parameters at the negotiated agreed-upon accuracy level at 505 may be conveyed in various formats to support efficient model adaptation and deployment. For example:
[0144] The UE 110 and the base station 120 may transfer the values of the AI / ML model parameters by exchanging lookup table identifiers (IDs). A plurality of lookup tables may be pre-defined, with each table assigned a unique lookup table ID and corresponding to different levels of accuracy for the same model parameters. By exchanging the lookup table ID and the corresponding indexes, the values of the AI / ML model parameters may be delivered and reconstructed with a few bits, as described in the example of lookup tables with reference to FIG. 4A-4C.
[0145] Alternatively, the values of the AI / ML model parameters may be directly transferred in the form of complete lookup tables. Each table is specifically defined to reflect values of the same model parameters at distinct accuracy levels, allowing the UE 110 to access values optimized for the negotiated level of accuracy.
[0146] Alternatively, the values of the AI / ML model parameters may also be transferred by referencing an approximation function. This may include an identifier (ID) for a specific function or polynomial, or a set of coefficients and / or exponents (variables) that define the function or polynomial. A single approximation function may be defined to span multiple accuracy levels, enabling the UE 110and / or the base station 120 to derive the values of the AI / ML model parameters by applying the function at the selected accuracy level. The lookup table may be dynamically generated based on an approximation function, or the values of the AI / ML model parameters may be directly computed from the approximation function, bypassing the need for a lookup table. In embodiments where lookup tables or lookup table IDs are transferred, the base station 120 may additionally inform the UE 110 regarding the application of interpolation between consecutive values listed in the table. This information may include the specified number of bits allocated for interpolation, which may range from 1 bit up to 8 or 16 bits, by way of example, as described in the example of lookup tables with reference to FIG. 4A-4C.
[0147] Alternatively, the base station 120 may transmit a plurality of AI / ML model configurations, including the corresponding parameters, to the UE 110 simultaneously. Upon receiving the configurations, the UE 110 may construct each AI / ML model based on the provided parameters. The UE 110 may hold these the plurality of models in a standby mode, awaiting further instructions (see 511 ). Once the models are prepared, the based station may subsequently issue a command indicating which specific AI / ML model the UE 110 should activate. This capability enables the network to dynamically select the most appropriate model for the current network conditions or application requirements without requiring an additional model / parameter transfer session. The approach optimizes the network efficiency by reducing latency associated with on-demand model transfer, as the UE 110 may immediately switch to a pre-loaded AI / ML model based on the indication of the base station. Furthermore, it allows the network to respond flexibly to varying channel conditions and computational requirements, as the plurality of model configurations stored at the UE facilitate quick adaptation to these factors.
[0148] At 506, upon completion of configuration, the AI / ML feature, functionality, or model may be activated using a unique identifier (ID) assigned within both the base station 120 and the UE 110. The base station 120 may transmit the ID to the UE 110 to ensure synchronized activation.
[0149] At 507, in cases where the UE 110 determines that the current AI / ML model is no longer adequate, applicable, or up to date, the UE 110 may initiate a request to re-negotiate with the base station 120 for an alternative AI / ML model. Such a request may be triggered if the UE 110 identifies incompatibilities with current network conditions, configurations, scenarios, or updated network-side models. The request message may be sent over the Physical Uplink Shared Channel (PUSCH) or Physical Uplink Control Channel (PUCCH) using pre-defined indicators or, alternatively, through dedicated signaling methods indicating whether it may accept AI / ML model parameters in either a lookup table format or an approximation function format (or other formats that can be used for this purpose).
[0150] At 508 (also 518), following the UE’s request or a network-initiated need for adaptation, thebase station 120 may update or change the AI / ML configurations and may further instruct the UE 110 to switch to a different AI / ML model type. The base station 120 and UE 110 may re-negotiate all configuration details, adjusting to the channel conditions or scenario, in alignment with the process outlined through 501 to 506.
[0151] At 509 (also 519), upon finalizing the configuration updates, the base station 120 may encode the values of the AI / ML model parameters based on the pre-determined lookup table(s) and / or approximation function(s) in accordance with the agreed accuracy level. The resulting indexes may then be sent to the UE 110. If interpolation was specified at 505, the base station 120 may indicate the values of the AI / ML model parameters by providing indexes, alongside an interpolation indication. This indication allows the selection of a value of the AI / ML parameter as an interpolated value between the indexed value and the subsequent value in the lookup table, thereby enhancing the accuracy of the transferred parameter beyond the precision directly represented by the lookup table.
[0152] At 510 (also 520), upon receiving the updated AI / ML configuration and transferred values of the AI / ML model parameters, the UE 110 may reconstruct or update the AI / ML model in accordance with the configuration instructions provided by the base station 120.
[0153] At 511 , as an alternative, the base station 120 may transmit multiple AI / ML model configurations and their associated parameters to the UE 110 simultaneously, utilizing the steps outlined in 501 through 509. Upon receipt, the UE 110 may construct each AI / ML model, holding them in standby until a specific model is designated for activation by the base station at 511 .
[0154] At 512, as another alternative, the UE 110 may independently select and indicate the AI / ML model to be activated and used, based on the configurations and parameters received from the base station 120.
[0155] At 513 to 514, the base station 120 may transmit reference signals, such as Channel State Information Reference Signals (CSI-RS), to the UE 110 to facilitate AI / ML model inference. The UE 110 may use these reference signals, or real-time signals from the base station 120, to perform inference using the configured AI / ML model(s) and prepares reports for transmission back to the base station 120.
[0156] At 515 to 517, the UE 110 may transmit the inference results from the AI / ML model(s), which may include outputs such as CSI feedback, predictions, or other relevant metrics, to the base station 120. The base station 120 may monitor the AI / ML model performance on the UE 110 by analyzing these reports and collects the UE-generated data (e.g., CSI reports) to serve as new data points for periodic fine-tuning of the AI / ML model, including updating model parameters as needed.
[0157] FIG. 6 shows a flowchart of an example method 600 implemented at a first apparatus in accordance with some example embodiments of the present disclosure. For the purpose of discussion, the method 600 will be described from the perspective of the terminal device 110 in FIG. 1.
[0158] At block 610, obtaining, from a second apparatus, accuracy level information indicating parameter quantization for an artificial intelligence and machine learning (AI / ML) model, the accuracy level information indicating at least one accuracy level.
[0159] At block 620, obtaining, from the second apparatus, indication information of at least one set of parameters of the AI / ML model corresponding to the at least one accuracy level.
[0160] At block 630, generating, the at least one set of parameters based on the indication information and the at least one accuracy level.
[0161] In some example embodiments, the method 600 further comprises: transmitting, to the second apparatus, first information associated with an AI / ML capabilities of the first apparatus; receiving, from the second apparatus, second information associated with a baseline AI / ML capabilities determined based on the first information; wherein the at least one accuracy level is determined based on the first information and the second information.
[0162] In some example embodiments, the first information comprises at least one of: supported model complexities; supported accuracies for quantization; or supported transfer formats.
[0163] In some example embodiments, the at least one accuracy level comprises the supported maximum accuracy for quantization.
[0164] In some example embodiments, the first information further comprises information of AI / ML capabilities supported by the first apparatus; and the second information comprises information of AI / ML capabilities supported by the second apparatus and requirement to the first apparatus; wherein the AI / ML model is determined based on the information of AI / ML capabilities supported by the first apparatus and the information of AI / ML capabilities supported by the second apparatus.
[0165] In some example embodiments, the indication information of the at least one set of parameters is obtained using one of the supported transfer formats.
[0166] In some example embodiments, the method 600 further comprises: generating, the at least one set of parameters based on the at least one lookup table.
[0167] In some example embodiments, the method 600 further comprises: generating, the at least one set of parameters based on the indexes of the at least one lookup table.
[0168] In some example embodiments, the method 600 further comprises: generating, the at least one set of parameters based on the variables of the at least one function.
[0169] In some example embodiments, the at least one lookup table is pre-defined or dynamically generated.
[0170] In some example embodiments, the at least one lookup table corresponding to the at least one accuracy level is generated by applying one of the following: an approximation equation; a linear equation with an interpolation function; or one or more distinct interpolation functions.
[0171] In some example embodiments, the at least one function comprises at least one interpolationfunction with configurable parameters as the variables comprised in the indication information of the at least one set of parameters.
[0172] In some example embodiments, the at least one set of parameters are quantized at a plurality of accuracy levels comprising the at least one accuracy level.
[0173] In some example embodiments, the first information and the second information are exchanged within the RRC_CONNECTED mode, or over Radio Resource Control (RRC) signaling before entering RRC_CONNECTED mode.
[0174] In some example embodiments, the first apparatus is or is comprised in a terminal device, and wherein the second apparatus is or is comprised in a network device.
[0175] FIG. 7 shows a flowchart of an example method 700 implemented at a second apparatus in accordance with some example embodiments of the present disclosure. For the purpose of discussion, the method 700 will be described from the perspective of the terminal device 110 in FIG. 1.
[0176] At block 710, transmitting, to a first apparatus, accuracy level information indicating parameter quantization for an artificial intelligence and machine learning (AI / ML) model, the accuracy level information indicating at least one accuracy level.
[0177] At block 720, transmitting, to the first apparatus, indication information of at least one set of parameters of the AI / ML model corresponding to the at least one accuracy level.
[0178] In some example embodiments, the method 700 further comprises: receiving, from the first apparatus, first information associated with an AI / ML capabilities of the first apparatus; transmitting, to the first apparatus, second information associated with a baseline AI / ML capabilities determined based on the first information; wherein the at least one accuracy level is determined based on the first information and the second information.
[0179] In some example embodiments, the first information comprises at least one of: supported model complexities; supported accuracies for quantization; or supported transfer formats.
[0180] In some example embodiments, the at least one accuracy level comprises the supported maximum accuracy for quantization.
[0181] In some example embodiments, the first information further comprises information of AI / ML capabilities supported by the first apparatus; and the second information comprises information of AI / ML capabilities supported by the second apparatus and requirement to the first apparatus; wherein the AI / ML model is determined based on the information of AI / ML capabilities supported by the first apparatus and the information of AI / ML capabilities supported by the second apparatus.
[0182] In some example embodiments, the indication information of the at least one set of parameters is transmitted using one of the supported transfer formats.
[0183] In some example embodiments, the indication information of the at least one set of parameters comprises at least one lookup table corresponding to the at least one accuracy level.
[0184] In some example embodiments, the indication information of the at least one set of parameters comprises indexes of at least one lookup table corresponding to the at least one accuracy level.
[0185] In some example embodiments, the indication information of the at least one set of parameters comprises variables of at least one function corresponding to the at least one accuracy level.
[0186] In some example embodiments, the at least one lookup table is pre-defined or dynamically generated.
[0187] In some example embodiments, the at least one lookup table corresponding to the at least one accuracy level is generated by applying one of the following: an approximation equation; a linear equation with an interpolation function; or one or more distinct interpolation functions.
[0188] In some example embodiments, the at least one function comprises at least one interpolation function with configurable parameters as the variables comprised in the indication information of the at least one set of parameters.
[0189] In some example embodiments, the at least one set of parameters are quantized at a plurality of accuracy levels comprising the at least one accuracy level.
[0190] In some example embodiments, the first information and the second information are exchanged within the RRC_CONNECTED mode, or over Radio Resource Control (RRC) signaling before entering RRC_CONNECTED mode.
[0191] In some example embodiments, the first apparatus is or is comprised in a terminal device, and wherein the second apparatus is or is comprised in a network device.
[0192] In some example embodiments, a first apparatus capable of performing any of the method 600 (for example, the terminal device 110 in FIG. 1 ) may comprise means for performing the respective operations of the method 600. The means may be implemented in any suitable form. For example, the means may be implemented in a circuitry or software module. The first apparatus may be implemented as or included in the terminal device 110 in FIG. 1 .
[0193] In some example embodiments, the first apparatus comprises means for obtaining, from a second apparatus, accuracy level information indicating parameter quantization for an artificial intelligence and machine learning (AI / ML) model, the accuracy level information indicating at least one accuracy level; means for obtaining, from the second apparatus, indication information of at least one set of parameters of the AI / ML model corresponding to the at least one accuracy level; and means for generating, the at least one set of parameters based on the indication information and the at least one accuracy level.
[0194] In some example embodiments, the first apparatus further comprises: means for transmitting, to the second apparatus, first information associated with an AI / ML capabilities of the first apparatus;means for receiving, from the second apparatus, second information associated with a baseline AI / ML capabilities determined based on the first information; wherein the at least one accuracy level is determined based on the first information and the second information.
[0195] In some example embodiments, the first information comprises at least one of: supported model complexities; supported accuracies for quantization; or supported transfer formats.
[0196] In some example embodiments, the at least one accuracy level comprises the supported maximum accuracy for quantization.
[0197] In some example embodiments, the first information further comprises information of AI / ML capabilities supported by the first apparatus; and the second information comprises information of AI / ML capabilities supported by the second apparatus and requirement to the first apparatus; wherein the AI / ML model is determined based on the information of AI / ML capabilities supported by the first apparatus and the information of AI / ML capabilities supported by the second apparatus.
[0198] In some example embodiments, the indication information of the at least one set of parameters is obtained using one of the supported transfer formats.
[0199] In some example embodiments, the first apparatus further comprises: means for generating, the at least one set of parameters based on the at least one lookup table.
[0200] In some example embodiments, the first apparatus further comprises: means for generating, the at least one set of parameters based on the indexes of the at least one lookup table.
[0201] In some example embodiments, the first apparatus further comprises: means for generating, the at least one set of parameters based on the variables of the at least one function.
[0202] In some example embodiments, the at least one lookup table is pre-defined or dynamically generated.
[0203] In some example embodiments, the at least one lookup table corresponding to the at least one accuracy level is generated by applying one of the following: an approximation equation; a linear equation with an interpolation function; or one or more distinct interpolation functions.
[0204] In some example embodiments, the at least one function comprises at least one interpolation function with configurable parameters as the variables comprised in the indication information of the at least one set of parameters.
[0205] In some example embodiments, the at least one set of parameters are quantized at a plurality of accuracy levels comprising the at least one accuracy level.
[0206] In some example embodiments, the first information and the second information are exchanged within the RRC_CONNECTED mode, or over Radio Resource Control (RRC) signaling before entering RRC_CONNECTED mode.
[0207] In some example embodiments, the first apparatus is or is comprised in a terminal device, and wherein the second apparatus is or is comprised in a network device.
[0208] In some example embodiments, a second apparatus capable of performing any of the method 700 (for example, the network device 120 in FIG. 1) may comprise means for performing the respective operations of the method 700. The means may be implemented in any suitable form. For example, the means may be implemented in a circuitry or software module. The second apparatus may be implemented as or included in the network device 1120 in FIG. 1 .
[0209] In some example embodiments, the second apparatus comprises means for transmitting, to a first apparatus, accuracy level information indicating parameter quantization for an artificial intelligence and machine learning (AI / ML) model, the accuracy level information indicating at least one accuracy level; and means for transmitting, to the first apparatus, indication information of at least one set of parameters of the AI / ML model corresponding to the at least one accuracy level.
[0210] I n some example embodiments, the second apparatus further comprises: means for receiving, from the first apparatus, first information associated with an AI / ML capabilities of the first apparatus; means for transmitting, to the first apparatus, second information associated with a baseline AI / ML capabilities determined based on the first information; wherein the at least one accuracy level is determined based on the first information and the second information.
[0211] In some example embodiments, the first information comprises at least one of: supported model complexities; supported accuracies for quantization; or supported transfer formats.
[0212] In some example embodiments, the at least one accuracy level comprises the supported maximum accuracy for quantization.
[0213] In some example embodiments, the first information further comprises information of AI / ML capabilities supported by the first apparatus; and the second information comprises information of AI / ML capabilities supported by the second apparatus and requirement to the first apparatus; wherein the AI / ML model is determined based on the information of AI / ML capabilities supported by the first apparatus and the information of AI / ML capabilities supported by the second apparatus.
[0214] In some example embodiments, the indication information of the at least one set of parameters is transmitted using one of the supported transfer formats.
[0215] In some example embodiments, the indication information of the at least one set of parameters comprises at least one lookup table corresponding to the at least one accuracy level.
[0216] In some example embodiments, the indication information of the at least one set of parameters comprises indexes of at least one lookup table corresponding to the at least one accuracy level.
[0217] In some example embodiments, the indication information of the at least one set of parameters comprises variables of at least one function corresponding to the at least one accuracy level.
[0218] In some example embodiments, the at least one lookup table is pre-defined or dynamicallygenerated.
[0219] In some example embodiments, the at least one lookup table corresponding to the at least one accuracy level is generated by applying one of the following: an approximation equation; a linear equation with an interpolation function; or one or more distinct interpolation functions.
[0220] In some example embodiments, the at least one function comprises at least one interpolation function with configurable parameters as the variables comprised in the indication information of the at least one set of parameters.
[0221] In some example embodiments, the at least one set of parameters are quantized at a plurality of accuracy levels comprising the at least one accuracy level.
[0222] In some example embodiments, the first information and the second information are exchanged within the RRC_CONNECTED mode, or over Radio Resource Control (RRC) signaling before entering RRC_CONNECTED mode.
[0223] In some example embodiments, the first apparatus is or is comprised in a terminal device, and wherein the second apparatus is or is comprised in a network device.
[0224] FIG. 8 is a simplified block diagram of a device 800 that is suitable for implementing example embodiments of the present disclosure. The device 800 may be provided to implement a communication device, for example, the terminal device 110 or the network device 120 as shown in FIG. 1. As shown, the device 800 includes one or more processors 810, one or more memories 820 coupled to the processor 810, and one or more communication modules 840 coupled to the processor 810.
[0225] The communication module 840 is for bidirectional communications. The communication module 840 has one or more communication interfaces to facilitate communication with one or more other modules or devices. The communication interfaces may represent any interface that is necessary for communication with other network elements. In some example embodiments, the communication module 840 may include at least one antenna.
[0226] The processor 810 may be of any type suitable to the local technical network and may include one or more of the following: general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs) and processors based on multicore processor architecture, as non-limiting examples. The device 800 may have multiple processors, such as an application specific integrated circuit chip that is slaved in time to a clock which synchronizes the main processor.
[0227] The memory 820 may include one or more non-volatile memories and one or more volatile memories. Examples of the non-volatile memories include, but are not limited to, a Read Only Memory (ROM) 824, an electrically programmable read only memory (EPROM), a flash memory, a hard disk, a compact disc (CD), a digital video disk (DVD), an optical disk, a laser disk, and other magneticstorage and / or optical storage. Examples of the volatile memories include, but are not limited to, a random-access memory (RAM) 822 and other volatile memories that will not last in the power-down duration.
[0228] A computer program 830 includes computer executable instructions that are executed by the associated processor 810. The instructions of the program 830 may include instructions for performing operations / acts of some example embodiments of the present disclosure. The program 830 may be stored in the memory, e.g., the ROM 824. The processor 810 may perform any suitable actions and processing by loading the program 830 into the RAM 822.
[0229] The example embodiments of the present disclosure may be implemented by means of the program 830 so that the device 800 may perform any process of the disclosure as discussed with reference to FIG. 2 to FIG. 7. The example embodiments of the present disclosure may also be implemented by hardware or by a combination of software and hardware.
[0230] In some example embodiments, the program 830 may be tangibly contained in a computer readable medium which may be included in the device 800 (such as in the memory 820) or other storage devices that are accessible by the device 800. The device 800 may load the program 830 from the computer readable medium to the RAM 822 for execution. In some example embodiments, the computer readable medium may include any types of non-transitory storage medium, such as ROM, EPROM, a flash memory, a hard disk, CD, DVD, and the like. The term “non-transitory,” as used herein, is a limitation of the medium itself (i.e. , tangible, not a signal) as opposed to a limitation on data storage persistency (e.g., RAM vs. ROM).
[0231] FIG. 9 shows an example of the computer readable medium 900 which may be in form of CD, DVD or other optical storage disk. The computer readable medium 900 has the program 830 stored thereon.
[0232] Generally, various embodiments of the present disclosure may be implemented in hardware or special purpose circuits, software, logic or any combination thereof. Some aspects may be implemented in hardware, and other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device. Although various aspects of embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representations, it is to be understood that the block, apparatus, system, technique or method described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
[0233] Some example embodiments of the present disclosure also provide at least one computer program product tangibly stored on a computer readable medium, such as a non-transitory computer readable medium. The computer program product includes computer-executable instructions, such asthose included in program modules, being executed in a device on a target physical or virtual processor, to carry out any of the methods as described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, or the like that perform particular tasks or implement particular abstract data types. The functionality of the program modules may be combined or split between program modules as desired in various embodiments. Machineexecutable instructions for program modules may be executed within a local or distributed device. In a distributed device, program modules may be located in both local and remote storage media.
[0234] Program code for carrying out methods of the present disclosure may be written in any combination of one or more programming languages. The program code may be provided to a processor or controller of a general-purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program code, when executed by the processor or controller, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may execute entirely on a machine, partly on the machine, as a stand-alone software package, partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0235] In the context of the present disclosure, the computer program code or related data may be carried by any suitable carrier to enable the device, apparatus or processor to perform various processes and operations as described above. Examples of the carrier include a signal, computer readable medium, and the like.
[0236] The computer readable medium may be a computer readable signal medium or a computer readable storage medium. A computer readable medium may include but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random-access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0237] Further, although operations are depicted in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Likewise, although several specific implementation details are contained in the above discussions, these should not be construed as limitations on the scope of the present disclosure, but rather as descriptions of features that may be specific to particular embodiments. Unless explicitly stated, certain features that are described in thecontext of separate embodiments may also be implemented in combination in a single embodiment. Conversely, unless explicitly stated, various features that are described in the context of a single embodiment may also be implemented in a plurality of embodiments separately or in any suitable subcombination.
[0238] Although the present disclosure has been described in languages specific to structural features and / or methodological acts, it is to be understood that the present disclosure defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
What is claimed is:1 . A first apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the first apparatus at least to: obtain, from a second apparatus, accuracy level information indicating parameter quantization for an artificial intelligence and machine learning (AI / ML) model, the accuracy level information indicating at least one accuracy level; obtain, from the second apparatus, indication information of at least one set of parameters of the AI / ML model corresponding to the at least one accuracy level; and generate, the at least one set of parameters based on the indication information and the at least one accuracy level.
2. The first apparatus of claim 1 , wherein the first apparatus is caused to: transmit, to the second apparatus, first information associated with an AI / ML capabilities of the first apparatus; receive, from the second apparatus, second information associated with a baseline AI / ML capabilities determined based on the first information; wherein the at least one accuracy level is determined based on the first information and the second information.
3. The first apparatus of claim 2, wherein the first information comprises at least one of: supported model complexities; supported accuracies for quantization; or supported transfer formats.
4. The first apparatus of claim 3, wherein the at least one accuracy level comprises the supported maximum accuracy for quantization.
5. The first apparatus of any of claims 1 to 4, wherein the first information further comprises information of AI / ML capabilities supported by the first apparatus; and the second information comprises information of AI / ML capabilities supported by the secondapparatus and requirement to the first apparatus; wherein the AI / ML model is determined based on the information of AI / ML capabilities supported by the first apparatus and the information of AI / ML capabilities supported by the second apparatus.
6. The first apparatus of any of claims 2 to 5, wherein the indication information of the at least one set of parameters is obtained using one of the supported transfer formats.
7. The first apparatus of claim 6, wherein the indication information of the at least one set of parameters comprises at least one lookup table corresponding to the at least one accuracy level, and the first apparatus is caused to: generate, the at least one set of parameters based on the at least one lookup table.
8. The first apparatus of claim 6, wherein the indication information of the at least one set of parameters comprises indexes of at least one lookup table corresponding to the at least one accuracy level, and the first apparatus is caused to: generate, the at least one set of parameters based on the indexes of the at least one lookup table.
9. The first apparatus of claim 6, wherein the indication information of the at least one set of parameters comprises variables of at least one function corresponding to the at least one accuracy level, and the first apparatus is caused to: generate, the at least one set of parameters based on the variables of the at least one function.
10. The first apparatus of any of claims 7 to 9, wherein the at least one lookup table is predefined or dynamically generated.
11. The first apparatus of any of claims 7 to 9, wherein the at least one lookup table corresponding to the at least one accuracy level is generated by applying one of the following: an approximation equation; a linear equation with an interpolation function; or one or more distinct interpolation functions.
12. The first apparatus of 9, wherein the at least one function comprises at least one interpolation function with configurable parameters as the variables comprised in the indicationinformation of the at least one set of parameters.
13. The first apparatus of any of claims 1 to 12, wherein the at least one set of parameters are quantized at a plurality of accuracy levels comprising the at least one accuracy level.
14. The first apparatus of any of claims 2 to 13, wherein the first information and the second information are exchanged within the RRC_CONNECTED mode, or over Radio Resource Control (RRC) signaling before entering RRC_CONNECTED mode.
15. The first apparatus of any of claims 2 to 14, wherein the first apparatus is or is comprised in a terminal device, and wherein the second apparatus is or is comprised in a network device.
16. A second apparatus, comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the second apparatus at least to: transmit, to a first apparatus, accuracy level information indicating parameter quantization for an artificial intelligence and machine learning (AI / ML) model, the accuracy level information indicating at least one accuracy level; and transmit, to the first apparatus, indication information of at least one set of parameters of the AI / ML model corresponding to the at least one accuracy level.
17. The second apparatus of claim 16, wherein the second apparatus is caused to: receive, from the first apparatus, first information associated with an AI / ML capabilities of the first apparatus; transmit, to the first apparatus, second information associated with a baseline AI / ML capabilities determined based on the first information; wherein the at least one accuracy level is determined based on the first information and the second information.
18. The second apparatus of claim 17, wherein the first information comprises at least one of: supported model complexities; supported accuracies for quantization; or supported transfer formats.
19. The second apparatus of claim 18, wherein the at least one accuracy level comprises the supported maximum accuracy for quantization.
20. The second apparatus of any of claims 16 to 19, wherein the first information further comprises information of AI / ML capabilities supported by the first apparatus; and the second information comprises information of AI / ML capabilities supported by the second apparatus and requirement to the first apparatus; wherein the AI / ML model is determined based on the information of AI / ML capabilities supported by the first apparatus and the information of AI / ML capabilities supported by the second apparatus.21 . The second apparatus of any of claims 17 to 20, wherein the indication information of the at least one set of parameters is transmitted using one of the supported transfer formats.
22. The second apparatus of claim 21 , wherein the indication information of the at least one set of parameters comprises at least one of: at least one lookup table corresponding to the at least one accuracy level; indexes of at least one lookup table corresponding to the at least one accuracy level; or variables of at least one function corresponding to the at least one accuracy level.
23. The second apparatus of claim 22, wherein the at least one lookup table is pre-defined or dynamically generated.
24. The second apparatus of claim 22, wherein the at least one lookup table corresponding to the at least one accuracy level is generated by applying one of the following: an approximation equation; a linear equation with an interpolation function; or one or more distinct interpolation functions.
25. The second apparatus of claim 22, wherein the at least one function comprises at least one interpolation function with configurable parameters as the variables comprised in the indication information of the at least one set of parameters.
26. The second apparatus of any of claims 16 to 25, wherein the at least one set of parametersare quantized at a plurality of accuracy levels comprising the at least one accuracy level.
27. The second apparatus of any of claims 17 to 26, wherein the first information and the second information are exchanged within the RRC_CONNECTED mode, or over Radio Resource Control (RRC) signaling before entering RRC_CONNECTED mode.
28. The second apparatus of any of claims 17 to 27, wherein the first apparatus is or is comprised in a terminal device, and wherein the second apparatus is or is comprised in a network device.
29. A method comprising: obtaining, from a second apparatus, accuracy level information indicating parameter quantization for an artificial intelligence and machine learning (AI / ML) model, the accuracy level information indicating at least one accuracy level. obtaining, from the second apparatus, indication information of at least one set of parameters of the AI / ML model corresponding to the at least one accuracy level. generating, the at least one set of parameters based on the indication information and the at least one accuracy level.
30. A method comprising: transmitting, to a first apparatus, accuracy level information indicating parameter quantization for an artificial intelligence and machine learning (AI / ML) model, the accuracy level information indicating at least one accuracy level. transmitting, to the first apparatus, indication information of at least one set of parameters of the AI / ML model corresponding to the at least one accuracy level.31 . A first apparatus comprising: means for obtaining, from a second apparatus, accuracy level information indicating parameter quantization for an artificial intelligence and machine learning (AI / ML) model, the accuracy level information indicating at least one accuracy level; means for obtaining, from the second apparatus, indication information of at least one set of parameters of the AI / ML model corresponding to the at least one accuracy level; and means for generating, the at least one set of parameters based on the indication information and the at least one accuracy level.
32. A second apparatus comprising: means for transmitting, to a first apparatus, accuracy level information indicating parameter quantization for an artificial intelligence and machine learning (AI / ML) model, the accuracy level information indicating at least one accuracy level; and means for transmitting, to the first apparatus, indication information of at least one set of parameters of the AI / ML model corresponding to the at least one accuracy level.
33. A computer readable medium comprising instructions stored thereon for causing an apparatus at least to perform the method of claim 31 or the method of claim 32.
34. A computer program comprising instructions for causing an apparatus at least to perform the method of claim 29 or the method of claim 30.