Quantization of data for machine learning model

By adaptively quantizing measurement data based on AI/ML model accuracy needs, the method addresses inefficiencies in current 3GPP standards, reducing energy and bandwidth consumption while optimizing resource usage in radio networks.

WO2025124796A1PCT designated stage expired Publication Date: 2025-06-19NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/081022
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-13
Filing Date
2024-11-04
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Current 3GPP standards use a fixed 7-bit quantization for UE measurements, which is inefficient for AI/ML applications, leading to increased energy and bandwidth consumption, and limited support for varying quantization levels based on model performance and resource availability.

Method used

Adaptive quantization of measurement data based on experimenting with different quantization levels to optimize machine learning model accuracy, allowing for varying quantization levels to be determined by network nodes and configured for user devices, thereby reducing overhead and enhancing resource utilization.

Benefits of technology

This approach reduces reporting overhead, decreases energy consumption at user equipment, and optimizes resource usage in radio networks, while maintaining acceptable inference accuracy for AI/ML applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024081022_19062025_PF_FP_ABST
    Figure EP2024081022_19062025_PF_FP_ABST
Patent Text Reader

Abstract

As an aspect, there is provided an apparatus suitable for determining at least one compressed quantization level for conveying measurement data for inference input of at least one machine learning model, wherein the determining is based on experimenting machine learning model inference accuracy with one or more candidates for the at least one compressed quantization level.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] QUANTIZATION OF DATA FOR MACHINE LEARNING MODEL

[0002] TECHNICAL FIELD

[0003] Various example embodiments relate generally to machine learning (ML) and / or artificial intelligence (Al) in association with communication.

[0004] BACKGROUND

[0005] Artificial intelligence (Al) may be broadly defined as getting computers to perform tasks mimicking human brain. Machine learning (ML) is one category of Al techniques: computer algorithms able to automatically improve their performance without explicit programming. Al algorithms were first conceived in the 1950’s but only in recent years Al / ML has become useful in vast area of real-world applications, partly due to advancements in computational power and in providing storage capacity for data.

[0006] Al / ML can help adjust and optimize radio access network (RAN) parameters and configurations using (real-time) monitoring and prediction of network performance, quality, and resource demand. Additionally, AI / ML can identify and diagnose degradation in network performance, as well as provide protection from cyberattacks. AI / ML is usable in energy saving, load balancing, mobility optimization, link adaptation and security just to mention but a few.

[0007] BRIEF DESCRIPTION

[0008] According to some aspects, there is provided the subject matter of the independent claims. Some further aspects are defined in the dependent claims. The embodiments that do not fall under the scope of the claims are to be interpreted as examples useful for understanding the disclosure.

[0009] LIST OF THE DRAWINGS

[0010] In the following, the invention will be described in greater detail with reference to the embodiments and the accompanying drawings, in which

[0011] Figure 1 presents an example of a network to which one or more embodiments are applicable;

[0012] Figure 2 illustrates a flow chart, according to some embodiments;

[0013] Figures 3 show a signalling chart, according to some embodiments; Figure 4 illustrate apparatuses, according to some embodiments. DESCRIPTION OF EMBODIMENTS

[0014] The following embodiments are exemplary. Although the specification may refer to “an”, “one”, or “some” embodiment(s) in several locations of the text, this does not necessarily mean that each reference is made to the same embodiment's), or that a particular feature only applies to a single embodiment. Single features of different embodiments may also be combined to provide other embodiments. For the purposes of the present disclosure, the phrases “at least one of A or B”, “at least one of A and B”, “A and / or B” means (A), (B), or (A and B). For the purposes of the present disclosure, the phrases “A or B” and “A and / or B” means (A), (B), or (A and B). For the purposes of the present disclosure, the phrase “A, B, and / or C” means (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C).

[0015] It shall be understood that although the terms “first” and “second” etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and similarly, a second element could be termed a first element, without departing from the scope of example embodiments.

[0016] Embodiments described may be implemented in or applied to any a radio system, such as one comprising at least one of the following radio access technologies (RATs): Worldwide Interoperability for Micro-wave Access (WiMAX), Global System for Mobile communications (GSM, 2G), GSM EDGE radio access Network (GERAN), General Packet Radio Service (GPRS), Universal Mobile Telecommunication System (UMTS, 3G) based on basic wideband-code division multiple access (W-CDMA), high-speed packet access (HSPA), Long Term Evolution (LTE), LTE-Advanced, and enhanced LTE (eLTE). Term ‘eLTE’ here denotes the LTE evolution that connects to a 5G core. LTE is also known as evolved UMTS terrestrial radio access (EUTRA) or as evolved UMTS terrestrial radio access network (EUTRAN). A term “resource” may refer to radio resources, such as a physical resource block (PRB), a radio frame, a subframe, a time slot, a sub band, a frequency region, a sub-carrier, a beam, etc. The term “transmission” and / or “reception” may refer to wirelessly transmitting and / or receiving via a wireless propagation channel on radio resources.

[0017] The embodiments are not, however, restricted to the systems / RATs given as an example but a person skilled in the art may apply the solution to other communication systems / networks provided with necessary properties. Some examples of a suitable communication networks include a 5G network and / or a 6G network. The 3GPP solution to 5G is referred to as New Radio (NR). 6G is envisaged to be a further development of 5G. NR has been envisaged to use multiple-input- multiple-output (M1M0) multi-antenna transmission techniques, more base stations or nodes than the current network deployments of LTE (a so-called small cell concept), including macro sites operating in co-operation with smaller local area access nodes and perhaps also employing a variety of radio technologies for better coverage and enhanced data rates. 5G will likely be comprised of more than one radio access technology / radio access network (RAT / RAN), each optimized for certain use cases and / or spectrum. 5G mobile communications may have a wider range of use cases and related applications including video streaming, augmented reality, different ways of data sharing and various forms of machine type applications, including vehicular safety, different sensors and real-time control. 5G is expected to have multiple radio interfaces, namely below 6GHz, cmWave and mmWave, and being integrable with existing legacy radio access technologies, such as the LTE.

[0018] As compared to the LTE, low latency applications and services in 5G may require bringing the content close(r) to the radio interface which may lead to local break out and multi-access edge computing (MEC). 5G enables analytics and knowledge generation to occur at the source of the data. This approach requires leveraging resources that may not be continuously connected to a network such as laptops, smartphones, tablets and sensors. MEC provides a distributed computing environment for application and service hosting. It also has the ability to store and process content in close proximity to cellular subscribers for faster response time. Edge computing covers a wide range of technologies such as wireless sensor networks, mobile data acquisition, mobile signature analysis, cooperative distributed peer-to-peer ad hoc networking and processing also classifiable as local cloud / fog computing and grid / mesh computing, dew computing, mobile edge computing, cloudlet, distributed data storage and retrieval, autonomic self-healing networks, remote cloud services, augmented and virtual reality, data caching, Internet of Things (massive connectivity and / or latency critical), critical communications (autonomous vehicles, traffic safety, real-time analytics, time-critical control, healthcare applications). Edge cloud may be brought into RAN by utilizing network function virtualization (NVF) and software defined networking (SDN). Using edge cloud may mean access node operations to be carried out, at least partly, in a server, host or node operationally coupled to a remote radio head or base station comprising radio parts. Network slicing allows multiple virtual networks to be created on top of a common shared physical infrastructure. The virtual networks are then customised to meet the specific needs of applications, services, devices, customers or operators.

[0019] In radio communications, node operations may be carried out, at least partly, in a central / centralized unit, CU, (e.g. server, host or node) operationally coupled to distributed unit, DU, (e.g. a radio head / node). It is also possible that node operations will be distributed among a plurality of servers, nodes or hosts. It should also be understood that the distribution of work between core network operations and base station operations may vary depending on implementation. Thus, 5G networks architecture may be based on a so-called CU-DU split. One gNB- CU controls several gNB-DUs. The term ‘gNB’ may correspond in 5G to the eNB in LTE. The gNBs (one or more) may communicate with one or more UEs. The gNB- CU (central node) may control a plurality of spatially separated gNB-DUs, acting at least as transmit / receive (Tx / Rx) nodes. In some embodiments, however, the gNB- DUs (also called DU) may comprise e.g. a radio link control (RLC), medium access control (MAC) layer and a physical (PHY) layer, whereas the gNB-CU (also called a CU) may comprise the layers above RLC layer, such as a packet data convergence protocol (PDCP) layer, a radio resource control (RRC) and an internet protocol (IP) layers. Other functional splits are possible too. It is considered that skilled person is familiar with the OS1 model and the functionalities within each layer.

[0020] As an example, a server or CU may generate a virtual network through which the server communicates with the radio node. In general, virtual networking may involve a process of combining hardware and software network resources and network functionality into a single, software-based administrative entity, a virtual network. Such virtual network may provide flexible distribution of operations between the server and the radio head / node. In practice, any digital signal processing task may be carried out in either the CU or the DU and the boundary where the responsibility is shifted between the CU and the DU may be selected according to implementation.

[0021] Some other possible technology advancements to be used are Software- Defined Networking (SDN), Big Data, and all-lP, to mention only a few non-limiting examples. For example, network slicing may be a form of virtual network architecture using the same principles behind software defined networking (SDN) and network functions virtualisation (NFV) in fixed networks. SDN and NFV may deliver greater network flexibility by allowing traditional network architectures to be partitioned into virtual elements that may be linked (also through software). Network slicing allows multiple virtual networks to be created on top of a common shared physical infrastructure. The virtual networks are then customised to meet the specific needs of applications, services, devices, customers or operators.

[0022] The plurality of gNBs (access points / network nodes), each comprising a CU and one or more DUs, may be connected to each other via the Xn interface over which the gNBs may negotiate. The gNBs may also be connected over next generation (NG) interfaces to a 5G core network (5GC), which may be a 5G equivalent for the core network of LTE. Such 5G CU-DU split architecture may be implemented using cloud / server so that the CU having higher layers locates in the cloud and the DU is closer to or comprises actual radio and antenna unit. There are similar plans ongoing for LTE / LTE-A / eLTE as well. When both eLTE and 5G will use similar architecture in a same cloud hardware (HW), the next step may be to combine software (SW) so that one common SW controls both radio access networks / technol- ogies (RAN / RAT). This may allow then new ways to control radio resources of both RANs. Furthermore, it may be possible to have configurations where the full protocol stack is controlled by the same HW and handled by the same radio unit as the CU.

[0023] It should also be understood that the distribution of labour between core network operations and base station operations may differ from that of the LTE or even be non-existent. Some other technology advancements probably to be used are Big Data and all-lP, which may change the way networks are being constructed and managed. 5G (or new radio, NR) networks are being designed to support multiple hierarchies, where MEC servers may be placed between the core and the base station or nodeB (gNB). It should be appreciated that MEC may be applied in 4G networks as well.

[0024] 5G may also utilize satellite communication to enhance or complement the coverage of 5G service, for example by providing backhauling. Possible use cases are providing service continuity for machine-to-machine (M2M) or Internet of Things (loT) devices or for passengers on board of vehicles, or ensuring service availability for critical communications, and future railway / maritime / aeronautical communications. Satellite communication may utilize geostationary earth orbit (GEO) satellite systems, but also low earth orbit (LEO) satellite systems, in particular mega-constellations (systems in which hundreds of (nano) satellites are deployed). Each satellite in the mega-constellation may cover several satellite-enabled network entities that create on-ground cells. The on-ground cells may be created through an on-ground relay node or by a gNB located on-ground or in a satellite.

[0025] The embodiments may be also applicable to narrow-band (NB) Inter- net-of-things (loT) systems which may enable a wide range of devices and services to be connected using cellular telecommunications bands. NB-loT is a narrowband radio technology designed for the Internet of Things (loT) and is one of technologies standardized by the 3rd Generation Partnership Project (3GPP). Other 3GPP loT technologies also suitable to implement the embodiments include machine type communication (MTC) and eMTC (enhanced Machine-Type Communication). NB-loT focuses specifically on low cost, long battery life, and enabling a large number of connected devices. The NB-loT technology is deployed “in-band” in spectrum allocated to Long Term Evolution (LTE) - using resource blocks within a normal LTE carrier, or in the unused resource blocks within a LTE carrier’s guard-band - or “standalone” for deployments in dedicated spectrum.

[0026] The embodiments may be also applicable to device-to-device (D2D), machine-to-machine, peer-to-peer (P2P) communications. The embodiments may be also applicable to vehicle-to-vehicle (V2V), vehicle-to-infrastructure (V21), in- frastructure-to-vehicle (12V), or in general to V2X or X2V communications.

[0027] Figure 1 illustrates an example of a communication system to which embodiments of the invention may be applied. The system may comprise a control node 110 providing one or more cells, such as cell 100, and a control node 112 providing one or more other cells, such as cell 102. Each cell may be, e.g., a macro cell, a micro cell, femto, or a pico cell, for example. In another point of view, the cell may define a coverage area or a service area of the corresponding access node. The control node 110, 112 may be an evolved Node B (eNB) as in the LTE and LTE-A, ng-eNB as in eLTE, gNB of 5G, or any other apparatus capable of controlling radio communication and managing radio resources within a cell. The control node 110, 112 may be called a base station, network node, or an access node.

[0028] The system may be a cellular communication system composed of a radio access network of access nodes, each controlling a respective cell or cells. The access node 110 may provide user equipment (UE) 120 (one or more UEs) with wireless access to other networks such as the Internet. The wireless access may comprise downlink (DL) communication from the control node to the UE 120 and uplink (UL) communication from the UE 120 to the control node.

[0029] Additionally, although not shown, one or more local area access nodes may be arranged such that a cell provided by the local area access node at least partially overlaps the cell of the access node 110 and / or 112. The local area access node may provide wireless access within a sub-cell. Examples of the sub-cell may include a micro, pico and / or femto cell. Typically, the sub-cell provides a hot spot within a macro cell. The operation of the local area access node may be controlled by an access node under whose control area the sub-cell is provided. In general, the control node for the small cell may be likewise called a base station, network node, or an access node.

[0030] There may be a plurality of UEs 120, 122 in the system. Each of them may be served by the same or by different control nodes 110, 112. The UEs 120, 122 may communicate with each other, in case D2D communication interface is established between them.

[0031] The term “terminal device” or “UE” refers to any end device that may be capable of wireless communication. By way of example rather than limitation, a terminal device may also be referred to as a communication device, user equipment (UE), a Subscriber Station CSS}, a Portable Subscriber Station, a Mobile Station (MS), or an Access Terminal (AT). The terminal device may include, but not limited to, a mobile phone, a cellular phone, a smart phone, voice over IP (VoIP) phones, wireless local loop phones, a tablet, a wearable terminal device, a personal digital assistant (PDA), portable computers, desktop computer, image capture terminal devices such as digital cameras, gaming terminal devices, music storage and playback appliances, vehicle-mounted wireless terminal devices, wireless endpoints, mobile stations, laptop-embedded equipment (LEE), laptop-mounted equipment (LME), USB dongles, smart devices, wireless customer-premises equipment (CPE), an Internet of Things (loT) device, a watch or other wearable, a head-mounted display (HMD), a vehicle, a drone, a medical device and applications (e.g., remote surgery), an industrial device and applications (e.g., a robot and / or other wireless devices operating in an industrial and / or an automated processing chain contexts), a consumer electronics device, a device operating on commercial and / or industrial wireless networks, and the like. In the following description, the terms “terminal device”, “communication device”, “terminal”, “user equipment” and “UE” may be used interchangeably.

[0032] In the case of multiple access nodes in the communication network, the access nodes may be connected to each other with an interface. LTE specifications call such an interface as X2 interface. For IEEE 802.11 network (i.e. wireless local area network, WLAN, WiFi), a similar interface may be provided between access points. An interface between an LTE access point and a 5G access point, or between two 5G access points may be called Xn. Other communication methods between the access nodes may also be possible. The access nodes 110 and 112 may be further connected via another interface to a core network 116 of the cellular communication system. The LTE specifications specify the core network as an evolved packet core (EPC), and the core network may comprise a mobility management entity (MME) and a gateway node. The MME may handle mobility of terminal devices in a tracking area encompassing a plurality of cells and handle signalling connections between the terminal devices and the core network. The gateway node may handle data routing in the core network and to / from the terminal devices. The 5G specifications specify the core network as a 5G core (5GC), and there the core network may comprise e.g. an access and mobility management function (AMF) and a user plane function / gateway (UPF), to mention only a few. The AMF may handle termination of non-access stratum (NAS) signalling, NAS ciphering & integrity protection, registration management, connection management, mobility management, access authentication and authorization, security context management. The UPF node may support packet routing & forwarding, packet inspection and QoS handling, for example.

[0033] It is envisioned that Al / ML will enable real-time analysis as well as automated operation and control in 5G and beyond radio access network (RAN). This requires the availability of data streamed from wireless devices in a timely manner, especially in extremely time-critical applications such as real-time video monitoring and extended reality (XR). This may be reflected in network architecture, such as by placing and moving ML agents to the required locations in the network, for example for data collection. User devices (mobile devices) may assist network in decision-making in resource management, thus a user device may act as an infrastructure resource.

[0034] As the network evolves to programmable and flexible cloud native implementation, Al / ML-based network automation will be used to simplify network management and optimization. It is expected that parts of the air interface, in particular signal processing algorithms, are supported and eventually even replaced with machine learning models. Thus, a 6G wireless communication standard will natively support an Al-based air interface.

[0035] Machine learning algorithms are usually classified into four different types: supervised learning, unsupervised learning, semi-supervised learning and reinforcement learning.

[0036] In supervised learning the algorithm learns from labelled data. For training, the algorithm receives input data and corresponding correct output labels. The algorithm is trained to predict accurate labels for new data.

[0037] In unsupervised learning the algorithm analyses unlabelled data. The aim is to discover patterns, relationships, or structures within the data. As an example, unsupervised learning algorithms make groups of similar data points.

[0038] Semi-supervised learning is a hybrid machine learning approach that combines labelled and unlabelled data for training. A limited amount of labelled data and a larger set of unlabelled data is used to improve training. This approach is useful when acquiring labelled data is expensive or time-consuming as is the case in many real-world applications. Semi-supervised learning techniques may be applied to various tasks, such as classification, regression, and anomaly detection, allowing models to make more accurate predictions and generalize better in real- world scenarios.

[0039] Reinforcement learning is a machine learning algorithm which learns from trial and error. An ML agent interacts with environment and learns from experience aiming to maximize cumulative rewards. The ML agent receives feedback through rewards or penalties based on its actions. The agent learns to take actions that lead to the most favourable outcomes over time. The algorithm adapts to changing environments, and aims to achieve long-term goals through a sequence of actions.

[0040] An example of an ML algorithm found applicable to adjust and optimize the radio access network parameters and settings is deep learning. Deep learning is a subset of machine learning algorithms using a neural network. N eural networks are also known as artificial neural networks (ANNs) or simulated neural networks (SNNs). Deep learning can be based on supervised, semi-supervised or unsupervised learning.

[0041] Artificial neural networks (ANNs) are comprised of an input layer, one or more hidden layers, and an output layer. Each node of a layer, or an artificial neuron, connects to another one and has an associated weight as well as a threshold value. If the output of an individual node is above the threshold value specified to this node, the node is activated and sending or passing data to the next layer of the neural network.

[0042] In the case the supervised learning is applied in the training of a neural network, the training is carried out by using examples, each of which contains a known "input" and "result", forming probability-weighted associations between them. The training comprises determining the difference between the output of the neural network (a prediction) for an input and a target output for the same input. The difference is called an error value. The neural network then adjusts its weighted associations according to a learning rule and using this error value. Successive adjustments make the neural network produce output that is approaching the target output. After a sufficient number of these adjustments, the training may be terminated based on certain criteria.

[0043] The recent advancements in the field of Al / ML show its potential in finding closer-to-optimal solutions and delivering enhanced performance for radio communication. Several in-network Al / ML functionalities employ measurements, e.g., reference signal received power (RSRP) measurements, as (part of) their input for various use cases. Some examples of these use cases are:

[0044] 1. Beam management, channel state information (CS1) compression and positioning enhancement under study in 3GPP RAN 1 WG.

[0045] 2. Mobility optimization, load balancing and network energy saving under study in 3GPP RAN3 WG. 1

[0046] Some Al / ML-based applications, such as load balancing, that require information from various user devices in same, nearby and / or different locations, such as cells or beams. In these cases, model inference is usually beneficial to be carried out at or by a network element (e.g., gNB, DU, CU, edge server, a user device acting or assisting in network management) that may collect measurement data reports from several user devices (comprising device able to make measurements in radio environment) and has enough computational power to process in many cases massive amount of data. The method and embodiments described herein are applicable to any Al / ML method including supervised, reinforcement and unsupervised learning. Additionally, it should be appreciated that the method and embodiments are suitable for use cases that benefit from data quantization including (but not limited to) beam management, channel state information (CSI) compression and positioning.

[0047] According to the current 3GPP standardization [TS 38.133], UE (user device) measurements, such as each measured RSRP value (LI and L3) are quantized with 7 bits. This fixed quantization does not fit for possible usage for Al / ML procedures in the best possible way, for example the total number of bits transmitted to a network element may increase rapidly when different beams, longer window size for measurements (in time-domain) and several user devices are involved as is the case in many real-world scenarios. This implies non-negligible overhead on energy (affecting UE battery-life) and bandwidth consumption.

[0048] When the reported measurement information is intended for being used as input data to an AI / ML model, the needed level of quantization of the data may need to be varied based on required model performance (e.g., accuracy), feature importance, and / or time or other resources usable for model inference.

[0049] In the following, by means of Figures 2 and 3, a method suitable for adapting a quantization level is disclosed. The method may be carried out by a network element or node, such as gNB, DU, CU, edge server or a user device when it assists in network operation related decision making or network management, or any suitable device. The apparatus may be called a network management device as well.

[0050] In block 200, at least one compressed quantization level is determined for conveying measurement data for inference input of at least one machine learning model, wherein the determining is based on experimenting machine learning model accuracy with one or more candidates for the at least one compressed quantization level.

[0051] For transmission of data samples from one node to another node in a communications network, sample values measured during sampling are quantized to produce a digital representation of the analogue signal that is measured. In short, quantization is a process, where continuous (infinite) values are mapped to a smaller set of discrete (finite) values. The number of these discrete values is the quantization level. As an example, a sampled value is approximated / quantized by the quantization codeword closest to it. Quantization levels maybe pre-determined and expressed in a table format along with a set of codewords. The closeness of the approximation between the sampled value and the quantized value that represents it depends on the number of quantization levels available: as simplified, the more quantization levels are provided, the better the approximation. Each quantization codeword is represented by a unique binary number comprised of bits. The number of quantization levels is thus related to the number of bits representing the quantization levels. The difference between the original and quantized signals is called quantization noise or quantization error. Quantization is an example of lossy encoding / compression: some information is lost when original, e.g., continuous values are approximated by a set with a limited number of codewords, i.e., quantization levels. To perform quantization, at least one quantization level (also referred to as number of clusters or codebook size), one set of codewords (also referred to as codebook or set of centroids) and one forward quantization rule (also known as mapping or assignment rule) are to be determined. These three parts may be designed separately or jointly. Determining the number of quantization levels is related to the rate-distortion theory and may be used directly in a communications system. Therefore, allowing a network node to set quantization levels via configuring one or more user devices translates directly into resource utilization optimization, e.g., radio and storage.

[0052] Scalar quantization (SQ) is taken here as an example of a quantization method. Different techniques may be applied to determine a quantization level, codebook, and forward quantization rule. As an example, a so-called elbow method may be applied to carry out a trial and error experiment with different quantization levels and to evaluate for each trial the ML model performance using an accuracy measure, such as comparing classification accuracies of the ML model when using i) the quantization level under trial and ii) a reference quantization level yielding to a quantization error, below a predefined threshold. Another example is relevance-based bit allocation in which trial and error experiment is carried out using different quantization levels and for each trial the ML model performance is evaluated using a distortion measure such as the Kullback-Leibler divergence or the conditional Kullback-Leibler divergence between two joint distributions over input and output of a ML model once with i) the quantized data using the quantization level under trial and once with ii) a reference quantization level yielding to a quantization error below a threshold. A Kullback-Leibler divergence measures the distance between two distributions. Here, one distribution is a result of using the quantized data with the quantization level under trial and another distribution is a result of using a reference (high-resolution) quantization. The second distribution is referred to as a true distribution or reference distribution). As another example, same methods may be used with other distortion measures capturing ML performance, e.g., other divergence, ML model accuracy for classifiers, etc.

[0053] Scalar quantization may have a uniform or non-uniform set of codewords, also referred to as uniform and non-uniform scalar quantization. Mid-riser and mid-tread uniform quantization are example for uniform codebooks. The dead zone quantization determines a set of non-uniform codewords, where a broad range of continuous values around zero are mapped to codeword zero, e.g., to reduce the impact of noise.

[0054] The forward quantization rule maybe any function capable of mapping continuous / original values to codewords / quantized values. As an example, Euclidean distance between a continuous value to all codewords is measured. The codeword corresponding to the minimum distance is selected to represent that continuous value. As another example, boundaries such as a minimum and maximum values may be defined in the continuous space to separate it into different clusters. Each cluster is represented by a codeword. A codeword is assigned as the quantized value depending on the cluster to which a continuous measurement or sample belongs according to boundaries set.

[0055] According to the current 3GPP standards, only a 7-bits uniform scalar quantization as specified in [TS 38.133] is used. More options could provide better support for a ML model training and inference and reduce signalling and energy consumption overhead. As an example, a user device may have the capability to carry out a uniform quantization using a new quantization level determined by a network node for measurements. Such capability comprises having i) a codebook design method such as mid-tread quantization to define codewords for the determined quantization level (or having the capability to receive and use a codebook from another node) and ii) a forward quantization rule to map the measurements to the codewords. In this case, user device applies the same uniform scalar quantization (SQ) with the (newly) determined quantization level to each of its measurements. As another example, a user device may have the capability to carry out uniform SQs given different quantization levels for different measurements, e.g., it carries out a uniform scalar quantization using a new quantization level determined by the network node for one or a group of its measurements while carrying out another uniform scalar quantization using another new quantization level determined by the network node for another measurement or another group of measurements.

[0056] Vector quantization (VQ) is taken herein as another example of a quantization method. Vector quantization may be seen as a multidimensional extension of a one-dimensional quantization scheme (scalar quantization). VQ involves quantizing more than one continuous or discrete measurement / sample into a finite set of representative vectors, known as a codebook or a set of code vectors or a set of codewords or a set of centroids. Dimensionality / number of elements of both original and quantized vector (code vector) are the same. Elements of the code vectors are quantized values of the input samples. The goal of vector quantization is to find a set of codewords while minimizing the distortion between the input original (e.g., continuous) data and the code vectors, thereby achieving a more compact representation of the data.

[0057] For VQ, different techniques may be applied to determine the quantization level, codebook, and forward quantization rule. As an example, a so-called elbow method may be applied to carry out trial and error experiment with different quantization levels and to evaluate for each trial the ML model performance using an accuracy measure, such as comparing classification accuracies of the ML model when using i) the quantization level under trial and ii) a reference quantization level yielding to a quantization error below a threshold. Another example to determine the quantization level is relevance-based bit allocation which carried out a search over different quantization levels and evaluates for each trial the ML model performance using a distortion measure, e.g., it compares the Kullback-Leibler divergence or the conditional Kullback-Leibler divergence between joint distributions over input and output of a ML model using i) the quantized input with quantization level under trial and ii) a reference quantization level yielding to a quantization error below a threshold. As another example, the same methods may be used with other distortion measures, e.g., other divergence, ML model accuracy for classifiers, etc.

[0058] A method called k-means is an example of VQ. The k-means cluster- ing / VQ method partitions a set of observations from original / continuous measurements (a dataset) into a given number of clusters (the given number of clusters is the quantization level) minimizing a distortion measure such as the sum of squared errors inside clusters (within-cluster sum of squared errors, WCSS). This is carried out by using an iterative algorithm. As an example, a set of codewords is initiated according to the quantization level. As an example, considering squared Euclidean distance as the forward quantization rule, points in the set of observations are mapped to initial codewords (the observations are mapped to the codeword with which they have the minimum squares Euclidean distance). Then, a new set of code vectors is defined, where each codeword is, e.g., the expected value of the points belonging to the same cluster as the code vector. This process is repeated until a termination condition is met, e.g., assignment / mapping of observations to clusters remains unchanged. As an example, cosine similarity is an alternative for Euclidean distance.

[0059] For a reference pattern or quantization level, currently standardized 7- bit quantization level may be used. Any other refence value may also be used.

[0060] In future generations of mobile networks such as 6G, the user device may have the capability to carry out quantization on different groups of its measurements and with different dimensionality (length of the original vector to be quantized or equivalently, length of the codeword) given different new quantization levels. A method for dimensionality reduction may be applied to data before applying the VQ (in general any quantization method). An example of a method for dimensionality reduction is the use of auto-encoders. Another example of a method for dimensionality reduction is principal component analysis.

[0061] After codebook design is finished (either SQ or VQ), during (an online) process of quantizing data, the codebook serves as a compressed representation of the original data. As an example, considering Euclidean distance as forward quantization rule, a data point may be replaced with the index of the closest code vector for transmission. In current communications networks, each index has a binary representation. Receiving entity of the index should be able to translate the index back into the quantized value or vector, i.e., the codeword. It should be ensured that a list of codewords and their corresponding binary representations are available at both user device and network node (or any other entity carrying out or assisting in network management).

[0062] For the purpose of the determination of compressed quantization level, a function or entity that is called a quantization pattern finder (QPF) may be applied. It may locate at a network node, such as gNB. As a general principle, the QPF uses different quantization patterns for experimenting and determining one or more (candidate) quantization patterns, usually ML model-wise, to be used in the network depending on network status as explained in further detail below.

[0063] In this context, a quantization pattern may be thought as a vector. Given that measurements are ordered in a specific way, elements of this vector represents the number of bits used to describe the respective measurement value or a group of measurement values (a certain number of bits represents a quantization level). The QPF experiments machine learning model performance with one or more candidate quantization patterns. Different quantization patterns may lead to different inference output accuracies as presented above.

[0064] In trial experiments, knowledge of the user device capabilities / op- tions, e.g., if it has the capability to do SQ and / or VQ may be used . This is explained more in the following (with regard to block 204).

[0065] In this example, (one or more) quantization pattern are identified which provides the least deteriorated accuracy or which provides a deterioration below a predetermined threshold as described above. Determined compressed quantization levels are those used in a (selected) quantization pattern from (candidate) quantization patterns, as simplified, since other decision criteria may exist, such as availability of radio resources, radio channel conditions, number of user devices served, delay requirements, energy saving and / or data processing capabilities. For example, when access to radio resources is not limited by heavy use (hot spot or rush time), a quantization pattern that uses more resources may be selected as well as in strong radio channel conditions. As another example, in the heavy use situation, a quantization pattern that uses less resources may be selected, as well for energy saving. Some deterioration in the accuracy may be acceptable as a trade-off.

[0066] The process may be carried out offline and only once as long as the ML model parameters obtained by training (e.g., neural network weights and bias values converged after last ML model training / updating event) remain unchanged. Hence, it will not influence or delay the inferences. The determination of the compressed quantization level may be started by detecting a triggering condition, such as (a change in) availability of radio resources, radio channel conditions, number of user devices served, delay requirements, energy saving and / or data processing capabilities. It is also possible to determine different quantization levels cell-spe- cifically, service-specifically, user device -specifically and / or network slice -specifically.

[0067] In response to at least one of the one or more quantization pattern candidates providing a suitable machine learning model accuracy, in block 200, information on a configuration for collecting measurement data is transmitted, block 202, to at least one user device. The configuration comprises the determined at least one compressed quantization level. The configuration may comprise a compressed quantization level for a specific measurement or a group of measurements. In addition, it may comprise any other needed configuration information, such as resource allocation for transmitting the collected measurement data, codebook along with binary representation of the codewords in case the codebook design is done by a network node. The configuration information may be transmitted as a radio resource control signalling. As stated above, the determined compressed quantization level is the one used in this (selected) quantization pattern, as simplified, since other decision criteria may exist, such as availability of radio resources, radio channel conditions, number of user devices served, delay requirements, energy saving and / or data processing capabilities, So the candidate providing the best or a predetermined inference accuracy may not be selected but a candidate providing a suitable, such as good enough accuracy when the circumstances are taken into account, may be selected instead. For example, when access to radio resources is not limited by heavy use (hot spot or rush time), a quantization pattern that uses more resources may be selected as well as in good radio channel conditions. As another example, in the heavy use situation, a quantization pattern that uses less resources may be selected, as well for energy saving. Some deterioration in the inference accuracy may be acceptable as a trade-off. For example, a gNB (DU, CU, edge server) selects one of the quantization patterns suggested by the QPF depending on radio resource status and Quality-of-Service (QoS) requirement of the user device(s). An example:

[0068] • In the case of abundant radio resources, the selected quantization pattern has the same inference accuracy as the 7-bits quantization

[0069] • In the case of scarce bandwidth or low QoS requirements, the selected quantization pattern has the best accuracy that can be served.

[0070] In block 204, a message comprising collected measurement data indicated as quantized using the determined at least one compressed quantization level is received from the at least one user device.

[0071] It should be appreciated that an option exists to transmit to the at least one user device a request for informing options for the compressed quantization level and select at least one of the informed options / capabilities. Examples of the options indicating a user device capability are scalar quantization and vector quantization on all or a subset of user device measurements. User device capabilities may comprise scalar quantization (such as the one currently in 3GPP standards) and / or vector quantization for all or part of measurement values or data. The search may be carried out considering (the informed or otherwise known options / capabilities). As an example, if a user device has the capability to carry out different (non-)uniform scalar quantization on each of its measurements, the QPF works on and selects quantization patterns stating the different quantization levels that should be used for each of the measurement. As an example, if a user device has a capability to carry out VQ on a first group of measurements and different non- uniform scalar quantization on each of the rest of the measurements, the QPF works on and selects quantization patterns with different quantization levels while the first element or quantization level states the quantization level for the VQ and the last elements indicate the quantization levels that should be used for each of the rest of the measurements.

[0072] As stated above, the process in block 200 may be carried out offline or in advance. In this case, measurement data may be received a plurality of times using the same quantization level determined by the network node (selected from QPF candidate quantization patterns).

[0073] By adjusting the quantization level, it is possible to decrease reporting overhead in terms of transmitting fewer number of bits. Additionally, it is possible to decrease energy consumption at the UE (prolonged battery life). Yet additionally, transferring massive amount of “high-resolution” measurement data, especially in dense UE deployments, reduces the availability of radio resources for other services. By adjusting the level of quantization, it is possible to enhance resource usage for services and robustness of the system in case of having scarce radio resources by increasing the probability of packet drop occurrence (note that in some cases if ML input is not delivered in a short period of time, it becomes obsolete, e.g., when making handover decisions).

[0074] In Figure 3, an example of signalling is presented. In this example, network node 300, after deciding to determine the quantization level (one or more) 304, transmits a capability request, 306, for one or more user devices (UEs), only one shown for the sake of clarity, UE 302. Different user devices may support different quantization options / capabilities. Radio conditions, services provided, data processing capabilities etc. may also differ. Therefore, also the quantization levels the network node selects / determines for a user device (or group of devices) may differ user device (or group) -wise.

[0075] UE 302 transmits information on which quantization options it is capable of supporting, 308. Then, network node 300 carries out the determination of quantization level(s), 310, and transmits to the one or more user devices the respective configuration for data collection 312. The configuration may be different for different user devices for many reasons, such as the service provided, radio conditions or the model used. User device 302 carries outthe measurements and quantization of the measurement data as according to the configuration, 314. Then the user device transmits the data as quantized to the network node 300, 316, that uses it as input for ML model inference 318. The network node 300, may also store the data and use it later, for example after user devices have transmitted enough data for inference of the ML model at issue that may take more than one transmission, or when there are enough processing resources. As another case, network node may use processing resources of edge cloud or core, or in the case of DU / CU division, when network node in the radio interface is a DU, CU for storing data and / or inference of the model.

[0076] By means of Figure 4, it is shown an example of an apparatus suitable for carrying outthe processes, embodiments and examples described above. Apparatus 10 comprises a control circuitry (CTRL) 12, such as at least one processor, and at least one memory 14 storing instructions that, when executed by the at least one processor, cause the apparatus at least to carry out any one of the above-described processes, embodiments and examples. The memory may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, flash memory, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The memory may comprise a database for storing data.

[0077] The apparatus 10 may be or comprised in in a network node, or distributed unit of a network node, centralized unit of a network node and / or in a radio unit of a network node, or any other apparatus acting or assisting in network resource management, such as an edge server or a user / terminal device of a communication system, e.g. a user terminal (UT), a computer (PC), a laptop, a tabloid computer, a cellular phone, a mobile phone, a communicator, a smart phone, a palm computer, a mobile transportation apparatus (such as a car), a household appliance, or any other communication apparatus, commonly called as UE in the description. Alternatively, the apparatus is comprised in such a terminal device. Further, the apparatus may be or comprise a module (to be attached to the UE) providing connectivity, such as a plug-in unit, an “USB dongle”, or any other kind of unit. The unit may be installed either inside the UE or attached to the UE with a connector or even wirelessly.

[0078] The apparatus may further comprise or have access to communication interface (TRX) or radio unit / head 16 comprising hardware and / or software for realizing communication connectivity according to one or more communication protocols. The TRX may provide the apparatus with communication capabilities with at least one user equipment, for example.

[0079] The apparatus may also comprise a user interface 18 comprising, for example, at least one keypad, a microphone, a touch display, a display, a speaker, etc. The user interface may be used to control the apparatus by the user.

[0080] The control circuitry 12 may comprise a circuitry 60 for determining quantization level(s), according to any of the embodiments.

[0081] It should be understood that for CU-DU (central unit - distributed unit) architecture the apparatus 10 may be or be comprised in a central unit (e.g. a control unit, an edge cloud server, a server) operatively coupled (e.g. via a wireless or wired network) to a distributed unit (e.g. a remote radio head / node). That is, the central unit (e.g. an edge cloud server) and the radio node may be stand-alone apparatuses communicating with each other via a radio path or via a wired connection. Alternatively, they may be in a same entity communicating via a wired connection, etc. The edge cloud or edge cloud server may serve a plurality of radio nodes or a radio access networks. In an embodiment, at least some of the described processes may be carried out by the central unit. In another embodiment, the apparatus may be instead comprised in the distributed unit, and at least some of the described processes may be carried out by the distributed unit. In an embodiment, the execution of at least some of the functionalities of the apparatus 10 may be shared between two physically separate devices (DU and CU) forming one operational entity. Therefore, the apparatus may be seen to depict the operational entity comprising one or more physically separate devices for executing at least some of the described processes. In an embodiment, the apparatus controls the execution of the processes, regardless of the location of the apparatus and regardless of where the processes / functions are carried out.

[0082] As an example of an apparatus suitable for carrying out the embodiments and examples described herein, it is put forward an apparatus, comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: determine at least one compressed quantization level for conveying measurement data for inference input of at least one machine learning model, wherein the determining is based on experimenting machine learning model inference accuracy with one or more candidates for the at least one compressed quantization level, in response to at least one of the one or more candidates providing a suitable machine learning model inference accuracy: transmit, to at least one user device, information on a configuration for collecting measurement data, wherein the configuration comprises the determined at least one compressed quantization level, and receive, from the at least one user device, a message comprising collected measurement data indicated as quantized using the determined at least one compressed quantization level.

[0083] As another example, an apparatus suitable for carrying out the embodiments and examples described herein, it is put forward an apparatus comprising: means (12, 14, 20) for determining at least one compressed quantization level for conveying measurement data for inference input of at least one machine learning model, wherein the determining is based on experimenting machine learning model inference accuracy with one or more candidates for the at least one compressed quantization level, in response to at least one of the one or more candidates providing a suitable machine learning model inference accuracy: means (16) for transmitting, to at least one user device, information on a configuration for collecting measurement data, wherein the configuration comprises the determined at least one compressed quantization level, and means (16) for receiving, from the at least one user device, a message comprising collected measurement data indicated as quantized using the determined at least one compressed quantization level.

[0084] As used in this application, the term ‘circuitry’ refers to all of the following: (a) hardware-only circuit implementations, such as implementations in only analog and / or digital circuitry, and (b) combinations of circuits and soft- ware (and / or firmware), such as (as applicable): (i) a combination of processor(s) or (ii) portions of processor(s) / software including digital signal processor(s), software, and memory(ies) that work together to cause an apparatus to carry out various functions, and (c) circuits, such as a microprocessor(s) or a portion of a micropro- cessor(s), that require software or firmware for operation, even if the software or firmware is not physically present. This definition of ‘circuitry’ applies to all uses of this term in this application. As a further example, as used in this application, the term ‘circuitry’ would also cover an implementation of merely a processor (or multiple processors) or a portion of a processor and its (or their) accompanying software and / or firmware. The term ‘circuitry’ would also cover, for example and if applicable to the particular element, a baseband integrated circuit or applications processor integrated circuit for a mobile phone or a similar integrated circuit in a server, a cellular network device, or another network device.

[0085] In an embodiment, at least some of the processes described may be carried out by an apparatus comprising corresponding means for carrying out at least some of the described processes. Some example means for carrying out the processes may include at least one of the following: detector, processor (including dual-core and multiple-core processors), digital signal processor, controller, receiver, transmitter, encoder, decoder, memory, RAM, ROM, software, firmware, display, user interface, display circuitry, user interface circuitry, user interface software, display software, circuit, antenna, antenna circuitry, and circuitry.

[0086] A term non-transitory, as used herein, is a limitation of the medium itself (i.e. tangible, not a signal) as opposed to a limitation on data storage persistency (e.g. RAM vs. ROM).

[0087] As used herein the term “means” is to be construed in singular form, i.e. referring to a single element, or in plural form, i.e. referring to a combination of single elements. Therefore, terminology “means for [carrying out A, B, C]”, is to be interpreted to cover an apparatus in which there is only one means for carrying out A, B and C, or where there are separate means for carrying out A, B and C, or partially or fully overlapping means for carrying out A, B, C. Further, terminology “means for carrying out A, means for carrying out B, means for performing C” is to be interpreted to cover an apparatus in which there is only one means for carrying out A, B and C, or where there are separate means for carrying out A, B and C, or partially or fully overlapping means for carrying out A, B, C.

[0088] The techniques and methods described herein may be implemented by various means. For example, these techniques may be implemented in hardware (one or more devices), firmware (one or more devices), software (one or more modules), or combinations thereof. For a hardware implementation, the apparatuses) of embodiments may be implemented within one or more applicationspecific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, micro-controllers, microprocessors, other electronic units designed to carry out the functions described herein, or a combination thereof. For firmware or software, the implementation may be carried out through modules of at least one chip set (e.g. procedures, functions, and so on) that carry out the functions described herein. The software codes may be stored in a memory unit and executed by processors. The memory unit may be implemented within the processor or externally to the processor. In the latter case, it maybe communicatively coupled to the processor via various means, as is known in the art. Additionally, the components of the systems described herein may be rearranged and / or complemented by additional components in order to facilitate the achievements of the various aspects, etc., described with regard thereto, and they are not limited to the precise configurations set forth in the given figures, as will be appreciated by one skilled in the art.

[0089] Embodiments as described may also be carried out in the form of a computer process defined by a computer program or portions thereof. Embodiments of the methods described may be carried out by executing at least one portion of a computer program comprising corresponding instructions. The computer program may be in source code form, object code form, or in some intermediate form, and it may be stored in some sort of carrier, which may be any entity or device capable of carrying the program. For example, the computer program may be stored on a computer program distribution medium readable by a computer or a processor. The computer program medium may be, for example but not limited to, a record medium, computer memory, read-only memory, electrical carrier signal, telecommunications signal, and software distribution package, for example. The computer program medium may be a non-transitory medium. Coding of software for carrying out the embodiments as shown and described is well within the scope of a person of ordinary skill in the art.

[0090] Clauses:

[0091] Clause 1: An apparatus, comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: determine at least one compressed quantization level for conveying measurement data for inference input of at least one machine learning model, wherein the determining is based on experimenting machine learning model inference accuracy with one or more candidates for the at least one compressed quantization level, in response to at least one of the one or more candidates providing a suitable machine learning model inference accuracy: transmit, to at least one user device, information on a configuration for quantizing measurement data, wherein the configuration comprises the determined at least one compressed quantization level, and receive, from the at least one user device, a message comprising collected measurement data indicated as quantized using the determined at least one compressed quantization level.

[0092] Clause 2: The apparatus of clausel, further comprising: detect a triggering condition for carrying out the determining the compressed quantization level, comprising: availability of radio resources, radio channel conditions, number of user devices served, delay requirements, energy saving and / or data processing capabilities.

[0093] Clause 3: The apparatus of clause 1 or 2, further comprising: transmit, to the at least one user device, a request for informing options for the compressed quantization level; receive, from the at least one user device, information on the options, and select at least one of the options to determine the at least one candidate for the compressed quantization level.

[0094] Clause 4: The apparatus of any preceding clause, wherein the determining the at least one compressed quantization level further comprises taking into consideration availability of radio resources, radio channel conditions, number of user devices served, delay requirements, energy saving and / or data processing and storage capabilities.

[0095] Clause 5. The apparatus of any preceding clause, wherein a compressed quantization level allowing the highest number of bits is used as a reference candidate in the experimenting the machine learning model inference accuracy.

[0096] Clause 6. The apparatus of clause 5, wherein the compressed quantization level allowing the highest number of bits is a value defined in a standard.

[0097] Clause 7. The apparatus of any preceding clause, wherein the determining the at least one compressed quantization level further comprises determining the at least one compressed quantization level cell-specifically, service-specifically, user device -specifically, network slice -specifically, and / or machine learning model -specifically.

[0098] Clause 8. The apparatus of any of the preceding clause wherein the apparatus is or is comprised in a network node, or distributed unit of a network node, centralized unit of a network node and / or in a radio unit of a network node, or any other apparatus acting or assisting in network resource management.

[0099] Clause 9: A method, comprising: determining at least one compressed quantization level for conveying measurement data for inference input of at least one machine learning model, wherein the determining is based on experimenting machine learning model inference accuracy with one or more candidates for the at least one compressed quantization level, in response to at least one of the one or more candidates providing a suitable machine learning model inference accuracy: transmitting, to at least one user device, information on a configuration for collecting measurement data, wherein the configuration comprises the determined at least one compressed quantization level, and receiving, from the at least one user device, a message comprising collected measurement data indicated as quantized using the determined at least one compressed quantization level. Clause 10: The method of clause 9, further comprising: detecting a triggering condition for carrying out the determining the compressed quantization level, comprising: availability of radio resources, radio channel conditions, number of user devices served, delay requirements, energy saving and / or data processing capabilities.

[0100] Clause 11: The method of clauses 9 or 10, further comprising: transmitting, to the at least one user device, a request for informing options for the compressed quantization level; receiving, from the at least one user device, information on the options, and selecting at least one of the options to determine the at least one candidate for the compressed quantization level.

[0101] Clause 12: The method of any preceding clause 9 to 11, wherein the determining the at least one compressed quantization level further comprises taking into consideration availability of radio resources, radio channel conditions, number of user devices served, delay requirements, energy saving and / or data processing capabilities.

[0102] Clause 13: The method of any preceding clause 9 to 12, wherein a compressed quantization level allowing the highest number of bits is used as a reference candidate in the experimenting the machine learning model inference accuracy.

[0103] Clause 14: The method of clause 13, wherein the compressed quantization level allowing the highest number of bits is a value defined in a standard.

[0104] Clause 15: The method of any preceding clause 9 to 14, wherein the determining the at least one compressed quantization level further comprises determining the at least one compressed quantization level cell-specifically, service- specifically, user device -specifically, network slice -specifically, and / or machine learning model -specifically.

[0105] Clause 16: A computer program product comprising program instructions which, when the program is executed by an apparatus, cause the apparatus to carry out the method according to any of clauses 9 to 15.

[0106] Clause 17: An apparatus, comprising means for carrying out the method according to any of clauses 9 tol5.

Claims

CLAIMS1. An apparatus, comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: determine at least one compressed quantization level for conveying measurement data for inference input of at least one machine learning model, wherein the determining is based on experimenting machine learning model inference accuracy with one or more candidates for the at least one compressed quantization level, in response to at least one of the one or more candidates providing a suitable machine learning model inference accuracy: transmit, to at least one user device, information on a configuration for collecting measurement data, wherein the configuration comprises the determined at least one compressed quantization level, and receive, from the at least one user device, a message comprising collected measurement data indicated as quantized using the determined at least one compressed quantization level.

2. The apparatus of claim 1, further comprising: detect a triggering condition for carrying out the determining the compressed quantization level, comprising: availability of radio resources, radio channel conditions, number of user devices served, delay requirements, energy saving and / or data processing capabilities.

3. The apparatus of claim 1 or 2, further comprising: transmit, to the at least one user device, a request for informing options for the compressed quantization level; receive, from the at least one user device, information on the options, and select at least one of the options to determine the at least one candidate for the compressed quantization level.

4. The apparatus of any preceding claim, wherein the determining the at least one compressed quantization level further comprises taking intoconsideration availability of radio resources, radio channel conditions, number of user devices served, delay requirements, energy saving and / or data processing capabilities.

5. The apparatus of any preceding claim, wherein a compressed quantization level allowing the highest number of bits is used as a reference candidate in the experimenting the machine learning model inference accuracy.

6. The apparatus of claim 5, wherein the compressed quantization level allowing the highest number of bits is a value defined in a standard.

7. The apparatus of any preceding claim, wherein the determining the at least one compressed quantization level further comprises determining the at least one compressed quantization level cell-specifically, service-specifically, user device -specifically, network slice -specifically, and / or machine learning model - specifically.

8. The apparatus of any of the preceding claim wherein the apparatus is or is comprised in a network node, or distributed unit of a network node, centralized unit of a network node and / or in a radio unit of a network node, or any other apparatus acting or assisting in network resource management.

9. A method, comprising: determining at least one compressed quantization level for conveying measurement data for inference input of at least one machine learning model, wherein the determining is based on experimenting machine learning model inference accuracy with one or more candidates for the at least one compressed quantization level, in response to at least one of the one or more candidates providing a suitable machine learning model inference accuracy: transmitting, to at least one user device, information on a configuration for collecting measurement data, wherein the configuration comprises the determined at least one compressed quantization level, and receiving, from the at least one user device, a message comprising collected measurement data indicated as quantized using the determined at least one compressed quantization level.

10. The method of claim 9, further comprising: detecting a triggering condition for carrying out the determining the compressed quantization level, comprising: availability of radio resources, radio channel conditions, number of user devices served, delay requirements, energy saving and / or data processing capabilities.

11. The method of claim 9 or 10, further comprising: transmitting, to the at least one user device, a request for informing options for the compressed quantization level; receiving, from the at least one user device, information on the options, and selecting at least one of the options to determine the at least one candidate for the compressed quantization level.

12. The method of any preceding claim 9 to 11, wherein the determining the at least one compressed quantization level further comprises taking into consideration availability of radio resources, radio channel conditions, number of user devices served, delay requirements, energy saving and / or data processing capabilities.

13. The method of any preceding claim 9 to 12, wherein a compressed quantization level allowing the highest number of bits is used as a reference candidate in the experimenting the machine learning model inference accuracy.

14. The method of claim 13, wherein the compressed quantization level allowing the highest number of bits is a value defined in a standard.

15. The method of any preceding claim 9 to 14, wherein the determining the at least one compressed quantization level further comprises determining the at least one compressed quantization level cell-specifically, service-specifically, user device -specifically, network slice -specifically, and / or machine learning model -specifically.

16. A computer program product comprising program instructionswhich, when the program is executed by an apparatus, cause the apparatus to carry out the method according to any of claims 9 to 15.

17. An apparatus, comprising means for carrying out the method ac- cording to any of claims 9 to 15.