Reporting of inference accuracies of quantized models for artificial intelligence or machine learning performance monitoring

By reporting inference accuracy levels of quantized AI/ML models, UEs enable network nodes to optimize AI/ML performance monitoring, addressing inefficiencies caused by unawareness of varying model accuracies across different UEs.

WO2025208320A1PCT designated stage Publication Date: 2025-10-09QUALCOMM INC +3
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/085460
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-02
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Network nodes are unaware of the varying inference accuracy levels of quantized AI/ML models deployed at different user equipment (UEs), leading to inefficient AI/ML performance monitoring.

Method used

UEs report an indication of the inference accuracy level of their quantized AI/ML models relative to a floating-point model, enabling network nodes to adjust AI/ML performance monitoring parameters accordingly.

Benefits of technology

Enhances the efficiency of AI/ML performance monitoring by allowing network nodes to tailor monitoring strategies based on the specific accuracy levels of quantized models used by UEs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024085460_09102025_PF_FP_ABST
    Figure CN2024085460_09102025_PF_FP_ABST
Patent Text Reader

Abstract

Various aspects of the present disclosure generally relate to wireless communication. In some aspects, a user equipment (UE) may transmit, to a network node, an indication of an inference accuracy level associated with a quantized artificial intelligence or machine learning (AI / ML) model for an AI / ML task, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to a floating-point AI / ML model for the AI / ML task. The UE may receive, from the network node, signaling associated with AI / ML performance monitoring based at least in part on the indication of the inference accuracy level. Numerous other aspects are described.
Need to check novelty before this filing date? Find Prior Art

Description

REPORTING OF INFERENCE ACCURACIES OF QUANTIZED MODELS FOR ARTIFICIAL INTELLIGENCE OR MACHINE LEARNING PERFORMANCE MONITORING

[0001] FIELD OF THE DISCLOSURE

[0002] Aspects of the present disclosure generally relate to wireless communication and specifically relate to techniques, apparatuses, and methods for reporting of inference accuracies of quantized models for artificial intelligence or machine learning performance monitoring.BACKGROUND

[0003] Wireless communication systems are widely deployed to provide various services that may include carrying voice, text, messaging, video, data, and / or other traffic. The services may include unicast, multicast, and / or broadcast services, among other examples. Typical wireless communication systems may employ multiple-access radio access technologies (RATs) capable of supporting communication with multiple users by sharing available system resources (for example, time domain resources, frequency domain resources, spatial domain resources, and / or device transmit power, among other examples) . Examples of such multiple-access RATs include code division multiple access (CDMA) systems, time division multiple access (TDMA) systems, frequency division multiple access (FDMA) systems, orthogonal frequency division multiple access (OFDMA) systems, single-carrier frequency division multiple access (SC-FDMA) systems, and time division synchronous code division multiple access (TD-SCDMA) systems.

[0004] The above multiple-access RATs have been adopted in various telecommunication standards to provide common protocols that enable different wireless communication devices to communicate on a municipal, national, regional, or global level. An example telecommunication standard is New Radio (NR) . NR, which may also be referred to as 5G, is part of a continuous mobile broadband evolution promulgated by the Third Generation Partnership Project (3GPP) . NR (and other mobile broadband evolutions beyond NR) may be designed to better support Internet of things (IoT) and reduced capability device deployments, industrial connectivity, millimeter wave (mmWave) expansion, licensed and unlicensed spectrum access, non-terrestrial network (NTN) deployment, sidelink and other device-to-device direct communication technologies (for example, cellular vehicle-to-everything (CV2X)  communication) , massive multiple-input multiple-output (MIMO) , disaggregated network architectures and network topology expansions, multiple-subscriber implementations, high-precision positioning, and / or radio frequency (RF) sensing, among other examples. As the demand for mobile broadband access continues to increase, further improvements in NR may be implemented, and other radio access technologies such as 6G may be introduced, to further advance mobile broadband evolution.SUMMARY

[0005] Some aspects described herein relate to a user equipment (UE) for wireless communication. The UE may include one or more memories and one or more processors coupled to the one or more memories. The one or more processors may be configured to transmit, to a network node, an indication of an inference accuracy level associated with a quantized artificial intelligence or machine learning (AI / ML) model for an AI / ML task, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to a floating-point AI / ML model for the AI / ML task. The one or more processors may be configured to receive, from the network node, signaling associated with AI / ML performance monitoring based at least in part on the indication of the inference accuracy level.

[0006] Some aspects described herein relate to a network node for wireless communication. The network node may include one or more memories and one or more processors coupled to the one or more memories. The one or more processors may be configured to receive an indication of an inference accuracy level associated with a quantized AI / ML model to be used by a UE for an AI / ML task, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to a floating-point AI / ML model for the AI / ML task. The one or more processors may be configured to transmit, to the UE, signaling associated with AI / ML performance monitoring based at least in part on the indication of the inference accuracy level.

[0007] Some aspects described herein relate to a server device for wireless communication. The server device may include one or more memories and one or more processors coupled to the one or more memories. The one or more processors may be configured to transmit, to a UE, a quantized AI / ML model for an AI / ML task, wherein the quantized AI / ML model is associated with a floating-point AI / ML model for the  AI / ML task. The one or more processors may be configured to transmit an indication of an inference accuracy level associated with the quantized AI / ML model, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to the floating-point AI / ML model.

[0008] Some aspects described herein relate to a method of wireless communication performed by a UE. The method may include transmitting, to a network node, an indication of an inference accuracy level associated with a quantized AI / ML model for an AI / ML task, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to a floating-point AI / ML model for the AI / ML task. The method may include receiving, from the network node, signaling associated with AI / ML performance monitoring based at least in part on the indication of the inference accuracy level.

[0009] Some aspects described herein relate to a method of wireless communication performed by a network node. The method may include receiving an indication of an inference accuracy level associated with a quantized AI / ML model to be used by a UE for an AI / ML task, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to a floating-point AI / ML model for the AI / ML task. The method may include transmitting, to the UE, signaling associated with AI / ML performance monitoring based at least in part on the indication of the inference accuracy level.

[0010] Some aspects described herein relate to a method of wireless communication performed by a server device. The method may include transmitting, to a UE, a quantized AI / ML model for an AI / ML task, wherein the quantized AI / ML model is associated with a floating-point AI / ML model for the AI / ML task. The method may include transmitting an indication of an inference accuracy level associated with the quantized AI / ML model, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to the floating-point AI / ML model.

[0011] Some aspects described herein relate to a non-transitory computer-readable medium that stores a set of instructions for wireless communication by a UE. The set of instructions, when executed by one or more processors of the UE, may cause the UE to transmit, to a network node, an indication of an inference accuracy level associated with a quantized AI / ML model for an AI / ML task, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to a  floating-point AI / ML model for the AI / ML task. The set of instructions, when executed by one or more processors of the UE, may cause the UE to receive, from the network node, signaling associated with AI / ML performance monitoring based at least in part on the indication of the inference accuracy level.

[0012] Some aspects described herein relate to a non-transitory computer-readable medium that stores a set of instructions for wireless communication by a network node. The set of instructions, when executed by one or more processors of the network node, may cause the network node to receive an indication of an inference accuracy level associated with a quantized AI / ML model to be used by a UE for an AI / ML task, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to a floating-point AI / ML model for the AI / ML task. The set of instructions, when executed by one or more processors of the network node, may cause the network node to transmit, to the UE, signaling associated with AI / ML performance monitoring based at least in part on the indication of the inference accuracy level.

[0013] Some aspects described herein relate to a non-transitory computer-readable medium that stores a set of instructions for wireless communication by a server device. The set of instructions, when executed by one or more processors of the server device, may cause the server device to transmit, to a UE, a quantized AI / ML model for an AI / ML task, wherein the quantized AI / ML model is associated with a floating-point AI / ML model for the AI / ML task. The set of instructions, when executed by one or more processors of the server device, may cause the server device to transmit an indication of an inference accuracy level associated with the quantized AI / ML model, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to the floating-point AI / ML model.

[0014] Some aspects described herein relate to an apparatus for wireless communication. The apparatus may include means for transmitting, to a network node, an indication of an inference accuracy level associated with a quantized AI / ML model for an AI / ML task, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to a floating-point AI / ML model for the AI / ML task. The apparatus may include means for receiving, from the network node, signaling associated with AI / ML performance monitoring based at least in part on the indication of the inference accuracy level.

[0015] Some aspects described herein relate to an apparatus for wireless communication. The apparatus may include means for receiving an indication of an inference accuracy level associated with a quantized AI / ML model to be used by a UE for an AI / ML task, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to a floating-point AI / ML model for the AI / ML task. The apparatus may include means for transmitting, to the UE, signaling associated with AI / ML performance monitoring based at least in part on the indication of the inference accuracy level.

[0016] Some aspects described herein relate to an apparatus for wireless communication. The apparatus may include means for transmitting, to a UE, a quantized AI / ML model for an AI / ML task, wherein the quantized AI / ML model is associated with a floating-point AI / ML model for the AI / ML task. The apparatus may include means for transmitting an indication of an inference accuracy level associated with the quantized AI / ML model, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to the floating-point AI / ML model.

[0017] Aspects of the present disclosure may generally be implemented by or as a method, apparatus, system, computer program product, non-transitory computer-readable medium, user equipment, base station, network node, network entity, wireless communication device, and / or processing system as substantially described with reference to, and as illustrated by, the specification and accompanying drawings.

[0018] The foregoing paragraphs of this section have broadly summarized some aspects of the present disclosure. These and additional aspects and associated advantages will be described hereinafter. The disclosed aspects may be used as a basis for modifying or designing other aspects for carrying out the same or similar purposes of the present disclosure. Such equivalent aspects do not depart from the scope of the appended claims. Characteristics of the aspects disclosed herein, both their organization and method of operation, together with associated advantages, will be better understood from the following description when considered in connection with the accompanying drawings.

[0019] While aspects and embodiments are described in this application by illustration to some examples, those skilled in the art will understand that additional implementations and use cases may come about in many different arrangements and scenarios. Innovations described herein may be implemented across many differing  platform types, devices, systems, shapes, sizes, packaging arrangements. For example, embodiments and / or uses may come about via integrated chip embodiments and other non-module-component based devices (e.g., end-user devices, vehicles, communication devices, computing devices, industrial equipment, retail / purchasing devices, medical devices, AI-enabled devices, etc. ) . While some examples may or may not be specifically directed to use cases or applications, a wide assortment of applicability of described innovations may occur. Implementations may range in spectrum from chip-level or modular components to non-modular, non-chip-level implementations and further to aggregate, distributed, or original equipment manufacturer (OEM) devices or systems incorporating one or more aspects of the described innovations. In some practical settings, devices incorporating described aspects and features may also necessarily include additional components and features for implementation and practice of claimed and described embodiments. For example, transmission and reception of wireless signals necessarily includes a number of components for analog and digital purposes (e.g., hardware components including antennas, RF-chains, power amplifiers, modulators, buffers, processor (s) , interleavers, adders / summers, etc. ) . It is intended that innovations described herein may be practiced in a wide variety of devices, chip-level components, systems, distributed arrangements, end-user devices, etc. of varying sizes, shapes, and constitution.BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The appended drawings illustrate some aspects of the present disclosure, but are not limiting of the scope of the present disclosure because the description may enable other aspects. Each of the drawings is provided for purposes of illustration and description, and not as a definition of the limits of the claims. The same or similar reference numbers in different drawings may identify the same or similar elements.

[0021] Fig. 1 is a diagram illustrating an example of a wireless communication network, in accordance with the present disclosure.

[0022] Fig. 2 is a diagram illustrating an example network node in communication with an example user equipment (UE) in a wireless network, in accordance with the present disclosure.

[0023] Fig. 3 is a diagram illustrating an example disaggregated base station architecture, in accordance with the present disclosure.

[0024] Fig. 4 is a diagram illustrating an example architecture of a functional framework for radio access network intelligence enabled by data collection, in accordance with the present disclosure.

[0025] Fig. 5 is a diagram illustrating an example of an artificial intelligence or machine learning (AI / ML) based beam management, in accordance with the present disclosure.

[0026] Fig. 6 is a diagram illustrating an example AI / ML architecture of a first wireless device in communication with a second wireless device, in accordance with the present disclosure.

[0027] Figs. 7-8 are diagrams illustrating examples associated with reporting of inference accuracies of quantized models for AI / ML performance monitoring, in accordance with the present disclosure.

[0028] Fig. 9 is a diagram illustrating an example process performed, for example, at a UE or an apparatus of a UE, in accordance with the present disclosure.

[0029] Fig. 10 is a diagram illustrating an example process performed, for example, at a network node or an apparatus of a network node, in accordance with the present disclosure.

[0030] Fig. 11 is a diagram illustrating an example process performed, for example, at a server device or an apparatus of a server device, in accordance with the present disclosure.

[0031] Figs. 12-14 are diagrams illustrating example apparatuses for wireless communication, in accordance with the present disclosure.DETAILED DESCRIPTION

[0032] Various aspects of the present disclosure are described hereinafter with reference to the accompanying drawings. However, aspects of the present disclosure may be embodied in many different forms and is not to be construed as limited to any specific aspect illustrated by or described with reference to an accompanying drawing or otherwise presented in this disclosure. Rather, these aspects are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art. One skilled in the art may appreciate that the scope of the disclosure is intended to cover any aspect of the disclosure disclosed herein, whether implemented independently of or in combination with any other aspect of the disclosure. For example, an apparatus may be implemented or a method may be  practiced using various combinations or quantities of the aspects set forth herein. In addition, the scope of the disclosure is intended to cover an apparatus having, or a method that is practiced using, other structures and / or functionalities in addition to or other than the structures and / or functionalities with which various aspects of the disclosure set forth herein may be practiced. Any aspect of the disclosure disclosed herein may be embodied by one or more elements of a claim.

[0033] Several aspects of telecommunication systems will now be presented with reference to various methods, operations, apparatuses, and techniques. These methods, operations, apparatuses, and techniques will be described in the following detailed description and illustrated in the accompanying drawings by various blocks, modules, components, circuits, steps, processes, or algorithms (collectively referred to as “elements” ) . These elements may be implemented using hardware, software, or a combination of hardware and software. Whether such elements are implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system.

[0034] Certain aspects and techniques as described herein may be implemented, at least in part, using an artificial intelligence (AI) program, such as a program that includes a machine learning (ML) or artificial neural network (ANN) model. An example AI or ML (AI / ML) model may include mathematical representations or define computing capabilities for making inferences from input data based on patterns or relationships identified in the input data. As used herein, the term “inferences” can include one or more of decisions, predictions, determinations, or values, which may represent outputs of the AI / ML model. The computing capabilities may be defined in terms of certain parameters of the AI / ML model, such as weights and biases. Weights may indicate relationships between certain input data and certain outputs of the AI / ML model, and biases are offsets which may indicate a starting point for outputs of the AI / ML model. An example AI / ML model operating on input data may start at an initial output based on the biases and then update its output based on a combination of the input data and the weights.

[0035] AI / ML models may be deployed in one or more devices (for example, network nodes and / or user equipments (UEs) ) and may be configured to enhance various aspects of a wireless communication system. For example, an AI / ML model may be trained to identify patterns or relationships in data corresponding to a network, a device, an air interface, or the like. An AI / ML model may support operational decisions relating to  one or more aspects associated with wireless communications devices, networks, or services. For example, an AI / ML model may be utilized for supporting or improving aspects such as signal coding / decoding, network routing, energy conservation, transceiver circuitry controls, frequency synchronization, timing synchronization, channel state estimation, channel equalization, channel state feedback, modulation, demodulation, device positioning, beamforming, load balancing, operations and management functions, security, etc.

[0036] In some examples, an AI / ML model may be trained for an AI / ML task at a network node or a server device (e.g., a model server) and transferred to one or more UEs to enable the UEs to use the AI / ML model to perform inference (e.g., prediction) for the AI / ML task. In order to increase efficiency (e.g., to enable an AI / ML model to be deployed at a UE) , a floating-point AI / ML model that is trained for an AI / ML task may be quantized resulting in a quantized AI / ML model for the AI / ML task. “Quantization” is a model size reduction technique that reduces the number of bits used to represent model parameters. For example, quantization may convert model weights (and / or other model parameters) from a high-precision (e.g., floating-point) representation to a lower-precision representation (e.g., a fixed value representation or an integer value representation) . Quantization of a model (e.g., a floating-point model) may reduce the memory usage and computational complexity the model.

[0037] In some examples, different UEs may have different UE-specific model quantization preferences or restrictions. For example, different UE AI / ML hardware accelerators may be associated with different quantization preferences or restrictions. This may lead to different UEs deploying different quantized AI / ML models, with different inference accuracy performance levels, for a same floating-point AI / ML model for an AI / ML task. A network node may perform AI / ML performance monitoring for the UEs deploying the different quantized AI / ML models for the AI / ML task. For example, the network node may schedule transmission of auxiliary reference signals (RSs) for prediction accuracy verification and / or configure inference error threshold values / number for triggering transmission of a performance alert message, among other examples. However, the network node may not be aware of the different model inference accuracy levels for the quantized AI / ML models deployed at the different UEs. This may lead to less efficient AI / ML performance monitoring by the network node, such as using too aggressive or conservative AI / ML performance monitoring parameters for some UEs.

[0038] Various aspects relate generally to reporting of inference accuracies of quantized AI / ML models. Some aspects more specifically relate to reporting of inference accuracy levels for quantized AI / ML models deployed at UEs for AI / ML performance monitoring. In some aspects, a UE may transmit, to a network node, an indication of an inference accuracy level associated with a quantized AI / ML model for an AI / ML task. The inference accuracy level may be indicative of a relative accuracy of the quantized AI / ML model with respect to a floating-point AI / ML model for the AI / ML task. The network node may transmit, and the UE may receive, signaling associated with AI / ML performance monitoring based at least in part on the indication of the inference accuracy level.

[0039] In some aspects, a server device may transmit, to a UE, a quantized AI / ML model to be used by the UE for an AI / ML task. The quantized AI / ML model may be associated with a floating-point AI / ML model for the AI / ML task. In some aspects, the server device may transmit, to a network node, an indication of an inference accuracy level associated with the quantized AI / ML model to be used by the UE for the AI / ML task. In some examples, the network node may transmit, and the UE may receive, signaling associated with AI / ML performance monitoring based at least in part on the indication of the inference accuracy level received at the network node from the server device.

[0040] Particular aspects of the subject matter described in this disclosure can be implemented to realize one or more of the following potential advantages. In some examples, by reporting, from the UE or the served device, the inference accuracy level of the quantized AI / ML model to be used by the UE, the described techniques can be used to make the network node aware of inference accuracy levels of different quantized AI / ML models used by different UEs, which may increase efficiency of AI / ML performance monitoring performed by the network node.

[0041] Multiple-access radio access technologies (RATs) have been adopted in various telecommunication standards to provide common protocols that enable wireless communication devices to communicate on a municipal, enterprise, national, regional, or global level. For example, 5G New Radio (NR) is part of a continuous mobile broadband evolution promulgated by the Third Generation Partnership Project (3GPP) . 5G NR supports various technologies and use cases including enhanced mobile broadband (eMBB) , ultra-reliable low-latency communication (URLLC) , massive machine-type communication (mMTC) , millimeter wave (mmWave) technology,  beamforming, network slicing, edge computing, Internet of Things (IoT) connectivity and management, and network function virtualization (NFV) .

[0042] As the demand for broadband access increases and as technologies supported by wireless communication networks evolve, further technological improvements may be adopted in or implemented for 5G NR or future RATs, such as 6G, to further advance the evolution of wireless communication for a wide variety of existing and new use cases and applications. Such technological improvements may be associated with new frequency band expansion, licensed and unlicensed spectrum access, overlapping spectrum use, small cell deployments, non-terrestrial network (NTN) deployments, disaggregated network architectures and network topology expansion, device aggregation, advanced duplex communication, sidelink and other device-to-device direct communication, IoT (including passive or ambient IoT) networks, reduced capability (RedCap) UE functionality, industrial connectivity, multiple-subscriber implementations, high-precision positioning, radio frequency (RF) sensing, and / or AI / ML, among other examples. These technological improvements may support use cases such as wireless backhauls, wireless data centers, extended reality (XR) and metaverse applications, meta services for supporting vehicle connectivity, holographic and mixed reality communication, autonomous and collaborative robots, vehicle platooning and cooperative maneuvering, sensing networks, gesture monitoring, human-brain interfacing, digital twin applications, asset management, and universal coverage applications using non-terrestrial and / or aerial platforms, among other examples. The methods, operations, apparatuses, and techniques described herein may enable one or more of the foregoing technologies and / or support one or more of the foregoing use cases.

[0043] Fig. 1 is a diagram illustrating an example of a wireless communication network 100 in accordance with the present disclosure. The wireless communication network 100 may be or may include elements of a 5G (or NR) network or a 6G network, among other examples. The wireless communication network 100 may include multiple network nodes 110, shown as a network node (NN) 110a, a network node 110b, a network node 110c, and a network node 110d. The network nodes 110 may support communications with multiple UEs 120, shown as a UE 120a, a UE 120b, a UE 120c, a UE 120d, and a UE 120e.

[0044] The network nodes 110 and the UEs 120 of the wireless communication network 100 may communicate using the electromagnetic spectrum, which may be  subdivided by frequency or wavelength into various classes, bands, carriers, and / or channels. For example, devices of the wireless communication network 100 may communicate using one or more operating bands. In some aspects, multiple wireless networks 100 may be deployed in a given geographic area. Each wireless communication network 100 may support a particular RAT (which may also be referred to as an air interface) and may operate on one or more carrier frequencies in one or more frequency ranges. Examples of RATs include a 4G RAT, a 5G / NR RAT, and / or a 6G RAT, among other examples. In some examples, when multiple RATs are deployed in a given geographic area, each RAT in the geographic area may operate on different frequencies to avoid interference with one another.

[0045] Various operating bands have been defined as frequency range designations FR1 (410 MHz through 7.125 GHz) , FR2 (24.25 GHz through 52.6 GHz) , FR3 (7.125 GHz through 24.25 GHz) , FR4a or FR4-1 (52.6 GHz through 71 GHz) , FR4 (52.6 GHz through 114.25 GHz) , and FR5 (114.25 GHz through 300 GHz) . Although a portion of FR1 is greater than 6 GHz, FR1 is often referred to (interchangeably) as a “Sub-6 GHz” band in some documents and articles. Similarly, FR2 is often referred to (interchangeably) as a “millimeter wave” band in some documents and articles, despite being different than the extremely high frequency (EHF) band (30 GHz through 300 GHz) , which is identified by the International Telecommunications Union (ITU) as a “millimeter wave” band. The frequencies between FR1 and FR2 are often referred to as mid-band frequencies, which include FR3. Frequency bands falling within FR3 may inherit FR1 characteristics or FR2 characteristics, and thus may effectively extend features of FR1 or FR2 into mid-band frequencies. Thus, “sub-6 GHz, ” if used herein, may broadly refer to frequencies that are less than 6 GHz, that are within FR1, and / or that are included in mid-band frequencies. Similarly, the term “millimeter wave, ” if used herein, may broadly refer to frequencies that are included in mid-band frequencies, that are within FR2, FR4, FR4-a or FR4-1, or FR5, and / or that are within the EHF band. Higher frequency bands may extend 5G NR operation, 6G operation, and / or other RATs beyond 52.6 GHz. For example, each of FR4a, FR4-1, FR4, and FR5 falls within the EHF band. In some examples, the wireless communication network 100 may implement dynamic spectrum sharing (DSS) , in which multiple RATs (for example, 4G / LTE and 5G / NR) are implemented with dynamic bandwidth allocation (for example, based on user demand) in a single frequency band. It is contemplated that the frequencies included in these operating bands (for example, FR1, FR2, FR3, FR4, FR4- a, FR4-1, and / or FR5) may be modified, and techniques described herein may be applicable to those modified frequency ranges.

[0046] A network node 110 may include one or more devices, components, or systems that enable communication between a UE 120 and one or more devices, components, or systems of the wireless communication network 100. A network node 110 may be, may include, or may also be referred to as an NR network node, a 5G network node, a 6G network node, a Node B, an eNB, a gNB, an access point (AP) , a transmission reception point (TRP) , a mobility element, a core, a network entity, a network element, a network equipment, and / or another type of device, component, or system included in a radio access network (RAN) .

[0047] A network node 110 may be implemented as a single physical node (for example, a single physical structure) or may be implemented as two or more physical nodes (for example, two or more distinct physical structures) . For example, a network node 110 may be a device or system that implements part of a radio protocol stack, a device or system that implements a full radio protocol stack (such as a full gNB protocol stack) , or a collection of devices or systems that collectively implement the full radio protocol stack. For example, and as shown, a network node 110 may be an aggregated network node (having an aggregated architecture) , meaning that the network node 110 may implement a full radio protocol stack that is physically and logically integrated within a single node (for example, a single physical structure) in the wireless communication network 100. For example, an aggregated network node 110 may consist of a single standalone base station or a single TRP that uses a full radio protocol stack to enable or facilitate communication between a UE 120 and a core network of the wireless communication network 100.

[0048] Alternatively, and as also shown, a network node 110 may be a disaggregated network node (sometimes referred to as a disaggregated base station) , meaning that the network node 110 may implement a radio protocol stack that is physically distributed and / or logically distributed among two or more nodes in the same geographic location or in different geographic locations. For example, a disaggregated network node may have a disaggregated architecture. In some deployments, disaggregated network nodes 110 may be used in an integrated access and backhaul (IAB) network, in an open radio access network (O-RAN) (such as a network configuration in compliance with the O-RAN Alliance) , or in a virtualized radio access network (vRAN) , also known as a cloud  radio access network (C-RAN) , to facilitate scaling by separating base station functionality into multiple units that can be individually deployed.

[0049] The network nodes 110 of the wireless communication network 100 may include one or more central units (CUs) , one or more distributed units (DUs) , and / or one or more radio units (RUs) . A CU may host one or more higher layer control functions, such as radio resource control (RRC) functions, packet data convergence protocol (PDCP) functions, and / or service data adaptation protocol (SDAP) functions, among other examples. A DU may host one or more of a radio link control (RLC) layer, a medium access control (MAC) layer, and / or one or more higher physical (PHY) layers depending, at least in part, on a functional split, such as a functional split defined by the 3GPP. In some examples, a DU also may host one or more lower PHY layer functions, such as a fast Fourier transform (FFT) , an inverse FFT (iFFT) , beamforming, physical random access channel (PRACH) extraction and filtering, and / or scheduling of resources for one or more UEs 120, among other examples. An RU may host RF processing functions or lower PHY layer functions, such as an FFT, an iFFT, beamforming, or PRACH extraction and filtering, among other examples, according to a functional split, such as a lower layer functional split. In such an architecture, each RU can be operated to handle over the air (OTA) communication with one or more UEs 120.

[0050] In some aspects, a single network node 110 may include a combination of one or more CUs, one or more DUs, and / or one or more RUs. Additionally or alternatively, a network node 110 may include one or more Near-Real Time (Near-RT) RAN Intelligent Controllers (RICs) and / or one or more Non-Real Time (Non-RT) RICs. In some examples, a CU, a DU, and / or an RU may be implemented as a virtual unit, such as a virtual central unit (VCU) , a virtual distributed unit (VDU) , or a virtual radio unit (VRU) , among other examples. A virtual unit may be implemented as a virtual network function, such as associated with a cloud deployment.

[0051] Some network nodes 110 (for example, a base station, an RU, or a TRP) may provide communication coverage for a particular geographic area. In the 3GPP, the term “cell” can refer to a coverage area of a network node 110 or to a network node 110 itself, depending on the context in which the term is used. A network node 110 may support one or multiple (for example, three) cells. In some examples, a network node 110 may provide communication coverage for a macro cell, a pico cell, a femto cell, or another type of cell. A macro cell may cover a relatively large geographic area (for  example, several kilometers in radius) and may allow unrestricted access by UEs 120 with service subscriptions. A pico cell may cover a relatively small geographic area and may allow unrestricted access by UEs 120 with service subscriptions. A femto cell may cover a relatively small geographic area (for example, a home) and may allow restricted access by UEs 120 having association with the femto cell (for example, UEs 120 in a closed subscriber group (CSG) ) . A network node 110 for a macro cell may be referred to as a macro network node. A network node 110 for a pico cell may be referred to as a pico network node. A network node 110 for a femto cell may be referred to as a femto network node or an in-home network node. In some examples, a cell may not necessarily be stationary. For example, the geographic area of the cell may move according to the location of an associated mobile network node 110 (for example, a train, a satellite base station, an unmanned aerial vehicle, or an NTN network node) .

[0052] The wireless communication network 100 may be a heterogeneous network that includes network nodes 110 of different types, such as macro network nodes, pico network nodes, femto network nodes, relay network nodes, aggregated network nodes, and / or disaggregated network nodes, among other examples. In the example shown in Fig. 1, the network node 110a may be a macro network node for a macro cell 130a, the network node 110b may be a pico network node for a pico cell 130b, and the network node 110c may be a femto network node for a femto cell 130c. Various different types of network nodes 110 may generally transmit at different power levels, serve different coverage areas, and / or have different impacts on interference in the wireless communication network 100 than other types of network nodes 110. For example, macro network nodes may have a high transmit power level (for example, 5 to 40 watts) , whereas pico network nodes, femto network nodes, and relay network nodes may have lower transmit power levels (for example, 0.1 to 2 watts) .

[0053] In some examples, a network node 110 may be, may include, or may operate as an RU, a TRP, or a base station that communicates with one or more UEs 120 via a radio access link (which may be referred to as a “Uu” link) . The radio access link may include a downlink and an uplink. “Downlink” (or “DL” ) refers to a communication direction from a network node 110 to a UE 120, and “uplink” (or “UL” ) refers to a communication direction from a UE 120 to a network node 110. Downlink channels may include one or more control channels and one or more data channels. A downlink control channel may be used to transmit downlink control information (DCI) (for example, scheduling information, reference signals, and / or configuration information)  from a network node 110 to a UE 120. A downlink data channel may be used to transmit downlink data (for example, user data associated with a UE 120) from a network node 110 to a UE 120. Downlink control channels may include one or more physical downlink control channels (PDCCHs) , and downlink data channels may include one or more physical downlink shared channels (PDSCHs) . Uplink channels may similarly include one or more control channels and one or more data channels. An uplink control channel may be used to transmit uplink control information (UCI) (for example, reference signals and / or feedback corresponding to one or more downlink transmissions) from a UE 120 to a network node 110. An uplink data channel may be used to transmit uplink data (for example, user data associated with a UE 120) from a UE 120 to a network node 110. Uplink control channels may include one or more physical uplink control channels (PUCCHs) , and uplink data channels may include one or more physical uplink shared channels (PUSCHs) . The downlink and the uplink may each include a set of resources on which the network node 110 and the UE 120 may communicate.

[0054] Downlink and uplink resources may include time domain resources (frames, subframes, slots, and / or symbols) , frequency domain resources (frequency bands, component carriers, subcarriers, resource blocks, and / or resource elements) , and / or spatial domain resources (particular transmit directions and / or beam parameters) . Frequency domain resources of some bands may be subdivided into bandwidth parts (BWPs) . A BWP may be a continuous block of frequency domain resources (for example, a continuous block of resource blocks) that are allocated for one or more UEs 120. A UE 120 may be configured with both an uplink BWP and a downlink BWP (where the uplink BWP and the downlink BWP may be the same BWP or different BWPs) . A BWP may be dynamically configured (for example, by a network node 110 transmitting a DCI configuration to the one or more UEs 120) and / or reconfigured, which means that a BWP can be adjusted in real-time (or near-real-time) based on changing network conditions in the wireless communication network 100 and / or based on the specific requirements of the one or more UEs 120. This enables more efficient use of the available frequency domain resources in the wireless communication network 100 because fewer frequency domain resources may be allocated to a BWP for a UE 120 (which may reduce the quantity of frequency domain resources that a UE 120 is required to monitor) , leaving more frequency domain resources to be spread across multiple UEs 120. Thus, BWPs may also assist in the implementation of lower- capability UEs 120 by facilitating the configuration of smaller bandwidths for communication by such UEs 120.

[0055] As described above, in some aspects, the wireless communication network 100 may be, may include, or may be included in, an IAB network. In an IAB network, at least one network node 110 is an anchor network node that communicates with a core network. An anchor network node 110 may also be referred to as an IAB donor (or “IAB-donor” ) . The anchor network node 110 may connect to the core network via a wired backhaul link. For example, an Ng interface of the anchor network node 110 may terminate at the core network. Additionally or alternatively, an anchor network node 110 may connect to one or more devices of the core network that provide a core access and mobility management function (AMF) . An IAB network also generally includes multiple non-anchor network nodes 110, which may also be referred to as relay network nodes or simply as IAB nodes (or “IAB-nodes” ) . Each non-anchor network node 110 may communicate directly with the anchor network node 110 via a wireless backhaul link to access the core network, or may communicate indirectly with the anchor network node 110 via one or more other non-anchor network nodes 110 and associated wireless backhaul links that form a backhaul path to the core network. Some anchor network node 110 or other non-anchor network node 110 may also communicate directly with one or more UEs 120 via wireless access links that carry access traffic. In some examples, network resources for wireless communication (such as time resources, frequency resources, and / or spatial resources) may be shared between access links and backhaul links.

[0056] In some examples, any network node 110 that relays communications may be referred to as a relay network node, a relay station, or simply as a relay. A relay may receive a transmission of a communication from an upstream station (for example, another network node 110 or a UE 120) and transmit the communication to a downstream station (for example, a UE 120 or another network node 110) . In this case, the wireless communication network 100 may include or be referred to as a “multi-hop network. ” In the example shown in Fig. 1, the network node 110d (for example, a relay network node) may communicate with the network node 110a (for example, a macro network node) and the UE 120d in order to facilitate communication between the network node 110a and the UE 120d. Additionally or alternatively, a UE 120 may be or may operate as a relay station that can relay transmissions to or from other UEs 120. A  UE 120 that relays communications may be referred to as a UE relay or a relay UE, among other examples.

[0057] The UEs 120 may be physically dispersed throughout the wireless communication network 100, and each UE 120 may be stationary or mobile. A UE 120 may be, may include, or may be included in an access terminal, another terminal, a mobile station, or a subscriber unit. A UE 120 may be, include, or be coupled with a cellular phone (for example, a smart phone) , a personal digital assistant (PDA) , a wireless modem, a wireless communication device, a handheld device, a laptop computer, a cordless phone, a wireless local loop (WLL) station, a tablet, a camera, a gaming device, a netbook, a smartbook, an ultrabook, a medical device, a biometric device, a wearable device (for example, a smart watch, smart clothing, smart glasses, a smart wristband, and / or smart jewelry, such as a smart ring or a smart bracelet) , an entertainment device (for example, a music device, a video device, and / or a satellite radio) , an XR device, a vehicular component or sensor, a smart meter or sensor, industrial manufacturing equipment, a Global Navigation Satellite System (GNSS) device (such as a Global Positioning System device or another type of positioning device) , a UE function of a network node, and / or any other suitable device or function that may communicate via a wireless medium.

[0058] A UE 120 and / or a network node 110 may include one or more chips, system-on-chips (SoCs) , chipsets, packages, or devices that individually or collectively constitute or comprise a processing system. The processing system includes processor (or “processing” ) circuitry in the form of one or multiple processors, microprocessors, processing units (such as central processing units (CPUs) , graphics processing units (GPUs) , neural processing units (NPUs) and / or digital signal processors (DSPs) ) , processing blocks, application-specific integrated circuits (ASIC) , programmable logic devices (PLDs) (such as field programmable gate arrays (FPGAs) ) , or other discrete gate or transistor logic or circuitry (all of which may be generally referred to herein individually as “processors” or collectively as “the processor” or “the processor circuitry” ) . One or more of the processors may be individually or collectively configurable or configured to perform various functions or operations described herein. A group of processors collectively configurable or configured to perform a set of functions may include a first processor configurable or configured to perform a first function of the set and a second processor configurable or configured to perform a  second function of the set, or may include the group of processors all being configured or configurable to perform the set of functions.

[0059] The processing system may further include memory circuitry in the form of one or more memory devices, memory blocks, memory elements or other discrete gate or transistor logic or circuitry, each of which may include tangible storage media such as random-access memory (RAM) or read-only memory (ROM) , or combinations thereof (all of which may be generally referred to herein individually as “memories” or collectively as “the memory” or “the memory circuitry” ) . One or more of the memories may be coupled (for example, operatively coupled, communicatively coupled, electronically coupled, or electrically coupled) with one or more of the processors and may individually or collectively store processor-executable code (such as software) that, when executed by one or more of the processors, may configure one or more of the processors to perform various functions or operations described herein. Additionally or alternatively, in some examples, one or more of the processors may be preconfigured to perform various functions or operations described herein without requiring configuration by software. The processing system may further include or be coupled with one or more modems (such as a Wi-Fi (for example, IEEE compliant) modem or a cellular (for example, 3GPP 4G LTE, 5G, or 6G compliant) modem) . In some implementations, one or more processors of the processing system include or implement one or more of the modems. The processing system may further include or be coupled with multiple radios (collectively “the radio” ) , multiple RF chains, or multiple transceivers, each of which may in turn be coupled with one or more of multiple antennas. In some implementations, one or more processors of the processing system include or implement one or more of the radios, RF chains or transceivers. The UE 120 may include or may be included in a housing that houses components associated with the UE 120 including the processing system.

[0060] Some UEs 120 may be considered machine-type communication (MTC) UEs, evolved or enhanced machine-type communication (eMTC) , UEs, further enhanced eMTC (feMTC) UEs, or enhanced feMTC (efeMTC) UEs, or further evolutions thereof, all of which may be simply referred to as “MTC UEs” . An MTC UE may be, may include, or may be included in or coupled with a robot, an uncrewed aerial vehicle, a remote device, a sensor, a meter, a monitor, and / or a location tag. Some UEs 120 may be considered IoT devices and / or may be implemented as NB-IoT (narrowband IoT) devices. An IoT UE or NB-IoT device may be, may include, or may be included in or  coupled with an industrial machine, an appliance, a refrigerator, a doorbell camera device, a home automation device, and / or a light fixture, among other examples. Some UEs 120 may be considered Customer Premises Equipment, which may include telecommunications devices that are installed at a customer location (such as a home or office) to enable access to a service provider's network (such as included in or in communication with the wireless communication network 100) .

[0061] Some UEs 120 may be classified according to different categories in association with different complexities and / or different capabilities. UEs 120 in a first category may facilitate massive IoT in the wireless communication network 100, and may offer low complexity and / or cost relative to UEs 120 in a second category. UEs 120 in a second category may include mission-critical IoT devices, legacy UEs, baseline UEs, high-tier UEs, advanced UEs, full-capability UEs, and / or premium UEs that are capable of URLLC, enhanced mobile broadband (eMBB) , and / or precise positioning in the wireless communication network 100, among other examples. A third category of UEs 120 may have mid-tier complexity and / or capability (for example, a capability between UEs 120 of the first category and UEs 120 of the second capability) . A UE 120 of the third category may be referred to as a reduced capacity UE ( “RedCap UE” ) , a mid-tier UE, an NR-Light UE, and / or an NR-Lite UE, among other examples. RedCap UEs may bridge a gap between the capability and complexity of NB-IoT devices and / or eMTC UEs, and mission-critical IoT devices and / or premium UEs. RedCap UEs may include, for example, wearable devices, IoT devices, industrial sensors, and / or cameras that are associated with a limited bandwidth, power capacity, and / or transmission range, among other examples. RedCap UEs may support healthcare environments, building automation, electrical distribution, process automation, transport and logistics, and / or smart city deployments, among other examples.

[0062] In some examples, two or more UEs 120 (for example, shown as UE 120a and UE 120e) may communicate directly with one another using sidelink communications (for example, without communicating by way of a network node 110 as an intermediary) . As an example, the UE 120a may directly transmit data, control information, or other signaling as a sidelink communication to the UE 120e. This is in contrast to, for example, the UE 120a first transmitting data in an UL communication to a network node 110, which then transmits the data to the UE 120e in a DL communication. In various examples, the UEs 120 may transmit and receive sidelink communications using peer-to-peer (P2P) communication protocols, device-to-device  (D2D) communication protocols, vehicle-to-everything (V2X) communication protocols (which may include vehicle-to-vehicle (V2V) protocols, vehicle-to-infrastructure (V2I) protocols, and / or vehicle-to-pedestrian (V2P) protocols) , and / or mesh network communication protocols. In some deployments and configurations, a network node 110 may schedule and / or allocate resources for sidelink communications between UEs 120 in the wireless communication network 100. In some other deployments and configurations, a UE 120 (instead of a network node 110) may perform, or collaborate or negotiate with one or more other UEs to perform, scheduling operations, resource selection operations, and / or other operations for sidelink communications.

[0063] In various examples, some of the network nodes 110 and the UEs 120 of the wireless communication network 100 may be configured for full-duplex operation in addition to half-duplex operation. A network node 110 or a UE 120 operating in a half-duplex mode may perform only one of transmission or reception during particular time resources, such as during particular slots, symbols, or other time periods. Half-duplex operation may involve time-division duplexing (TDD) , in which DL transmissions of the network node 110 and UL transmissions of the UE 120 do not occur in the same time resources (that is, the transmissions do not overlap in time) . In contrast, a network node 110 or a UE 120 operating in a full-duplex mode can transmit and receive communications concurrently (for example, in the same time resources) . By operating in a full-duplex mode, network nodes 110 and / or UEs 120 may generally increase the capacity of the network and the radio access link. In some examples, full-duplex operation may involve frequency-division duplexing (FDD) , in which DL transmissions of the network node 110 are performed in a first frequency band or on a first component carrier and transmissions of the UE 120 are performed in a second frequency band or on a second component carrier different than the first frequency band or the first component carrier, respectively. In some examples, full-duplex operation may be enabled for a UE 120 but not for a network node 110. For example, a UE 120 may simultaneously transmit an UL transmission to a first network node 110 and receive a DL transmission from a second network node 110 in the same time resources. In some other examples, full-duplex operation may be enabled for a network node 110 but not for a UE 120. For example, a network node 110 may simultaneously transmit a DL transmission to a first UE 120 and receive an UL transmission from a second UE 120 in  the same time resources. In some other examples, full-duplex operation may be enabled for both a network node 110 and a UE 120.

[0064] In some examples, the UEs 120 and the network nodes 110 may perform MIMO communication. “MIMO” generally refers to transmitting or receiving multiple signals (such as multiple layers or multiple data streams) simultaneously over the same time and frequency resources. MIMO techniques generally exploit multipath propagation. MIMO may be implemented using various spatial processing or spatial multiplexing operations. In some examples, MIMO may support simultaneous transmission to multiple receivers, referred to as multi-user MIMO (MU-MIMO) . Some RATs may employ advanced MIMO techniques, such as mTRP operation (including redundant transmission or reception on multiple TRPs) , reciprocity in the time domain or the frequency domain, single-frequency-network (SFN) transmission, or non-coherent joint transmission (NC-JT) .

[0065] In some aspects, the UE 120 may include a communication manager 140. As described in more detail elsewhere herein, the communication manager 140 may transmit, to a network node, an indication of an inference accuracy level associated with a quantized AI / ML model for an AI / ML task, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to a floating-point AI / ML model for the AI / ML task; and receive, from the network node, signaling associated with AI / ML performance monitoring based at least in part on the indication of the inference accuracy level. Additionally, or alternatively, the communication manager 140 may perform one or more other operations described herein.

[0066] In some aspects, the network node 110 may include a communication manager 150. As described in more detail elsewhere herein, the communication manager 150 may receive an indication of an inference accuracy level associated with a quantized AI / ML model to be used by a UE for an AI / ML task, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to a floating-point AI / ML model for the AI / ML task; and transmit, to the UE, signaling associated with AI / ML performance monitoring based at least in part on the indication of the inference accuracy level.

[0067] In some aspects, as described in more detail elsewhere herein, the communication manager 150 may transmit, to a UE, a quantized AI / ML model for an AI / ML task, wherein the quantized AI / ML model is associated with a floating-point  AI / ML model for the AI / ML task; and transmit an indication of an inference accuracy level associated with the quantized AI / ML model, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to the floating-point AI / ML model. Additionally, or alternatively, the communication manager 150 may perform one or more other operations described herein.

[0068] As indicated above, Fig. 1 is provided as an example. Other examples may differ from what is described with regard to Fig. 1.

[0069] Fig. 2 is a diagram illustrating an example network node 110 in communication with an example UE 120 in a wireless network in accordance with the present disclosure.

[0070] As shown in Fig. 2, the network node 110 may include a data source 212, a transmit processor 214, a transmit (TX) MIMO processor 216, a set of modems 232 (shown as 232a through 232t, where t ≥ 1) , a set of antennas 234 (shown as 234a through 234v, where v ≥ 1) , a MIMO detector 236, a receive processor 238, a data sink 239, a controller / processor 240, a memory 242, a communication unit 244, a scheduler 246, and / or a communication manager 150, among other examples. In some configurations, one or a combination of the antenna (s) 234, the modem (s) 232, the MIMO detector 236, the receive processor 238, the transmit processor 214, and / or the TX MIMO processor 216 may be included in a transceiver of the network node 110. The transceiver may be under control of and used by one or more processors, such as the controller / processor 240, and in some aspects in conjunction with processor-readable code stored in the memory 242, to perform aspects of the methods, processes, and / or operations described herein. In some aspects, the network node 110 may include one or more interfaces, communication components, and / or other components that facilitate communication with the UE 120 or another network node.

[0071] The terms “processor, ” “controller, ” or “controller / processor” may refer to one or more controllers and / or one or more processors. For example, reference to “a / the processor, ” “a / the controller / processor, ” or the like (in the singular) should be understood to refer to any one or more of the processors described in connection with Fig. 2, such as a single processor or a combination of multiple different processors. Reference to “one or more processors” should be understood to refer to any one or more of the processors described in connection with Fig. 2. For example, one or more processors of the network node 110 may include transmit processor 214, TX MIMO processor 216, MIMO detector 236, receive processor 238, and / or controller / processor  240. Similarly, one or more processors of the UE 120 may include MIMO detector 256, receive processor 258, transmit processor 264, TX MIMO processor 266, and / or controller / processor 280.

[0072] In some aspects, a single processor may perform all of the operations described as being performed by the one or more processors. In some aspects, a first set of (one or more) processors of the one or more processors may perform a first operation described as being performed by the one or more processors, and a second set of (one or more) processors of the one or more processors may perform a second operation described as being performed by the one or more processors. The first set of processors and the second set of processors may be the same set of processors or may be different sets of processors. Reference to “one or more memories” should be understood to refer to any one or more memories of a corresponding device, such as the memory described in connection with Fig. 2. For example, operation described as being performed by one or more memories can be performed by the same subset of the one or more memories or different subsets of the one or more memories.

[0073] For downlink communication from the network node 110 to the UE 120, the transmit processor 214 may receive data ( “downlink data” ) intended for the UE 120 (or a set of UEs that includes the UE 120) from the data source 212 (such as a data pipeline or a data queue) . In some examples, the transmit processor 214 may select one or more MCSs for the UE 120 in accordance with one or more channel quality indicators (CQIs) received from the UE 120. The network node 110 may process the data (for example, including encoding the data) for transmission to the UE 120 on a downlink in accordance with the MCS (s) selected for the UE 120 to generate data symbols. The transmit processor 214 may process system information (for example, semi-static resource partitioning information (SRPI) ) and / or control information (for example, CQI requests, grants, and / or upper layer signaling) and provide overhead symbols and / or control symbols. The transmit processor 214 may generate reference symbols for reference signals (for example, a cell-specific reference signal (CRS) , a demodulation reference signal (DMRS) , or a channel state information (CSI) reference signal (CSI-RS) ) and / or synchronization signals (for example, a primary synchronization signal (PSS) or a secondary synchronization signals (SSS) ) .

[0074] The TX MIMO processor 216 may perform spatial processing (for example, precoding) on the data symbols, the control symbols, the overhead symbols, and / or the reference symbols, if applicable, and may provide a set of output symbol  streams (for example, T output symbol streams) to the set of modems 232. For example, each output symbol stream may be provided to a respective modulator component (shown as MOD) of a modem 232. Each modem 232 may use the respective modulator component to process (for example, to modulate) a respective output symbol stream (for example, for orthogonal frequency division multiplexing (OFDM) ) to obtain an output sample stream. Each modem 232 may further use the respective modulator component to process (for example, convert to analog, amplify, filter, and / or upconvert) the output sample stream to obtain a time domain downlink signal. The modems 232a through 232t may together transmit a set of downlink signals (for example, T downlink signals) via the corresponding set of antennas 234.

[0075] A downlink signal may include a DCI communication, a MAC control element (MAC-CE) communication, an RRC communication, a downlink reference signal, or another type of downlink communication. Downlink signals may be transmitted on a PDCCH, a PDSCH, and / or on another downlink channel. A downlink signal may carry one or more transport blocks (TBs) of data. A TB may be a unit of data that is transmitted over an air interface in the wireless communication network 100. A data stream (for example, from the data source 212) may be encoded into multiple TBs for transmission over the air interface. The quantity of TBs used to carry the data associated with a particular data stream may be associated with a TB size common to the multiple TBs. The TB size may be based on or otherwise associated with radio channel conditions of the air interface, the MCS used for encoding the data, the downlink resources allocated for transmitting the data, and / or another parameter. In general, the larger the TB size, the greater the amount of data that can be transmitted in a single transmission, which reduces signaling overhead. However, larger TB sizes may be more prone to transmission and / or reception errors than smaller TB sizes, but such errors may be mitigated by more robust error correction techniques.

[0076] For uplink communication from the UE 120 to the network node 110, uplink signals from the UE 120 may be received by an antenna 234, may be processed by a modem 232 (for example, a demodulator component, shown as DEMOD, of a modem 232) , may be detected by the MIMO detector 236 (for example, a receive (Rx) MIMO processor) if applicable, and / or may be further processed by the receive processor 238 to obtain decoded data and / or control information. The receive processor 238 may provide the decoded data to a data sink 239 (which may be a data pipeline, a data queue,  and / or another type of data sink) and provide the decoded control information to a processor, such as the controller / processor 240.

[0077] The network node 110 may use the scheduler 246 to schedule one or more UEs 120 for downlink or uplink communications. In some aspects, the scheduler 246 may use DCI to dynamically schedule DL transmissions to the UE 120 and / or UL transmissions from the UE 120. In some examples, the scheduler 246 may allocate recurring time domain resources and / or frequency domain resources that the UE 120 may use to transmit and / or receive communications using an RRC configuration (for example, a semi-static configuration) , for example, to perform semi-persistent scheduling (SPS) or to configure a configured grant (CG) for the UE 120.

[0078] One or more of the transmit processor 214, the TX MIMO processor 216, the modem 232, the antenna 234, the MIMO detector 236, the receive processor 238, and / or the controller / processor 240 may be included in an RF chain of the network node 110. An RF chain may include one or more filters, mixers, oscillators, amplifiers, analog-to-digital converters (ADCs) , and / or other devices that convert between an analog signal (such as for transmission or reception via an air interface) and a digital signal (such as for processing by one or more processors of the network node 110) . In some aspects, the RF chain may be or may be included in a transceiver of the network node 110.

[0079] In some examples, the network node 110 may use the communication unit 244 to communicate with a core network and / or with other network nodes. The communication unit 244 may support wired and / or wireless communication protocols and / or connections, such as Ethernet, optical fiber, common public radio interface (CPRI) , and / or a wired or wireless backhaul, among other examples. The network node 110 may use the communication unit 244 to transmit and / or receive data associated with the UE 120 or to perform network control signaling, among other examples. The communication unit 244 may include a transceiver and / or an interface, such as a network interface.

[0080] The UE 120 may include a set of antennas 252 (shown as antennas 252a through 252r, where r ≥ 1) , a set of modems 254 (shown as modems 254a through 254u, where u ≥ 1) , a MIMO detector 256, a receive processor 258, a data sink 260, a data source 262, a transmit processor 264, a TX MIMO processor 266, a controller / processor 280, a memory 282, and / or a communication manager 140, among other examples. One or more of the components of the UE 120 may be included in a housing 284. In some aspects, one or a combination of the antenna (s) 252, the modem (s) 254, the MIMO  detector 256, the receive processor 258, the transmit processor 264, or the TX MIMO processor 266 may be included in a transceiver that is included in the UE 120. The transceiver may be under control of and used by one or more processors, such as the controller / processor 280, and in some aspects in conjunction with processor-readable code stored in the memory 282, to perform aspects of the methods, processes, or operations described herein. In some aspects, the UE 120 may include another interface, another communication component, and / or another component that facilitates communication with the network node 110 and / or another UE 120.

[0081] For downlink communication from the network node 110 to the UE 120, the set of antennas 252 may receive the downlink communications or signals from the network node 110 and may provide a set of received downlink signals (for example, R received signals) to the set of modems 254. For example, each received signal may be provided to a respective demodulator component (shown as DEMOD) of a modem 254. Each modem 254 may use the respective demodulator component to condition (for example, filter, amplify, downconvert, and / or digitize) a received signal to obtain input samples. Each modem 254 may use the respective demodulator component to further demodulate or process the input samples (for example, for OFDM) to obtain received symbols. The MIMO detector 256 may obtain received symbols from the set of modems 254, may perform MIMO detection on the received symbols if applicable, and may provide detected symbols. The receive processor 258 may process (for example, decode) the detected symbols, may provide decoded data for the UE 120 to the data sink 260 (which may include a data pipeline, a data queue, and / or an application executed on the UE 120) , and may provide decoded control information and system information to the controller / processor 280.

[0082] For uplink communication from the UE 120 to the network node 110, the transmit processor 264 may receive and process data ( “uplink data” ) from a data source 262 (such as a data pipeline, a data queue, and / or an application executed on the UE 120) and control information from the controller / processor 280. The control information may include one or more parameters, feedback, one or more signal measurements, and / or other types of control information. In some aspects, the receive processor 258 and / or the controller / processor 280 may determine, for a received signal (such as received from the network node 110 or another UE) , one or more parameters relating to transmission of the uplink communication. The one or more parameters may include a reference signal received power (RSRP) parameter, a received signal strength  indicator (RSSI) parameter, a reference signal received quality (RSRQ) parameter, a CQI parameter, or a transmit power control (TPC) parameter, among other examples. The control information may include an indication of the RSRP parameter, the RSSI parameter, the RSRQ parameter, the CQI parameter, the TPC parameter, and / or another parameter. The control information may facilitate parameter selection and / or scheduling for the UE 120 by the network node 110.

[0083] The transmit processor 264 may generate reference symbols for one or more reference signals, such as an uplink DMRS, an uplink sounding reference signal (SRS) , and / or another type of reference signal. The symbols from the transmit processor 264 may be precoded by the TX MIMO processor 266, if applicable, and further processed by the set of modems 254 (for example, for DFT-s-OFDM or CP-OFDM) . The TX MIMO processor 266 may perform spatial processing (for example, precoding) on the data symbols, the control symbols, the overhead symbols, and / or the reference symbols, if applicable, and may provide a set of output symbol streams (for example, U output symbol streams) to the set of modems 254. For example, each output symbol stream may be provided to a respective modulator component (shown as MOD) of a modem 254. Each modem 254 may use the respective modulator component to process (for example, to modulate) a respective output symbol stream (for example, for OFDM) to obtain an output sample stream. Each modem 254 may further use the respective modulator component to process (for example, convert to analog, amplify, filter, and / or upconvert) the output sample stream to obtain an uplink signal.

[0084] The modems 254a through 254u may transmit a set of uplink signals (for example, R uplink signals or U uplink symbols) via the corresponding set of antennas 252. An uplink signal may include a UCI communication, a MAC-CE communication, an RRC communication, or another type of uplink communication. Uplink signals may be transmitted on a PUSCH, a PUCCH, and / or another type of uplink channel. An uplink signal may carry one or more TBs of data. Sidelink data and control transmissions (that is, transmissions directly between two or more UEs 120) may generally use similar techniques as were described for uplink data and control transmission, and may use sidelink-specific channels such as a physical sidelink shared channel (PSSCH) , a physical sidelink control channel (PSCCH) , and / or a physical sidelink feedback channel (PSFCH) .

[0085] One or more antennas of the set of antennas 252 or the set of antennas 234 may include, or may be included within, one or more antenna panels, one or more  antenna groups, one or more sets of antenna elements, or one or more antenna arrays, among other examples. An antenna panel, an antenna group, a set of antenna elements, or an antenna array may include one or more antenna elements (within a single housing or multiple housings) , a set of coplanar antenna elements, a set of non-coplanar antenna elements, or one or more antenna elements coupled with one or more transmission or reception components, such as one or more components of Fig. 2. As used herein, “antenna” can refer to one or more antennas, one or more antenna panels, one or more antenna groups, one or more sets of antenna elements, or one or more antenna arrays. “Antenna panel” can refer to a group of antennas (such as antenna elements) arranged in an array or panel, which may facilitate beamforming by manipulating parameters of the group of antennas. “Antenna module” may refer to circuitry including one or more antennas, which may also include one or more other components (such as filters, amplifiers, or processors) associated with integrating the antenna module into a wireless communication device.

[0086] In some examples, each of the antenna elements of an antenna 234 or an antenna 252 may include one or more sub-elements for radiating or receiving radio frequency signals. For example, a single antenna element may include a first sub-element cross-polarized with a second sub-element that can be used to independently transmit cross-polarized signals. The antenna elements may include patch antennas, dipole antennas, and / or other types of antennas arranged in a linear pattern, a two-dimensional pattern, or another pattern. A spacing between antenna elements may be such that signals with a desired wavelength transmitted separately by the antenna elements may interact or interfere constructively and destructively along various directions (such as to form a desired beam) . For example, given an expected range of wavelengths or frequencies, the spacing may provide a quarter wavelength, a half wavelength, or another fraction of a wavelength of spacing between neighboring antenna elements to allow for the desired constructive and destructive interference patterns of signals transmitted by the separate antenna elements within that expected range.

[0087] The amplitudes and / or phases of signals transmitted via antenna elements and / or sub-elements may be modulated and shifted relative to each other (such as by manipulating phase shift, phase offset, and / or amplitude) to generate one or more beams, which is referred to as beamforming. The term “beam” may refer to a directional transmission of a wireless signal toward a receiving device or otherwise in a  desired direction. “Beam” may also generally refer to a direction associated with such a directional signal transmission, a set of directional resources associated with the signal transmission (for example, an angle of arrival, a horizontal direction, and / or a vertical direction) , and / or a set of parameters that indicate one or more aspects of a directional signal, a direction associated with the signal, and / or a set of directional resources associated with the signal. In some implementations, antenna elements may be individually selected or deselected for directional transmission of a signal (or signals) by controlling amplitudes of one or more corresponding amplifiers and / or phases of the signal (s) to form one or more beams. The shape of a beam (such as the amplitude, width, and / or presence of side lobes) and / or the direction of a beam (such as an angle of the beam relative to a surface of an antenna array) can be dynamically controlled by modifying the phase shifts, phase offsets, and / or amplitudes of the multiple signals relative to each other.

[0088] Different UEs 120 or network nodes 110 may include different numbers of antenna elements. For example, a UE 120 may include a single antenna element, two antenna elements, four antenna elements, eight antenna elements, or a different number of antenna elements. As another example, a network node 110 may include eight antenna elements, 24 antenna elements, 64 antenna elements, 128 antenna elements, or a different number of antenna elements. Generally, a larger number of antenna elements may provide increased control over parameters for beam generation relative to a smaller number of antenna elements, whereas a smaller number of antenna elements may be less complex to implement and may use less power than a larger number of antenna elements. Multiple antenna elements may support multiple-layer transmission, in which a first layer of a communication (which may include a first data stream) and a second layer of a communication (which may include a second data stream) are transmitted using the same time and frequency resources with spatial multiplexing.

[0089] While blocks in Fig. 2 are illustrated as distinct components, the functions described above with respect to the blocks may be implemented in a single hardware, software, or combination component or in various combinations of components. For example, the functions described with respect to the transmit processor 264, the receive processor 258, and / or the TX MIMO processor 266 may be performed by or under the control of the controller / processor 280.

[0090] Fig. 3 is a diagram illustrating an example disaggregated base station architecture 300 in accordance with the present disclosure. One or more components of  the example disaggregated base station architecture 300 may be, may include, or may be included in one or more network nodes (such one or more network nodes 110) . The disaggregated base station architecture 300 may include a CU 310 that can communicate directly with a core network 320 via a backhaul link, or that can communicate indirectly with the core network 320 via one or more disaggregated control units, such as a Non-RT RIC 350 associated with a Service Management and Orchestration (SMO) Framework 360 and / or a Near-RT RIC 370 (for example, via an E2 link) . The CU 310 may communicate with one or more DUs 330 via respective midhaul links, such as via F1 interfaces. Each of the DUs 330 may communicate with one or more RUs 340 via respective fronthaul links. Each of the RUs 340 may communicate with one or more UEs 120 via respective RF access links. In some deployments, a UE 120 may be simultaneously served by multiple RUs 340.

[0091] Each of the components of the disaggregated base station architecture 300, including the CUs 310, the DUs 330, the RUs 340, the Near-RT RICs 370, the Non-RT RICs 350, and the SMO Framework 360, may include one or more interfaces or may be coupled with one or more interfaces for receiving or transmitting signals, such as data or information, via a wired or wireless transmission medium.

[0092] In some aspects, the CU 310 may be logically split into one or more CU user plane (CU-UP) units and one or more CU control plane (CU-CP) units. A CU-UP unit may communicate bidirectionally with a CU-CP unit via an interface, such as the E1 interface when implemented in an O-RAN configuration. The CU 310 may be deployed to communicate with one or more DUs 330, as necessary, for network control and signaling. Each DU 330 may correspond to a logical unit that includes one or more base station functions to control the operation of one or more RUs 340. For example, a DU 330 may host various layers, such as an RLC layer, a MAC layer, or one or more PHY layers, such as one or more high PHY layers or one or more low PHY layers. Each layer (which also may be referred to as a module) may be implemented with an interface for communicating signals with other layers (and modules) hosted by the DU 330, or for communicating signals with the control functions hosted by the CU 310. Each RU 340 may implement lower layer functionality. In some aspects, real-time and non-real-time aspects of control and user plane communication with the RU (s) 340 may be controlled by the corresponding DU 330.

[0093] The SMO Framework 360 may support RAN deployment and provisioning of non-virtualized and virtualized network elements. For non-virtualized network  elements, the SMO Framework 360 may support the deployment of dedicated physical resources for RAN coverage requirements, which may be managed via an operations and maintenance interface, such as an O1 interface. For virtualized network elements, the SMO Framework 360 may interact with a cloud computing platform (such as an open cloud (O-Cloud) platform 390) to perform network element life cycle management (such as to instantiate virtualized network elements) via a cloud computing platform interface, such as an O2 interface. A virtualized network element may include, but is not limited to, a CU 310, a DU 330, an RU 340, a non-RT RIC 350, and / or a Near-RT RIC 370. In some aspects, the SMO Framework 360 may communicate with a hardware aspect of a 4G RAN, a 5G NR RAN, and / or a 6G RAN, such as an open eNB (O-eNB) 380, via an O1 interface. Additionally or alternatively, the SMO Framework 360 may communicate directly with each of one or more RUs 340 via a respective O1 interface. In some deployments, this configuration can enable each DU 330 and the CU 310 to be implemented in a cloud-based RAN architecture, such as a vRAN architecture.

[0094] The Non-RT RIC 350 may include or may implement a logical function that enables non-real-time control and optimization of RAN elements and resources, AI / ML workflows including model training and updates, and / or policy-based guidance of applications and / or features in the Near-RT RIC 370. The Non-RT RIC 350 may be coupled to or may communicate with (such as via an A1 interface) the Near-RT RIC 370. The Near-RT RIC 370 may include or may implement a logical function that enables near-real-time control and optimization of RAN elements and resources via data collection and actions via an interface (such as via an E2 interface) connecting one or more CUs 310, one or more DUs 330, and / or an O-eNB with the Near-RT RIC 370.

[0095] In some aspects, to generate AI / ML models to be deployed in the Near-RT RIC 370, the Non-RT RIC 350 may receive parameters or external enrichment information from external servers. Such information may be utilized by the Near-RT RIC 370 and may be received at the SMO Framework 360 or the Non-RT RIC 350 from non-network data sources or from network functions. In some examples, the Non-RT RIC 350 or the Near-RT RIC 370 may tune RAN behavior or performance. For example, the Non-RT RIC 350 may monitor long-term trends and patterns for performance and may employ AI / ML models to perform corrective actions via the SMO Framework 360 (such as reconfiguration via an O1 interface) or via creation of RAN management policies (such as A1 interface policies) .

[0096] As indicated above, Fig. 3 is provided as an example. Other examples may differ from what is described with regard to Fig. 3.

[0097] The network node 110, the controller / processor 240 of the network node 110, the UE 120, the controller / processor 280 of the UE 120, the CU 310, the DU 330, the RU 340, or any other component (s) of Figs. 1, 2, or 3 may implement one or more techniques or perform one or more operations associated with reporting of inference accuracies of quantized models for AI / ML performance monitoring, as described in more detail elsewhere herein. For example, the controller / processor 240 of the network node 110, the controller / processor 280 of the UE 120, any other component (s) of Fig. 2, the CU 310, the DU 330, or the RU 340 may perform or direct operations of, for example, process 900 of Fig. 9, process 1000 of Fig. 10, process 1100 of Fig. 11, or other processes as described herein (alone or in conjunction with one or more other processors) . In some aspects, the server device described herein is the network node 110, is included in the network node 110, or includes one or more components of the network node 110 shown in Fig. 2. The memory 242 may store data and program codes for the network node 110, the network node 110, the CU 310, the DU 330, or the RU 340. The memory 282 may store data and program codes for the UE 120. In some examples, the memory 242 or the memory 282 may include a non-transitory computer-readable medium storing a set of instructions (for example, code or program code) for wireless communication. The memory 242 may include one or more memories, such as a single memory or multiple different memories (of the same type or of different types) . The memory 282 may include one or more memories, such as a single memory or multiple different memories (of the same type or of different types) . For example, the set of instructions, when executed (for example, directly, or after compiling, converting, or interpreting) by one or more processors of the network node 110, the UE 120, the CU 310, the DU 330, or the RU 340, may cause the one or more processors to perform process 900 of Fig. 9, process 1000 of Fig. 10, process 1100 of Fig. 11, or other processes as described herein. In some examples, executing instructions may include running the instructions, converting the instructions, compiling the instructions, and / or interpreting the instructions, among other examples.

[0098] In some aspects, a UE (e.g., the UE 120) includes means for transmitting, to a network node, an indication of an inference accuracy level associated with a quantized AI / ML model for an AI / ML task, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to a floating-point AI / ML  model for the AI / ML task; and / or means for receiving, from the network node, signaling associated with AI / ML performance monitoring based at least in part on the indication of the inference accuracy level. The means for the UE to perform operations described herein may include, for example, one or more of communication manager 140, antenna 252, modem 254, MIMO detector 256, receive processor 258, transmit processor 264, TX MIMO processor 266, controller / processor 280, or memory 282.

[0099] In some aspects, a network node (e.g., the network node 110) includes means for receiving an indication of an inference accuracy level associated with a quantized AI / ML model to be used by a UE for an AI / ML task, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to a floating-point AI / ML model for the AI / ML task; and / or means for transmitting, to the UE, signaling associated with AI / ML performance monitoring based at least in part on the indication of the inference accuracy level. The means for the network node to perform operations described herein may include, for example, one or more of communication manager 150, transmit processor 214, TX MIMO processor 216, modem 232, antenna 234, MIMO detector 236, receive processor 238, controller / processor 240, memory 242, or scheduler 246.

[0100] In some aspects, a server device includes means for transmitting, to a UE, a quantized AI / ML model for an AI / ML task, wherein the quantized AI / ML model is associated with a floating-point AI / ML model for the AI / ML task; and / or means for transmitting an indication of an inference accuracy level associated with the quantized AI / ML model, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to the floating-point AI / ML model. In some aspects, the means for the server device to perform operations described herein may include, for example, one or more of communication manager 150, transmit processor 214, TX MIMO processor 216, modem 232, antenna 234, MIMO detector 236, receive processor 238, controller / processor 240, memory 242, or scheduler 246.

[0101] As indicated above, Fig. 3 is provided as an example. Other examples may differ from what is described with regard to Fig. 3.

[0102] Fig. 4 is a diagram illustrating an example architecture 400 of a functional framework for RAN intelligence enabled by data collection, in accordance with the present disclosure. In some scenarios, the functional framework for RAN intelligence may be enabled by further enhancement of data collection through use cases and / or examples. For example, principles or algorithms for RAN intelligence enabled by  AI / ML and the associated functional framework (e.g., the AI functionality and / or the input / output of the component for AI enabled optimization) have been utilized or studied to identify the benefits of AI enabled RAN through possible use cases (e.g., beam management, energy saving, load balancing, mobility management, and / or coverage optimization, among other examples) . In one example, as shown by the architecture 400, a functional framework for RAN intelligence may include multiple logical entities, such as a model training host 402, a model inference host 404, data sources 406, and an actor 408.

[0103] The model inference host 404 may be configured to run an AI / ML model based on inference data provided by the data sources 406, and the model inference host 404 may produce an output (e.g., a prediction) with the inference data input to the actor 408. The actor 408 may be an element or an entity of a core network or a RAN. For example, the actor 408 may be a UE, a network node, base station (e.g., a gNB) , a CU, a DU, and / or an RU, among other examples. In addition, the actor 408 may also depend on the type of tasks performed by the model inference host 404, type of inference data provided to the model inference host 404, and / or type of output produced by the model inference host 404. For example, if the output from the model inference host 404 is associated with beam management, the actor 408 may be a UE, a DU or an RU. In some examples, the model inference host 404 may be hosted on the actor 408. For example, a UE may be the actor 408 and may host the model inference host 404. In some aspects, a UE (e.g., the actor 408) may be a data source 406. For example, the UE may perform a measurement (e.g., an NR measurement) , may input the measurement to the AI / ML model at the model inference host 404 (or may provide the measurement to the model inference host 404) , and may act based on an output of the AI / ML model.

[0104] After the actor 408 receives an output from the model inference host 404, the actor 408 may determine whether to act based on the output. For example, if the actor 408 is a DU or an RU and the output from the model inference host 404 is associated with beam management, the actor 408 may determine whether to change / modify a Tx / Rx beam based on the output. If the actor 408 determines to act based on the output, in some examples, the actor 408 may indicate the action to at least one subject of action 410. For example, if the actor 408 determines to change / modify a Tx / Rx beam for a communication between the actor 408 and the subject of action 410 (e.g., a UE) , then the actor 408 may transmit a beam (re-) configuration or a beam switching indication to the subject of action 410. The actor 408 may modify its Tx / Rx beam based on the beam  (re-) configuration, such as switching to a new Tx / Rx beam or applying different parameters for a Tx / Rx beam, among other examples. As another example, the actor 408 may be a UE, and the output from the model inference host 404 may be associated with beam management. For example, the output may be one or more predicted measurement values for one or more beams. The actor 408 (e.g., a UE) may determine that a measurement report (e.g., a layer 1 (L1) RSRP report) is to be transmitted to a network node.

[0105] The data sources 406 may also be configured for collecting data that is used as training data for training an ML model or as inference data for feeding an ML model inference operation. For example, the data sources 406 may collect data from one or more core network and / or RAN entities, which may include the actor 408 or the subject of action 410, and provide the collected data to the model training host 402 for ML model training. In some aspects, the model training host 402 may be co-located with the model inference host 404 and / or the actor 408. For example, the actor 408 or the subject of action 410 may provide performance feedback associated with the beam configuration to the data sources 406, where the performance feedback may be used by the model training host 402 for monitoring or evaluating the ML model performance, such as whether the output (e.g., prediction) provided to the actor 408 is accurate. In some examples, if the output provided by the actor 408 is inaccurate (or the accuracy is below an accuracy threshold) , then the model training host 402 may determine to modify or retrain the ML model used by the model inference host, such as via an ML model deployment / update.

[0106] As indicated above, Fig. 4 is provided as an example. Other examples may differ from what is described with regard to Fig. 4.

[0107] Fig. 5 is a diagram illustrating an example 500 of AI / ML based beam management, in accordance with the present disclosure. As shown in Fig. 5, an AI / ML model 510 may be deployed at or on a UE 120. For example, a model inference host (such as a model inference host) may be deployed at, or on, a UE 120. The AI / ML model 510 may enable the UE 120 to determine one or more inferences or predictions based on data input to the AI / ML model 510.

[0108] For example, as shown by reference number 515, an input to the AI / ML model 510 may include measurements associated with a first set of beams. For example, a network node 110 may transmit one or more signals using respective beams from the first set of beams. The UE 120 may perform measurements (e.g., L1 RSRP  measurements, L1 signal to interference plus noise (SINR) measurements, or other measurements) of the first set of beams to obtain a first set of measurements. For example, each beam, from the first set of beams, may be associated with one or more measurements performed by the UE 120. The UE 120 may input the first set of measurements (e.g., L1 RSRP measurement values or L1 SINR measurement values) into the AI / ML model 510 along with information associated with the first set of beams and / or a second set of beams, such as a beam direction (e.g., spatial direction) , beam width, beam shape, and / or other characteristics of the respective beams from the first set of beams and / or the second set of beams.

[0109] As shown by reference number 520, the AI / ML model 510 may output one or more predictions. The one or more predictions may include predicted measurement values (e.g., predicted L1 RSRP measurement values or predicted SINR measurement values) associated with the second set of beams. This may reduce a quantity of beam measurements that are performed by the UE 120, thereby conversing power of the UE 120 and / or network resources that would have otherwise been used to measure all beams included in the first set of beams and the second set of beams. This type of prediction may be referred to as a codebook based spatial domain selection or prediction.

[0110] As another example, an output of the AI / ML model 510 may include a point-direction, an angle of departure (AoD) , and / or an angle of arrival (AoA) of a beam included in the second set of beams. This type of prediction may be referred to as a non-codebook based spatial domain selection or prediction. As another example, multiple measurement report or values, collected at different points in time, may be input to the AI / ML model 510. This may enable the AI / ML model 510 to output codebook based and / or non-codebook based predictions for a measurement value, an AoD, and / or an AoA, among other examples, of a beam at a future time. The output (s) of the AI / ML model 510, as described herein, may facilitate initial access procedures, secondary cell group (SCG) setup procedures, beam refinement procedures, link quality or interference adaptation procedure, beam failure and / or beam blockage predictions, and / or radio link failure predictions, among other examples.

[0111] In some examples, the first set of beams (e.g., that are measured) may be referred to as Set B beams and the second set of beams (e.g., that are associated with predicted measurements) may be referred to as Set A beams. In some examples, the first set of beams (e.g., the Set B beams) may be a subset of the second set of beams  (e.g., the Set A beams) . In some other examples, the first set of beams and the second set of beams may be different beams and / or may be mutually exclusive sets. For example, the first set of beams (e.g., the Set B beams) may include wide beams (e.g., unrefined beams or beams having a beam width that satisfies a first threshold) and the second set of beams (e.g., the Set A beams) may include narrow beams (e.g., refined beams or beams having a beam width that satisfies a second threshold) . In one example, the AI / ML model 510 may perform spatial-domain beam predictions for beams included in the Set A beams based on measurement results of beams included in the Set B beams. As another example, the AI / ML model 510 may perform temporal beam prediction for beams included in the Set A beams based on historic measurement results of beams included in the Set B beams.

[0112] As indicated above, Fig. 5 is provided as an example. Other examples may differ from what is described with regard to Fig. 5.

[0113] Fig. 6 is a diagram illustrating an example AI / ML architecture 600 of a first wireless device 602 in communication with a second wireless device 604, in accordance with the present disclosure. In some aspects, the first wireless device 602 may be a UE (e.g., UE 120) , and the second wireless device 604 may be a network node (e.g., network node 110) . Note that the example AI / ML architecture of first wireless device 602 may be applied to second wireless device 604, and vice versa.

[0114] The first wireless device 602 may be, or may include, a chip, SoC, chipset, package, or device that includes one or more processors, processing blocks or processing elements (collectively “processor 610” ) and one or more memory blocks or elements (collectively “memory 620” ) . Processor 610 may be coupled to transceiver 640, which includes RF circuitry 642 coupled to antennas 646 via an interface 644, for transmitting or receiving signals.

[0115] One or more AI / ML models 630 (collectively “AI / ML model 630” ) may be stored in the memory 620 and accessible to the processor (s) 610. Individual or groups of AI / ML models 630 may be associated with respective model identifiers (IDs) . In some aspects, different AI / ML models 630, which may optionally be associated with different model identifiers, may have different characteristics. One or more AI / ML models 630 may be selected based on respective features, characteristics, or applications, as well as characteristics or conditions of first wireless device 602 (such as, a power state, a mobility state, a battery reserve, a temperature, etc. ) . For example, the AI / ML models 630 may have different inference data and output pairings (such as, different types of  inference data producing different types of output) , different levels of accuracies associated with the predictions, different latencies associated with producing the predictions, different ML model sizes, different coefficients, different parameters, etc.

[0116] The processor (s) 610 may deploy the AI / ML models 630 to produce respective output data based on input data. As an example, the AI / ML model 630 may take measurements of one or more reference signals (such as, corresponding to one or more wide beams) as input to predict a channel characteristic (e.g., beam measurement) associated with one or more different reference signals (such as, corresponding to one or more narrow beams within each of the one or more wide beams, one or more other wide beams, or one or more narrow beams outside the one or more wide beams, among other examples) . The input data may include, for example, measurements of one or more reference or pilot signals, such as a CQI, a signal-to-noise ratio (SNR) , an SINR, a signal-to-noise-plus-distortion ratio (SNDR) , an RSSI, an RSRP, an RSRQ, and / or a block error rate (BLER) , among other examples. The output data may include, for example, compressed CSI feedback or one or more predicted measurements (or characteristics) of one or more reference or pilot signals.

[0117] As shown in Fig. 6, the first wireless device 602 and / or the second wireless device 604 may communicate with a model server 650. In some aspects, the model server 650 may perform various AI / ML management tasks for first wireless device 602 and / or second wireless device 604. For example, the model server 650 may host various types and / or versions of AI / ML models 630 for the first wireless device 602 and / or the second wireless device 604 to download. In some examples, the model server 650 may monitor and evaluate the performance of the AI / ML model 630. Additionally, or alternatively, in some examples (e.g., where the first wireless device 602 is a UE and the second wireless device is a network node) , the second wireless device 604 may perform AI / ML performance monitoring to monitor and evaluate the performance of the AI / ML model 630 used by the first wireless device 602. In some examples, the model server 650 may transmit signals or provide indications / instructions to activate or deactivate the use of a particular AI / ML model 630 at the first wireless device 602 or the second wireless device 604. In some examples, the model server 650 may switch to a different AI / ML model 630 being used at the first wireless device 602 or the second wireless device 604, and the model server 650 may provide such an instruction to the respective first wireless device 602 or second wireless device 604. The model server 650 may operate as a model training host (such as model training host  402 discussed in connection with Fig. 4) and update the AI / ML model 630 using training data. In some cases, the model server 650 may operate as a data source (such as data source 406 discussed in connection with Fig. 4) to collect and host training data, inference data, performance feedback, etc., associated with AI / ML model 630.

[0118] In some examples, an AI / ML model (e.g., the AI / ML model 630) may be trained for an AI / ML task (e.g., beam management / prediction) at a network node or a model server (e.g., the model server 650) and transferred to one or more UEs to enable the one or more UEs to use the AI / ML model to perform inference (e.g., prediction) for the AI / ML task. In order to increase efficiency (e.g., to deploy an AI / ML model at a UE) , a floating-point AI / ML model for an AI / ML task may be quantized resulting in a quantized AI / ML model for the AI / ML task. “Quantization” is a model size reduction technique that reduces the number of bits used to represent model parameters. For example, quantization may convert model weights (and / or other model parameters) from a high-precision (e.g., floating-point) representation to a lower-precision representation (e.g., a fixed value representation or an integer value representation) . Quantization of a model (e.g., a floating-point model) may reduce the memory usage and computational complexity the model.

[0119] In some examples, different UEs may have different UE-specific model quantization preferences or restrictions. For example, different UE AI / ML hardware accelerators may have different quantization preferences or restrictions. This may lead to different UEs deploying different quantized AI / ML models, with different inference accuracy performance levels, for a same floating-point AI / ML model for an AI / ML task. Examples of quantization preferences or restrictions that may determine the quantization applied to a floating-point AI / ML model to generate a corresponding quantized AI / ML model may include metrics (e.g., clustering, K-means, entropy, uniformed, non-uniformed, etc. ) to determine floating-to-fixed point representations, quantization bit-width and schemes (e.g., fixed-point, integer, etc. ) , component-specific preferences (e.g., whether activation functions should be quantized, a preference of fixed-point 16 instead of integer 16 for pooling layers, etc. ) . In some examples, different RF circuitry vendors’ chipsets may be associated with different preferences or restrictions regarding quantization schemes. In some examples, even for chipsets from the same RF circuitry vendor, different original equipment manufacturers (OEMs) may have different tailored quantization preferences or restrictions, for example, due to specific power or thermal budget considerations. In addition, there may be various  modules or types of hardware accelerators associated with different quantization preferences or restrictions, even from a same RF circuitry vendor.

[0120] In some aspects, the model server 650 may provide the quantized UE-side AI / ML model to the UE. In some examples, the model server 650 may be a third-party server (e.g., a server associated with an entity other than an operator of a network node and / or wireless communication network to which the UE connects) . For example, model servers (e.g., model server 650) associated with different vendors (e.g., UE and / or RF circuitry vendors) may provide quantized UE-side AI / ML models associated with a certain AI / ML task, functionality, feature, or model ID (e.g., an AI / ML task / functionality / feature / model ID associated with a wireless communication standard, such as a 3GPP standard) . As an example, a model server (e.g., model server 650) may store a trained floating-point AI / ML model for an AI / ML task of wide-to-narrow spatial beam prediction (e.g., via L1-RSRP prediction) with a model ID associated with 8 Set B beams and 32 Set A beams. A UE may communicate with the model server to report quantization preferences and / or restrictions associated with the UE. For example, the model server may be a third-party server (e.g., associated with a vendor of the AI / ML hardware accelerator of the UE) , and the UE may communicate with the model server based at least in part on proprietary protocols not specified in a wireless communication standard (e.g., a 3GPP standard) . The model server may then prepare a quantized AI / ML model (e.g., by quantizing the floating-point AI / ML model) for the UE to download, following the UE’s quantization preferences and / or restrictions. In some aspects, the associated model inference accuracy level (e.g., in terms of average L1-RSRP prediction accuracy in comparison with the floating-point AI / ML model, or in comparison with actually measured L1-RSRPs, which can be calculated by the model server as well) for the quantized AI / ML model may also be downloadable together with the quantized model.

[0121] Different UEs may deploy different quantized AI / ML models associated with the same AI / ML task, and the different quantized AI / ML models have different initial model inference accuracy levels (e.g., due to the different quantization schemes used for the different AI / ML models) . A network node may perform AI / ML performance monitoring for the UEs deploying the different quantized AI / ML models for the AI / ML task. For example, the network node may schedule transmission of auxiliary RSs for prediction accuracy verification and / or configure inference error threshold values / number for triggering transmission of a performance alert message, among other  examples. However, the network node may not be aware of the different initial model inference accuracy levels for the quantized AI / ML models deployed at the different UEs. This may lead to less efficient AI / ML performance monitoring, such as using too aggressive or conservative AI / ML performance monitoring parameters for some UEs.

[0122] In some aspects described herein, a UE may report, to a network node, an inference accuracy level associated with a quantized AI / ML model to be used by the UE for an AI / ML task. In some aspects described herein, a server device (e.g., the model server 650) that provides a quantized AI / ML model to a UE may report, to a network node, an inference accuracy level associated with the quantized AI / ML model. In this way, the network node may be made aware of inference accuracy levels of different quantized AI / ML models used by different UEs, which may increase efficiency of AI / ML performance monitoring performed by the network node.

[0123] As indicated above, Fig. 6 is provided as an example. Other examples may differ from what is described with respect to Fig. 6.

[0124] Fig. 7 is a diagram illustrating an example 700 associated with reporting of inference accuracies of quantized models for AI / ML performance monitoring, in accordance with the present disclosure. As shown in Fig. 7, example 700 includes communication between a network node 110, a UE 120, and a model server 650. In some aspects, the model server 650 may be a third-party server that is associated with an entity other than an operator of the network node 110 (e.g., an operator of a wireless communication network associated with the network node 110) . In some other aspects, the model server may be associated with a same operator as the network node 110.

[0125] As shown in Fig. 7, and by reference number 705, the UE 120 may obtain a quantized AI / ML model for an AI / ML task. The quantized AI / ML model may be associated with a floating-point AI / ML model for the AI / ML task. For example, the quantized AI / ML model may be generated by applying quantization to the floating-point AI / ML model. That is, the quantized AI / ML model may be quantized version of the floating-point AI / ML model for the AI / ML task. The quantized AI / ML model may have the same inputs and output as the floating-point AI / ML model (e.g., inputs and outputs associated with the AI / ML task) .

[0126] In some aspects, the AI / ML task may be a beam measurement or beam management task. For example, the AI / ML task may be associated with predicting beam measurements for a set of beams. In such examples, the quantized AI / ML model (and the floating-point AI / ML model) may input beam measurements (e.g., L1-RSRP  measurements or other beam measurements) for a first set of beams (e.g., Set B beams) and may output predicted beam measurements (e.g., predicted L1-RSRP measurements or other beam measurements) for a second set of beams (e.g., Set A beams) . In another example, the AI / ML task may be associated predicting a number of top beams of a set of beams. In such examples, the quantized AI / ML model (and the floating-point AI / ML model) may input beam measurements (e.g., L1-RSRP measurements or other beam measurements) for a first set of beams (e.g., Set B beams) and may output identifiers (e.g., beam identifiers, RS identifiers, or the like) for a number (N) beams of a second set of beams (e.g., Set A beams) predicted to have highest / best beam measurements (e.g., the top N beams predicted to have the highest L1-RSRP measurements) . In some other aspects, the quantized AI / ML model may be associated with other AI / ML tasks.

[0127] In some aspects, as shown in Fig. 7, the UE 120 may obtain (e.g., receive or download) the quantized AI / ML model from the model server 650. In such examples, the model server 650 may transmit, and the UE 120 may receive, the quantized AI / ML model for the AI / ML task. The model server 650 may manage the floating-point AI / ML model for the AI / ML task and the quantized AI / ML model for the AI / ML task. For example, the model server 650 may store the floating-point AI / ML model for the AI / ML task. In some examples, the floating-point AI / ML model and / or the AI / ML task (e.g., an AI / ML functionality) may be associated with a logical model ID. That is, the logical model ID may identify an AI / ML logical model, which may be defined as a model with certain inputs and output (e.g., corresponding to an AI / ML task) . In some aspects, the model server 650 may generate the quantized AI / ML model in connection with communication with the UE 120. For example, the UE 120 may communicate with the model server 650 to request a quantized AI / ML model for the AI / ML task and to indicate one or more quantization preferences and / or restrictions associated with the UE 120. The model server 650 may then generate / prepare the quantized AI / ML model for the AI / ML task by performing quantization of the floating-point AI / ML model in accordance with the one or more quantization preferences and / or restrictions indicated by the UE 120. In some other aspects, the model server 650 may store the floating-point AI / ML model for the AI / ML tasks and multiple quantized AI / ML models for the AI / ML task that are generated (e.g., by the model server 650) by performing quantization of the floating-point AI / ML model using different quantization schemes and / or parameters. In this case, the UE 120 may select a quantized AI / ML model from the multiple quantized AI / ML models for the AI / ML task that are stored at the model  server 650. In some aspects, the UE 120 may communicate with the model server 650 to obtain the quantized AI / ML model via a protocol that is transparent to the network node 110 and / or a wireless communications standard (e.g., a 3GPP standard) used for communications between the UE 120 and the network node 110.

[0128] Although Fig. 7 shows the UE 120 obtaining the quantized AI / ML model from the model server 650, in some other aspects, the UE 120 may not obtain the quantized AI / ML model from the model server 650. In some aspects, the UE 120 may obtain the quantized AI / ML model via the network node 110. For example, the UE 120 may receive the quantized AI / ML model from the network node 110 or from another network entity (e.g., a core network device) via the network node 110. In some aspects, the UE 120 may obtain the quantized AI / ML model for the AI / ML task by performing quantization of the floating-point AI / ML model for the AI / ML task. For example, the UE 120 may receive the floating-point AI / ML model from the model server 650, the network node 110, or another network entity, or the UE 120 may train the floating-point AI / ML model for the AI / ML task, and the UE 120 may perform quantization of the floating-point AI / ML model to generate the quantized AI / ML model.

[0129] As further shown in Fig. 7, and by reference number 710, the UE 120 may transmit, and the network node 110 may receive, an indication of an inference accuracy level for the quantized AI / ML model. In some aspects, the inference accuracy level may be indicative of a relative accuracy of the quantized AI / ML model with respect to the floating-point AI / ML model for the AI / ML task (e.g., the corresponding floating-point AI / ML model that was quantized to generate the quantized AI / ML model) . In some aspects, the UE 120 may transmit the indication of the inference accuracy level for the quantized AI / ML model before the UE 120 is activated (e.g., by the network node 110) with the quantized AI / ML model and / or the AI / ML task. For example, the UE 120 may transmit the indication of the inference accuracy level for the quantized AI / ML model (e.g., together with an indication of a model ID associated with the AI / ML task) during initial access as UE capability information. In some aspects, the inference accuracy level may be an initial inference accuracy level for the quantized AI / ML model (e.g., prior to AI / ML performance monitoring performed by the network node 110) .

[0130] In some aspects, a definition for the inference accuracy level may be pre-configured / pre-defined for the UE 120 (e.g., specified in a wireless communication standard) or configured / indicated via configuration information transmitted by the  network node 110 and received by the UE 120 (e.g., via RRC signaling, a MAC-CE, or DCI) . In such examples, different definitions of inference accuracy (e.g., different measurements of the relative accuracy of the quantized AI / ML model with respect to the floating-point AI / ML model) may be pre-configured / pre-defined or configured / indicated for different AI / ML tasks (e.g., different AI / ML functionalities or logical models) . In some aspects, a plurality of candidate inference accuracy levels may be associated with the AI / ML task, and the indication of the inference accuracy level may indicate a selected candidate inference accuracy level from the plurality of candidate inference accuracy levels. In some examples, the plurality of candidate inference accuracy levels associated with the AI / ML task may be pre-configured / pre-defined (e.g., specified in a wireless communication standard) . In some other examples, the network node 110 may transmit, and the UE 120 may receive, configuration information indicating the plurality of candidate inference accuracy levels for the AI / ML task. In some aspects, inference accuracy levels may be different for different model IDs associated with the same AI / ML task (e.g., the same AI / ML functionality) . In some examples, the plurality of candidate inference accuracy levels may be associated with different model IDs, and the selected candidate inference accuracy level (indicated in the indication of the inference accuracy level) may be associated with a model ID of the quantized AI / ML model. In some examples, the plurality of candidate inference accuracy levels may include a plurality of model IDs (for a same AI / ML task) corresponding to different inference accuracy levels, and the UE 120 may indicate a model ID, from the plurality of model IDs, that corresponds to an inference accuracy level of the quantized AI / ML model.

[0131] In some aspects, the inference accuracy level may be based at least in part on statistical differences between certain model output features provided by the floating-point AI / ML model for the AI / ML task and the quantized AI / ML model to be used for the AI / ML task by the UE 120, and / or between certain model output features provided by a genie-aided detector / model (e.g., ground truth values for the certain model output features) for the AI / ML task and the quantized AI / ML model to be used for the AI / ML task by the UE 120 in association with a certain dataset (e.g., a reference dataset) including the model inputs and the output features provided by the genie-aided model. In some aspects, output values for one or more target output features determined by applying the floating-point AI / ML model for the AI / ML task to the reference dataset may be within an error value X of ground truth values (e.g., values determined via a  genie-aided detector) that are included in the reference dataset for the one or more target output features . For example, the error value X may be an error tolerance for the floating-point AI / ML model for the AI / ML task that is specified in a wireless communication standard. In some examples, X may be an error tolerance that is applied for each of the one or more target features, an error tolerance for an average error over the one or more target output features, an error tolerance for a total error over the one or more target output features, or an error tolerance for a maximum error among the one or more target output features, among other examples. In some aspects, the inference accuracy level for the quantized AI / ML model may be a difference value Y between output values for the one or more target output features determined by applying the quantized AI / ML model to the reference dataset and the output values for the one or more target features determined by applying the floating-point AI / ML model to the reference dataset. For example, the difference value Y may be an average difference between the output values of the quantized AI / ML model and the output values of the floating-point AI / ML model over the one or more target output features, the total difference between the output values of the quantized AI / ML model and the output values of the floating-point AI / ML model over the one or more target output features, a maximum difference between the output values of the quantized AI / ML model and the corresponding output values of the floating-point AI / ML model over the one or more target output features, or a vector of differences between the output values of the quantized AI / ML model and the corresponding output values of the floating-point AI / ML model for each of the one or more target output features, among other examples.

[0132] In some aspects, in a case in which the AI / ML task is associated with predicting beam measurements (e.g., L1-RSRP measurements) for a set of beams (e.g., Set A beams) , the inference accuracy level may be based at least in part on differences between first beam measurements (e.g., first L1-RSRP measurements) predicted by the quantized AI / ML model for a subset of the set of beams (e.g., a subset of the Set A beams) based on model inputs (e.g., beam measurements for set B beams) included in a reference dataset and second beam measurements predicted by the floating-point AI / ML model for the subset of the set of beams (e.g., the subset of the Set A beams) based on the model inputs included in the reference dataset. For example, the quantized AI / ML model (and the floating-point AI / ML model) may input L1-RSRP measurements for Set B beams (e.g., a set of wide beams) and output L1-RSRP measurements for Set A beams (e.g., a set of narrow beams) . The L1-RSRP prediction accuracy between  predicted L1-RSRP values determined by the floating-point AI / ML model for the reference dataset and ground truth L1-RSRP values (e.g., determined via a genie-aided detector) for the top-N (e.g., N = 4 in one example) Set A beams (e.g., N beams with the highest predicted L1-RSRP values in the Set A beams) may be within ±X dB (where X is an error tolerance that may be specified in a wireless communication standard) . The top-N Set A beams that are used for comparing the predicted L1-RSRP values output by the floating-point AI / ML model and the ground truth L1-RSRP values may be the ground truth top-N beams identifier (e.g., via a genie-aided detector) in the reference dataset. The value of N may be pre-defined or pre-configured (e.g., specified in a wireless communication standard) In this example, the inference accuracy level for the quantized AI / ML model may be a difference value Y between the predicted L1-RSRP values determined by the quantized AI / ML model for the reference dataset and the predicted L1-RSRP values determined by the floating-point AI / ML model for the reference dataset for the top-N Set A beams (e.g., the ground truth top-N Set A beams identified in the reference dataset) . For example, the output L1-RSRP values predicted for the top-N Set A beams (e.g., the ground truth top-N Set A beams) by the quantized AI / ML model may be within ±Y dB of the corresponding L1-RSRP values predicted for the top-N Set A beams by the floating-point AI / ML model. For example, Y may be an average L1-RSRP difference over the top-N Set A beams, a total L1-RSRP difference over the top-N Set A beams, a maximum L1-RSRP difference over the top-N Set A beams, or a vector of L1-RSRP differences for the top-N Set A beams. In some aspects, the indication of the inference accuracy level for the quantized AI / ML model may be or include an indication of Y. In some examples, the indication of the inference accuracy level for the quantized AI / ML model may include an indication of Y and an indication of a model ID associated with the AI / ML task.

[0133] In some aspects, in a case in which the AI / ML task is associated with predicting a number of top beams of a set of beams (e.g., identifying which N Set A beams are the top-N beams) , the inference accuracy level may be based at least in part on a percentage of the top beams accurately predicted by the quantized AI / ML model based on model inputs included in the reference dataset. For example, the inference accuracy level may be based at least in part on the percentage of the top beams accurately predicted by the quantized AI / ML model based on model inputs included in the reference dataset relative to the top beams predicted by the floating-point AI / ML model based on the model inputs included in the dataset. In this example, the reference  accuracy level Y may be the percentage of the top-N Set A beams accurately predicted by applying the quantized AI / ML model to the reference dataset, as compared with the top-N Set A beams predicted by applying the floating-point AI / ML model to the reference dataset. In some aspects, this value of Y may be indicated in the indication of the inference accuracy level for the quantized AI / ML model.

[0134] In some aspects, the model server 650 may indicate the inference accuracy level to the UE 120 (e.g., when the UE 120 obtains the quantized AI / ML model from the model server 650) . For example, the model server 650 may transmit, and the UE 120 may receive, an indication of a value of Y for the quantized AI / ML model. The model server 650 may transmit the indication of the value of Y for the quantized AI / ML model to the UE 120 together with the quantized AI / ML model. In such examples, the model server 650 may determine the inference accuracy level (e.g., the value of Y) for the quantized AI / ML model by applying the quantized AI / ML model and the floating-point AI / ML model to the reference dataset. In some other aspects, the UE 120 may determine the inference accuracy level (e.g., the value of Y) for the quantized AI / ML model, for example by applying the quantized AI / ML model to the reference dataset and comparing the model outputs of the quantized AI / ML model for the reference dataset to model outputs of the floating-point AI / ML model for the reference dataset (e.g., which may be received from the model server 650) .

[0135] The reference dataset used for determining the inference accuracy may include model inputs and ground truth (e.g., genie-aided) model output values. In some aspects, the reference dataset used for determining the inference accuracy level may be pre-defined (e.g., specified in a wireless communication standard) . For example, a wireless communication standard may specify different reference datasets for different AI / ML tasks (e.g., different AI / ML functionalities) . In some aspects, the reference dataset may be indicated via signaling transmitted by the network node 110. In some examples, the network node 110 may signal the dataset to one or more party entities (e.g., model servers) where quantized AI / ML models may be downloaded by UEs. For example, the network node 110 may transmit, and the model server 650 may receive, the reference dataset. Additionally, or alternatively, the network node 110 may configure the UE 120 may the reference dataset. For example, the network node 110 may transmit, and the UE 120 may receive, the reference dataset. In some other aspects, one or more third-party entities (e.g., model servers) may signal reference datasets used by the third-party entities to the network node 110, together with model IDs (e.g., logical model IDs)  corresponding to the AI / ML tasks associated with the reference datasets. For example, the model server 650 may transmit, and the network node 110 may receive, the reference dataset. In this case, the model server 650 may transmit an indication of a model ID associated with the floating-point AI / ML model and / or AI / ML task together with the reference dataset.

[0136] As further shown in Fig. 7, and by reference number 715, the network node 110 may transmit, and the UE 120 may receive, signaling associated with AI / ML performance monitoring based at least in part on the indication of the inference accuracy level for the quantized AI / ML model. In some aspects, the network node 110 may determine one or more AI / ML performance monitoring parameters for the UE 120 based at least in part on the indicated inference accuracy level for the quantized AI / ML model to be used by the UE 120, and the network node 110 may transmit, to the UE 120, signaling that indicates configuration information including the one or more AI / ML performance monitoring parameters and / or scheduling information based at least in part on the one or more AI / ML performance parameters.

[0137] In some aspects, the network node 110 may schedule transmission of RSs (e.g., auxiliary RSs) for prediction accuracy verification for the UE 120 based at least in part on the indicated inference accuracy level for the quantized AI / ML model. For example, the network node 110 may schedule the transmission, to the UE 120, of the RSs for the prediction accuracy verification more or less frequently (e.g., as compared to a default scheduling frequency and / or as compared to RSs scheduled for other UEs) based at least in part on the indicated inference accuracy level for the quantized AI / ML model. In such examples, the signaling associated with the AI / ML performance monitoring may include signaling that indicates the scheduling of the RSs to be transmitted to the UE 120 and / or the transmission, from the network node 110 to the UE 120, of the RSs for the prediction accuracy verification.

[0138] In some aspects, the network node 110 may determine / configure an error threshold value for the UE 120 to be used to determine whether a performance alert message should be sent to the UE 120. For example, the network node 110 may determine / configure higher or lower error threshold values for different UEs associated with an AI / ML task based at least in part on the respective inference accuracy levels for the quantized AI / ML models to be used by the different UEs. In such examples, the signaling associated with the AI / ML performance monitoring may include signaling that indicates or configures the error threshold value for the UE 120 and / or a performance  alert message transmitted to the UE 120 based at least in part on the error threshold value determined / configured for the UE 120.

[0139] As indicated above, Fig. 7 is provided as an example. Other examples may differ from what is described with respect to Fig. 7.

[0140] Fig. 8 is a diagram illustrating an example 800 associated with reporting of inference accuracies of quantized models for AI / ML performance monitoring, in accordance with the present disclosure. As shown in Fig. 8, example 800 includes communication between a network node 110, a UE 120, and a model server 650. In some aspects, the model server 650 may be a third-party server that is associated with an entity other than an operator of the network node 110 (e.g., an operator of a wireless communication network associated with the network node 110) . In some other aspects, the model server may be associated with a same operator as the network node 110.

[0141] As shown in Fig. 8, and by reference number 805, the model server 650 may transmit, and the UE 120 may receive, a quantized AI / ML model for an AI / ML task. For example, the UE 120 may download the quantized AI / ML model from the model server 650. The quantized AI / ML model may be associated with a floating-point AI / ML model for the AI / ML task. For example, the quantized AI / ML model may be generated by applying quantization to the floating-point AI / ML model. That is, the quantized AI / ML model may be quantized version of the floating-point AI / ML model for the AI / ML task. The quantized AI / ML model may have the same inputs and output as the floating-point AI / ML model (e.g., inputs and outputs associated with the AI / ML task) .

[0142] In some aspects, the AI / ML task may be a beam measurement or beam management task. For example, the AI / ML task may be associated with predicting beam measurements for a set of beams. In such examples, the quantized AI / ML model (and the floating-point AI / ML model) may input beam measurements (e.g., L1-RSRP measurements or other beam measurements) for a first set of beams (e.g., Set B beams) and may output predicted beam measurements (e.g., predicted L1-RSRP measurements or other beam measurements) for a second set of beams (e.g., Set A beams) . In another example, the AI / ML task may be associated predicting a number of top beams of a set of beams. In such examples, the quantized AI / ML model (and the floating-point AI / ML model) may input beam measurements (e.g., L1-RSRP measurements or other beam measurements) for a first set of beams (e.g., Set B beams) and may output identifiers (e.g., beam identifiers, RS identifiers, or the like) for a number (N) beams of a second  set of beams (e.g., Set A beams) predicted to have highest / best beam measurements (e.g., the top N beams predicted to have the highest L1-RSRP measurements) . In some other aspects, the quantized AI / ML model may be associated with other AI / ML tasks.

[0143] The model server 650 may manage the floating-point AI / ML model for the AI / ML task and the quantized AI / ML model for the AI / ML task. For example, the model server 650 may store the floating-point AI / ML model for the AI / ML task. In some examples, the floating-point AI / ML model and / or the AI / ML task (e.g., an AI / ML functionality) may be associated with a logical model ID. That is, the logical model ID may identify an AI / ML logical model, which may be defined as a model with certain inputs and output (e.g., corresponding to an AI / ML task) . In some aspects, the model server 650 may generate the quantized AI / ML model in connection with communication with the UE 120. For example, the UE 120 may communicate with the model server 650 to request a quantized AI / ML model for the AI / ML task and to indicate one or more quantization preferences and / or restrictions associated with the UE 120. The model server 650 may then generate / prepare the quantized AI / ML model for the AI / ML task by performing quantization of the floating-point AI / ML model in accordance with the one or more quantization preferences and / or restrictions indicated by the UE 120. In some other aspects, the model server 650 may store the floating-point AI / ML model for the AI / ML tasks and multiple quantized AI / ML models for the AI / ML task that are generated (e.g., by the model server 650) by performing quantization of the floating-point AI / ML model using different quantization schemes and / or parameters. In this case, the UE 120 may select a quantized AI / ML model from the multiple quantized AI / ML models for the AI / ML task that are stored at the model server 650. In such examples, each of the multiple quantized AI / ML models for the AI / ML task may be associated with a respective model ID (e.g., a respective physical model ID) . For example, the model server 650 may store respective quantized AI / ML models corresponding to multiple physical model IDs that are associated with the same AI / ML task (e.g., AI / ML functionality) and / or the same logical model ID. In some aspects, the UE 120 may communicate with the model server 650 to obtain the quantized AI / ML model via a protocol that is transparent to the network node 110 and / or a wireless communications standard (e.g., a 3GPP standard) used for communications between the UE 120 and the network node 110.

[0144] As further shown in Fig. 8, and by reference number 810, the model server 650 may transmit, and the network node 110 may receive, an indication of an inference  accuracy level for the quantized AI / ML model to be used by the UE 120. In some aspects, the inference accuracy level may be indicative of a relative accuracy of the quantized AI / ML model with respect to the floating-point AI / ML model for the AI / ML task (e.g., the corresponding floating-point AI / ML model that was quantized to generate the quantized AI / ML model) . In some aspects, the inference accuracy level may be an initial inference accuracy level for the quantized AI / ML model (e.g., prior to AI / ML performance monitoring performed by the network node 110) . In some aspects, the model server 650 may transmit, to the network node 110, an indication of the inference accuracy level for the quantized AI / ML model transmitted to (e.g., downloaded by) the UE 120. In such examples, the model server 650 may also transmit an indication of a model ID (e.g., a physical model ID) associated with the quantized AI / ML model and / or an indication of a logical model ID associated with the floating-point AI / ML model and / or the AI / ML task. For example, the model server 650 may transmit the indication of the model ID (e.g., the physical model ID) to register the quantized AI / ML model with the network node 110 as an AI / ML model for the AI / ML task.

[0145] In some aspects, the model server 650 may transmit, to the network node 110, indications of respective inference accuracy levels for a plurality of quantized AI / ML models for the AI / ML task (e.g., including the quantized AI / ML model to be used by the UE 120) . In such examples, the model server 650 may also transmit, to the network node 110, an indication of a respective model ID (e.g., physical model ID) for each of the plurality of quantized AI / ML models (e.g., to register the plurality of quantized AI / ML models with the network node 110 for the AI / ML task) . For example, the plurality of quantized AI / ML models may be a plurality of quantized AI / ML models for the AI / ML task (e.g., associated with the same logical model ID) that are stored on and / or have been generated by the model server 650. In some examples, the model server 650 may also indicate, to the network node 110, the logical model ID corresponding to the AI / ML task or floating-point AI / ML model associated with the plurality of quantized AI / ML models. In some examples, for each of one floating-point AI / ML models (corresponding to respective AI / ML tasks) managed by (e.g., stored on) the model server 650, the model server 650 may transmit, to the network node 110, an indication of the corresponding logical model ID and respective indications of physical model IDs and inference accuracy levels for one or more quantized AI / ML models associated with that logical model ID.

[0146] In some aspects, the model server 650 may transmit the indication of the inference accuracy level (and the physical model ID) for the quantized AI / ML model to be used by the UE 120 to the network node 110 after transmitting the quantized AI / ML model to the UE 120. For example, the model server 650 may transmit the indication of the inference accuracy level (and the physical model ID) for the quantized AI / ML model to the network node 110 in response to the UE 120 downloading the quantized AI / ML model from the model server 650 (or in response to generating the quantized AI / ML model in a case in which the model server 650 generates the quantized AI / ML model for the UE 120) . In some other aspects, the model server 650 may transmit the indication of the inference accuracy level (and the physical model ID) for the quantized AI / ML model to be used by the UE 120 to the network node 110 before transmitting the quantized AI / ML model to the UE 120. For example, in a case in which the model server 650 stores multiple quantized AI / ML models for an AI / ML task that are downloadable by UEs, the model server 650 may transmit the indications of the inference accuracy levels and the physical model IDs for the quantized AI / ML models to the network node 110 in connection with generating the quantized AI / ML models and independent of when the UE 120 downloads the quantized AI / ML model. In some aspects, the model server 650 may transmit the indications of the inference accuracy levels and physical model IDs for one or more quantized AI / ML models to multiple network nodes (e.g., including the network node 110) . In some aspects, the network node 110 may receive indications of inference accuracy levels and physical model IDs for quantized AI / ML models from multiple model servers (e.g., multiple third-party model servers) .

[0147] In some aspects, a definition for the inference accuracy level may be pre-configured / pre-defined (e.g., specified in a wireless communication standard) or configured / indicated via signaling from the network node 110. In such examples, different definitions of inference accuracy (e.g., different measurements of the relative accuracy of the quantized AI / ML model with respect to the floating-point AI / ML model) may be pre-configured / pre-defined or configured / indicated for different AI / ML tasks (e.g., different AI / ML functionalities or logical models) . In some examples in which the definition for the inference accuracy level (e.g., for an AI / ML task) is configured / indicated by the network node 110, the model server 650 may receive the signaling from the network node 110. In other examples, in which the definition for the inference accuracy level (e.g., for an AI / ML task) is configured / indicated by the  network node 110, the UE 120 may receive the signaling from the network node and indicate the definition for the inference accuracy level to the model server 650 (e.g., when communicating with the model server 650 to obtain the quantized AI / ML model) . In some aspects, a plurality of candidate inference accuracy levels may be associated with the AI / ML task, and the indication of the inference accuracy level may indicate a selected candidate inference accuracy level from the plurality of candidate inference accuracy levels. In some examples, the plurality of candidate inference accuracy levels associated with the AI / ML task may be pre-configured / pre-defined (e.g., specified in a wireless communication standard) . In some other examples, the network node 110 may transmit, and the model server 650 and / or the UE 120 may receive, configuration information indicating the plurality of candidate inference accuracy levels for the AI / ML task. In some aspects, inference accuracy levels may be different for different model IDs associated with the same AI / ML task (e.g., the same AI / ML functionality) . In some examples, the plurality of candidate inference accuracy levels may be associated with different model IDs, and the selected candidate inference accuracy level (indicated in the indication of the inference accuracy level) may be associated with a model ID of the quantized AI / ML model.

[0148] In some aspects, the inference accuracy level may be based at least in part on statistical differences between certain model output features provided by the floating-point AI / ML model for the AI / ML task and the quantized AI / ML model to be used for the AI / ML task by the UE 120, and / or between certain model output features provided by a genie-aided detector / model (e.g., ground truth values for the certain model output features) for the AI / ML task and the quantized AI / ML model to be used for the AI / ML task by the UE 120 in association with a certain dataset (e.g., a reference dataset) including the model inputs and the output features provided by the genie-aided model. In some aspects, as discussed above in connection with Fig. 7, output values for one or more target output features determined by applying the floating-point AI / ML model for the AI / ML task to the reference dataset may be within an error value X of ground truth values (e.g., values determined via a genie-aided detector) that are included in the reference dataset for the one or more target output features . For example, the error value X may be an error tolerance for the floating-point AI / ML model for the AI / ML task that is specified in a wireless communication standard. In some aspects, as discussed above in connection with Fig. 7, the inference accuracy level for the quantized AI / ML model may be a difference value Y between output values for the one  or more target output features determined by applying the quantized AI / ML model to the reference dataset and the output values for the one or more target features determined by applying the floating-point AI / ML model to the reference dataset.

[0149] In some aspects, in a case in which the AI / ML task is associated with predicting beam measurements (e.g., L1-RSRP measurements) for a set of beams (e.g., Set A beams) , the inference accuracy level may be based at least in part on differences between first beam measurements (e.g., first L1-RSRP measurements) predicted by the quantized AI / ML model for a subset of the set of beams (e.g., a subset of the Set A beams) based on model inputs (e.g., beam measurements for set B beams) included in a reference dataset and second beam measurements predicted by the floating-point AI / ML model for the subset of the set of beams (e.g., the subset of the Set A beams) based on the model inputs included in the reference dataset. For example, the quantized AI / ML model (and the floating-point AI / ML model) may input L1-RSRP measurements for Set B beams (e.g., a set of wide beams) and output L1-RSRP measurements for Set A beams (e.g., a set of narrow beams) . The L1-RSRP prediction accuracy between predicted L1-RSRP values determined by the floating-point AI / ML model for the reference dataset and ground truth L1-RSRP values (e.g., determined via a genie-aided detector) for the top-N (e.g., N = 4 in one example) Set A beams (e.g., N beams with the highest predicted L1-RSRP values in the Set A beams) may be within ±X dB (where X is an error tolerance that may be specified in a wireless communication standard) . The top-N Set A beams that are used for comparing the predicted L1-RSRP values output by the floating-point AI / ML model and the ground truth L1-RSRP values may be the ground truth top-N beams identifier (e.g., via a genie-aided detector) in the reference dataset. The value of N may be pre-defined or pre-configured (e.g., specified in a wireless communication standard) In this example, the inference accuracy level for the quantized AI / ML model may be a difference value Y between the predicted L1-RSRP values determined by the quantized AI / ML model for the reference dataset and the predicted L1-RSRP values determined by the floating-point AI / ML model for the reference dataset for the top-N Set A beams (e.g., the ground truth top-N Set A beams identified in the reference dataset) . For example, the output L1-RSRP values predicted for the top-N Set A beams (e.g., the ground truth top-N Set A beams) by the quantized AI / ML model may be within ±Y dB of the corresponding L1-RSRP values predicted for the top-N Set A beams by the floating-point AI / ML model. For example, Y may be an average L1-RSRP difference over the top-N Set A beams, a total L1-RSRP difference  over the top-N Set A beams, a maximum L1-RSRP difference over the top-N Set A beams, or a vector of L1-RSRP differences for the top-N Set A beams. In some aspects, the indication of the inference accuracy level for the quantized AI / ML model may be or include an indication of Y. In some examples, the indication of the inference accuracy level for the quantized AI / ML model that is transmitted from the model server 650 to the network node 110 may include an indication of Y and an indication of a logical model ID associated with the AI / ML task. In such examples, the indication of the inference accuracy level that is transmitted from the model server 650 to the network node 110 may also include an indication of a physical model ID associated with the quantized AI / ML model (e.g., to register the physical model ID with the network node 110 for the AI / ML task) .

[0150] In some aspects, in a case in which the AI / ML task is associated with predicting a number of top beams of a set of beams (e.g., identifying which N Set A beams are the top-N beams) , the inference accuracy level may be based at least in part on a percentage of the top beams accurately predicted by the quantized AI / ML model based on model inputs included in the reference dataset. For example, the inference accuracy level may be based at least in part on the percentage of the top beams accurately predicted by the quantized AI / ML model based on model inputs included in the reference dataset relative to the top beams predicted by the floating-point AI / ML model based on the model inputs included in the dataset. In this example, the reference accuracy level Y may be the percentage of the top-N Set A beams accurately predicted by applying the quantized AI / ML model to the reference dataset, as compared with the top-N Set A beams predicted by applying the floating-point AI / ML model to the reference dataset. In some aspects, this value of Y may be indicated in the indication of the inference accuracy level for the quantized AI / ML model that is transmitted from the model server 650 to the network node 110.

[0151] In some aspects, the model server 650 may determine the inference accuracy level (e.g., the value of Y) for the quantized AI / ML model by applying the quantized AI / ML model and the floating-point AI / ML model to the reference dataset. The reference dataset used for determining the inference accuracy may include model inputs and ground truth (e.g., genie-aided) model output values. In some aspects, the reference dataset used for determining the inference accuracy level may be pre-defined (e.g., specified in a wireless communication standard) . For example, a wireless communication standard may specify different reference datasets for different AI / ML  tasks (e.g., different AI / ML functionalities) . In some aspects, the reference dataset may be indicated via signaling transmitted by the network node 110. For example, the network node 110 may transmit, and the model server 650 may receive, the reference dataset. In some examples, the network node 110 may signal the dataset to one or more third-party entities (e.g., the model server 650 and / or one or more other models servers) where quantized AI / ML models may be downloaded by UEs. In some other aspects, the model server 650 may transmit, and the network node 110 may receive, the reference dataset. In this case, the model server 650 may transmit an indication of a logical model ID associated with the floating-point AI / ML model and / or AI / ML task together with the reference dataset. For example, one or more third-party entities (e.g., including the model server 650 and / or one or more other model servers) may signal reference datasets used by the third-party entities to the network node 110, together with model IDs (e.g., logical model IDs) corresponding to the AI / ML tasks associated with the reference datasets.

[0152] As further shown in Fig. 8, and by reference number 815, the UE 120 may transmit, and the network node 110 may receive, and indication of a model ID (e.g., a physical model ID) associated with the quantized AI / ML model to be used by the UE 120 (e.g., the quantized AI / ML model received at the UE 120 from the model server 650) . The network node 110 may identify the inference accuracy level of the quantized AI / ML model to be used by the UE 120 based at least in part on the indication of the physical model ID received from the UE 120 and the indication of the inference accuracy level corresponding to that physical model ID received from the model server 650. In some examples, the UE 120 may also transmit, to the network node 110, an indication of the logical model ID corresponding to the AI / ML task and / or the floating-point AI / ML model associated with the quantized AI / ML model. In some aspects, the UE 120 may transmit the indication of the physical model ID associated with the quantized AI / ML model to the network node 110 before the UE 120 is activated (e.g., by the network node 110) with the quantized AI / ML model and / or the AI / ML task. For example, the UE 120 may transmit the indication of the physical model ID associated with the quantized AI / ML model (e.g., together with an indication of the logical model ID associated with the AI / ML task) during initial access as UE capability information.

[0153] As further shown in Fig. 8, and by reference number 820, the network node 110 may transmit, and the UE 120 may receive, signaling associated with AI / ML performance monitoring based at least in part on the indication of the inference  accuracy level for the quantized AI / ML model. In some aspects, the network node 110 may determine one or more AI / ML performance monitoring parameters for the UE 120 based at least in part on the indicated inference accuracy level for the quantized AI / ML model to be used by the UE 120, and the network node 110 may transmit, to the UE 120, signaling that indicates configuration information including the one or more AI / ML performance monitoring parameters and / or scheduling information based at least in part on the one or more AI / ML performance parameters.

[0154] In some aspects, the network node 110 may schedule transmission of RSs (e.g., auxiliary RSs) for prediction accuracy verification for the UE 120 based at least in part on the indicated inference accuracy level for the quantized AI / ML model. For example, the network node 110 may schedule the transmission, to the UE 120, of the RSs for the prediction accuracy verification more or less frequently (e.g., as compared to a default scheduling frequency and / or as compared to RSs scheduled for other UEs) based at least in part on the indicated inference accuracy level for the quantized AI / ML model. In such examples, the signaling associated with the AI / ML performance monitoring may include signaling that indicates the scheduling of the RSs to be transmitted to the UE 120 and / or the transmission, from the network node 110 to the UE 120, of the RSs for the prediction accuracy verification.

[0155] In some aspects, the network node 110 may determine / configure an error threshold value for the UE 120 to be used to determine whether a performance alert message should be sent to the UE 120. For example, the network node 110 may determine / configure higher or lower error threshold values for different UEs associated with an AI / ML task based at least in part on the respective inference accuracy levels for the quantized AI / ML models to be used by the different UEs. In such examples, the signaling associated with the AI / ML performance monitoring may include signaling that indicates or configures the error threshold value for the UE 120 and / or a performance alert message transmitted to the UE 120 based at least in part on the error threshold value determined / configured for the UE 120.

[0156] As indicated above, Fig. 8 is provided as an example. Other examples may differ from what is described with respect to Fig. 8.

[0157] Fig. 9 is a diagram illustrating an example process 900 performed, for example, at a UE or an apparatus of a UE, in accordance with the present disclosure. Example process 900 is an example where the apparatus or the UE (e.g., UE 120)  performs operations associated with reporting of inference accuracies of quantized models for AI / ML performance monitoring.

[0158] As shown in Fig. 9, in some aspects, process 900 may include transmitting, to a network node, an indication of an inference accuracy level associated with a quantized AI / ML model for an AI / ML task, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to a floating-point AI / ML model for the AI / ML task (block 910) . For example, the UE (e.g., using transmission component 1204 and / or communication manager 1206, depicted in Fig. 12) may transmit, to a network node, an indication of an inference accuracy level associated with a quantized AI / ML model for an AI / ML task, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to a floating-point AI / ML model for the AI / ML task, as described above.

[0159] As further shown in Fig. 9, in some aspects, process 900 may include receiving, from the network node, signaling associated with AI / ML performance monitoring based at least in part on the indication of the inference accuracy level (block 920) . For example, the UE (e.g., using reception component 1202 and / or communication manager 1206, depicted in Fig. 12) may receive, from the network node, signaling associated with AI / ML performance monitoring based at least in part on the indication of the inference accuracy level, as described above.

[0160] Process 900 may include additional aspects, such as any single aspect or any combination of aspects described below and / or in connection with one or more other processes described elsewhere herein.

[0161] In a first aspect, the indication of the inference accuracy level indicates a selected candidate inference accuracy level, from a plurality of candidate inference accuracy levels associated with the AI / ML task.

[0162] In a second aspect, alone or in combination with the first aspect, the plurality of candidate inference accuracy levels are associated with different model identifiers, and the selected candidate inference accuracy level is associated with a model identifier of the quantized AI / ML model.

[0163] In a third aspect, alone or in combination with one or more of the first and second aspects, process 900 includes receiving, from the network node, configuration information indicating the plurality of candidate inference accuracy levels associated with the AI / ML task.

[0164] In a fourth aspect, alone or in combination with one or more of the first through third aspects, the AI / ML task is associated with predicting beam measurements for a set of beams, and the inference accuracy level is based at least in part on differences between first beam measurements predicted by the quantized AI / ML model for a subset of the set of beams based on model inputs included in a dataset and second beam measurements predicted by the floating-point AI / ML model for the subset of the set of beams based on the model inputs included in the dataset.

[0165] In a fifth aspect, alone or in combination with one or more of the first through fourth aspects, the subset of the set of beams includes a number of ground truth top beams associated with the dataset.

[0166] In a sixth aspect, alone or in combination with one or more of the first through fifth aspects, the first beam measurements include first L1-RSRP measurements, and the second beam measurements includes second L1-RSRP measurements.

[0167] In a seventh aspect, alone or in combination with one or more of the first through sixth aspects, the AI / ML task is associated with predicting a number of top beams, of a set of beams, and the inference accuracy level is based at least in part on a percentage of the top beams accurately predicted by the quantized AI / ML model based on model inputs included in a dataset relative to the top beams predicted by the floating-point AI / ML model based on the model inputs included in the dataset.

[0168] Although Fig. 9 shows example blocks of process 900, in some aspects, process 900 may include additional blocks, fewer blocks, different blocks, or differently arranged blocks than those depicted in Fig. 9. Additionally, or alternatively, two or more of the blocks of process 900 may be performed in parallel.

[0169] Fig. 10 is a diagram illustrating an example process 1000 performed, for example, at a network node or an apparatus of a network node, in accordance with the present disclosure. Example process 1000 is an example where the apparatus or the network node (e.g., network node 110) performs operations associated with reporting of inference accuracies of quantized models for AI / ML performance monitoring.

[0170] As shown in Fig. 10, in some aspects, process 1000 may include receiving an indication of an inference accuracy level associated with a quantized AI / ML model to be used by a UE for an AI / ML task, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to a floating-point AI / ML model for the AI / ML task (block 1010) . For example, the network node (e.g., using reception component 1302 and / or communication manager 1306, depicted in Fig.  13) may receive an indication of an inference accuracy level associated with a quantized AI / ML model to be used by a UE for an AI / ML task, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to a floating-point AI / ML model for the AI / ML task, as described above.

[0171] As further shown in Fig. 10, in some aspects, process 1000 may include transmitting, to the UE, signaling associated with AI / ML performance monitoring based at least in part on the indication of the inference accuracy level (block 1020) . For example, the network node (e.g., using transmission component 1304 and / or communication manager 1306, depicted in Fig. 13) may transmit, to the UE, signaling associated with AI / ML performance monitoring based at least in part on the indication of the inference accuracy level, as described above.

[0172] Process 1000 may include additional aspects, such as any single aspect or any combination of aspects described below and / or in connection with one or more other processes described elsewhere herein.

[0173] In a first aspect, receiving the indication of the inference accuracy level includes receiving the indication of the inference accuracy level from the UE.

[0174] In a second aspect, alone or in combination with the first aspect, receiving the indication of the inference accuracy level includes receiving the indication of the inference accuracy level from a server device associated with the floating-point AI / ML model and the quantized AI / ML model.

[0175] In a third aspect, alone or in combination with one or more of the first and second aspects, receiving the indication of the inference accuracy level from the server device associated with the floating-point AI / ML model and the quantized AI / ML model includes receiving, from the server device, indications of respective inference accuracy levels for a plurality of quantized AI / ML models for the AI / ML task, wherein the plurality of quantized AI / ML models includes the quantized AI / ML model.

[0176] In a fourth aspect, alone or in combination with one or more of the first through third aspects, process 1000 includes receiving, from the UE, an indication of a model identifier associated with the quantized AI / ML model.

[0177] In a fifth aspect, alone or in combination with one or more of the first through fourth aspects, the indication of the inference accuracy level indicates a selected candidate inference accuracy level, from a plurality of candidate inference accuracy levels associated with the AI / ML task.

[0178] In a sixth aspect, alone or in combination with one or more of the first through fifth aspects, the plurality of candidate inference accuracy levels are associated with different model identifiers, and the selected candidate inference accuracy level is associated with a model identifier of the quantized AI / ML model.

[0179] In a seventh aspect, alone or in combination with one or more of the first through sixth aspects, process 1000 includes transmitting, to the UE, configuration information indicating the plurality of candidate inference accuracy levels associated with the AI / ML task.

[0180] In an eighth aspect, alone or in combination with one or more of the first through seventh aspects, the AI / ML task is associated with predicting beam measurements for a set of beams, and the inference accuracy level is based at least in part on differences between first beam measurements predicted by the quantized AI / ML model for a subset of the set of beams based on model inputs included in a dataset and second beam measurements predicted by the floating-point AI / ML model for the subset of the set of beams based on the model inputs included in the dataset.

[0181] In a ninth aspect, alone or in combination with one or more of the first through eighth aspects, the subset of the set of beams includes a number of ground truth top beams associated with the dataset.

[0182] In a tenth aspect, alone or in combination with one or more of the first through ninth aspects, the first beam measurements include first L1-RSRP measurements, and the second beam measurements includes second L1-RSRP measurements.

[0183] In an eleventh aspect, alone or in combination with one or more of the first through tenth aspects, the AI / ML task is associated with predicting a number of top beams, of a set of beams, and the inference accuracy level is based at least in part on a percentage of the top beams accurately predicted by the quantized AI / ML model based on model inputs included in a dataset relative to the top beams predicted by the floating-point AI / ML model based on the model inputs included in the dataset.

[0184] Although Fig. 10 shows example blocks of process 1000, in some aspects, process 1000 may include additional blocks, fewer blocks, different blocks, or differently arranged blocks than those depicted in Fig. 10. Additionally, or alternatively, two or more of the blocks of process 1000 may be performed in parallel.

[0185] Fig. 11 is a diagram illustrating an example process 1100 performed, for example, at a server device or an apparatus of a server device, in accordance with the present disclosure. Example process 1100 is an example where the apparatus or the  server device (e.g., model server 650) performs operations associated with reporting of inference accuracies of quantized models for AI / ML performance monitoring.

[0186] As shown in Fig. 11, in some aspects, process 1100 may include transmitting, to a UE, a quantized AI / ML model for an AI / ML task, wherein the quantized AI / ML model is associated with a floating-point AI / ML model for the AI / ML task (block 1110) . For example, the server device (e.g., using transmission component 1404 and / or communication manager 1406, depicted in Fig. 14) may transmit, to a UE, a quantized AI / ML model for an AI / ML task, wherein the quantized AI / ML model is associated with a floating-point AI / ML model for the AI / ML task, as described above.

[0187] As further shown in Fig. 11, in some aspects, process 1100 may include transmitting an indication of an inference accuracy level associated with the quantized AI / ML model, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to the floating-point AI / ML model (block 1120) . For example, the server device (e.g., using transmission component 1404 and / or communication manager 1406, depicted in Fig. 14) may transmit an indication of an inference accuracy level associated with the quantized AI / ML model, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to the floating-point AI / ML model, as described above.

[0188] Process 1100 may include additional aspects, such as any single aspect or any combination of aspects described below and / or in connection with one or more other processes described elsewhere herein.

[0189] In a first aspect, transmitting the indication of the inference accuracy level includes transmitting the indication of the inference accuracy level to a network node.

[0190] In a second aspect, alone or in combination with the first aspect, transmitting the indication of the inference accuracy level to the network node includes transmitting, to the network node, indications of respective inference accuracy levels for a plurality of quantized AI / ML models for the AI / ML task, wherein the plurality of quantized AI / ML models includes the quantized AI / ML model.

[0191] In a third aspect, alone or in combination with one or more of the first and second aspects, transmitting the indication of the inference accuracy level includes transmitting the indication of the inference accuracy level to the UE.

[0192] In a fourth aspect, alone or in combination with one or more of the first through third aspects, the indication of the inference accuracy level indicates a selected  candidate inference accuracy level, from a plurality of candidate inference accuracy levels associated with the AI / ML task.

[0193] In a fifth aspect, alone or in combination with one or more of the first through fourth aspects, the plurality of candidate inference accuracy levels are associated with different model identifiers, and the selected candidate inference accuracy level is associated with a model identifier of the quantized AI / ML model.

[0194] In a sixth aspect, alone or in combination with one or more of the first through fifth aspects, the AI / ML task is associated with predicting beam measurements for a set of beams, and the inference accuracy level is based at least in part on differences between first beam measurements predicted by the quantized AI / ML model for a subset of the set of beams based on model inputs included in a dataset and second beam measurements predicted by the floating-point AI / ML model for the subset of the set of beams based on the model inputs included in the dataset.

[0195] In a seventh aspect, alone or in combination with one or more of the first through sixth aspects, the subset of the set of beams includes a number of ground truth top beams associated with the dataset.

[0196] In an eighth aspect, alone or in combination with one or more of the first through seventh aspects, the first beam measurements include first L1-RSRP measurements, and the second beam measurements includes second L1-RSRP measurements.

[0197] In a ninth aspect, alone or in combination with one or more of the first through eighth aspects, the AI / ML task is associated with predicting a number of top beams, of a set of beams, and the inference accuracy level is based at least in part on a percentage of the top beams accurately predicted by the quantized AI / ML model based on model inputs included in a dataset relative to the top beams predicted by the floating-point AI / ML model based on the model inputs included in the dataset.

[0198] Although Fig. 11 shows example blocks of process 1100, in some aspects, process 1100 may include additional blocks, fewer blocks, different blocks, or differently arranged blocks than those depicted in Fig. 11. Additionally, or alternatively, two or more of the blocks of process 1100 may be performed in parallel.

[0199] Fig. 12 is a diagram illustrating an example apparatus 1200 for wireless communication, in accordance with the present disclosure. The apparatus 1200 may be a UE, or a UE may include the apparatus 1200. In some aspects, the apparatus 1200 includes a reception component 1202, a transmission component 1204, and / or a  communication manager 1206, which may be in communication with one another (for example, via one or more buses and / or one or more other components) . In some aspects, the communication manager 1206 is the communication manager 140 described in connection with Fig. 1. As shown, the apparatus 1200 may communicate with another apparatus 1208, such as a UE, a network node (such as a CU, a DU, an RU, or a base station) or a server device, using the reception component 1202 and the transmission component 1204.

[0200] In some aspects, the apparatus 1200 may be configured to perform one or more operations described herein in connection with Figs. 4-8. Additionally, or alternatively, the apparatus 1200 may be configured to perform one or more processes described herein, such as process 900 of Fig. 9, or a combination thereof. In some aspects, the apparatus 1200 and / or one or more components shown in Fig. 12 may include one or more components of the UE described in connection with Fig. 2. Additionally, or alternatively, one or more components shown in Fig. 12 may be implemented within one or more components described in connection with Fig. 2. Additionally, or alternatively, one or more components of the set of components may be implemented at least in part as software stored in one or more memories. For example, a component (or a portion of a component) may be implemented as instructions or code stored in a non-transitory computer-readable medium and executable by one or more controllers or one or more processors to perform the functions or operations of the component.

[0201] The reception component 1202 may receive communications, such as reference signals, control information, data communications, or a combination thereof, from the apparatus 1208. The reception component 1202 may provide received communications to one or more other components of the apparatus 1200. In some aspects, the reception component 1202 may perform signal processing on the received communications (such as filtering, amplification, demodulation, analog-to-digital conversion, demultiplexing, deinterleaving, de-mapping, equalization, interference cancellation, or decoding, among other examples) , and may provide the processed signals to the one or more other components of the apparatus 1200. In some aspects, the reception component 1202 may include one or more antennas, one or more modems, one or more demodulators, one or more MIMO detectors, one or more receive processors, one or more controllers / processors, one or more memories, or a combination thereof, of the UE described in connection with Fig. 2.

[0202] The transmission component 1204 may transmit communications, such as reference signals, control information, data communications, or a combination thereof, to the apparatus 1208. In some aspects, one or more other components of the apparatus 1200 may generate communications and may provide the generated communications to the transmission component 1204 for transmission to the apparatus 1208. In some aspects, the transmission component 1204 may perform signal processing on the generated communications (such as filtering, amplification, modulation, digital-to-analog conversion, multiplexing, interleaving, mapping, or encoding, among other examples) , and may transmit the processed signals to the apparatus 1208. In some aspects, the transmission component 1204 may include one or more antennas, one or more modems, one or more modulators, one or more transmit MIMO processors, one or more transmit processors, one or more controllers / processors, one or more memories, or a combination thereof, of the UE described in connection with Fig. 2. In some aspects, the transmission component 1204 may be co-located with the reception component 1202 in one or more transceivers.

[0203] The communication manager 1206 may support operations of the reception component 1202 and / or the transmission component 1204. For example, the communication manager 1206 may receive information associated with configuring reception of communications by the reception component 1202 and / or transmission of communications by the transmission component 1204. Additionally, or alternatively, the communication manager 1206 may generate and / or provide control information to the reception component 1202 and / or the transmission component 1204 to control reception and / or transmission of communications.

[0204] The transmission component 1204 may transmit, to a network node, an indication of an inference accuracy level associated with a quantized AI / ML model for an AI / ML task, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to a floating-point AI / ML model for the AI / ML task. The reception component 1202 may receive, from the network node, signaling associated with AI / ML performance monitoring based at least in part on the indication of the inference accuracy level.

[0205] The reception component 1202 may receive, from the network node, configuration information indicating the plurality of candidate inference accuracy levels associated with the AI / ML task.

[0206] The number and arrangement of components shown in Fig. 12 are provided as an example. In practice, there may be additional components, fewer components, different components, or differently arranged components than those shown in Fig. 12. Furthermore, two or more components shown in Fig. 12 may be implemented within a single component, or a single component shown in Fig. 12 may be implemented as multiple, distributed components. Additionally, or alternatively, a set of (one or more) components shown in Fig. 12 may perform one or more functions described as being performed by another set of components shown in Fig. 12.

[0207] Fig. 13 is a diagram illustrating an example apparatus 1300 for wireless communication, in accordance with the present disclosure. The apparatus 1300 may be a network node, or a network node may include the apparatus 1300. In some aspects, the apparatus 1300 includes a reception component 1302, a transmission component 1304, and / or a communication manager 1306, which may be in communication with one another (for example, via one or more buses and / or one or more other components) . In some aspects, the communication manager 1306 is the communication manager 150 described in connection with Fig. 1. As shown, the apparatus 1300 may communicate with another apparatus 1308, such as a UE or a network node (such as a CU, a DU, an RU, or a base station) , using the reception component 1302 and the transmission component 1304.

[0208] In some aspects, the apparatus 1300 may be configured to perform one or more operations described herein in connection with Figs. 4-8. Additionally, or alternatively, the apparatus 1300 may be configured to perform one or more processes described herein, such as process 1000 of Fig. 10, or a combination thereof. In some aspects, the apparatus 1300 and / or one or more components shown in Fig. 13 may include one or more components of the network node described in connection with Fig. 2. Additionally, or alternatively, one or more components shown in Fig. 13 may be implemented within one or more components described in connection with Fig. 2. Additionally, or alternatively, one or more components of the set of components may be implemented at least in part as software stored in one or more memories. For example, a component (or a portion of a component) may be implemented as instructions or code stored in a non-transitory computer-readable medium and executable by one or more controllers or one or more processors to perform the functions or operations of the component.

[0209] The reception component 1302 may receive communications, such as reference signals, control information, data communications, or a combination thereof, from the apparatus 1308. The reception component 1302 may provide received communications to one or more other components of the apparatus 1300. In some aspects, the reception component 1302 may perform signal processing on the received communications (such as filtering, amplification, demodulation, analog-to-digital conversion, demultiplexing, deinterleaving, de-mapping, equalization, interference cancellation, or decoding, among other examples) , and may provide the processed signals to the one or more other components of the apparatus 1300. In some aspects, the reception component 1302 may include one or more antennas, one or more modems, one or more demodulators, one or more MIMO detectors, one or more receive processors, one or more controllers / processors, one or more memories, or a combination thereof, of the network node described in connection with Fig. 2. In some aspects, the reception component 1302 and / or the transmission component 1304 may include or may be included in a network interface. The network interface may be configured to obtain and / or output signals for the apparatus 1300 via one or more communications links, such as a backhaul link, a midhaul link, and / or a fronthaul link.

[0210] The transmission component 1304 may transmit communications, such as reference signals, control information, data communications, or a combination thereof, to the apparatus 1308. In some aspects, one or more other components of the apparatus 1300 may generate communications and may provide the generated communications to the transmission component 1304 for transmission to the apparatus 1308. In some aspects, the transmission component 1304 may perform signal processing on the generated communications (such as filtering, amplification, modulation, digital-to-analog conversion, multiplexing, interleaving, mapping, or encoding, among other examples) , and may transmit the processed signals to the apparatus 1308. In some aspects, the transmission component 1304 may include one or more antennas, one or more modems, one or more modulators, one or more transmit MIMO processors, one or more transmit processors, one or more controllers / processors, one or more memories, or a combination thereof, of the network node described in connection with Fig. 2. In some aspects, the transmission component 1304 may be co-located with the reception component 1302 in one or more transceivers.

[0211] The communication manager 1306 may support operations of the reception component 1302 and / or the transmission component 1304. For example, the  communication manager 1306 may receive information associated with configuring reception of communications by the reception component 1302 and / or transmission of communications by the transmission component 1304. Additionally, or alternatively, the communication manager 1306 may generate and / or provide control information to the reception component 1302 and / or the transmission component 1304 to control reception and / or transmission of communications.

[0212] The reception component 1302 may receive an indication of an inference accuracy level associated with a quantized AI / ML model to be used by a UE for an AI / ML task, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to a floating-point AI / ML model for the AI / ML task. The transmission component 1304 may transmit, to the UE, signaling associated with AI / ML performance monitoring based at least in part on the indication of the inference accuracy level.

[0213] The reception component 1302 may receive, from the UE, an indication of a model identifier associated with the quantized AI / ML model.

[0214] The transmission component 1304 may transmit, to the UE, configuration information indicating the plurality of candidate inference accuracy levels associated with the AI / ML task.

[0215] The number and arrangement of components shown in Fig. 13 are provided as an example. In practice, there may be additional components, fewer components, different components, or differently arranged components than those shown in Fig. 13. Furthermore, two or more components shown in Fig. 13 may be implemented within a single component, or a single component shown in Fig. 13 may be implemented as multiple, distributed components. Additionally, or alternatively, a set of (one or more) components shown in Fig. 13 may perform one or more functions described as being performed by another set of components shown in Fig. 13.

[0216] Fig. 14 is a diagram illustrating an example apparatus 1400 for wireless communication, in accordance with the present disclosure. The apparatus 1400 may be a server device, or a server device may include the apparatus 1400. In some aspects, the apparatus 1400 includes a reception component 1402, a transmission component 1404, and / or a communication manager 1406, which may be in communication with one another (for example, via one or more buses and / or one or more other components) . In some aspects, the communication manager 1406 is the communication manager 150 described in connection with Fig. 1. As shown, the apparatus 1400 may communicate  with another apparatus 1408, such as a UE or a network node (such as a CU, a DU, an RU, or a base station) , using the reception component 1402 and the transmission component 1404.

[0217] In some aspects, the apparatus 1400 may be configured to perform one or more operations described herein in connection with Figs. 4-8. Additionally, or alternatively, the apparatus 1400 may be configured to perform one or more processes described herein, such as process 1100 of Fig. 11, or a combination thereof. In some aspects, the apparatus 1400 and / or one or more components shown in Fig. 14 may include one or more components of the server device described in connection with Fig. 2. Additionally, or alternatively, one or more components shown in Fig. 14 may be implemented within one or more components described in connection with Fig. 2. Additionally, or alternatively, one or more components of the set of components may be implemented at least in part as software stored in one or more memories. For example, a component (or a portion of a component) may be implemented as instructions or code stored in a non-transitory computer-readable medium and executable by one or more controllers or one or more processors to perform the functions or operations of the component.

[0218] The reception component 1402 may receive communications, such as reference signals, control information, data communications, or a combination thereof, from the apparatus 1408. The reception component 1402 may provide received communications to one or more other components of the apparatus 1400. In some aspects, the reception component 1402 may perform signal processing on the received communications (such as filtering, amplification, demodulation, analog-to-digital conversion, demultiplexing, deinterleaving, de-mapping, equalization, interference cancellation, or decoding, among other examples) , and may provide the processed signals to the one or more other components of the apparatus 1400. In some aspects, the reception component 1402 may include one or more antennas, one or more modems, one or more demodulators, one or more MIMO detectors, one or more receive processors, one or more controllers / processors, one or more memories, or a combination thereof, of the server device described in connection with Fig. 2.

[0219] The transmission component 1404 may transmit communications, such as reference signals, control information, data communications, or a combination thereof, to the apparatus 1408. In some aspects, one or more other components of the apparatus 1400 may generate communications and may provide the generated communications to  the transmission component 1404 for transmission to the apparatus 1408. In some aspects, the transmission component 1404 may perform signal processing on the generated communications (such as filtering, amplification, modulation, digital-to-analog conversion, multiplexing, interleaving, mapping, or encoding, among other examples) , and may transmit the processed signals to the apparatus 1408. In some aspects, the transmission component 1404 may include one or more antennas, one or more modems, one or more modulators, one or more transmit MIMO processors, one or more transmit processors, one or more controllers / processors, one or more memories, or a combination thereof, of the server device described in connection with Fig. 2. In some aspects, the transmission component 1404 may be co-located with the reception component 1402 in one or more transceivers.

[0220] The communication manager 1406 may support operations of the reception component 1402 and / or the transmission component 1404. For example, the communication manager 1406 may receive information associated with configuring reception of communications by the reception component 1402 and / or transmission of communications by the transmission component 1404. Additionally, or alternatively, the communication manager 1406 may generate and / or provide control information to the reception component 1402 and / or the transmission component 1404 to control reception and / or transmission of communications.

[0221] The transmission component 1404 may transmit, to a UE, a quantized AI / ML model for an AI / ML task, wherein the quantized AI / ML model is associated with a floating-point AI / ML model for the AI / ML task. The transmission component 1404 may transmit an indication of an inference accuracy level associated with the quantized AI / ML model, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to the floating-point AI / ML model.

[0222] The number and arrangement of components shown in Fig. 14 are provided as an example. In practice, there may be additional components, fewer components, different components, or differently arranged components than those shown in Fig. 14. Furthermore, two or more components shown in Fig. 14 may be implemented within a single component, or a single component shown in Fig. 14 may be implemented as multiple, distributed components. Additionally, or alternatively, a set of (one or more) components shown in Fig. 14 may perform one or more functions described as being performed by another set of components shown in Fig. 14.

[0223] The following provides an overview of some Aspects of the present disclosure:

[0224] Aspect 1: A method of wireless communication performed by a user equipment (UE) , comprising: transmitting, to a network node, an indication of an inference accuracy level associated with a quantized artificial intelligence or machine learning (AI / ML) model for an AI / ML task, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to a floating-point AI / ML model for the AI / ML task; and receiving, from the network node, signaling associated with AI / ML performance monitoring based at least in part on the indication of the inference accuracy level.

[0225] Aspect 2: The method of Aspect 1, wherein the indication of the inference accuracy level indicates a selected candidate inference accuracy level, from a plurality of candidate inference accuracy levels associated with the AI / ML task.

[0226] Aspect 3: The method of Aspect 2, wherein the plurality of candidate inference accuracy levels are associated with different model identifiers, and wherein the selected candidate inference accuracy level is associated with a model identifier of the quantized AI / ML model.

[0227] Aspect 4: The method of any of Aspects 2-3, further comprising: receiving, from the network node, configuration information indicating the plurality of candidate inference accuracy levels associated with the AI / ML task.

[0228] Aspect 5: The method of any of Aspects 1-4, wherein the AI / ML task is associated with predicting beam measurements for a set of beams, and wherein the inference accuracy level is based at least in part on differences between first beam measurements predicted by the quantized AI / ML model for a subset of the set of beams based on model inputs included in a dataset and second beam measurements predicted by the floating-point AI / ML model for the subset of the set of beams based on the model inputs included in the dataset.

[0229] Aspect 6: The method of Aspect 5, wherein the subset of the set of beams includes a number of ground truth top beams associated with the dataset.

[0230] Aspect 7: The method of any of Aspects 5-6, wherein the first beam measurements include first layer 1 reference signal received power (L1-RSRP) measurements, and wherein the second beam measurements includes second L1-RSRP measurements.

[0231] Aspect 8: The method of any of Aspects 1-4, wherein the AI / ML task is associated with predicting a number of top beams, of a set of beams, and wherein the inference accuracy level is based at least in part on a percentage of the top beams  accurately predicted by the quantized AI / ML model based on model inputs included in a dataset relative to the top beams predicted by the floating-point AI / ML model based on the model inputs included in the dataset.

[0232] Aspect 9: A method of wireless communication performed by a network node, comprising: receiving an indication of an inference accuracy level associated with a quantized artificial intelligence or machine learning (AI / ML) model to be used by a user equipment (UE) for an AI / ML task, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to a floating-point AI / ML model for the AI / ML task; and transmitting, to the UE, signaling associated with AI / ML performance monitoring based at least in part on the indication of the inference accuracy level.

[0233] Aspect 10: The method of Aspect 9, wherein receiving the indication of the inference accuracy level comprises: receiving the indication of the inference accuracy level from the UE.

[0234] Aspect 11: The method of Aspect 9, wherein receiving the indication of the inference accuracy level comprises: receiving the indication of the inference accuracy level from a server device associated with the floating-point AI / ML model and the quantized AI / ML model.

[0235] Aspect 12: The method of Aspect 11, wherein receiving the indication of the inference accuracy level from the server device associated with the floating-point AI / ML model and the quantized AI / ML model comprises: receiving, from the server device, indications of respective inference accuracy levels for a plurality of quantized AI / ML models for the AI / ML task, wherein the plurality of quantized AI / ML models includes the quantized AI / ML model.

[0236] Aspect 13: The method of Aspect 12, further comprising: receiving, from the UE, an indication of a model identifier associated with the quantized AI / ML model.

[0237] Aspect 14: The method of any of Aspects 9-13, wherein the indication of the inference accuracy level indicates a selected candidate inference accuracy level, from a plurality of candidate inference accuracy levels associated with the AI / ML task.

[0238] Aspect 15: The method of Aspect 14, wherein the plurality of candidate inference accuracy levels are associated with different model identifiers, and wherein the selected candidate inference accuracy level is associated with a model identifier of the quantized AI / ML model.

[0239] Aspect 16: The method of any of Aspects 14-15, further comprising: transmitting, to the UE, configuration information indicating the plurality of candidate inference accuracy levels associated with the AI / ML task.

[0240] Aspect 17: The method of any of Aspects 9-16, wherein the AI / ML task is associated with predicting beam measurements for a set of beams, and wherein the inference accuracy level is based at least in part on differences between first beam measurements predicted by the quantized AI / ML model for a subset of the set of beams based on model inputs included in a dataset and second beam measurements predicted by the floating-point AI / ML model for the subset of the set of beams based on the model inputs included in the dataset.

[0241] Aspect 18: The method of Aspect 17, wherein the subset of the set of beams includes a number of ground truth top beams associated with the dataset.

[0242] Aspect 19: The method of any of Aspects 17-18, wherein the first beam measurements include first layer 1 reference signal received power (L1-RSRP) measurements, and wherein the second beam measurements includes second L1-RSRP measurements.

[0243] Aspect 20: The method of any of Aspects 9-16, wherein the AI / ML task is associated with predicting a number of top beams, of a set of beams, and wherein the inference accuracy level is based at least in part on a percentage of the top beams accurately predicted by the quantized AI / ML model based on model inputs included in a dataset relative to the top beams predicted by the floating-point AI / ML model based on the model inputs included in the dataset.

[0244] Aspect 21: A method of wireless communication performed by a server device, comprising: transmitting, to a user equipment (UE) , a quantized artificial intelligence or machine learning (AI / ML) model for an AI / ML task, wherein the quantized AI / ML model is associated with a floating-point AI / ML model for the AI / ML task; and transmitting an indication of an inference accuracy level associated with the quantized AI / ML model, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to the floating-point AI / ML model.

[0245] Aspect 22: The method of Aspect 21, wherein transmitting the indication of the inference accuracy level comprises: transmitting the indication of the inference accuracy level to a network node.

[0246] Aspect 23: The method of Aspect 22, wherein transmitting the indication of the inference accuracy level to the network node comprises: transmitting, to the network node, indications of respective inference accuracy levels for a plurality of quantized AI / ML models for the AI / ML task, wherein the plurality of quantized AI / ML models includes the quantized AI / ML model.

[0247] Aspect 24: The method of Aspect 21, wherein transmitting the indication of the inference accuracy level comprises: transmitting the indication of the inference accuracy level to the UE.

[0248] Aspect 25: The method of any of Aspects 21-24, wherein the indication of the inference accuracy level indicates a selected candidate inference accuracy level, from a plurality of candidate inference accuracy levels associated with the AI / ML task.

[0249] Aspect 26: The method of Aspect 25, wherein the plurality of candidate inference accuracy levels are associated with different model identifiers, and wherein the selected candidate inference accuracy level is associated with a model identifier of the quantized AI / ML model.

[0250] Aspect 27: The method of any of Aspects 21-26, wherein the AI / ML task is associated with predicting beam measurements for a set of beams, and wherein the inference accuracy level is based at least in part on differences between first beam measurements predicted by the quantized AI / ML model for a subset of the set of beams based on model inputs included in a dataset and second beam measurements predicted by the floating-point AI / ML model for the subset of the set of beams based on the model inputs included in the dataset.

[0251] Aspect 28: The method of Aspect 27, wherein the subset of the set of beams includes a number of ground truth top beams associated with the dataset.

[0252] Aspect 29: The method of any of Aspects 27-28, wherein the first beam measurements include first layer 1 reference signal received power (L1-RSRP) measurements, and wherein the second beam measurements includes second L1-RSRP measurements.

[0253] Aspect 30: The method of any of Aspects 21-26, wherein the AI / ML task is associated with predicting a number of top beams, of a set of beams, and wherein the inference accuracy level is based at least in part on a percentage of the top beams accurately predicted by the quantized AI / ML model based on model inputs included in a dataset relative to the top beams predicted by the floating-point AI / ML model based on the model inputs included in the dataset.

[0254] Aspect 31: An apparatus for wireless communication at a device, the apparatus comprising one or more processors; one or more memories coupled with the one or more processors; and instructions stored in the one or more memories and executable by the one or more processors to cause the apparatus to perform the method of one or more of Aspects 1-30.

[0255] Aspect 32: An apparatus for wireless communication at a device, the apparatus comprising one or more memories and one or more processors coupled to the one or more memories, the one or more processors configured to cause the device to perform the method of one or more of Aspects 1-30.

[0256] Aspect 33: An apparatus for wireless communication, the apparatus comprising at least one means for performing the method of one or more of Aspects 1-30.

[0257] Aspect 34: A non-transitory computer-readable medium storing code for wireless communication, the code comprising instructions executable by one or more processors to perform the method of one or more of Aspects 1-30.

[0258] Aspect 35: A non-transitory computer-readable medium storing a set of instructions for wireless communication, the set of instructions comprising one or more instructions that, when executed by one or more processors of a device, cause the device to perform the method of one or more of Aspects 1-30.

[0259] Aspect 36: A device for wireless communication, the device comprising a processing system that includes one or more processors and one or more memories coupled with the one or more processors, the processing system configured to cause the device to perform the method of one or more of Aspects 1-30.

[0260] Aspect 37: An apparatus for wireless communication at a device, the apparatus comprising one or more memories and one or more processors coupled to the one or more memories, the one or more processors individually or collectively configured to cause the device to perform the method of one or more of Aspects 1-30.

[0261] The foregoing disclosure provides illustration and description but is not intended to be exhaustive or to limit the aspects to the precise forms disclosed. Modifications and variations may be made in light of the above disclosure or may be acquired from practice of the aspects.

[0262] As used herein, the term “component” is intended to be broadly construed as hardware or a combination of hardware and at least one of software or firmware. “Software” shall be construed broadly to mean instructions, instruction sets, code, code  segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executables, threads of execution, procedures, or functions, among other examples, whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise. As used herein, a “processor” is implemented in hardware or a combination of hardware and software. It will be apparent that systems or methods described herein may be implemented in different forms of hardware or a combination of hardware and software. The actual specialized control hardware or software code used to implement these systems or methods is not limiting of the aspects. Thus, the operation and behavior of the systems or methods are described herein without reference to specific software code, because those skilled in the art will understand that software and hardware can be designed to implement the systems or methods based, at least in part, on the description herein. A component being configured to perform a function means that the component has a capability to perform the function, and does not require the function to be actually performed by the component, unless noted otherwise.

[0263] As used herein, “satisfying a threshold” may, depending on the context, refer to a value being greater than the threshold, greater than or equal to the threshold, less than the threshold, less than or equal to the threshold, equal to the threshold, or not equal to the threshold, among other examples.

[0264] As used herein, a phrase referring to “at least one of” a list of items refers to any combination of those items, including single members. As an example, “at least one of: a, b, or c” is intended to cover a, b, c, a + b, a + c, b + c, and a + b + c, as well as any combination with multiples of the same element (for example, a + a, a + a + a, a + a + b, a + a + c, a + b + b, a + c + c, b + b, b + b + b, b + b + c, c + c, and c + c + c, or any other ordering of a, b, and c) .

[0265] No element, act, or instruction used herein should be construed as critical or essential unless explicitly described as such. Also, as used herein, the articles “a” and “an” are intended to include one or more items and may be used interchangeably with “one or more. ” Further, as used herein, the article “the” is intended to include one or more items referenced in connection with the article “the” and may be used interchangeably with “the one or more. ” Furthermore, as used herein, the terms “set” and “group” are intended to include one or more items and may be used interchangeably with “one or more. ” Where only one item is intended, the phrase “only one” or similar language is used. Also, as used herein, the terms “has, ” “have, ” “having, ” and similar  terms are intended to be open-ended terms that do not limit an element that they modify (for example, an element “having” A may also have B) . Further, the phrase “based on” is intended to mean “based on or otherwise in association with” unless explicitly stated otherwise. Also, as used herein, the term “or” is intended to be inclusive when used in a series and may be used interchangeably with “and / or, ” unless explicitly stated otherwise (for example, if used in combination with “either” or “only one of” ) . It should be understood that “one or more” is equivalent to “at least one. ”

[0266] Even though particular combinations of features are recited in the claims or disclosed in the specification, these combinations are not intended to limit the disclosure of various aspects. Many of these features may be combined in ways not specifically recited in the claims or disclosed in the specification. The disclosure of various aspects includes each dependent claim in combination with every other claim in the claim set.

Claims

1.A user equipment (UE) for wireless communication, comprising:one or more memories; andone or more processors, coupled to the one or more memories, configured to cause the UE to:transmit, to a network node, an indication of an inference accuracy level associated with a quantized artificial intelligence or machine learning (AI / ML) model for an AI / ML task, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to a floating-point AI / ML model for the AI / ML task; andreceive, from the network node, signaling associated with AI / ML performance monitoring based at least in part on the indication of the inference accuracy level.2.The UE of claim 1, wherein the indication of the inference accuracy level indicates a selected candidate inference accuracy level, from a plurality of candidate inference accuracy levels associated with the AI / ML task.3.The UE of claim 2, wherein the plurality of candidate inference accuracy levels are associated with different model identifiers, and wherein the selected candidate inference accuracy level is associated with a model identifier of the quantized AI / ML model.4.The UE of claim 2, wherein the one or more processors are further configured to cause the UE to:receive, from the network node, configuration information indicating the plurality of candidate inference accuracy levels associated with the AI / ML task.5.The UE of claim 1, wherein the AI / ML task is associated with predicting beam measurements for a set of beams, and wherein the inference accuracy level is based at least in part on differences between first beam measurements predicted by the quantized AI / ML model for a subset of the set of beams based on model inputs included in a  dataset and second beam measurements predicted by the floating-point AI / ML model for the subset of the set of beams based on the model inputs included in the dataset.6.The UE of claim 5, wherein the subset of the set of beams includes a number of ground truth top beams associated with the dataset.7.The UE of claim 5, wherein the first beam measurements include first layer 1 reference signal received power (L1-RSRP) measurements, and wherein the second beam measurements includes second L1-RSRP measurements.8.The UE of claim 1, wherein the AI / ML task is associated with predicting a number of top beams, of a set of beams, and wherein the inference accuracy level is based at least in part on a percentage of the top beams accurately predicted by the quantized AI / ML model based on model inputs included in a dataset relative to the top beams predicted by the floating-point AI / ML model based on the model inputs included in the dataset.9.A network node for wireless communication, comprising:one or more memories; andone or more processors, coupled to the one or more memories, configured to cause the network node to:receive an indication of an inference accuracy level associated with a quantized artificial intelligence or machine learning (AI / ML) model to be used by a user equipment (UE) for an AI / ML task, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to a floating-point AI / ML model for the AI / ML task; andtransmit, to the UE, signaling associated with AI / ML performance monitoring based at least in part on the indication of the inference accuracy level.10.The network node of claim 9, wherein the one or more processors, to cause the network node to receive the indication of the inference accuracy level, are configured to cause the network node to:receive the indication of the inference accuracy level from the UE.11.The network node of claim 9, wherein the one or more processors, to cause the network node to receive the indication of the inference accuracy level, are configured to cause the network node to:receive the indication of the inference accuracy level from a server device associated with the floating-point AI / ML model and the quantized AI / ML model.12.The network node of claim 11, wherein the one or more processors, to cause the network node to receive the indication of the inference accuracy level from the server device associated with the floating-point AI / ML model and the quantized AI / ML model, are configured to cause the network node to:receive, from the server device, indications of respective inference accuracy levels for a plurality of quantized AI / ML models for the AI / ML task, wherein the plurality of quantized AI / ML models includes the quantized AI / ML model.13.The network node of claim 12, wherein the one or more processors are further configured to cause the network node to:receive, from the UE, an indication of a model identifier associated with the quantized AI / ML model.14.The network node of claim 9, wherein the indication of the inference accuracy level indicates a selected candidate inference accuracy level, from a plurality of candidate inference accuracy levels associated with the AI / ML task.15.The network node of claim 14, wherein the plurality of candidate inference accuracy levels are associated with different model identifiers, and wherein the selected candidate inference accuracy level is associated with a model identifier of the quantized AI / ML model.16.The network node of claim 14, wherein the one or more processors are further configured to cause the network node to:transmit, to the UE, configuration information indicating the plurality of candidate inference accuracy levels associated with the AI / ML task.17.The network node of claim 9, wherein the AI / ML task is associated with predicting beam measurements for a set of beams, and wherein the inference accuracy level is based at least in part on differences between first beam measurements predicted by the quantized AI / ML model for a subset of the set of beams based on model inputs included in a dataset and second beam measurements predicted by the floating-point AI / ML model for the subset of the set of beams based on the model inputs included in the dataset.18.The network node of claim 17, wherein the subset of the set of beams includes a number of ground truth top beams associated with the dataset.19.The network node of claim 17, wherein the first beam measurements include first layer 1 reference signal received power (L1-RSRP) measurements, and wherein the second beam measurements includes second L1-RSRP measurements.20.The network node of claim 9, wherein the AI / ML task is associated with predicting a number of top beams, of a set of beams, and wherein the inference accuracy level is based at least in part on a percentage of the top beams accurately predicted by the quantized AI / ML model based on model inputs included in a dataset relative to the top beams predicted by the floating-point AI / ML model based on the model inputs included in the dataset.21.A server device for wireless communication, comprising:one or more memories; andone or more processors, coupled to the one or more memories, configured to cause the server device to:transmit, to a user equipment (UE) , a quantized artificial intelligence or machine learning (AI / ML) model for an AI / ML task, wherein the quantized AI / ML model is associated with a floating-point AI / ML model for the AI / ML task; andtransmit an indication of an inference accuracy level associated with the quantized AI / ML model, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to the floating-point AI / ML model.22.The server device of claim 21, wherein the one or more processors, to cause the server device to transmit the indication of the inference accuracy level, are configured to cause the server device to:transmit the indication of the inference accuracy level to a network node.23.The server device of claim 22, wherein the one or more processors, to cause the server device to transmit the indication of the inference accuracy level to the network node, are configured to cause the server device to:transmit, to the network node, indications of respective inference accuracy levels for a plurality of quantized AI / ML models for the AI / ML task, wherein the plurality of quantized AI / ML models includes the quantized AI / ML model.24.The server device of claim 21, wherein the one or more processors, to cause the server device to transmit the indication of the inference accuracy level, are configured to cause the server device to:transmit the indication of the inference accuracy level to the UE.25.The server device of claim 21, wherein the indication of the inference accuracy level indicates a selected candidate inference accuracy level, from a plurality of candidate inference accuracy levels associated with the AI / ML task.26.The server device of claim 21, wherein the AI / ML task is associated with predicting beam measurements for a set of beams, and wherein the inference accuracy level is based at least in part on differences between first beam measurements predicted by the quantized AI / ML model for a subset of the set of beams based on model inputs included in a dataset and second beam measurements predicted by the floating-point AI / ML model for the subset of the set of beams based on the model inputs included in the dataset.27.The server device of claim 26, wherein the subset of the set of beams includes a number of ground truth top beams associated with the dataset.28.The server device of claim 27, wherein the first beam measurements include first layer 1 reference signal received power (L1-RSRP) measurements, and wherein the second beam measurements includes second L1-RSRP measurements.29.The server device of claim 21, wherein the AI / ML task is associated with predicting a number of top beams, of a set of beams, and wherein the inference accuracy level is based at least in part on a percentage of the top beams accurately predicted by the quantized AI / ML model based on model inputs included in a dataset relative to the top beams predicted by the floating-point AI / ML model based on the model inputs included in the dataset.30.A method of wireless communication performed by a user equipment (UE) , comprising:transmitting, to a network node, an indication of an inference accuracy level associated with a quantized artificial intelligence or machine learning (AI / ML) model for an AI / ML task, wherein the inference accuracy level is indicative of a relative accuracy of the quantized AI / ML model with respect to a floating-point AI / ML model for the AI / ML task; andreceiving, from the network node, signaling associated with AI / ML performance monitoring based at least in part on the indication of the inference accuracy level.

Citation Information

Patent Citations

  • Air interface test method and system based on AI / ML time domain beam prediction

    CN117241312A

  • Model monitoring method and device, communication equipment, communication system and storage medium

    CN117546509A

  • Enhancing wireless communication efficiency in 5G / 6G networks through AI / ML model management and deployment

    CN117750356A

  • Monitoring method and wireless communication device

    WO2023245515A1