Adaptive inference time for AI / ML on PHY-assisted signaling-physical layer

By calculating and signaling the inference time of AI/ML models in wireless communication networks, the problem of large differences in processing time among different devices is solved, the performance of the network and UE is optimized, and more efficient AI/ML model processing is achieved.

CN121444512APending Publication Date: 2026-01-30FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480037692.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-06
Filing Date
2024-02-29
Publication Date
2026-01-30

AI Technical Summary

Technical Problem

In wireless communication networks, the processing time of AI/ML models varies greatly, which prevents faster networks and UEs from fully utilizing their potential for better performance in terms of latency. The processing time defined by the existing 3GPP standard is the worst case and cannot meet the actual needs of different devices.

Method used

By using AI/ML models in wireless communication networks, employing function-based lifecycle management and model ID-based lifecycle management, and combining neural network attributes and hardware characteristics, signaling inference time is calculated to optimize processing time.

Benefits of technology

The processing time of AI/ML models in wireless communication networks has been optimized, avoiding worst-case latency, improving the performance of the network and UE, and meeting the actual needs of different devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121444512A_ABST
    Figure CN121444512A_ABST
Patent Text Reader

Abstract

Embodiments provide an apparatus of a wireless communication network that uses one or more artificial intelligence / machine learning (AI / ML) models for one or more use cases, where the apparatus determines an inference time of one or more of the AI / ML models to be used in one or more network entities of the wireless communication network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this application relate to the field of wireless communication, and more specifically, to wireless communication using communication-related models, such as models at the physical layer (PHY). Some embodiments relate to signaling associated with such models and / or the use or training of such models. Background Technology

[0002] Figure 1 is a schematic diagram of an example of a terrestrial wireless network 100. As shown in Figure 1(a), the terrestrial wireless network 100 includes a core network 102 and one or more radio access networks RAN1, RAN2, ..., RANN. Figure 1(b) is a schematic diagram of an example of a radio access network RANn, which may include one or more base stations gNB1 to gNB5, each base station serving a specific area around the base station schematically represented by corresponding cells 1061 to 1065. Base stations are provided to serve users within the cell. The term base station BS refers to gNB in ​​5G networks, eNB in ​​UMTS / LTE / LTE-A / LTE-A Pro, or simply BS in other mobile communication standards. Users can be fixed or mobile devices. The wireless communication system can also be accessed by mobile or fixed IoT devices connected to base stations or users. Mobile devices or IoT devices can include physical devices, ground-based vehicles such as robots or cars, aircraft such as manned or unmanned aerial vehicles (UAVs), the latter also referred to as drones, buildings and other items or devices embedded therein with electronics, software, sensors, actuators, etc., and network connections that enable these devices to collect and exchange data across existing network infrastructure. Figure 1(b) shows an exemplary view of five cells; however, RANn may include more or fewer such cells, and RANn may also include only one base station. Figure 1(b) shows two users UE1 and UE2, also referred to as user equipment (UEs), in cell 1062 and served by base station gNB2. Another user UE3 is shown in cell 1064 served by base station gNB4. Arrows 1081, 1082, and 1083 schematically represent uplink / downlink connections used for sending data from users UE1, UE2, and UE3 to base stations gNB2 and gNB4, or for sending data from base stations gNB2 and gNB4 to users UE1, UE2, and UE3. Furthermore, Figure 1(b) shows two IoT devices 1101 and 1102 in cell 1064, which can be fixed or mobile devices. IoT device 1101 accesses the wireless communication system via base station gNB4 to receive and transmit data, as schematically indicated by arrow 1121. IoT device 1102 accesses the wireless communication system via user UE3, as schematically indicated by arrow 1122. The corresponding base stations gNB1 to gNB5 can be connected to the core network 102, for example, via the S1 interface and the corresponding backhaul links 1141 to 1145, which is schematically indicated in Figure 1(b) by arrows pointing to the "core". The core network 102 can be connected to one or more external networks.Furthermore, some or all of the corresponding base stations gNB1 to gNB5 can be connected to each other, for example via the S1 or X2 interface or the XN interface in the NR, via the corresponding backhaul links 1161 to 1165, which are schematically represented in Figure 1(b) by arrows pointing to “gNBs”.

[0003] For data transmission, a physical resource grid can be used. A physical resource grid can include a set of resource elements to which various physical channels and physical signals are mapped. For example, physical channels can include physical downlink, uplink, and sidelink shared channels (PDSCH, PUSCH, PSSCH) carrying user-specific data, also known as downlink, uplink, and sidelink payload data; physical broadcast channels (PBCH) carrying, for example, Master Information Blocks (MIBs); physical downlink shared channels (PDSCH) carrying, for example, System Information Blocks (SIBs); and physical downlink, uplink, and sidelink control channels (PDCCH, PUCCH, PSSCH) carrying, for example, downlink control information (DCI), uplink control information (UCI), and sidelink control information (SCI). For uplink, physical channels, or more precisely, according to 3GPP transport channels, can further include physical random access channels (PRACH or RACH) used by the UE to access the network once the UE is synchronized and has acquired the MIB and SIB. Physical signals can include reference signals or symbols (RS), synchronization signals, etc. Resource grids can include frames or radio frames that have a specific duration in the time domain and a given bandwidth in the frequency domain. Frames can have a predetermined number of subframes of a predetermined length, for example, 1 millisecond. Depending on the cyclic prefix (CP) length, each subframe can include one or more time slots of 12 or 14 OFDM symbols. All OFDM symbols can be used for DL ​​or UL or only a subset, for example, when utilizing shortened transmission time intervals (sTTI) or a mini-time slot / non-time slot-based frame structure that includes only a few OFDM symbols.

[0004] Wireless communication systems can be any single-tone or multi-carrier system using frequency division multiplexing, such as orthogonal frequency division multiplexing (OFDM) systems, orthogonal frequency division multiple access (OFDMA) systems, or any other IFFT-based signal with or without CP, such as DFT-s-OFDM. Other waveforms can be used, such as non-orthogonal waveforms for multiple access, such as filter bank multicarrier (FBMC), generalized frequency division multiplexing (GFDM), orthogonal time-frequency spatial modulation (OTFS), or universal filtered multicarrier (UFMC). Wireless communication systems can operate, for example, according to LTE-Advanced pro standards or NR (5G) (New Radio) standards or IEEE 802.11 (WiFi) standards, such as IEEE 802.11 ax.

[0005] The wireless network or communication system depicted in Figure 1 can be a heterogeneous network with different coverage networks. For example, each macro cell includes macro base stations, such as a macro cell network of base stations gNB1 to gNB5, and small cell base stations (not shown in Figure 1), such as a network of femtocells or picocells.

[0006] In addition to the terrestrial wireless networks described above, there are also non-terrestrial wireless communication networks, including spaceborne transceivers such as satellites, and / or airborne transceivers such as unmanned aerial vehicle systems. Non-terrestrial wireless communication networks or systems can operate in a similar manner to the terrestrial systems described above with reference to Figure 1, for example, according to the LTE-Advanced Pro specification or the new NR (5G) radio standard.

[0007] In mobile communication networks, such as those described above with reference to Figure 1, such as LTE or 5G / NR networks, there may be UEs that communicate directly with each other via one or more sidelink (SL) channels, for example, using the PC5 interface. UEs that communicate directly with each other via sidelinks may include vehicles communicating directly with other vehicles (V2V communication) or vehicles communicating with other entities in the wireless communication network, such as roadside entities like traffic lights, traffic signs, or pedestrians (V2X communication). Other UEs may not be vehicle-related UEs and may include any of the devices described above. Such devices may also communicate directly with each other using SL channels (D2D communication).

[0008] When considering two UEs communicating directly with each other via a sidelink, the two UEs can be served by the same base station, allowing the base station to provide sidelink resource allocation configuration or assistance to the UEs. For example, the two UEs can be within the coverage area of ​​one of the base stations depicted in Figure 1. This is called the "in-coverage" scenario. Another scenario is called the "out-of-coverage" scenario. Note that "out-of-coverage" does not mean that the two UEs are not within one of the cells depicted in Figure 1, but rather that these UEs...

[0009] - It may not be connected to the base station, for example, they are not in an RRC connected state, so that the UE does not receive any sidelink resource allocation configuration or assistance from the base station, and / or

[0010] - It can connect to the base station; however, for one or more reasons, the base station may not provide the UE with sidelink resource allocation configuration or assistance, and / or

[0011] - It can connect to base stations that may not support NR V2X services, such as GSM, UMTS, and LTE base stations.

[0012] When considering two UEs communicating directly with each other via a sidelink, for example using a PC5 interface, one UE can also connect to a BS and relay information from the BS to the other UE via the sidelink interface. Relaying can be performed in the same frequency band (in-band relay) or in other frequency bands (out-of-band relay). In the first case, communication on the UE and communication on the sidelink can be decoupled using different time slots, as in a Time Division Duplex (TDD) system.

[0013] Figure 2 This is a schematic diagram of a scenario where two UEs that communicate directly with each other are connected to the coverage area of ​​a base station. The base station gNB has a coverage area schematically represented by circle 200, which essentially corresponds to the cell schematically represented in Figure 1. The UEs that communicate directly with each other include a first vehicle 202 and a second vehicle 204, both located within the coverage area 200 of the base station gNB. Vehicles 202 and 204 are both connected to the base station gNB, and they are also directly connected to each other via the PC5 interface. The scheduling and / or interference management of V2V services are assisted by the gNB via control signaling on the Uu interface, which is the radio interface between the base station and the UE. In other words, the gNB provides SL resource allocation configuration or assistance to the UEs, and the gNB allocates resources to be used for V2V communication on the sidelink. This configuration is also referred to as Mode 1 configuration in NR V2X or Mode 3 configuration in LTE V2X.

[0014] Figure 3 This is a schematic diagram of an out-of-coverage scenario where UEs communicating directly with each other are not connected to a base station, although they may be physically within the cell of the wireless communication network. Alternatively, some or all of the directly communicating UEs may be connected to a base station, but the base station does not provide SL resource allocation configuration or assistance. Three vehicles 206, 208, and 210 are shown communicating directly with each other via a sidelink, for example, using a PC5 interface. V2V service scheduling and / or interference management are based on algorithms implemented between the vehicles. This configuration is also referred to as Mode 2 configuration in NR V2X or Mode 4 configuration in LTE V2X. As described above, as an out-of-coverage scenario... Figure 3 The scenario in the text does not necessarily mean that the corresponding Mode 2 UE (in NR) or Mode 4 UE (in LTE) is outside the coverage area of ​​the base station. Rather, it means that the corresponding Mode 2 UE (in NR) or Mode 4 UE (in LTE) is not served by the base station, is not connected to the base station in the coverage area, or is connected to the base station but does not receive SL resource allocation configuration or assistance from the base station. Therefore, the following situation may exist: Figure 2 Within the coverage area 200 shown, in addition to NR mode 1 or LTE mode 3 UEs 202 and 204, there are also NR mode 2 or LTE mode 4 UEs 206, 208 and 210.

[0015] Of course, the first vehicle 202 may also be covered by the gNB, i.e., connected to the gNB via Uu, while the second vehicle 204 is not covered by the gNB and is only connected to the first vehicle 202 via the PC5 interface, or the second vehicle is connected to the first vehicle 202 via the PC5 interface but connected to another gNB via Uu. This will be from Figure 4 and Figure 5 It became clear during the discussion.

[0016] Figure 4 This is a schematic diagram of a scenario where two UEs communicate directly with each other, with only one UE connected to the base station. The base station gNB has a coverage area schematically represented by circle 200, which essentially corresponds to the cell schematically represented in Figure 1. The UEs communicating directly with each other include a first vehicle 202 and a second vehicle 204, with only the first vehicle 202 within the coverage area 200 of the base station gNB. The two vehicles 202 and 204 are directly connected to each other via a PC5 interface.

[0017] Figure 5 This is a schematic diagram of a scenario where two UEs communicate directly with each other, and the two UEs are connected to different base stations. The first base station gNB1 has a coverage area schematically represented by a first circle 2001, and the second base station gNB2 has a coverage area schematically represented by a second circle 2002. The UEs communicating directly with each other include a first vehicle 202 and a second vehicle 204. The first vehicle 202 is located within the coverage area 2001 of the first base station gNB1 and is connected to the first base station gNB1 via a Uu interface. The second vehicle 204 is located within the coverage area 2002 of the second base station gNB2 and is connected to the second base station gNB2 via a Uu interface.

[0018] For wireless communication systems as described above, machine learning schemes for various use cases, such as beam prediction, CSI prediction, CSI compression, and positioning, are discussed in 3GPP RAN1, while machine learning schemes for mobility and network enhancement are discussed in 3GPP RAN2 and RAN. However, integrating such schemes into 5G systems is not straightforward. In particular, AI / ML schemes can have very different complexities, and furthermore, UE capabilities can vary significantly between different vendors and devices. This introduces the problem that processing times for different AI / ML networks on different devices can vary considerably. However, the processing times currently defined in 3GPP standards always consider worst-case performance. In the case of AI / ML, this would mean that faster networks and faster UEs do not necessarily benefit from their better performance in terms of latency.

[0019] Therefore, there is a need to enhance the use of AI / ML models in wireless communication networks.

[0020] It should be noted that the information in the foregoing sections is only used to enhance the understanding of the background of the present invention, and therefore may include information that does not constitute prior art and is known to those skilled in the art. Attached Figure Description

[0021] Embodiments of the present invention are described herein with reference to the accompanying drawings.

[0022] Figure 1 shows a schematic diagram of an example of a wireless communication system;

[0023] Figure 2 This is a schematic diagram of a scenario where UEs that communicate directly with each other are connected to the base station within its coverage area;

[0024] Figure 3 This is a schematic diagram of an out-of-coverage scenario where UEs that communicate directly with each other do not receive SL resource allocation configuration or assistance from the base station;

[0025] Figure 4 This is a schematic diagram of a partial out-of-coverage scenario, in which some UEs that are communicating directly with each other do not receive SL resource allocation configuration or assistance from the base station;

[0026] Figure 5 This is a schematic diagram of a scenario where UEs that communicate directly with each other are connected to different base stations within the coverage area;

[0027] Figure 6 This is a diagram illustrating the worst-case processing time in a wireless communication scenario;

[0028] Figure 7 A schematic diagram of a typical model of a neural network in conjunction with an embodiment is shown;

[0029] Figure 8 This is a schematic diagram of a wireless communication system according to an embodiment, including transceivers such as base stations or repeaters, and multiple communication devices such as UEs.

[0030] Figure 9a shows a schematic diagram of signaling between the gNB and the UE according to an embodiment;

[0031] Figure 9b shows a schematic diagram of signaling between the first UE and the second UE according to an embodiment;

[0032] Figure 10 The illustration shows a schematic diagram of the task solved by the embodiments described herein, such as a possible mapping of AI / ML functions to one or more AI / ML processors;

[0033] Figures 11a-d show schematic block diagrams of embodiments for training and transmitting models according to the embodiments; and

[0034] Figure 12Examples of computer systems on which the units or modules and steps of the method described in the present invention can be performed are shown. Detailed Implementation

[0035] In the following description, the same or equivalent elements or elements having the same or equivalent functions are indicated by the same or equivalent reference numerals.

[0036] In the following description, numerous details are set forth to provide a more thorough explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention can be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the invention. Furthermore, unless otherwise specifically stated, features of the different embodiments described below can be combined with each other.

[0037] like Figure 6 As shown, the processing time defined in the specification, such as processing time 1002, is the worst-case processing time. This is because the processing time needs to be defined as the time required for the UE to provide feedback based on the previous processing 1006 or to perform the action indicated by 1004. Therefore, the processing time defined in the specification must be implemented by all devices and algorithms / methods 1008; otherwise, some devices may not be able to respond accordingly.

[0038] Figure 7 A schematic diagram of a typical model of a neural network 700 is shown, which has inputs x1 to x... p The input layer, hidden layer, and output layers y0 to y0 are... q Output layer 1016.

[0039] Model, model training, and model inference

[0040] The embodiments particularly relate to model training, which is the process of adapting a model to so-called training data. A model can first be described by its structure, i.e., multiple interconnected layers, see [reference needed]. Figure 7 Each layer can be described by its input size IS (the number of values ​​entering the layer), output size OS (the number of values ​​leaving the layer), and layer type, such as fully connected, convolutional, etc. Additionally, auxiliary layers such as Sigmoid, ReLU, Dropout, BatchNorm, etc., can exist. Each of these layers can describe a mathematical operation with IS-dimensional input and OS-dimensional output.

[0041] Typically, the parameters (weights) of such neural layers are not fixed before training. However, they can be randomly initialized using a uniform distribution or other initialization procedures, such as Kaiming or He initialization. The training process involves finding the weights that minimize a certain loss function on a so-called training set.

[0042] The training set can include samples that can be collected by the UE itself, the network, or provided by another entity. Using these samples, the training process can involve learning algorithms such as stochastic descent, Adam, calibrated Adam, etc., to optimize the model's weights. An unoptimized model can be referred to as untrained, and a model that has been optimized using a training set can be referred to as a trained model.

[0043] After model training, model inference can be performed. Model inference refers to feeding some unknown samples into the trained model and obtaining the model's output, in order to perform further actions based on that output. Therefore, inference time can be defined as the time spent by the trained model generating the output data from the input data. This may also include latency caused by preprocessing or postprocessing required when using an AI / ML model.

[0044] Regarding the implementation of AI / ML models in wireless communication scenarios, two different approaches can be identified for integrating AI / ML-based methods into the 3GPP framework.

[0045] AI / ML-based lifecycle management (LCM)

[0046] Function-based LCM anticipates that the actual AI model or algorithm is transparent to the network. Therefore, the network may only know that the UE supports a particular function or feature, without knowing which model the UE actually uses to implement that function. In this case, the network is primarily responsible for activating and deactivating a specific AI function. The selection or generation of the model is internal to the UE.

[0047] Lifecycle Management (LCM) Based on AI / ML Model IDs

[0048] Model ID-based LCM uses a central unit where all models in use are registered. Each registered model is uniquely identified by a model ID. The model ID may indicate only the model's structure or also its weights. Additionally, it can link to one or more training datasets that have been used or can be used for a particular model.

[0049] The examples involve two methods.

[0050] Embodiments of the present invention can be found in Figures 1 to 12. Figure 5Implemented in the wireless communication system or network shown, which includes transceivers such as base stations, gNBs or access points (APs) or repeaters, and multiple communication devices such as user equipment, UEs or stations, STAs.

[0051] Implementations may rely on the use of, for example, in such wireless communication systems or networks Figure 7 The model shown is an AI / ML model, and the different processing times used or required can be addressed based on different implementations and / or different computing capabilities, resulting in... Figure 6 The situation indicated in the document is intended to address, at least partially, the drawbacks of worst-case handling time.

[0052] Figure 8 This is a schematic diagram of a wireless communication system including transceiver 200, such as a base station or repeater, and multiple communication devices 2021 to 202n, such as UEs. UEs can communicate directly with each other via a wireless communication link or channel 203, such as a radio link (e.g., using a PC5 interface (side link)). Furthermore, the transceiver and UE 202 can communicate via a wireless communication link or channel 204, such as a radio link (e.g., using a uU interface). Transceiver 200 may include one or more antennas ANT or an antenna array with multiple antenna elements, signal processor 200a, and transceiver unit 200b. UE 202 may include one or more antennas ANT or an antenna array with multiple antennas, processors 202a1 to 202an, and transceiver (e.g., receiver and / or transmitter) units 202b1 to 202bn. Base station 200 and / or one or more UEs 202 can operate according to the teachings of the invention described herein.

[0053] The embodiments propose solutions, for example, implementing one or more methods and / or devices and / or network architectures and auxiliary signaling to implement AI / ML methods for different use cases, such as CSI prediction, CSI compression, HARQ prediction, AI localization, beam prediction, beam adaptation and / or mobility enhancement in 5G NR systems.

[0054] Some embodiments relate to aspects such as what the network entity is, what attributes the hardware and / or software and / or network involve, what the hardware accelerator unit is, or which parts of the model to be processed may be involved. Such definitions, as in the rest of the aspects described herein, apply to other aspects without any limitation.

[0055] Some embodiments are described in conjunction with sections 1 through 6. Although described in sections, these sections describe the basic invention from different perspectives, such that the details described herein can be combined with each other without limitation, and the details described in conjunction with some implementations in one section are also valid for the embodiments described in other sections without limitation.

[0056] 1. Calculation of inference time

[0057] One aspect of the embodiments described herein relates to the calculation of inference time.

[0058] In one embodiment, an apparatus for a wireless communication network is provided that uses one or more artificial intelligence / machine learning (AI / ML) models for one or more use cases, wherein the apparatus is used to determine the inference time of one or more AI / ML models to be used in one or more network entities of the wireless communication network. Alternatively or additionally, the AI / ML model may be a general optimizer, an unknown (network / 3GPP unknown) algorithm, a neural network, and / or a solver. Typically, an AI / ML model can be a general term for an entity with specific inputs and outputs that solves a specific problem. Although such an entity may sometimes be considered a black box, there are ways to define such models.

[0059] In this embodiment, inference time includes the time required to fully or partially process the AI / ML model, and inference time is provided in the form of absolute time or offset value.

[0060] In one embodiment, inference time can be represented by one or more of the following:

[0061] - seconds, milliseconds, microseconds, nanoseconds; multiples of these time units, such as (x * nanoseconds), number of time slots, number of subframes, number of OFDM symbols, number of cycles,

[0062] - Offset value, indicating at least one of the following: offset time relative to a reference time, such as a reference time provided by a navigation system, such as GPS; offset relative to the start of a frame; or offset relative to the frame structure (such as the Physical Downlink Control Channel (PDCCH)) or synchronization signal (such as the primary synchronization sequence (PSS), secondary synchronization sequence (SSS), or sidelink synchronization sequence transmitted via the sidelink broadcast channel (PSBCH)).

[0063] In this embodiment, inference time includes the time required to partially process the AI / ML model, where this portion is a part of the AI / ML model to be processed; wherein the AI / ML model includes unprocessed parts. This can be understood as only processing a portion of the model in some use cases or some AI / ML models. In these cases, other parts are not processed. Therefore, the unprocessed parts may not contribute significantly to the processing time.

[0064] In this embodiment, the inference time of the AI / ML model is determined using an inference time model that uses at least one or more first attributes of the AI / ML model and / or one or more second attributes of network entities that are to be used in at least a portion of the AI / ML model to calculate the inference time.

[0065] In this embodiment, each AI / ML model includes a specific neural network, and the network entity includes specific hardware for implementing that specific neural network.

[0066] One or more first attributes of the AI / ML model include one or more attributes of the neural network, and one or more second attributes of the network entity include one or more attributes of the hardware.

[0067] In this embodiment, the properties of the neural network include one or more of the following:

[0068] - The number of layers in a neural network,

[0069] - The depth of the neural network, for example, the number of layers that must be executed sequentially.

[0070] - The number of specific operations, such as floating-point operations, multiplication, addition, integer operations, Boolean operations, exponential functions, etc.

[0071] - The width of the layers in a neural network, such as the input size IS and / or output size OS.

[0072] - The types of layers in a neural network, such as convolutional layers, activation layers, batch normalization layers, or fully connected layers, and

[0073] Hardware attributes include one or more of the following:

[0074] - The number of hardware accelerator units, such as the number of graphics processing units (GPUs), tensor processing units (TPUs), or tensor cores;

[0075] - Processor speed, such as floating-point operations per second (FLOPS), additions per second, multiplications per second, and integer operations per second;

[0076] - The number of processor cores,

[0077] -The type of core processing,

[0078] - A combination of processing cores, such as x GPU cores and y tensor cores.

[0079] - Memory size

[0080] -Memory speed,

[0081] - Memory type

[0082] - Memory architecture.

[0083] A hardware accelerator unit may be or may include one or more physical or logical units. For example, power is measured in the number of standardized accelerator units.

[0084] In this embodiment, the AI / ML models used in the wireless communication network are uniquely numbered and identifiable, and the device uses one or more of the following to determine the inference time of the supported AI / ML model identifier (ID):

[0085] - Processing time for supported AI / ML model IDs

[0086] - The number of multiple or a set of supported AI / ML models to be processed in parallel or sequentially.

[0087] In this embodiment, the AI / ML models used in the wireless communication network are uniquely numbered and identifiable;

[0088] The device is used to determine the inference time of at least one specific supported AI / ML model, which can be run as a separate AI / ML model within a use case model; and / or

[0089] The device is used to determine the inference time of at least one set of supported AI / ML models, which can run simultaneously in use cases.

[0090] In an embodiment, the specific AI / ML model to be used in a network entity is inferred by identifying a certain feature or function supported by the network entity. For example, n-bit CSI feedback inference uses a specific AI / ML model that implements a precoding engine, or n-bit SINR feedback inference uses a specific AI / ML model that implements a switching function.

[0091] In this embodiment, the device includes network entities that use AI / ML models, such as:

[0092] -User Equipment (UE), or

[0093] - Remote UE, or

[0094] - Relay UE, or

[0095] -Radio Access Network (RAN) entities, such as gNBs or roadside units (RSUs), or

[0096] - Core network (CN) entities, such as Access and Mobility Functions (AMF) or Location Management Functions (LMF).

[0097] and / or

[0098] The device is separate from one or more network entities that use AI / ML models. For example, the device may contain another network entity from a wireless communication network, or an entity from a network different from the wireless communication network, such as the Internet.

[0099] In an embodiment, the device is used to indicate whether an AI / ML model is available or unavailable on a network entity, and / or to fall back to the default procedure if the determined inference time of an AI / ML model is equal to or less than the predefined or (pre)configured processing time of one or more operations of the use case in which the AI / ML model is used.

[0100] Regarding indicating a model as unavailable, even if the inference time is below a threshold (which may occur when the device is able to process the model faster than a predefined threshold), for example, the processing conflicts with another model, causing the AI / ML processor to be occupied / blocked, thus preventing the UE from processing that model in parallel with another configured model. Other scenarios are also not excluded; for example, the UE might intend to perform this calculation on a different processor to save power, rather than using its AI / ML processor.

[0101] In this embodiment, the device communicates via a side link, and the processing time is configured in the resource pool configuration (RP).

[0102] In an embodiment, the device is used to indicate the inference time of a certain AI / ML model or AI / ML function to the network and / or network entities and / or gNB.

[0103] In the embodiments, use cases include one or more of the following:

[0104] - Channel State Information (CSI) prediction,

[0105] -CSI compression,

[0106] - Hybrid Automatic Repeat Request (HARQ) forecasting,

[0107] -Location of user equipment

[0108] - Beam management,

[0109] - Beam prediction

[0110] - Beam adaptation,

[0111] -Enhanced mobility

[0112] -SINR prediction

[0113] -SL resource allocation,

[0114] -SL sensing,

[0115] - Toggle (HO), or conditional toggle (CHO).

[0116] -Discover.

[0117] In an embodiment, the apparatus is used to indicate inference time to one or more user equipments (UEs) communicating via a sidelink SL.

[0118] In this embodiment, the device is set up

[0119] - RAN entities, such as gNB or RSU, are used to align inference time among multiple UEs operating in Mode 1, or

[0120] -SL UE, or in a remote UE, or

[0121] -In the relay UE, or

[0122] - In multiple UEs, for coordinating inference time via sidelinks when running in mode 1 or mode 2, for example...

[0123] ○ During the SL synchronization and / or SL discovery and / or SL connection establishment phases, for example, during the transmission of the Physical Side Link Broadcast Channel (PSBCH), or

[0124] ○ Using signaling via the Physical Side Link Control Channel (PSCCH),

[0125] ○ Use signaling embedded in the Physical Side Link Shared Channel (PSSCH),

[0126] Feedback switching is used via the Physical Side Link Feedback Channel (PSFCH).

[0127] According to an embodiment, a method is provided for operating an apparatus for a wireless communication network that uses one or more artificial intelligence / machine learning (AI / ML) models for one or more use cases, the method comprising determining the inference time of one or more AI / ML models to be used in one or more network entities of the wireless communication network.

[0128] Inference time, i.e., the processing time required to execute a machine learning algorithm / method, can be calculated at the UE or gNB. The calculation can be based on certain rules or formulas that include one or more of the following parameters:

[0129] • The number of layers in a neural network

[0130] • The depth of the neural network, for example, the number of layers that must be executed sequentially.

[0131] • Layer width, such as input size (IS) and output size (OS).

[0132] • Layer types, such as convolutional layers, fully connected layers, etc.

[0133] • The number of hardware accelerator units, such as the number of GPUs, TPUs, tensor cores, and other units. The value exchanged for this purpose can be based on the number of real-valued model parameters and / or the number of real-valued operations.

[0134] • Processor speed, for example, FLOPS

[0135] • Number of processor cores

[0136] • The type of processing core,

[0137] • A combination of processing cores, for example, x GPU cores and y tensor cores.

[0138] • Memory size, memory speed, memory type, memory architecture

[0139] • Supported model IDs, for example, when AI / ML models are uniquely numbered and identifiable.

[0140] • Processing time for model ID

[0141] • Model IDs or model groups that can be processed in parallel or sequentially.

[0142] • Supported features or functional identifiers that can infer the specific AI / ML engine / model / pattern to be used. For example, n-bit CSI feedback can infer the use of a specific AI / ML precoding engine, and n-bit SINR feedback can infer a specific AI / ML switching function.

[0143] 2. Signaling for inference time

[0144] One aspect of the embodiments described herein relates to signaling for inference time, such as inference time calculated as described above.

[0145] According to an embodiment, a user equipment (UE) of a wireless communication network is provided, which uses one or more artificial intelligence / machine learning (AI / ML) models for one or more use cases.

[0146] The UE uses one or more AI / ML models, and

[0147] The UE sends a signal to the wireless communication network to notify the UE of the inference time required to execute one or more AI / ML models.

[0148] According to an embodiment, the UE signals the inference time to at least one of the gNB, the UE, and the relay UE.

[0149] According to the embodiment, UE

[0150] - In response to one or more AI / ML models transmitting data from network entities in a wireless communication network to the UE, or

[0151] - In response to the activation of one or more AI / ML models and / or AI / ML functions from a network entity in the wireless communication network to the UE, or

[0152] - In response to a request from a network entity in the wireless communication network, for example, if the UE has one or more AI / ML models pre-configured, or after one or more AI / ML models have been transmitted to the UE, or

[0153] - When accessing a wireless communication network, if the UE is pre-configured with one or more AI / ML models, for example, along with signaling of the UE's capabilities,

[0154] Send a signal to notify the reasoning time.

[0155] According to an embodiment, the network entity of the wireless communication network that transmits AI / ML models or requests inference time includes one or more of the following:

[0156] - Another UE, or a relay UE, or a remote UE,

[0157] - Radio access network (RAN) entities, such as gNBs or roadside units (RSUs).

[0158] - Core network (CN) entities, such as Access and Mobility Functions (AMF) or Location Management Functions (LMF).

[0159] According to an embodiment, inference time includes the time required to process the AI / ML model in whole or in part, and the inference time is provided in the form of an absolute value or an offset value.

[0160] According to an embodiment, inference time can be provided by one or more of the following:

[0161] - Seconds, milliseconds, microseconds, nanoseconds; multiples of these time units, such as (x * seconds / milliseconds / microseconds / nanoseconds), number of time slots, number of subframes, number of OFDM symbols, number of cycles,

[0162] - Offset value, indicating at least one of the following: offset time relative to a reference time, such as a reference time provided by a navigation system, such as GPS; offset relative to the start of a frame; or offset relative to the frame structure (such as the Physical Downlink Control Channel (PDCCH)) or synchronization signal (such as the primary synchronization sequence (PSS), secondary synchronization sequence (SSS), or sidelink synchronization sequence transmitted via the sidelink broadcast channel (PSBCH)).

[0163] According to an embodiment, inference time includes a portion of the time required to process the AI / ML model, wherein the portion is a part of the AI / ML model to be processed; wherein the AI / ML model includes a portion that is not processed.

[0164] According to the embodiment, UE

[0165] - For example, using an inference time model to determine inference time, which uses at least one or more attributes of the AI / ML model and one or more attributes of the UE, or

[0166] - Receive inference time from a wireless communication network, for example, from the apparatus of any of the above embodiments, or from a network entity including the apparatus of any of the above embodiments, such as a RAN entity of a CN entity, or from another UE, for example, via a side link interface (also known as PC5).

[0167] According to an embodiment, the UE signals the number of instances of a specific AI / ML model that the UE can process in parallel and / or the number of AI / ML models.

[0168] According to an embodiment, the UE selects from a set of configured or pre-configured inference times the inference time of a specific AI / ML model that the UE can implement when executing a specific AI / ML model, and the inference time of that specific AI / ML model to be signaled. That is, the embodiment covers different instances and / or different models that operate the same model sequentially, simultaneously, or in parallel.

[0169] According to an embodiment, inference time is at least a portion of the processing time required to process a particular AI / ML model.

[0170] According to an embodiment, the UE only signals the inference time of a particular AI / ML model to the wireless communication network if the inference time allows the execution of the particular AI / ML model in accordance with the processing time constraints associated with the use case using the particular AI / ML model.

[0171] According to an embodiment, the inference time of a specific AI / ML model is associated with a specific AI / ML model identifier (ID) or function, and the UE only reports the AI / ML model ID if the UE can meet the processing time constraints.

[0172] According to an embodiment, the wireless communication network uses one or more artificial intelligence / machine learning (AI / ML) models for one or more use cases.

[0173] The UE executes one or more AI / ML models to perform one or more specific operations.

[0174] The UE signals the wireless communication network to inform it of the complexity or capacity it can perform, enabling the use of a specific AI / ML model to perform the specific operation within a predefined processing time associated with that operation.

[0175] In response to signaling, the UE receives one or more AI / ML models that it can execute from the wireless communication network to perform a specific operation according to a predefined processing time.

[0176] According to the embodiments, complexity or capacity involves at least one of the following:

[0177] - The number of layers in the neural network of the AI / ML model.

[0178] - The depth of the neural network in the AI / ML model, for example, the number of layers that must be executed sequentially.

[0179] - The number of specific operations, such as floating-point operations, multiplication, addition, integer operations, Boolean operations, and exponential functions.

[0180] - The width of the layers in the neural network of the AI / ML model, such as the input size (IS) and / or output size (OS).

[0181] - The type of layers in the neural network of the AI / ML model, such as convolutional layers, activation layers, batch normalization layers, or fully connected layers, and

[0182] - The number of hardware accelerator units in the UE, such as the number of graphics processing units (GPUs), tensor processing units (TPUs), or tensor cores;

[0183] -UE's processor speed, such as floating-point operations per second (FLOPS), additions per second, multiplications per second, and integer operations per second.

[0184] - The number of processor cores,

[0185] -The type of core processing,

[0186] - A combination of processing cores, such as x GPU cores and y tensor cores.

[0187] -UE's memory size

[0188] -UE's memory speed,

[0189] - The type of memory used by the UE

[0190] -UE's memory architecture.

[0191] As described above, such a hardware accelerator unit can be at least one physical unit and / or logical unit. For example, power can be measured by the number of standardized accelerator units.

[0192] According to an embodiment, if the currently used or requested AI / ML model cannot meet the predefined processing time, the UE receives information from the wireless communication network indicating whether to roll back the AI / ML model or to perform a rollback procedure as required.

[0193] Alternatively, the UE may be (pre-)configured to use a fallback process if the currently used or requested AI / ML model cannot meet the processing time requirements.

[0194] For example, pre-configuration can involve one or more of the following:

[0195] -Specified in the specifications upon which the operation of the wireless communication network is based.

[0196] - Pre-configuration, such as via semi-static configuration as part of higher-layer signaling (such as MAC, RRC, or SIB), or specific AI / ML control channels or AI / ML protocols.

[0197] - Pre-loaded by the manufacturer's factory; and / or

[0198] - Configured or indicated by lower-level signaling (SCI or DCI).

[0199] According to an embodiment, a method is provided for operating a user equipment (UE) of a wireless communication network, wherein the wireless communication network uses one or more artificial intelligence / machine learning (AI / ML) models for one or more use cases, the method comprising:

[0200] Using one or more AI / ML models, and

[0201] Signal the wireless communication network to notify the inference time required to execute one or more AI / ML models.

[0202] According to an embodiment, a method is provided for operating a user equipment (UE) of a wireless communication network, wherein the UE executes one or more AI / ML models to perform one or more specific operations, and the wireless communication network uses one or more artificial intelligence / machine learning (AI / ML) models for one or more use cases, the method comprising:

[0203] Signaling to the wireless communication network the complexity or capacity that the UE can perform, enabling the execution of a specific operation using a specific AI / ML model within a predefined processing time associated with that operation, and

[0204] In response to this signaling, one or more AI / ML models that the UE can execute are received from the wireless communication network for performing specific operations according to a predefined processing time.

[0205] Regarding the described embodiments, neural networks can vary considerably in terms of complexity. Furthermore, the computing power of the device can also experience high variability. Currently, the specification has limited capacity to represent this. Specifically, the current 5G specification supports two different PDSCH processing times based on the UE's capabilities, which is the time required for complete decoding. Similar processing times also exist for the minimum time before PUSCH preparation, the expected DFI (Downlink Feedback Indicator), or the minimum gap between DCI / PDCCH and PDSCH. The UE can signal to the network which processing times it supports during initial access. Based on this, the network can choose one of the PDSCH processing times. However, in the case of neural networks, it depends not only on the UE's own capabilities but also on the actual network, which may be unknown to the UE during initial access (e.g., because the network transmits it at a later stage). Additionally, the network may not know exactly what capabilities the UE possesses. In this case, see [link to relevant documentation]. Figure 6 Therefore, a processing time needs to be selected so that the UE is expected to meet the requirement. Thus, prior to this invention, computational assumptions were made for the worst-case scenario.

[0206] To address this issue, the embodiments provide auxiliary signals that indicate the expected or test inference time required for the UE to execute the neural network, see Figures 9a and 9b.

[0207] Figures 9a-b illustrate schematic signaling between the gNB and the UE in Figure 9a and between the two UEs in Figure 9b, such as auxiliary signaling between the gNB and the UE or between UEs. This signaling can be provided in response to neural network transmissions from the network to the UE, or it can be explicitly requested by the network, for example, using signal 12 from the gNB to the UE / from one UE to another and / or vice versa.

[0208] Information 14 may indicate at least one of the following: model parameters, model structure, model ID that identifies the corresponding model, and function ID that identifies the corresponding function.

[0209] For example, the UE can provide signal 16 to indicate whether the UE includes and / or will provide or retain the required capability and / or indicate the correct or incorrect reception of signal 14.

[0210] Using signal 18, the UE can report inference performance, such as processing time, the number of parallel transmissions, etc.

[0211] Inference time can be the total time required for the entire process or a portion of the process. Furthermore, it can be determined by actual execution and measurement time, or it can be calculated based on a latency model, as detailed above regarding the calculation of inference time. Inference time can be provided in milliseconds, microseconds, nanoseconds, the number of time slots or OFDM symbols, or the number of cycles, or as an offset value.

[0212] In the embodiments, or as a different operating mode or following a different process, the UE may also send the number of parallel AI / ML instances that the UE can process.

[0213] In an embodiment, the UE can select from a set of (pre)configured processing times the processing time it is likely capable of achieving.

[0214] In the embodiments, processing time may be associated with a specific model ID / function, and the UE will only report whether it can or cannot execute the specific model ID / function if the UE can also meet the processing time constraints, i.e., the model is available and / or unavailable.

[0215] In the embodiment, the UE reports the complexity / capacity it can perform within a specific processing time.

[0216] In this embodiment, if a given UE cannot meet the processing time, the gNB can also indicate to the UE the fallback method to use. This may be the case if the UE is interrupted by further processing, or if the UE is required to perform DRX to save power.

[0217] 3. Auxiliary signaling

[0218] One aspect of the embodiments described herein relates to auxiliary signaling, such as the auxiliary signaling of section 2.

[0219] According to an embodiment, a user equipment (UE) of a wireless communication network is provided, which uses one or more artificial intelligence / machine learning (AI / ML) models for one or more use cases.

[0220] The UE is configured or pre-configured with multiple AI / ML models for performing one or more specific operations, and

[0221] Depending on one or more standards, in order to perform one or more specific operations, the UE:

[0222] - Switch from the first AI / ML model to the second AI / ML model, or

[0223] -Disable one or more of multiple AI / ML models, or

[0224] - Switch from non-AI / ML mode to AI / ML mode, or

[0225] - Switch from AI / ML mode to non-AI / ML mode, or

[0226] - Switch from the current operating mode to the new operating mode.

[0227] The non-AI / ML mode mentioned above refers to signal processing of data by a processor that does not use an AI / ML engine, uses a hardware-accelerated AI / ML engine, or uses software-based AI / ML processing to perform special operations.

[0228] According to an embodiment, the UE is configured or pre-configured with multiple AI / ML models of varying complexity for performing specific operations, and

[0229] Depending on one or more criteria, the user interface (UE) switches from a first AI / ML model to a second AI / ML model to perform a specific operation. The second AI / ML model has a complexity that is lower or higher than that of the first AI / ML model.

[0230] According to an embodiment, one or more standards include one or more of the following:

[0231] - Receiver conditions, such as the reference signal received power (RSRP) and signal-to-interference-plus-noise ratio (SINR), can cause changes in receiver conditions to lead to switching between AI / ML models trained for different SINR values ​​or SINR ranges.

[0232] -UE's battery level

[0233] -UE heat level,

[0234] - Changes in inference time, such as due to additional models executed in parallel.

[0235] - Handling changes in processing time requirements, such as switching to URLLC mode.

[0236] - Changes in packet load, such as buffer state.

[0237] - Changes in bandwidth and / or the number of active carriers

[0238] - Power saving operation

[0239] -The semantics of the data, for example, message types, such as emergency messages.

[0240] - QoS key performance indicators (KPIs), such as packet reception rate (PRR).

[0241] - Signaling from the gNB or another UE, such as a command to switch to another model.

[0242] According to an embodiment, the UE is configured or pre-configured with multiple AI / ML models to be executed in parallel for performing one or more specific operations, and

[0243] If the UE determines that its computing power is insufficient to operate multiple AI / ML models in parallel, the UE will disable one or more of the multiple AI / ML models.

[0244] As described above, the order in which computing capacity or capabilities are deactivated can be determined by the UE or can be configured based on priority. That is, according to the embodiment, the UE deactivates one or more of multiple AI / ML models according to a deactivation order determined by the UE or, for example, based on priority (pre)configuration.

[0245] According to an embodiment, if the UE determines that its computing power is insufficient to operate a certain AI / ML model, the UE switches from the current operating mode to a new operating mode. The new operating mode enables the UE to execute the AI / ML model according to the expected performance, such as the processing time required for the operation performed by the UE using the AI / ML model.

[0246] According to an embodiment, the new operating mode makes the input size (IS) of the AI / ML model smaller than the IS of the current operating mode, thereby achieving predefined transmit and / or receive performance within a given smaller ε of the configured or pre-configured performance range while obtaining processing results faster. For example, the size of IS can be reduced or made smaller without significantly degrading performance. For example, the performance degradation is kept within a certain ε. The parameter ε may be related to or represent the maximum permissible error range or deviation. According to an embodiment, this value can be obtained by comparing the model with another model or algorithm. According to an embodiment, ε is the deviation of the time average, representing the degradation of model performance. The actual value of ε can be (pre)configured.

[0247] According to an embodiment, the UE switches to a new PHY or MAC mode, for example, a PHY or MAC mode with fewer transmit and / or receive antennas than the current PHY or MAC mode.

[0248] According to the embodiment, the UE sends a signal to the network entity of the wireless communication network to notify it to switch from a first AI / ML model to a second AI / ML model, or to disable one or more of a plurality of AI / ML models, or to switch from the current operating mode to a new operating mode. The network entity of the wireless communication network includes one or more of the following:

[0249] -Another UE, or a remote UE, or a relay UE,

[0250] - Radio access network (RAN) entities, such as gNBs or roadside units (RSUs).

[0251] - Core network (CN) entities, such as Access and Mobility Functions (AMF) or Location Management Functions (LMF).

[0252] According to the embodiment, in order to signal to the RAN or CN entity, the UE uses uplink control information (UCI), MAC control element (MAC CE), radio resource control information element (RRC IE), SL control information (SCI), first and / or second phase SCI and / or auxiliary information message (AIM) or any other higher-layer signaling to signal the handover / deactivation.

[0253] According to an embodiment, in order to send a signal to another UE, the UE

[0254] - During the initial access phase, for example, within the transmission of the Physical Side Link Broadcast Channel (PSBCH), or

[0255] -Use signaling via the Physical Side Link Control Channel (PSCCH),

[0256] -Use signaling embedded in the Physical Side Link Shared Channel (PSSCH),

[0257] - Use feedback switching via the Physical Side Link Feedback Channel (PSFCH)

[0258] Send a signal to notify the user to switch or deactivate.

[0259] According to an embodiment, a method is provided for operating a user equipment (UE) of a wireless communication network, the UE being configured or pre-configured with multiple AI / ML models for performing one or more specific operations, the wireless communication network using one or more artificial intelligence / machine learning (AI / ML) models for one or more use cases, the method comprising:

[0260] Perform one or more specific operations by executing the following:

[0261] - Switch from the first AI / ML model to the second AI / ML model, or

[0262] -Disable one or more of multiple AI / ML models, or

[0263] - Switch from non-AI / ML mode to AI / ML mode, or

[0264] - Switch from AI / ML mode to non-AI / ML mode, or

[0265] - Switch from the current operating mode to the new operating mode.

[0266] With the aid of auxiliary signaling, the UE can be (pre-)configured with multiple AI / ML methods of varying complexity. It can then switch to a more complex or simpler method based on indications, reception conditions such as RSRP, SINR, and / or battery level, or another trigger. If such a handover is decided at the UE, the UE can indicate the handover to the gNB using UCI, MAC CE, or RRC IE, or any other higher-layer signaling.

[0267] Furthermore, the UE can determine that its computing power is insufficient to operate multiple AI operations in parallel. In this case, the UE can instruct the deactivation or activation of certain AI operations.

[0268] In addition, or as an alternative, if the processing power at the UE is insufficient to perform a certain AI operation, the UE can also switch back to a different PHY or MAC mode, for example, with fewer transmit and / or receive antennas, in case the smaller input of the AI ​​operation will result in a faster processing result, and this will still achieve certain transmit and / or receive performance, or at least within a given small ε within the (pre)configured performance interval.

[0269] In embodiments, signaling can be extended for UEs communicating via the sidelink (SL). This depends on the operating mode, such as mode 1 or mode 2. In mode 1, the gNB can align inference time with the UE that wants to communicate in direct mode. In mode 2, the UE must coordinate inference time itself via sidelink control signaling. Here, this can be indicated during the initial access phase, for example, within a PSBCH transmission, or using signaling embedded in the data channel (PSSCH) via the sidelink control channel (PSCCH), or signaling transmitted within a feedback exchange via the PSFCH.

[0270] 4. Multiple models

[0271] One aspect of the embodiments described herein relates to operating multiple models, such as a set of models, sequentially or at least in parallel.

[0272] According to an embodiment, a user equipment (UE) of a wireless communication network is provided, which uses one or more artificial intelligence / machine learning (AI / ML) models for one or more use cases.

[0273] The UE is configured or pre-configured with multiple AI / ML models for performing one or more specific operations.

[0274] The UE has an AI / ML model processing circuit, which has one or more constraints that allow only a certain number of AI / ML models to be executed.

[0275] According to an embodiment, the UE maps the processing of multiple AI / ML models to the AI / ML model processing circuit, taking into account the constraints of the AI / ML model processing circuit and / or the input received from the wireless communication network.

[0276] According to the embodiment, the AI / ML model processing circuit constraints include:

[0277] -UE's AI / ML model processing circuitry has only one AI / ML accelerator.

[0278] - The UE's AI / ML model processing circuitry has two or more AI / ML accelerators, where the specific processing power of the two or more AI / ML accelerators depends on certain processing capabilities. For example, it depends on whether the two or more AI / ML accelerators have the same or different processing capabilities. For instance, in the case where the AI / ML model processing circuitry includes a high-performance Tensor Processing Unit (TPU) and low-performance cores such as a Graphics Processing Unit (GPU) or a Central Processing Unit (CPU), the AI / ML model is mapped to two or more AI / ML accelerators.

[0279] - Definition of processing time, for example, processing time can include

[0280] ○ Loading one or more AI / ML models plus processing one or more AI / ML models,

[0281] ○ Loading one or more AI / ML models, processing one or more AI / ML models, and updating one or more AI / ML models.

[0282] According to an embodiment, when the UE processes more than one AI / ML model on only one processor, the UE signals the network entity of the wireless network to notify which algorithm or function of the AI / ML model requires longer processing time. This is because the two AI / ML models / functions share the same processing unit. One option is for the processing unit to prioritize processing one model, thus satisfying the inference time requirement for the first model, but requiring a longer inference time for the second model. Another option is for the processing unit to share processing power equally, therefore requiring a longer inference time when both models are executed simultaneously.

[0283] According to an embodiment, the UE is configured to receive signaling from a network entity of the wireless communication network, the signaling indicating...

[0284] - First, calculate the preference for which AI / ML model, or

[0285] - A priority list of multiple AI / ML models, e.g., first, second, third... which AI / ML model to compute.

[0286] According to an embodiment, when the UE switches processing from the current AI / ML model to a new AI / ML model, the UE signals the network entity of the wireless communication network to notify it of the duration of the reconfiguration.

[0287] According to an embodiment, in response to a request from a network entity in the wireless communication network, the UE switches processing from the current AI / ML model to a new AI / ML model, and

[0288] In response to a request or in response to a trigger, the UE sends one or more of the following to the wireless communication network:

[0289] - A confirmation message indicating that the loading of the new AI / ML model has been successfully completed.

[0290] - A conflict message indicating that loading a new AI / ML model is impossible, for example, along with a possible fallback AI / ML model that needs to be used or can be configured.

[0291] - An update message indicating the duration of computation for a new AI / ML model that may require additional processing time and / or the duration of computation for an additional (e.g., old) AI / ML model, for example, because changing the model may change computational complexity and / or may require additional processing time.

[0292] According to the embodiments described herein, the UE may have a trigger, which may be internal or external. For example, the trigger may be related to at least one of the following: changes in signal quality, changes in mobility, changes in location or altitude (e.g., in the case of the UE being a drone), changes in available battery power, the state of the UE (e.g., stationary, changing indoors, changing outdoors), changes in frequency band (e.g., FR1->FR2 or vice versa), or others.

[0293] According to an embodiment, the UE signals a network entity of the wireless communication network to notify which of the multiple AI / ML models requires how much processing power.

[0294] For example, the UE can signal to the network entity how many of its AI / ML processing units and / or memory space and / or which AI / ML processing units it needs, allowing the network entity to instruct the UE which AI / ML algorithm combination it should use for a particular computation and / or how to partition its algorithms. Alternatively or supplementarily, the UE can indicate which AI / ML algorithms use what percentage or number of AI / ML processing units / memory, for example:

[0295] AI / ML Algorithm 1 -> 20% AI / ML Units, 15% Memory

[0296] • AI / ML algorithm 2 -> 35% AI / ML unit, 25% memory

[0297] • AI / ML Algorithm 3 -> 80% AI / ML Units, 45% Memory

[0298] In such an embodiment, models or algorithms 1 and 2 can run or be processed together, while models 2 and 3 will exceed the hardware capabilities of the UE.

[0299] The solutions mentioned above and in this article can be combined without limitation, such as combining functions or functions that change over time, such as with changes in operating modes.

[0300] According to an embodiment, the network entity of the wireless communication network includes one or more of the following:

[0301] -Another UE,

[0302] -Remote UE,

[0303] -Relay UE,

[0304] - Radio access network (RAN) entities, such as gNBs or roadside units (RSUs).

[0305] - Core network (CN) entities, such as Access and Mobility Functions (AMF) or Location Management Functions (LMF).

[0306] According to an embodiment, a method is provided for operating a user equipment (UE) of a wireless communication network, wherein the UE is configured or pre-configured with multiple AI / ML models for performing one or more specific operations, and the wireless communication network uses one or more artificial intelligence / machine learning (AI / ML models) to handle one or more use cases. The method includes:

[0307] One or more constraints of the UE-based AI / ML model processing circuit can be applied to execute only a certain number of AI / ML models.

[0308] For example, in embodiments involving multiple models, the UE may have limited processing capabilities, such as only one (or a limited number, but higher) AI / ML unit. In this case, running more than one AI / ML function simultaneously may require a long processing time, or may not be feasible at all. Therefore, embodiments propose optimizations for mapping or configuring specific AI / ML functions to AI / ML processing units in certain ways.

[0309] The following constraints may apply:

[0310] ·UE has only one AI / ML accelerator.

[0311] ·UE has multiple AI / ML accelerators,

[0312] How to map multiple functions to different accelerators, which may have different capabilities, so the mapping depends on the specific function to be computed and the available processing capabilities:

[0313] ○The same ability,

[0314] ○ Different capabilities, such as high-performance (TPU = Tensor Processing Unit) and low-performance cores (Graphics Processing Unit, GPU / Central Processing Unit (CPU)).

[0315] • Processing time definition: It may include model loading + model processing + model updating.

[0316] Figure 10 A schematic representation of the task solved by the embodiments described herein is shown, for example, a possible mapping from AI / ML functions to an AI / ML processor. At least one AI / ML function 221 to 22 n The set is mapped or distributed to m AI / ML processors or accelerators 241 to 24 m , n≥1, where this distribution is particularly favorable for (n+m)>2.

[0317] In the following description relating to the embodiments, signaling related to the embodiments described herein is provided:

[0318] ■ The embodiments relate to signaling in cases where the UE must perform computations of more than one function on only one processor. For example, the UE may indicate which algorithm to execute or the longer processing time required to compute the function.

[0319] ■ The implementation involves signaling from a BS or gNB or another UE: for example, which function preference is calculated first, or a priority list of a given number of functions, such as first, second, third, etc., which function is calculated.

[0320] ■ The example involves model switching time: Loading different models into the TPU / GPU may take some time to configure a specific AI / ML core with given input parameters.

[0321] ○The UE signals the network / another UE to notify the duration of the reconfiguration.

[0322] ○ Two-way communication mechanism: The network instructs the UE to prepare for model loading, and the UE sends...

[0323] ■ Confirmation message upon successful loading

[0324] ■ Conflict message: When loading fails, there may be a fallback AI / ML model that is being used or configured.

[0325] ■Update message: The UE signals the network to notify the computation duration of the new model and / or the computation duration of additional models that may require additional processing time, such as the old model.

[0326] ■ General capability signaling from UE to gNB or from network to UE, such as which AI / ML requires how much processing power.

[0327] 5. Model training

[0328] Figure 11a shows a schematic block diagram of an example model training 52 according to an embodiment, which can be performed outside the network, for example, using a cloud 54 or an external data center. According to the embodiment, the model 56 obtained using training data 58 can then be packaged and transmitted to a network, such as network 100 or a different network. In this case, feedback from the UE can be collected and used, for example, to retrain / update the network in the cloud 54.

[0329] Figure 11b shows a schematic block diagram according to an embodiment, illustrating training 52 performed in a network, such as network 100 or a different network. In this case, feedback 62 from the UE can be used in the training process 52 and / or to improve network 100. Model 56 can then be packaged and transmitted to the UE.

[0330] Figure 11c shows a schematic block diagram illustrating online training that can be performed in the network and / or on the UE. In this case, the entire or a portion of the network can be trained, or, as shown in Figure 11c, a pre-trained network 56p can be used, with only a few layers 64 fine-tuned for the current location / situation or use case. This training can be performed once, periodically, or triggered as needed. In another embodiment, the model can be used for inference later or simultaneously.

[0331] Figure 11d illustrates a schematic block diagram showing the model split across more than one entity, such as cloud / Internet 54, core network (RAN) 66, and / or UE entity 68. In this case, training and / or inference can be performed entirely or partially on one or more of the following entities: input data, training data 58, feedback data 62, weight update data (e.g., forward and / or backward propagation), intermediate data, and / or output data. In another embodiment, portions of model 56 can be transferred or updated between entities 54, 66, and / or 68.

[0332] One aspect of the embodiments described herein relates to model training.

[0333] According to an embodiment, a user equipment (UE) of a wireless communication network is provided, which uses one or more artificial intelligence / machine learning (AI / ML models) for one or more use cases.

[0334] The UE is configured or pre-configured with one or more AI / ML models for performing one or more specific operations, and

[0335] UE uses a training set to train AI / ML models.

[0336] According to an embodiment, the UE trains an AI / ML model while connected to a wireless communication network.

[0337] According to the embodiments, the UE changes its connection mode to a training mode or an evaluation mode, such as RRC_TRAINING or RRC_EVALUATION mode, or a different RRC mode, such as the UE switching to RRC_INACTIVE or RRC_IDLE mode while training an AI / ML model, or another connection mode, such as DRX mode or paging mode.

[0338] The basic idea is that the UE can use a certain amount (e.g., all its available processing power / battery) for model training and will suppress its indirect access to the network to send or receive data, for example, similar to DRX mode. For this purpose, the UE can use, for example, RRC_INACTIVE mode. Alternatively, an AI training mode (RRC_TRAINING) can be defined. Optionally, in this mode, the UE can still listen for certain messages, for example, to maintain timing or connectivity with the network. Thus, if it has completed model training, it can immediately send to network entities with the correct timing advance and power control values. Furthermore, if the gNB wants to terminate model training at the UE if it takes too long, or if it has other data to transmit, such as high-priority messages to the UE, or if the UE is receiving data from the gNB or another UE, the UE in RRC_INACTIVE or RRC_TRAING mode can still respond to high-priority messages, such as emergency messages or interrupt signals.

[0339] According to the embodiments, the AI / ML model trained by the UE is either an untrained AI / ML model or a pre-trained AI / ML model that needs to be improved or updated. For example, if the UE does not have sufficient processing power, has limited battery power, or is busy computing another AI / ML model, the AI / ML model can be pre-trained by another network entity or a core network entity and sent to the UE, which will then update the model by only using a training set that is still needed.

[0340] According to the implementation example, the training set is

[0341] -A complete training set designed to train AI / ML models from scratch, or

[0342] - A portion of the training set, intended for fine-tuning pre-trained AI / ML models, or

[0343] - When retraining the model using an updated training set relative to the initial training set, additional training samples are added to improve model performance.

[0344] According to the embodiment, UE

[0345] - Train AI / ML models using predefined training procedures or training sets; for example, you can define training procedures and / or training sets, and

[0346] -from

[0347] ○One or more measurements performed by the UE, and / or

[0348] ○ Network entities in wireless communication networks or network entities different from wireless communication networks, such as databases in the Internet.

[0349] Obtain the training set.

[0350] For example, in the above scenario, some parts of the training may depend on the radio channel, such as channel state information (CSI), like SINR, or on the receiver configuration, such as a receiver configured to receive multiple radio streams, or on a specific procedure or processing running on the UE, such as a HARQ procedure or the number of retransmissions. This measurement or data may be uniquely available at the UE, allowing the UE to measure the information used.

[0351] According to an embodiment, one or more of the following can be applied to the training time. The training time can be...

[0352] - (pre-)configured, or

[0353] -The UE sends a signal to the network entity of the wireless communication network to notify it of the training time, or

[0354] -The network sends a signal to the UE to notify it of the training time.

[0355] Training time is the time required / allocated by the UE to train an AI / ML model using the training set.

[0356] According to the embodiment, during the training of the AI / ML model, the UE uses

[0357] - Non-AI rollback process, and / or

[0358] - Enter training mode, for example, with reduced connectivity, and / or

[0359] - A trained version of the AI / ML model.

[0360] According to the embodiment, the UE sends a signal to notify the network entity of the wireless communication network.

[0361] - Estimated training time required for the AI / ML model, and / or

[0362] - The completion of AI / ML model training, optionally including an indication of which AI / ML model was trained when using more than one AI / ML model, or

[0363] - An interruption signal, which stops or interrupts training. In this case, the UE can also signal the reason, such as overheating or being busy with other AI / ML training.

[0364] According to an embodiment, the UE sends an interruption signal to the network entity of the wireless communication network. The interruption signal indicates that the UE should stop training or interrupt training and / or indicates the reason for stopping or interrupting, such as overheating or being busy with other AI / ML training.

[0365] According to an embodiment, the network entity of the wireless communication network includes one or more of the following:

[0366] -Another UE,

[0367] -Remote UE,

[0368] -Relay UE,

[0369] - Radio access network (RAN) entities, such as gNBs or roadside units (RSUs).

[0370] - Core network (CN) entities, such as Access and Mobility Functions (AMF) or Location Management Functions (LMF).

[0371] According to an embodiment, a method is provided for operating a user equipment (UE) of a wireless communication network, the UE being configured or pre-configured with one or more AI / ML models for performing one or more specific operations, the wireless communication network using one or more artificial intelligence / machine learning (AI / ML) models for one or more use cases, the method including training the AI / ML models using a training set.

[0372] According to an embodiment, model training can be performed online, i.e., on the fly. In this training mode, the UE collects a training set on the fly from its latest measurements and learns these processes using predefined training procedures. This allows the model to be trained from scratch or to improve / update an already pre-trained model. In an alternative scenario, the training set can be provided by a network or another external entity, such as a database. The training set can be a complete training set designed to train the model from scratch, or it can be an update of the training set. The UE can follow the following process:

[0373] ■ Training time: This is the time required for the UE to train based on a specific training set. This time can be configured by the specification or the network (pre-)configured. It can also be a formula, for example, a larger training set requires more training time. Furthermore, it can also be signaled by the UE to the network / gNB.

[0374] ■ During the training period: As long as the training period has not elapsed, the network / gNB assumes that the AI ​​model is not yet ready. This may mean that only non-AI fallback procedures are applied during this period. In another embodiment, the UE may apply an already trained AI model, but not the updated model. The updated model is only used after the training period has elapsed.

[0375] ■ Exchange of model training time: The UE can signal the estimated training time to the network / gNB.

[0376] ■ Signal the UE when model training is complete and for which models, for example, when considering more than one model.

[0377] ■ Signal to stop or interrupt training, for example, by using an interrupt signal. In this case, the UE may also optionally signal a reason, such as overheating or being busy with other AI / ML training.

[0378] 6. Self-benchmarking

[0379] One aspect of the embodiments described herein relates to self-benchmarking of such functionality.

[0380] According to an embodiment, an apparatus for a wireless communication network is provided, which uses one or more artificial intelligence / machine learning (AI / ML) models for one or more use cases.

[0381] The device determines the performance of one or more AI / ML models used by one or more network entities in a wireless communication network to perform one or more specific operations.

[0382] According to an embodiment, if it is determined that an AI / ML model does not perform as expected, such as if the AI / ML model produces worse performance than a non-AI / ML model method used to perform a specific operation or below a certain threshold, the apparatus causes the network entity to modify the method used to perform the specific operation.

[0383] According to an embodiment, in order to modify the method for performing a specific operation, the apparatus causes the network entity to:

[0384] - Switch from a specific AI / ML model to another AI / ML model for performing a specific operation, or

[0385] - Report performance to another network entity, or

[0386] -Disable a specific AI / ML model and apply a non-AI / ML model method that performs a specific operation, or

[0387] - Switch from the current operating mode to the new operating mode, or

[0388] - Switch to training, testing, or evaluation mode.

[0389] According to an embodiment, the device includes network entities using AI / ML models, for example,

[0390] -User Equipment (UE), or

[0391] - Remote UE, or

[0392] - Relay UE, or

[0393] - Radio access network (RAN) entities, such as gNBs or roadside units (RSUs).

[0394] - Core network (CN) entities, such as Access and Mobility Functions (AMF) or Location Management Functions (LMF).

[0395] Alternatively or otherwise, the device is separated from one or more network entities that use AI / ML models. For example, the device includes another network entity of a wireless communication network or an entity of a network different from the wireless communication network, such as an entity of the Internet.

[0396] According to an embodiment, the apparatus includes a user equipment (UE), which uses one or more AI / ML models to perform one or more specific operations, and monitors the performance of one or more AI / ML models and provides performance metrics, and / or

[0397] The UE provides reports on performance metrics to the wireless communication network, and / or

[0398] When it is determined that an AI / ML model does not perform as expected, for example, when an AI / ML model produces worse performance than a non-AI / ML model method used to perform a specific operation, UE

[0399] - Switch from a specific AI / ML model to another AI / ML model for performing a specific operation, or

[0400] -Disable a specific AI / ML model and apply a non-AI / ML model method that performs a specific operation, or

[0401] - Switch from the current operating mode to the new operating mode, or

[0402] - Switch to training, testing, or evaluation mode.

[0403] According to the embodiment, UE

[0404] - In response to a request from the wireless communication network, and / or

[0405] - In response to one or more pre-configured conditions, and / or

[0406] - Periodically, where the periodicity can be pre-configured according to specifications or can be configured by the wireless communication network.

[0407] Provide reports on performance metrics.

[0408] According to an embodiment, the UE provides a report on performance metrics in response to one or more pre-configured conditions, which include one or more of the following:

[0409] - Packet error rate (PER), e.g., high PER or low PER,

[0410] -Bit error rate (BER)

[0411] Decoding failed.

[0412] - Radio link failure (RLF)

[0413] - At least one beam recovery procedure has been performed or is currently being performed.

[0414] - At least one performance metric, such as the mean square error of the compressed model against actual measurements and throughput.

[0415] According to an embodiment, the report is associated with a test window that collects the required data for the report, and the test window has multiple configuration parameters that are pre-configured according to specifications and / or configured by the wireless communication network.

[0416] According to an embodiment, the multiple configuration parameters include one or more of the following:

[0417] - Window size, which defines the time required to collect the data for the report. The window size can be indicated in terms of duration, such as seconds, milliseconds, microseconds, nanoseconds, number of time slots, number of subframes, number of OFDM symbols, or number of cycles.

[0418] - One or more parameters indicating the time and / or frequency resources of the test signal or the type of test sequence used.

[0419] - Periodicity of one or more test windows,

[0420] - One or more performance metrics to be measured and reported during the test window, where performance metrics may include one or more error metrics such as mean squared error, cross-entropy loss, absolute error, and throughput.

[0421] According to an embodiment, the UE is configured or pre-configured with thresholds for one or more error or performance metrics, and if one, a certain number, or all of the thresholds are exceeded, a specific AI / ML model is switched on / off / modified and / or the operating mode is switched and / or a report is triggered. Modifying the AI / ML model may refer to updating model weights, adding / replacing certain layers, or training / fine-tuning the model.

[0422] According to an embodiment, the apparatus includes a RAN entity serving a user equipment (UE), such as a gNB or RSU, wherein the UE uses one or more AI / ML models to perform one or more specific operations, and the RAN entity monitors the performance of one or more AI / ML models performed by the UE and provides performance metrics.

[0423] The RAN entity receives baseline data from the UE, and performance metrics are determined based on the baseline data.

[0424] If it is determined that an AI / ML model is not performing as expected, such as when the AI / ML model produces worse performance than a non-AI / ML model method used to perform a specific operation, the RAN entity causes the UE to...

[0425] - Switch from a specific AI / ML model to another AI / ML model for performing a specific operation, or

[0426] - Modify specific AI / ML models, for example, by updating weights or changing some adaptive / fine-tuning layers, or

[0427] -Disable a specific AI / ML model and apply a non-AI / ML model method that performs a specific operation, or

[0428] - Switch to training, testing, or evaluation mode.

[0429] - Switch from the current operating mode to the new operating mode.

[0430] According to an embodiment, the device obtains baseline data from a test window, which can be defined relative to a reference time and / or space and / or frequency. The baseline data may include one or more of the following:

[0431] -Additional measurement signal,

[0432] - Possibly more complex different models; and / or

[0433] - Traditional process.

[0434] According to an embodiment, the report includes one or more performance metrics, such as throughput, reconstruction error, for example, mean absolute or squared reconstruction error of CSI, SINR difference, number of retransmissions, number of ACK / NACK, and ACK-NACK ratio.

[0435] According to an embodiment, a method is provided for operating an apparatus for a wireless communication network that uses one or more artificial intelligence / machine learning (AI / ML) models for one or more use cases. The method includes determining the performance of one or more AI / ML models used in one or more network entities of the wireless communication network to perform one or more specific operations.

[0436] Using AI models in practice can present several challenges. For example, an AI model might perform much worse than expected. This could be due to, for instance, a mismatch between the training set and real-world field data. It could also be that the AI ​​model fails to generalize. In these cases, performance may be significantly worse compared to existing fallback mechanisms. According to embodiments, the device can compare the performance of one or more AI / ML models against a fallback mechanism and / or any other threshold that may be dynamic, defined, or predefined. Therefore, performance must be monitored, and if performance is insufficient, decommissioning the AI ​​must be considered.

[0437] Performance monitoring can be performed at the UE or the network / gNB. If performed at the gNB, the UE can report the baseline data obtained from the fallback mechanism to the gNB. If performed at the UE, the UE can report errors / performance metrics to the network / gNB.

[0438] The report can be initiated in the following circumstances:

[0439] • Requests from gNB / network, and / or

[0440] • Periodically, where periodicity can be configured by the specification or gNB / network (pre-)configured, and / or

[0441] • Triggered by performance / error metrics exceeding a specific threshold.

[0442] Alternatively or additionally, each report can be associated with a test window, in which the data required for the report is collected. This test window can have multiple configuration parameters, pre-configured by the specification and / or gNB / network:

[0443] • Window size, the time required to collect the data for the report, such as the duration in milliseconds, seconds, time slots, frames, subframes, or OFDM symbols.

[0444] • Error / Performance Indicators

[0445] There may be multiple error metrics, such as mean squared error, cross-entropy loss, absolute error, throughput, etc.

[0446] The network can configure one or more error / performance metrics for the UE, which are measured and reported to the network during the test window.

[0447] For use cases such as CSI prediction, the gNB does not need to know whether the UE is using AI or a fallback mechanism. In such cases, the UE can also autonomously decide to switch back to the fallback mechanism if performance is insufficient. The UE can be (pre)configured with thresholds for one or more error / performance metrics, and switches to the fallback mechanism if one, a specific number, or all thresholds are exceeded. Thresholds and / or error / performance metrics can be configured according to model / model ID / AI functionality and / or globally.

[0448] Embodiments of this disclosure particularly relate to wireless communication systems, such as 3GPP systems or WiFi systems, including user equipment (UE) and / or devices according to any of the preceding claims.

[0449] According to embodiments, the user equipment (UE) or device or wireless communication network of any of the preceding claims can be designated as

[0450] UEs include one or more of the following: power-limited UEs; or handheld UEs, such as those used by pedestrians and referred to as vulnerable road users (VRUs); or pedestrian UEs (P-UEs); or personal or handheld UEs used by public safety personnel and emergency responders and referred to as public safety UEs (PS-UEs); or IoT UEs, such as sensors, actuators, or UEs provided in a campus network that perform repetitive tasks and request input from gateway nodes at periodic intervals; or mobile terminals; or stationary terminals; or cell IoT-UEs; or SL UEs; or vehicle UEs; or vehicle group leader UEs (GL-UEs); or dispatch UEs (S-UEs); or IoT or narrowband IoT ( NB-IoT devices; or ground-based vehicles; or aircraft; or unmanned aerial vehicles; or mobile base stations; or roadside units (RSUs); or buildings; or any other item or device that provides network connectivity enabling the item / device to communicate using a wireless communication network, such as sensors or actuators; or any other item or device that provides network connectivity enabling the item / device to communicate using a sidelink of a wireless communication network, such as sensors or actuators; or Wi-Fi devices, stations (STAs), access points (APs), nodes, or mesh nodes; or mesh points; or mesh APs; or any network entity with sidelink capability, and

[0451] The network entities of the wireless communication system include one or more of the following:

[0452] - Base stations, such as macro cell base stations, or small cell base stations, or central units of base stations, or distributed units of base stations, or integrated access and backhaul (IAB) nodes, or Wi-Fi devices such as access points (APs) or mesh nodes (mesh APs).

[0453] -Roadside Unit (RSU)

[0454] -UE, such as SL UE, or group leader UE (GL-UE) or relay UE,

[0455] - Remote wireless head,

[0456] - Core network entities, such as Access and Mobility Management Function (AMF), Service Management Function (SMF), or Mobile Edge Computing (MEC) entities.

[0457] - Such as network slicing in the context of NR or 5G core network.

[0458] - Any Transmitter Point (TRP) that enables an item or device to communicate using a wireless communication network, wherein the item or device is provided with network connectivity to communicate using a wireless communication network.

[0459] The various elements and features of this invention can be implemented in hardware using analog and / or digital circuitry, in software, by instructions executed by one or more general-purpose or special-purpose processors, or as a combination of hardware and software. For example, embodiments of this invention can be implemented in the environment of a computer system or another processing system. Figure 12 An example of a computer system 500 is shown. Units or modules, and the steps of methods performed by these units, can be executed on one or more computer systems 500. The computer system 500 includes one or more processors 502, such as dedicated or general-purpose digital signal processors. The processors 502 are connected to a communication infrastructure 504, such as a bus or network. The computer system 500 includes main memory 506, such as random access memory (RAM), and secondary memory 508, such as a hard disk drive and / or a removable storage drive. The secondary memory 508 can allow computer programs or other instructions to be loaded into the computer system 500. The computer system 500 may also include a communication interface 510 to allow software and data to be transferred between the computer system 500 and external devices. This communication can be in the form of electronic, electromagnetic, optical, or other signals that can be processed by the communication interface. This communication can use wires or cables, optical fibers, telephone lines, cellular telephone links, RF links, and other communication channels 512.

[0460] The terms "computer program medium" and "computer-readable medium" are used to refer generally to tangible storage media, such as removable storage units or hard disks installed in hard disk drives. These computer program products are means for providing software to computer system 500. The computer program, also known as computer control logic, is stored in main memory 506 and / or secondary storage 508. The computer program can also be received via communication interface 510. When executed, the computer program enables computer system 500 to implement the present invention. Specifically, when executed, the computer program enables processor 502 to implement the processes of the present invention, such as any of the methods described herein. Thus, such a computer program can represent the controller of computer system 500. In the case of implementing this disclosure using software, the software can be stored in the computer program product and loaded into computer system 500 using a removable storage drive or an interface such as communication interface 510.

[0461] The implementation in hardware or software can be executed using digital storage media, such as cloud storage, floppy disks, DVDs, Blu-rays, CDs, ROMs, PROMs, EPROMs, EEPROMs, or flash memory, which store electronically readable control signals that cooperate (or are capable of cooperating with) a programmable computer system to execute the corresponding methods. Therefore, the digital storage media can be computer-readable.

[0462] Some embodiments of the invention include a data carrier having electronically readable control signals, which is capable of cooperating with a programmable computer system to perform one of the methods described herein.

[0463] Typically, embodiments of the present invention can be implemented as a computer program product having program code that, when run on a computer, is operable to perform one of these methods. The program code may, for example, be stored on a machine-readable medium.

[0464] Other embodiments include a computer program stored on a machine-readable medium for performing one of the methods described herein. In other words, therefore, one embodiment of the method of the present invention is a computer program having program code that, when run on a computer, performs one of the methods described herein.

[0465] Therefore, another embodiment of the method of the present invention is a data carrier (or digital storage medium, or computer-readable medium) including a computer program recorded thereon for performing one of the methods described herein. Thus, another embodiment of the method of the present invention is a data stream or signal sequence representing a computer program for performing one of the methods described herein. The data stream or signal sequence may, for example, be configured to be transmitted via a data communication connection, such as via the Internet. Another embodiment includes a processing device, such as a computer or programmable logic device, configured or adapted to perform one of the methods described herein. Another embodiment includes a computer on which a computer program for performing one of the methods described herein is installed.

[0466] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, the field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, these methods are preferably performed by any hardware device.

[0467] The above embodiments are merely illustrative of the principles of the invention. It should be understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. Therefore, the invention is limited only by the scope of the forthcoming patent claims, and not by the specific details presented through the description and explanation of the embodiments herein.

[0468] abbreviation

[0469]

[0470]

Claims

1. An apparatus of a wireless communication network, the wireless communication network using one or more artificial intelligence / machine learning, AI / ML, models for one or more use cases, wherein the apparatus determines an inference time of one or more of the AI / ML models to be used in one or more network entities of the wireless communication network.

2. The apparatus of claim 1, wherein the inference time comprises a time required to process the AI / ML model completely or partially, the inference time being provided in the form of an absolute time or an offset value.

3. The apparatus of claim 2, wherein the inference time is provided in the form of one or more of the following: - seconds, milliseconds, microseconds, nanoseconds; multiples of these time units, number of time slots, number of subframes, number of OFDM symbols, number of periods, - an offset value indicating an offset time relative to a reference time, e.g. a reference time provided by a navigation system, e.g. GPS; an offset relative to a start of a frame; or an offset relative to a frame structure, such as a physical downlink control channel, PDCCH, or a synchronization signal, e.g. a primary synchronization sequence, PSS, or a secondary synchronization sequence, SSS, or a sidelink synchronization sequence transmitted via a physical sidelink broadcast channel, PSBCH, at least one of the group.

4. The apparatus of claim 2 or 3, wherein the inference time comprises a time required to process the AI / ML model partially, wherein the part is a part of the AI / ML model to be processed; wherein the AI / ML model comprises a part not to be processed.

5. The apparatus of any one of the preceding claims, wherein the inference time of an AI / ML model is determined using an inference time model, the inference time model calculating the inference time using one or more first properties of the AI / ML model and / or one or more second properties of the network entity for which at least a part of the AI / ML model is to be used.

6. The apparatus of claim 5, wherein each of the AI / ML models comprises a specific neural network, and the network entity comprises specific hardware for implementing the specific neural network, and the one or more first properties of the AI / ML model comprise one or more properties of the neural network, and the one or more second properties of the network entity comprise one or more properties of the hardware.

7. The apparatus of claim 6, wherein the properties of the neural network comprise one or more of the following: - a number of layers of the neural network, - a depth of the neural network, e.g. a number of layers that have to be executed sequentially, - a number of specific operations, e.g. floating point operations, multiplication, addition, integer operations, Boolean operations, exponential function, - a width of a layer of the neural network, e.g. an input size, IS, and / or an output size, OS, - a type of a layer of the neural network, e.g. a convolutional layer, an activation layer, a batch normalization, or a fully connected layer, and the properties of the hardware comprise one or more of the following: - the number of hardware accelerator units, e.g. the number of graphics processing units, GPUs, or the number of tensor processing units, TPUs, or the number of tensor cores, - the processor speed, e.g. the number of floating point operations per second, FLOPS, the number of additions per second, the number of multiplications per second, the number of integer operations per second, - the number of processor cores, - the type of processing cores, - the combination of processing cores, e.g. x number of GPU cores and y number of tensor cores, - the memory size, - the memory speed, - the type of memory, - the memory architecture.

8. The apparatus according to any one of the preceding claims, wherein the AI / ML models used in the wireless communication network are uniquely numbered and identifiable, and the apparatus determines the inference time of the supported AI / ML model identification, ID, using one or more of the following: - the processing time of the supported AI / ML model ID, - the number or set of supported AI / ML models to be processed in parallel or sequentially.

9. The apparatus according to any one of the preceding claims, wherein the AI / ML models used in the wireless communication network are uniquely numbered and identifiable, wherein the apparatus determines the inference time of at least one specific supported AI / ML model that can operate as an individual AI / ML in a use case model; and / or wherein the apparatus determines the inference time of at least one set of supported AI / ML models that can be operated simultaneously for the use case.

10. The apparatus according to any one of the preceding claims, wherein a specific AI / ML model to be used in a network entity is inferred from the identification of a specific feature or function supported by the network entity, e.g. n-bit CSI feedback inferring a specific AI / ML model implementing a precoding engine, or n-bit SINR feedback inferring a specific AI / ML model implementing a handover function.

11. The apparatus according to any one of the preceding claims, wherein the apparatus comprises a network entity using AI / ML models, e.g. - a user equipment, UE, or - a remote UE, or - a relay UE, or - a radio access network, RAN, entity, like a gNB or a road side unit, RSU, or - a core network, CN, entity, like an access and mobile function, AMF, or a location management function, LMF, and / or the apparatus is separate from one or more network entities using AI / ML models, e.g. the apparatus comprises another network entity of the wireless communication network or an entity of a network different from the wireless communication network, like an entity of the internet.

12. The apparatus according to any one of the preceding claims, wherein the apparatus indicates that a specific AI / ML model is available or not available on a specific network entity, and / or falls back to a default procedure if the determined inference time of a specific AI / ML model is equal to or less than a predefined or (pre-)configured processing time of one or more operations of a use case using the specific AI / ML model.

13. The apparatus of claim 12, wherein the apparatus communicates via sidelink, and wherein the processing time is configured in a resource pool configuration, RP.

14. The apparatus of any one of the preceding claims, wherein the apparatus indicates an inference time of a specific AI / ML model or AI / ML function to the network and / or network entity and / or gNB.

15. The apparatus of any one of the preceding claims, wherein the use case comprises one or more of the following: - Channel State Information, CSI, prediction - CSI compression, - Hybrid Automatic Repeat Request, HARQ, prediction, - Positioning of a user equipment, - Beam management, - Beam prediction, - Beam adaptation, - Mobility enhancement, - SINR prediction, - SL resource allocation, - SL sensing, - Handover, HO, or conditional handover, CHO, - Discovery.

16. The apparatus of any one of the preceding claims, wherein the apparatus indicates the inference time to one or more user equipment, UE, communicating via sidelink, SL.

17. The apparatus of claim 16, the apparatus being arranged in - a RAN entity, like a gNB or RSU, for aligning inference time between a plurality of UEs operating in mode 1, or - a SL UE, or remote UE, or - a relay UE, or - a plurality of UEs for coordinating inference time over sidelink when operating in mode 1 or mode 2, e.g. o during SL synchronization and / or SL discovery and / or SL connection setup phase, e.g. in transmission of a Physical Sidelink Broadcast Channel, PSBCH, or o using signaling via a Physical Sidelink Control Channel, PSCCH, o using signaling embedded in a Physical Sidelink Shared Channel, PSSCH, o using feedback exchange via a Physical Sidelink Feedback Channel, PSFCH.

18. A user equipment, UE, of a wireless communication network using one or more Artificial Intelligence / Machine Learning, AI / ML, models for one or more use cases, wherein the UE uses one or more of the AI / ML models, and wherein the UE signals an inference time required by the UE to execute one or more of the AI / ML models to the wireless communication network.

19. The user equipment, UE, of claim 18, wherein the UE signals the inference time to at least one of a gNB, a UE, and a relay UE.

20. The user equipment, UE, of claim 18 or 19, wherein the UE - in response to a transmission of one or more of the AI / ML models from a network entity of the wireless communication network to the UE, or - in response to an activation of one or more of the AI / ML models and / or AI / ML functions from a network entity of the wireless communication network to the UE, or - in response to a request from a network entity of the wireless communication network, e.g. in case the UE is preconfigured with one or more AI / ML models or after one or more AI / ML models are transmitted to the UE, or - when accessing the wireless communication network, in case the UE is pre-configured with one or more AI / ML models, e.g. together with the signaling of the UE capabilities, signaling the inference time.

21. The user device, UE, of claim 20, wherein the network entity of the wireless communication network that transmits the AI / ML model or requests the inference time comprises one or more of: - a further UE, or a relay UE, or a remote UE, - a radio access network, RAN, entity, like a gNB or a road side unit, RSU, - a core network, CN, entity, like an access and mobility function, AMF, or a location management function, LMF.

22. The user device, UE, of any one of claims 10 to 21, wherein the inference time comprises a time needed to process the AI / ML model completely or partially, the inference time being provided in form of an absolute time or an offset value.

23. The apparatus of claim 22, wherein the inference time is provided in form of one or more of: - seconds, milliseconds, microseconds, nanoseconds; multiples of these time units, number of slots, number of subframes, number of OFDM symbols, number of periods, - an offset value indicating an offset time relative to a reference time, e.g. a reference time provided by a navigation system, e.g. GPS; an offset relative to a start of a frame; or an offset relative to a frame structure, such as a physical downlink control channel, PDCCH, or a synchronization signal, e.g. a primary synchronization sequence, PSS, or a secondary synchronization sequence, SSS, or a sidelink synchronization sequence transmitted via a physical sidelink broadcast channel, PSBCH, of at least one of the group.

24. The apparatus of claim 22 or 23, wherein the inference time comprises a time needed to process the AI / ML model partially, wherein the part is a part of the AI / ML model to be processed; wherein the AI / ML model comprises a part that is not processed.

25. The user device, UE, of any one of claims 18 to 24, wherein the UE: - determines the inference time, e.g. using an inference time model that uses at least one or more properties of the AI / ML model and one or more properties of the UE, or - receives the inference time from the wireless communication network, e.g. from an apparatus of any one of claims 1 to 17, or from a network entity comprising an apparatus of any one of claims 1 to 17, like a RAN entity of a CN entity, or from another UE, e.g. via a sidelink interface, also referred to as PC5.

26. The user device, UE, of any one of claims 11 to 15, wherein the UE signals a number of instances of a specific AI / ML model and / or a number of AI / ML models that the UE is capable of processing in parallel.

27. The user device, UE, of any one of claims 11 to 16, wherein the UE selects the inference time for a specific AI / ML model to be signaled from a set of configured or pre-configured inference times that the UE is capable of achieving when executing the specific AI / ML model.

28. The user device, UE, of any one of claims 18 to 27, wherein the inference time is at least a portion of a processing time required to process the particular AI / ML model.

29. The user device, UE, of any one of claims 11 to 28, wherein the UE signals an inference time of the particular AI / ML model to the wireless communication network only if the inference time allows performing the particular AI / ML model according to a processing time constraint associated with the use case using the particular AI / ML model.

30. The user device, UE, of claim 29, wherein the inference time of the particular AI / ML model is associated with a particular AI / ML model identification, ID, or functionality, and the UE reports the AI / ML model ID only if the UE is able to meet the processing time constraint.

31. A user device, UE, of a wireless communication network using one or more artificial intelligence / machine learning, AI / ML, models for one or more use cases, wherein the UE performs one or more of the AI / ML models to be used for performing one or more particular operations, wherein the UE signals a complexity or capacity that the UE is able to perform to the wireless communication network such that a particular operation is performed using a particular AI / ML model within a predefined processing time associated with the particular operation, and wherein, in response to the signaling, the UE receives from the wireless communication network one or more of the AI / ML models that the UE is able to perform for performing the particular operation according to the predefined processing time.

32. The user device, UE, of claim 31, wherein the complexity or capacity relates to at least one of: - a number of layers of a neural network of an AI / ML model, - a depth of a neural network of an AI / ML model, e.g., a number of layers that have to be executed sequentially, - a number of certain operations, e.g., floating point operations, multiplications, additions, integer operations, Boolean operations, exponential functions, - a width of a layer of a neural network of an AI / ML model, e.g., an input size, IS, and / or an output size, OS, - a type of a layer of a neural network of an AI / ML model, e.g., a convolutional layer, an activation layer, a batch normalization, or a fully connected layer, and - a number of hardware accelerator units of the UE, e.g., a number of graphics processing units, GPUs, or a number of tensor processing units, TPUs, or a number of tensor cores, - a processor speed of the UE, e.g., a number of floating point operations per second, FLOP, a number of additions per second, a number of multiplications per second, a number of integer operations per second, - a number of processing cores, - a type of processing cores, - a combination of processing cores, e.g., x number of GPU cores and y number of tensor cores, - a memory size of the UE, - a memory speed of the UE, - a memory type of the UE, - a memory architecture of the UE.

33. The user device, UE, of any one of claims 18 to 32, wherein in case the AI / ML model currently used or requested to be used does not meet the predefined processing time, the UE receives from the wireless communication network a fallback AI / ML model or information indicating to proceed according to a fallback procedure to be used, or wherein the UE is (pre-)configured to use a fallback procedure in case the processing time cannot be met by the AI / ML model currently used or requested to be used.

34. A user device, UE, of a wireless communication network using one or more artificial intelligence / machine learning, AI / ML, models for one or more use cases, wherein the UE is configured or pre-configured with one or more AI / ML models for performing one or more specific operations, and wherein the UE trains the AI / ML model using a training set.

35. The user device, UE, of claim 34, wherein the UE trains the AI / ML model while being connected to the wireless communication network.

36. The user device, UE, of claim 34 or 35, wherein the UE changes its connected mode to, - a training mode or evaluation mode, e.g. RRC TRAINING or RRC EVALUATION mode, or - a different RRC mode, e.g. RRC Inactive or RRC Idle mode when training an AI / ML model, or - another connected mode, e.g. DRX mode, paging mode.

37. The user device, UE, of any one of claims 34 to 36, wherein the AI / ML model the UE trains is an untrained AI / ML model or a pre-trained AI / ML model to be improved or updated.

38. The user device, UE, of any one of claims 34 to 37, wherein the training set is - a complete training set aiming at training an AI / ML model from scratch, or - a partial training set aiming at fine-tuning a pre-trained AI / ML model, or - an updated training set, updated with respect to an initial training set, adding additional training samples to improve the model performance when retraining the model in combination with the initial training set.

39. The user device, UE, of any one of claims 34 to 38, wherein the UE - trains the AI / ML model using a predefined training procedure or training set, and - obtains the training set from o one or more measurements performed by the UE, and / or o a network entity of the wireless communication network or an entity of a different network than the wireless communication network, like a database in the internet.

40. The user device, UE, of any one of claims 34 to 39, wherein - the training time is (pre-)configured, or - the UE signals the training time to a network entity of the wireless communication network, or - the network signals the training time to the UE, - the training time is the time the UE needs / allocates to train the AI / ML model using the training set. ​ 41. The user device, UE, of any one of claims 34 to 40, wherein during training of the AI / ML model, the UE uses - a non-AI fallback procedure, and / or - enters a training mode, e.g. with reduced connectivity, and / or - a trained version of the AI / ML model.

42. The user device, UE, of any one of claims 34 to 41, wherein the UE signals to a network entity of the wireless communication network - an estimated time needed for training the AI / ML model, and / or - a completion of training of the AI / ML model, optionally together with an indication of which AI / ML models were trained, or - or a stop signal, which stops or interrupts training.

43. The user device, UE, of any one of claims 34 to 42, wherein the UE signals to a network entity of the wireless communication network a stop signal, which indicates to stop or interrupt training and / or indicates a reason for the stop or interruption, e.g. overheating, busy with other AI / ML training.

44. The user device, UE, of any one of claims 34 to 43, wherein the network entity of the wireless communication network comprises one or more of: - a further UE, - a remote UE, - a relay UE, - a radio access network, RAN, entity, like a gNB or a road side unit, RSU, - a core network, CN, entity, like an access and mobility function, AMF, or a location management function, LMF.

45. A wireless communication system, like a third generation partnership project, 3GPP, system or a WiFi system, comprising a user device, UE, and / or apparatus according to any one of the preceding claims.

46. The user device, UE, or apparatus or wireless communication network of any one of the preceding claims, the UE comprising one or more of: a power limited UE; or a handheld UE, like a UE used by a pedestrian and referred to as a vulnerable road user, VRU; or a pedestrian UE, P-UE; or a body or handheld UE used by public safety and emergency personnel and referred to as a public safety UE, PS-UE; or an IoT UE, e.g. a sensor, actuator or UE provided in a campus network performing a repetitive task and requiring input from a gateway node at periodic intervals; or a mobile terminal; or a stationary terminal; or a cell IoT-UE; or a SL UE; or a vehicle UE; or a group leader UE, GL-UE; or a scheduling UE, S-UE; or an IoT or narrowband IoT, NB-IoT, device; or a ground based vehicle; or an aircraft; or an unmanned aircraft; or a mobile base station; or a road side unit, RSU; or a building. ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ wherein ​ ​ or any other item or device, e.g. a sensor or actuator, provided with network connectivity enabling the item / device to communicate using the wireless communication network; or any other item or device, e.g. a sensor or actuator, provided with network connectivity enabling the item / device to communicate using a sidelink of the wireless communication network; or a Wi-Fi device, a station (STA), an access point (AP), a node or mesh node; or a mesh point; or a Mesh AP; or any sidelink-capable network entity, and wherein the network entities of the wireless communication system comprise one or more of: - a base station, like a macro cell base station, or a small cell base station, or a central unit of a base station, or a distributed unit of a base station, or an integrated access and backhaul, IAB, node, or a Wi-Fi device like an access point (AP) or a mesh node (mesh AP); - a road side unit, RSU, - a UE, like a SL UE, or a group leader UE, GL-UE, or a relay UE, - a remote radio head, - a core network entity, like an access and mobility management function, AMF, or a service management function, SMF, or a mobile edge computing, MEC, entity, - a network slice, like in the context of NR or 5G core, - any transmission / reception point, TRP, enabling an item or device to communicate using the wireless communication network, the item or device being provided with network connectivity to communicate using the wireless communication network.

47. A method for operating an apparatus of a wireless communication network, the wireless communication network using one or more artificial intelligence / machine learning, AI / ML, models for one or more use cases, the method comprising: determining an inference time of one or more of the AI / ML models to be used in one or more network entities of the wireless communication network.

48. A method for operating a user equipment, UE, of a wireless communication network, the wireless communication network using one or more artificial intelligence / machine learning, AI / ML, models for one or more use cases, the method comprising: using one or more of the AI / ML models, and signaling to the wireless communication network an inference time required for executing one or more of the AI / ML models.

49. A method for operating a user equipment, UE, of a wireless communication network, the UE executing one or more of AI / ML models to be used for performing one or more specific operations, the wireless communication network using one or more artificial intelligence / machine learning, AI / ML, models for one or more use cases, the method comprising: signaling to the wireless communication network a complexity or capacity the UE is capable of executing, such that the specific operation is performed using a specific AI / ML model within a predefined processing time associated with the specific operation, and in response to the signaling, receiving from the wireless communication network one or more of the AI / ML models the UE is capable of executing for performing the specific operation according to the predefined processing time.

50. A method for operating a user equipment (UE) of a wireless communication network, the UE being configured or preconfigured with one or more AI / ML models for performing one or more specific operations, the wireless communication network using one or more artificial intelligence / machine learning (AI / ML) models for one or more use cases, the method comprising: training the AI / ML model using a training set.

51. A non-transitory computer program product comprising a computer-readable medium storing instructions for performing the method of any one of claims 47-50 when executed on a computer.