Multi-exit neural networks for controllable accuracy and distributed training / inference of ai / ML models for mobile communication

WO2026104637A1PCT designated stage Publication Date: 2026-05-21FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
Filing Date
2025-11-14
Publication Date
2026-05-21

Smart Images

  • Figure EP2025083093_21052026_PF_FP_ABST
    Figure EP2025083093_21052026_PF_FP_ABST
Patent Text Reader

Abstract

An apparatus of a wireless communication system according to an embodiment is provided. The apparatus comprises a storage (210) having stored thereon a model (215) comprising a plurality of sub-models and / or having stored thereon one or more submodels of the plurality of sub-models of the model (215), wherein the model (215) is an artificial intelligence / machine learning model. Moreover, the apparatus comprises a processing unit (220) configured for executing the model (215) and / or the one or more sub-models of the model (215) to produce an output of the model (215) depending on an input of the model (215).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Multi-Exit Neural Networks for Controllable Accuracy and Distributed Training / lnference of AI / ML Models for Mobile Communication

[0002] Description

[0003] The present invention relates to multi-exit neural networks, and, in particular, to multi-exit neural networks for controllable accuracy and distributed training / inference of AI / ML models for mobile communication.

[0004] Fig. 1 is a schematic representation of an example of a terrestrial wireless network 100 including, as is shown in Fig. 1(a), the core network and one or more radio access networks RANi, RAN2, ... RANN (RAN = Radio Access Network). Fig. 1(b) is a schematic representation of an example of a radio access network RANnthat may include one or more base stations gNBi to gNBs (gNB = next generation Node B), each serving a specific area surrounding the base station schematically represented by respective cells IO61 to IO65. The base stations are provided to serve users within a cell. The one or more base stations may serve users in licensed and / or unlicensed bands. The term base station, BS, refers to a gNB in 5G networks, an eNB in UMTS / LTE / LTE-A / LTE-A Pro, or just a BS in other mobile communication standards. A user may be a stationary device or a mobile device. The wireless communication system may also be accessed by mobile or stationary loT (Internet of Things) devices which connect to a base station or to a user. The mobile devices or the loT devices may include physical devices, ground based vehicles, such as robots or cars, aerial vehicles, such as manned or unmanned aerial vehicles, UAVs, the latter also referred to as drones, buildings and other items or devices having embedded therein electronics, software, sensors, actuators, or the like as well as network connectivity that enables these devices to collect and exchange data across an existing network infrastructure. Fig. 1(b) shows an exemplary view of five cells, however, the RANnmay include more or less such cells, and RANnmay also include only one base station. Fig. 1(b) shows two users UEi and UE2, (UE = User Equipment) also referred to as user equipment, UE, that are in cell IO62 and that are served by base station gNB2. Another user UE3 is shown in cell IO64 which is served by base station gNB4. The arrows IO81, IO82 and IO83 schematically represent uplink / downlink connections for transmitting data from a user UE1, UE2 and UE3 to the base stations gNB2, gNB4 or for transmitting data from the base stations gNB2, gNB4 to the users UE1, UE2, UE3. This may be realized on licensed bands or on unlicensed bands. Further, Fig. 1(b) shows two loT devices 110i and HO2 in cell IO64, which may be stationary or mobile devices. The loT device 110i accesses the wireless communication system via the base station gNB4 to receive and transmit data as schematically represented by arrow 112i. The loT device HO2 accesses

[0005] FH241106PEP-2025373084.DOCX the wireless communication system via the user UE3 as is schematically represented by arrow 1122. The respective base stations gNBi to gNBs may be connected to the core network 102, e.g. via the S1 interface, via respective backhaul links 114i to 114s, which are schematically represented in Fig. 1(b) by the arrows pointing to “core”. The core network 102 may be connected to one or more external networks. The external network may be the Internet or a private network, such as an intranet or any other type of campus networks, e.g. a private WiFi or 4G or 5G mobile communication system. Further, some or all of the respective base stations gNBi to gNBs may be connected, e.g. via the S1 or X2 interface or the XN interface in NR (New Radio), with each other via respective backhaul links 1161 to 1165, which are schematically represented in Fig. 1(b) by the arrows pointing to “gNBs”. A sidelink channel allows direct communication between UEs, also referred to as device-to-device, D2D (Device to Device), communication. The sidelink interface in 3GPP (3G Partnership Project) is named PC5 (Proximity-based Communication 5).

[0006] For data transmission a physical resource grid may be used. The physical resource grid may comprise a set of resource elements to which various physical channels and physical signals are mapped. For example, the physical channels may include the physical downlink, uplink and sidelink shared channels, PDSCH (Physical Downlink Shared Channel), PUSCH (Physical Uplink Shared Channel), PSSCH (Physical Sidelink Shared Channel), carrying user specific data, also referred to as downlink, uplink and sidelink payload data, the physical broadcast channel, PBCH (Physical Broadcast Channel), carrying for example a master information block, MIB, and one or more of a system information block, SIB, one or more sidelink information blocks, SLIBs, if supported, the physical downlink, uplink and sidelink control channels, PDCCH (Physical Downlink Control Channel), PUCCH (Physical Uplink Control Channel), PSCCH (Physical Sidelink Control Channel), the downlink control information, DCI, the uplink control information, UCI, and the sidelink control information, SCI, and physical sidelink feedback channels, PSFCH (Physical sidelink feedback channel), carrying PC5 feedback responses. Note, the sidelink interface may support a 2-stage SCI (Speech Call Items). This refers to a first control region comprising some parts of the SCI, and, optionally, a second control region, which comprises a second part of control information.

[0007] For the uplink, the physical channels may further include the physical random-access channel, PRACH (Packet Random Access Channel) or RACH (Random Access Channel), used by UEs for accessing the network once a UE synchronized and obtained the MIB and SIB. The physical signals may comprise reference signals or symbols, RS, synchronization signals and the like. The resource grid may comprise a frame or radio frame having a certain duration in the time domain and having a given bandwidth in the

[0008] FH241106PEP-2025373084.DOCX frequency domain. The frame may have a certain number of subframes of a predefined length, e.g. 1ms. Each subframe may include one or more slots of 12 or 14 OFDM symbols (OFDM = Orthogonal Frequency-Division Multiplexing) depending on the cyclic prefix, CP, length. A frame may also include of a smaller number of OFDM symbols, e.g. when utilizing a shortened transmission time interval, sTTI (slot or subslot transmission time interval), or a mini-slot / non-slot-based frame structure comprising just a few OFDM symbols.

[0009] The wireless communication system may be any single-tone or multicarrier system using frequency-division multiplexing, like orthogonal frequency-division multiplexing, OFDM, or orthogonal frequency-division multiple access, OFDMA (Orthogonal frequency-division multiple access), or any other IFFT-based signal (IFFT = Inverse Fast Fourier Transformation) with or without CP, e.g. DFT-s-OFDM (DFT = discrete Fourier transform). Other waveforms, like non-orthogonal waveforms for multiple access, e.g. filter-bank multicarrier, FBMC, generalized frequency division multiplexing, GFDM, or universal filtered multi carrier, UFMC, may be used. The wireless communication system may operate, e.g., in accordance with the LTE-Advanced pro standard, or the 5G or NR, New Radio, standard, or the NR-U, New Radio Unlicensed, standard.

[0010] The wireless network or communication system depicted in Fig. 1 may be a heterogeneous network having distinct overlaid networks, e.g., a network of macro cells with each macro cell including a macro base station, like base stations gNBi to gNBs, and a network of small cell base stations, not shown in Fig. 1, like femto or pico base stations. In addition to the above described terrestrial wireless network also non-terrestrial wireless communication networks, NTN, exist including spaceborne transceivers, like satellites, and / or airborne transceivers, like unmanned aircraft systems. The non-terrestrial wireless communication network or system may operate in a similar way as the terrestrial system described above with reference to Fig. 1, for example in accordance with the LTE- Advanced Pro standard or the 5G or NR, new radio, standard.

[0011] In mobile communication networks, for example in a network like that described above with reference to Fig. 1, like an LTE or 5G / NR network, there may be UEs that communicate directly with each other over one or more sidelink, SL, channels, e.g., using the PC5 / PC3 interface or WiFi direct. UEs that communicate directly with each other over the sidelink may include vehicles communicating directly with other vehicles, V2V communication, vehicles communicating with other entities of the wireless communication network, V2X communication, for example roadside units, RSUs, or roadside entities, like traffic lights, traffic signs, or pedestrians. An RSU may have a functionality of a BS or of a

[0012] FH241106PEP-2025373084.DOCX UE, depending on the specific network configuration. Other UEs may not be vehicular related UEs and may comprise any of the above-mentioned devices. Such devices may also communicate directly with each other, D2D communication, using the SL channels.

[0013] In a wireless communication network, like the one depicted in Fig. 1, it may be desired to locate a UE with a certain accuracy, e.g., determine a position of the UE in a cell. Several positioning approaches are known, like satellite-based positioning approaches, e.g., autonomous and assisted global navigation satellite systems, A-GNSS, such as GPS, mobile radio cellular positioning approaches, e.g., observed time difference of arrival, OTDOA, and enhanced cell ID, E-CID, or combinations thereof.

[0014] In the context of 5G Framework for 3GPP AI / ML functionalities (Al: artificial intelligence; ML: machine-learning) will be identified, managed, and utilized for UE-side models and UE-part of two-sided models.

[0015] Regarding known wireless communication systems, only common models and additionally fine tuning is studied.

[0016] In the following, the term “functionality” may, e.g., be defined, for example, according to its definition in TR38.843: According to the definition there, in functionality-based LCM, network indicates activation / deactivation / fallback / switching of AI / ML functionality via 3GPP signaling (e.g., RRC, MAC-CE, DOI). Models may not be identified at the Network, and UE may perform model-level LCM. Whether and how much awareness / interaction NW should have about model-level LCM requires further study. For functionality identification, there may be either one or more than one functionality defined within an AI / ML-enabled feature, whereby AI / ML-enabled Feature refers to a Feature where AI / ML may be used. Note: UE may have one AI / ML model for the functionality, or UE may have multiple AI / ML models for the functionality.

[0017] In model-l D-based LCM, models are identified at the Network, and Network / UE may activate / deactivate / select / switch individual AI / ML models via model ID.

[0018] When it comes to functionality / model management, 3GPP has made the following agreements / discussions in TR38.843 and in ongoing Release 19 regarding AI / ML for NR air interface:

[0019] According to the agreement for LCM for UE-sided model / Common LCM framework / signaling (RAN2, Rel-19), for UE-sided model, for the functionality

[0020] FH241106PEP-2025373084.DOCX management, the “network decision, network-initiated” AI / ML management is supported as a baseline. The following can be considered further “UE autonomous, decision reported to the network”, “Network decision, U E-initiated” (i.e. proactive approach). Moreover, “UE-autonomous, UE’s decision is not reported to the network” is not considered for Rel-19

[0021] According to TR38843, methods to assess / monitor the applicability and expected performance of an inactive model / functionality, including the following examples for the purpose of activation / selection / switching of UE-side models / UE-part of two-sided models / functionalities (if applicable) are an assessment / monitoring based on the additional conditions associated with the model / functionality, an assessment / monitoring based on input / output data distribution, an assessment / monitoring using the inactive model / functionality for monitoring purpose and measuring the inference accuracy, and an assessment / monitoring based on past knowledge of the performance of the same model / functionality (e.g., based on other UEs).

[0022] The above indicates how the RAN1 Rel-18 was thinking about selection of most suitable model / functionality.

[0023] In the agreement RAN1, Rel-19 CSI-compression is considered. The assumption is that the decoder is trained and then different UE vendors can develop their respective encoders. In our case, this process is different, since the UE part(s) / submodels can also be generated first, and a subset of the model (a number of sub-models in the “middle”) can be trained jointly and / or can be known to both sides (UE and NW).

[0024] From RAN1 perspective, for UE part of two-sided model, further study the following example of Ml-Option2 (including the feasibility / necessity) would be appreciated.

[0025] In Al-Example2-1, step A, a dataset is transferred from the NW / NW-side to UE / UE-side via standardized signaling. It should be noted that RAN1 study of Step A only focuses on RAN1 aspect of the dataset transfer from NW to UE. Other solution for dataset exchange is out of RAN1 scope.

[0026] In step B, a UE part of two-sided model(s) is(are) developed based on at least the above dataset.

[0027] In step C, the UE reports information of its UE part of two-sided model(s) corresponding to the above dataset to the NW. As for further study, explanations would be appreciated how

[0028] FH241106PEP-2025373084.DOCX a model ID is determined / assigned for each AI / ML model (including relationship between dataset and model ID). It should be noted that some step(s) may not be needed for Ml-Option 2.

[0029] Moreover, it should be noted that the above example is based on the assumption of NW-first training. It is separate discussion for the assumption of UE-first training.

[0030] Furthermore, it should be noted that the study should consider the impact on inter-vendor collaboration, at least including complexity, performance, interoperability in RAN4 / testing related aspects and feasibility.

[0031] It would be appreciated for further study whether / how to consider UE-side additional condition(s) for the dataset

[0032] Multiple exit points in early exit networks correspond to different sub-networks with varying scale and complexity. They provide the option to terminate / exit the execution of a neural network earlier with a usable result, see, e.g., Fig. 4.

[0033] It would be appreciated, if improved concepts for distributed artificial intelligence / machine learning models would be provided.

[0034] The object of the present invention is to provide improved concepts for wireless communication systems. The object of the present invention is solved by the subject¬ matter of the independent claims. Particular embodiments are provided in the dependent claims.

[0035] An apparatus of a wireless communication system according to an embodiment is provided. A model comprises a plurality of sub-models, wherein the model is an artificial intelligence / machine learning model. Moreover, the apparatus is configured for executing the model and / or one or more sub-models of the plurality of sub-models of the model to produce an output of the model depending on an input of the model.

[0036] According to an embodiment, the model and / or the one or more sub-models may, e.g., be stored in the apparatus.

[0037] In an embodiment, the apparatus may, e.g., comprise a storage having stored thereon the model comprising the plurality of sub-models and / or having stored thereon the one or more sub-models of the plurality of sub-models of the model. Moreover, the apparatus

[0038] FH241106PEP-2025373084.DOCX may, e.g., comprise a processing unit configured for executing the model and / or the one or more sub-models of the model to produce the output of the model depending on the input of the model.

[0039] Furthermore, an apparatus of a wireless communication system according to an embodiment is provided. The apparatus comprises a storage having stored thereon a model comprising a plurality of sub-models and / or having stored thereon one or more sub¬ models of the plurality of sub-models of the model, wherein the model is an artificial intelligence / machine learning model. Moreover, the apparatus comprises a processing unit configured for executing the model and / or the one or more sub-models of the model to produce an output of the model depending on an input of the model.

[0040] According to an embodiment, the apparatus may, e.g., comprise at least one exit point of the model configured to provide an intermediate output associated with a functionality of the wireless communication system.

[0041] In an embodiment, the apparatus may, e.g., comprise an activation control unit configured to activate and / or control at least one of the plurality of sub-models of the model depending on requirements and / or measurements within the wireless communication system.

[0042] Moreover, a method of a wireless communication system according to an embodiment is provided, wherein a model comprises a plurality of sub-models, wherein the model is an artificial intelligence / machine learning model. The method comprises executing the model and / or one or more sub-models of the plurality of sub-models of the model to produce an output of the model depending on an input of the model.

[0043] Moreover, a method of a wireless communication system according to an embodiment is provided. An apparatus of the wireless communication system comprises a storage having stored thereon a model comprising a plurality of sub-models and / or having stored thereon one or more sub-models of the plurality of sub-models of the model, wherein the model is an artificial intelligence / machine learning model. The method comprises executing the model and / or the one or more sub-models of the model, by a processing unit of the apparatus, to produce an output of the model depending on an input of the model.

[0044] FH241106PEP-2025373084.DOCX Furthermore, a computer program for implementing the method the above-described embodiment is provided, when the computer program is executed on a computer or signal processor.

[0045] Some embodiments are based on the finding that a potential suitability and expected performance of an AI / ML model should not (only) be associated with static properties, like, e.g., the encapsulating functionality or a pre-defined area, but should instead be individually determined based on the measurements available at each point in time. This implies that depending on the measured input and a performance requirement / constraint, a different model or functionality would need to be applied. Larger / complex models would be required for challenging tasks, while smaller / simpler models for easier settings. Functionalities that facilitate more measurements / signaling / resources should only be activated for challenging tasks.

[0046] Some embodiments provide a dynamic / fluid approach for the functionality (and model) selection / management.

[0047] In the following, embodiments of the present invention are described in more detail with reference to the figures, in which:

[0048] Fig. 1 illustrates a schematic representation of an example of a terrestrial wireless network.

[0049] Fig. 2 illustrates an apparatus of a wireless communication system according to an embodiment.

[0050] Fig. 3 illustrates different configurations of a backbone neural network comprising of two sub-models according to an embodiment.

[0051] Fig. 4 illustrates an early exit network illustration.

[0052] Fig. 5 illustrates an example of a computer system on which units or modules as well as the steps of the methods described in accordance with the inventive approach may execute.

[0053] Fig. 6 illustrates a modular joint source, channel coding and modulation (JSCCM) implementation according to an embodiment.

[0054] FH241106PEP-2025373084.DOCX In the following, particular embodiments are described with reference to the drawings.

[0055] An apparatus of a wireless communication system according to an embodiment is provided. A model 215 comprises a plurality of sub-models, wherein the model 215 is an artificial intelligence / machine learning model. Moreover, the apparatus is configured for executing the model 215 and / or one or more sub-models of the plurality of sub-models of the model 215 to produce an output of the model 215 depending on an input of the model 215.

[0056] According to an embodiment, the model 215 and / or the one or more sub-models may, e.g., be stored in the apparatus.

[0057] In an embodiment, the apparatus may, e.g., comprise a storage 210 having stored thereon the model 215 comprising the plurality of sub-models and / or having stored thereon the one or more sub-models of the plurality of sub-models of the model 215. Moreover, the apparatus may, e.g., comprise a processing unit 220 configured for executing the model 215 and / or the one or more sub-models of the model 215 to produce the output of the model 215 depending on the input of the model 215.

[0058] Such an embodiment is depicted in Fig. 2, where a particular apparatus of a wireless communication system according to an embodiment is illustrated

[0059] The apparatus of Fig. 2 comprises a storage 210 having stored thereon a model 215 comprising a plurality of sub-models and / or having stored thereon one or more sub¬ models of the plurality of sub-models of the model 215, wherein the model 215 is an artificial intelligence / machine learning model.

[0060] Moreover, the apparatus of Fig. 2 comprises a processing unit 220 configured for executing the model 215 and / or the one or more sub-models of the model 215 to produce an output of the model 215 depending on an input of the model 215.

[0061] According to an embodiment, the apparatus may, e.g., comprise at least one exit point 216 of the model 215 configured to provide an intermediate output associated with a functionality of the wireless communication system. This at least one exit point 216 may, e.g., be referred to as an early exit or early exit point.

[0062] FH241106PEP-2025373084.DOCX In an embodiment, the apparatus may, e.g., further be configured to determine the capability of supporting at least one exit point 216 of the model 215 and provide an output related to this determined capability in response to a request.

[0063] In an embodiment, the apparatus may, e.g., comprise an activation control unit 230 configured to activate and / or control at least one of the plurality of sub-models of the model 215 depending on requirements and / or measurements within the wireless communication system.

[0064] According to an embodiment, the activation control unit 230 may, e.g., further be configured to activate or de-activate at least one of the plurality of sub-models of the model 215 based on a configuration message comprising parameters received from a higher layer configuration of the wireless communication system.

[0065] In an embodiment, the activation control unit 230 may, e.g., further be configured to activate or de-activate at least one of the plurality of sub-models of the model (215) based on a configuration message comprising parameters received from a higher layer configuration of the wireless communication system.

[0066] According to an embodiment, the apparatus may, e.g., be configured to conduct a measurement and / or to report on a measurement to produce the intermediate output associated with the functionality of the wireless communication system.

[0067] In an embodiment, the apparatus may, e.g., be configured to conduct a measurement and / or to report a measurement with respect to a signal transmitted within the wireless communication system. And / or, the apparatus may, e.g., be configured to conduct a measurement and / or to report a measurement with respect to a beam used for transmitting and / or receiving data within the wireless communication system. And / or, the apparatus may, e.g., be configured to conduct a measurement and / or to report a measurement with respect to a position of the apparatus or of another device of the wireless communication system and / or the apparatus may, e.g., be configured to conduct a measurement or a set of measurements and / or to utilise a measurement to predict / estimate a second measurement or event.

[0068] According to an embodiment, the apparatus may, e.g., be configured to performed or estimate or predict the measurement to select a configuration from a set of configurations provided by the network entity, wherein the configuration indicates at least one cell (e.g., a primary cell) for the apparatus.

[0069] FH241106PEP-2025373084.DOCX In an embodiment, the apparatus may, e.g., be configured to log the measurement performed or estimated or predicted by the apparatus, wherein in response to request from the network or in response to the logged measurements or information exceeding a certain size, the apparatus is configured to transfer the recorded data to at least one network entity.

[0070] According to an embodiment, the apparatus may, e.g., be configured to conduct a prediction and / or a measurement to determine if one cell is stronger than another cell in future time; and / or the apparatus may, e.g., be configured to conduct a prediction and / or a measurement of a radio link failure within a certain time window; and / or the apparatus may, e.g., be configured to conduct a detection of a radio link failure.

[0071] According to an embodiment, the model 215 may, e.g., comprise two or more exit points as the at least one exit point 216, wherein each of the two or more exit points may, e.g., be configured to provide an intermediate output associated with a functionality of the wireless communication system.

[0072] In an embodiment, the intermediate output of each of the two or more exit points may, e.g., be associated with a different functionality being different from the functionality which is associated with the intermediate output of any other exit point of the two or more exit points.

[0073] According to an embodiment, the apparatus may, e.g., be configured to conduct a measurement and / or to report on a measurement to produce the intermediate output associated with the functionality of the wireless communication system.

[0074] In an embodiment, the apparatus may, e.g., be configured to conduct a measurement and / or to report a measurement with respect to a signal transmitted within the wireless communication system. And / or, the apparatus may, e.g., be configured to conduct a measurement and / or to report a measurement with respect to a beam used for transmitting and / or receiving data within the wireless communication system. And / or, the apparatus may, e.g., be configured to conduct a measurement and / or to report a measurement with respect to a position of the apparatus or of another device of the wireless communication system.

[0075] According to an embodiment, the model 215 comprises two or more exit points as the at least one exit point 216, wherein each of the two or more exit points may, e.g., be

[0076] FH241106PEP-2025373084.DOCX configured to provide an intermediate output associated with a functionality of the wireless communication system.

[0077] In an embodiment, the intermediate output of each of the two or more exit points may, e.g., be associated with a different functionality being different from the functionality which is associated with the intermediate output of any other exit point of the two or more exit points.

[0078] According to an embodiment, the apparatus may, e.g., be a user equipment.

[0079] In an embodiment, the model 215 and / or the one or more sub-models of the plurality of sub-models of the model 215 being stored in the storage 210 of the user equipment are trained at the user equipment individually for the user equipment.

[0080] According to an embodiment, a first sub-model of the plurality of sub-models of the model 215 may, e.g., be stored in the storage 210 of the user equipment. A second sub-model of the plurality of sub-models of the model 215 may, e.g., be stored in a storage of a network entity of the wireless communication system.

[0081] In an embodiment, the second sub-model may, e.g., be trained at the network entity.

[0082] According to an embodiment, the model 215 comprises a first group of one or more sub¬ models being stored in the storage 210 of the user equipment. The model 215 comprises a second group of one or more sub-models being stored in a storage of the network entity.

[0083] In an embodiment, the first group of one or more sub-models may, e.g., be trained at the user equipment without sharing implementation information with the network entity. The second group of one or more sub-models may, e.g., be trained at the network entity without sharing implementation information with the user equipment.

[0084] According to an embodiment, the model 215 comprises a third group of one or more sub¬ models being shared and / or being trained to be shared with the user equipment and with the network entity.

[0085] In an embodiment, the user equipment may, e.g., be configured to inform the network entity on one or more of its capabilities (e.g., on how many sub-models it can execute due to hardware / latency / battery / etc. constraints),

[0086] FH241106PEP-2025373084.DOCX According to an embodiment, as an example of said hardware / computation constraints, the AI / ML sub-models may, e.g., be associated with different computational requirements. The processing unit 220 may be embodied as or comprise one or more processors to meet these requirements. Such processors may, e.g., comprise an APU (APU: Al Processing Unit) and / or, e.g., a CPU. The APU may, e.g., comprise several processing units, such as Google®’ s Tensor Process Units (TPUs), Qualcomm®’ s Neural Processing Units (NPUs), Apple®’s Neural Engine (ANE), Samsung®’s Exynos, etc. The UE may, e.g., report different computational requirements associated with different AI / ML sub-models.

[0087] According to an embodiment, the user equipment may, e.g., be adapted to be configured by the network entity to load a specific number of sub-models for inference depending on performance requirements (e.g., worst-case or average-case performance requirements), wherein said sub-models and associated early exits are configured to provide outputs for same or different functionalities (e.g., depending on the model properties and the configuration from the network).

[0088] In an embodiment, to be configured by the network entity to load a specific number of sub¬ models, the user equipment may, e.g., be configured to receive from the network entity one or more of the following:

[0089] a model identifier identifying the model 215,

[0090] a number of sub-models to be loaded,

[0091] an identifier or an index of an early exit,

[0092] a flag indicating if only the output of the early exit or also an output of a sub-model shall be provided.

[0093] According to an embodiment, the apparatus may, e.g., be a network entity.

[0094] In an embodiment, the model 215 may, e.g., be a first model. The apparatus comprises an input interface for receiving irreversible processed information from another apparatus of the wireless communication system, wherein the irreversible processed information may, e.g., be an output of another model or may, e.g., be a processed output derived from the output of the other model, wherein the other model may, e.g., be another artificial intelligence / machine learning model of said other apparatus or of a further apparatus of

[0095] FH241106PEP-2025373084.DOCX the wireless communication system, wherein the output of the other model results from feeding confidential input into the other model, wherein the confidential input may, e.g., be not reconstructable from the irreversible processed information. The input interface may, e.g., be configured to feed the irreversible processed information into the first model 215 or into at least one of the one or more sub-models of the first model 215.

[0096] According to an embodiment, the model 215 comprises at least two sub-models. A first sub-model and a second sub-model of the at least two sub-models are arranged such that a sub-model output of the first sub-model may, e.g., be fed into a second sub-model as a first part of an input of the second sub-model. The apparatus comprises an input interface being configured to receive further input information from another apparatus of the wireless communication system, and being configured to feed the further input information as a second part of the input of the second sub-model into the second sub-model.

[0097] In an embodiment, the input interface may, e.g., be configured for receiving irreversible processed information as the further input information from said other apparatus, wherein the irreversible processed information may, e.g., be an output of another model or may, e.g., be a processed output derived from the output of the other model, wherein the other model may, e.g., be another artificial intelligence / machine learning model of said other apparatus or of a further apparatus of the wireless communication system, wherein the output of the other model results from feeding confidential input into the other model, wherein the confidential input may, e.g., be not reconstructable from the irreversible processed information. The input interface may, e.g., be configured to feed the irreversible processed information as the second part of the input of the second sub-model into the second sub-model.

[0098] According to an embodiment, the model 215 may, e.g., be a first model. The apparatus comprises an input model interface for receiving an external output of a further model being a further artificial intelligence / machine learning model of a further device of the wireless communication system or for receiving a processed output being derived from the external output. The apparatus may, e.g., be configured to feed the external output or the processed output or data being derived from the output or from the processed output as input into the model 215 or into a sub-model of the plurality of sub-models of the model 215.

[0099] In an embodiment, the model 215 can be trained with input data being independent from the further model.

[0100] FH241106PEP-2025373084.DOCX According to an embodiment, the external output or the processed output may, e.g., be irreversible such that confidential input which has been fed into the further model may, e.g., be not reconstructable from the external output or from the processed output.

[0101] In an embodiment, the model 215 may, e.g., be a first model. The apparatus may, e.g., be configured to provide a model output of the first model 215 to another device of the wireless communication system for providing said other device with input for a further model being a further artificial intelligence / machine learning model of said other device, wherein a model output of said further model may, e.g., be employed for conducting a functionality of the wireless communication system.

[0102] According to an embodiment, the first model 215 can be trained with output data being independent from the further model.

[0103] In an embodiment, the model output of the first model 215 which is provided to the further model may, e.g., be irreversible such that model input that has been fed into the first model 215 to obtain the model output of the first model 215 may, e.g., be not reconstructable from the model output of the first model 215.

[0104] According to an embodiment, the first apparatus comprises a battery for providing energy for the apparatus. The apparatus may, e.g., be configured to decide whether or not to provide the model output of the first model 215 to said other device depending on a state of the battery.

[0105] In an embodiment, the apparatus may, e.g., be configured to decide whether or not to provide the model output of the first model 215 to said other device depending on a signalling overhead.

[0106] According to an embodiment, the apparatus may, e.g., be configured to receive an external model output from the other device, wherein the external model output may, e.g., be a model output of the further model of the other device in response to inputting the model output of the first model 215 into the further model of the other device.

[0107] In an embodiment, the apparatus may, e.g., be configured to conduct a functionality of the wireless communication system using the external model output.

[0108] According to an embodiment, the apparatus may, e.g., be configured to control a quality or a reliability of the model output of the first model 215 using the external model output.

[0109] FH241106PEP-2025373084.DOCX In an embodiment, the other device may, e.g., be a network component of the wireless communication system.

[0110] According to an embodiment, the apparatus may, e.g., be a user equipment of the wireless communication system, and the other device may, e.g., be another user equipment of the wireless communication system.

[0111] In an embodiment, the apparatus may, e.g., be a network entity of the wireless communication system, and the other device may, e.g., be another user equipment of the wireless communication system.

[0112] According to an embodiment, the first model 215 and / or the further model may, e.g., be available (e.g., is common) for a plurality of devices of the wireless communication system.

[0113] In an embodiment, the first model 215 may, e.g., be not available for any other apparatus of the wireless communication system. Or, the further model may, e.g., be not available for the apparatus.

[0114] According to an embodiment, the apparatus may, e.g., be configured to execute a first task using a model output of the model 215. The apparatus may, e.g., be configured to execute a second task, being different from the first task, using the intermediate output being provided at the at least one exit point 216 of the model 215. At least one of the first task and the second task comprises conducing a functionality of the wireless communication system.

[0115] In an embodiment, at least one of the first task and the second task may, e.g., be a task to transmit data and / or to receive data and / or to determine a position of the apparatus or of another apparatus of the wireless communication system and / or to control a beam management, and / or to determine target primary cell and / or time for handover and / or a configuration of serving cells comprising primary and secondary cells to be applied and / or indicate to higher layers to initiate handover to different cell or to a cell on a different RAT, and / or to determine a configuration for conditional handover or to select the handover base station directly.

[0116] According to an embodiment, the apparatus may, e.g., be configured to execute the first task using the model output without using the intermediate output being provided at the at

[0117] FH241106PEP-2025373084.DOCX least one exit point 216 of the model 215. And / or, the apparatus may, e.g., be configured to execute the second task using the intermediate output being provided at the at least one exit point 216 of the model 215 without using the model output.

[0118] In an embodiment, the intermediate output being provided at the at least one exit point 216 of the model 215 may, e.g., be an output of a first sub-model of the model 215. The model output of the model 215 may, e.g., be an output of a second sub-model of the model 215. The model 215 and / or the second sub-model may, e.g., be trained using model training output depending on the first task. The first sub-model may, e.g., be trained using model training output depending on the second task.

[0119] According to an embodiment, the intermediate output being provided at the at least one exit point 216 may, e.g., be an output of a first one of a plurality of layers of the model 215. The model output of the model 215 may, e.g., be an output of a second one of the plurality of layers of the model 215, being different from the first one of the plurality of layers.

[0120] In an embodiment, the storage 210 of the apparatus has stored therein a first one of the plurality of sub-models of the model 215. A further apparatus of the wireless communication system comprises a second one of the plurality of sub-models of the model 215.

[0121] According to an embodiment, the apparatus may, e.g., be configured to provide an output of the first one of the plurality of sub-models of the model 215 to the further apparatus for being input into the second one of the plurality of sub-models of the model 215.

[0122] In an embodiment, the apparatus may, e.g., be configured to receive an output of the second one of the plurality of sub-models of the model 215 from the further apparatus. The apparatus may, e.g., be configured to conduct a functionality of the wireless communication system using the output of the second one of the plurality of sub-models of the model 215.

[0123] According to an embodiment, the storage 210 of the apparatus has stored therein only a first one of the plurality of sub-models of the model 215. A further apparatus of the wireless communication system comprises a second one of the plurality of sub-models of the model 215. The apparatus may, e.g., be configured to provide an output of the first one of the plurality of sub-models to the further apparatus for being input into the second one of the plurality of sub-models.

[0124] FH241106PEP-2025373084.DOCX In an embodiment, the model 215 may, e.g., be a neural network.

[0125] According to an embodiment, the neural network may, e.g., be a CNN or may, e.g., be a LSTM or may, e.g., be a Bayesian NN or comprises transformer layers.

[0126] In an embodiment, the apparatus may, e.g., be configured to determine a confidence in a prediction of the model output and / or a confidence in a prediction of the intermediate output being provided at the at least one exit point 216.

[0127] According to an embodiment, the activation control unit 230 may, e.g., be configured to activate one or more subsequent layers following a layer where the model output is outputted depending on the confidence in a prediction of the model output and / or depending on the confidence in the prediction of the intermediate output being provided at the at least one exit point 216.

[0128] In an embodiment, the model 215 and / or the one-or more sub-models of the plurality of sub-models of the model 215 have been trained using a loss function.

[0129] According to an embodiment, the loss function may, e.g., be implemented as a MSE, or as a top-k accuracy, or as a F1 -score.

[0130] In an embodiment, the apparatus may, e.g., be configured to employ the intermediate output being provided at the at least one exit point 216 for determining an AoA and / or a ToA and / or a LOS / NLOS classification.

[0131] According to an embodiment, the apparatus may, e.g., be configured to employ the intermediate output being provided at the at least one exit point 216 for determining one or more beams with a highest RSRP among a plurality of beams.

[0132] In an embodiment, the apparatus may, e.g., be configured to employ the intermediate output being provided at the at least one exit point 216 to predict one or more beams with a highest RSRP of wide beams among the plurality of beams. The apparatus may, e.g., be configured to employ the intermediate output being provided at the at least one exit point 216 to predict one or more beams with a highest RSRP of narrow beams among the plurality of beams. Moreover, the apparatus may, e.g., be configured to employ the intermediate output being provided at the at least one exit point 216 may, e.g., be an

[0133] FH241106PEP-2025373084.DOCX output of a layer of the model 215 preceding a another layer of the model 215 which outputs the second model output.

[0134] According to an embodiment, the apparatus may, e.g., be configured to employ the intermediate output to predict a time to radio link failure (e.g. with the serving cell). And / or, the apparatus may, e.g., be configured to employ the intermediate output to predict measurements (e.g., RSRP, RSRQ, SINR) for the available base stations, wherein the apparatus may, e.g., be configured to employ the model output of the model 215 to determine a configuration for conditional handover or to select the handover base station directly.

[0135] In an embodiment, the plurality of sub-models of the model 215 comprise a first sub¬ model and a second sub-model. The apparatus may, e.g., be configured to determine a coarse position of the apparatus depending on the intermediate output being provided at the at least one exit point 216. Moreover, the apparatus may, e.g., be configured to feed an output of the first sub-model as an input into the second sub-model to obtain a model output of the model 215 from an output of the second sub-model. Furthermore, the apparatus may, e.g., be configured to determine a more precise position of the apparatus compared to the coarse position depending on the output of the second sub-model.

[0136] According to an embodiment, the apparatus may, e.g., be configured to determine the more precise position of the apparatus depending on the output of the second sub-model and depending on another additional positioning method.

[0137] In an embodiment, a sub-model of the model 215 which provides the intermediate output at the at least one exit point 216 may, e.g., be updated by another entity using in-device fine-tuning.

[0138] According to an embodiment, the apparatus may, e.g., be configured to provide the intermediate output of the at least one exit point 216 to one or more network entities to perform inference.

[0139] In an embodiment, the intermediate output being provided at the at least one exit point 216 comprise one or more weights of the model 215 being a neural network.

[0140] According to an embodiment, the apparatus may, e.g., be configured to indicate its capability to provide optimized models to another device of the wireless communication system.

[0141] FH241106PEP-2025373084.DOCX In an embodiment, the capability may, e.g., be limited to certain areas.

[0142] Moreover, a system according to an embodiment is provided.

[0143] The system comprises an apparatus according to one of the above-described embodiments.

[0144] Moreover, the system comprises the other apparatus described above and / or the further apparatus described above and / or the further device described above, and / or the other device described above.

[0145] In the following, particular embodiments of the present invention are described.

[0146] At first, a high-level overview is provided.

[0147] As loading / activating different models depending on the input measurements is practically infeasible due to hardware limitations in the UE (or NW), we utilize a modified form of early exit (EE) neural networks (NNs) as the basis for our solution.

[0148] As seen in Fig. 4, an EE neural network (setup) has typically the following properties:

[0149] The backbone NN model is organized in sub-models. Each sub-model can contain an arbitrary number of NN layers of different types (e.g., CNN, LSTM, Bayesian NN, transformer layers, etc.)

[0150] The outputs of selected branching points of the backbone network are fed to smaller neural networks (side-branches) and an early exit (output) is produced. This output is the same as the final network output (e.g., the UE position for positioning or the top-K beams for beam management), but with reduced accuracy. Typically, accuracy of early exits increases from left to right.

[0151] The different sub-models can be trained or activated for inference in different entities (e.g., first ml run at the UE, the following m2 run at the gNB and the final ones at the LMF).

[0152] For classification tasks, each early exit can provide its confidence in its prediction. This implies that later sub-models can be activated in a sequential manner if and only if they

[0153] FH241106PEP-2025373084.DOCX are required. So, the decision on which sub-models to utilize and to which early exit to terminate the inference process is made automatically, depending on the specific input in the overall NN.

[0154] As will be described in more detail below, we have the following improvements / alterations, focusing on mobile radio communication (3gpp, 5G+) application:

[0155] Inputs from different entities (that are not part of the model input in the left - see Fig. 4), can be passed at sub-models. (This is illustrated as “additional information for sub-model K-1” in Fig. 4). This way, each “part” (collection of sub-models that are executed in the UE, gNB or NW) can be private from the other side and use privileged information / inputs.

[0156] The backbone NN can be trained in a way that maintains the proprietary requirements of the sub-models running at the UE and the sub-models running at the NW (so these are trained separately).

[0157] Part of the neural network model (for example the “middle” part, but not limited to this) can be common, so it can run anywhere. This allows flexibility of execution, even when the other parts of the model are proprietary.

[0158] Each branching point can be trained to provide outputs that have contextual meaning, so the implementation of disjoint training is easier.

[0159] Different early exits can correspond to different functionalities (in the TR38.843 sense).

[0160] Implications on the various AI / ML model / functionality life cycle management (TR38.843) steps:

[0161] Inference can flexibly consider the following constraints: Latency (for example, can treat inherently “no delay,” “low delay” and “delay tolerant” positioning requirements set in TS22.071); performance requirements; UE / gNB / NW hardware / battery / capability / load / etc; energy savings requirements; signaling overhead (e.g., can execute more sub-models to transmit less information).

[0162] Model monitoring and management may, e.g., implement parallel inference and monitoring with small overhead. It may, e.g., be run at the UE to provide output of early exit due to latency requirements, to execute more sub-models with some delay and evaluate the accuracy of the early exit output. It may, e.g., be run at the UE to provide

[0163] FH241106PEP-2025373084.DOCX output of early exit due to latency requirements. Moreover, it may, e.g., be run at the UE to execute more sub-models at the NW and evaluate the accuracy of the early exit output that runs at the UE; to provide outputs that belong to different functionalities at the same time. Moreover, it may, e.g., be run at the UE to enables low-latency decision and implementation to switch functionalities.

[0164] Training / re-training / fine-tuning may, e.g., comprise (repeat) training in different sides, respecting proprietary nature of (sub-) models. It may, e.g., moreover comprise fine-tuning by adding sub-model(s), for example, initial sub-models trained to the NW to extended / fine-tuned (and compiled, pruned, etc.) for each UE separately by UE vendors; and, for example, initial sub-models trained through UE vendor collaboration to obtain extended / fine-tuned in the NW (e.g., for adapting to specific areas).

[0165] In the following, the points above are described in more detail.

[0166] At first, training aspects of embodiments, in particular, training (or fine-tuning) of particular embodiments within a 3GPP scenario is described.

[0167] An aspect of embodiments employs loss functions and sequential training or fine-tuning.

[0168] The structure of the early exit model can be decided from the beginning. The output from each exit point of each side-branch calculates its own loss (based on loss functions suitable for the task, like MSE, top-k accuracy, F1-score, etc.) and these are combined [average, sum or weighted sum] to a single loss calculation.

[0169] In a complementary approach, a baseline (initial) early exit model is pre-trained and it is kept fixed. New sub-models with accompanying early exits are added and only these are further trained [e.g., for further accuracy, fine-tuning to new environment, etc.]

[0170] According to a first example of an embodiment, direct positioning (NW -> UE) is considered. NW-side trains a “global” early-exit positioning model, which can achieve an average level of accuracy. The model or several sub-models and a training dataset is shared with UE vendors. Each vendor adds sub-models that are tailored (e.g., compiled / compressed) for their underlying hardware. This way, both accuracy and computation / memory / latency / hardware capability constraints from the UE side can be addressed.

[0171] FH241106PEP-2025373084.DOCX According to a second example of another embodiment, Beam management (UE -> NW) is considered. UE vendors develop a beam management model. It has good average performance, but cannot address every combination of geometry / radio conditions. A NW-side adds sub-models that are tailored for the conditions of each specific area, thus achieving better accuracy in these areas.

[0172] According to a third example of another embodiment, UE mobility (e.g. handover, beam¬ selection, beam-switching) is considered (UE -> NW). UE vendors develop a mobility management model that is able to predict the time to radio link failure (RLF), but cannot determine the next best BS for handover. A NW-side adds sub-models (which may be developed either by the network vendor and / or UE vendor) that are tailored for the BS topology and conditions of each specific area and are able to predict subsequent measurements (e.g., RSRP, RSRQ, SINR) for the available (for HO) BSs or provide the configuration for conditional handover (CHO). The sub-models may be executed (i.e. inference operation) at the network side and / or at the UE side. For example, the network may indicate an identifier for a model applicable to the area (for example, using a RRC message, such as RRC_Setup or RRC_Reconfiguration message) or using system information (SI) messages. A UE may compare the identifier(s) and / or validity condition(s) of the model(s) it currently has available for inference, with the identifier of the model indicated by the network and / or the validity information, and the UE may initiate an update operation either with the RAN network or towards the OTT server or perform model switching.

[0173] Now, training of sub-models and proprietary properties according to embodiments is described.

[0174] Sub-models and side-branches can be trained end-to-end and all entities (UE, NW, gNB, etc.) have knowledge of the entire early exit model.

[0175] Some sub-models can be trained at the UE-side and others at the gNB / NW, without shared knowledge on proprietary sensor measurements, implementation details or assistance information.

[0176] Examples of private information at UE-side comprise: sensor information: UE speed, orientation, camera, etc.; and / or antenna blocking patterns; and / or a position estimate using side-link or landmarks.

[0177] Examples of private information at NW-side comprise beam shape / beamforming matrix.

[0178] FH241106PEP-2025373084.DOCX For additional flexibility, some sub-models [e.g., sub-models 1 to ml] are trained at the UE-side (without sharing implementation information) with the NW, others [e.g., sub¬ models m2 to the final sub-model of the backbone neural network] are trained at the NW (without sharing implementation information) and a number of sub-models [e.g., sub¬ models ml to m2 in the “middle”] are trained to be shared with UE and NW. The last ones can be deployed at either entity, depending on the given constraints and requirements.

[0179] In either case, sub-model outputs at branching points that define the data exchange between the proprietary sub-models, can be trained to be as small as possible, to reduce signaling burden.

[0180] Complementary, sub-model outputs at branching point that define the data exchange between the proprietary sub-models, are trained to provide contextual information [e.g., early clusters for positioning]

[0181] Now, a placement of early exits according to embodiments is described.

[0182] Early exits (outputs from branching points and side-branches) can be placed according to resource constraints and energy consumption requirements, for different UE[ / gNB] capabilities.

[0183] Alternatively, early exits are placed according to the increased performance (or reduction of uncertainty) they achieve, compared to previous early exits.

[0184] Adapting the early exit concept to 3GPP requirements / discussions, early exits can be placed according to their ability to provide outputs that address different functionalities [according to TR38.843],

[0185] A first example relates to positioning: For positioning, early exits to the left can provide predictions on assisted positioning required outputs (e.g., AoA, ToA, LOS / NLOS classification, etc.), intermediate exits can provide coarse (e.g., in a grid) position estimate, while last exits (and final output) can provide direct position estimate with different levels of accuracy.

[0186] It should be noted that this could be NW executing the earlier parts of the model and UE continuing the execution of sub-models for better position estimation, possibly including its local sensor measurements as inputs to these sub-models.

[0187] FH241106PEP-2025373084.DOCX Taking an example of UE-based positioning, the UE may be provided by a sub-model which is configured to estimate direct position of a UE. As discussed above, there may be one or more early exits, wherein each sub-model the information (e.g. UE position) improves. The branch point output may be exposed by the UE to higher layers (e.g. applications) or generic sub-models that can be loaded to the UE (e.g. a Side-branch) or to a network entity.

[0188] Fig. 3 illustrates different configurations of a backbone neural network comprising of two sub-models (sub-model 1, and sub-model 2) according to an embodiment. The first early exit point (e.g., Early Exit 1) may provide a coarse UE position (Output 1). After processing the sub-model 2, the final exit may provide a finer UE location. There may be possibility of having one or more side-branches between sub-model 1 and sub-model 2, where the information at the branch point may be fed to either a ML model that is fine¬ tuned to the environment or the information at branch point (e.g. NN weights) are exposed to application layer. The exposed information can be combined with ground truth labels obtained from an independent source (e.g. different from the RAT dependent positioning method). Assuming that the lifecycle management of both sub-models (1 and 2) and early exit 1 in this example are maintained by a first vendor (e.g. the UE vendor). The LCM of the early exit 1b may be done by the same or a different vendor (e.g. the operator or a service provider). By exposing the information at the branching point, the second vendor may utlise additional information (e.g. ground truth obtained by alternative source, like camera) to improve the ML model of Early Exit 1b. This enables the early exit 1b to be better adapted to the environment or utilize additional information available to improve the position estimate.

[0189] In case of UE-based positioning, the ML model Early Exit 1b, may be updated at the UE from the first vendor through in-device fine-tuning, over the OTT server or from the LMF. Likewise, the ML model Early Exit 1b, may form part of an application layer, where an application running in an UE may obtain / read information from the branching point and determine the UE location at the application layer with the model trained for Early Exit 1b. Alternatively, the UE may be configured to provide the branching point output to one or more network entities to perform inference.

[0190] Likewise, in case of NW-based methods (e.g. based on measurements of the NW entities), the network entities may process one or more sub-models. The output of the branching point may be exposed to external clients and / or an application at the UE. The

[0191] FH241106PEP-2025373084.DOCX UE may be able to process it further with additional information available at the UE side to obtain independent output.

[0192] The Early-Exit 1b (in above three figures) could be seen as adaptation layer, whose lifecycle may be managed separate to the existing sub-models. For example, the Early- Exit 1b may be adapted to perform precise positioning in a certain warehouse setting, whereas the remainder of the backbone neural network may be generic in nature (e.g. provided by device vendor(s)). For the ML model pertaining to Early Exit 1b, labels may additionally be collected specific to this environment using high precision external sources (e.g. using camera-based detection such as QR codes, landmarks, features, pattern recognition, high-precision reference systems etc). The Early-exit 1b may be trained by a different vendor (e.g. a service provider) and serves as an adaptation to the generic model which may be trained by the device vendor (e.g. UE-vendor, NW-vendor).

[0193] The output of branching point may be neural network weights, or they may be represent certain information. By exposing the output of the branching point to application layer and / or a second entity (e.g. a UE or a network entity), the output may be logged or associated with other labels.

[0194] A UE may be able to indicate its capability to provide optimized models from third party to a second entity. The capability may be limited to certain areas.

[0195] Another example relates to beam management: For beam management, early exits to the left can predict top-K (beams with highest RSRP) of wide beams in a hierarchical codebook, intermediate exits can predict top-K (beams with highest RSRP) of narrow beams in a hierarchical codebook, while last exits (and final output) can predict the actual RSRP values for these top-K narrow beams (e.g., to ensure that the selected beams satisfy a throughput constraint)

[0196] A further example relates to mobility: For mobility, early exits to the left can predict the time to radio link failure (e.g. with the serving cell) with increasing accuracy, intermediate exits can predict measurements (e.g., RSRP, RSRQ, SINR) for the available (for HO) BSs (e.g. target cells or one or more frequency-layers) and last exits (and final output) can provide the configuration for conditional handover (CHO) or select the HO BS directly.

[0197] In other words, the early exits from sub-models at the left may predict measurements in simple scenarios (like intra-frequency measurements) or inter-frequency measurement (e.g. in the same band), whereas the intermediate layers may be able to predict

[0198] FH241106PEP-2025373084.DOCX parameters in an increasingly complex scenario (like inter-frequency measurements in a different frequency range) (e.g. serving cell is in FR1 and target cell is in FR2) and / or different quasi-colocation conditions (e.g. target cell located in a different site and different frequency layer compared to source cells and / or cells where measurement is performed). In some examples, the intermediate results of processing one or more sub¬ models is further reported to the network, wherein the network entity processes further to determine suitable handover cells (e.g. the sub-model at the network side may take into account the expected network load across cells and / or across frequency layers and / or across different RAT technologies, e.g. LTE).

[0199] For example, an early or intermediate sub-model may predict one or more mobility parameters (e.g. RSRP, RSRQ, SINR or RLF) in the same frequency layer (e.g. same carrier frequency and / or bandwidth) and a subsequent sub-model may predict mobility parameters in a different frequency layer (e.g. frequency range different to the first frequency layer) and / or with different quasi-colocation conditions (e.g. the cells of the serving cell and / or target cells).

[0200] In another example, an early or intermediate sub-model may support a joint source and channel coding (JSCC) task, combined with a classical (non AI / ML) quantization and modulation function and a subsequent sub-model may encapsulate the joint quantization and modulation functions. This way, the entire model supports a joint source, channel coding and modulation (JSCCM) task. A “mirroring” setting is envisioned for the NW-sided AI / ML model.

[0201] Fig. 6 illustrates a modular joint source, channel coding and modulation (JSCCM) implementation according to an embodiment. In particular, Fig. 6 shows an end-to-end communication pipeline for Channel State Information (CSI) feedback that supports two functionalities: an AI / ML Joint Source-Channel Coding with Modulation (JSCCM) path, and an AI / ML CSI compression and classical quantization / modulation path. The top half is the UE-side (encoder / transmitter) and the bottom half is the NW-side (decoder / receiver).

[0202] At the UE side:

[0203] Depending on the activated functionality, the upper or lower pipeline is fed with the CSI input. In the upper path, the AI / ML JSCC sub-model (encoder) maps the CSI into a channel-ready latent representation and the AI / ML quantization and modulation sub¬ model turns the latent into discrete symbols and modulates them. The lower, hybrid classical path, consists of an AI / ML CSI compression encoder (as defined in 3GPP Rel.

[0204] FH241106PEP-2025373084.DOCX 19 and Rel. 20), followed by classical quantization and modulation blocks. The side¬ branch representing the early exit EE1 enables the utilization of the classical quantization and modulation blocks after the AI / ML JSCC sub-model.

[0205] The process is “mirrored” at the NW side for the de-quantization and decoding tasks.

[0206] Based on an early exit output, the UE may choose to report the output according to the upper branch or the lower branch. In other words, the UE may choose to utilize classical (legacy) quantization and modulation approaches or AI / ML based quantization and modulation. The UE may indicate this information to the network entity by adding additional information, which is associated with the information transmitted by the UE, so that the receiving side knows whether the UE utilized the classical (legacy) quantisation and modulation approach and the receiver can then utilize either the upper chain or the lower chain in the receiver to obtain the reconstructed CSI.

[0207] The information included by the UE may be a bit, a series of bits, flag, enumeration or similar. It may be carried in UCI report, MAC-CE or even in RRC messages (e.g. UAI information indicating whether the first branch or the second branch is applicable - in other words, if the feature is implemented with classical functionality or with AI / ML functionality.

[0208] The indication of the classical functionality or AI / ML functionality used by the UE may be valid for a subframe, a radio frame on which the indication is available or associated with (e.g. valid for x subframes after receiving). Alternatively, it may be triggered by MAC-CE and may be valid until it is cancelled or reconfigured or a timer runs out. Finally, the feature may be applicable after x time units after reporting unless explicitly cancelled or RLF occurs or a timer expires.

[0209] Now, inter-vendor training collaboration according to embodiments is described.

[0210] Training models between vendors and / or training different (proprietary) parts of a model from different vendors in different sides, have already been discussed in 3GPP within the CSI-compression study item in Release 19. This means that training early exit models according to embodiments in a setting with different vendors involved is feasible, by adapting any of these (high-level) approaches.

[0211] FH241106PEP-2025373084.DOCX According to embodiments, the training dataset should additionally contain the output of the early exits and the output of sub-models that is used for continuing computation on the other side.

[0212] A variation that is not captured in the options below is the case where a part of the sub¬ models (with earlier exits) is trained by vendor A in a proprietary manner, another part of the sub-models (with later exits and final model output) is trained by vendor B, and several sub-models (e.g., “in the middle” but not limited to this) are open to both vendors and can be executed in either / any side. In this case, the sub-model outputs and early exit results have to be the same, but the sub-models implementation can differ, depending on the hardware platform.

[0213] The output of the early exits and the output of sub-models that is used for continuing computation on the other side, but how the sub-models are implemented is different for the UE and a GPU-enabled gNB or the NW.

[0214] To alleviate / resolve the issues related to inter-vendor training collaboration of AI / ML-based CSI compression using two-sided model, the following options may, e.g., be employed:

[0215] Option 1 comprises a fully standardized reference model (structure and parameters).

[0216] Option 2 comprises a standardized dataset.

[0217] Option 3 comprises a standardized reference model structure and / or a parameter exchange between a NW-side and a UE-side.

[0218] Option 4 comprises a standardized data / dataset format and / or Dataset exchange between NW-side and UE-side.

[0219] Option 5 comprises a standardized model format and / or a reference model exchange between NW-side and UE-side.

[0220] For CSI compression using two-sided model use case, considered AI / ML model training collaborations may, e.g., comprise:

[0221] Type 1 exhibits a joint training of the two-sided model at a single side / entity, e.g., UE- sided or Network-sided.

[0222] FH241106PEP-2025373084.DOCX Type 2 exhibits a joint training of the two-sided model at network side and UE side, respectively.

[0223] Type 3 exhibits a separate training at network side and UE side, where the UE-side CSI generation part and the NW-side CSI reconstruction part are trained by UE side and network side, respectively.

[0224] It should be noted that joint training means the generation model and reconstruction model should be trained in the same loop for forward propagation and backward propagation. Joint training could be done both at single node or across multiple nodes (e.g., through gradient exchange between nodes).

[0225] Moreover, it should be noted that separate training may, e.g., include sequential training starting with UE side training, or sequential training starting with NW side training

[0226] Furthermore, it should be noted that training collaboration Type 2 over the air interface for model training (not including model update) is concluded to be deprioritized in Rel-18 SI.

[0227] For Type 2 (Joint training of the two-sided model at network side and UE side, respectively), note that joint training includes both simultaneous training and sequential training, in which the pros and cons could be discussed separately. Further, note that Type 2 sequential training starts with NW side training.

[0228] In the following, inference according to embodiments is considered.

[0229] Early-exits enabled UE-sided, NW-sided or cooperative inference in more than one entity provides the flexibility to achieve the required / requested level of performance on a task (e.g., positioning or beam management), while accounting for several constraints, for example:

[0230] latency (for example, can treat inherently “no delay,” “low delay” and “delay tolerant” positioning requirements set in TS22.071)

[0231] UE / gNB / NW hardware / battery / capability / load / etc.

[0232] Energy savings requirements (e.g., prefer earlier exits to save UE battery / energy)

[0233] FH241106PEP-2025373084.DOCX Signaling overhead (e.g., can execute more sub-models to transmit less information).

[0234] Depending on how the early exit model was trained, the initial sub-models might be executed in the UE (or in the NW). In both cases, a new type of functionality [cooperative inference functionality] needs to be defined, in which:

[0235] The UE (or NW) may, e.g., configured to transmit either a result (from an early exit) or an intermediate result (from an early exit) and the respective sub-model outputs to continue execution at the NW (or UE).

[0236] Alternatively, the UE (or NW) is configured to transmit intermediate results and the respective sub-model outputs, that provide outputs that correspond to different functionalities. The other side (NW or UE) can: i) use the results as-is; ii) continue execution of the early exit model; iii) utilize the plurality of information on a different algorithm.

[0237] According to a first example of such an embodiment, in positioning, the UE may, e.g., transmit to the NW LOS / NLOS information (extracted from earlier sub-models) and estimate on x,y,z position (extracted from later sub-models). The NW / LMF can utilize this information to achieve a better positioning accuracy.

[0238] According to a second example of such an embodiment, in beam management, the UE transmits to the NW top-K wide beam prediction (extracted from earlier sub-models) and top-M narrow beam prediction (extracted from later sub-models). The NW can utilize this information to predict which beams would most likely satisfy a throughput criterion.

[0239] Now, inference in the same or different entities according to embodiments is described.

[0240] In the most basic configuration, the UE notifies the NW on the availability of the early exit model and pre-loads all (or a maximum possible number of) sub-models. It performs inference up to the early exit that satisfies the performance constraint.

[0241] The UE may, e.g., inform the NW on its capability (on how many sub-models it can execute due to hardware / latency / battery / etc. constraints). The NW configures the UE to load a specific number of sub-models for inference, based on [worst- or average-case] performance requirements. These sub-models and the associated early exits can provide

[0242] FH241106PEP-2025373084.DOCX outputs to the same or different functionalities, depending on the model properties and the configuration from the NW.

[0243] Complementary, the UE may, e.g., inform the NW on its capability, and the NW configures how many sub-models the UE will load (and which outputs of branching points or early exit outputs the UE will provide). The NW loads the remaining sub-models for cooperative inference [NW knows environment properties and manages UE-side model]. The UE-side sub-models and the associated early exits can provide outputs to the same or different functionalities, depending on the model properties and the configuration from the NW.

[0244] The NW can configure this in a message containing at least one of the following:

[0245] • The model ID (of the full model);

[0246] • The number of sub-models to be loaded;

[0247] • The ID or index of an early exit;

[0248] • A flag indicating if only the output of the early exit or also the output of a sub-model should be provided;

[0249] • A combination of the above.

[0250] A first example according to an embodiment relates to positioning: For positioning, early exits to the left can provide predictions on assisted positioning required outputs (e.g., AoA, ToA, LOS / NLOS classification, etc.), intermediate exits can provide coarse (e.g., in a grid) position estimate, while last exits (and final output) can provide direct position estimate with different levels of accuracy. It should be noted that this could be the NW executing the earlier parts of the model and UE continuing the execution of sub-models for better position estimation, possibly including its local sensor measurements as inputs to these sub-models.

[0251] A second example according to another embodiment relates to beam management. For beam management, early exits to the left can predict top-K (beams with highest RSRP) of wide beams in a hierarchical codebook, intermediate exits can predict top-K (beams with highest RSRP) of narrow beams in a hierarchical codebook, while last exits (and final output) can predict the actual RSRP values for these top-K narrow beams (e.g., to ensure that the selected beams satisfy a throughput constraint).

[0252] A third example relates to mobility: For mobility, early exits to the left can predict the time to radio link failure (e.g. with the serving cell) with increasing accuracy, intermediate exits can predict measurements (e.g., RSRP, RSRQ, SINR) for the available (for HO) BSs (e.g.

[0253] FH241106PEP-2025373084.DOCX target cells or one or more frequency-layers) and last exits (and final output) can provide the configuration for conditional handover (CHO) or select the HO BS directly.

[0254] Alternatively, the UE informs the NW on the necessity to load additional sub-models for cooperative inference, due to environment / input properties [NW does not know the environment properties and the UE triggers cooperative inference]. In this case, the UE can determine the number of sub-models, e.g., by:

[0255] Predicting the required number of sub-models, due to experience of inference in similar environments (e.g., encapsulated in the training data). Confidence of prediction of early exits is low.

[0256] Monitoring indicates that performance (e.g., QoS) is lower than expected / required.

[0257] Regarding (privileged / proprietary) assistance information to sub-models, same information as in training may, e.g., be employed. Examples of private information at the UE-side may, e.g., comprise sensor information: UE speed, orientation, camera, etc.; and / or antenna blocking patterns; and / or a position estimate using side-link or landmarks. Examples of private information at the NW-side may, e.g., comprise a beam shape and / or a beamforming matrix.

[0258] Regarding model monitoring and management, here are two aspects: A first aspect relates to how does the UE / NW learn which sub-models to pre-load in each case (model management). A second aspect relates to how actual model performance monitoring is realized.

[0259] Regarding learning which sub-models to pre-load / configure, according to an embodiment, the UE may, e.g., perform inference without any information of environment properties from the NW. It reports to the NW on the number of sub-models required in each session to achieve the required level of performance. In case performance was not achieved, it reports to the NW information from the monitoring analytic that detected the inadequate performance level (e.g., reports on accuracy gap, possible uncertainty of early exits, etc.). This way, the NW progressively develops an estimator on which sub-models need to be activated in each case.

[0260] In a different approach, initially the NW runs the early exit model locally. Once it learns an estimator on which sub-models need to be activated in each case, it enables the cooperative inference functionality.

[0261] FH241106PEP-2025373084.DOCX Regarding learning a seamless selection / switching of functionalities, in such a complementary approach, the UE is configured to transmit intermediate results and the respective sub-model outputs, that provide outputs that correspond to different functionalities. This enables the implementation of a low-latency functionality management decision, on which functionalities to activate or switch between.

[0262] According to some embodiments, cooperative / progressive monitoring may, e.g., be employed.

[0263] The concept of early exit networks allows for a form of cooperative / progressive monitoring approach, not possible with other AI / ML model types.

[0264] Progressive monitoring may, e.g., comprise to run at the UE / NW and provide output of early exit due to latency requirements, and to execute more sub-models with some delay and evaluate the accuracy of the early exit output.

[0265] Cooperative / progressive monitoring may, e.g., comprise to run at the UE (NW) and provide output of early exit due to latency requirements, and to execute more sub-models at the NW (UE) and evaluate the accuracy of the early exit output.

[0266] According to some embodiments:

[0267] In a first aspect, inputs from different entities (that are not part of the model input in the left), can be passed at sub-models. (This is illustrated as “additional information for sub¬ model K-1” in Fig. 4). This way, each “part” (collection of sub-models that are executed in the UE, gNB or NW) can be private from the other side and use privileged information / inputs.

[0268] In a second aspect, the backbone NN can be trained in a way that maintains the proprietary requirements of the sub-models running at the UE and the sub-models running at the NW (so these are trained separately).

[0269] In a third aspect, a part of the neural network model (“middle” part) can be common, so it can run anywhere. This allows flexibility of execution, even when the other parts of the model are proprietary.

[0270] FH241106PEP-2025373084.DOCX In a fourth aspect, different early exits can correspond to different functionalities, including the examples there. This affects both training and inference (e.g., what the UE is configured to transmit).

[0271] Implications on the various AI / ML model / functionality life cycle management (TR38.843) steps may, e.g., comprise one or more of the following:

[0272] Inference:

[0273] ■ Cooperative inference functionality

[0274] ■ can flexibly consider the following constraints:

[0275] • Energy savings requirements

[0276] • Signaling overhead (e.g., can execute more sub-models to transmit less information).

[0277] Cooperative / progressive monitoring:

[0278] ■ parallel inference and monitoring with small overhead:

[0279] • run at the UE -> provide output of early exit due to latency requirements -> execute more sub-models at the NW and evaluate the accuracy of the early exit output that runs at the UE.

[0280] T raining / re-training / fine-tuning:

[0281] ■ fine-tuning by adding sub-model(s):

[0282] • Initial sub-models trained to the NW -> extended / fine-tuned (and compiled, pruned, etc.) for each UE separately by UE vendors

[0283] • Initial sub-models trained through UE vendor collaboration -> extended / fine-tuned in the NW (e.g., for adapting to specific areas).

[0284] Although some aspects of the described concept have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or a device corresponds to a method step or a feature of a method step. Analogously, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus.

[0285] Various elements and features of the present invention may be implemented in hardware using analog and / or digital circuits, in software, through the execution of instructions by one or more general purpose or special-purpose processors, or as a combination of hardware and software. For example, embodiments of the present invention may be

[0286] FH241106PEP-2025373084.DOCX implemented in the environment of a computer system or another processing system. Fig.

[0287] 5 illustrates an example of a computer system 600. The units or modules as well as the steps of the methods performed by these units may execute on one or more computer systems 600. The computer system 600 includes one or more processors 602, like a special purpose or a general-purpose digital signal processor. The processor 602 is connected to a communication infrastructure 604, like a bus or a network. The computer system 600 includes a main memory 606, e.g., a random-access memory, RAM, and a secondary memory 608, e.g., a hard disk drive and / or a removable storage drive. The secondary memory 608 may allow computer programs or other instructions to be loaded into the computer system 600. The computer system 600 may further include a communications interface 610 to allow software and data to be transferred between computer system 600 and external devices. The communication may be in the from electronic, electromagnetic, optical, or other signals capable of being handled by a communications interface. The communication may use a wire or a cable, fiber optics, a phone line, a cellular phone link, an RF link and other communications channels 612.

[0288] The terms “computer program medium” and “computer readable medium” are used to generally refer to tangible storage media such as removable storage units or a hard disk installed in a hard disk drive. These computer program products are means for providing software to the computer system 600. The computer programs, also referred to as computer control logic, are stored in main memory 606 and / or secondary memory 608. Computer programs may also be received via the communications interface 610. The computer program, when executed, enables the computer system 600 to implement the present invention. In particular, the computer program, when executed, enables processor 602 to implement the processes of the present invention, such as any of the methods described herein. Accordingly, such a computer program may represent a controller of the computer system 600. Where the disclosure is implemented using software, the software may be stored in a computer program product and loaded into computer system 600 using a removable storage drive, an interface, like communications interface 610.

[0289] The implementation in hardware or in software may be performed using a digital storage medium, for example cloud storage, a floppy disk, a DVD, a Blue-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate or are capable of cooperating with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium may be computer readable.

[0290] FH241106PEP-2025373084.DOCX A storage medium for a machine learning model can hold the final trained model and / or critical components required for the model execution. This includes the model architecture, which defines the structure of the model, such as the number and type of layers and how these layers are connected. The storage medium can contain the weights, which are the learned parameters from the training process. These weights determine how the model processes input data to generate predictions. Additionally, hyperparameters like learning rate, and number of epochs can be stored to ensure the model can be retrained or fine-tuned under the same conditions. Additionally, input and output formats are included as well as preprocessing and postprocessing instructions to standardize data before and after it passes through the model.

[0291] Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.

[0292] Generally, embodiments of the present invention may be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer. The program code may for example be stored on a machine readable carrier.

[0293] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier. In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0294] A further embodiment of the inventive methods is, therefore, a data carrier or a digital storage medium, or a computer-readable medium comprising, recorded thereon, the computer program for performing one of the methods described herein. A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals may for example be configured to be transferred via a data communication connection, for example via the Internet. A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein. A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0295] FH241106PEP-2025373084.DOCX In some embodiments, a programmable logic device, for example a field programmable gate array, may be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware apparatus.

[0296] The above described embodiments are merely illustrative for the principles of the present invention. It is understood that modifications and variations of the arrangements and the details described herein are apparent to others skilled in the art. It is the intent, therefore, to be limited only by the scope of the impending patent claims and not by the specific details presented by way of description and explanation of the embodiments herein.

[0297] FH241106PEP-2025373084.DOCX ABBREVIATIONS

[0298]

[0299] FH241106PEP-2025373084.DOCX

[0300]

[0301] FH241106PEP-2025373084.DOCX

[0302]

[0303] FH241106PEP-2025373084.DOCX

Claims

Claims1. An apparatus of a wireless communication system, wherein a model (215) comprises a plurality of sub-models, wherein the model (215) is an artificial intelligence / machine learning model,wherein the apparatus is configured for executing the model (215) and / or one or more sub-models of the plurality of sub-models of the model (215) to produce an output of the model (215) depending on an input of the model (215).

2. An apparatus according to claim 1, wherein the model (215) and / or the one or more sub-models are stored in the apparatus.

3. An apparatus according to claim 1 or 2,wherein the apparatus comprises a storage (210) having stored thereon the model (215) comprising the plurality of sub-models and / or having stored thereon the one or more sub-models of the plurality of sub-models of the model (215),wherein the apparatus comprises a processing unit (220) configured for executing the model (215) and / or the one or more sub-models of the model (215) to produce the output of the model (215) depending on the input of the model (215).

4. An apparatus of a wireless communication system, wherein the apparatus comprises:a storage (210) having stored thereon a model (215) comprising a plurality of sub¬ models and / or having stored thereon one or more sub-models of the plurality of sub-models of the model (215), wherein the model (215) is an artificial intelligence / machine learning model, anda processing unit (220) configured for executing the model (215) and / or the one or more sub-models of the model (215) to produce an output of the model (215) depending on an input of the model (215).

5. An apparatus according to one of the preceding claims,FH241106PEP-2025373084.DOCXwherein the apparatus is configured to conduct a measurement and / or to report on a measurement to produce the intermediate output associated with the functionality of the wireless communication system.

6. An apparatus according to one of the preceding claims,wherein the apparatus is configured to conduct a measurement and / or to report a measurement with respect to a signal transmitted within the wireless communication system, and / orwherein the apparatus is configured to conduct a measurement and / or to report a measurement with respect to a beam used for transmitting and / or receiving data within the wireless communication system, and / orwherein the apparatus is configured to conduct a measurement and / or to report a measurement with respect to a position of the apparatus or of another device of the wireless communication system, and / orwherein the apparatus is configured to conduct a measurement or a set of measurements and / or to utilise a measurement to predict / estimate a second measurement or event.

7. An apparatus according to claim 6,wherein the apparatus is configured to performed or estimate or predict the measurement to select a configuration from a set of configurations provided by the network entity, wherein the configuration indicates at least one cell (e.g., a primary cell) for the apparatus.

8. An apparatus according to claim 6 or 7,wherein the apparatus is configured to log the measurement performed or estimated or predicted by the apparatus, wherein in response to request from the network or in response to the logged measurements or information exceeding a certain size, the apparatus is configured to transfer the recorded data to at least one network entity.

9. An apparatus according to one of claims 6 to 8,FH241106PEP-2025373084.DOCXwherein the apparatus is configured to conduct a prediction and / or a measurement to determine if one cell is stronger than another cell in future time; and / orwherein the apparatus is configured to conduct a prediction and / or a measurement of a radio link failure within a certain time window; and / orwherein the apparatus is configured to conduct a detection of a radio link failure.

10. An apparatus according to one of the preceding claims,wherein the apparatus is a user equipment.

11. An apparatus according to claim 10,wherein the model (215) and / or the one or more sub-models of the plurality of sub¬ models of the model (215) are trained at the user equipment individually for the user equipment.

12. An apparatus according to claim 10 or 11 ,wherein a first sub-model of the plurality of sub-models of the model (215) is stored in the user equipment, andwherein a second sub-model of the plurality of sub-models of the model (215) is stored in a storage of a network entity of the wireless communication system.

13. An apparatus according to claim 10, further depending on claim 3 or 4,wherein the model (215) and / or the one or more sub-models of the plurality of sub¬ models of the model (215) being stored in the storage (210) of the user equipment are trained at the user equipment individually for the user equipment.

14. An apparatus according to claim 10 or 11, further depending on claim 3 or 4,wherein a first sub-model of the plurality of sub-models of the model (215) is stored in the storage (210) of the user equipment, andFH241106PEP-2025373084.DOCXwherein a second sub-model of the plurality of sub-models of the model (215) is stored in a storage of a network entity of the wireless communication system.

15. An apparatus according to claim 10,wherein the second sub-model is trained at the network entity.

16. An apparatus according to claim 10 or 11 ,wherein the model (215) comprises a first group of one or more sub-models being stored in the user equipment,wherein the model (215) comprises a second group of one or more sub-models being stored in a storage of the network entity.

17. An apparatus according to claim 10 or 13, further depending on claim 3 or 4,wherein the model (215) comprises a first group of one or more sub-models being stored in the storage (210) of the user equipment,wherein the model (215) comprises a second group of one or more sub-models being stored in a storage of the network entity.

18. An apparatus according to claim 16 or 17,wherein the first group of one or more sub-models is trained at the user equipment without sharing implementation information with the network entity,wherein the second group of one or more sub-models is trained at the network entity without sharing implementation information with the user equipment.

19. An apparatus according to one of claims 16 to 18,wherein the model (215) comprises a third group of one or more sub-models being shared and / or being trained to be shared with the user equipment and with the network entity.FH241106PEP-2025373084.DOCX20. An apparatus according to one of claims 10 to 19,wherein the user equipment is configured to inform the network entity on one or more of its capabilities (e.g., on how many sub-models it can execute due to hardware / latency / battery / etc. constraints).

21. An apparatus according to claim 20,wherein the user equipment is adapted to be configured by the network entity to load a specific number of sub-models for inference depending on performance requirements (e.g., worst-case or average-case performance requirements), wherein said sub-models and associated early exits are configured to provide outputs for same or different functionalities (e.g., depending on the model properties and the configuration from the network).

22. An apparatus according to claim 21,wherein, to be configured by the network entity to load a specific number of sub¬ models, the user equipment is configured to receive from the network entity one or more of the following:a model identifier identifying the model (215),a number of sub-models to be loaded,an identifier or an index of an early exit,a flag indicating if only the output of the early exit or also an output of a sub-model shall be provided.

23. An apparatus according to one of claims 1 to 9,wherein the apparatus is a network entity.

24. An apparatus according to one of the preceding claims,wherein the model (215) is a first model,FH241106PEP-2025373084.DOCXwherein the apparatus comprises an input interface for receiving irreversible processed information from another apparatus of the wireless communication system, wherein the irreversible processed information is an output of another model or is a processed output derived from the output of the other model, wherein the other model is another artificial intelligence / machine learning model of said other apparatus or of a further apparatus of the wireless communication system, wherein the output of the other model results from feeding confidential input into the other model, wherein the confidential input is not reconstructable from the irreversible processed information,wherein the input interface is configured to feed the irreversible processed information into the first model (215) or into at least one of the one or more sub¬ models of the first model (215).

25. An apparatus according to one of claims 1 to 23,wherein the model (215) comprises at least two sub-models,wherein a first sub-model and a second sub-model of the at least two sub-models are arranged such that a sub-model output of the first sub-model is fed into a second sub-model as a first part of an input of the second sub-model,wherein the apparatus comprises an input interface being configured to receive further input information from another apparatus of the wireless communication system, and being configured to feed the further input information as a second part of the input of the second sub-model into the second sub-model.

26. An apparatus according to claim 25,wherein the input interface is configured for receiving irreversible processed information as the further input information from said other apparatus, wherein the irreversible processed information is an output of another model or is a processed output derived from the output of the other model, wherein the other model is another artificial intelligence / machine learning model of said other apparatus or of a further apparatus of the wireless communication system, wherein the output of the other model results from feeding confidential input into the other model, wherein the confidential input is not reconstructable from the irreversible processed information,FH241106PEP-2025373084.DOCXwherein the input interface is configured to feed the irreversible processed information as the second part of the input of the second sub-model into the second sub-model.

27. An apparatus according to one of the preceding claims,wherein the model (215) is a first model,wherein the apparatus comprises an input model interface for receiving an external output of a further model being a further artificial intelligence / machine learning model of a further device of the wireless communication system or for receiving a processed output being derived from the external output, andwherein the apparatus is configured to feed the external output or the processed output or data being derived from the output or from the processed output as input into the model (215) or into a sub-model of the plurality of sub-models of the model (215).

28. An apparatus according to claim 27,wherein the model (215) can be trained with input data being independent from the further model.

29. An apparatus according to claim 27 or 28,wherein the external output or the processed output is irreversible such that confidential input which has been fed into the further model is not reconstructable from the external output or from the processed output.

30. An apparatus according to one of the preceding claims,wherein the model (215) is a first model,wherein the apparatus is configured to provide a model output of the first model (215) to another device of the wireless communication system for providing said other device with input for a further model being a further artificial intelligence / machine learning model of said other device, wherein a model outputFH241106PEP-2025373084.DOCXof said further model is to be employed for conducting a functionality of the wireless communication system.

31. An apparatus according to claim 30,wherein the first model (215) can be trained with output data being independent from the further model.

32. An apparatus according to claim 30 or 31 ,wherein the model output of the first model (215) which is provided to the further model is irreversible such that model input that has been fed into the first model (215) to obtain the model output of the first model (215) is not reconstructable from the model output of the first model (215).

33. An apparatus according to one of claim 30 to 32,wherein the first apparatus comprises a battery for providing energy for the apparatus,wherein the apparatus is configured to decide whether or not to provide the model output of the first model (215) to said other device depending on a state of the battery.

34. An apparatus according to one of claims 30 to 33,wherein the apparatus is configured to decide whether or not to provide the model output of the first model (215) to said other device depending on a signalling overhead.

35. An apparatus according to one of claims 30 to 34,wherein the apparatus is configured to receive an external model output from the other device, wherein the external model output is a model output of the further model of the other device in response to inputting the model output of the first model (215) into the further model of the other device.

36. An apparatus according to claim 35,FH241106PEP-2025373084.DOCXwherein the apparatus is configured to conduct a functionality of the wireless communication system using the external model output.

37. An apparatus according to claim 35 or 36,wherein the apparatus is configured to control a quality or a reliability of the model output of the first model (215) using the external model output.

38. An apparatus according to one of claims 30 to 37,wherein the other device is a network component of the wireless communication system.

39. An apparatus according to one of claims 30 to 38,wherein the apparatus is a user equipment of the wireless communication system, and the other device is another user equipment of the wireless communication system.

40. An apparatus according to one of claims 30 to 38,wherein the apparatus is a network entity of the wireless communication system, and the other device is another user equipment of the wireless communication system.

41. An apparatus according to one of claims 26 to 40,wherein the first model (215) and / or the further model is available (e.g., is common) for a plurality of devices of the wireless communication system.

42. An apparatus according to claim 41,wherein the first model (215) is not available for any other apparatus of the wireless communication system, orwherein the further model is not available for the apparatus.FH241106PEP-2025373084.DOCX43. An apparatus according to one of the preceding claims,wherein the apparatus further comprises at least one exit point (216) of the model (215) configured to provide an intermediate output associated with a functionality of the wireless communication system.

44. An apparatus according to one of claims 1 to 42,wherein the apparatus is further configured to determine the capability of supporting at least one exit point (216) of the model (215) and provide an output related to this determined capability in response to a request.

45. An apparatus according to claim 43 or 44,wherein the model (215) comprises two or more exit points as the at least one exit point (216), wherein each of the two or more exit points is configured to provide an intermediate output associated with a functionality of the wireless communication system.

46. An apparatus according to claim 45,wherein the intermediate output of each of the two or more exit points is associated with a different functionality being different from the functionality which is associated with the intermediate output of any other exit point of the two or more exit points.

47. An apparatus according to one of claims 43 to 46,wherein the apparatus is configured to execute a first task using a model output of the model (215),wherein the apparatus is configured to execute a second task, being different from the first task, using the intermediate output being provided at the at least one exit point (216) of the model (215),wherein at least one of the first task and the second task comprises conducing a functionality of the wireless communication system.FH241106PEP-2025373084.DOCX48. An apparatus according to claim 47,wherein at least one of the first task and the second task is a task to transmit data and / or to receive data and / or to determine a position of the apparatus or of another apparatus of the wireless communication system, and / or to control a beam management, and / or to determine target primary cell and / or time for handover and / or a configuration of serving cells comprising primary and secondary cells to be applied and / or indicate to higher layers to initiate handover to different cell or to a cell on a different RAT, and / or to determine a configuration for conditional handover or to select the handover base station directly.

49. An apparatus according to claim 47 or 48,wherein the apparatus is configured to execute the first task using the model output without using the intermediate output being provided at the at least one exit point (216) of the model (215), and / orwherein the apparatus is configured to execute the second task using the intermediate output being provided at the at least one exit point (216) of the model (215) without using the model output.

50. An apparatus according to one of claims 47 to 49,wherein the intermediate output being provided at the at least one exit point (216) of the model (215) is an output of a first sub-model of the model (215),wherein the model output of the model (215) is an output of a second sub-model of the model (215),wherein the model (215) and / or the second sub-model is trained using model training output depending on the first task, andwherein the first sub-model is trained using model training output depending on the second task.

51. An apparatus according to one of claims 47 to 50,FH241106PEP-2025373084.DOCXwherein the intermediate output being provided at the at least one exit point (216) is an output of a first one of a plurality of layers of the model (215), andwherein the model output of the model (215) is an output of a second one of the plurality of layers of the model (215), being different from the first one of the plurality of layers.

52. An apparatus according to claim 51,wherein the apparatus has stored therein a first one of the plurality of sub-models of the model (215),wherein a further apparatus of the wireless communication system comprises a second one of the plurality of sub-models of the model (215).

53. An apparatus according to claim 51, further depending on claim 3 or 4,wherein the storage (210) of the apparatus has stored therein a first one of the plurality of sub-models of the model (215),wherein a further apparatus of the wireless communication system comprises a second one of the plurality of sub-models of the model (215).

54. An apparatus according to claim 52 or 53,wherein the apparatus is configured to provide an output of the first one of the plurality of sub-models of the model (215) to the further apparatus for being input into the second one of the plurality of sub-models of the model (215).

55. An apparatus according to claim 54,wherein the apparatus is configured to receive an output of the second one of the plurality of sub-models of the model (215) from the further apparatus,wherein the apparatus is configured to conduct a functionality of the wireless communication system using the output of the second one of the plurality of sub¬ models of the model (215).FH241106PEP-2025373084.DOCX56. An apparatus according to one of claims 43 to 55,wherein the apparatus is configured to determine a confidence in a prediction of the model output and / or a confidence in a prediction of the intermediate output being provided at the at least one exit point (216).

57. An apparatus according to claim 56,wherein the apparatus further comprises an activation control unit (230) configured to activate and / or control at least one of the plurality of sub-models of the model (215) depending on requirements and / or measurements within the wireless communication system,wherein the activation control unit (230) is configured to activate one or more subsequent layers following a layer where the model output is outputted depending on the confidence in a prediction of the model output and / or depending on the confidence in the prediction of the intermediate output being provided at the at least one exit point (216).

58. An apparatus according to one of claims 43 to 57,wherein the apparatus further comprises an activation control unit (230) configured to activate and / or control at least one of the plurality of sub-models of the model (215) depending on requirements and / or measurements within the wireless communication system.

59. An apparatus according to one of claims 43 to 58,wherein the apparatus is configured to employ the intermediate output being provided at the at least one exit point (216) for determining an AoA and / or a ToA and / or a LOS / NLOS classification.

60. An apparatus according to one of claims 43 to 59,wherein the apparatus is configured to employ the intermediate output being provided at the at least one exit point (216) for determining one or more beams with a highest RSRP among a plurality of beams.FH241106PEP-2025373084.DOCX61. An apparatus according to claim 60,wherein the apparatus is configured to employ the intermediate output being provided at the at least one exit point (216) to predict one or more beams with a highest RSRP of wide beams among the plurality of beams,wherein the apparatus is configured to employ the intermediate output being provided at the at least one exit point (216) to predict one or more beams with a highest RSRP of narrow beams among the plurality of beams, andwherein the apparatus is configured to employ the intermediate output being provided at the at least one exit point (216) is an output of a layer of the model (215) preceding a another layer of the model (215) which outputs the second model output.

62. An apparatus according to one of claims 43 to 61 ,wherein the plurality of sub-models of the model (215) comprise a first sub-model and a second sub-model,wherein the apparatus is configured to determine a coarse position of the apparatus depending on the intermediate output being provided at the at least one exit point (216),wherein the apparatus is configured to feed an output of the first sub-model as an input into the second sub-model to obtain a model output of the model (215) from an output of the second sub-model,wherein the apparatus is configured to determine a more precise position of the apparatus compared to the coarse position depending on the output of the second sub-model.

63. An apparatus according to claim 62,wherein the apparatus is configured to determine the more precise position of the apparatus depending on the output of the second sub-model and depending on another additional positioning method.FH241106PEP-2025373084.DOCX64. An apparatus according to claim 62 or 63,wherein a sub-model of the model (215) which provides the intermediate output at the at least one exit point (216) is updated by another entity using in-device fine- tuning.

65. An apparatus according to one of claims 62 to 64,wherein the apparatus is configured to provide the intermediate output of the at least one exit point (216) to one or more network entities to perform inference.

66. An apparatus according to one of claims 43 to 65,wherein the intermediate output being provided at the at least one exit point (216) comprise one or more weights of the model (215) being a neural network.

67. An apparatus according to one of claims 43 to 66,wherein the model (215) is trained depending on an output of the at least one exit point (216).

68. An apparatus according to claim 67,wherein the model (215) is trained further depending on an output of a sub-model that continues computation with data that is output at the least one exit point (216).

69. An apparatus according to one of claims 1 to 42,wherein the apparatus has stored therein only a first one of the plurality of sub¬ models of the model (215),wherein a further apparatus of the wireless communication system comprises a second one of the plurality of sub-models of the model (215),wherein the apparatus is configured to provide an output of the first one of the plurality of sub-models to the further apparatus for being input into the second one of the plurality of sub-models.FH241106PEP-2025373084.DOCX70. An apparatus according to one of claims 1 to 42, further depending on claim 3 or 4,wherein the storage (210) of the apparatus has stored therein only a first one of the plurality of sub-models of the model (215),wherein a further apparatus of the wireless communication system comprises a second one of the plurality of sub-models of the model (215),wherein the apparatus is configured to provide an output of the first one of the plurality of sub-models to the further apparatus for being input into the second one of the plurality of sub-models.

71. An apparatus according to one of claims 1 to 42,wherein the apparatus further comprises an activation control unit (230) configured to activate and / or control at least one of the plurality of sub-models of the model (215) depending on requirements and / or measurements within the wireless communication system.

72. An apparatus according to claim 57 or 58 or 71,wherein the activation control unit (230) is further configured to activate or de¬ activate at least one of the plurality of sub-models of the model (215) based on a configuration message comprising parameters received from a higher layer configuration of the wireless communication system.

73. An apparatus according to one of the preceding claims,wherein the model (215) is a neural network.

74. An apparatus according to claim 73,wherein the neural network is a CNN or is a LSTM or is a Bayesian NN or comprises transformer layers.

75. An apparatus according to one of the preceding claims,FH241106PEP-2025373084.DOCXwherein the model (215) and / or the one-or more sub-models of the plurality of sub¬ models of the model (215) have been trained using a loss function.

76. An apparatus according to claim 75,wherein the loss function is implemented as a MSE, or as a top-k accuracy, or as a F1 -score.

77. An apparatus according to one of the preceding claims,wherein the apparatus is configured to employ the intermediate output to predict a time to radio link failure (e.g. with the serving cell); and / orwherein the apparatus is configured to employ the intermediate output to predict measurements (e.g., RSRP, RSRQ, SINR) for the available base stations, wherein the apparatus is configured to employ the model output of the model (215) to determine a configuration for conditional handover or to select the handover base station directly.

78. An apparatus according to one of the preceding claims,wherein the apparatus is configured to indicate its capability to provide optimized models to another device of the wireless communication system.

79. An apparatus according to claim 78,wherein the capability is limited to certain areas.

80. An apparatus according to one of the preceding claims,wherein the apparatus comprises one or more processors,wherein the one or more processors comprise an Artificial Intelligence Processing Unit (APU).

81. An apparatus according to one of the preceding claims further depending on claim 10, wherein the apparatus, being the user equipment of claim 10, is configured to transmit information on one or more computational requirements of at least one ofFH241106PEP-2025373084.DOCXthe one or more sub-models, e.g., to a network entity of the wireless communication system.

82. An apparatus according to one of the preceding claims,wherein the apparatus is configured to implement a joint source and channel coding (JSCC) task.

83. An apparatus according to claim 82,wherein the apparatus is configured to employ the model and / or at least one sub¬ model of the one or more sub-models to implement the joint source and channel coding task.

84. An apparatus according to claim 83,wherein the model and / or the at least one sub-model implements a compression functionality.

85. An apparatus according to one of claims 82 to 84,wherein the apparatus is configured to employ a non AI / ML quantization function and / or a non AI / ML modulation function.

86. An apparatus according to claim 85,wherein a subsequent sub-model of the apparatus is configured to encapsulate the joint quantization and modulation functions.

87. An apparatus according to one of claims 82 to 86,wherein the joint source and channel coding task relates to source and channel coding of Channel State Information.

88. An apparatus according to one of claims 82 to 87,FH241106PEP-2025373084.DOCXwherein the apparatus is configured to activate and deactivate different functionalities of the joint source and channel coding task at different points in time.

89. An apparatus according to claim 88,wherein the different functionalities are located in different branches of a processing chain.

90. An apparatus according to claim 88 or 89,wherein one or more indicators indicates which of the different functionalities are activated and / or which of the different functionalities are deactivated.

91. An apparatus according to claim 90,wherein one or more indicators comprise at least one of a bit, a series of bits, a flag, and a enumeration .

92. An apparatus according to one of the preceding claims,wherein an UCI report or a MAC-CE or a RRC message comprises the one or more indicators.

93. An apparatus according to one of claims 90 to 92,wherein a functionality triggered for activation or deactivation by one or more indicators are valid for a subframe or are valid for a radio frame on which the indication is available or associated with.

94. An apparatus according to one of claims 90 to 92,wherein at least one of the one or more indicators is a MAC-CE and a functionality triggered by the MAC-CE is valid until it is cancelled or reconfigured or a timer runs out.

95. An apparatus according to one of claims 90 to 92,FH241106PEP-2025373084.DOCXwherein a functionality triggered for activation or deactivation by one or more indicators are valid for a certain time duration and / or for a number of time units.

96. A system comprising:an apparatus according to one of the preceding claims, andthe other apparatus of claim 24, wherein the apparatus implements an apparatus according to claim 24; or the other apparatus of claim 25, wherein the apparatus implements an apparatus according to claim 25; and / or,the further device of claim 27, wherein the apparatus implements an apparatus according to claim 27; and / orthe other device of claim 30, wherein the apparatus implements an apparatus according to claim 30; and / orthe further apparatus of claim 52 or 53, wherein the apparatus implements an apparatus according to claim 52 or 53; or the further apparatus of claim 69 or 70, wherein the apparatus implements an apparatus according to claim 69 or 70.

97. A method of a wireless communication system, wherein a model (215) comprises a plurality of sub-models, wherein the model (215) is an artificial intelligence / machine learning model, wherein the method comprises:executing the model (215) and / or one or more sub-models of the plurality of sub¬ models of the model (215) to produce an output of the model (215) depending on an input of the model (215).

98. A method of a wireless communication system, wherein an apparatus of the wireless communication system comprises a storage (210) having stored thereon a model (215) comprising a plurality of sub-models and / or having stored thereon one or more sub-models of the plurality of sub-models of the model (215), wherein the model (215) is an artificial intelligence / machine learning model, wherein the method comprises:FH241106PEP-2025373084.DOCXexecuting the model (215) and / or the one or more sub-models of the model (215), by a processing unit (220) of the apparatus, to produce an output of the model (215) depending on an input of the model (215).

99. A computer program for implementing the method of claim 97 or 98 when being executed on a computer or signal processor.FH241106PEP-2025373084.DOCX