Apparatus and method for validation, training, testing and monitoring for artificial intelligence / machine learning models and functionalities

WO2026202381A2PCT designated stage Publication Date: 2026-10-01FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2026/059023
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-28
Filing Date
2026-03-27
Publication Date
2026-10-01

Smart Images

  • Figure EP2026059023_01102026_PF_FP_ABST
    Figure EP2026059023_01102026_PF_FP_ABST
Patent Text Reader

Abstract

An apparatus for validating at least one AI / ML model or functionality according to an embodiment is provided. The apparatus comprises a validation module (120) configured to conduct an evaluation by evaluating a performance and / or an applicability of the at least one AI / ML model or functionality depending on classes and / or attributes and / or performance criteria. The validation module (120) is configured to generate a validation output indicating whether the at least one AI / ML model or functionality meet one or more predefined criteria depending on the evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Apparatus and Method for Validation, Training, Testing and Monitoring for Artificial Intelligence / Machine Learning Models and Functionalities

[0002] Description

[0003] The present invention relates to post deployment validation and testing in the field of artificial intelligence and machine learning, and to an apparatus and method for validation, training, testing and monitoring for artificial intelligence / machine learning models and functionalities.

[0004] Fig. 9A and Fig. 9B are schematic representations of an example of a terrestrial wireless network 100 including, as is shown in Fig. 9A, the core network and one or more radio access networks RAN1, RAN2, …RANN(RAN = Radio Access Network). Fig. 9B is a schematic representation of an example of a radio access network RANnthat may include one or more base stations gNB1to gNB5(gNB = next generation Node B), each serving a specific area surrounding the base station schematically represented by respective cells 1061to 1065. The base stations are provided to serve users within a cell. The one or more base stations may serve users in licensed and / or unlicensed bands. The term base station, BS, refers to a gNB in 5G or 6G networks, an eNB in UMTS / LTE / LTE-A / LTE-A Pro, or just a BS in other mobile communication standards. A user may be a stationary device or a mobile device. The wireless communication system may also be accessed by mobile or stationary loT (Internet of Things) devices which connect to a base station or to a user. The mobile devices or the loT devices may include physical devices, ground based vehicles, such as robots or cars, aerial vehicles, such as manned or unmanned aerial vehicles, UAVs, the latter also referred to as drones, buildings and other items or devices having embedded therein electronics, software, sensors, actuators, or the like as well as network connectivity that enables these devices to collect and exchange data across an existing network infrastructure. Fig. 9B shows an exemplary view of five cells, however, the RANnmay include more or less such cells, and RANnmay also include only one base station. Fig. 9B shows two users UE1 and UE2, (UE = User Equipment) also referred to as user equipment, UE, that are in cell 1062and that are served by base station gNB2. Another user UE3 is shown in cell 1064which is served by base station gNB4. The arrows 1081, 1082and 1083schematically represent uplink / downlink connections for transmitting data from a user UE1, UE2 and UE3 to the base stations gNB2, gNB4or for transmitting data from the base stations gNB2, gNB4to the users UE1, UE2, UE3. This may be realized on licensed bands or on unlicensed bands. Further, Fig. 9B shows two IoT devices 1101and 1102in cell 1064, which may be stationary or mobile devices. The IoT device 1101accesses the wireless communication system via the base

[0005] FH260314PCT-2026106843. DOCXstation gNB4 to receive and transmit data as schematically represented by arrow 1121. The IoT device 1102accesses the wireless communication system via the user UE3as is schematically represented by arrow 1122. The respective base stations gNB1to gNB5may be connected to the core network 102, e.g. via the S1 interface, via respective backhaul links 1141to 1145, which are schematically represented in Fig. 9B by the arrows pointing to “core”. The core network 102 may be connected to one or more external networks. The external network may be the Internet or a private network, such as an intranet or any other type of campus networks, e.g. a private WiFi or 4G or 5G or 6G mobile communication system. Further, some or all of the respective base stations gNB1to gNB5may be connected, e.g. via the S1 or X2 interface or the XN interface in NR (New Radio), with each other via respective backhaul links 1161to 1165, which are schematically represented in Fig. 9B by the arrows pointing to “gNBs”. A sidelink channel allows direct communication between UEs, also referred to as device-to-device, D2D (Device to Device), communication. The sidelink interface in 3GPP (3G Partnership Project) is named PC5 (Proximity-based Communication 5).

[0006] For data transmission a physical resource grid may be used. The physical resource grid may comprise a set of resource elements to which various physical channels and physical signals are mapped. For example, the physical channels may include the physical downlink, uplink and sidelink shared channels, PDSCH (Physical Downlink Shared CHannel), PUSCH (Physical Uplink Shared Channel), PSSCH (Physical Sidelink Shared Channel), carrying user specific data, also referred to as downlink, uplink and sidelink payload data, the physical broadcast channel, PBCH (Physical Broadcast Channel), carrying for example a master information block, MIB, and one or more of a system information block, SIB, one or more sidelink information blocks, SLIBs, if supported, the physical downlink, uplink and sidelink control channels, PDCCH (Physical Downlink Control Channel), PUCCH (Physical Uplink Control CHannel), PSCCH (Physical Sidelink Control Channel), the downlink control information, DCI, the uplink control information, UCI, and the sidelink control information, SCI, and physical sidelink feedback channels, PSFCH (Physical sidelink feedback channel), carrying PC5 feedback responses. Note, the sidelink interface may support a 2-stage SCI (Speech Call Items). This refers to a first control region comprising some parts of the SCI, and, optionally, a second control region, which comprises a second part of control information.

[0007] For the uplink, the physical channels may further include the physical random-access channel, PRACH (Packet Random Access Channel) or RACH (Random Access Channel), used by UEs for accessing the network once a UE synchronized and obtained the MIB and SIB. The physical signals may comprise reference signals or symbols, RS,

[0008] FH260314PCT-2026106843. DOCXsynchronization signals and the like. The resource grid may comprise a frame or radio frame having a certain duration in the time domain and having a given bandwidth in the frequency domain. The frame may have a certain number of subframes of a predefined length, e.g. 1ms. Each subframe may include one or more slots of 12 or 14 OFDM symbols (OFDM = Orthogonal Frequency-Division Multiplexing) depending on the cyclic prefix, CP, length. A frame may also include of a smaller number of OFDM symbols, e.g. when utilizing a shortened transmission time interval, sTTI (slot or subslot transmission time interval), or a mini-slot / non-slot-based frame structure comprising just a few OFDM symbols.

[0009] The wireless communication system may be any single-tone or multicarrier system using frequency-division multiplexing, like orthogonal frequency-division multiplexing, OFDM, or orthogonal frequency-division multiple access, OFDMA (Orthogonal frequency-division multiple access), or any other IFFT-based signal (IFFT = Inverse Fast Fourier Transformation) with or without CP, e.g. DFT-s-OFDM (DFT = discrete Fourier transform). Other waveforms, like non-orthogonal waveforms for multiple access, e.g. filter-bank multicarrier, FBMC, generalized frequency division multiplexing, GFDM, or universal filtered multi carrier, UFMC, may be used. The wireless communication system may operate, e.g., in accordance with the LTE-Advanced pro standard, or the 5G or 6G or NR, New Radio, standard, or the NR-U, New Radio Unlicensed, standard.

[0010] The wireless network or communication system depicted in Fig. 9A and Fig. 9B may be a heterogeneous network having distinct overlaid networks, e.g., a network of macro cells with each macro cell including a macro base station, like base stations gNBi to gNBs, and a network of small cell base stations, not shown in Fig. 9A and Fig. 9B, like femto or pico base stations. In addition to the above described terrestrial wireless network also nonterrestrial wireless communication networks, NTN, exist including spaceborne transceivers, like satellites, and / or airborne transceivers, like unmanned aircraft systems. The non-terrestrial wireless communication network or system may operate in a similar way as the terrestrial system described above with reference to Fig. 9A and Fig. 9B, for example in accordance with the LTE-Advanced Pro standard or the 5G or 6G or NR, new radio, standard.

[0011] In mobile communication networks, for example in a network like that described above with reference to Fig. 9A and Fig. 9B, like an LTE or 5G or 6G / NR network, there may be UEs that communicate directly with each other over one or more sidelink, SL, channels, e.g., using the PC5 / PC3 interface or WiFi direct. UEs that communicate directly with each other over the sidelink may include vehicles communicating directly with other vehicles,

[0012] FH260314PCT-2026106843. DOCXV2V communication, vehicles communicating with other entities of the wireless communication network, V2X communication, for example roadside units, RSUs, or roadside entities, like traffic lights, traffic signs, or pedestrians. An RSU may have a functionality of a BS or of a UE, depending on the specific network configuration. Other UEs may not be vehicular related UEs and may comprise any of the above-mentioned devices. Such devices may also communicate directly with each other, D2D communication, using the SL channels.

[0013] In a wireless communication network, like the one depicted in Fig. 9A and Fig. 9B, it may be desired to locate a UE with a certain accuracy, e.g., determine a position of the UE in a cell. Several positioning approaches are known, like satellite-based positioning approaches, e.g., autonomous and assisted global navigation satellite systems, A-GNSS, such as GPS, mobile radio cellular positioning approaches, e.g., observed time difference of arrival, OTDOA, and enhanced cell ID, E-CID, or combinations thereof.

[0014] In the context of 5G or 6G Framework for 3GPP AI / ML functionalities (Al: artificial intelligence; ML: machine-learning) will be identified, managed, and utilized for UE-side models and UE-part of two-sided models.

[0015] Regarding known wireless communication systems, only common models and additionally fine tuning is studied.

[0016] In the following, the term “functionality” may, e.g., be defined, for example, according to its definition in TR38.843: According to the definition there, in functionality-based LCM, network indicates activation / deactivation / fallback / switching of AI / ML functionality via 3GPP signaling (e.g., RRC, MAC-CE, DCI). Models may not be identified at the Network, and UE may perform model-level LCM. Whether and how much awareness / interaction NW should have about model-level LCM requires further study. For functionality identification, there may be either one or more than one functionality defined within an AI / ML-enabled feature, whereby AI / ML-enabled Feature refers to a Feature where AI / ML may be used. Note: UE may have one AI / ML model for the functionality, or UE may have multiple AI / ML models for the functionality.

[0017] In model-ID-based LCM, models are identified at the Network, and Network / UE may activate / deactivate / select / switch individual AI / ML models via model ID.

[0018] FH260314PCT-2026106843. DOCXWhen it comes to functionality / model management, 3GPP has made the following agreements / discussions in TR38.843 and in ongoing Release 19 regarding AI / ML for NR air interface:

[0019] According to the agreement for LCM for UE-sided model I Common LCM framework / signaling (RAN2, Rel-19), for UE-sided model, for the functionality management, the “network decision, network-initiated” AI / ML management is supported as a baseline. The following can be considered further “UE autonomous, decision reported to the network”, “Network decision, UE-initiated” (i.e. proactive approach). Moreover, “UE-autonomous, UE’s decision is not reported to the network” is not considered for Rel-19

[0020] According to TR38843, methods to assess / monitor the applicability and expected performance of an inactive model / functionality, including the following examples for the purpose of activation / selection / switching of UE-side models / UE-part of two-sided models / functionalities (if applicable) are an assessment / monitoring based on the additional conditions associated with the model / functionality, an assessment / monitoring based on input / output data distribution, an assessment / monitoring using the inactive model / functionality for monitoring purpose and measuring the inference accuracy, and an assessment / monitoring based on past knowledge of the performance of the same model / functionality (e.g., based on other UEs).

[0021] The above indicates how the RAN1 Rel-18 was thinking about selection of most suitable model / functionality.

[0022] In RAN1 Rel-19, CSI-compression is considered. The assumption is that the decoder is trained and then different UE vendors can develop their respective encoders. This process may, e.g., be different, since the UE part(s) / submodels can also be generated first, and a subset of the model (a number of sub-models in the “middle”) can be trained jointly and / or can be known to both sides (UE and NW).

[0023] From RAN1 perspective, for UE part of two-sided model, further study the following example of Ml-Option2 (including the feasibility / necessity) would be appreciated.

[0024] In Al-Example2-1, step A, a dataset is transferred from the NW / NW-side to UE / UE-side via standardized signaling. It should be noted that RAN1 study of Step A only focuses on RAN1 aspect of the dataset transfer from NW to UE. Other solution for dataset exchange is out of RAN1 scope.

[0025] FH260314PCT-2026106843. DOCXIn step B, a UE part of two-sided model(s) is(are) developed based on at least the above dataset.

[0026] In step C, the UE reports information of its UE part of two-sided model(s) corresponding to the above dataset to the NW. As for further study, explanations would be appreciated how a model ID is determined / assigned for each AI / ML model (including relationship between dataset and model ID). It should be noted that some step(s) may not be needed for Ml-Option 2.

[0027] Moreover, it should be noted that the above example is based on the assumption of NW-first training. It is separate discussion for the assumption of UE-first training.

[0028] Furthermore, it should be noted that the study should consider the impact on inter-vendor collaboration, at least including complexity, performance, interoperability in RAN4 / testing related aspects and feasibility.

[0029] It would be appreciated for further study whether / how to consider UE-side additional condition(s) for the dataset.

[0030] When a (general) AI / ML model has been trained, a series of offline (e.g., in a test chamber or using static datasets) and online tests (in the field) are carried out in order to determine its generalization / robustness capabilities. Generalization means how capable the model is to provide reliable inference outputs (predictions) for inputs (data) it has not encountered during training, while robustness is associated with the model’s ability to react to (reject) adversarial inputs.

[0031] Outside the 5G+ / 6G domain, for example on autonomous driving systems, large-scale recommender systems or machine translation applications, pre-deployment tests of AI / ML models include evaluation on carefully selected offline test datasets, followed by a series of tests where the model is operating in the target environment, such as:

[0032] • Dark launch or shadow mode tests: here the model is used for inference in parallel to a different model or a non-AI / ML method. Inference results (model outputs) are not used in the actual service but are collected and evaluated offline. If the model performs better compared to the current (AI / ML-based on non-AI / ML) solution, then the model is deployed and used for inference, otherwise the data collection and training phases are repeated.

[0033] FH260314PCT-2026106843. DOCX• A / B testing, where different versions of an AI / ML model are provided to different inference requests (e.g., different users) and a new model is deployed only when its performance is superior compared to other models under some statistical significance criteria.

[0034] These approaches are not suited for the pre- / post-deployment monitoring, validation and testing of AI / ML models for 5G+ / 6G applications. This can be seen intuitively in Fig. 10.

[0035] Fig. 10 illustrates areas with changing conditions that affect the performance of AI / ML models for positioning (top) or beam management (bottom).

[0036] On the top row, an area (e.g., a large warehouse) where an AI / ML model can perform positioning is shown. Different sub-areas have different properties, depending on the specific geometry. But the lower sub-area’s properties can vary during the day or between days (e.g., depending on the number of trucks that are stationed there for (un-)loading. The same is true for the area depicted in the bottom image. Here, a base station is serving the UEs (cell phones of people sitting in the cafe), but is occasionally (for example, for a few minutes three times per hour during the day) blocked by a bus.

[0037] If one generalizes these simple examples, it becomes clear that pre-deployment testing of 5G+ / 6G AI / ML models cannot cover every possible combination of environment geometry (e.g., affecting the LOS / NLOS conditions), radio conditions (e.g., SNR / SINR conditions) and UE movement patterns (e.g., adding Doppler effects into the mix). There will always be combinations of these parameters that do not occur too often (called also corner cases or rare events) and are under-represented (or not represented at all) in the data, as they are individual for each target location / area (defined here in the sense of TS 23.032), leading to a poor AI / ML model generalization. Dark launch / shadow mode or A / B testing practices are unable to address this problem, since one would need to monitor (or predeploy) the candidate models for a very long time in several areas, before collecting a representative dataset of such cases.

[0038] Within 3GPP, RAN4 has agreed that non-static scenarios / configurations are required for proper testing of UE-sided AI / ML models and has tentatively agreed that new / updated models can be tested in deployed UEs, before their usage (in a similar functionality as the shadow mode / dark launch and A / B testing described before). It has been also agreed, that UE-sided model testing / validation in the field provides also insights on the suitability

[0039] FH260314PCT-2026106843. DOCXof AI / ML functionalities (as defined in TR38.843). In this way, evaluating an AI / ML model allows direct evaluation of the corresponding functionalitie(s) that this model supports.

[0040] RAN4#110-bis Agreement:

[0041] • Both static and non-static scenarios / configurations could be needed for Al testing o RAN4 will further discuss how to use them case by case

[0042] ■ FFS whether to use static scenarios / configurations as baseline. o Refine the definitions of static and non-static scenarios / configurations based on two bullets below

[0043] ■ Static: channel model and SNR settings are fixed and do not change over the test, specific channel realizations may be dynamic ■ Non-static: Non-static scenarios / configuration can be further considered in application to use cases. The details of models are FFS and may include non-stationary SNR and other conditions.

[0044] Issue 1-2: Post deployment testing

[0045] RAN4#114 agreement

[0046] When operating in the field, two aspects of UE operation may impact performance:

[0047] • Update or fine tuning or addition / removal of models, if applicable, resulting in a change of functionality which may impact functionality performance

[0048] • Data drift / mismatch between the conditions encountered by the UE in the field and the training data, which may impact functionality performance.

[0049] For dealing with potential changes in the performance of functionalities two options may be available. The options are not mutually exclusive:

[0050] • OPTION 1: Conduct the validation of a change in functionality before its deployment / activation in already deployed UEs

[0051] o Validation takes into account the UE hardware in which the model is to be deployed / activated.

[0052] o FFS whether the validation takes place in the field (i.e., on each individual device) or in the lab conditions (i.e., per device type / model) o FFS whether validation can be done at the device along with inference for another active functionality(ies)

[0053] o FFS on the feasibility

[0054] FH260314PCT-2026106843. DOCXo FFS on one possibility for consideration for option 1 is to capture model functionality input (and if needed other test data such as ground truth) during conformance testing. This stored data can later be used to validate new or updated models.

[0055] o Other aspects not precluded.

[0056] • OPTION 2: Using performance monitoring and LCM procedures

[0057] o Performance monitoring will be designed in other groups

[0058] o RAN4 may consider the need and feasibility of requirements and tests to ensure consistency and accuracy of monitoring metrics or other monitoring related data sent from the UE, and set requirements as feasible / needed. o FFS on monitoring can be used for managing fallback, changes in functionality, functionality update functionality switching functionality transfer, if applicable

[0059] FFS: RAN4 needs to clarify whether changed functionalities that did not pass conformance testing are expected to be activated at the device.

[0060] For dealing with drift / mismatch between the conditions encountered by the UE in the field and the training data for the model, which may impact functionality performance, monitoring is needed.

[0061] • Performance monitoring will be designed in other groups

[0062] • RAN4 may consider the need and feasibility of requirements and tests to ensure consistency and accuracy of monitoring metrics or other monitoring related data sent from the UE, and set requirements as feasible / needed.

[0063] • Monitoring can be used for managing fallback, change in functionality, model functionality update / functionality model switching / functionality model transfer, if applicable

[0064] RAN4#114 way forward (R4-2502953)

[0065] The following table presents the description of options 1 and 2 and sub-options as already agreed.

[0066] A new name for “option 1” and “option 2” is proposed to reduce ambiguity when discussing in the future.

[0067] FH260314PCT-2026106843. DOCXThe descriptions of the options are the same as the already agreed description.

[0068] The “further potential clarifications” box captures issues that have been raised that could further clarify the options. It does not represent any agreement on the issues, or that all are relevant, but is intended to stimulate input and discussion to future meetings.

[0069] The “significant issues” line is intended to capture issues that have been raised that should be answered in order to make a decision.

[0070] RAN4#110-bis Agreement:

[0071] • Both static and non-static scenarios / configurations could be needed for Al testing o RAN4 will further discuss how to use them case by case

[0072] ■ FFS whether to use static scenarios / configurations as baseline. o Refine the definitions of static and non-static scenarios / configurations based on two bullets below

[0073] ■ Static: channel model and SNR settings are fixed and do not change over the test, specific channel realizations may be dynamic

[0074] ■ Non-static: Non-static scenarios / configuration can be further considered in application to use cases. The details of models are FFS and may include non-stationary SNR and other conditions.

[0075] Issue 1-2: Post deployment testing

[0076] RAN4#114 agreement

[0077] When operating in the field, two aspects of UE operation may impact performance:

[0078] • Update or fine tuning or addition / removal of models, if applicable, resulting in a change of functionality which may impact functionality performance

[0079] • Data drift / mismatch between the conditions encountered by the UE in the field and the training data, which may impact functionality performance.

[0080] For dealing with potential changes in the performance of functionalities two options may be available. The options are not mutually exclusive:

[0081] • OPTION 1: Conduct the validation of a change in functionality before its deployment / activation in already deployed UEs

[0082]

[0083] o Validation takes into account the UE hardware in which the model is to be

[0084] FH260314PCT-2026106843. DOCXdeployed / activated.

[0085] o FFS whether the validation takes place in the field (i.e., on each individual device) or in the lab conditions (i.e., per device type / model) o FFS whether validation can be done at the device along with inference for another active functionality(ies)

[0086] o FFS on the feasibility

[0087] o FFS on one possibility for consideration for option 1 is to capture model functionality input (and if needed other test data such as ground truth) during conformance testing. This stored data can later be used to validate new or updated models.

[0088] o Other aspects not precluded.

[0089] • OPTION 2: Using performance monitoring and LCM procedures

[0090] o Performance monitoring will be designed in other groups

[0091] o RAN4 may consider the need and feasibility of requirements and tests to ensure consistency and accuracy of monitoring metrics or other monitoring related data sent from the UE, and set requirements as feasible / needed.

[0092] o FFS on monitoring can be used for managing fallback, changes in functionality, model functionality update / model functionality switching / model functionality transfer, if applicable

[0093] FFS: RAN4 needs to clarify whether changed functionalities that did not pass conformance testing are expected to be activated at the device.

[0094] For dealing with drift / mismatch between the conditions encountered by the UE in the field and the training data for the model, which may impact functionality performance, monitoring is needed.

[0095] • Performance monitoring will be designed in other groups

[0096] • RAN4 may consider the need and feasibility of requirements and tests to ensure consistency and accuracy of monitoring metrics or other monitoring related data sent from the UE, and set requirements as feasible / needed. • Monitoring can be used for managing fallback, change in functionality, model functionality update / functionality model switching / functionality model transfer, if applicable

[0097] RAN4#114 way forward (R4-2502953)

[0098]

[0099] The following table presents the description of options 1 and 2 and sub-options as

[0100] FH260314PCT-2026106843. DOCXalready agreed.

[0101] A new name for “option 1” and “option 2” is proposed to reduce ambiguity when discussing in the future.

[0102] The descriptions of the options are the same as the already agreed description.

[0103] The “further potential clarifications” box captures issues that have been raised that could further clarify the options. It does not represent any agreement on the issues, or that all are relevant, but is intended to stimulate input and discussion to future meetings.

[0104] The “significant issues” line is intended to capture issues that have been raised that should be answered in order to make a decision.

[0105] Option 1 Option 2

[0106] New name Pre-activation functionality / model Post deployment update testing management based on LCM

[0107] Description Conduct the validation of a change in Using performance Al functionality before its monitoring and LCM deployment / activation in already procedures deployed UEs • Performance • Validation takes into account monitoring will be the UE hardware in which the model is designed in other to be deployed / activated. groups

[0108] • RAN4 may consider the need and feasibility of requirements and tests to ensure consistency and accuracy of monitoring metrics or other monitoring related data sent from the UE, and set requirements as feasible / needed. Sub-options Possible collection of input data to a

[0109] model during conformance testing for

[0110] use later on.

[0111] Further potential The following proposals have been The following proposals clarifications (not made for option 1 but not discussed, have been made for

[0112]

[0113]

[0114]

[0115] resolved) o Whether the testing of a new option 1 but not

[0116] FH260314PCT-2026106843. DOCXmodel is on a device or in a lab. discussed.

[0117] o Whether models that have been • Whether the metric updated post conformancetesting used for should be kept in a device. performance testing o Whether parallel operating of and the metric used one model and (in-device) testing of for monitoring can another can be assumed. be aligned.

[0118] o Whether signaling from a UE of • Whether monitoring changes / updates are needed. used for post• Whether there is any relation to deployment testing GCF procedures. should be NW • Whether it is possible to sided only.

[0119] differentiate “major” and “minor”

[0120] model changes and only apply

[0121] option 1 to major changes.

[0122] o Definition of Fine-tuning and

[0123] Model update

[0124] ■ Model update can

[0125] have a performance

[0126] impact.

[0127] ■ Fine-tuning might

[0128] have little

[0129] performance impact.

[0130] • Hybrid Validation approach

[0131] o A hybrid approach

[0132] integrating both preactivation

[0133] functionality / model update

[0134] testing (Option 1) and postdeployment management

[0135] based on LCM (Option 2)

[0136] aims to provide a balanced

[0137] validation strategy.

[0138] Significant issues Whether the option is reasonable or Whether LCM would cause too much complexity and monitoring can actually test time burden when introducing unambiguously identify

[0139]

[0140] updates. individual model

[0141] FH260314PCT-2026106843. DOCXperformance with reasonable complexity (considering other variations due to e.g. channel, interference, scheduling, other Al models etc.) Whether LCM monitoring can have RAN4 requirements to ensure consistent monitoring between different UEs.

[0142] RAN4#117 agreement

[0143] The topic has been extensively discussed also in Rel. 20 and has been deemed important for future 6G implementations also.

[0144] Agreement:

[0145] • The following clarification is provided to align companies’ understanding on predeployment conformance and post-deployment enhancement options.

[0146] Ways to guarantee When to perform Where to Comment AI / ML performance perform

[0147] (1) Pre-deployment | Before cell-phone Testing lab for Same as | conformance test | shipped into market conformance existing | testing conformance ] testing | (2) Post-deployment | After cell-phone FFS in UE Similar as | pre-activation | shipped into vendors’ lab or product | functionality test | market, but before testing lab testing ] (Option 1 in Rel-19 new AI / ML

[0148] discussion) | functionality

[0149] l activated

[0150]

[0151]

[0152]

[0153] (3) Post-deployment After new AI / ML FFS in-field

[0154] FH260314PCT-2026106843. DOCXpost-activation functionality practical

[0155] functionality testing activated network

[0156] [based on environment

[0157] performance

[0158] monitoring]

[0159] (Similar to Option 2

[0160] in Rel-19 discussion)

[0161] . FFS the following options for post-deployment enhancement in 6G study:

[0162] • Post-deployment pre-activation functionality test (Option 1 in Rel-19 discussion)

[0163] • Post-deployment post-activation functionality testing based on performance monitoring (Option 2 in Rel-19 discussion)

[0164]

[0165] Extending this idea, WO2023146749 suggests the utilization of a verification dataset, which can be downloaded at the UE (or sent to a UE model reference repository) and enable the verification of the UE-sided model (after its compilation and compression at the UE device). The results of this verification process are communicated to the NW, in a process that allows verification of UE-sided models in any supported area for which verification data are available.

[0166] WO2024199858 defines a verification time window, within which verification data for the UE-sided AI / ML model are collected. The model is evaluated under a set of verification KPIs and their respective thresholds and a verification report is transmitted to the NW.

[0167] Even though not discussed in 3GPP Rel. 18 and Rel. 19, several use cases (e.g., beam management or handover management) could be addressed using reinforcement learning algorithms:

[0168] Reinforcement Learning (RL): A process of training an AI / ML model (policy) to interact with an environment and take actions (model’s output) based on the environment’s current state (model’s input), with the goal of maximizing the expected cumulative reward (feedback signal). For the AI / ML model (policy) training, direct interaction with the environment, available logged data from the environment, or a combination of both can be used.

[0169] FH260314PCT-2026106843. DOCXThe Reinforcement Learning (RL) research field is quite broad.

[0170] • In “classical / online” RL and multi-armed bandits algorithms, an agent (AI / ML model) interacts with an environment and through training learns to improve (i.e., select outputs / actions that maximize the expected cumulative reward).

[0171] • In model-based RL, the agent learns a model of the environment using data collected during training and uses this model for online planning. In cases where a model of the environment is known, but its dimensionality hinders online planning in finite time, tree-based planners are utilized.

[0172] • In Imitation Learning, an agent learns to mimic the demonstrated behaviour (selected actions in the observations) either utilizing offline, logged data or by querying the expert interactively.

[0173] • In Inverse RL, an agent is tasked to examine offline, logged data from several (expert and non-expert) explicit demonstrations or interactions during nominal operation to approximate the underlying reward function. This reward function is in turn utilized to train an agent to solve the underlying application.

[0174] • In Offline RL, the reward function is available, and the agent utilizes offline, logged data from the environment containing previous (expert and non-expert) interactions during nominal operation, aiming at learning a behaviour that achieves better performance from the best performing interactions in the dataset.

[0175] It is clear from this discussion that RL algorithms are equally (or even more) challenging to test, while some aspects of them (e.g. bandits or online RL) can only be tested in the field after their deployment in the target devices.

[0176] Beyond the challenges associated with reinforcement learning algorithms, fine-tuning or re-training of AI / ML models in the field, especially when using mostly locally collected data, introduces additional risks that can cause a previously well-performing model to degrade or fail. These risks include, but are not limited to:

[0177] Overfitting, where the model memorizes the local training data instead of learning generalizable patterns, resulting in strong performance on the training data but poor performance under new or unseen conditions;

[0178] FH260314PCT-2026106843. DOCXUnderfitting, where the model fails to learn meaningful patterns due to, for example, insufficient local data volume or inadequate model capacity, leading to poor performance across all conditions;

[0179] Catastrophic forgetting, where new learning conducted during fine-tuning or re-training destroys previously acquired knowledge, causing the model to lose its ability to handle scenarios it could previously address reliably;

[0180] Label noise and label quality degradation, where in-field, “live” label collection / construction can lead to noisy, incorrect, or inconsistent labels that can corrupt the learning process and lead to unreliable models;

[0181] Limited / Local data: where the radio environment is highly dynamic and the data available for local fine-tuning or re-training may be limited in volume and biased towards recently observed conditions.

[0182] Furthermore, it should be noted that an AI / ML functionality, as perceived by the network (e.g., in the sense of functionality-based LCM as defined in TR 38.843), may be implemented by more than one AI / ML models at the UE side. Such implementations may rely on proper switching or routing between smaller, specialized models models, for example through conditional architecture structures such as mixture-of-experts networks, early exit networks, or slimmable neural networks.

[0183] In these architectures, a decision or routing module determines which sub-model(s) or which portion of the network is activated for a given input. As the fine-tuning (or retraining) of these modules is subseptible to the same risks as the AI / ML models mentioned above, the overall AI / ML functionality performance, as observed by the network, can be significantly reduced, even if all sub-models remain individually functional. Since such internal architectural details are transparent to the specification and to the network, performance degradation originating from routing or switching failures cannot be detected or addressed through conventional means, further motivating the need for robust post-deployment validation and testing mechanisms.

[0184] In addition to the above model-level risks, a practical constraint on post-deployment validation and testing arises from the limited processing resources available at the UE. In the 3GPP Rel-19 and Rel-20 framework, dedicated processing unit pools (CPU, 2 and CPU, 3) have been introduced for AI / ML-related CSI processing, separate from the

[0185] FH260314PCT-2026106843. DOCXbaseline CPU pool used for conventional CSI computations. The number of simultaneous AI / ML processing operations a UE can support is bounded by the capacity of these pools. Not all UE device categories are expected to support both pools: CPU, 3 is anticipated to be available only in high-end devices, while lower-tier devices may only support CPU, 2. This has direct implications for post-deployment monitoring, validation and testing:

[0186] If the UE must simultaneously perform inference for one or more active AI / ML functionalities (e.g., beam prediction or CSI prediction in normal operation), the remaining processing capacity available for monitoring, validation or testing is further reduced.

[0187] Even for UEs supporting both CPU, 2 and CPU, 3, the allocation of these pools to different AI / ML report types (e.g., beam prediction inference vs. CSI prediction inference vs. monitoring reports) is subject to UE capability, meaning that the actual resources available for testing at any given time depend on the concurrent AI / ML workload.

[0188] Shadow evaluation, where a candidate AI / ML model runs inference in parallel with the currently active model without its outputs being used in the actual service, inherently requires the UE to execute two models simultaneously, effectively doubling the demand on the AI / ML-dedicated processing pools. For UEs that only support CPU, 2, this parallel execution may be infeasible altogether, as the pool capacity may be insufficient to sustain both the active functionality inference and the shadow inference concurrently, rendering shadow evaluation either severely degraded or entirely impractical on such devices.

[0189] Consequently, the ability of a UE to perform adequate post-deployment monitoring, validation and testing is not only a function of the quality and representativeness of the monitoring / validation / test data but also a function of the UE's available AI / ML processing resources. This further motivates the need for a class-aware validation framework that allows the network to prioritize testing of the most critical environment (subset of) states, adapt the testing granularity to the UE's capabilities, and schedule testing in coordination with active AI / ML inference tasks.

[0190] It should further be noted that the validation and testing framework described here is not limited to post-deployment evaluation of fully trained or updated models. It is applicable (and essential) as a step of an online training or fine-tuning process conducted at the UE or at an external entity. In such a process, the model may periodically pause its training or fine-tuning cycle to validate the current, intermediate version of the model against representative validation / test data from one or more environment classes or (subset of) states, before continuing with further training iterations.

[0191] FH260314PCT-2026106843. DOCXNaturally, performing such intermediate validation during training further increases the demand on the UE's AI / ML-dedicated processing resources (CPU, 2 and, if available, CPU, 3), as the UE must interleave inference for validation purposes with the ongoing training computations, reinforcing the need for a configurable and network-coordinated approach to validation scheduling.

[0192] In all cases, the number of classes or (sub-set of) states to be tested in a given validation cycle may be determined depending on the available AI / ML processing resources at the UE (e.g., the number of unoccupied processing units in the CPU, 2 and, if supported, CPU, 3 pools), such that when processing resources are limited, the network may configure the UE to test only a prioritized subset of classes, with the remaining classes being deferred to subsequent validation cycles.

[0193] WO 2023 / 146749 A1 discloses a method for wireless communication by a first network device includes receiving a machine learning model

[0194] WO 2024 / 199858 A1 discloses embodiments enable reliability assessment of a machine learning model for mobility management related predictions.

[0195] However, it would be appreciated, if approaches would be provided that reliably ensure that an AI / ML functionality / model will perform as expected in a target area, for all possible combinations of environment geometry, radio conditions and UE movement patterns that might emerge at any given point in time.

[0196] It would therefore be appreciated, if improved concepts for post deployment validation and testing for artificial intelligence / machine learning models and functionalities would be provided.

[0197] The object of the present invention is to provide improved concepts for post deployment validation and testing for artificial intelligence / machine learning models and functionalities. The object of the present invention is solved by the subject-matter of the independent claims. Particular embodiments are provided in the dependent claims.

[0198] An apparatus for validating at least one AI / ML model or functionality according to an embodiment is provided. The apparatus comprises a validation module configured to conduct an evaluation by evaluating a performance and / or an applicability of the at least one AI / ML model or functionality depending on classes and / or attributes and / or

[0199] FH260314PCT-2026106843. DOCXperformance criteria. The validation module is configured to generate a validation output indicating whether the at least one AI / ML model or functionality meet one or more predefined criteria depending on the evaluation.

[0200] Moreover, an apparatus of a wireless communication system according to an embodiment is provided. A configuration, being obtained by the apparatus and / or being stored in the apparatus, comprises a plurality of datasets, wherein each of the plurality of datasets comprises one or more attributes and / or one or more classes; wherein the one or more attributes and / or the one or more classes depend on a property and / or a state of the wireless communication system and / or depend on a property and / or a state of a user equipment and / or a network entity of the wireless communication system. The apparatus comprises a processor; wherein. For each dataset of one or more datasets of the plurality of datasets, the processor is configured to execute at least one AI / ML model or functionality with model input data, which depends on said dataset, to obtain model output data for each of the one or more datasets.

[0201] Moreover, a method for validating at least one AI / ML model or functionality according to an embodiment is provided. The method comprises:

[0202] Conducting an evaluation by evaluating a performance and / or an applicability of the at least one AI / ML model or functionality depending on classes and / or attributes and / or performance criteria. And:

[0203] Generating a validation output indicating whether the at least one AI / ML model or functionality meet one or more predefined criteria depending on the evaluation.

[0204] Furthermore, a method for a wireless communication system according to an embodiment is provided. A configuration, being obtained by an apparatus and / or being stored in the apparatus, comprises a plurality of datasets, wherein each of the plurality of datasets comprises one or more attributes and / or one or more classes; wherein the one or more attributes and / or the one or more classes depend on a property and / or a state of the wireless communication system and / or depend on a property and / or a state of a user equipment and / or a network entity of the wireless communication system. The apparatus comprises a processor. For each dataset of one or more datasets of the plurality of datasets, the processor executes at least one AI / ML model or functionality with model input data, which depends on said dataset, to obtain model output data for each of the one or more datasets.

[0205] FH260314PCT-2026106843. DOCXFurthermore, a computer program according to an embodiment for implementing the above-described method, when being executed on a computer or signal processor is provided.

[0206] In the following, embodiments of the present invention are described in more detail with reference to the figures, in which:

[0207] Fig. 1 illustrates an apparatus for validating at least one AI / ML model or functionality according to an embodiment, which comprises a validation module.

[0208] Fig. 2 illustrates an apparatus for validating at least one AI / ML model or functionality according to another embodiment, which further comprises a configuration interface.

[0209] Fig. 3 illustrates an example overview of embodiments for a simplified AI / ML positioning task.

[0210] Fig. 4A, 4B illustrate a (Semi-)supervised machine learning dataset (Fig. 4A), and a reinforcement learning dataset (Fig. 4B).

[0211] Fig. 5 illustrates an example depiction of 3GPP network with core network depicted using SBI according to an embodiment.

[0212] Fig. 6 illustrates a validation of an AI / ML model at an external server retrieving information from the NW. UE transfers its model to the OTT server for validation according to an embodiment.

[0213] Fig. 7 illustrates a validation of an AI / ML model at an external server retrieving information from the NW. Model is available at the OTT server and is downloaded at the UE, once verified / validated according to an embodiment.

[0214] Fig. 8 illustrates a validation of an AI / ML model at an external server, with training data received from the NW and model output provided to NW for verification according to an embodiment.

[0215] FH260314PCT-2026106843. DOCXFig. 9A, 9B illustrate a schematic representation of an example of a terrestrial wireless network.

[0216] Fig. 10 illustrates areas with changing conditions that affect the performance of AI / ML models for positioning (top) or beam management (bottom).

[0217] Fig. 11 illustrates an example of a computer system on which units or modules as well as the steps of the methods described in accordance with the inventive approach may execute.

[0218] Fig. 1 illustrates an apparatus for validating at least one AI / ML model or functionality according to an embodiment.

[0219] The apparatus comprises a validation module 120 configured to conduct an evaluation by evaluating a performance and / or an applicability of the at least one AI / ML model or functionality depending on classes and / or attributes and / or performance criteria; and configured to generate a validation output indicating whether the at least one AI / ML model or functionality meet one or more predefined criteria depending on the evaluation.

[0220] According to an embodiment, the classes and / or attributes and / or performance criteria may, e.g., be stored in a memory. The validation module may, e.g., be configured to access the classes and / or attributes and / or performance criteria during the evaluation.

[0221] In an embodiment, the apparatus may, e.g., configured to obtain and / or process a configuration comprising the classes and / or attributes and / or performance criteria associated with a dataset.

[0222] Fig. 2 illustrates an apparatus for validating at least one AI / ML model or functionality according to another embodiment. The apparatus of Fig. 2 further comprises a configuration interface 110. This (optional) configuration interface 110 may, e.g., be configured to receive the configuration.

[0223] According to an embodiment, the configuration may, e.g., comprise attributes related to the input data of the at least one AI / ML model or functionality. And / or, the configuration may, e.g., comprise attributes resulting from the processing of data used as input for the at least one AI / ML model or functionality. And / or, the configuration may, e.g., may, e.g.,

[0224] FH260314PCT-2026106843. DOCXcomprise attributes not utilized by the AI / ML functionality / model but characterizing the state of the environment.

[0225] In an embodiment, at least one of the classes may, e.g., be a dataset comprising input and / or output data and / or one or more performance metrics and / or one or more performance thresholds, for example, wherein the dataset may, e.g., correspond to a substate.

[0226] According to an embodiment, each of the classes and / or each of the datasets may be associated with an Associated ID. The Associated ID represents one or more network configurations or conditions under which training data were collected for the training of the least one AI / ML model or functionality, without revealing information on the actual network implementation. In this way, the network can communicate to the UE or to other entities which environment state or configuration a given class or dataset corresponds to, while preserving the confidentiality of proprietary network deployment details.

[0227] In an embodiment, the validation module 120 may, e.g., be configured to generate the validation output such that the validation output may, e.g., comprise:

[0228] an output / prediction for each test measurement of one of all AI / ML models supporting a specified functionality; and / or

[0229] at least one performance indicator per dataset or per class; and / or

[0230] a pass / fail indicator per dataset or per class depending on a performance threshold; and / or

[0231] a single performance indicator for a plurality of datasets depending on how frequently a specific state appears and / or depending on a number of datapoints in each class.

[0232] According to an embodiment, the validation module 120 may, e.g., be configured to generate the validation output such that the validation output may, e.g., comprise at least one performance indicator, wherein the at least one performance indicator may, e.g., comprise information on AI / ML beam management and / or information on AI / ML direct positioning and / or information on AI / ML assisted positioning and / or information on CSI measurement and reporting and / or information on mobility.

[0233] FH260314PCT-2026106843. DOCXIn an embodiment, the at least one performance indicator may, e.g., comprise the information on AI / ML beam management and may, e.g., indicate:

[0234] a beam prediction accuracy by comparing the prediction results and a beam measurements from a resource set / resources; and / or

[0235] L1-RSRP difference information depending on an actual measurement of the L1- RSRP of one or more of Top K predicted beam, and L1-RSRP measurements from a resource set / resources for monitoring; and / or

[0236] RSRP difference information between the predicted RSRP and measured L1- RSRP of corresponding beam(s) of a resource set / resources for monitoring; and / or

[0237] a model prediction uncertainty / confidence; and / or

[0238] a pass / fail test, indicating performance achieved compared to a configured threshold.

[0239] According to an embodiment, the validation module 120 may, e.g., be configured to generate the validation output depending on an associated margin wherein the margin represents a specified tolerance range for acceptable prediction deviation.

[0240] In an embodiment, the at least one performance indicator may, e.g., comprise the information on AI / ML direct positioning and may, e.g., indicate:

[0241] a mean squared error between a predicted position and a label; and / or

[0242] an absolute average or maximum error between a predicted position and a label; and / or

[0243] a model prediction uncertainty or confidence, and / or

[0244] a pass / fail test indicating performance achieved compared to a configured threshold.

[0245] According to an embodiment, the at least one performance indicator may, e.g., comprise the information on AI / ML assisted positioning and may, e.g., indicate:

[0246] FH260314PCT-2026106843. DOCXa classification metrics for Line-of-sight and / or Non-line-of-sight, for example, an accuracy or an F1-score; and / or

[0247] a model prediction uncertainty and / or a model prediction confidence; and / or

[0248] a pass / fail test indicating performance achieved compared to a configured threshold.

[0249] In an embodiment, the at least one performance indicator may, e.g., comprise the information on CSI measurement and reporting and may, e.g., indicate:

[0250] a difference between a predicted CQI and a measured CQI for a given test configuration, and / or a difference between a predicted CQI and expected CQI for a given test configuration; and / or

[0251] a difference between a predicted PMI and a UE-reported PMI report for a given test configuration, and / or a difference between a predicted PMI and an expected PMI report for a given test configuration; and / or

[0252] a difference between a predicted RI and a UE-reported RI report for a given test configuration, and / or a difference between a predicted RI and an expected RI report for a given test configuration; and / or

[0253] a stability of one or more predicted CSI components under static scenarios, and / or a latency of one or more predicted CSI components under mobility scenarios; and / or

[0254] a performance improvement and / or performance degradation, e.g., depending on a difference between a predicted CSI component and a measured CSI component or depending on a difference between a predicted component and a gNB-selected component. For example, if a BLER of below 10% is obtained on using predicted CQI, using CQI+1 may cause the BLER to raise over 10%.

[0255] According to an embodiment, the at least one performance indicator may, e.g., comprise the information on mobility and indicates an L1-RSRP prediction accuracy and / or an L3-RSRP prediction accuracy.

[0256] FH260314PCT-2026106843. DOCXIn an embodiment, the validation module 120 may, e.g., be configured to mark the at least one AI / ML model or functionality as validated or not per dataset or per class depending on the evaluation.

[0257] According to an embodiment, the validation module 120 may, e.g., be configured to generate the validation output, such that:

[0258] the validation output may, e.g., comprise a pass / fail result for the functionality / model to be applied per dataset (class); and / or

[0259] the validation output may, e.g., comprise information on a gap between a performance threshold and an achieved performance indicator; and / or

[0260] the validation output may, e.g., comprise a signal to an AI / ML model or functionality monitoring and management framework, informing on a gap between the performance threshold and an achieved performance indicator.

[0261] In an embodiment, the validation module 120 may, e.g., be configured to mark the at least one AI / ML model or functionality as validated per dataset or class for a specific performance level, depending on the evaluation.

[0262] According to an embodiment, the validation module 120 may, e.g., be configured to detect clusters from a dataset depending on attributes data indicating data on the attributes.

[0263] In an embodiment, the validation module 120 may, e.g., be configured to construct the dataset for the at least one AI / ML model or functionality by

[0264] constructing a lower-dimensional space utilizing the attributes data,

[0265] detecting clusters in the lower-dimensional space,

[0266] selecting, from each of the clusters, a number of samples that are used as a test set.

[0267] According to an embodiment, the apparatus may, e.g., be configured to record the dataset with samples selected to represent the environment.

[0268] FH260314PCT-2026106843. DOCXIn an embodiment, the validation module 120 may, e.g., be configured to evaluate the performance of the AI / ML model or functionality in different states of a cell or of an area.

[0269] According to an embodiment, the validation module 120 may, e.g., be configured to provide samples for areas where at least two of the clusters overlap.

[0270] In an embodiment, the validation module 120 may, e.g., be configured to detect the clusters by employing an unsupervised learning method or by employing a weak or semisupervision learning method or by employing a prototypical network.

[0271] According to an embodiment, the validation module 120 may, e.g., be configured to employ a clustering or classification algorithm for determining

[0272] a number of different states an environment, for example, an area or a cell, can be in,

[0273] a number of data points which represent one of the different states,

[0274] a change of an environment comprising a situation where a state, e.g. being represented by a class, becomes obsolete, or when a new state, e.g. being represented by a class, is formed.

[0275] In an embodiment, the apparatus may, e.g., be configured to obtain representative data from one or more or all possible states of an environment, when receiving a new AI / ML model or functionality. The validation module 120 may, e.g., be configured to generate the validation output for the new AI / ML model or functionality.

[0276] According to an embodiment, the validation module 120 may, e.g., be configured to value the AI / ML model or functionality as either suitable or unsuitable for a specific area and / or a specific cell and / or a specific zone and / or for a sub-set of states.

[0277] In an embodiment, the at least one AI / ML model or functionality is an AI / ML model or functionality for a wireless communication system.

[0278] According to an embodiment, the attributes comprise one or more of:

[0279] an RSRP measurement pattern of an initial set of beams,

[0280] FH260314PCT-2026106843. DOCXa LOS / NLOS indication,

[0281] an entering of specified sub-areas in area / cell,

[0282] SNR / SINR measurements,

[0283] Doppler effect measurements,

[0284] information on whether a user is transitioning from indoor to outdoor environment and vice-versa,

[0285] a network load,

[0286] one or more BLER values,

[0287] information on HARQ retransmissions,

[0288] information on an interference detected by a UE.

[0289] In an embodiment, the attributes comprise a reinforcement learning or bandits policy existence and magnitude of uncertainty.

[0290] According to an embodiment, the validation module 120 may, e.g., be configured to detect an existence and / or a magnitude of uncertainty in a reinforcement learning or bandits policy

[0291] by detecting a deviation of ensemble predictions for ensemble policies; and / or

[0292] by conducting approximate counts on observation-action pairs in the data, indicating that some of the actions have not been selected often enough for some input observations; and / or

[0293] by conducting a method that learns a next-observation-predictor, which can estimate the policy uncertainty via calculating a prediction error on the observations; and / or

[0294] FH260314PCT-2026106843. DOCXby conducting a method, where a neural network is trained to predict the features of the observations, which are generated by a fixed random network, for example, by conducting a random-network distillation method.

[0295] In an embodiment, the apparatus may, e.g., be configured to monitor one or more events associated with the at least one AI / ML model or functionality.

[0296] According to an embodiment, the validation module 120 may, e.g., be configured to determine if an update of the classes is to be conducted by monitoring one or more events associated with the at least one AI / ML model or functionality.

[0297] In an embodiment, the one or more events comprise information on AI / ML Beam Management, and / or information on positioning and / or information on CSI measurement and reporting and / or information on mobility.

[0298] According to an embodiment, the one or more events comprise information on AI / ML Beam Management and comprise an indication indicating

[0299] that a best beam for moving UE does not change smoothly, and / or

[0300] that a measured RSRP value is lower than a first threshold value, and / or

[0301] that a QoS value is lower than a second threshold value, and / or

[0302] a QoE value is lower than a third threshold value.

[0303] In an embodiment, the one or more events comprise information on positioning and comprise an indication indicating

[0304] that a label-based or label-free monitoring indicates model quality below a threshold,

[0305] a heavy NLOS detection.

[0306] According to an embodiment, the one or more events comprise information on mobility and comprise an indication indicating

[0307] an increased number of failures, for example, as indicated by MDT, and / or

[0308] FH260314PCT-2026106843. DOCXping pong handover effects.

[0309] In an embodiment, the one or more events comprise information on reinforcement learning-based beam management and / or mobility and comprise an indication indicating

[0310] a high regret or an increased exploration, for example, for bandit algorithms, or

[0311] a large variance of rewards per episode or increased level of surprise for exploration-directed algorithms, and / or

[0312] that a QoS value is lower than a threshold value, and / or

[0313] that a QoE value is lower than another threshold value.

[0314] According to an embodiment, the one or more events comprise information on CSI measurement and reporting and comprise an indication indicating

[0315] an increased number of BLER detected by the UE; and / or

[0316] a change in Rl, CQI values, which do not change smoothly, for example, due to a sudden change in mobility parameters or a sudden blockage or a partial blockages; and / or

[0317] a detection of interference from other UEs or other network entities; and / or

[0318] a change in polarization, for example, due to rotation.

[0319] In an embodiment, the one or more events comprise information on mobility and comprise an indication indicating

[0320] a mismatch between a predicted L1-RSRP and a measured L1-RSRP; and / or

[0321] a mismatch between a predicted L3-RSRP and a measured L3-RSRP; and / or

[0322] a mismatch between a predicted handover failure and an actual occurrence of radio link failure; and / or

[0323] FH260314PCT-2026106843. DOCXa time-mismatch between a predicted handover failure and an actual occurrence of a radio link failure, and / or

[0324] a mismatch between a predicted measurement event and a true measurement event occurring at a UE; and / or

[0325] a time-mismatch between a predicted measurement event and a true measurement event occurring at the UE.

[0326] According to an embodiment, the one or more events comprise

[0327] a measured RSRP and / or a predicted RSRP improves above a threshold; and / or

[0328] a measured RSRP and / or a predicted RSRP falls below a threshold; and / or

[0329] a measured RSRP and / or a predicted RSRP raises above a first threshold, but remains below a second threshold; and / or

[0330] a measured RSRP and / or a predicted RSRP falls below a third threshold but remains above a fourth threshold; and / or

[0331] an RLF occurs earlier than a threshold value for a predicted RLF time.

[0332] In an embodiment, the network entity may, e.g., be configured to update the AI / ML model or functionality depending on the validation output to obtain an updated AI / ML model or functionality.

[0333] According to an embodiment, the AI / ML model or functionality is a neural network.

[0334] In an embodiment, the neural network is a CNN or is a LSTM or is a Bayesian NN or may, e.g., comprise transformer layers.

[0335] According to an embodiment, the AI / ML model or functionality has been trained using reinforcement learning.

[0336] In an embodiment, the AI / ML model or functionality has been trained using a loss function.

[0337] FH260314PCT-2026106843. DOCXAccording to an embodiment, the loss function is implemented as a MSE, or as a top-k accuracy, or as a F1-score.

[0338] According to an embodiment, the apparatus may, e.g., be an apparatus of a wireless communication system.

[0339] In an embodiment, the apparatus may, e.g., be a network entity of a wireless communication system.

[0340] According to an embodiment, an OTT-server executes the AI / ML model or functionality with a provided dataset, obtains an output of the AI / ML model and sends the output of the AI / ML model back to the network entity for validation. The network entity may, e.g., be configured to receive the output of the AI / ML model from the OTT-server. The validation module 120 of the network entity may, e.g., be configured to conduct the evaluation to generate the validation output.

[0341] In an embodiment, the network entity may, e.g., be configured to receive the AI / ML model or functionality, being a trained model or functionality, from an OTT-server. The network entity may, e.g., be configured to execute the AI / ML model or functionality. The validation module 120 of the network entity may, e.g., be configured to generate the validation output depending on the executing of the AI / ML model or functionality.

[0342] According to an embodiment, the network entity may, e.g., be configured to determine network analytics and / or whether the model has been verified / validated.

[0343] In an embodiment, the network entity may, e.g., be configured to indicate a successful verification or validation of the AI / ML model or functionality to the OTT-server and / or may, e.g., be configured to indicate one or more applicability conditions of the AI / ML model or functionality to the OTT-server.

[0344] According to an embodiment, the network entity may, e.g., be configured to update the AI / ML model or functionality depending on the validation output to obtain an updated AI / ML model or functionality.

[0345] In an embodiment, the network entity may, e.g., be configured to transmit the updated AI / ML model or functionality to a UE of the wireless communication system.

[0346] According to an embodiment, the apparatus may, e.g., be an OTT-server.

[0347] FH260314PCT-2026106843. DOCXIn an embodiment, the OTT-server may, e.g., be configured to receive input data for the AI / ML model or functionality and an expected output of the AI / ML model or functionality from a network entity of the wireless communication system or from a user equipment of the wireless communication system. The validation module 120 of the OTT-server may, e.g., be configured to execute the AI / ML model or functionality with the input data, may, e.g., be configured to compare output data of the AI / ML model or functionality with the expected output, and may, e.g., be configured to generate the validation output depending on the output data and the expected output.

[0348] According to an embodiment, the OTT-server may, e.g., be configured to receive the input data from the network entity. The OTT-server may, e.g., be configured to transmit the validation output to the network entity or to another network entity of a wireless communication system.

[0349] In an embodiment, the OTT-server may, e.g., be configured to receive the input data from the user equipment. The OTT-server may, e.g., be configured to transmit the validation output to the user equipment or to a network entity of a wireless communication system.

[0350] According to an embodiment, the OTT-server may, e.g., be configured to receive update information on an update of the AI / ML model or functionality from a UE of a wireless communication system. The OTT-server may, e.g., be configured to receive a dataset from a network entity of the wireless communication system. The validation module 120 of the OTT-server may, e.g., be configured to validate the model using the dataset from the network entity depending on the update information from the UE to generate the validation output.

[0351] In an embodiment, the UE trains and / or fine-tunes the AI / ML model or functionality. The OTT-server may, e.g., be configured to receive the AI / ML model or functionality or at least one model parameter of the AI / ML model or functionality from the UE.

[0352] According to an embodiment, the OTT-server may, e.g., be configured to request validation data for validating the AI / ML model or functionality from the network entity. The OTT-server may, e.g., be configured to receive the validation data from the network entity.

[0353] In an embodiment, the OTT-server may, e.g., be configured to update the AI / ML model or functionality depending on the validation output to obtain an updated AI / ML model or functionality.

[0354] FH260314PCT-2026106843. DOCXAccording to an embodiment, the OTT server may, e.g., be configured to transmit the updated AI / ML model or functionality to a UE of the wireless communication system.

[0355] In an embodiment, the apparatus may, e.g., be a UE of a wireless communication system.

[0356] According to an embodiment, the validation module 120 may, e.g., be configured to execute the AI / ML model or functionality with input data, is configured to compare output data of the AI / ML model or functionality with an expected output, and may, e.g., be configured to generate the validation output depending on the output data and the expected output.

[0357] In an embodiment, the apparatus may, e.g., be configured to receive the input data for the AI / ML model or functionality and the expected output of the AI / ML model or functionality from a network entity of the wireless communication system.

[0358] Moreover, an apparatus of a wireless communication system according to an embodiment is provided.

[0359] A configuration, being obtained by the apparatus and / or being stored in the apparatus, comprises a plurality of datasets, wherein each of the plurality of datasets comprises one or more attributes and / or one or more classes; wherein the one or more attributes and / or the one or more classes depend on a property and / or a state of the wireless communication system and / or depend on a property and / or a state of a user equipment and / or a network entity of the wireless communication system.

[0360] The apparatus comprises a processor. For each dataset of one or more datasets of the plurality of datasets, the processor is configured to execute at least one AI / ML model or functionality with model input data, which depends on said dataset, to obtain model output data for each of the one or more datasets.

[0361] According to an embodiment, the plurality of datasets may, e.g., comprise information on and / or depend on one or more of the following:

[0362] one or more measurement configurations,

[0363] one or more measurements,

[0364] one or more predictions,

[0365] aggregated data,

[0366] FH260314PCT-2026106843. DOCXpost-processed data,

[0367] time stamps,

[0368] correlations with other measurements,

[0369] one or more input data,

[0370] one or more output data,

[0371] relation of input and output data,

[0372] functional behavior between input and output.

[0373] In an embodiment, the plurality of datasets may, e.g., comprise information on and / or depend on one or more of the following:

[0374] aggregation of data,

[0375] post-processing of measurement data,

[0376] time stamping of measurements,

[0377] correlations with other measurements,

[0378] a relation of input and output data,

[0379] functional behavior between input and output,

[0380] logging, e.g., logging of measurement data and / or prediction data,

[0381] reporting, e.g., reporting of measurement data and / or prediction data.

[0382] According to an embodiment, each of the one or more datasets may, e.g., be associated with secondary data, wherein the second data comprises location data and / or timing data and / or UE or network entity ID data.

[0383] In an embodiment, the secondary data comprises location data indicating a position of a UE or a network entity in an environment; and / or the secondary data comprises timing data indicates when the information in the one or more attributes and / or in the one or more classes of a dataset has been obtained; and / or the secondary data comprises UE or network entity data indicating the UE or on the network entity to which information in the dataset relates.

[0384] According to an embodiment, each of the one or more datasets comprises the secondary data.

[0385] In an embodiment, a dataset of the one or more datasets may, e.g., further comprise a performance criteria and / or ground truth information, wherein the apparatus is configured to determine a performance result for the dataset depending on the model output data of

[0386] FH260314PCT-2026106843. DOCXthe dataset and depending on the performance criteria and / or the ground truth information of the dataset.

[0387] According to an embodiment, the apparatus comprises an interface, wherein the interface may, e.g., be configured to receive the one or more datasets. And / or, the interface may, e.g., be configured to output and / or transmit model output information comprising the model output data of the one or more datasets and / or depending on the model output data of the one or more datasets.

[0388] In an embodiment, the interface may, e.g., be configured to receive the one or more datasets from a network entity of the wireless communication system or from an OTT-server.

[0389] According to an embodiment, the interface may, e.g., be configured to transmit the model output information to a network entity of the wireless communication system or to an OTT-server.

[0390] In an embodiment, the interface may, e.g., be configured to transmit the model output information to said network entity of the wireless communication system or to said OTT-server for validating the AI / ML model or functionality.

[0391] According to an embodiment, said network entity or said OTT-server may, e.g., be an apparatus according to one of the above-described embodiments.

[0392] In an embodiment, the apparatus may, e.g., be configured to validate the AI / ML model or functionality using the model output data of the one or more datasets.

[0393] According to an embodiment, the apparatus may, e.g., be a user equipment.

[0394] In an embodiment, the apparatus may, e.g., be configured to conduct filtering based on the attributes and / or performance criteria, e.g., providing output based on the classes of the configuration.

[0395] According to an embodiment, the apparatus may, e.g., be configured to filter measurement data and / or prediction data depending on the attributes and / or depending on the performance criteria.

[0396] FH260314PCT-2026106843. DOCXIn an embodiment, the apparatus may, e.g., be configured to train and / or to fine-tune the AI / ML model or functionality.

[0397] According to an embodiment, the interface of the apparatus may, e.g., be configured to receive the validation output from the network entity or from the OTT-server.

[0398] In an embodiment, the apparatus may, e.g., be configured to update or fine-tune the AI / ML model or functionality depending on the validation output.

[0399] According to an embodiment, the interface of the apparatus may, e.g., be configured to receive an updated version of the AI / ML model or functionality from the network entity or from the OTT-server. The apparatus may, e.g., be configured to employ the updated version of the AI / ML model or functionality.

[0400] In an embodiment, the network entity may, e.g., be configured to receive data from one or more other network entities of the wireless communication system. The network entity may, e.g., be configured to analyse the data to obtain an analysing result and to transmit the analysing result of to a UE or to another network entity of the wireless communication system or to an entity outside of the wireless communication system.

[0401] According to an embodiment, the network entity may, e.g., be configured to receive raw measurements from a UE or a group of UEs of the wireless communication system and / or from one or more RAN nodes of the wireless communication system, which pertain to a UE or a group of UEs of the wireless communication system. The validation of the network entity may, e.g., be configured to generate the validation output depending on the raw measurements.

[0402] In an embodiment, the network entity may, e.g., be configured to group information from one or more other network entities of the wireless communication system, and may, e.g., be configured to link them together with an identifier, for example, the identifier being a dataset identifier, UE-identifier, timestamp, analysis identifier or a combination thereof.

[0403] According to an embodiment, the information may, e.g., comprise ground truth labels, for example, position information and / or velocity information and / or handover failure information.

[0404] In an embodiment, the network entity, e.g. a NWDAF, may, e.g., be configured to obtain the information by subscribing to the one or more other network entities, e.g. a gNB and / or

[0405] FH260314PCT-2026106843. DOCXan AF and / or an LMF, or by subscribing to an intermediate network node, e.g. a LMF and / or O& M and / or a gNB and / or an AMF, which collects the information.

[0406] According to an embodiment, the network entity may, e.g., be configured to perform a data analytics function to obtain analytics, for example, for detecting various patterns, such as a novelty detection and / or a change in performance and / or a change in attributes and / or a change in classes.

[0407] In an embodiment, the network entity may, e.g., be configured to provide the analytics to one or more internal network functions (e.g. LMF, PCF, AF) or one or more external functions, such as an OTT-server interacting with core network via NEF. Or, the network entity may, e.g., be configured to provide the analytics to external servers, e.g. interacting via NEF.

[0408] According to an embodiment, the network entity may, e.g., be configured to provide the analytics to external servers, interacting via the NEF. The NEF enables an external server to query a network storage for data with certain attributes.

[0409] In an embodiment, an AF or the NEF is able to subscribe to analytics corresponding to certain parameters, for example, comprising an attribute and / or a UE-ID and / or a vendor-ID and / or an area ID and / or a model ID, wherein the AF or the NEF requests to retrieve a dataset, for example, corresponding to the attribute and / or to the UE-ID and / or to the vendor ID and / or to the areaID and / or to the modelID.

[0410] According to an embodiment, the network entity may, e.g., be configured to store the analytics may in a network repository, e.g. in UDR.

[0411] In an embodiment, an analysed dataset is associated with one or more attributes.

[0412] According to an embodiment, the network entity may, e.g., be configured to provide a trigger based on events, for example unseen pattern detected or indicating to a network node, external node or a UE that the environment has changed, such that another network entity of the wireless communication system is informed to initiate a data collection procedure with a UE and / or a group of UEs and / or a RAN node and / or a group of RAN nodes of the wireless communication system.

[0413] In an embodiment, the network entity, e.g. a NWDAF, may, e.g., be configured to detect a change in the environment and may, e.g., be configured to indicate an event to a second

[0414] FH260314PCT-2026106843. DOCXnetwork entity of the wireless communication system, e.g. gNB, so that the second network entity is informed to initiate a data collection procedure from one or more UEs of the wireless communication system, e.g. retrieving logs stored by the one or more UEs, for example, the second network entity indicates to a group of UEs via system information broadcast or paging that the second network entity is requesting certain information to be reported.

[0415] According to an embodiment, the OTT-server may, e.g., be configured to classify data into different classes, wherein the different classes are identified by attributes, wherein a performance indicator is associated with at least one if the attributes. The OTT-server may, e.g., be configured to receive and / or to transmit a result of data analytics to a network entity of the wireless communication system for verification.

[0416] In an embodiment, the UE may, e.g., be configured to receive information on one or more attributes and their values for which the AI / ML model or functionality at the UE is valid.

[0417] According to an embodiment, the UE may, e.g., be configured to indicate a validity of the AI / ML model or functionality or that the AI / ML model or functionality is no longer valid, when one or more attributes are no longer valid.

[0418] In an embodiment, the UE may, e.g., be configured to indicate the validity of the AI / ML model or functionality or that the AI / ML model or functionality is no longer valid by updating a UAI or by using a ProvideCapabilityMessage indicating the change in applicability of the AI / ML model or functionality used by the UE.

[0419] Moreover, a system according to an embodiment is provided. The system comprises a network entity and a user equipment. The network entity or the user equipment implements an apparatus according to one of the above-described embodiments.

[0420] Furthermore, a system according to an embodiment is provided. The system comprises a network entity, a user equipment, and an OTT-server. The network entity or the user equipment or the OTT-server implements an apparatus according to one of the abovedescribed embodiments.

[0421] According to an embodiment, the UE may, e.g., be configured to report a key or a token to the network entity, for example when registering its applicable capabilities at the network. The network may, e.g., be configured to check, depending on the key or the token,

[0422] FH260314PCT-2026106843. DOCXwhether the AI / ML model or functionality can be activated at the network or if the AI / ML model or functionality cannot be activated, e.g., needs to be further tested, for example, with additional data.

[0423] In an embodiment, the UE may, e.g., be configured to report its capabilities. A RAN node, being implemented at the network entity or at another network entity, may, e.g., be configured to check whether the AI / ML model or functionality can be activated by the UE for given applicability conditions, for example, for a current state of the environment.

[0424] According to an embodiment, the network may, e.g., be configured to provide further test data, and may, e.g., be configured to request results from the UE. The UE may, e.g., be configured to perform inference on the further test data and may, e.g., be configured to report the results. The network entity may, e.g., be configured to activate the AI / ML model or functionality and / or may, e.g., be configured to provide a new key or a new token to be registered with the AI / ML model or functionality.

[0425] In an embodiment, the network entity may, e.g., be configured to store fine-tuned model information, e.g. UE-context information, in an NG-RAN node and / or in an AMF and / or in an UDM and / or in an UDR.

[0426] In the following, particular embodiments of the present invention are described.

[0427] A core idea is that each area / cell / zone served by one or more TRPs / gNBs can be in a different state, described by a set of distinct properties that hold for this specific situation in space and time. These properties depend on the combination of geometrical (e.g., temporal blockages), radio (e.g., propagation conditions, SINR) and UE (e.g., moving speed and moving pattern) properties.

[0428] We define a set of attributes (resembling the notion of features in AI / ML terminology) which separately or combined provide a clear indication on which state the environment currently is. There are different categories of attributes:

[0429] • Attributes related to the input data of the AI / ML functionality / model. For example, the differences between the RSRP pattern of Set B beams (for the same UE location) could indicate the existance or not of a (temporal) blocker.

[0430] FH260314PCT-2026106843. DOCX• Attributes resulting from the processing of data used as AI / ML functionality / model input. For example, determining the LOS / NLOS conditions is an indication on how “challenging” specific sub-areas are for AI / ML positioning, due to termporal and fixed blockers.

[0431] • Attributes not utilized by the AI / ML functionality / model but still crucial for characterizing the state of the environment, such as time information (to capture differences due to time of day, workdays vs weekends, seasonality effects, etc.), SNR / SINR measurements, Doppler effects due to UE movement, information on whether a user is transitioning from indoor to outdoor environment and vice-versa, and network load.

[0432] A clustering or classification algorithm, utilizing any combination of selected attributes is able to determine:

[0433] • The number of different states the environment (e.g., specific area / cell) can be in;

[0434] • How many datapoints represent each state (i.e., how frequent each state occurs);

[0435] • Environment changes, since some states (represented by classes) would become obsolete and new states will be formed.

[0436] For the application, states imply different “versions” the (radio) environment can be in. By looking at collections of conditions (we call these attributes in this application) we can infer the status (“version”) of the (radio) environment (e.g., SINR measurements, Doppler, network load, etc.).

[0437] Having this clustering, enables thorough UE-sided AI / ML functionality / model monitoring, validation and testing. Once a new / updated model enters the area / cell / zone, representative data from all possible environment states are provided to the UE for inference. The UE, depending on the configuration, either calculates the requested KPIs (per dataset) and reports the results back to the NW, either reports the model outputs directly. In an alternative version, the NW communicates both the dataset and the related attributes it utilizes to distinguish between states. This way, the UE can build its own classifier or, at minimum, is able to map its own models to different attribute values, thus enhancing its functionality or model-based LCM operations (e.g., model section / (de-) activation / switching and fallback or update of supported functionalities in the UE capability report) for a specific area.

[0438] FH260314PCT-2026106843. DOCXConsequently, the AI / ML functionality / model can be validated as suitable for the specific area / cell / zone or for a sub-set of states.

[0439] Furthermore, each class or dataset may be associated with an Associated ID. An Associated ID is used to express network configurations or conditions during data collection for AI / ML model training, without disclosing implementation-specific details of the network. By linking each class to a distinct Associated ID, the validation, testing, and monitoring framework can operate in a vendor- and implementation-agnostic manner: the network indicates the relevant Associated ID(s) to the UE or to an external entity, enabling them to identify which model or functionality applies to the current conditions without exposing the actual network deployment. This also enables the UE to report monitoring, validation or testing results per Associated ID, allowing the network to assess AI / ML model performance on a per-configuration basis.

[0440] An example overview of embodiments is shown in Fig. 3, for a simplified AI / ML positioning task. In particular, Fig. 3 illustrates a dataset acquisition architecture for a direct AI / ML positioning task. Top left: An area consists of a large unobstructed sub-are (green), and a series of racks. The sub-area around smaller racks is characterized by similar radio conditions (as indicated by the shaded blue region), while the sub-area around the larger rack is different. Bottom left: after the selection of attributes (and possible projection to a lower dimensional space), different classes representing the different states of the environment are formed. Right: when required, data are sampled per class under different sampling / selection schemes.

[0441] Despite a model being fine-tuned with area-specific data, there is a chance that its prior functionality is not affected, i.e., when continual learning, multi-task learning, active learning or similar approaches are utilized for the further training / adaptation of the model. In this case, the training entity of the AI / ML model can notify on the training method and the post-deployment testing can be conducted with a fewer data, targeting specific classes, or not at all.

[0442] According to an embodiment, an apparatus is provided. The apparatus comprises a processing unit configured to validate one or more AI / ML models or functionalities; a configuration interface 110 that receives a configuration, wherein the configuration comprises a set of [predefined] classes, attributes, or performance criteria associated with a dataset; and a validation module 120 configured to evaluate the [performance / applicability] of the AI / ML models based on the [predefined] classes,

[0443] FH260314PCT-2026106843. DOCXattributes, or performance criteria; and generate a validation output indicating whether the AI / ML models meet the predefined criteria.

[0444] In the following, particular embodiments of the validation module 120 are described.

[0445] E.g., a validation report reports an information related to the validation.

[0446] Different validation information can be reported, depending on the validation apparatus configuration:

[0447] • The output / prediction for each test measurement (input) of one of all AI / ML models supporting a specified functionality.

[0448] • At least one KPI per dataset (or class). The KPIs need not be the same for all datasets. For example, for classes with many samples the KPI could be average accuracy, but for classes with fewer samples, we might require the outputs to all test samples.

[0449] • If a performance / KPI threshold is configured, a PASS / FAIL indicator per dataset (or class).

[0450] • A single KPI if different weighting per dataset (for example, depending on how frequently the specific state appears I number of datapoints in each class) and / or KPI is supported. This also protects proprietary requirements if the apparatus is the UE, since the reported KPI is the weighted sum of all KPIs calculated.

[0451] Examples of KPIs may, e.g., be:

[0452] • AI / ML Beam Management

[0453] o Top 1 or Top K beam prediction accuracy (with or without margin) by comparing the prediction results and the Top 1 or Top K beam based on the measurements from a resource set / resources for monitoring o The L1-RSRP difference information based on actual measurement of the L1-RSRP of one or more of Top K predicted beam, and L1-RSRP measurements from a resource set / resources for monitoring o The RSRP difference information between the predicted RSRP and measured L1-RSRP of corresponding beam(s) of a resource set / resources for monitoring

[0454] FH260314PCT-2026106843. DOCXo Model prediction uncertainty / confidence

[0455] o Test PASS / FAIL, indicating performance achieved compared to a configured threshold

[0456] • AI / ML Direct Positioning

[0457] o Mean squared error (MSE) between predicted position and label o Absolute average / max error between predicted position and label o Model prediction uncertainty / confidence

[0458] o Test PASS / FAIL, indicating performance achieved compared to a configured threshold

[0459] • AI / ML Assisted Positioning

[0460] o Classification metrics (e.g., accuracy, F1-score) for LOS / NLOS

[0461] o Model prediction uncertainty / confidence

[0462] o Test PASS / FAIL, indicating performance achieved compared to a configured threshold

[0463] • CSI Measurement and Reporting

[0464] o CSI components include - CQI (channel quality indicator), PMI (precoding matric indicator), CRI (CSI-RS resource indicator), SSBRI (SS / PBCH resource block indicator), LI (layer indicator), Rl (rank indicator).

[0465] o Accuracy between predicted CSI vs. measured CSI report across one or more CSI components.

[0466] ■ Difference between Predicted CQI vs. measured CQI I expected CQI for the given test configuration.

[0467] ■ Difference between Predicted PMI vs. UE-reported PMI or expected PMI report for the given configuration.

[0468] ■ Difference between Predicted Rl vs. UE-reported Rl or expected Rl report for the given configuration.

[0469] o Stablity of predicted CSI components under static scenarios, (or latency of predicted CSI components under mobility scenarios).

[0470] o Performance improvement / degradation on utlising predicted CSI component vs. measured CSI component (or vs. gNB selected component).

[0471] ■ For example, if BLER of below 10% is obtained on using predicted CQI, using CQI+1 shall cause the BLER to raise over 10%.

[0472] In the following, validation actions are considered.

[0473] FH260314PCT-2026106843. DOCXThere are several actions that can follow the validation output of the AI / ML functionality / model and can be individually applied or in combination:

[0474] • Functionality / Model marked as validated or not per dataset / class, depending on the testing / validation outcome

[0475] • If a KPI / performance threshold is provided in the configuration:

[0476] o Validation output could be PASS / FAIL for the functionality / model to be applied per dataset (class)

[0477] o Validation output could be the gap between threshold and achieved KPI o Validation output could be a signal to the functionality / model monitoring and management framework, informing on the gap in KPI vs threshold. From there, LCM processes for functionality (e.g., more intense monitoring, deactivation, etc.) or model (e.g., more intense monitoring, data collection for re-training or fine-tuning) could be initiated

[0478] • Functionality / Model marked as validated per dataset / class for specific performance level, depending on the testing / validation outcome (KPI / performance values)

[0479] Validation procedure: Construction of the classifier

[0480] Embodiments can be applied to models trained both with (semi-)supervised and reinforcement learning algorithms. Different types of algorithms could be used for the same task. For example, beam management can be seen as a supervised learning task, where a labeled dataset is available during training or as a reinforcement learning task, where a bandit or a fully-fledged RL algorithm solves the problem online. The difference of interest here, as shown in Fig. 4A, 4B, would be on the datasets collected:

[0481] • (semi-)supervised learning datasets, with features x and - optional - labels y for prediction tasks;

[0482] reinforcement learning datasets, with observations o, actions a, rewards r, initial states / , termination indications T, cost / constraint violations c, and goals / targets g for decision making tasks.

[0483] FH260314PCT-2026106843. DOCXIn particular, Fig. 4A and 4B illustrate a (Semi-)supervised machine learning dataset (Fig.

[0484] 4A), and a reinforcement learning dataset (Fig. 4B). Shaded columns indicate optional entries.

[0485] For training the classifier, several different options exist:

[0486] • Unsupervised learning methods learn pattern exclusively from unlabeled data.

[0487] Challenges of unsupervised learning are scalability of methods due to their computational expensiveness, high sensitivity to hyperparameters and interpretability. Furthermore, traditional metrics (accuracy or precision) are unavailable. Hence, evaluation often relies on subjective measures (cluster compactness), such as PCA ort-SNE methods.

[0488] • In weak- or semi-supervision learning, only a small portion of the data is labeled.

[0489] However, it is assumed that labeled and unlabeled data share similar distributions. Hence, cluster evaluation techniques (cluster assumption) from unsupervised learning can be used. Constructing labeled evaluation sets faces the challenge of quality of labeled data (i.e., noisy, irrelevant, or mislabeled data), and small labeled data size. Hence, a combination of cluster and accuracy metrics must be defined.

[0490] Prototypical networks: The goal of prototypical learning is to create a (neural network) classifier that can classify new, unseen data points by comparing them to representative prototypes of each class. Given that data is often limited, prototypical networks operate under the assumption that a classifier should exhibit a simple inductive bias. This method is grounded in the concept that there exists an embedding space where points naturally cluster around a single prototype representation for each class. However, such methods rely on prototypes, which cannot capture the true class characteristics when they are noisy or unrepresentative. Consequently, for each new class prototype the classifier adapts to, a new test set must be defined that best represents the adapted class.

[0491] In the following, as an example, we suggest a clustering-based method to construct a dataset to for AI / ML functionality / model which can be adapted to any unsupervised, semisupervised and prototypical learning approach:

[0492] • First, a dataset must be available or must be recorded with samples selected that best represent the environment. From this dataset, a lower-dimensional space is constructed utilizing the attributes data. Methods to transform the high-dimensional input data onto lower dimensions range from linear techniques such as principal

[0493] FH260314PCT-2026106843. DOCXcomponent analysis (PCA) and linear discriminant analysis (LDA) to non-linear techniques such as uniform manifold approximation and projection (UMAP) and isometric mapping (Isomap). An established method is t-distributed stochastic neighbor embedding.

[0494] • In this (e.g., 2D) lower-dimensional space, clusters are detected (for example an open large-scale line-of-sight area forms one cluster, while a multipath are with high non-line-of-sight effects form another cluster). From each cluster, a (small) number of samples are chosen equally distributed and are used as the test set for each environment. Hence, it can be evaluated in which state of the cell / area the functionality / model performance is low / high.

[0495] • In case some clusters overlap (for example, in large industrial warehouses and factory roads, repetitive pattern of repeating objects can lead to similar multipath effects and overlapping clusters in the lower-dimensional space), samples from these overlapping parts of the classes need also be provided during the functionality / model validation or testing.

[0496] In the following, particular example attributes used by the classifier are considered:

[0497] For tasks related to (semi-)supervised learning, like AI / ML beam management, AI / ML direct and assisted positioning, CSI measurement and reporting and AI / ML mobility addressed in Rel. 18 and Rel. 19 of 3GPP discussions, relevant attributes are the following:

[0498] • RSRP measurement pattern of an initial set of beams (Set B beams in Rel. 18 / 19 in 3GPP) (AI / ML beam management and AI / ML mobility)

[0499] • LOS / NLOS indication (AI / ML positioning and beam management)

[0500] • Entering specified sub-areas in area / cell (AI / ML positioning)

[0501] • SNR / SINR measurements (ALL)

[0502] • Doppler effect measurements (ALL)

[0503] • Information on whether a user is transitioning from indoor to outdoor environment and vice-versa (ALL)

[0504] • Network load (ALL)

[0505] • BLER values (beam management, mobility, CSI reporting)

[0506] • HARQ retransmissions (beam management, CSI reporting)

[0507] • Interference detected by the UE.

[0508] FH260314PCT-2026106843. DOCXRLF occurs within a finite window of predicted RLF time.

[0509] For the reinforcement learning use cases, example attributes may, e.g., be:

[0510] • SNR / SINR measurements

[0511] • Doppler effect measurements

[0512] • Information on whether a user is transitioning from indoor to outdoor environment and vice-versa

[0513] • Network load

[0514] • Policy existence and magnitude of uncertainty (explained below).

[0515] Policy / agent uncertainty stems from the fact that some aspects of the environment (i.e., cell / area) are under-explored (e.g., some states are not so common or different actions / decisions have not been tried for some input observations) or that some important aspects of the environment cannot be observed directly (making the state definition ambiguously). There are several ways to detect the existence and magnitude of uncertainty in the policy, for example:

[0516] • Deviation of ensemble predictions for ensemble policies;

[0517] • Approximate counts on observation-action pairs in the data, indicating that some of the actions have not been selected often enough for some input observations; • Method that learn a next-state-predictor can estimate the policy uncertainty via calculating a prediction error on the observed states

[0518] • Methods like random-network distillation (RND), where a neural network is trained to predict the features of the observations, which are generated by a fixed random network

[0519] In the following, updates of classes in the classifier are considered:

[0520] If there are any changes in the environment, some of the classes might become obsolete and some new classes might be born, for which data need to be collected. In turn, the classifier needs to be fine-tuned. These changes are usually detected by specific monitoring events associated with an already deployed AI / ML functionality / model (e.g., frequent functionality / model fallbacks or switches).

[0521] For (semi-)supervised ML models, such events may, e.g., be:

[0522] AI / ML Beam Management

[0523] FH260314PCT-2026106843. DOCXo Best beam for moving UE does not change smoothly, which an indication of heavy NLOS conditions

[0524] o The measured RSRP values are low

[0525] o Low QoS / QoE

[0526] • Positioning:

[0527] o Label-based or label-free monitoring indicate low-model quality

[0528] o Heavy NLOS detection

[0529] • Mobility:

[0530] o Increased number of failures, as indicated by the UE logs ( e.g collected as part of M DT)

[0531] o Ping pong handover effects

[0532] o Events triggered based on predicted or measured events

[0533] ■ RSRP (measured or predicted) improves above a threshold.

[0534] ■ RSRP (measured or predicted) falls below a threshold.

[0535] ■ RSRP (measured or predicted) raises above a threshold but remains below a second threshold.

[0536] ■ RSRP (measured or predicted) falls above a threshold but remains above a second threshold.

[0537] ■ RLF occurs earlier than a threshold value from the predicted RLF time.

[0538] ■ RLF occurs later than a threshold value from the predicted RLF time (e.g. when the UE has not taken action based on the predicted value).

[0539] ■ RLF occurs earlier than a first threshold value but later than a second threshold value from the predicted RLF time.

[0540] ■ RLF occurs later than a threshold value from the predicted RLF time (e.g. when the UE has not taken action based on the predicted value).

[0541] ■ RLF occurs earlier than a first threshold value but later than a second threshold value from the predicted RLF time (e.g. when the UE has not taken action based on the predicted value).

[0542] o RLF occurs within a finite window of predicted RLF time.

[0543] CSI prediction or reporting

[0544] FH260314PCT-2026106843. DOCXo Increased number of BLER detected by the UE

[0545] o Change in Rl, CQI values that don’t change smoothly (e.g. due to sudden change in mobility parameters or sudden blockage or partial blockages). o Propagation conditions

[0546] o Detection of interference from other UEs or other network entities o Change in polarization (e.g. due to rotation).

[0547] For RL models, such events may, e.g., be:

[0548] • AI / ML Beam Management AND mobility

[0549] o High regret, increased exploration, etc. for bandit algorithms

[0550] o Large variance of rewards per episode, increased level of surprise for exploration-directed algorithms

[0551] o Low QoS / QoE

[0552] In the following, aspects of particular embodiments, e.g., relating to a system Architecture, and, e.g., relating to a configuration interface(s) are considered.

[0553] Fig. 5 illustrates an example depiction of 3GPP network with core network depicted using SBI according to an embodiment.

[0554] In particular, Fig. 5 depicts a mobile communication network, where the connection in core network (CN) is depicted using service-based interface (SBI). Not all entities are shown, but a subset of network functionalities that are needed to explain the concept presented in this invention.

[0555] A UE (e.g. UE1) may establish connectivity with the network, by registering itself with a network node (e.g. AMF in case of 5GS). The UE may achieve connectivity with another network entity which may be placed within a trusted domain of the mobile network operator or outside the trusted domain. In Fig. 5, an example of scenario is depicted where the UE server is depicted as OTT and is placed outside the trusted domain. The UE server is then able to influence, provide data or information, or retrieve data and information to the network by using the network exposure function. Alternatively, the OTT server may be provisioned within the trusted zone of a network operator and may have an application function (AF) directly connected to the SBI of the core network. In one way or other, the AF is able to provide and / or receive data, provide and / or retrieve analytics,

[0556] FH260314PCT-2026106843. DOCXretrieve event notification from a second network node (e.g. the NWDAF). There may also be provision to save data within network nodes, such as UDR (unified data repository).

[0557] A UE OTT-server is the server with which the UE can establish a connection to, and the UE can interact with the OTT-server to get updates and / or exchange information transparent to the mobile network to which UE is registered with. The OTT server may be located inside the trusted region or may be outside the trusted region of a core network. An OTT server may also serve as a UE-model repository, a UE-model validation server, or a software management server (over which model is provisioned or updated) for the UE.

[0558] In some examples, the UE model may be trained at the OTT-server, which may be located inside (i.e. access to CN via AF) or outside (i.e. access via NEF) the core network of the mobile operator.

[0559] In some scenarios, the UE model may be trained and / or fine-tuned at the UE.

[0560] The trained model may be validated at the UE, at an OTT-server or at a NWDAF function. E.g., the following variants may, e.g., exist:

[0561] 1) The OTT-server interacts with the core network, to obtain datasets for training and / or monitoring.

[0562] a. The OTT-server executes the model with the provided dataset, obtains the output and sends the output back to the network entity for validation. OR b. The OTT-server receives the input data and the expected output. The OTT- server executes the model, compares the expected output with the output data of the model, determines one or more KPI and sends the result back to the network.

[0563] 2) The OTT-server provides a ML-model to the core network, to get it validated.

[0564] a. The OTT provides the trained model to the core network and gets it executed there.

[0565] b. The network entity determines the network analytics and whether the model has been verified / validated.

[0566] c. The network entity may indicate successful verification / validation to the model server and may further indicate one or more applicability conditions.

[0567] FH260314PCT-2026106843. DOCX3) The OTT-server receives updates from the UE regarding the model updates and validates the model with using the received dataset from a network function.

[0568] a. The UE trains or fine-tunes an existing model.

[0569] b. The UE transfers the model or at least one model parameter to the OTT server.

[0570] c. The OTT server interacts with at least one network network node to obtain validation data for the UE model.

[0571] d. The OTT-server validates the performance of the model using the validation data.

[0572] Fig. 6 illustrates a validation of an AI / ML model at an external server retrieving information from the NW. UE transfers its model to the OTT server for validation according to an embodiment.

[0573] Fig. 7 illustrates a validation of an AI / ML model at an external server retrieving information from the NW. Model is available at the OTT server and is downloaded at the UE, once verified / validated according to an embodiment.

[0574] Fig. 8 illustrates a validation of an AI / ML model at an external server, with training data received from the NW and model output provided to NW for verification according to an embodiment.

[0575] In some examples, a network node (e.g. NWDAF) may be configured to have an analysing functionality. The network node may subscribe to one or more network nodes to receive data or a set of data. The data may be collected from multiple network nodes and associated at the network node using timestamps or with additional processing (e.g. deriving associations based on one or more information from a network node and / or a UE). The network node may analyse and provide the results of analytics to UE, another network node or an entity outside the mobile network (interacting with the mobile network node using the network exposure function).

[0576] A network node may collect the raw measurements from a UE or a group of UEs and / or raw measurement from one or more RAN nodes pertaining to a UE or a group of UEs. A network node may group the information collected from one or more nodes, linking them together with an identifier. The identifier may be a data-set identifier, UE-identifier, timestamp, analysis identifier or it may be based on a combination of any of the above. The information may additionally (and optionally) further consist of ground truth labels (e.g. position, velocity, handover failures... etc). The network node (e.g. NWDAF) may

[0577] FH260314PCT-2026106843. DOCXobtain the information by subscribing to each node (e.g. gNB, AF, LMF) or by subscribing to an intermediate network node (e.g. LMF, O& M, gNB, AMF) that collects the information.

[0578] The network node may be equipped with analyser module, which performs data analytics function. The data analytics may detect various patterns (such as novelty detection, change in performance, change in attributes, change in classes, etc.).

[0579] In some examples, the network node may provide a trigger based on events, such as unseen pattern detected or indicating to a network node, external node or a UE that the environment has changed. A second network node may initiate data collection procedure with a UE and / or a group of UEs and / or a RAN node and / or a group of RAN nodes.

[0580] For example, if a network node (e.g. NWDAF) detects change in environment and indicates an event notifying a second network node (e.g. LMF), the LMF may initiate data collection from a UE or a group of UEs or RAN nodes or a combination of all above.

[0581] Similarly, if a network node (e.g. NWDAF) detects change in environment and indicates an event notifying a second network node (e.g. gNB), then the gNB may initiate data collection procedures from one or more UEs (e.g. retrieving the logs stored by the UE). In such scenario, the gNB may also indicate to a group of UEs via system information broadcast or paging that the gNB is requesting certain information to be reported.

[0582] The network function may expose the analytics to internal network functions (e.g. LMF, PCF, AF) or external function such as an OTT-server interacting with core network via NEF.

[0583] Likewise, in some examples, the node outside the mobile network (e.g. OTT server) be equipped with an analyser to classify the data into different classes. The classes may be identified with attributes, and to at least one attribute there may be at least one KPI attached. The OTT server may receive and / or send the result of data analytics to the network function for verification.

[0584] In some implementations, the analysed dataset may be stored in a network repository (e.g. in UDR) or it may be exposed to external servers (e.g. interacting via NEF). The analysed dataset may be associated with one or more attributes.

[0585] In some examples, the NEF may enable the external server to query the network storage for data with certain attributes. For example, a query may request the network to provide

[0586] FH260314PCT-2026106843. DOCXdata that are representative of a certain area (e.g. within a certain group of cells, RAN area, or paging area), data that represent certain characteristics (e.g. NLOS heavy).

[0587] A UE may receive a configuration about one or more attributes and their values for which the model at the UE is valid. A UE may indicate the validity of the model or that the model is no longer valid, when one or more attributes are no longer valid. This may be indicated by updating the UAI or ProvideCapabilityMessage, indicating the change in applicability of the ML model used by the UE.

[0588] An AF or a NEF may be able to subscribe to analytics corresponding to certain parameters, this may include an attribute, UE-ID, vendor-ID, area ID, model ID and so forth. Likewise, it request to retrieve a dataset corresponding to an attribute, or a UE-ID or a vendor ID, arealD or modellD.

[0589] In the following, signaling aspects according to particular embodiments are described.

[0590] The idea is to get a key / token from the network, when a model has passed certain conformance tests in the repository or at the UE. The UE reports the key to the network, (for example) when registering its applicable capabilities at the network. With the provided key, the network can check whether this model can be activated at the network or if the model needs to be further tested (with additional data) etc.

[0591] This can lead to the following steps:

[0592] • UE reporting its capabilities, where the capabilities enabled by ML are additionally indicated by a key or a token indicating the version of the ML model used by the UE.

[0593] • The RAN node checks whether the functionality can be activated by the UE, for the given applicability conditions (current state of the environment).

[0594] • The NW may provide further test data, and request results from the UE.

[0595] • The UE performs inference on the provided test data and report results.

[0596] • The NW activates the model and / or provides a new key / token to be registered with the model.

[0597] Saving of fine-tuning information at NW side may, e.g., comprise that fine-tuned model information may be UE-context information, and may be stored in NG-RAN node, AMF, UDM, UDR.

[0598] FH260314PCT-2026106843. DOCXAlthough some aspects of the described concept have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or a device corresponds to a method step or a feature of a method step. Analogously, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus.

[0599] Various elements and features of the present invention may be implemented in hardware using analog and / or digital circuits, in software, through the execution of instructions by one or more general purpose or special-purpose processors, or as a combination of hardware and software. For example, embodiments of the present invention may be implemented in the environment of a computer system or another processing system. Fig.

[0600] 11 illustrates an example of a computer system 600. The units or modules as well as the steps of the methods performed by these units may execute on one or more computer systems 600. The computer system 600 includes one or more processors 602, like a special purpose or a general-purpose digital signal processor. The processor 602 is connected to a communication infrastructure 604, like a bus or a network. The computer system 600 includes a main memory 606, e.g., a random-access memory, RAM, and a secondary memory 608, e.g., a hard disk drive and / or a removable storage drive. The secondary memory 608 may allow computer programs or other instructions to be loaded into the computer system 600. The computer system 600 may further include a communications interface 610 to allow software and data to be transferred between computer system 600 and external devices. The communication may be in the from electronic, electromagnetic, optical, or other signals capable of being handled by a communications interface. The communication may use a wire or a cable, fiber optics, a phone line, a cellular phone link, an RF link and other communications channels 612.

[0601] The terms “computer program medium” and “computer readable medium” are used to generally refer to tangible storage media such as removable storage units or a hard disk installed in a hard disk drive. These computer program products are means for providing software to the computer system 600. The computer programs, also referred to as computer control logic, are stored in main memory 606 and / or secondary memory 608. Computer programs may also be received via the communications interface 610. The computer program, when executed, enables the computer system 600 to implement the present invention. In particular, the computer program, when executed, enables processor 602 to implement the processes of the present invention, such as any of the methods described herein. Accordingly, such a computer program may represent a controller of the computer system 600. Where the disclosure is implemented using software, the software

[0602] FH260314PCT-2026106843. DOCXmay be stored in a computer program product and loaded into computer system 600 using a removable storage drive, an interface, like communications interface 610.

[0603] The implementation in hardware or in software may be performed using a digital storage medium, for example cloud storage, a floppy disk, a DVD, a Blue-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate or are capable of cooperating with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium may be computer readable.

[0604] A storage medium for a machine learning model can hold the final trained model and / or critical components required for the model execution. This includes the model architecture, which defines the structure of the model, such as the number and type of layers and how these layers are connected. The storage medium can contain the weights, which are the learned parameters from the training process. These weights determine how the model processes input data to generate predictions. Additionally, hyperparameters like learning rate, and number of epochs can be stored to ensure the model can be retrained or fine-tuned under the same conditions. Additionally, input and output formats are included as well as preprocessing and postprocessing instructions to standardize data before and after it passes through the model.

[0605] Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.

[0606] Generally, embodiments of the present invention may be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer. The program code may for example be stored on a machine readable carrier.

[0607] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier. In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0608] FH260314PCT-2026106843. DOCXA further embodiment of the inventive methods is, therefore, a data carrier or a digital storage medium, or a computer-readable medium comprising, recorded thereon, the computer program for performing one of the methods described herein. A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals may for example be configured to be transferred via a data communication connection, for example via the Internet. A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein. A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0609] In some embodiments, a programmable logic device, for example a field programmable gate array, may be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware apparatus.

[0610] The above described embodiments are merely illustrative for the principles of the present invention. It is understood that modifications and variations of the arrangements and the details described herein are apparent to others skilled in the art. It is the intent, therefore, to be limited only by the scope of the impending patent claims and not by the specific details presented by way of description and explanation of the embodiments herein.

[0611] FH260314PCT-2026106843. DOCXABBREVIATIONS

[0612] Abbreviation Definition

[0613] 3GPP third generation partnership project 5GC 5G core network

[0614] BS base station

[0615] CHO conditional handover

[0616] CSI-RS channel state information reference signal DMRS demodulation reference signal

[0617] DOA direction of arrival

[0618] E-CID enhanced cell ID

[0619] eNB evolved node b

[0620] E-SMLC evolved serving mobile location center.

[0621] E-UTRA evolved UMTS terrestrial radio access gNB next generation node-b

[0622] GPS Global Positioning System

[0623] HO handover

[0624] LMF location management function

[0625] LMU location measurement unit

[0626] LPP LTE positioning protocol

[0627] LTE Long-term evolution

[0628] NG next generation

[0629] ng-eNB next generation eNB

[0630] NG-RAN either a gNB or an ng-eNB

[0631] NR new radio

[0632] NRPPa new radio positioning protocol a OTDOA observe time difference of arrival

[0633] PRS positioning reference signal

[0634] PTRS phase tracking reference signal

[0635] RLF radio link failure

[0636] QCL quasi colocation

[0637]

[0638] RAN radio access network

[0639] FH260314PCT-2026106843. DOCXRP reception point

[0640] RSTD reference signal time difference

[0641] RTOA relative time of arrival

[0642] RTT round trip time

[0643] SA Standalone

[0644] SRS sounding reference signal

[0645] TDM Time Domain Multiplexing

[0646] TOF time of flight

[0647] TRP transmission reception point

[0648] RS reference signal

[0649] QCL quasi co-located

[0650] AoA Angle of Arrival

[0651] AoD Angle of Departure

[0652] PAS Power Angular Spectrum

[0653] NR New Radio

[0654] gNB next generation node-b

[0655] GPS Global Positioning System

[0656] LMF location management function

[0657] LMU location measurement unit

[0658] LPP LTE positioning protocol

[0659] LTE Long-term evolution

[0660] NG next generation

[0661] ng-eNB next generation eNB

[0662] NG-RAN either a gNB or an ng-eNB

[0663] NR new radio

[0664] CA carrier aggregation

[0665] CAM Cooperative Awareness Message

[0666] DAS distributed antenna systems

[0667] DL Downlink

[0668] FL Frequency layer

[0669] FC Frequency component. This is either a BWP of a wideband carrier or GNSS Global navigation satellite system

[0670]

[0671] OOC Out-Of-Coverage

[0672] FH260314PCT-2026106843. DOCXPSFCH Physical Sidelink Feedback Channel

[0673] P-UE Pedestrian UE: should not be limited to pedestrians, but represents any UE RS Reference signal

[0674] RE resource elements

[0675] SINR Signal to interference and noise ratio

[0676] SL Sidelink

[0677] SPRS, SP- Sidelink positioning reference signals

[0678] V2X Vehicle to anything

[0679] VRU Vulnerable road user

[0680] V-UE Vehicular UE

[0681] BWP Bandwidth Part

[0682] TEG Timing Error Group

[0683] ZC Zadoff-Chu sequence

[0684] UE User equipment

[0685] UL Uplink

[0686] Uu (interface) Interface between UE

[0687] ToA Time of Arrival

[0688] TDOA Time Difference of Arrival

[0689] LOS Line of sight

[0690] PRU Positioning reference unit

[0691] ToF Time of flight

[0692] CSI Channel state information

[0693] CQI Channel quality indicator

[0694] PMI Precoding matric indicator

[0695] CRI CSI-RS resource indicator

[0696] SSBRI SS / PBCH resource block indicator

[0697] LI Layer indicator

[0698]

[0699] RI Rank indicator

[0700] FH260314PCT-2026106843. DOCX

Claims

Claims1. An apparatus for validating at least one AI / ML model or functionality,wherein the apparatus comprises a validation module (120) configured to conduct an evaluation by evaluating a performance and / or an applicability of the at least one AI / ML model or functionality depending on classes and / or attributes and / or performance criteria; and configured to generate a validation output indicating whether the at least one AI / ML model or functionality meet one or more predefined criteria depending on the evaluation.

2. An apparatus according to claim 1,wherein the classes and / or attributes and / or performance criteria are stored in a memory, wherein the validation module (120) is configured to access the classes and / or attributes and / or performance criteria during the evaluation.

3. An apparatus according to claim 1 or 2,wherein the apparatus is configured to obtain and / or process a configuration comprising the classes and / or attributes and / or performance criteria associated with a dataset.

4. An apparatus according to claim 3,wherein the apparatus comprises a configuration interface (110) configured to receive the configuration.

5. An apparatus according to claim 3 or 4,wherein the configuration comprises attributes related to the input data of the at least one AI / ML model or functionality; and / orwherein the configuration comprises attributes resulting from the processing of data used as input for the at least one AI / ML model or functionality; and / orwherein the configuration comprises attributes not utilized by the AI / ML functionality / model but characterizing the state of the environment.FH260314PCT-2026106843. DOCX6. An apparatus according to one of the preceding claims,wherein the validation module (120) is configured to generate the validation output such that the validation output comprises:an output / prediction for each test measurement of one of all AI / ML models supporting a specified functionality; and / orat least one performance indicator per dataset or per class; and / ora pass / fail indicator per dataset or per class depending on a performance threshold; and / ora single performance indicator for a plurality of datasets depending on how frequently a specific state appears and / or depending on a number of datapoints in each class.

7. An apparatus according to claim 6,wherein the validation module (120) is configured to generate the validation output such that the validation output comprises at least one performance indicator, wherein the at least one performance indicator comprises information on AI / ML beam management and / or information on AI / ML direct positioning and / or information on AI / ML assisted positioning and / or information on CSI measurement and reporting and / or information on mobility.

8. An apparatus according to claim 7,wherein the at least one performance indicator comprises the information on AI / ML beam management and indicates:a beam prediction accuracy by comparing the prediction results and a beam measurements from a resource set / resources; and / orL1-RSRP difference information depending on an actual measurement of the L1- RSRP of one or more of Top K predicted beam, and L1-RSRP measurements from a resource set / resources for monitoring; and / orFH260314PCT-2026106843. DOCXRSRP difference information between the predicted RSRP and measured L1- RSRP of corresponding beam(s) of a resource set / resources for monitoring; and / ora model prediction uncertainty / confidence; and / ora pass / fail test, indicating performance achieved compared to a configured threshold.

9. An apparatus according to claim 8,wherein the validation module (120) is configured to generate the validation output depending on an associated margin wherein the margin represents a specified tolerance range for acceptable prediction deviation.

10. An apparatus according to one of claims 7 to 9,wherein the at least one performance indicator comprises the information on AI / ML direct positioning and indicates:a mean squared error between a predicted position and a label; and / oran absolute average or maximum error between a predicted position and a label; and / ora model prediction uncertainty or confidence, and / ora pass / fail test indicating performance achieved compared to a configured threshold.

11. An apparatus according to one of claims 7 to 10,wherein the at least one performance indicator comprises the information on AI / ML assisted positioning and indicates:a classification metrics for Line-of-sight and / or Non-line-of-sight, for example, an accuracy or an F1-score; and / orFH260314PCT-2026106843. DOCXa model prediction uncertainty and / or a model prediction confidence; and / ora pass / fail test indicating performance achieved compared to a configured threshold.

12. An apparatus according to one of claims 7 to 11,wherein the at least one performance indicator comprises the information on CSI measurement and reporting and indicates:a difference between a predicted CQI and a measured CQI for a given test configuration, and / or a difference between a predicted CQI and expected CQI for a given test configuration; and / ora difference between a predicted PMI and a UE-reported PMI report for a given test configuration, and / or a difference between a predicted PMI and an expected PMI report for a given test configuration; and / ora difference between a predicted RI and a UE-reported RI report for a given test configuration, and / or a difference between a predicted RI and an expected RI report for a given test configuration; and / ora stability of one or more predicted CSI components under static scenarios, and / or a latency of one or more predicted CSI components under mobility scenarios; and / ora performance improvement and / or performance degradation, e.g., depending on a difference between a predicted CSI component and a measured CSI component or depending on a difference between a predicted component and a gNB-selected component.

13. An apparatus according to one of claims 7 to 12,wherein the at least one performance indicator comprises the information on mobility and indicates an L1-RSRP prediction accuracy and / or an L3-RSRP prediction accuracy.

14. An apparatus according to one of the preceding claims,FH260314PCT-2026106843. DOCXwherein the validation module (120) is configured to mark the at least one AI / ML model or functionality as validated or not per dataset or per class depending on the evaluation.

15. An apparatus according to one of the preceding claims,wherein the validation module (120) is configured to generate the validation output, such that:the validation output comprises a pass / fail result for the functionality / model to be applied per dataset (class); and / orthe validation output comprises information on a gap between a performance threshold and an achieved performance indicator; and / orthe validation output comprises a signal to an AI / ML model or functionality monitoring and management framework, informing on a gap between the performance threshold and an achieved performance indicator.

16. An apparatus according to claim 15,wherein the validation module (120) is configured to mark the at least one AI / ML model or functionality as validated per dataset or class for a specific performance level, depending on the evaluation.

17. An apparatus according to one of the preceding claims,wherein the validation module (120) is configured to detect clusters from a dataset depending on attributes data indicating data on the attributes.

18. An apparatus according to claim 17,wherein the validation module (120) is configured to construct the dataset for the at least one AI / ML model or functionality byconstructing a lower-dimensional space utilizing the attributes data,FH260314PCT-2026106843. DOCXdetecting clusters in the lower-dimensional space,selecting, from each of the clusters, a number of samples that are used as a test set.

19. An apparatus according to claim 18,wherein the apparatus is configured to record the dataset with samples selected to represent the environment.

20. An apparatus according to one of claims 17 to 19,wherein the validation module (120) is configured to evaluate the performance of the AI / ML model or functionality in different states of a cell or of an area.

21. An apparatus according to one of claims 17 to 20,wherein the validation module (120) is configured to provide samples for areas where at least two of the clusters overlap.

22. An apparatus according to one of claims 17 to 21,wherein the validation module (120) is configured to detect the clusters by employing an unsupervised learning method or by employing a weak or semisupervision learning method or by employing a prototypical network.

23. An apparatus according to one of claims 17 to 22,wherein the validation module (120) is configured to employ a clustering or classification algorithm for determininga number of different states an environment, for example, an area or a cell, can be in,a number of data points which represent one of the different states,FH260314PCT-2026106843. DOCXa change of an environment comprising a situation where a state, e.g. being represented by a class, becomes obsolete, or when a new state, e.g. being represented by a class, is formed.

24. An apparatus according to one of the preceding claims,wherein the apparatus is configured to obtain representative data from one or more or all possible states of an environment, when receiving a new AI / ML model or functionality, andwherein the validation module (120) is configured to generate the validation output for the new AI / ML model or functionality.

25. An apparatus according to claim 24,wherein the validation module (120) is configured to value the AI / ML model or functionality as either suitable or unsuitable for a specific area and / or a specific cell and / or a specific zone and / or for a sub-set of states.

26. An apparatus according to one of the preceding claims,wherein the at least one AI / ML model or functionality is an AI / ML model or functionality for a wireless communication system.

27. An apparatus according to one of the preceding claims:wherein the attributes comprise one or more of:an RSRP measurement pattern of an initial set of beams,a LOS / NLOS indication,an entering of specified sub-areas in area / cell,SNR / SINR measurements,Doppler effect measurements,FH260314PCT-2026106843. DOCXinformation on whether a user is transitioning from indoor to outdoor environment and vice-versa,a network load,one or more BLER values,information on HARQ retransmissions,information on an interference detected by a UE.

28. An apparatus according to one of the preceding claims:wherein the attributes comprise a reinforcement learning or bandits policy existence and magnitude of uncertainty.

29. An apparatus according to one of the preceding claims,wherein the validation module (120) is configured to detect an existence and / or a magnitude of uncertainty in a reinforcement learning or bandits policyby detecting a deviation of ensemble predictions for ensemble policies; and / orby conducting approximate counts on observation-action pairs in the data, indicating that some of the actions have not been selected often enough for some input observations; and / orby conducting a method that learns a next-observation-predictor, which can estimate the policy uncertainty via calculating a prediction error on the observations; and / orby conducting a method, where a neural network is trained to predict the features of the observations, which are generated by a fixed random network, for example, by conducting a random-network distillation method.

30. An apparatus according to one of the preceding claims,FH260314PCT-2026106843. DOCXwherein the apparatus is configured to monitor one or more events associated with the at least one AI / ML model or functionality.

31. An apparatus according to one of claims 1 to 29,wherein the validation module (120) is configured to determine if an update of the classes is to be conducted by monitoring one or more events associated with the at least one AI / ML model or functionality.

32. An apparatus according to claim 30 or 31,wherein the one or more events comprise information on AI / ML Beam Management, and / or information on positioning and / or information on CSI measurement and reporting and / or information on mobility.

33. An apparatus according to claim 32,wherein the one or more events comprise information on AI / ML Beam Management and comprise an indication indicatingthat a best beam for moving UE does not change smoothly, and / orthat a measured RSRP value is lower than a first threshold value, and / orthat a QoS value is lower than a second threshold value, and / orthat a QoE value is lower than a third threshold value.

34. An apparatus according to claim 32 or 33,wherein the one or more events comprise information on positioning and comprise an indication indicatingthat a label-based or label-free monitoring indicates model quality below a threshold,a heavy NLOS detection.FH260314PCT-2026106843. DOCX35. An apparatus according to one of claims 32 to 34,wherein the one or more events comprise information on mobility and comprise an indication indicatingan increased number of failures, for example, as indicated by MDT, and / orping pong handover effects.

36. An apparatus according to one of claims 32 to 35,wherein the one or more events comprise information on reinforcement learningbased beam management and / or mobility and comprise an indication indicatinga high regret or an increased exploration, for example, for bandit algorithms, ora large variance of rewards per episode or increased level of surprise for exploration-directed algorithms, and / orthat a QoS value is lower than a threshold value, and / ora QoE value is lower than another threshold value.

37. An apparatus according to one of claims 32 to 36,wherein the one or more events comprise information on CSI measurement and reporting and comprise an indication indicatingan increased number of BLER detected by the UE; and / ora change in Rl, CQI values, which do not change smoothly, for example, due to a sudden change in mobility parameters or a sudden blockage or a partial blockages; and / ora detection of interference from other UEs or other network entities; and / ora change in polarization, for example, due to rotation.FH260314PCT-2026106843. DOCX38. An apparatus according to one of claims 32 to 37,wherein the one or more events comprise information on mobility and comprise an indication indicatinga mismatch between a predicted L1-RSRP and a measured L1-RSRP; and / ora mismatch between a predicted L3-RSRP and a measured L3-RSRP; and / ora mismatch between a predicted handover failure and an actual occurrence of radio link failure; and / ora time-mismatch between a predicted handover failure and an actual occurrence of a radio link failure, and / ora mismatch between a predicted measurement event and a true measurement event occurring at a UE; and / ora time-mismatch between a predicted measurement event and a true measurement event occurring at the UE.

39. An apparatus according to claim 38,wherein the one or more events comprisea measured RSRP and / or a predicted RSRP improves above a threshold; and / ora measured RSRP and / or a predicted RSRP falls below a threshold; and / ora measured RSRP and / or a predicted RSRP raises above a first threshold, but remains below a second threshold; and / ora measured RSRP and / or a predicted RSRP falls below a third threshold but remains above a fourth threshold; and / oran RLF occurs earlier than a threshold value for a predicted RLF time.

40. An apparatus according to one of the preceding claims,FH260314PCT-2026106843. DOCXwherein the network entity is configured to update the AI / ML model or functionality depending on the validation output to obtain an updated AI / ML model or functionality.

41. An apparatus according to one of the preceding claims,wherein the AI / ML model or functionality is a neural network.

42. An apparatus according to claim 41,wherein the neural network is a CNN or is a LSTM or is a Bayesian NN or comprises transformer layers.

43. An apparatus according to one of the preceding claims,wherein the AI / ML model or functionality has been trained using reinforcement learning.

44. An apparatus according to one of the preceding claims,wherein the AI / ML model or functionality has been trained using a loss function.

45. An apparatus according to claim 44,wherein the loss function is implemented as a MSE, or as a top-k accuracy, or as a F1-score.

46. An apparatus according to one of the preceding claims,wherein at least one of the classes is a dataset comprising input and / or output data and / or one or more performance metrics and / or one or more performance thresholds, for example, wherein the dataset corresponds to a substate.

47. An apparatus according to one of the preceding claims,wherein each of the classes and / or each of a plurality of datasets is associated with an Associated ID, wherein the Associated ID identifies one or more networkFH260314PCT-2026106843. DOCXconfigurations or conditions under which training data for the at least one AI / ML model or functionality were collected.

48. An apparatus according to claim 47,wherein the Associated ID does not reveal information on the actual network implementation.

49. An apparatus according to one of the preceding claims,wherein the apparatus is an apparatus of a wireless communication system.

50. An apparatus according to one of the preceding claims,wherein the apparatus is a network entity of a wireless communication system.

51. An apparatus according to claim 50,wherein an OTT-server executes the AI / ML model or functionality with a provided dataset, obtains an output of the AI / ML model and sends the output of the AI / ML model back to the network entity for validation,wherein the network entity is configured to receive the output of the AI / ML model from the OTT-server, andwherein the validation module (120) of the network entity is configured to conduct the evaluation to generate the validation output.

52. An apparatus according to claim 50,wherein the network entity is configured to receive the AI / ML model or functionality, being a trained model or functionality, from an OTT-server,wherein the network entity is configured to execute the AI / ML model or functionality,FH260314PCT-2026106843. DOCXwherein the validation module (120) of the network entity is configured to generate the validation output depending on the executing of the AI / ML model or functionality.

53. An apparatus according to claim 52,wherein the network entity is configured to determine network analytics and / or whether the model has been verified / validated.

54. An apparatus according to claim 52 or 53,wherein the network entity is configured to indicate a successful verification or validation of the AI / ML model or functionality to the OTT-server and / or is configured to indicate one or more applicability conditions of the AI / ML model or functionality to the OTT-server.

55. An apparatus according to one of claims 50 to 54,wherein the network entity is configured to update the AI / ML model or functionality depending on the validation output to obtain an updated AI / ML model or functionality.

56. An apparatus according to claim 55,wherein the network entity is configured to transmit the updated AI / ML model or functionality to a UE of the wireless communication system.

57. An apparatus according to one of claims 1 to 49,wherein the apparatus is an OTT-server.

58. An apparatus according to claim 57,wherein the OTT-server is configured to receive input data for the AI / ML model or functionality and an expected output of the AI / ML model or functionality from a network entity of the wireless communication system or from a user equipment of the wireless communication system,FH260314PCT-2026106843. DOCXwherein the validation module (120) of the OTT-server is configured to execute the AI / ML model or functionality with the input data, is configured to compare output data of the AI / ML model or functionality with the expected output, and is configured to generate the validation output depending on the output data and the expected output.

59. An apparatus according to claim 58,wherein the OTT-server is configured to receive the input data from the network entity, andwherein the OTT-server is configured to transmit the validation output to the network entity or to another network entity of a wireless communication system.

60. An apparatus according to claim 58,wherein the OTT-server is configured to receive the input data from the user equipment, andwherein the OTT-server is configured to transmit the validation output to the user equipment or to a network entity of a wireless communication system.

61. An apparatus according to claim 57,wherein the OTT-server is configured to receive update information on an update of the AI / ML model or functionality from a UE of a wireless communication system, andwherein the OTT-server is configured to receive a dataset from a network entity of the wireless communication system,wherein the validation module (120) of the OTT-server is configured to validate the model using the dataset from the network entity depending on the update information from the UE to generate the validation output.

62. An apparatus according to claim 61,wherein the UE trains and / or fine-tunes the AI / ML model or functionality,FH260314PCT-2026106843. DOCXwherein the OTT-server is configured to receive the AI / ML model or functionality or at least one model parameter of the AI / ML model or functionality from the UE.

63. An apparatus according to one of claims 57 to 62,wherein the OTT-server is configured to request validation data for validating the AI / ML model or functionality from the network entity, andwherein the OTT-server is configured to receive the validation data from the network entity.

64. An apparatus according to one of claims 57 to 63,wherein the OTT-server is configured to update the AI / ML model or functionality depending on the validation output to obtain an updated AI / ML model or functionality.

65. An apparatus according to claim 64,wherein the OTT server is configured to transmit the updated AI / ML model or functionality to a UE of the wireless communication system.

66. An apparatus according to one of claims 1 to 49,wherein the apparatus is a UE of a wireless communication system.

67. An apparatus according to claim 66,wherein the validation module (120) of the apparatus is configured to execute the AI / ML model or functionality with input data, is configured to compare output data of the AI / ML model or functionality with an expected output, and is configured to generate the validation output depending on the output data and the expected output.

68. An apparatus according to claim 67,FH260314PCT-2026106843. DOCXwherein the apparatus is configured to receive the input data for the AI / ML model or functionality and the expected output of the AI / ML model or functionality from a network entity of the wireless communication system.

69. An apparatus of a wireless communication system, wherein a configuration, being obtained by the apparatus and / or being stored in the apparatus, comprises a plurality of datasets, wherein each of the plurality of datasets comprises one or more attributes and / or one or more classes; wherein the one or more attributes and / or the one or more classes depend on a property and / or a state of the wireless communication system and / or depend on a property and / or a state of a user equipment and / or a network entity of the wireless communication system,wherein the apparatus comprises a processor; wherein, for each dataset of one or more datasets of the plurality of datasets, the processor is configured to execute at least one AI / ML model or functionality with model input data, which depends on said dataset, to obtain model output data for each of the one or more datasets.

70. An apparatus according to claim 69,wherein the plurality of datasets comprise information on and / or depend on one or more of the following:one or more measurement configurations,one or more measurements,one or more predictions,aggregated data,post-processed data,time stamps,correlations with other measurements,one or more input data,one or more output data,relation of input and output data,functional behavior between input and output.

71. An apparatus according to one of claims 69 or 70,wherein the plurality of datasets comprise information on and / or depend on one or more of the following:FH260314PCT-2026106843. DOCXaggregation of data,post-processing of measurement data,time stamping of measurements,correlations with other measurements,a relation of input and output data,functional behavior between input and output,logging, e.g., logging of measurement data and / or prediction data,reporting, e.g., reporting of measurement data and / or prediction data.

72. An apparatus according to one of claims 69 to 71,wherein each of the one or more datasets is associated with secondary data, wherein the second data comprises location data and / or timing data and / or UE or network entity ID data.

73. An apparatus according to claim 72,wherein the secondary data comprises location data indicating a position of a UE or a network entity in an environment; and / orwherein the secondary data comprises timing data indicates when the information in the one or more attributes and / or in the one or more classes of a dataset has been obtained; and / orwherein the secondary data comprises UE or network entity data indicating the UE or on the network entity to which information in the dataset relates.

74. An apparatus according to claim 72 or 73,wherein each of the one or more datasets comprises the secondary data.

75. An apparatus according to one of claims 69 to 74,wherein a dataset of the one or more datasets further comprises a performance criteria and / or ground truth information, wherein the apparatus is configured to determine a performance result for the dataset depending on the model outputFH260314PCT-2026106843. DOCXdata of the dataset and depending on the performance criteria and / or the ground truth information of the dataset.

76. An apparatus according to one of claims 69 to 75,wherein the apparatus comprises an interface, wherein the interface is configured to receive at least one of the one or more datasets; and / orwherein the interface is configured to output and / or transmit model output information comprising the model output data of the one or more datasets and / or depending on the model output data of the one or more datasets.

77. An apparatus according to claim 76,wherein the interface is configured to receive said at least one of the one or more datasets from a network entity of the wireless communication system or from an OTT-server.

78. An apparatus according to claim 76 or 77,wherein the interface is configured to transmit the model output information to a network entity of the wireless communication system or to an OTT-server.

79. An apparatus according to claim 78wherein the interface is configured to transmit the model output information to said network entity of the wireless communication system or to said OTT-server for validating the AI / ML model or functionality.

80. An apparatus according to claim 79,wherein said network entity or said OTT-server is an apparatus according to one of claims 1 to 65.

81. An apparatus according to one of claims 69 to 80,wherein the apparatus is configured to validate the AI / ML model or functionality using the model output data of the one or more datasets.FH260314PCT-2026106843. DOCX82. An apparatus according to claims 71,wherein the apparatus implements an apparatus according to one of claims 1 to 68.

83. An apparatus according to one of claims 69 to 82,wherein the apparatus is a user equipment.

84. An apparatus according to one of claims 69 to 83,wherein the apparatus is configured to conduct filtering based on the attributes and / or performance criteria, e.g., providing output based on the classes of the configuration.

85. An apparatus according to claim 84,wherein the apparatus is configured to filter measurement data and / or prediction data depending on the attributes and / or depending on the performance criteria.

86. An apparatus according to one of claims 69 to 85,wherein the apparatus is configured to train and / or to fine-tune the AI / ML model or functionality.

87. An apparatus according to one of claims 69 to 86, further depending on claim 78 or 79,wherein the interface of the apparatus is configured to receive the validation output from the network entity or from the OTT-server.

88. An apparatus according to claim 87,wherein the apparatus is configured to update or fine-tune the AI / ML model or functionality depending on the validation output.FH260314PCT-2026106843. DOCX89. An apparatus according to one of claims 69 to 88, further depending on claim 78 or 79,wherein the interface of the apparatus is configured to receive an updated version of the AI / ML model or functionality from the network entity or from the OTT-server,wherein the apparatus is configured to employ the updated version of the AI / ML model or functionality.

90. An apparatus according to one of claims 50 to 56,wherein the network entity is configured to receive data from one or more other network entities of the wireless communication system,wherein the network entity is configured to analyse the data to obtain an analysing result and to transmit the analysing result of to a UE or to another network entity of the wireless communication system or to an entity outside of the wireless communication system.

91. An apparatus according to one of claims 50 to 56 or according to claim 90,wherein the network entity is configured to receive raw measurements from a UE or a group of UEs of the wireless communication system and / or from one or more RAN nodes of the wireless communication system, which pertain to a UE or a group of UEs of the wireless communication system,wherein the validation of the network entity is configured to generate the validation output depending on the raw measurements.

92. An apparatus according to one of claims 50 to 56 or according to claim 90 or 91,wherein the network entity is configured to group information from one or more other network entities of the wireless communication system, and is configured to link them together with an identifier, for example, the identifier being a dataset identifier, UE4dentifier, timestamp, analysis identifier or a combination thereof.

93. An apparatus according to claim 92,FH260314PCT-2026106843. DOCXwherein the information comprises ground truth labels, for example, position information and / or velocity information and / or handover failure information.

94. An apparatus according to claim 92 or 93,wherein the network entity, e.g. a NWDAF, is configured to obtain the information by subscribing to the one or more other network entities, e.g. a gNB and / or an AF and / or an LMF, or by subscribing to an intermediate network node, e.g. a LMF and / or O& M and / or a gNB and / or an AMF, which collects the information.

95. An apparatus according to one of claims 50 to 56 or according to one of claims 92 to 94,wherein the network entity is configured to perform a data analytics function to obtain analytics, for example, for detecting various patterns, such as a novelty detection and / or a change in performance and / or a change in attributes and / or a change in classes.

96. An apparatus according to claim 95,wherein the network entity is configured to provide the analytics to one or more internal network functions (e.g. LMF, PCF, AF) or one or more external functions, such as an OTT-server interacting with core network via NEF, orwherein the network entity is configured to provide the analytics to external servers, e.g. interacting via NEF.

97. An apparatus according to claim 96,wherein the network entity is configured to provide the analytics to external servers, interacting via the NEF,wherein the NEF enables an external server to query a network storage for data with certain attributes.

98. An apparatus according to claim 97,FH260314PCT-2026106843. DOCXwherein an AF or the NEF is able to subscribe to analytics corresponding to certain parameters, for example, comprising an attribute and / or a UE-ID and / or a vendor-ID and / or an area ID and / or a model ID, wherein the AF or the NEF requests to retrieve a dataset, for example, corresponding to the attribute and / or to the UE-ID and / or to the vendor ID and / or to the areaID and / or to the modelID.

99. An apparatus according to one of claims 95 to 98,wherein the network entity is configured to store the analytics may in a network repository, e.g. in UDR.

100. An apparatus according to one of claims 95 to 99,wherein an analysed dataset is associated with one or more attributes.

101. An apparatus according to one of claims 50 to 56 or according to one of claims 90 to 100,wherein the network entity is configured to provide a trigger based on events, for example unseen pattern detected or indicating to a network node, external node or a UE that the environment has changed, such that another network entity of the wireless communication system is informed to initiate a data collection procedure with a UE and / or a group of UEs and / or a RAN node and / or a group of RAN nodes of the wireless communication system.

102. An apparatus according to one of claims 50 to 56 or according to one of claims 90 to 101,wherein the network entity, e.g. a NWDAF, is configured to detect a change in the environment and is configured to indicate an event to a second network entity of the wireless communication system, e.g. gNB, so that the second network entity is informed to initiate a data collection procedure from one or more UEs of the wireless communication system, e.g. retrieving logs stored by the one or more UEs, for example, the second network entity indicates to a group of UEs via system information broadcast or paging that the second network entity is requesting certain information to be reported.

103. An apparatus according to one of claims 57 to 65,FH260314PCT-2026106843. DOCXwherein the OTT-server is configured to classify data into different classes, wherein the different classes are identified by attributes, wherein a performance indicator is associated with at least one if the attributes,wherein the OTT-server is configured to receive and / or to transmit a result of data analytics to a network entity of the wireless communication system for verification.

104. An apparatus according to one of the preceding claims,wherein the apparatus is configured to receive information on one or more attributes and their values for which the AI / ML model or functionality at the UE is valid.

105. An apparatus according to claim 104,wherein the apparatus is configured to indicate a validity of the AI / ML model or functionality or that the AI / ML model or functionality is no longer valid, when one or more attributes are no longer valid.

106. An apparatus according to claim 105,wherein the apparatus is configured to indicate the validity of the AI / ML model or functionality or that the AI / ML model or functionality is no longer valid by updating a UAI or by using a ProvideCapabilityMessage indicating the change in applicability of the AI / ML model or functionality used by the UE.

107. A system comprising:a network entity, anda user equipment,wherein the network entity or the user equipment implements an apparatus according to one of claims 1 to 49.

108. A system comprising:FH260314PCT-2026106843. DOCXa network entity,a user equipment, andan OTT-server,wherein the network entity or the user equipment or the OTT-server implements an apparatus according to one of claims 1 to 49.

109. A system according to claim 107 or 108,wherein the UE is configured to report a key or a token to the network entity, for example when registering its applicable capabilities at the network, andthe network is configured to check, depending on the key or the token, whether the AI / ML model or functionality can be activated at the network or if the AI / ML model or functionality cannot be activated, e.g., needs to be further tested, for example, with additional data.

110. A system according to one of claims 107 to 109,wherein the UE is configured to report its capabilities,wherein a RAN node, being implemented at the network entity or at another network entity, is configured to check whether the AI / ML model or functionality can be activated by the UE for given applicability conditions, for example, for a current state of the environment.

111. A system according to claim 110,wherein the network is configured to provide further test data, and is configured to request results from the UE,wherein the UE is configured to perform inference on the further test data and is configured to report the results,FH260314PCT-2026106843. DOCXwherein the network entity is configured to activate the AI / ML model or functionality and / or is configured to provide a new key or a new token to be registered with the AI / ML model or functionality.

112. A system according to one of claims 107 to 111,wherein the network entity is configured to store fine-tuned model information, e.g. UE-context information, in an NG-RAN node and / or in an AMF and / or in an UDM and / or in an UDR.

113. A system comprising:an apparatus according to one of claims 69 to 89, andan apparatus for validating an AI / ML model or functionality,wherein the apparatus according to one of claims 67 to 87 is configured to provide the model output data of the one or more datasets or model information data depending on the model output data of the one or more datasets to the apparatus for validating the AI / ML model or functionality, andwherein the apparatus for validating the AI / ML model or functionality uses the model output data of the one or more datasets or the model information data for validating the AI / ML model or functionality.

114. A system according to claim 113,wherein the apparatus for validating the AI / ML model or functionality is an apparatus according to one of claims 1 to 68.

115. A method for validating at least one AI / ML model or functionality, wherein the method comprises:conducting an evaluation by evaluating a performance and / or an applicability of the at least one AI / ML model or functionality depending on classes and / or attributes and / or performance criteria; andFH260314PCT-2026106843. DOCXgenerating a validation output indicating whether the at least one AI / ML model or functionality meet one or more predefined criteria depending on the evaluation.

116. A method for a wireless communication system, wherein a configuration, being obtained by an apparatus and / or being stored in the apparatus, comprises a plurality of datasets, wherein each of the plurality of datasets comprises one or more attributes and / or one or more classes; wherein the one or more attributes and / or the one or more classes depend on a property and / or a state of the wireless communication system and / or depend on a property and / or a state of a user equipment and / or a network entity of the wireless communication system,wherein the apparatus comprises a processor; wherein, for each dataset of one or more datasets of the plurality of datasets, the processor executes at least one AI / ML model or functionality with model input data, which depends on said dataset, to obtain model output data for each of the one or more datasets.

117. A computer program for implementing the method of claim 115 or 116, when being executed on a computer or signal processor.FH260314PCT-2026106843. DOCX