Model performance prediction device and method in communication network supporting AI / ML
By evaluating the metrics and activation overhead of AI/ML models and optimizing their usage, the problem of improper model switching in AI/ML-supported communication networks was resolved, resulting in more efficient network performance and resource utilization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-13
- Publication Date
- 2026-04-07
AI Technical Summary
In AI/ML-supported communication networks, existing technologies struggle to effectively manage and switch between different models to meet performance requirements under varying regions and conditions, leading to network performance instability and resource waste.
An apparatus and method for a wireless communication system are provided, which determine whether to activate or switch an AI/ML model by identifying and evaluating metrics of the AI/ML model, taking into account benefits and activation costs, in order to optimize the use of the model.
It improves the stability of network performance and the efficiency of resource utilization, ensures efficient switching and activation of models under different regions and conditions, and reduces resource waste.
Smart Images

Figure CN121816728A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of wireless communication systems or networks, in particular to a communication network supporting AI / ML, and more specifically to an apparatus and method for model performance prediction in a communication network supporting AI / ML. BACKGROUND
[0002] Figure 15 is a schematic representation of an example of a terrestrial wireless network 100, as shown in Figure 15(a), the terrestrial wireless network 100 comprising a core network and one or more radio access networks RAN1, RAN2,... RAN N (RAN = Radio Access Network). Figure 15(b) is a schematic representation of an example of a radio access network RAN n (RAN = Radio Access Network). Figure 15(b) is a schematic representation of an example of a radio access network RAN n may comprise one or more base stations gNB1 to gNB5 (gNB = next generation Node B), each serving a particular area surrounding the base station, schematically represented by the respective cells 1061 to 1065. The base stations are used to serve users within the cells. One or more base stations can serve users in licensed and / or unlicensed bands. The term base station BS refers to a gNB in 5G networks, to an eNB in UMTS / LTE / LTE-A / LTE-A Pro, and simply to a BS in other mobile communication standards. Users can be fixed devices or mobile devices. The wireless communication system can also be accessed by mobile or fixed IoT (Internet of Things) devices that connect to the base stations or to the users. The mobile devices or IoT devices can comprise physical devices, ground-based vehicles (robots or cars), aerial vehicles (such as manned aircrafts or unmanned aircraft vehicles (UAV), the latter also known as drones), buildings and other items of this nature having embedded therein electronics, software, sensors, actuators, etc., and connectivity to a network enabling them to collect and exchange data with each other or with other devices in the network. Figure 15(b) shows an exemplary view of five cells, but the RAN n may comprise more or less such cells, and the RAN nIt is also possible to include only one base station. Figure 15(b) shows two users UE1 and UE2 (UE = User Equipment), also known as user devices UE, in cell 1062 served by base station gNB2. Another user UE3 is shown in cell 1064 served by base station gNB4. Arrows 1081, 1082 and 1083 schematically represent uplink / downlink connections for transmitting data from users UE1, UE2 and UE3 to base stations gNB2, gNB4 or for transmitting data from base stations gNB2, gNB4 to users UE1, UE2, UE3, which can be implemented in the licensed frequency band or the unlicensed frequency band. Furthermore, Figure 15(b) shows two loT devices 1101 and 1102 in cell 1064, which can be stationary or mobile devices. loT device 1101 accesses the wireless communication system via base station gNB4, as shown by arrow 1121 for receiving and transmitting data. loT device 1102 accesses the wireless communication system via user UE3, as schematically shown by arrow 1122. Each base station gNB1 to gNB5 can be connected to the core network 102 via respective backhaul links 1141 to 1145, e.g., via the S1 interface, as schematically indicated by the arrows pointing to "core" in Figure 15(b). The core network 102 can be connected to one or more external networks. An external network can be the Internet or a private network, such as an intranet or any other type of campus network, e.g., a private WiFi or 4G or 5G mobile communication system. Furthermore, some or all of the base stations gNB1 to gNB5 can be connected to each other, e.g., via the S1 or X2 interface or the XN interface in NR (New Radio), via respective backhaul links 1161 to 1165, as schematically indicated by the arrows pointing to gNB s in Figure 15(b). The sidelink channel allows direct communication between UEs (also known as D2D (Device to Device) communication). The sidelink interface in 3GPP (Third Generation Partnership Project) is named PC5 (Proximity-based Communication 5).
[0003] Data transmission can use a physical resource grid. The physical resource grid can comprise a grid of resource elements, each grid constituting a time-frequency "tile". Various physical channels and physical signals can be mapped on these resource elements. For example, physical channels can include physical downlink, uplink and sidelink shared channels: PDSCH (Physical Downlink Shared CHannel), PUSCH (Physical Uplink Shared Channel), PSSCH (Physical Sidelink Shared Channel), carrying user-specific data (also known as downlink, uplink and sidelink payload data); and physical broadcast channels PBCH (Physical Broadcast Channel) (e.g. carrying a master information block (MIB), and one or more system information blocks (SIBs), one or more sidelink information blocks (SLIBs), if supported), physical downlink, uplink and sidelink control channels: PDCCH (Physical Downlink Control Channel), PUCCH (Physical Uplink Control CHannel), PSCCH (Physical Sidelink Control Channel); and downlink control information (DCI), uplink control information (UCI) and sidelink control information (SCI), and physical sidelink feedback channel: PSFCH (Physical sidelink feedback channel) (carrying PC5 feedback responses). Note that the sidelink interface can support 2-stage SCI (speech call item). This refers to a first control region comprising certain parts of the SCI, and optionally a second control region comprising a second part of control information.
[0004] For the uplink, the physical channels can further include a physical random access channel, PRACH (Packet Random Access Channel) or RACH (Random Access Channel), for use by UEs to synchronize and access the network after obtaining the MIB and SIBs. The physical signals can include reference signals or symbols (RSs), synchronization signals, etc. The resource grid can include frames or radio frames having a certain duration (e.g., 10 milliseconds) in the time domain and a certain bandwidth in the frequency domain. The frames can have a certain number of subframes of predetermined length (e.g., 1 millisecond), each subframe including one or more slots of 12 or 14 OFDM symbols (Orthogonal Frequency-Division Multiplexing), the exact number depending on the cyclic prefix (CP) length. The frames can also include a smaller number of OFDM symbols, e.g., when employing shortened transmission time intervals, sTTIs (slot or subslot transmission time intervals), or a mini-slot / non-slot frame structure including only a few OFDM symbols.
[0005] The wireless communication system can be any single-tone or multi-carrier system using frequency division multiplexing, such as the orthogonal frequency-division multiplexing (OFDM) or orthogonal frequency-division multiple access (OFDMA), or other IFFT (Inverse Fast Fourier Transformation)-based signal, with or without CP, e.g. DFT-s-OFDM (DFT, Discrete Fourier Transformation)-spread-OFDM. Other waveforms, such as non-orthogonal waveforms like filter bank multicarrier (FBMC), generalized frequency division multiplexing (GFDM) or universal filtered multi-carrier (UFMC) can also be used. The wireless communication system can operate, for example, according to the LTE-Advanced pro standard, or the 5G or NR (New Radio) standards, or the NR-U (New Radio-Unlicensed) standard.
[0006] The wireless network or communication system depicted in Figure 15 can be a heterogeneous network with different overlay networks, e.g., a macro cell network, each macro cell including a macro base station, like base station gNB1 to gNB5, and a network of small base stations, like femto or pico base stations, not shown in Figure 15. In addition to the above-described terrestrial wireless networks, there are also non-terrestrial wireless communication networks (NTN) including spaceborne transceivers, like satellites, and / or airborne transceivers, like unmanned aircraft systems. Non-terrestrial wireless communication networks or systems can operate in a similar way as the terrestrial systems described with reference to Figure 15, e.g., following the LTE-Advanced Pro standard or the 5G or NR (new radio) standard.
[0007] In a mobile communication network, e.g., an LTE or 5G / NR network as described with reference to Figure 15, UEs can communicate directly via one or more sidelink (SL) channels, e.g., using a PC5 / PC3 interface or WiFi Direct. UEs communicating directly with each other via the sidelink can include vehicles communicating directly with other vehicles (V2V communication), vehicles communicating with other entities of the wireless communication network, e.g., roadside units (RSUs) or roadside entities like traffic lights, traffic signs, or pedestrians (V2X communication). Depending on the specific network configuration, RSUs can have the functionality of a BS or a UE. Other UEs can not be vehicle-related UEs and can include any of the devices described above. Such devices can also communicate directly with each other using the SL channel (D2D communication).
[0008] In a wireless communication network, like the one shown in Figure 15, it can be required to position a UE with a certain accuracy, e.g., to determine the position of the UE in a cell. There are multiple positioning methods known, like satellite-based positioning methods (like GPS and other autonomous and assisted global navigation satellite systems (A-GNSS)), mobile radio cellular positioning methods (like observed time difference of arrival (OTDOA) and enhanced cell ID (E-CID) or combinations thereof).
[0009] In the general AI / ML 3GPP SI, one of the envisioned ways to achieve generalization of AI / ML solutions is to enable switching between different models.
[0010] It would be highly valuable to have improved concepts for a communication network supporting AI / ML. SUMMARY
[0011] According to an embodiment, an apparatus of a wireless communication system is provided. The apparatus is configured to determine a metric of an AI / ML model and / or a function thereof out of one or more inactive AI / ML models, wherein the one or more inactive AI / ML models are suitable for supporting a task of a user equipment and / or a network entity of the wireless communication system, the apparatus being the user equipment or being different from the user equipment; wherein the apparatus is configured to determine the metric of the AI / ML model and / or the function thereof such that the metric considers a benefit of employing the AI / ML model and / or the function thereof and such that the metric considers an activation overhead for activating the AI / ML model and / or the function thereof. Further, the apparatus is configured to determine whether to activate the AI / ML model and / or the function thereof depending on the metric of the AI / ML model and / or the function thereof.
[0012] Further, according to another embodiment, an apparatus of a wireless communication system is provided. The apparatus is configured to activate an AI / ML model and / or a function thereof out of one or more AI / ML models; wherein the one or more AI / ML models are suitable for supporting a task of a user equipment and / or a network entity of the wireless communication system; wherein the apparatus is the user equipment or is different from the user equipment; wherein whether to activate the AI / ML model and / or the function thereof depends on a metric of the AI / ML model and / or the function thereof. The metric considers a benefit of employing the AI / ML model and / or the function thereof and wherein the metric considers an activation overhead for activating the AI / ML model and / or the function thereof.
[0013] Further, according to another embodiment, a user equipment of a wireless communication system is provided. The user equipment is configured to receive information about an output of an AI / ML model and / or an output of a function thereof out of one or more AI / ML models from another apparatus of the wireless communication system, wherein the one or more AI / ML models are suitable for supporting a task of a user equipment and / or a network entity of the wireless communication system. Whether the AI / ML model and / or the function thereof is activated by the other apparatus depends on a metric of the AI / ML model and / or the function thereof, wherein the metric considers a benefit of employing the AI / ML model and / or the function thereof and wherein the metric considers an activation overhead for activating the AI / ML model and / or the function thereof.
[0014] Further, a method for a wireless communication system is provided according to embodiments. The method comprises determining a metric of an AI / ML model and / or a function thereof of one or more non-active AI / ML models, wherein the one or more non-active AI / ML models are suitable for supporting a task of a user equipment and / or a network entity of the wireless communication system, wherein the method is performed by the user equipment or by a device of the wireless communication system different from the user equipment; wherein the determining the metric of the AI / ML model and / or the function thereof is such that the metric considers a benefit of employing the AI / ML model and / or the function thereof, and such that the metric considers an activation overhead for activating the AI / ML model and / or the function thereof. Further, the method comprises determining whether to activate the AI / ML model and / or the function thereof depending on the metric of the AI / ML model and / or the function thereof.
[0015] Further, a method for a wireless communication system is provided according to embodiments. The method comprises activating an AI / ML model and / or a function thereof of one or more AI / ML models; wherein the one or more AI / ML models are suitable for supporting a task of a user equipment and / or a network entity of the wireless communication system; wherein the method is performed by the user equipment or by a device of the wireless communication system different from the user equipment; wherein whether to activate the AI / ML model and / or the function thereof depends on a metric of the AI / ML model and / or the function thereof. The metric considers a benefit of employing the AI / ML model and / or the function thereof, and wherein the metric considers an activation overhead for activating the AI / ML model and / or the function thereof.
[0016] Further, a method for a wireless communication system is provided according to embodiments. The method comprises receiving, by a user equipment, from another device of the wireless communication system, information about an output of an AI / ML model and / or an output of a function thereof of one or more AI / ML models, wherein the one or more AI / ML models are suitable for supporting a task of a user equipment and / or a network entity of the wireless communication system. Whether the AI / ML model and / or the function thereof is activated by the other device depends on a metric of the AI / ML model and / or the function thereof, wherein the metric considers a benefit of employing the AI / ML model and / or the function thereof, and wherein the metric considers an activation overhead for activating the AI / ML model and / or the function thereof.
[0017] Further, a computer program is provided according to embodiments for implementing one of the above-described methods, when the program is executed by a computer or a signal processor.
[0018] Figure 1 An example set of models trained for different regions is shown in Fig. 3.
[0019] Figure 1 Different regions / scenarios where model selection / switching is required according to embodiments are shown. The shaded regions represent model overlap.
[0020] Here, the mobile UE always starts from area A and activates model A. When the UE moves to the shaded area, it needs to decide whether to continue using the specific model or switch to another model. Especially in the top-right corner scenario, switching from model A to model B can not be sufficient to cope with, and a new model needs to be trained for the shaded area.
[0021] As Figure 2 illustrated, different models can be defined not only for spatially different areas but also for different values of SNR, UE speed, etc.
[0022] Figure 2 Different operation conditions / scenarios are illustrated where model selection / switching is needed according to embodiments. Shaded areas represent model overlap.
[0023] Figure 2 a illustrates different cells where model / feature selection / switching is needed according to embodiments.
[0024] In particular, Figure 2 a illustrates an example of mobility management under the 3GPP discussion framework. When a UE moves from the coverage area of a (source) cell towards the coverage area of one or more (target) cells (possibly with overlap), a decision needs to be made on which target gNB the UE is connected to, a process known as handover.
[0025] Multiple handover management mechanisms have been defined, including “legacy” handover, conditional handover (CHO), LTM, and DAPS. They share the commonality that the UE must be ensured to be efficiently and seamlessly connected to the target gNB, avoiding any QoS degradation or radio link interruption.
[0026] When AI / ML support is incorporated into the 3GPP discussion (e.g., for scenarios of beam management, CSI compression / prediction, and positioning, etc.), the process becomes more complex. For example, assume Figure 1 the UE illustrated in FIG. 1 has an active AI / ML model performing beam management. It is assumed that the active model achieves the expected performance targets / constraints in some areas (e.g., cell 1 and cell 2), but in other areas (e.g., cell 3), a different model (or functionality) needs to be employed to achieve the performance targets. In this case, in addition to the measurement indicators of regular handover management (e.g., received RSRP), the prediction of the estimated performance of the model / feature in the target cell is also crucial.
[0027] Finally, even for the same area and conditions, different implementations of a model achieving the same functionality (presented on the network with the same model ID) can be available. Here, the monitoring entity can decide to switch between models according to, for example, expected performance versus model complexity (seeFigure 3 ).
[0028] In some cases, the monitoring entity can require certain configured downlink reference signals for the UE to receive or certain configured uplink reference signals for the UE to transmit, which can be in addition or in replacement of the reference signals the UE is currently transmitting or receiving. Similarly, in some cases, the non-active ML model can require certain configured downlink reference signals for the UE to receive or certain configured uplink reference signals for the UE to transmit, which can be in addition or in replacement of the reference signals the UE is currently transmitting / receiving. In such cases, the UE can request the network to transmit certain configured reference signals for the network device to transmit.
[0029] In some cases, the request for the UE to transmit or receive can be sent by the network:
[0030] 1) The network entity indicates to the UE through higher layer signaling mechanism (e.g. RRC signaling or LPP signaling) one or more configured reference signals the UE can expect to receive and / or transmit.
[0031] 2) The network entity indicates to the UE through reconfiguration or lower layer trigger to initiate the reception or transmission of such signals. For example, the new configuration can be provided by RRC-Reconfiguration or a new LPP message ProvideAssistanceData. Or, a MAC-CE or physical layer DCI or sidelink DCI can be used to indicate the UE to receive or transmit such signaling. Or, a MAC-CE can be used to switch the non-active model and indicate the UE to receive or transmit the signals again.
[0032] In some cases, the UE can decide (due to its internal implementation or triggered by higher layer or triggered by another entity) to receive additional downlink reference signals from at least one network entity. The UE can need to perform measurements on certain downlink reference signals for inference or monitoring purposes. But if the network does not actively send the reference signals the UE needs, the UE can:
[0033] 1) Request the network entity to transmit the reference signals, which indicates an identifier identifying the configuration of the reference signals from the multiple reference signal configurations that have been indicated to the UE.
[0034] Or:
[0035] Request the network entity to transmit the reference signals and / or request the NW to provide the UE with the on-demand configuration of the reference signals the UE can receive, which indicates at least one parameter describing the reference signals, e.g. periodicity, bandwidth, subcarrier spacing, spatial direction, etc.
[0036] 2) Perform measurements on the reference signals transmitted by the network.
[0037] 3) Monitor the performance of the activated ML model and / or at least one non-activated ML model.
[0038] The network can have provided configuration and / or assistance data indicating that a reference signal for AI / ML monitoring can need to be requested for a specific configuration. BRIEF DESCRIPTION OF DRAWINGS
[0039] Embodiments of the application are described in more detail below with reference to the accompanying drawings, in which:
[0040] Figure 1 Different areas / scenarios where model selection / switching is needed according to embodiments are shown.
[0041] Figure 2 Different operating conditions / scenarios where model selection / switching is needed according to embodiments are shown.
[0042] Figure 2 a Different cells where model / function selection / switching is needed according to embodiments are shown.
[0043] Figure 3 Model / function relationships in 3GPP are shown.
[0044] Figure 4 Model switching scenarios are shown.
[0045] Figure 5 Performance gains and costs of computing model activation or switching decisions according to embodiments are shown.
[0046] Figure 6 Data collected from multiple UEs in the same area at different points in time for model selection / activation / deactivation / switching according to embodiments are shown.
[0047] Figure 7 A scenario where model A is activated in overlapping AB areas according to embodiments is shown.
[0048] Figure 8 A second scenario where model Z is activated in overlapping AB areas according to embodiments is shown.
[0049] Figure 9 Model switching operations in bidirectional operation according to embodiments are shown.
[0050] Figure 10 An example representation of a 3GPP network is shown, with representative functional blocks depicted.
[0051] Figure 11 Transmission of multiple inference devices in a wireless communication system according to embodiments is shown.
[0052] Figure 12A transmission from an inference device in a wireless communication system according to an embodiment is shown.
[0053] Figure 13 A mechanism is shown that combines labeled data from inference data with ground truth data from one or more sources to obtain labeled data for training a model and / or a functional and / or performance metric.
[0054] Figure 14 A flowchart of providing one or more AI / ML models from a network to a user equipment according to another embodiment is shown.
[0055] Fig. 15 shows a schematic representation of an example of a terrestrial wireless network.
[0056] Figure 16 An example of a computer system is shown in which units or modules and steps according to the methods of the present application can be executed. DETAILED DESCRIPTION
[0057] An apparatus of a wireless communication system is provided according to an embodiment.
[0058] The apparatus is configured to determine a metric of an AI / ML model and / or a function thereof of one or more inactive AI / ML models, wherein the one or more inactive AI / ML models are suitable for supporting a task of a user equipment and / or a network entity of the wireless communication system, the apparatus being the user equipment or being different from the user equipment; wherein the apparatus is configured to determine the metric of the AI / ML model and / or the function thereof such that the metric takes into account a benefit of employing the AI / ML model and / or the function thereof, and such that the metric takes into account an activation overhead for activating the AI / ML model and / or the function thereof.
[0059] Further, the apparatus is configured to determine whether to activate the AI / ML model and / or the function thereof in dependence on the metric of the AI / ML model and / or the function thereof.
[0060] According to an embodiment, the apparatus can be configured to, for example, activate the AI / ML model if the apparatus has determined that the AI / ML model should be activated.
[0061] In an embodiment, the apparatus can be, for example, a user equipment; and the apparatus can be configured to, for example, perform the task employing the AI / ML model if the apparatus has determined that the AI / ML model should be activated.
[0062] According to embodiments, the apparatus can e.g. be different from the user equipment; and the apparatus can e.g. be configured to send information to the user equipment to activate the AI / ML model if the apparatus has determined that the AI / ML model should be activated. Or, the apparatus can e.g. be configured to send information to another apparatus of the wireless communication system to activate the AI / ML model if the apparatus has determined that the AI / ML model should be activated.
[0063] According to embodiments, the term “function” can e.g. refer to a particular configuration, input or output of an AI / ML model within the apparatus. For example, the term can relate to at least one function comprising an AI / ML model. Each function can e.g. contain one or more models and can e.g. be distinguished from another function by at least one of configuration, input or output. For example, models with the same configuration, input and output can e.g. be considered as part of the same function. The configuration within a function can e.g. encompass various elements like network signaling configuration, training configuration, monitoring configuration, reporting configuration and other related parameters.
[0064] In embodiments involving operations of the same function, the solution can e.g. focus on managing AI / ML models within the same function. The apparatus can e.g. determine or can e.g. be configured to determine a metric of AI / ML models between multiple models within the same function, where the AI / ML models within the same function can e.g. share a common configuration, input and output. The apparatus can e.g. be further configured to evaluate the benefit of employing the AI / ML models within the same function and the activation overhead required for each model. Based on the determined metric, the apparatus can e.g. decide whether to activate or deactivate the AI / ML models within the same function, taking into account the overall benefit and overhead involved.
[0065] In another embodiment, the solution can e.g. involve managing multiple interconnected AI / ML models within the same function. The apparatus can e.g. determine or can e.g. be configured to determine a metric of a set of interconnected AI / ML models within the same function, where the interconnected models cooperate to perform a particular task. The apparatus can e.g. be further configured to consider the benefit of employing the interconnected AI / ML models and the activation overhead required for each individual model. Based on the metric, taking into account the collective benefit and overhead involved with the interconnected models, the apparatus can e.g. decide whether to activate or deactivate the set of interconnected AI / ML models within the same function. Activation or deactivation of any individual AI / ML model within the set can e.g. affect the overall performance and functionality of the interconnected models.
[0066] In another embodiment involving operations between different functions, the solution can for example support operations between different functions, where at least one function can for example be a function supporting AI / ML features. The apparatus can for example determine or can for example be configured to determine a metric of AI / ML models and / or their functions within a current function, where the current function does not include a function supporting AI / ML features. The apparatus can for example evaluate a benefit of employing AI / ML models and / or functions within the current function and an activation overhead required for each model and / or function. The apparatus can for example determine a metric of AI / ML models and / or their functions within a target function, where the target function includes at least one function supporting AI / ML features. The apparatus can for example evaluate a benefit of employing AI / ML models and / or functions within the target function and an activation overhead required for each model and / or function; and can for example decide whether to activate the target function based on the determined metrics and the evaluated benefits and activation overheads, while taking into account an overall improvement and overhead involved in activating AI / ML feature enhanced functions.
[0067] In another embodiment, the solution can for example support operations between different functions, where at least one function can for example be a function supporting AI / ML features. The apparatus can for example determine or can for example be configured to determine a metric of AI / ML models and / or their functions within a current function, where the current function can for example include a function supporting AI / ML features.
[0068] In an embodiment, the apparatus can for example be configured to determine a metric of AI / ML models and / or their functions such that the metric takes into account a benefit and / or cost of deactivating a currently employed AI / ML model and / or its function.
[0069] According to an embodiment, the apparatus can for example be configured to determine a metric of AI / ML models and / or their functions such that the metric takes into account a benefit and / or cost of switching from a currently employed AI / ML model and / or its function to an AI / ML model and / or its function.
[0070] In an embodiment, the apparatus can for example be configured to determine, from a metric of AI / ML models and / or their functions, whether an AI / ML model and / or its function can be activated or whether a non-AI / ML (e.g. traditional) function should be employed, for example as a fallback.
[0071] According to an embodiment, an activation overhead for activating an AI / ML model and / or its function includes one or more of:
[0072] - a computational cost, for example a number of processing cycles, a number of multiplications, etc.,
[0073] - signaling cost, e.g. amount of data of signaling messages to be exchanged between the user equipment and the units of the wireless communication system,
[0074] - activation time required for activating the model, etc.,
[0075] - increase of latency,
[0076] - monitoring cost,
[0077] - combination of the above.
[0078] According to embodiments, the activation overhead for activating an AI / ML model and / or its functionality can for example comprise a monitoring cost, wherein the monitoring cost can for example depend on the availability of ground truth labels and / or PRUs, and / or can for example depend on the frequency of measuring all beams in a codebook in beam management.
[0079] Regarding the monitoring cost, one of the costs related to the LCM of an AI / ML model is the monitoring cost. This actually means that the model performance monitoring can cause certain overhead due to measurements / signaling or even resource availability, e.g. availability of PRUs of ground truth labels in positioning, or frequent measurements of all beams in a codebook in beam management.
[0080] Different models / functionality can have different monitoring requirements. For example, a monitoring configuration for monitoring the input / output of a model to determine whether it is close to the training data distribution can have much lower overhead than a monitoring configuration that supports frequent measurements of all beams in a codebook.
[0081] Therefore, even if a model / functionality is predicted to be the best performing model / functionality, it can not be selected to be activated due to strict monitoring requirements; or only be activated under certain monitoring configurations.
[0082] In embodiments, the apparatus is configured to determine the metric of the AI / ML model and / or its functionality in dependence on at least one of:
[0083] - information about the currently activated functionality / model and its properties,
[0084] - information about the performance or related QoS of the currently activated functionality / model,
[0085] - information about the cell ID and / or area ID and / or dataset ID,
[0086] - potential performance requirements and / or cost limitations,
[0087] - input data of the currently employed AI / ML model,
[0088] - measurements related to the applicable conditions of the function, e.g. SNR level, UE speed, Doppler, beam codebook type, PRS identity, model pair information for the bi-directional model, and / or e.g. network synchronization error, and / or e.g. UE / gNB RX and TX timing error,
[0089] - information about alarms from other model monitoring entities, and / or general monitoring metric computation results from other model monitoring entities,
[0090] - information about the amount of time the currently active model has been active,
[0091] - high level features / post-processing information about the UE state, e.g. UE orientation / position / speed, predicted future UE trajectory,
[0092] - assistance information from the network, about general properties of the radio environment, issues reported by other UEs.
[0093] With respect to the aspect of using cell ID / area ID / data set ID as input to the estimator, the model can be trained, e.g., by mixing data from collected data sets of various cells / areas. The reasonable assumption is that as long as the radio / environment properties do not change, the model achieves the expected performance in these cells / areas. Thus, the cell ID / area ID / data set ID as input to the estimator can indicate whether the model (and the supported function) can achieve the expected performance within a certain performance target / constraint range.
[0094] According to an embodiment, the one or more AI / ML models comprise two or more AI / ML models.
[0095] In an embodiment, the apparatus can be, e.g., a user equipment. The apparatus can be, e.g., configured to activate which of the at least two AI / ML models according to a network element of the wireless communication system, to select and / or activate one of the two or more AI / ML models, and / or to switch from one of the two or more AI / ML models to another one of the two or more AI / ML models.
[0096] According to an embodiment, the apparatus can be, e.g., a user equipment. The apparatus can be, e.g., configured to receive information about the rules from a network element of the wireless communication system, wherein the information relates to selecting and / or activating one of the two or more AI / ML models, and / or switching from one of the two or more AI / ML models to another one of the two or more AI / ML models.
[0097] In an embodiment, the apparatus can for example be a user equipment. The apparatus can for example be configured to request, from a network element of a wireless communication system, permission to select and / or activate one of two or more AI / ML models, and / or to switch from one of the two or more AI / ML models to another one of the two or more AI / ML models. Further, the apparatus can for example be configured to select and / or activate one of the two or more AI / ML models, and / or to switch from one of the two or more AI / ML models to another one of the two or more AI / ML models, upon receiving the permission from the network element.
[0098] According to an embodiment, the apparatus can for example be a user equipment. The apparatus can for example be configured to select and / or activate one of two or more AI / ML models, and / or to switch from one of the two or more AI / ML models to another one of the two or more AI / ML models, in accordance with selection information received from a network element of a wireless communication system.
[0099] In an embodiment, the apparatus can for example be configured to determine, in accordance with a current location of the user equipment, a performance of one or more AI / ML models supporting a task of the user equipment and / or a network entity.
[0100] According to an embodiment, each of the two or more AI / ML models can for example be applicable to a geographical area. The apparatus can for example be configured to, if the current location of the user equipment is located in a geographical area to which two of the two or more AI / ML models are applicable, determine whether to activate a first or a second of the two AI / ML models by determining a metric of each of the two AI / ML models. The metric of the AI / ML model considers a benefit of employing the AI / ML model and / or a function thereof, and wherein the metric considers an activation overhead for activating the AI / ML model and / or the function thereof.
[0101] In an embodiment, the apparatus can for example be a user equipment. The apparatus can for example be configured to determine a metric of each of the two AI / ML models, if the apparatus has determined that it is located in a geographical area to which the two AI / ML models are applicable.
[0102] According to an embodiment, the apparatus can for example be different from a user equipment. The apparatus can for example be configured to receive, from the user equipment, information that the user equipment is located in a geographical area to which the two AI / ML models are applicable. Further, the apparatus can for example be configured to determine a metric of each of the two AI / ML models in response to receiving the information.
[0103] In embodiments, the apparatus can for example be configured to determine the metric of each of the two or more AI / ML models and / or the metric of the functionality thereof in dependence on a characteristic of a current environment of the user equipment.
[0104] According to embodiments, the apparatus can for example be configured to determine the metric of each of the two or more AI / ML models and / or the metric of the functionality thereof in dependence on a characteristic of the user equipment and / or in dependence on a characteristic of the network entity.
[0105] In embodiments, the apparatus can for example be configured to determine the metric of each of the two or more AI / ML models and / or the metric of the functionality thereof in dependence on a state of a battery power of the user equipment and / or in dependence on an activated battery power saving mode.
[0106] According to embodiments, the apparatus can for example be configured to determine the metric of each of the two or more AI / ML models and / or the metric of the functionality thereof in dependence on a transmission characteristic of a transmission between the user equipment and the network and / or in dependence on a transmission characteristic of a transmission between the user equipment and another user equipment and / or in dependence on a radio environment property.
[0107] In embodiments, the apparatus can for example be configured to determine the metric of each of the two or more AI / ML models and / or the metric of the functionality thereof in dependence on a current state of the user equipment and in dependence on one or more possible future states of the user equipment.
[0108] According to embodiments, the apparatus can for example be configured to determine the metric of each of the two or more AI / ML models and / or the metric of the functionality thereof in dependence on two or more possible future actions of the user equipment.
[0109] In embodiments, the apparatus can for example be configured to determine the metric of each of the two or more AI / ML models and / or the metric of the functionality thereof in dependence on a reward function that returns a real value indicative of a performance of one of the two or more AI / ML models (e.g., when the user equipment is in a current state and one of one or more future states, performing one of two or more possible future actions).
[0110] According to embodiments, the reward function returns one of the following values:
[0111] - for beam management, a value indicative of a performance, e.g., a top K accuracy of the AI / ML model or a system throughput achieved with the selected beam,
[0112] - for CSI compression, a value indicating performance, e.g. indicating throughput or similarity between decoder output and target CSI,
[0113] - for direct / assisted positioning, e.g. a value indicating prediction accuracy, e.g. evaluated by a PRU capable of generating ground truth labels.
[0114] In embodiments, the apparatus can be configured to determine the metric for each of the two or more AI / ML models and / or for the functionality thereof according to a cost function, the cost function taking into account a measure of an overhead or computational cost for activating a particular AI / ML model of the one or more AI / ML models, and / or a measure of an overhead or computational cost for activating a functionality of the particular AI / ML model, and / or a measure of an overhead or computational cost for switching from a current AI / ML model of the one or more AI / ML models to another AI / ML model of the one or more AI / ML models.
[0115] According to embodiments, the apparatus can be configured to determine the metric for each of the two or more AI / ML models and / or for the functionality thereof according to a reward function and according to a cost function.
[0116] In embodiments, the apparatus can be configured to determine the metric for each of the two or more AI / ML models and / or for the functionality thereof by determining a linear combination of the reward function and the cost function.
[0117] According to embodiments, the reward function returns a value that penalizes switching from one of the two or more AI / ML models to another of the two or more AI / ML models.
[0118] In embodiments, the reward function returns a value that penalizes repeatedly switching from one of the two or more AI / ML models to another of the two or more AI / ML models.
[0119] According to embodiments, the cost function returns a value that penalizes switching from one of the two or more AI / ML models to another of the two or more AI / ML models.
[0120] In embodiments, the cost function returns a value that penalizes repeatedly switching from one of the two or more AI / ML models to another of the two or more AI / ML models.
[0121] According to embodiments, each of the one or more AI / ML models is implemented by one or more neural networks.
[0122] In embodiments, the task can for example be a positioning task of the user equipment and / or of the network entity.
[0123] According to embodiments, the task can for example be a management task or a configuration task of the user equipment and / or of the network entity, for example a beam management task of the user equipment and / or of the network entity.
[0124] According to embodiments, the task can for example be an encoding task of the user equipment and / or of the network entity, or a compression task of the user equipment and / or of the network entity, for example a task for compressing channel state information.
[0125] Further, according to another embodiment a device of a wireless communication system is provided.
[0126] The device is configured to activate an AI / ML model of one or more AI / ML models and / or a function thereof; wherein the one or more AI / ML models are suitable for supporting a task of a user equipment and / or of a network entity of the wireless communication system; wherein the device is the user equipment or different from the user equipment; wherein whether to activate the AI / ML model and / or the function thereof depends on a metric of the AI / ML model and / or a metric of the function thereof.
[0127] The metric takes into account a benefit of employing the AI / ML model and / or the function thereof, and wherein the metric takes into account an activation overhead for activating the AI / ML model and / or the function thereof.
[0128] According to embodiments, the device implements the device according to one of the above embodiments.
[0129] In embodiments, the device does not implement the device according to one of the above embodiments, but the device can for example be configured to receive information about an AI / ML model of the one or more AI / ML models to be activated from the device according to one of the above embodiments.
[0130] According to embodiments, the device can for example be the user equipment.
[0131] In embodiments, the device is for example different from the user equipment, but the device can for example be configured to provide an output from the AI / ML model to the user equipment and / or to the network entity to support the user equipment and / or the network entity to perform the task.
[0132] Further, according to another embodiment a user equipment of a wireless communication system is provided.
[0133] The user equipment is configured to receive information about an output of an AI / ML model of one or more AI / ML models and / or a function thereof from another device of the wireless communication system, wherein the one or more AI / ML models are suitable for supporting a task of the user equipment and / or of a network entity of the wireless communication system.
[0134] Whether the AI / ML model and / or its functionality is activated by the further apparatus depends on a metric of the AI / ML model and / or a metric of its functionality, wherein the metric takes into account a benefit of employing the AI / ML model and / or its functionality, and wherein the metric takes into account an activation overhead for activating the AI / ML model and / or its functionality.
[0135] According to an embodiment, the further apparatus can be, for example, an apparatus according to one of the above embodiments.
[0136] Further, a wireless communication system is provided, which comprises an apparatus according to one of the above embodiments and a user equipment.
[0137] According to an embodiment, the wireless communication system further comprises a further apparatus according to one of the above embodiments.
[0138] In an embodiment, the user equipment can be, for example, a user equipment according to one of the above embodiments.
[0139] Further, a wireless communication system is provided, which comprises a first apparatus according to one of the above embodiments and a second apparatus according to one of the above embodiments. The first apparatus is configured to select and / or activate one of one or more AI / ML models according to the second apparatus and / or to select and / or activate one of one or more AI / ML models and / or to switch from one of one or more AI / ML models to another one of one or more AI / ML models according to a switch from one of one or more AI / ML models to another one of one or more AI / ML models.
[0140] Before providing further embodiments of the present application, some background information is provided.
[0141] In the 5G AI / ML framework, the life cycle management (LCM) referred to in the field of machine learning covers the whole process of end-to-end development, deployment and maintenance of machine learning models. The process contains multiple stages such as data preparation, model training, testing, deployment, monitoring and maintenance. For the proposed solution, the focus is on the LCM stages related to landmark utilization.
[0142] Data collection defined in the 3GPP framework refers to the process of collecting data by network nodes, management entities or UEs for the purpose of AI / ML model training, data analysis and inference. Data collection and preparation involve data collection and preparation by UEs, networks or outside the network (such as non-3GPP entities). The data is used to train machine learning models offline or in real time.
[0143] Model training is defined as the process of learning the input / output relationship by data driven approach to train the AI / ML model and obtain the trained AI / ML model for inference: This phase trains the machine learning model using the prepared data, involving the selection of the right algorithm and optimization of the model’s performance.
[0144] Model validation is defined as a sub-process of training that assesses the quality of the AI / ML model by using a dataset different from the one used to train the model, which helps to select model parameters that generalize beyond the dataset used for model training.
[0145] Model testing is defined as a sub-process of training that assesses the performance of the final AI / ML model using a dataset different from the one used for model training and validation. Unlike AI / ML model validation, the testing phase does not presuppose subsequent model adjustment.
[0146] Model monitoring is defined as the process of monitoring the inference performance of the AI / ML model. The model needs to be continuously monitored after deployment to detect any performance degradation or errors. This phase involves tracking model performance (measure) indicators, detecting data drift, and retraining the model if necessary.
[0147] Model maintenance: The model needs to be continuously maintained and updated to ensure its performance remains optimal. This phase involves retraining the model with new data, upgrading the algorithm, and improving the architecture.
[0148] An activated model can for example be understood as the (AI / ML) model currently used for inference.
[0149] A non-activated model can for example be understood as all available models that the UE can use.
[0150] A candidate model for activation can for example be understood as a model that potentially has performance equivalent to or better than the currently activated model.
[0151] According to the 5G framework, function identification can for example be understood as the process / method to identify AI / ML functions for common understanding between the NW and the UE.
[0152] The UE can have a single AI / ML model for a function, or the UE can have multiple AI / ML models for a function (see for example Figure 3 Functions A, B, Z in ). Models for the same function can have different structures, their input / output / auxiliary information configuration can be the same.
[0153] Alternatively, a more complex model trained by data from multiple sites can implement more than one function (see for example Figure 3The model selection / activation / deactivation / switching / back-off indication is based on the model ID. The model ID is used to identify the model in the LCM procedure.
[0154] For the function-based UE part / UE-side model LCM procedure, the UE can provide a single AI / ML function activation / deactivation / switching / back-off indication by one approach. The UE can receive assistance data to enable the function. Another approach is that the terminal can receive a single AI / ML function activation / deactivation / switching / back-off indication from a second entity, e.g., a coordination entity. For the latter approach, the second entity as the network indicates the AI / ML function activation / deactivation / back-off / switching via 3GPP signaling, e.g., RRC, MAC-CE, DCI.
[0155] Figure 3 The model / function relationship in 3GPP is shown.
[0156] For example, when the network needs to identify the UE AI / ML model, the AI / ML model can have a model ID associated with the configuration of at least part of the AI / ML operation. In the LCM procedure based on the model ID, the model selection / activation / deactivation / switching / back-off indication is based on the model ID.
[0157] The model description information or meta information refers to the supplementary information about the model provided in the model identification process. The model description information can include a list of applicable AI / ML-enabled features and / or applicable conditions for the model. The conditions may, for example, include one / more applicable functions, applicable RRC configurations, model pairing ID.
[0158] The same function can be implemented with different models. For example, in some embodiments, if the AI / ML model is implemented in the form of a neural network, the neural network may, for example, include at least one of a fully connected layer, a pooling layer, and a convolutional layer. According to some embodiments, a dense network may, for example, be employed. In some embodiments, weight pruning and / or node pruning may, for example, be employed.
[0159] To enable model activation / selection / switching, the necessity, feasibility, and potential (specification) impact of methods to evaluate the applicability of non-active AI / ML models / functions have been discussed, including the following examples:
[0160] Evaluation by comparing the applicability conditions of the model with the current conditions. These “static” conditions, such as region ID, SNR level, UE speed, etc.
[0161] Evaluation based on input data distribution. This can be: Note that the measured data seems to be the training data used to train the model, so the model should perform well in this scenario.
[0162] Evaluation by model monitoring, using non-active models for monitoring, measuring their inference accuracy / system performance. This can be understood as: loading non-active models, providing them with input data and recording their output. If a non-active model performs better than the currently active model (based on the output of all models), switch to that model.
[0163] One open question is: how to mitigate the system performance impact that this can cause (if any). The problem is: without running all models in parallel, how can one predict whether a non-active model is more applicable?
[0164] There can be multiple drawbacks in the way of evaluating by using non-active models for monitoring and measuring inference accuracy / system performance:
[0165] Parallel loading of non-active models and performing inference can cause high computational overhead for the UE.
[0166] Determining whether a non-active model is better than the active model can require ground truth labels for monitoring.
[0167] If the UE is in a sub-area or condition set for which more than one model is applicable, it can get stuck in a constant model switching dilemma.
[0168] Loading non-active models that perform the same task (can be activated under the same applicable conditions) but have different implementations (e.g. Transformers vs. CNN, small / large models, etc.) or belong to different functions (and thus can require different inputs) is still challenging.
[0169] The following provides specific embodiments.
[0170] Embodiments relate to evaluation of the applicability of non-active AI / ML models / functions.
[0171] Instead of loading candidate non-active models in parallel, these approaches construct estimators: for a selected non-active function / model, the estimator predicts the expected gain of activating that function / model, taking into account the model’s expected performance / QoS and the cost of selecting / activating / deactivating / switching to the candidate function / model.
[0172] It can be observed that, on the basis of a trade-off between performance / cost, model selection / activation / deactivation / switching can be achieved in multiple ways, for example:
[0173] • For example, the model with the best performance can always be used, regardless of its associated cost.
[0174] • Performance requirements can be provided, for example. A performance estimator can be queried, for example, and a list of candidate models meeting the performance constraints compiled therefrom. The finally activated model can be the one with the lowest expected cost on the list, for example, according to a cost estimator.
[0175] • Maximum acceptable cost can be provided, for example. A cost estimator can be queried, for example, and a list of candidate models meeting the cost constraints compiled therefrom. The finally activated model can be the one with the highest expected performance on the list, for example, according to a performance estimator.
[0176] An overview of embodiments is provided below.
[0177] According to embodiments, a method of predicting / estimating expected benefit of a machine learning (ML) model / function by an inference device is provided, the device supporting multiple ML models / functions. The method comprises:
[0178] • selecting a set of applicable models / functions from the supported ML models / functions; and
[0179] • estimating expected performance and / or cost of the selected models based on or instead of the activated models, or for initial model selection;
[0180] • deciding on activation or switching of ML models / functions by the device; or
[0181] • providing information on the estimated expected benefit to a second device to enable the second device to make model activation / switching decisions for the inference device and / or other inference devices with similar functionality.
[0182] The expected benefit of some embodiments depends on short, medium and long term factors, which avoids both failures and high cost switching while improving performance.
[0183] According to embodiments, the inference device can be unilateral or bilateral, for example.
[0184] In some embodiments, the selection step can be performed by the device or the NW or both, for example.
[0185] According to some embodiments, the estimation step can be performed by using data of successful and failed model switching (“successful in the next cycle”), for example.
[0186] In embodiments, the problem can be modelled as an MDP over multiple partitions to trade off cost / risk, for example.
[0187] According to embodiments, information / QoS / policies for operation over multiple partitions can be received to train or configure the weight data, for example.
[0188] Technical details of some specific embodiments are provided below.
[0189] An operational example is first described.
[0190] It is assumed that there are Figure 4 Example regions are shown. The following assumptions are made in this example:
[0191] 1. For simplicity, it is assumed in this example that all models complete training and inference on the UE.
[0192] 2. Example functions can be:
[0193] ° Beam management model (i.e. select best serving beam out of X total beams without measuring all beams).
[0194] ° Positioning model (i.e. estimate UE’s location).
[0195] 3. Training properties of available AI / ML models:
[0196] ° Model A is trained using data from sub-region A (+ some border data with sub-region B), so there is some buffer / overlap.
[0197] ° Model B is trained using data from sub-region B (+ some border data with sub-region A), so there is some buffer / overlap.
[0198] ° Model Z is a general model trained using multi-site data.
[0199] 4. Model complexity:
[0200] ° Sub-region B is more “challenging” than sub-region A (e.g. more obstacles, reflections, etc.). This means that model B can need different inputs than model A (e.g. need more beam measurements for beam management, or need additional AI / ML model to classify LOS / NLOS positioning). This also means that model B can have a much larger number of parameters than model A.
[0201] ° Model Z is trained using data from multiple sub-regions across different sites. Since it is not optimized for a specific sub-region, it has strong generalization capability, but has a much larger number of parameters than models A and B.
[0202] 5. Model performance:
[0203] ° Models A and B are “specialized” models that have the best performance in their respective sub-regions.
[0204] ° Model Z has lower performance but covers the entire region.
[0205] ° Another non-AIML model is available with the worst performance compared to the available AI / ML models. This model is sometimes referred to as the fallback model, which can be switched to at any time when the AI / ML model performance degrades (e.g., multiple temporary blockage points occur or permanent major changes in the environment geometry).
[0206] ° Taking positioning as an example, assume the performance of each model is as follows:
[0207] ■ Model A -> accuracy < 20 cm in sub-area A
[0208] ■ Model B -> accuracy < 20 cm in sub-area B
[0209] ■ Model Z -> accuracy < 2 m in the whole area
[0210] ■ Non-AIML model -> accuracy < 5 m in the whole area
[0211] In another example, equivalent KPIs for BM can be considered, for example.
[0212] 6. Model selection / activation / deactivation / switching cost:
[0213] ° If observing UE trajectories, a model switching mechanism needs to be established when using models A and B to ensure high-level performance throughout the UE trajectory.
[0214] ° Different cost levels of scenarios are as follows:
[0215] ■ Example #1: Model A and Model B support the same function. This means that model switching occurs within the same function, and the selection / activation / deactivation / switching cost is low -> the input / output / auxiliary information configuration is the same, and no coordination with the gNB is required.
[0216] ■ Example #2: Model A and Model B support different functions (e.g., support different sets A / set B of BM). This means that the function (configuration) also needs to be switched, resulting in an increased selection / activation / deactivation / switching cost -> different input / output / auxiliary information configurations can be required, which must be coordinated with the gNB.
[0217] ■ Example #3: Model A and / or B are not stored at the UE device. In this case, the model needs to be downloaded from the NW or acquired through the user plane before activation.
[0218] In this example, a reasonable model / function selection / activation / deactivation / switching mechanism should have the following characteristics:
[0219] • The performance of each available model (A, B, Z) and the fallback performance of the location where the UE is located should be predictable.
[0220] • One should be able to decide at the UE location indicated by the black circle which sub-model A or B is the best choice.
[0221] • One should be able to predict when the relative model performance can change when the UE is located at the position indicated by the red square.
[0222] • One should be able to consider the cost of model / functional selection / activation / deactivation / switching and make the corresponding decision in scenarios where both models can be used (e.g., when the UE is located at the position indicated by the grey triangle).
[0223] Figure 4 Model switching scenarios are shown. Dashed lines indicate the applicability of each model.
[0224] According to the 3GPP standard, the task in all scenarios is to switch to a model Y with better performance before the performance of the currently used model X degrades to an unacceptable level, so the switching decision must be timely and robust. The standard achieves this goal by:
[0225] 1. Clearly defining when (under what conditions) a model is suitable for activation / use.
[0226] 2. Based on the current conditions / measurements, "loading" the available candidate non-activated models, running them in parallel with the currently running model, and evaluating whether their performance is better than the activated model.
[0227] Multiple models with overlapping coverage ranges should be developed (e.g., Model 1 is suitable for SNR < 20 dB, Model 2 is suitable for SNR > 10 dB) to allow for imperfections in model selection / switching, thereby enhancing the robustness of the model selection and switching mode.
[0228] This solution has the following problems:
[0229] 1. Parallel loading of non-activated models and performing inference can bring high computational overhead to the user equipment (UE).
[0230] 2. Determining whether one or more non-activated models is better than the activated model can require obtaining ground truth labels required for monitoring.
[0231] 3. If the UE is in a sub-area or set of conditions where more than one model can be applied, it can fall into the dilemma of continuous model switching.
[0232] 4. Loading non-activated models that perform the same task (can be activated under the same applicable conditions) but have different implementations (e.g., Transformers vs. CNN, small / large models, etc.) or belong to different functions (and thus can require different inputs) still presents challenges.
[0233] According to some embodiments:
[0234] 1. Instead of loading candidate non-active models and running them in parallel, construct an estimator that, for a selected non-active function / model, predicts the expected gain of activating this function / model, taking into account the expected performance / QoS of the model as well as the cost of selecting / activating / deactivating / switching to the candidate function / model.
[0235] 2. The estimator should not only provide a one-step prediction (i.e. an immediate performance / cost estimate), but also encode in the prediction the long-term performance / cost trade-off of activating a particular model, taking into account the short-term requirements for model (re-)switching based on the available models and radio environment properties.
[0236] To achieve the goal of “encoding in the prediction the long-term performance / cost trade-off of activating a particular model, taking into account the short-term requirements for model (re-)switching based on the available models and radio environment properties”, we model the optimal selection / activation / deactivation / switching of AI / ML models as e.g. a Markov Decision Process (MDP). A Markov Decision Process (MDP) is a framework for modeling sequential decision problems under uncertainty.
[0237] For example, a MDP can be defined by:
[0238] • A set of actions A - i.e. all executable actions / decisions. In this case, the decision to select / activate (or switch to) a particular AI / ML function / model, or to select an activated fallback / non-AI / ML function / model. The expected gain of each model (potential action) is the output of the estimator.
[0239] • A set of states S - encompassing all available measurements that can provide the estimator with information to make the right decision. In short, it is not enough to know only the information of the currently used model A to predict whether another model would be better. In this case, examples of inputs / measures that can be used in arbitrary combination are:
[0240] ° Information about the currently activated function / model (if a function / model is activated) and its properties.
[0241] ° Information about the performance (or related QoS) of the currently activated function / model (if a function / model is activated).
[0242] ° Potential performance requirements and / or cost limitations.
[0243] ° Input data used by the AI / ML model (last X time steps, if the model is already activated).
[0244] ° Measurements related to the applicable conditions of the function (e.g. SNR level, UE speed, Doppler, beam codebook type, PRS identity, model pair information for the bi-directional model, network synchronization error, UE / gNB RX and TX timing error, etc.).
[0245] ° Information about alerts from other model monitoring entities (may indicate performance degradation of the active AI / ML model) and / or general monitoring metrics computation results.
[0246] ° Information about the amount of time the currently active model has been active.
[0247] • Information about alerts from other model monitoring entities (may indicate performance degradation of the active AI / ML model) and / or general monitoring metrics computation results.
[0248] ° Information about the amount of time the currently active model has been active.
[0249] ° High-level features / post-processing information about the UE state. For example, UE orientation / position / speed, predicted future UE trajectory, etc.
[0250] ° Assistance information from the network, about general properties of the radio environment, issues reported by other UEs.
[0251] ° Cell ID / Zone ID / Dataset ID.
[0252] • A reward function R, to evaluate the quality of a decision - when we measure and take an action , we get a real number that reflects the effect of the action. In this problem, the reward function contains both performance and cost components (these are not inputs to the estimator, but rather used for its training / programming):
[0253] ° (Instantaneous) cost of selecting / activating / deactivating / switching to an AI / ML model .
[0254] ■ This cost is controllable when switching between models of the same function, as the models should have the same input / assistance information requirements, and no explicit signaling / coordination between UE and NW is needed.
[0255] ■The cost can increase significantly when switching between models of different functionalities, as models can have input / assistance information requirements and need explicit signaling / coordination between UE and NW. In some cases (e.g., different functionalities of beam management support different Set A / Set B), the newly activated model can need the UE to provide a new set of beam measurements as input. Note here that additional cost terms can be introduced (e.g., complexity cost proportional to the AI / ML model size, energy cost proportional to the model usage energy, overhead cost if the model is not stored at the UE and needs to be downloaded, etc.).
[0256] • (Instantaneous) model performance (P) This can be the positioning accuracy in a positioning use case, the throughput, or the best beam prediction accuracy in a beam management use case, etc. The key detail is when this performance metric is computed: to know if the switch was successful, the (average) performance of the newly activated model needs to be monitored until it is deactivated / switched. Figure 5 The decision cost / performance metric computation occasion is illustrated. In a variant of this (instantaneous) model performance metric, the model performance can be marked as suboptimal if a model switch or fallback operation happens shortly after (in short, if a model performs well for 5 time steps and then the performance drops so much that a switch to a better model or to a fallback / non-AI / ML solution is needed, then the actual performance is insufficient).
[0257] • Discount factor to determine the future reach of the action impact (e.g., does the activated model bring a long-term performance boost? Does it need more model switches - possibly in the short term?).
[0258] • Transition model T - describes how the world evolves (physical laws, future radio environment evolution, how the UE moves, etc.).
[0259] Figure 5 The performance benefit and cost of computing a model activation or switch decision are illustrated, according to an embodiment.
[0260] In an embodiment, it is assumed that a model selection / activation / deactivation / switch strategy (also referred to as strategy ) is available. This strategy will take as input the state / measurements and select a model (take an action) .
[0261] This scheme has two estimators (denoted and , corresponding to performance and cost). These functions take as input the state / measurements and predict the performance and cost of the current selection / activation / deactivation / switch to a particular model And follow the subsequent selection / activation / deactivation / switching strategy. At that time, the long-term benefits generated ( and long-term selection / activation / deactivation / switching costs ( For example, consider Figure 6 The data comes from different UEs. A series of status, actions, and performance / cost KPIs are collected here. Estimator and The training mechanism encodes the following characteristics: at time points Model selection not only affects the time point The state / performance / cost of the system can be assessed, and long-term effects may also occur (e.g., eventually requiring a switch to a different model, so starting with that model might be a better decision). This can be achieved, for example, by training the Q-function using future performance / cost values from the entire UE trajectory / experience (which can be weighted to increase the importance / confidence of immediate results).
[0262] Typical algorithms here could be Monte Carlo estimation or TD learning.
[0263] Figure 6 The illustration shows model selection / activation / deactivation / switching data collected from multiple UDs within the same area at different times (e.g., different dates, time periods, etc.) according to an embodiment.
[0264] The final part of the solution is how to choose a suitable model. Several options exist here:
[0265] Always use the model with the best performance, regardless of its associated costs.
[0266] Provide performance requirements Query estimator This process compiles a list of candidate models expected to meet the performance constraints. The final activated model is selected from this list based on... The model with the lowest expected cost.
[0267] • Provide the maximum acceptable cost Query estimator And compile a list of candidate models expected to meet the cost constraints. The final activated model is the one in the list based on... The model with the highest expected performance.
[0268] The following will present specific examples based on the embodiments.
[0269] by Figure 4The two sub-regions shown are used as an example. Assume we map model positioning accuracy from [0.0m, 10.0m] to [1, 0], i.e. a perfect model has a performance of 1 and a model with a positioning accuracy worse than 10m has a performance of zero. Further, assume that in a direct positioning task, the initial estimate of model performance is as follows:
[0270] • Model A -> accuracy < 20cm in sub-region A. In scaled performance in [0, 1], assume performance is .
[0271] • Model B -> accuracy < 20cm in sub-region B. In scaled performance in [0, 1], assume performance is .
[0272] • Model Z -> accuracy < 2m in entire region. In scaled performance in [0, 1], assume performance is .
[0273] • Non-AIML model -> accuracy < 5m in entire region. In scaled performance in [0, 1], assume performance is .
[0274] Further, assume we have collected data from multiple UEs as shown in Figure 6 . Assume that model B does not perform as expected, with a positioning accuracy level < 80cm instead of < 20cm. In scaled performance in [0, 1], this performance is . Based on the data and construction / training of the and estimators, we derive:
[0275] • Model Z has the same performance in any region.
[0276] • Fallback (non-AI / ML) has the same performance in all regions.
[0277] • Model Z is applicable in the entire region, so there is no selection / activation / deactivation / switching cost.
[0278] • Fallback (non-AI / ML) is applicable in the entire region, so there is no selection / activation / deactivation / switching cost.
[0279] • Use model A in the region where model A is valid (state indicates model A was trained with this data), with an expected performance of 0.98.
[0280] • Use model A (state) in regions where model A is invalid. (Instructing Model B to be trained using these data), with an expected performance of 0.0. Its performance may be better than 0.0, but this cannot be verified due to a lack of UE data (Model A was not used in region B during training / deployment).
[0281] · Use Model B (state) in the regions where Model B is valid. (This indicates that Model B was trained using these data), with an expected performance of 0.92 (this is indicated by the new data after deployment, not the initial estimate).
[0282] · Use Model B (state) in regions where Model B is invalid. (This indicates that model A was trained using these data), with an expected performance of 0.0. Its performance could potentially be better than 0.0, but this cannot be verified due to a lack of UE data (model B was not used in region A during training / deployment).
[0283] · Assume that when the UE is located in the left sub-region A ( , Figure 4 The middle circle mark or the right sub-region B ( , Figure 4 When the region is marked with a circle, its (normalized) cost for selection / activation / deactivation / switching is 0.0 because the switching in this region is not recorded in the data. Furthermore, the expected performance of model A in sub-region B is 0.0, and vice versa; therefore, there is no significance for model selection / activation / deactivation / switching in this region.
[0284] · Assume that when the UE is located in the overlapping area of A and B ( , Figure 4 When the triangle marker is in the middle, the expected selection / activation / deactivation / switching cost (regardless of which model is currently activated) is high (0.7) because the region needs to be switched frequently to maintain high performance, based on available data.
[0285] Now consider two scenarios: Scenario #1, where the cost of selection / activation / deactivation / switching is acceptable and optimal positioning accuracy is required; and Scenario #2, where the cost of selection / activation / deactivation / switching is greater than... That is unacceptable.
[0286] The first scenario #1 (optimal performance, see embodiment) will now be described according to an example. Figure 7 ).
[0287] • In the left sub-region A, use model A because it is the best performing model and is not expected to be switched.
[0288] • In the left sub-region A, model A is used because it is the best performing model and no switch is expected.
[0289] • In the overlapping sub-region, model A is used. Both models (A and B) have the same expected cost ( ), but model A has higher expected performance.
[0290] Figure 7 A first scenario (#1) is shown according to an embodiment, where model A is activated in the overlapping AB region. The solid lines indicate the model applied at each UE location, selected based on the performance / cost trade-off of and respectively.
[0291] A second scenario #2 is now described according to another embodiment (avoiding high cost, see Figure 8 ).
[0292] • In the left sub-region A, model A is used because it is the best performing model and no switch is expected.
[0293] • In the right sub-region B, model B is used because it is the best performing model and no switch is expected.
[0294] • In the overlapping sub-region, model Z is used. Both models (A and B) have the same expected cost ( ), which is higher than the maximum acceptable cost ( ). The sub-optimal model is model Z, which has a cost of zero, so this model is activated in the overlapping region.
[0295] • Note that when the UE approaches the overlapping region (box indicating the UE location in sub-region A) where it can switch to model Z, the expected selection / activation / deactivation / switching cost is not zero ( ), but does not exceed 0.5, because the model to be activated is model Z, which does not require a switch (it covers the entire region).
[0296] Similarly, the expected performance is not the performance of model A, but lower ( ). This is because a switch to model Z is possible, and model Z has a lower performance than model A, so the average expected performance is lower.
[0297] Figure 8 A second scenario (#2) is shown according to an embodiment, where model Z is activated in the overlapping AB region. The solid lines indicate the model applied at each UE location, selected based on the performance / cost trade-off of and respectively.
[0298] The AI / ML use cases supported by the embodiments will be shown in detail below.
[0299] An AI / ML positioning use case according to embodiments is first described.
[0300] For 5G positioning, two AI / ML approaches are considered: direct positioning and assisted positioning. Model inference can be performed at the UE or at the network side. In direct AI / ML positioning, an AI / ML model infers the UE position directly with channel observation data collected from signal measurements, such as signal power, channel impulse response (CIR), time-of-arrival (ToA) estimates, and angle-of-arrival (AoA) estimates. If the device has multiple antennas, measurements for each antenna or beam can be provided. Additionally, assistance information is also supported, such as the UE reporting its velocity based on internal sensors.
[0301] In AI / ML assisted positioning, an AI / ML model pre-processes the measurements, and the position is computed by other algorithms. The AI / ML model provides new or enhanced measurements, such as LOS / NLOS identification, AoA estimates, ToA estimates, measurement quality / reliability information, correction values, and measurement classification. For example, the model can identify specular or diffuse reflections in the measurements.
[0302] An example of a function switch between sub-area A and sub-area B can be implemented, for example, as follows:
[0303] • Function #1 of sub-area A supports CIR as an AI / ML model input. When switching to the more challenging sub-area B, the AI / ML model needs to utilize the velocity information of the UE in addition to the CIR, i.e., another function needs to be activated and a coordination between the UE and the NW needs to be implemented.
[0304] An AI / ML beam management use case is now described according to embodiments.
[0305] AI / ML based beam management, case 1: Spatial domain DL beam prediction for beam set A based on measurement results of beam set B. Case 2: Temporal DL beam prediction for beam set A based on historical measurement results of beam set B. AI / ML inputs for both cases can be: RSRP or CIR measurements based on set B, or RSRP measurements based on set B and assistance information (Tx and / or Rx beam shape information (such as Tx and / or Rx beam pattern, Tx and / or Rx beam boresight direction (azimuth and elevation), 3dB beamwidth, etc.), expected Tx and / or Rx beam parameters needed for prediction (such as expected Tx and / or Rx angles, Tx and / or Rx beam IDs needed for prediction), UE position information, UE direction information, Tx beam usage information, UE orientation information, etc.
[0306] The function switching from sub-area A to sub-area B (assuming a 64-beam codebook) and vice versa is as follows:
[0307] ■Sub-area A’s function #1 supports a set B of 10 beams (meaning 10 beams need to be measured for RSRP before prediction) and a set A of 54 beams. For function #1, the model for sub-area A needs to predict a single best beam from the 54 beams that are not measured. When switching to the more challenging sub-area B, the AI / ML model needs to measure more beams in set B (e.g., 16 beams) and predict (and measure) the top 5 best beams from the remaining 48 beams in set A. This different configuration corresponds to different function requirements, which need to be activated and coordinated between the UE and the NW.
[0308] The AI / ML CSI compression use case is now described according to an embodiment.
[0309] AI / ML based on a bidirectional model in CSI compression. In a bidirectional model, the UE and the network jointly perform a paired AI / ML model for inference: the first part of the inference is performed by the UE and the remaining part is performed by the network, or vice versa.
[0310] CSI compression using machine learning (ML) involves training a model to learn and compress channel state information (CSI) from the original channel or precoding matrix. The compressed CSI is transmitted from the UE to the NW, where it is decompressed and used for beamforming and other functions. Different compression models can be used depending on the available payload size and network configuration. The matching relationship between the input-CSI-NW and output-CSI-UE options needs to be studied to ensure model training accuracy, performance optimization, and AI / ML energy saving effects. In the CSI compression use case using a bidirectional model, the NW configures the maximum payload size. The UE selects the rank and CSI generation model within the maximum payload limit configured by the network. In an alternative option, the NW configures a list of model IDs and the maximum payload size, and the UE selects the rank and CSI generation model from the configured list while being constrained by the maximum payload configured by the network. In a third option, the NW configures the model ID for the UE to use, and the UE will use the corresponding CSI generation model configured by the NW.
[0311] Figure 9 Model switching operations performed in bidirectional operation according to an embodiment are shown.
[0312] In a bidirectional model according to Figure 9 Model A1 is active on the UE side and model C1 is active on the NW side. The decision of the UE to switch, activate, or select a model depends on the NW function.
[0313] In one aspect, since the functionality of model C1 is applicable to both A1 and A2, the UE can switch based on the evaluation of model A2. The NW and the UE exchange the supported functionalities by both parties. The UE can be configured by the NW with rules to facilitate model switching. In an alternative, the UE can be allowed to freely activate / switch the model (A1 or A2) when the output and the NW functionality are not affected.
[0314] In the same example, if the UE predicts that it is beneficial to activate model A3, the UE requests to the NW to activate model A3. The UE can also provide the expected benefit information to the NW.
[0315] In a different aspect, the estimator according to the proposed solution can be optimally aligned between the NW and the UE. In this case. The expected benefit is jointly determined by the NW and the UE through actions or / and states or / and rewards.
[0316] The estimator output according to embodiments is now described.
[0317] It should be noted here that the estimator output is the same in all supported use cases: an estimate of the expected performance of the AI / ML model ( ) and / or an estimate of the expected selection / activation / deactivation / switching cost ( ). Note that there is a confidence interval in these estimates (hence indicating how certain / robust the estimator’s prediction is).
[0318] The capabilities and NW signaling according to embodiments will be described below.
[0319] In one aspect, the apparatus (UE or BS) provides to the NW the number of supported functionalities by the UE. The UE provides to the network the number of supported models within each functionality. Wherein the apparatus is used to evaluate at least one functionality or model for selection, activation, deactivation, or switching.
[0320] In a related aspect, the apparatus as a UE will receive from the network an indication of the preferred or applicable model or / and functionality. Wherein, the UE will select the model to be evaluated based on the indication and evaluate the expected performance or a parameter reflecting the selected performance.
[0321] In a second aspect, the UE can receive a configuration message from the NW. The configuration message includes information enabling the UE to evaluate the performance benefit from one or more models. The information can include information on QoS, information on configuration, or both, to enable the UE to set the optimal state, action, and / or reward.
[0322] Figure 10 An example representation of a 3GPP network depicting representative functional blocks is shown.
[0323] In particular, Figure 10Components of a 3GPP wireless communication system (or 5G System (5GS)) are depicted. The system consists of a user equipment (UE), an access network (AN), a core network (CN) and a data network (DN). The UE can register with the AMF using an NG-RAN node (e.g. gNB) using a 3GPP defined radio access technology (e.g. NR), or by using a non-3GPP access method (e.g. WiFi) through a non-3GPP interworking function (N3IWF).
[0324] The core network contains one or more functions that can interact with each other using a so-called service-based architecture over interfaces. As an example, an AMF can send a message to an LMF over an Nlmf interface, while an LMF can send a message to an AMF over a Namf interface.
[0325] A brief description of network entities / functions is provided below to simplify the working principle of the core network. In the core network, the AMF (access and mobility function) is a network function that sends control plane signaling of the core network to the UE. The UE registers with the network through the AMF. The AMF manages the mobility of user equipment and handles access authentication and authorization in the 5G core network. The LMF (location management function) is responsible for determining the UE location by interacting with one or more network functions and / or access network nodes and / or UEs, and providing the location to a location service client, which can be another application function (AF) in the access network or in an external network, AMF, UE, entity. The network exposure function (NEF) exposes services, capabilities and / or information to external applications and third parties in a secure and controlled manner. Application functions in the data network can access information in the 5GS directly through the NEF or based on a service-based architecture. The network repository function (NRF) provides a centralized network function information repository for the 5G core network, facilitating the discovery and access of available network functions and their capabilities. The charging function (CHF) manages the charging and settlement of user services in the 5G core network, covering data usage, service subscription and payment authorization. The policy control function (PCF) is responsible for controlling and managing policy decisions and execution related to quality of service (QoS), network resources and user access in the 5G core network. The unified data management (UDM) stores and manages user-related data, such as user profiles and authentication credentials in the 5G core network. The unified data repository (UDR) is a central repository for user-related data in the 5G core network, including subscription information and session information. The network data analytics function (NWDAF) collects and analyzes network data to provide decision support for network optimization, quality of service (QoS) improvement and resource allocation in the 5G core network. The authentication server function (AUSF) handles authentication and security-related functions in the 5G core network, including generating authentication vectors and verifying user identity. Of course, in addition to the functions discussed above, there are other application functions in the network, and the above functions are only intended to outline the possible role of the application function, not to limit the scope of its functions.
[0326] Figure 11 Transmissions from multiple inference devices in a wireless communication system are shown, according to an embodiment.
[0327] Figure 10 The OTT server can be an entity managed by the network, or can be a third party server (e.g. from a vendor). The OTT can take information from the 5GS for training the estimator based on the performance of the inference device and / or the output of the estimator and / or ground truth and / or additional information. It can take the information through the network exposure interface. For example, a specific UE vendor ‘A’ can run an inference model at the UE that is loaded to the UE through the application layer (5G user plane data transmitted through the 3GPP access network, non-3GPP access network, or directly through external data connection (e.g. standalone WiFi)). The vendor can be interested in one or more parameters that are used to train the estimator features (e.g.: UE location computed by the network, RSRP of the received signal, or Block Error Rate (BLER)). The vendor can subscribe to certain information (e.g. real values like UE location, RSRP, BLER...etc.) through the NEF of the network. The received data can be correlated with the data from different sources with a unique mechanism (e.g. adding a timestamp) that enables the OTT to align the UE received information with the network received information, to train the model / feature or to train the estimator prediction model performance.
[0328] Figure 12 Transmissions from inference devices in a wireless communication system are shown, according to an embodiment.
[0329] A first case according to an embodiment involves a situation where the inference model is located at the UE, and the estimator is also located at the UE:
[0330] A vendor or network can provide multiple models for a specific feature. However, the UE can only store a subset of the models. The number of models (K) that the UE stores can be limited by the UE capability. The estimator model can be provided to the UE for estimating the performance of L models, where L > K. As an example, the UE can store a generic model and k specific models. The generic model can provide a coarse positioning result over a city range, while the specific models can provide higher resolution positioning. The estimator can predict the model with the best performance, which can not be stored in the UE. Then, the UE needs to make a request to the network and / or OTT server or other UEs in proximity (e.g. through sidelink) to download and activate the model.
[0331] Figure 13 A mechanism is shown, according to an embodiment, to combine the labeled data in the inference data with the ground truth from one or more sources to obtain labeled data for training the model and / or feature and / or performance indicators.
[0332] In an example, the model can be stored in the OTT server and / or in a network entity. The model delivery can be subject to entity authorization and subscription limitations. Thus, when the UE makes a request to the network, the network entity (e.g. LMF) can need to interact with the UDM / AUSF to check whether the requested model is authorized and / or within the scope of the UE subscription. In addition, the network entity can need to interact with other network functions (like URF and / or NWDAF and / or UDM and / or UDR). If the model exists in the network and the UE has subscription rights and / or authorization to use the model, the network function provides the model to the UE. Otherwise, the network function can signal a fallback solution to the UE and / or indicate a fallback model and / or provide a fallback model.
[0333] The AI / ML model can be trained, for example, at the core network (e.g. NWDAF or other application function) and stored in an entity of the core network or application function (like UDR). Alternatively, the model can be trained, for example, at an entity outside the core network (e.g. at a computing system outside the 5G core network) and can be transferred, for example, to the 5G core network. One implementation is that an application function in the 5G core network allows an external entity storing a trained model to publish the model to an application function within the core network, for example, through a NEF interface. Alternatively, an AF inside the CN can subscribe, for example, to an entity outside the core network (e.g. a server in the data network) and obtain the model by interacting with the server in the data network. As a further alternative, the model can be transferred, for example, through an operation and maintenance interface.
[0334] The network function can interact, for example, with the UDM or a second network function, for example, to determine authorization and / or subscription to request certain data from the UE, and / or to provide data collected from the UE to a third network function or client or server. For example, the UE can store a privacy profile, for example, or one or more network entities can store a privacy profile and / or authorization. For example, on the premise of authorization, this information can be provided to an entity outside the 5G core network and / or to an entity outside the network function that has already obtained this information. As an example, measurements obtained by a network entity associated with the transmission to the UE and / or measurements or information reported by the UE can be transferred to an external client (e.g. a server in the data network) upon authorization. For example, if the UE has refused to share location-related information to an external client, the external client cannot obtain training data from the UE. The privacy profile can be stored in an AF (like AMF or UDM or UDR or AUSF).
[0335] The model delivery can be subject to UE subscription and / or authorization.
[0336] Another embodiment relates to a second case where the inference model is located at least at the UE and the estimator is located at the network (e.g. LMF).
[0337] Similar to the above case, the network can have a specific function model that is more performant that can not be stored at the UE. Given the current conditions, the estimator function at the network side can estimate that other models can be more suitable for the UE, or the network can predict that other models can be needed in the near future. The network can determine that it is more advantageous to store the model at the UE in advance. For example, when a vehicle is driving on a highway or an AGV is running in an industrial workshop, the network can predict the next location and its corresponding model by analyzing the location data.
[0338] When a new model will be a more performant model, the network entity can determine that it can be needed to load the new model at the UE in advance.
[0339] Figure 14 A flowchart is shown of providing one or more AI / ML models from a network to a user equipment according to another embodiment.
[0340] Although some aspects of the described concepts are described in the context of an apparatus, it is clear that these aspects also correspond to a description of the corresponding method - in which each module or device corresponds to a method step or a feature of a method step. Likewise, aspects described in the context of a method step also correspond to a description of a module, item or feature of a corresponding apparatus.
[0341] Various elements and features of the present application can be physically implemented in hardware, in software or in a combination of both software and hardware. For example, embodiments of the present application can be implemented in an environment of a computer system or other processing system. Figure 16 An example of a computer system 600 is shown. The units or modules and the steps of the methods performed by these units can be executed on one or more computer systems 600. The computer system 600 includes one or more processors 602, such as a special purpose or a general purpose digital signal processor. The processor 602 is connected to a communication infrastructure 604, such as a bus or network. The computer system 600 includes a main memory 606, such as a random access memory (RAM), and a secondary memory 608, such as a hard disk drive and / or a removable storage drive. The secondary memory 608 can be used for storing computer programs or other instructions that are loaded into the computer system 600. The computer system 600 can further include a communication interface 610 for communicating with external devices, such as software and data. The communication can take place in the form of electronic, electromagnetic, optical or other signals capable of being processed by the communication interface. The communication can use wires or cables, optical fibers, phone lines, cellular phone links, RF links, and other communication channels 612.
[0342] The terms "computer program medium" and "computer readable medium" generally refer to tangible storage media, such as removable storage units or hard disk drives, installed in hard disk drive units. These computer program products are means for providing software to the computer system 600. The computer programs (also known as computer control logic) are stored in the main memory 606 and / or the secondary memory 508. Computer programs can also be received via the communications interface 610. Such computer programs, when executed, enable the computer system 600 to implement the present application. In particular, the computer programs, when executed, enable the processor 602 to implement the processes of the present application, such as any of the methods described herein. Accordingly, such computer programs can be
[0343] A hardware or software implementation can be accomplished with a digital storage medium, such as a cloud storage, a floppy disc, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium can be computer readable.
[0344] Some embodiments according to the application comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.
[0345] Generally, embodiments of the present application can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer. The program code can be stored on a machine readable carrier.
[0346] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein. In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods when the computer program is executed on a computer.
[0347] A further embodiment of the inventive method is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals can for example be configured to be transferred via a data communication connection, for example via the Internet. A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted for performing one of the methods described herein. A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
[0348] In some embodiments, a programmable logic device (for example a field programmable gate array) can be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array can cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.
[0349] The above-described embodiments are merely representative of the principles of the present application. It will be apparent to those skilled in the art that modifications and variations can be made to the structures and details disclosed herein without departing from the spirit and scope of the application. Thus, the present application is not limited to the particular details discussed herein.
[0350] Abbreviations
[0351] Abbreviations Definitions 3GPP Third Generation Partnership Project 5GC 5G Core Network BS Base Station CSI-RS Channel State Information Reference Signal DMRS Demodulation Reference Signal DOA Direction of Arrival E-CID Enhanced Cell ID eNB Evolved Node B E-SMLC Evolved Serving Mobile Location Center E-UTRA Evolved UMTS Terrestrial Radio Access gNB Next Generation Node B GPS Global Positioning System LMF Location Management Function LMU Location Measurement Unit LPP LTE Positioning Protocol LTE Long Term Evolution NG Next Generation ng-eNB Next Generation eNB NG-RAN gNB or ng-eNB NR New Radio NRPPa New Radio Positioning Protocol A OTDOA Observed Time Difference of Arrival PRS Positioning Reference Signal PTRS Phase Tracking Reference Signal QCL Quasi Co-Location RAN Radio Access Network RP Reception Point RSTD Reference Signal Time Difference RTOA Relative Time of Arrival RTT Round Trip Time SA Single SRS Sounding Reference Signal TDM Time Division Multiplexing TOF Time of Flight TRP Transmission Reception Point RS Reference Signal QCL Quasi Co-Location AoA Angle of Arrival AoD Angle of Departure PAS Power Angular Spectrum NR New Radio gNB Next Generation Node B GPS Global Positioning System LMF Location Management Function LMU Location Measurement Unit LPP LTE Positioning Protocol LTE Long Term Evolution NG Next Generation ng-eNB Next Generation eNB NG-RAN gNB or ng-eNB NR New Radio CA Carrier Aggregation CAM Cooperative Awareness Message DAS Distributed Antenna System DL Downlink FL Frequency Layer FC Frequency Component. This is a BWP of a wideband carrier. GNSS Global Navigation Satellite System OOC Out of Coverage PSFCH Physical Sidelink Feedback Channel P-UE Pedestrian UE: Should not be limited to pedestrians, but rather to any UE RS Reference Signal RE Resource Element SINR Signal to Interference plus Noise Ratio SL Sidelink SPRS, SP-PRS Sidelink Positioning Reference Signal V2X Vehicle to Everything VRU Vulnerable Road User V-UE Vehicle User Equipment BWP Bandwidth Part TEG Timing Error Group ZC (sequence) Zadoff-Chu sequence UE User Equipment UL Uplink Uu (interface) Interface between UEs ToA Time of Arrival TDOA Time Difference of Arrival LOS Line of Sight PRU Positioning Reference Unit ToF Time of Flight
Claims
1. A device for a wireless communication system, in, The apparatus is configured to determine metrics for one or more inactive AI / ML models and / or their functionality, wherein the one or more inactive AI / ML models are suitable for supporting tasks of network entities in user equipment and / or wireless communication systems, and the apparatus is a user equipment or otherwise; wherein the apparatus is configured to determine metrics for the AI / ML models and / or their functionality such that the metrics take into account the benefits of employing the AI / ML models and / or their functionality, and such that the metrics take into account the activation overhead of activating the AI / ML models and / or their functionality. The device is configured to determine whether to activate the AI / ML model and / or its functions based on a metric of the AI / ML model and / or its functionality.
2. The apparatus according to claim 1, The device is configured to activate the AI / ML model if it has been determined that the AI / ML model should be activated.
3. The apparatus according to claim 2, The device is a user equipment; and the device is configured to perform a task using an AI / ML model if it is determined that an AI / ML model should be activated.
4. The apparatus according to claim 1, The device described therein is different from the user equipment; and the device is configured to: if it has been determined that the AI / ML model should be activated, transmit information to the user equipment to activate the AI / ML model; or The device is configured to transmit information to other devices in the wireless communication system to activate the AI / ML model if it has been determined that the AI / ML model should be activated.
5. The apparatus according to any one of the preceding claims, in, The functions include specific configurations, inputs, or outputs of the AI / ML model within the device.
6. The apparatus according to any one of the preceding claims, The device is configured to determine metrics for AI / ML models across multiple models within the same function, wherein the AI / ML models within the same function share common configurations, inputs, and outputs. The device is configured to evaluate the benefits of using AI / ML models within the same functionality and the activation overhead required for each model. The device is configured to determine whether to activate or deactivate an AI / ML model within the same function based on a defined metric and the overall benefits and costs involved.
7. The apparatus according to any one of the preceding claims, in, The device is configured to determine a metric for an interconnected set of AI / ML models within the same function, wherein the interconnected models collaborate to perform a specific task. The device is configured to analyze the benefits of using interconnected AI / ML models and the activation overhead required for each individual model. The device is configured to determine whether to activate or deactivate a set of interconnected AI / ML models within the same function, based on metrics and the collective benefits and costs involved in utilizing interconnected AI / ML models.
8. The apparatus according to any one of the preceding claims, in, The device is configured to determine a metric for an AI / ML model and / or its functionality within a current function, wherein the current function does not include features that support AI / ML characteristics. The device is configured to evaluate the benefits of employing AI / ML models and / or functions within the current functionality, as well as the activation overhead required for each model and / or function. The device is configured to determine a metric for an AI / ML model and / or its functionality within a target function, wherein the target function includes at least one function that supports AI / ML features. The device is configured to evaluate the benefits of employing AI / ML models and / or features within a target function, as well as the activation overhead required for each model and / or feature; and the device is configured to determine whether to activate the target function based on determined metrics and evaluated benefits and activation overhead, as well as the total benefits and overhead involved in activating features supporting AI / ML features.
9. The apparatus according to any one of the preceding claims, in, The device is configured to determine a measure of the AI / ML model and / or its functionality within a current function, wherein the current function includes features that support AI / ML characteristics.
10. The apparatus according to any one of the preceding claims, in, The device is configured to determine a metric for an AI / ML model and / or its functionality, such that the metric takes into account the benefits and / or costs of deactivating the currently employed AI / ML model and / or its currently employed functionality.
11. The apparatus according to any one of the preceding claims, in, The device is configured to determine a metric for an AI / ML model and / or its functionality, such that the metric takes into account the benefits and / or costs of switching from the currently used AI / ML model and / or its functionality to the AI / ML model and / or its functionality.
12. The apparatus according to any one of the preceding claims, in, The device is configured to determine, based on metrics of the AI / ML model and / or its functionality, whether to activate the AI / ML model and / or its functionality, or whether to use non-AI / ML (such as conventional) functionality, for example, as a fallback.
13. The apparatus according to any one of the preceding claims, The activation overhead for activating AI / ML models and / or their functions includes one or more of the following: - Calculate the cost, such as the number of processing cycles, the number of multiplications, etc. - Signaling costs, such as the amount of data in signaling messages that need to be exchanged between user equipment and units of the wireless communication system. - Activation time required to activate the model, etc. -Increased latency, - Monitoring costs, - A combination of the above items.
14. The apparatus according to claim 13, in, The activation overhead for activating AI / ML models and / or their functionality includes monitoring costs, which depend on the availability of PRUs for ground truth labels in positioning, and / or on the frequency of measurements of all beams in the codebook in beam management.
15. The apparatus according to any one of the preceding claims, in, The device is configured to determine a metric for an AI / ML model and / or its functionality based on at least one of the following: Information about the currently active function / model and its attributes. Information regarding the performance or related QoS of the currently active feature / model. Information regarding cell ID and / or area ID and / or dataset ID. Potential performance requirements and / or cost constraints, The input data for the currently used AI / ML models, Measurements related to the applicability of the function, such as SNR level, UE speed, Doppler, beam codebook type, PRS identifier, model pairing information of the two-sided model, and / or, for example, network synchronization error, and / or, for example, UE / gNB RX and TX timing errors. Information about alarms from other model-monitored entities, and / or results of general monitoring metric calculations from other model-monitored entities. Information about the amount of time the currently active model has been activated. Advanced features / post-processing information about the UE state, such as UE orientation / location / velocity, and predicted future UE trajectory. Supplemental information from the network, including general properties of the radio environment and other issues reported by the UE.
16. The apparatus according to any one of the preceding claims, in, One or more AI / ML models include two or more AI / ML models.
17. The apparatus according to claim 16, in, The device is a user equipment. The device is configured to select and / or activate one of the two or more AI / ML models, and / or switch from one of the two or more AI / ML models to another of the two or more AI / ML models, depending on which of the network units of the wireless communication system activates.
18. The apparatus according to claim 16 or 17, in, The device is a user equipment. The device is configured to receive information about rules from a network unit of a wireless communication system, wherein the information relates to selecting and / or activating one of the two or more AI / ML models, and / or switching from one of the two or more AI / ML models to another of the two or more AI / ML models.
19. The apparatus according to any one of claims 16 to 18, in, The device is a user equipment. The device is configured to request permission from a network element of a wireless communication system to select and / or activate one of the two or more AI / ML models, and / or switch from one of the two or more AI / ML models to another of the two or more AI / ML models; and The device is configured to, upon receiving permission from the network unit, select and / or activate one of the two or more AI / ML models, and / or switch from one of the two or more AI / ML models to the other of the two or more AI / ML models.
20. The apparatus according to any one of claims 16 to 19, in, The device is a user equipment. The device is configured to select and / or activate one of the two or more AI / ML models, and / or switch from one of the two or more AI / ML models to another of the two or more AI / ML models, based on selection information received from a network unit of a wireless communication system.
21. The apparatus according to any one of claims 16 to 20, in, The device is configured to determine the performance of one or more AI / ML models for supporting tasks of the user equipment and / or network entities, based on the current location of the user equipment.
22. The apparatus according to claim 21, in, The two or more AI / ML models are each applicable to a geographic region; and The device is configured to: if the user device's current location is within a geographic area applicable to two of the two or more AI / ML models, determine whether to activate the first or second of the two AI / ML models by determining the metric of each AI / ML model. For each of the two AI / ML models, the metrics of the AI / ML model take into account the benefits of using the AI / ML model and / or its functionality, and the metrics take into account the activation costs of activating the AI / ML model and / or its functionality.
23. The apparatus according to claim 22, The device mentioned is a user equipment, and The device is configured to determine the metric for each of the two AI / ML models if it has been determined that it is located in a geographic region to which the two AI / ML models apply.
24. The apparatus according to claim 22, The device described therein is different from user equipment, and in, The device is configured to receive information from the user equipment regarding the geographic region where the two AI / ML models are applicable, and The device is configured to determine a metric for each of the two AI / ML models in response to receiving the information.
25. The apparatus according to any one of claims 16 to 24, in, The device is configured to determine a metric for each of the two or more AI / ML models and / or their functionality, based on the characteristics of the current environment of the user device.
26. The apparatus according to any one of claims 16 to 25, in, The device is configured to determine a metric for each of the two or more AI / ML models and / or their functionality, based on the characteristics of the user equipment and / or the characteristics of the network entity.
27. The apparatus according to claim 26, in, The device is configured to determine a metric for each of the two or more AI / ML models and / or their functionality based on the state of the user device’s battery power and / or based on an activated battery power saving mode.
28. The apparatus according to any one of claims 16 to 27, in, The apparatus is configured to determine a metric of each of the two or more AI / ML models and / or its functionality based on the transmission characteristics of the transmission between the user equipment and the network, and / or based on the transmission characteristics of the transmission between the user equipment and another user equipment, and / or based on radio environment attributes.
29. The apparatus according to any one of claims 16 to 28, in, The device is configured to determine a metric for each of the two or more AI / ML models and / or their functionality based on the current state of the user device and one or more possible future states of the user device.
30. The apparatus according to any one of claims 16 to 29, in, The device is configured to determine a metric for each of the two or more AI / ML models and / or their functionality based on two or more possible future actions of the user device.
31. The apparatus according to any one of claims 16 to 30, in, The device is configured to determine each of the two or more AI / ML models and / or a metric of its functionality based on a reward function that returns a real value indicating the performance of one of the two or more AI / ML models. For example, when the user device is in one of the current state and one or more future states, perform one of two or more possible future actions.
32. The apparatus according to claim 31, in, The reward function returns one of the following values: For beam management, the value indicates performance, such as the top K accuracies of an AI / ML model or the system throughput achieved using the selected beam. For CSI compression, values indicating performance, such as throughput or the similarity between the decoder output and the target CSI, are used. For direct / assisted positioning, values indicating prediction accuracy are derived, for example, from PRU evaluations that can generate true ground labels.
33. The apparatus according to any one of claims 16 to 32, in, The apparatus is configured to determine a metric for each of the two or more AI / ML models and / or its functionality based on a cost function that considers the overhead or computational cost of activating a particular AI / ML model among the one or more AI / ML models, and / or the overhead or computational cost of activating the functionality of a particular AI / ML model, and / or the overhead or computational cost of switching from the current AI / ML model among the one or more AI / ML models to another AI / ML model among the one or more AI / ML models.
34. The apparatus according to claim 31 or 32 and claim 33, in, The device is configured to determine a metric for each of the two or more AI / ML models and / or their functionality based on a reward function and a cost function.
35. The apparatus according to claim 34, in, The device is configured to determine a metric of the functionality of each of the two or more AI / ML models and / or its features by determining a linear combination of a reward function and a cost function.
36. The apparatus according to any one of claims 31 to 35, in, The reward function returns a penalty value for switching from one of the two or more AI / ML models to another.
37. The apparatus according to claim 36, in, The reward function returns a value that penalizes repeated switching from one of the two or more AI / ML models to another.
38. The apparatus according to any one of claims 34 to 37, and further according to claim 33. The cost function returns a penalty value for switching from one of the two or more AI / ML models to another.
39. The apparatus according to claim 38, The cost function returns a value that penalizes repeated switching from one of the two or more AI / ML models to another.
40. The apparatus according to any one of the preceding claims, in, Each of the one or more AI / ML models is implemented by one or more neural networks.
41. The apparatus according to any one of the preceding claims, in, The task is the location task of the user equipment and / or the network entity.
42. The apparatus according to any one of the preceding claims, in, The task is a management or configuration task of the user equipment and / or the network entity, such as a beam management task of the user equipment and / or the network entity.
43. The apparatus according to any one of the preceding claims, in, The task is an encoding task of the user equipment and / or the network entity, or a compression task of the user equipment and / or the network entity, such as a task for compressing channel state information.
44. An apparatus for a wireless communication system, in, The device is configured to activate one or more AI / ML models and / or their functionalities; wherein the one or more AI / ML models are adapted to support the tasks of network entities in user equipment and / or wireless communication systems; wherein the device is a user equipment or the device is different from a user equipment; wherein whether the AI / ML model and / or its functionalities are activated depends on the measurement of the AI / ML model and / or its functionalities. The metric described therein takes into account the benefits of employing an AI / ML model and / or its functionality, and the metric described therein takes into account the activation overhead of activating the AI / ML model and / or its functionality.
45. The apparatus according to claim 44, The device described herein implements the device according to any one of claims 1 to 43.
46. The apparatus according to claim 44, The device described therein does not implement the device according to any one of claims 1 to 43. The device is configured to receive information about an AI / ML model among one or more AI / ML models to be activated from the device according to any one of claims 1 to 43.
47. The apparatus according to any one of claims 44 to 46, wherein the apparatus is a user equipment.
48. The apparatus according to any one of claims 44 to 46, The device described is not user equipment. The device is configured to provide output from an AI / ML model to a user device and / or a network entity to support the user device and / or the network entity in performing tasks.
49. A user equipment for a wireless communication system, in, The user equipment is configured to receive information from other devices of the wireless communication system regarding the output of one or more AI / ML models and / or their functions, wherein the one or more AI / ML models are adapted to support the tasks of the user equipment and / or network entities of the wireless communication system; Whether an AI / ML model and / or its functionality has been activated by other devices depends on a metric of the AI / ML model and / or its functionality, wherein the metric takes into account the benefits of using the AI / ML model and / or its functionality, and wherein the metric takes into account the activation overhead of activating the AI / ML model and / or its functionality.
50. The user equipment according to claim 49, in, Other devices are any one of claims 44 to 46 or the device according to claim 48.
51. A wireless communication system, comprising: The apparatus according to any one of claims 1 to 43, and User equipment.
52. The wireless communication system according to claim 51, The wireless communication system further includes the apparatus according to any one of claims 44 to 46 or according to claim 48.
53. The wireless communication system according to claim 51 or 52, in, The user equipment is the user equipment as described in claim 49 or 50.
54. A wireless communication system, comprising: The first device according to any one of claims 1 to 43, and The second device according to any one of claims 1 to 43, The first device is configured to select and / or activate one of one or more AI / ML models based on the selection and / or activation of one of one or more AI / ML models based on switching from one of one or more AI / ML models to another of one or more AI / ML models.
55. A method for a wireless communication system, in, The method includes determining metrics for one or more inactive AI / ML models and / or their functionality, wherein the one or more inactive AI / ML models are suitable for supporting tasks of network entities in a user equipment and / or wireless communication system, wherein the method is performed by a user equipment or by means other than a user equipment in the wireless communication system; wherein the determination of metrics for AI / ML models and / or their functionality considers the benefits of employing AI / ML models and / or their functionality, and considers the activation overhead for activating AI / ML models and / or their functionality. The method includes determining whether to activate an AI / ML model and / or its functions based on metrics of the AI / ML model and / or its functionality.
56. A method for a wireless communication system, in, The method includes activating one or more AI / ML models and / or their functions; wherein the one or more AI / ML models are adapted to support tasks of a user equipment and / or a network entity of the wireless communication system; wherein the method is performed by a user equipment or by a device in the wireless communication system other than a user equipment; wherein whether the AI / ML model and / or its functions are activated depends on a metric of the AI / ML model and / or its functions. The metric described therein takes into account the benefits of using an AI / ML model and / or its functionality, and the metric described therein takes into account the activation overhead of activating the AI / ML model and / or its functionality.
57. A method for a wireless communication system, in, The method includes receiving information from other devices of a wireless communication system regarding the output of one or more AI / ML models and / or the output of their functions, wherein the one or more AI / ML models are adapted to support the tasks of the user equipment and / or network entities of the wireless communication system; Whether an AI / ML model and / or its functionality has been activated by other devices depends on a metric of the AI / ML model and / or its functionality, wherein the metric takes into account the benefits of using the AI / ML model and / or its functionality, and wherein the metric takes into account the activation overhead of activating the AI / ML model and / or its functionality.
58. A computer program, when executed by a computer or signal processor, for implementing the method as described in any one of claims 55 to 57.