Method and apparatus for providing ai and ml media service in a wireless communication system
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2024-09-26
- Publication Date
- 2026-04-22
AI Technical Summary
Current wireless communication systems face challenges in efficiently delivering and processing artificial intelligence (AI) and machine learning (ML) media services, particularly in terms of compatibility between user equipment (UE) devices and network providers, processing power limitations on UE devices, and privacy concerns related to data transmission.
The proposed method and apparatus involve a communication method in a wireless communication system that enables the delivery and management of AI/ML models between a network and user equipment (UE), allowing for AI/ML media services. This includes negotiating a configuration for AI split inference, establishing data pipelines for AI model delivery and intermediate data transmission, and performing split inference processes while ensuring privacy by determining split points based on privacy policies.
This solution enables efficient delivery and processing of AI/ML media services by optimizing AI model delivery and inference processes, addressing compatibility and processing power limitations, and ensuring privacy by controlling data transmission across the network and UE.
Smart Images

Figure KR2024014609_10042025_PF_FP_ABST
Abstract
Description
METHOD AND APPARATUS FOR PROVIDING AI AND ML MEDIA SERIVCE IN A WIRELESS COMMUNICATION SYSTEM
[0001] The present disclosure relate generally to a wireless communication system, and more particularly to a method and apparatus for presenting artificial intelligence (AI) and machine learning (ML) media services in a wireless communication system.
[0002] 5th generation (5G) mobile communication technologies define broad frequency bands such that high transmission rates and new services are possible, and can be implemented not only in “Sub 6 gigahertz (GHz)” bands such as 3.5GHz, but also in “Above 6GHz” bands referred to as millimeter wave (mmWave) including 28GHz and 39GHz. In addition, it has been considered to implement 6th generation (6G) mobile communication technologies (referred to as Beyond 5G systems) in terahertz (THz) bands (for example, 95GHz to 3THz bands) in order to accomplish transmission rates fifty times faster than 5G mobile communication technologies and ultra-low latencies one-tenth of 5G mobile communication technologies.
[0003] At the beginning of the development of 5G mobile communication technologies, in order to support services and to satisfy performance requirements in connection with enhanced Mobile BroadBand (eMBB), Ultra Reliable Low Latency Communications (URLLC), and massive Machine-Type Communications (mMTC), there has been ongoing standardization regarding beamforming and massive multi input multi output (MIMO) for mitigating radio-wave path loss and increasing radio-wave transmission distances in mmWave, supporting numerologies (for example, operating multiple subcarrier spacings) for efficiently utilizing mmWave resources and dynamic operation of slot formats, initial access technologies for supporting multi-beam transmission and broadbands, definition and operation of BandWidth Part (BWP), new channel coding methods such as a Low Density Parity Check (LDPC) code for large amount of data transmission and a polar code for highly reliable transmission of control information, L2 pre-processing, and network slicing for providing a dedicated network specialized to a specific service.
[0004] Currently, there are ongoing discussions regarding improvement and performance enhancement of initial 5G mobile communication technologies in view of services to be supported by 5G mobile communication technologies, and there has been physical layer standardization regarding technologies such as Vehicle-to-everything (V2X) for aiding driving determination by autonomous vehicles based on information regarding positions and states of vehicles transmitted by the vehicles and for enhancing user convenience, New Radio Unlicensed (NR-U) aimed at system operations conforming to various regulation-related requirements in unlicensed bands, new radio (NR) user equipment (UE) Power Saving, Non-Terrestrial Network (NTN) which is UE-satellite direct communication for providing coverage in an area in which communication with terrestrial networks is unavailable, and positioning.
[0005] Moreover, there has been ongoing standardization in air interface architecture / protocol regarding technologies such as Industrial Internet of Things (IIoT) for supporting new services through interworking and convergence with other industries, Integrated Access and Backhaul (IAB) for providing a node for network service area expansion by supporting a wireless backhaul link and an access link in an integrated manner, mobility enhancement including conditional handover and Dual Active Protocol Stack (DAPS) handover, and two-step random access for simplifying random access procedures (2-step random access channel (RACH) for NR). There also has been ongoing standardization in system architecture / service regarding a 5G baseline architecture (for example, service based architecture or service based interface) for combining Network Functions Virtualization (NFV) and Software-Defined Networking (SDN) technologies, and Mobile Edge Computing (MEC) for receiving services based on UE positions.
[0006] As 5G mobile communication systems are commercialized, connected devices that have been exponentially increasing will be connected to communication networks, and it is accordingly expected that enhanced functions and performances of 5G mobile communication systems and integrated operations of connected devices will be necessary. To this end, new research is scheduled in connection with eXtended Reality (XR) for efficiently supporting Augmented Reality (AR), Virtual Reality (VR), Mixed Reality (MR) and the like, 5G performance improvement and complexity reduction by utilizing Artificial Intelligence (AI) and Machine Learning (ML), AI service support, metaverse service support, and drone communication.
[0007] Furthermore, such development of 5G mobile communication systems will serve as a basis for developing not only new waveforms for providing coverage in terahertz bands of 6G mobile communication technologies, multi-antenna transmission technologies such as Full Dimensional MIMO (FD-MIMO), array antennas and large-scale antennas, metamaterial-based lenses and antennas for improving coverage of terahertz band signals, high-dimensional space multiplexing technology using Orbital Angular Momentum (OAM), and Reconfigurable Intelligent Surface (RIS), but also full-duplex technology for increasing frequency efficiency of 6G mobile communication technologies and improving system networks, AI-based communication technology for implementing system optimization by utilizing satellites and Artificial Intelligence (AI) from the design stage and internalizing end-to-end AI support functions, and next-generation distributed computing technology for implementing services at levels of complexity exceeding the limit of UE operation capability by utilizing ultra-high-performance communication and computing resources.
[0008] The above information is presented as background information only to assist with an understanding of the disclosure. No determination has been made, and no assertion is made, as to whether any of the above might be applicable as prior art with regard to the disclosure.
[0009] Based on the above discussion, various embodiments of the present disclosure are to present a method and apparatus for providing artificial intelligence (AI) and machine learning (ML) media services in a wireless communication system.
[0010] According to an aspect of an exemplary embodiment, there is provided a communication method in a wireless communication.
[0011] Various embodiments of the present disclosure may present a method and apparatus for presenting artificial intelligence (AI) and machine learning (ML) media services in a wireless communication system.
[0012] Effects obtainable in the present disclosure are not limited to the effects mentioned above, and other effects not mentioned may be clearly understood by those skilled in the art from the description below.
[0013] FIG. 1 illustrates the overall 5G media streaming architecture according to various embodiments of the present disclosure.
[0014] FIG. 2 illustrates the 5G media streaming general architecture according to various embodiments of the present disclosure.
[0015] FIG. 3 illustrates a procedure for media downlink streaming according to various embodiments of the present disclosure.
[0016] FIG. 4 illustrates a baseline procedure describing the establishment of a unicast media downlink streaming session, according to various embodiments of the present disclosure.
[0017] FIG. 5 illustrates an example of an AI / ML media service scenario where an AI / ML model must be delivered from a network to a UE, according to various embodiments of the present disclosure.
[0018] FIG. 6 illustrates an example of a scenario where an AI model is delivered to a UE and where media is also streamed to the UE, according to various embodiments of the present disclosure.
[0019] FIG. 7 illustrates an example of a scenario where inferencing required for an AI media service is split between a network and a UE, according to various embodiments of the present disclosure.
[0020] FIG. 8 illustrates a basic architecture for a split AI inferencing scenario, according to various embodiments of the present disclosure.
[0021] FIG. 9 illustrates another basic architecture for a split AI inferencing scenario, according to various embodiments of the present disclosure.
[0022] FIG. 10 illustrates an example of an AI for media (AI4Media) architecture, according to various embodiments of the present disclosure.
[0023] FIG. 11 illustrates an embodiment of this disclosure showing a procedure for the delivery of an AI model with configurations between the network and UE, according to various embodiments of the present disclosure.
[0024] FIG. 12 illustrates a scenario of split inferencing where certain data required to be sent between the UE and network may raise privacy concerns, according to various embodiments of the present disclosure.
[0025] FIG. 13 illustrates an embodiment of this invention as a scenario which can avoid the exposure of privacy sensitive information outside the UE device, by keeping private data local o the UE device, according to various embodiments of the present disclosure.
[0026] FIG. 14 illustrates an AI model with multiple possible split points at the different layers of the AI model, according to various embodiments of the present disclosure.
[0027] FIG. 15 illustrates an embodiment of this disclosure showing a procedure for the delivery of an AI model with configurations between the network and UE, according to various embodiments of the present disclosure.
[0028] FIG. 16 illustrates an embodiment of this disclosure showing a detail procedure for the delivery of an AI model with configurations between the network and UE, according to various embodiments of the present disclosure.
[0029] FIG. 17 illustrates an embodiment of this disclosure showing a detail procedure for the delivery of an AI model with configurations between the network and UE, according to various embodiments of the present disclosure.
[0030] FIG. 18 illustrates a signalling message format for a step 1 of Fig. 16, according to various embodiments of the present disclosure.
[0031] FIG. 19 illustrates a signalling message format for a step 3 of Fig. 16, according to various embodiments of the present disclosure.
[0032] FIG. 20 illustrates a signalling message format for a step 6 of Fig. 16, according to various embodiments of the present disclosure.
[0033] FIG. 21 illustrates a signalling message format for a step 8 of Fig. 16, according to various embodiments of the present disclosure.
[0034] FIG. 22 illustrates a signalling message format for a step 9 of Fig. 16, according to various embodiments of the present disclosure.
[0035] According to various embodiments of the present disclosure, it presents a method performed by a user equipment (UE) in a wireless communication system, the method comprising: receiving, from a first network entity, information regarding at least one artificial intelligence (AI) model for a service, identifying AI data inferencing capabilities for the AI model, requesting an AI split inference, negotiating a configuration for the service and a split point of the AI split inference; establishing an AI model data deliver pipeline for delivering the AI model; establishing an intermediate data deliver pipeline for delivering intermediated data, performing the split inference process using the AI model data deliver pipeline and the intermediate data deliver pipeline, transmitting a status report for service, and updating the configuration based on the status report.
[0036] The information regarding at least one AI model includes a uniform resource locator (URL) to obtain a list the at least one AI model.
[0037] The method may further include transmitting a feedback message based on a predefined time interval or a predefined threshold, and updating the configuration and the split point based on the feedback message.
[0038] The feedback message includes at least one of information on a network capability parameter, information on a latency and information on an AI model parameter.
[0039] The split point of the AI split inference is determined based on a privacy police.
[0040] According to various embodiments of the present disclosure, it presents a method performed by a first network entity in a wireless communication system, the method comprising, transmitting, to a user equipment (UE), information regarding at least one artificial intelligence (AI) model for a service, identifying AI data inferencing capabilities for the AI model, requesting an AI split inference, negotiating a configuration for the service and a split point of the AI split inference, establishing an AI model data deliver pipeline for delivering the AI model, establishing an intermediate data deliver pipeline for delivering intermediated data, performing the split inference process using the AI model data deliver pipeline and the intermediate data deliver pipeline, receiving a status report for service, and updating the configuration based on the status report.
[0041] The information regarding at least one AI model includes a uniform resource locator (URL) to obtain a list the at least one AI model.
[0042] The method may further include receiving a feedback message based on a predefined time interval or a predefined threshold, and updating the configuration and the split point based on the feedback message.
[0043] The feedback message includes at least one of information on a network capability parameter, information on a latency and information on an AI model parameter.
[0044] The split point of the AI split inference is determined based on a privacy police.
[0045] According to another embodiment of the disclosure, a user equipment (UE) in a wireless communication system, the UE comprising, at least one transceiver; and at least one processor operatively coupled with the at least one transceiver, wherein the at least one processor is configured to: receive, from a first network entity, information regarding at least one artificial intelligence (AI) model for a service, identify AI data inferencing capabilities for the AI model, request an AI split inference, negotiate a configuration for the service and a split point of the AI split inference, establish an AI model data deliver pipeline for delivering the AI model, establish an intermediate data deliver pipeline for delivering intermediated data, perform the split inference process using the AI model data deliver pipeline and the intermediate data deliver pipeline, transmit a status report for service, and update the configuration based on the status report.
[0046] The information regarding at least one AI model includes a uniform resource locator (URL) to obtain a list the at least one AI model.
[0047] The at least one processor is further configured to transmit a feedback message based on a predefined time interval or a predefined threshold, and update the configuration and the split point based on the feedback message.
[0048] The feedback message includes at least one of information on a network capability parameter, information on a latency and information on an AI model parameter.
[0049] The split point of the AI split inference is determined based on a privacy police.
[0050] According to another embodiment of the disclosure, A first network entity in a wireless communication system, the first network entity comprising: at least one transceiver, and at least one processor operatively coupled with the at least one transceiver, wherein the at least one processor is configured to: transmit, to a user equipment (UE), information regarding at least one artificial intelligence (AI) model for a service, identify AI data inferencing capabilities for the AI model, request an AI split inference, negotiate configuration for the service and a split point of the AI split inference, establish an AI model data deliver pipeline for delivering the AI model, establish an intermediate data deliver pipeline for delivering intermediated data, perform the split inference process using the AI model data deliver pipeline and the intermediate data deliver pipeline, receive a status report for service, and update the configuration based on the status report.
[0051] The information regarding at least one AI model includes a uniform resource locator (URL) to obtain a list the at least one AI model.
[0052] The at least one processor is further configured to receive a feedback message based on a predefined time interval or a predefined threshold, and update the configuration and the split point based on the feedback message.
[0053] The feedback message includes at least one of information on a network capability parameter, information on a latency and information on an AI model parameter.
[0054] The split point of the AI split inference is determined based on a privacy police.
[0055] Hereinafter, the operation principle of the disclosure will be described in detail with reference to the accompanying drawings. In the following description of the disclosure, a detailed description of known functions or configurations incorporated herein will be omitted when it is determined that the description may make the subject matter of the disclosure unnecessarily unclear. The terms which will be described below are terms defined in consideration of the functions in the disclosure, and may be different according to users, intentions of the users, or customs. Therefore, the definitions of the terms should be made based on the contents throughout the specification.
[0056] In the following description, terms for identifying access nodes, terms referring to network entities, terms referring to messages, terms referring to interfaces between network entities, terms referring to various identification information, and the like are illustratively used for the sake of descriptive convenience. Therefore, the disclosure is not limited by the terms as used below, and other terms referring to subjects having equivalent technical meanings may be used.
[0057] In the following description, a base station is an entity that allocates resources to terminals, and may be at least one of a gNode B, an eNode B, a Node B, a base station (BS), a wireless access unit, a base station controller, and a node on a network. A terminal may include a user equipment (UE), a mobile station (MS), a cellular phone, a smartphone, a computer, or a multimedia system capable of performing communication functions. In the disclosure, a “downlink (DL)” refers to a radio link via which a base station transmits a signal to a terminal, and an “uplink (UL)” refers to a radio link via which a terminal transmits a signal to a base station. Further, in the following description, LTE or LTE-A systems may be described by way of example, but the embodiments of the disclosure may also be applied to other communication systems having similar technical backgrounds or channel types. Examples of such communication systems may include 5th generation mobile communication technologies (5G, new radio, and NR) developed beyond LTE-A, and in the following description, the “5G” may be the concept that covers the exiting LTE, LTE-A, or other similar services. In addition, based on determinations by those skilled in the art, the embodiments of the disclosure may also be applied to other communication systems through some modifications without significantly departing from the scope of the disclosure. Herein, it will be understood that each block of the flowchart illustrations, and combinations of blocks in the flowchart illustrations, can be implemented by computer program instructions.
[0058] These computer program instructions can be provided to a processor of a general-purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart block or blocks. These computer program instructions may also be stored in a computer usable or computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer usable or computer-readable memory produce an article of manufacture including instruction means that implement the function specified in the flowchart block or blocks. The computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions that execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.
[0059] Furthermore, each block of the flowchart illustrations may represent a module, segment, or portion of code, which includes one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks may occur out of the order. For example, two blocks shown in succession may in fact be executed substantially concurrently or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. As used in the embodiments of the disclosure, the term “unit” refers to a software element or a hardware element, such as a Field Programmable Gate Array (FPGA) or an Application Specific Integrated Circuit (ASIC), which performs a predetermined function. However, the “unit” does not always have a meaning limited to software or hardware. The “unit” may be constructed either to be stored in an addressable storage medium or to execute one or more processors. Therefore, the “unit” includes, for example, software elements, object-oriented software elements, class elements or task elements, processes, functions, properties, procedures, sub-routines, segments of a program code, drivers, firmware, micro-codes, circuits, data, database, data structures, tables, arrays, and parameters. The elements and functions provided by the “unit” may be either combined into a smaller number of elements, or a “unit”, or divided into a larger number of elements, or a “unit”. Moreover, the elements and “units” or may be implemented to reproduce one or more CPUs within a device or a security multimedia card. Furthermore, the “unit” in the embodiments may include one or more processors.
[0060] In the following description, the disclosure will be described using terms and names defined in the 5GS and NR standards, which are the latest standards specified by the 3rd generation partnership project (3GPP) group among the existing communication standards, for the convenience of description. However, the disclosure is not limited by these terms and names, and may be applied in the same way to systems that conform other standards. In particular, the disclosure may be applied to the 3GPP 5GS / NR (5th generation mobile communication standards).
[0061] According to various embodiments of the present disclosure, 5G network systems for multimedia, architectures and procedures for AI / ML model transfer and delivery over 5G, AI / ML model transfer and delivery over 5G for AI enhanced multimedia services, split AI / ML inferencing between UE and 5G network, split AI / ML inference configuration, negotiating multi-split AI / ML inference between entities are provided.
[0062] Artificial Intelligence (AI) is a general concept defining the capability for a system to act based on 2 major conditions:
[0063] - The context in which a task has to be done, meaning the value or state of different input parameters.
[0064] - The past experience of achieving the same task with different parameter values and the record of potential success with each parameter value.
[0065] Machine Learning (ML) is often described as a subset of AI, in which an application has the capacity to learn from the past experience. This learning feature usually starts with an initial training phase so as to ensure a minimum level of performance when it is placed into service.
[0066] Recently, AI / ML has been introduced and generalized in media related applications, ranging from legacy applications such as image classification, speech / face recognition, to more recent ones such as video quality enhancement. As research into this field matures, more and more complex AI / ML-based applications requiring higher computational processing can be expected; such processing involves dealing with significant amounts of data not only for the inputs and outputs into the AI / ML models, but also for the increasing data size and complexity of the AI / ML models themselves. This growing amount of AI / ML related data, together with a need for supporting processing intensive mobile applications (such as VR, AR / MR, gaming, and more), highlights the importance of handling certain aspects of AI / ML processing by the server over 5G system, in order to meet the required latency requirements of various applications.
[0067] Current implementations of AI / ML are mainly proprietary solutions, enabled via applications without compatibility with other market solutions. In order to support AI / ML for multimedia applications over 5G, AI / ML models should support compatibility between UE devices and application providers from different MNOs. Not only this, but AI / ML model delivery for AI / ML media services should support media context, UE status, and network status based selection and delivery of the AI / ML model. The processing power of UE devices is also a limitation for AI / ML media services, since next generation media services, such as AR, are typically consumed on lightweight, low processing power devices, such as AR glasses, for which long battery life is also a major design hurdle / limitation.
[0068] This invention extends the current frameworks and architectures for 5G media streaming (5GMS) in order to support AI / ML media services, with several embodiments of different possible architectures, each enabling:
[0069] - The delivery of AI / ML models from the network to the UE for multimedia services.
[0070] - The selection, configuration, and management of said AI / ML models and their delivery by newly defined network and UE entities, which can consider the 5G network status, cloud / edge AI media inferencing capabilities and functions, UE processing / runtime status and / or capability and functions, and media characteristics, as the input for these decisions related to AI / ML model delivery, and AI / ML model split delivery decisions.
[0071] - The data and information shared between the network and UE, used for the discovery of capabilities and functions, as well as the negotiation and configuration of the split AI media inference process. Such information may include information related to network and / or UE processing capabilities, and / or information describing / or originating from the characteristics of the AI model to be delivered and used for the AI / ML media service.
[0072] - Taking into concern privacy aspects of the different data components for the split AI / ML media service, and enabling a multiple split configurations to avoid sending privacy sensitive information, including the calibration of multiple split points
[0073] - On successful capability / function discovery, as well as negotiation for the split configuration, the delivery of the AI model may occur in a progressive download manner, where the AI model data is divided into independently deliverable and inference-able subsets in advance. This delivery mechanism may allow for: simultaneous inferencing to start before the complete download of the AI model data, dynamic configuration of split inferencing without further data delivery (since the UE already receives the whole AI model), and efficient partial updates of the AI model during the AI media service (since certain parts for update can be easily identified and replaced, using the granularity of the AI model data subset).
[0074] - Enabling the configuration of dynamic split decisions, wherein the split configuration for AI / ML can be re-negotiated and updated depending on the certain feedback mechanisms which may be defined by thresholds or other related factors
[0075] The following is enabled by this invention:
[0076] Network capability / function / status, UE capability / function / status and multimedia context driven AI / ML model selection and AI / ML model split inference decision / negotiation, delivery and management between network and UE for AI multimedia services
[0077] Figure 1 shows the overall 5G Media Streaming Architecture in TS 26.501, representing the specified 5GMS functions within the 5GS as defined TS 23.501.
[0078] Figure 2 shows the 5G Media Streaming General Architecture from TS 26.501, identifying which media streaming functional entities and interfaces are specified within the specification.
[0079] Figure 4 shows, for reference, the baseline procedure describing the establishment of a unicast media downlink streaming session as defined by TS 26.501.
[0080] Figure 5 shows a simple AI / ML media service scenario where an AI / ML model is required to be delivered from the network to the UE (end device). Upon receiving the AI model, the UE device performs the inferencing of the model, feeding the relevant media as an input into the AI model.
[0081] A typical example:
[0082] - John is in Seoul for his summer vacation, and he is in Jamsil wanting to visit Lotte Tower for sightseeing. John cannot read Korean, and finds it difficult to navigate his way to Lotte Tower.
[0083] - John takes out his mobile phone (UE), and opens an augmented reality navigation service on it. His network operator provides the service via 5G, and through the analysis of different information, a suitable AI model is delivered to his mobile phone. Such information includes information available from the network, such as John’s UE’s location, his charging policy, network availability and conditions (bandwidth, latency) etc, his UE’s processing capabilities and status, as well as the media properties which will be used as the input to the AI model.
[0084] - Once the AI model is delivered to John’s phone, the AR navigation service initiates the camera on the phone to capture the John’s surroundings.
[0085] - The captured video from the phone’s camera is fed as the input into the AI model, and the AI model inferencing is initiated.
[0086] - The output of the AI model provides direction labels (such as navigation arrows) which are shown as overlays in the phone’s screen live camera in order to guide John to Lotte Tower. Road signs in Korean are also overlayed by English labels output from the AI model.
[0087] Figure 6 shows a scenario where an AI model is delivered to the UE, and also where media (such as video) is also streamed to the UE. In the UE, the streamed video is fed as an input into the received AI model for processing.
[0088] The AI model may perform any media related processing, for example: video upscaling, video quality enhancement, vision applications such as object recognition, facial recognition, etc.
[0089] A simple description of the required steps is:
[0090] - Service provisioning and announcement of AI media service
[0091] - Service access information acquisition
[0092] -- Including possible request / subscription of AI model by UE or network (which task UE wants to perform, takes into account media requirements, network status parameters, UE status parameters, network or UE selects suitable AI model), building or ingesting an adapted model if not already available, and model selection
[0093] - Requesting the start of the AI data / media delivery
[0094] - Delivering the AI data (and possibly media data) for AI media inferencing:
[0095] -- Session(s) establishment(s)
[0096] -- Delivery of AI model from network to UE
[0097] -- Configure media session downlink
[0098] -- Stream media from network
[0099] -- AI media inference in UE
[0100] Figure 7 shows a scenario where the inferencing required for the AI media service is split between the network and UE. A portion of the AI model to be inferenced on the UE is delivered from network to the UE. Another portion of the AI model to be inferenced in the network is provisioned by the network to an entity which performs the inferencing in the network. The media for inferencing is firstly provisioned and ingested by the network to the network inferencing entity, when feeds the media as an input into the network portion of the AI model. The output of the network side inference (intermediate data) is then sent to the UE, which received this intermediate data and feeds it as an input into the UE side portion of the AI model, hence completing the inference of the whole model.
[0101] In this scenario, the split decision and configuration is negotiated between the UE and the network, and a simple description of the required steps is (see figures 10 and 11 for concrete details):
[0102] - Service provisioning and announcement of AI media service
[0103] - Service access information acquisition
[0104] -- Including possible request / subscription of AI model by UE or network (which task UE wants to perform, takes into account media requirements, network status parameters, UE status parameters, network or UE selects suitable AI model), building or ingesting an adapted model if not already available, and model selection
[0105] - Discovering cloud / edge AI media inferencing capabilities and functions
[0106] - Requesting AI split inferencing, either by the network or the UE, for an AI split inference service
[0107] - Discovering client AI media inferencing capabilities and functions
[0108] - Negotiating splitting the AI media inference process
[0109] - Starting the inference process in the server
[0110] - Acknowledging the split and providing the AI data split inferencing access information
[0111] - Acknowledging the split configuration
[0112] - Requesting the start of the AI data / media delivery
[0113] - Delivering the AI data (and possibly media data) for AI media inferencing:
[0114] -- Session(s) establishment(s)
[0115] -- Delivery of split AI model from network to UE
[0116] -- Configure media session downlink
[0117] -- Stream media from network
[0118] -- AI media inference in UE
[0119] In one split configuration example, an AI model service may consist of a core portion, as well as a task specific portion (e.g. traffic sign recognition task, or facial recognition task), where the core portion of the AI model is common to multiple possible tasks. In this case, the split configuration may coincide the core and task portions in a manner such that the network performs the inference of the core portion of the model, and the UE (receives and) performs the inference of the task portion of the model.
[0120] Figure 8 shows a basic architecture for a split AI inferencing scenario, where the media source originates from the network, and as such, where the split inferencing occurs first in the network, then subsequently in the UE. This architecture shows logical functions related to user plane data, in particular AI data.
[0121] AI data may include:
[0122] - AI model data (i.e. the data related to the structure of the AI model, including the number of layers, the weights and biases for the nodes and links between the layers etc)
[0123] - Intermediate data, which is the output of a first split inference, typically required to be delivered to a second device or entity, as the input to a subsequent second split inference. Intermediate data may have media characteristics.
[0124] - Inference output data, which is the output of an AI inference process. Depending on the nature of the AI media inferencing for the given Ai media service, this inference output data may include: labels for identifying recognition like tasks from media, actual media data such as video and / or audio, XR related data such as 3D models or any other possible inference output.
[0125] In figure 8, an AI model repository in the network provides corresponding network AI model subsets and UE AI model subsets to the network (to the network AI model inference engine) and UE (via the AI model delivery function, 5G system, UE AI model access function, to the UE AI model inference engine) respectively.
[0126] On network split inferencing by the network AI model inference engine, the output intermediate data is delivered to the UE as the input to the UE AI model inference engine (via the intermediate delivery and access functions). The final inference output data is then consumed within the UE at the data destination.
[0127] Figure 9 shows another basic architecture for a split AI inferencing scenario different to figure 8, where the media source originates in the UE device, and as such, where the split inferencing occurs first in the UE device, then subsequently in the UE. Typically after the final inference in the network, the inference output data is then also sent to the UE for consumption. This architecture shows logical functions related to user plane data, in particular AI data.
[0128] AI data may include:
[0129] - AI model data (i.e. the data related to the structure of the AI model, or AI model topology information, including the number of layers, the weights and biases for the nodes and links between the layers etc)
[0130] - Intermediate data, which is the output of a first split inference, typically required to be delivered to a second device or entity, as the input to a subsequent second split inference. Intermediate data may have media characteristics.
[0131] - Inference output data, which is the output of an AI inference process. Depending on the nature of the AI media inferencing for the given Ai media service, this inference output data may include: labels for identifying recognition like tasks from media, actual media data such as video and / or audio, XR related data such as 3D models or any other possible inference output.
[0132] The logical user plane functions in figure 9 have similar functionality to those described under figure 8.
[0133] Figure 10 shows and AI for media (AI4Media) architecture which identifies the various functional entities and interfaces for enabling AI model delivery for media services in this invention.
[0134] 5GAI AF:An Application Function similar to that defined in TS 23.501 clause 6.2.10, dedicated to AI media services. Typically provides various control functions to the AI Data Session Handler on the UE and / or to the 5GAI Application Provider. It may interact with other 5GC network functions, such as a Data Collection Proxy (DCP) function entity (which interacts with the AI / ML Endpoint and / or 3GPP Core Network to collect information required for the 5GAI AF). The DCP may or may not include NWDAF function / functionality. The 5GAI AF may contain logical subfunctions such as an AI Capability Manager, which handles the negotiation and handling of capability related data and decision in the network, and also between the network and UE.
[0135] 5GAI AS:An Application Server dedicated to AI media services, which hosts 5G AI media (sub)functions, such as the AI Data Delivery / access function and AI Inference Engine. The 5GAI AS typically supports AI model hosting by ingesting AI models from an AI Media Application Provider, and egesting models to other network functions for network inferencing, such as the Media AS. In addition to those described above, the 5GAI AS may also contain Media AS functionalities. The 5GAI AS may also contain an AI Inference Engine subfunction which performs full or partial inferencing on the network.
[0136] 5GAI Media Application Provider:External application, with content-specific media functionality, and / or AI-specific media functionality (AI model creation, splitting, updating etc.).
[0137] The 5GAI Client in the UE contains:
[0138] AI Data Session Handler:a function on the UE that communications with the 5GAI AF in order to establish, control and support the delivery of an AI model session, and / or a media session, and may perform additional functions such as consumption and QoE metrics collection and reporting. The AI Data Session Handler may expose APIs that can be used by the 5GAI Aware Application. It may contain logical subfunctions such as an AI Capability Manager, which handles the negotiation and handling of capability related data and decision internally in the UE, and also between the UE and network.
[0139] AI Data Handler:a function on the UE that communicates with the AI AS in order to download / stream (or even upload) the AI model data, and may provide APIs to the 5GAI Aware Application for AI model inferencing, and to the AI Data Session Handler for AI model session control in the UE, and also the subfunctions AI Data Access Function for accessing AI model data such as topology data and or AI model parameters (weights, biases), and AI Inference Engine for inferencing in the UE.
[0140] In another embodiment of this invention, the AI inference engine in the UE may exist outside the AI Data Handler. It may also exist in another function in the UE.
[0141] In another embodiment of this invention, the AI engine in the network may exist outside the 5GAI AS.
[0142] Figure 11 is an embodiment of this disclosure showing a procedure for the delivery of an AI model with configurations between the network and UE such that the AI model can be delivered in a manner described by figure 7.
[0143] Specifically:
[0144] 1. Service provisioning and announcement of AI media service, in particular between the 5GAI AF (application function) and the 5GAI application provider.
[0145] 2. Service access information acquisition
[0146] - Required AI model for the service is known in service access information (AI model known)
[0147] - During this step, the available or required AI model(s) for the service can be made known to the UE, by means of information made available via a URL link pointing to a file or manifest which may last such available AI models
[0148] - The received information may already contain AI model specific information, such as:the size of the AI model network, includingthe number of layers contained in the AI model structure, the number of nodes and links in each layer, the complexity of each layer in the AI model (i.e. the number of free parameters), the possible split points for the model for split inferencing, and also the AI model target inference delay.
[0149] - Additional steps here for model request / subscribe, building / ingesting adapted model if not available, model selection
[0150] 3. Discovering cloud / edge and client AI media inferencing capabilities and functions
[0151] 4. Requesting AI split inference
[0152] 5. Negotiate splitting the AI media inference process
[0153] - A split point may be decided during this stage, and the requirements for such a split point decision maybe that such as the total AI model target inference delay (or latency) for the service. In order to decide a split point, data received from steps 2, 3, 4 and 5 may be used for various calculations on deciding a split point, either in the UE or in the network, or both.
[0154] - Once a split point is decided, the configuration for the delivery of the split AI model may occur during this step, or, alternatively, it may occur when configuring the delivery pipelines for step 10. As embodiments of this disclosure, such configurations may include:
[0155] -- Configuration: static_split or dynamic_split;
[0156] Whether the split configuration is static during the service, or may be changed dynamically depending on factors in / during the service.
[0157] -- This step is further divided into additional embodiments as defined later in this invention
[0158] 6. Acknowledge split and provide the AI data split inferencing access info
[0159] 7. Acknowledge the split
[0160] 8. Request the start of AI data / media delivery
[0161] 9. The UE (5GAI client) requests the start of the AI data delivery from the network
[0162] 10~14. For steps 10~14, the configuration of the AI model and data delivery pipelines may include the same parameters, procedures and configurations as described in step 5.
[0163] 15. The UE reports its AI status to the network
[0164] 16. The AI status (in particular split inference related status) on the network side is also reported to the AF.
[0165] 17. The network related AI status report is sent to the UE.
[0166] 18. The media status is also aggregated by the AI data session handler.
[0167] 19. In step 19, an update of the split configuration (e.g. changing the split point for split inferencing) may occur. The control signalling of this split point re-configuration (or dynamic configuration) may utilize the metadata as described in step 5.
[0168] Steps 15 to 19 are also further described as additional embodiments in this invention.
[0169] Figure 12 shows the scenarios of split inferencing where certain data required to be sent between the UE and network (e.g. edge) may raise privacy concerns.
[0170] In scenario 1, the media source within the UE generates media (e.g. video or audio), and the media is then passed into the partial AI / ML model, as the 1stAI / ML inference on the UE. The output of this partial AI / ML model inference, also defined as intermediate data in this invention, is then sent via 5G to a network entity. At the network entity, which contains the second part of the partial AI / ML model, the intermediate data is fed into this second partial AI / ML model for the 2ndAI / ML inference in the network entity. The 1stand 2ndinferences using the first and second partial AI / ML models, at the UE and network entity respectively, completes the AI inference of the media data, outputting the result data. Depending on the AI media service, this result data may for example be media data, or other forms of output such as text labels, bounding boxes, or coordinates etc in the case of gesture recognition. The result data is then sent from the network back to the UE device, where it is consumed.
[0171] In scenario 2, the media data is sent via 5G to the network entity for the 1stinference using a first partial AI / ML model, where the output intermediate data is sent from the network entity back to the UE device. On receiving the intermediate data, it is used by the UE to perform the 2ndinference, using the second partial AI / ML model. The resulting data is then consumer on the same UE device.
[0172] In both scenarios, a complete AI / ML model is split into two different partial AI / ML models, which are separately inferenced in the UE and network entity (or vice versa).
[0173] Also for both scenarios, there is a concern of privacy, since the delivery of privacy sensitive data is required between the UE and network entity (or vice versa). In scenario 1, the result data from the final inference of the AI model may contain UE user specific information which is not suitable to entities other than the UE device. As seen in figure 12, in scenario 1, the result data not only susceptible to privacy issues when being delivered to the UE device, but it is made available to the network since it is created by the network entity.
[0174] In scenario 2, likewise the raw (or encoded) media data is directly sent to the network entity by the UE device. For many AI / ML media applications, the resulting output of the AI / ML inference is data which is not media data, meaning that much of the information contained in media data is obsolete, and is information which can be susceptible to privacy issues.
[0175] Figure 13 shows an embodiment of this invention as a scenario which can avoid the exposure of privacy sensitive information outside the UE device, by keeping private data local o the UE device.
[0176] Instead of splitting the AI model into two partial models as in the scenarios show in figure 12, the scenario in figure 13 splits the AI model into 3 partial models, such that a total of 3 partial inferences takes place, twice on the UE device, and once in the network entity.
[0177] Through such a scenario, only intermediate data is shared between the UE and network entity (not that the intermediate data output from the 1stinference is different data to the intermediate data output from the 2ndinference). In this manner, the original media data, and the final result data is also kept local to the UE. With this scenario embodiment, even if the intermediate data is somehow made know to a third party entity, without the knowledge of the AI model used, the split point at which the partial AI model was split at, and without the actual partial AI model data, it is difficult for any third party to obtain either the original media data, or the result data for the AI media service.
[0178] Figure 14 shows an AI model with multiple possible split points at the different layers of the AI model. Possible split points depend on the AI model structure or architecture, and not all layers may be possible split points. In the context of the scenarios in figure 12 and 13, the AI model in figure 14 maybe split into 2 partial models at any of the given split points (for figure 12), or it may be split into 3 partial models at any of the given split points (by selecting 2 split points, for figure 13).
[0179] Figure 15 is an embodiment of this invention, wherein step 5 of figure 11 is expanded into more detailed steps, notably steps 5 to 9 in figure 15.
[0180] After requesting the AI split inference in step 4:
[0181] Step 5
[0182] - The AI Data Session Handler and 5GAI AF negotiate the split service configuration. The split service configuration is dependent on the nature of the AI media service itself, for example, whether the media data originates in the UE or the network, whether the first partial inference occurs in the UE or the network, the direction of delivery of the intermediate data (uplink or downlink), and whether output data is required to be sent for monitoring or feedback purposes (e.g. sending output data to the network even though it is the UE device which consumes the output), or whether other feedback data is required to be delivered. As embodiments of this invention, the split service configurations typically correspond to the scenarios as described in figure 12.
[0183] Step 6
[0184] - Once the split service configuration is negotiated and agreed between the AI Data Session Handler and the 5GAI AF, the 5GAI Client and AI Data Session Handler may either check privacy related policies for the service or application, either predefined by service policies, or by user selection decisions.
[0185] Step 7
[0186] - Upon deciding a decision which requires privacy enablement, the AI Data Session Handler may request for a multi-split configuration, based on the initial negotiated split service configuration in step 5. As an embodiment of this invention, the multi-split configuration typically corresponds to the scenario as described in figure 13.
[0187] Step 8
[0188] - Upon receiving the request, the 5GAI AF either accepts the requested multi-split configuration, or renegotiates the configuration accordingly. It may also be possible to renegotiate a multi-split configuration which contains more than 2 split-points (e.g. more than 3 partial AI models), where partial inferences may be performed by different network entities, or other peer device entities.
[0189] Step 9
[0190] - On deciding and agreeing the multi-split configuration, the AI Data Session Handler and 5GAI AF then negotiate decide the exact split points to be used for the multi-split inference service. This corresponds to deciding the split points (out of the possible split points which typically exist at certain layers in the AI model as shown in figure 14) where the AI model is divided into 3 or more partial models.
[0191] - In an embodiment of this invention, the negotiation and decision of these split points can further be described in several additional procedures:
[0192] -- The exchange of parameters between the UE and network.Since the decision of the exact split points to be used is dependent on a number of different factors and parameters which may be exchanged between the UE and network, namely:
[0193] --- AI model parameters
[0194] ---- Model size
[0195] ---- Number of layers
[0196] ----Input size
[0197] ----Output size
[0198] ----Layer-wise intermediate data size ( or at split points specified )
[0199] ----Language ( OS )
[0200] ---Network performance parameters
[0201] ---- QoS
[0202] ---- Network bandwidth
[0203] ---Device performance parameters
[0204] ---- Split-point layer hardware performance ( processing time )
[0205] ---- Generic hardware performance
[0206] ---- Dynamic Load in System ( e.g. CPU usage )
[0207] -- The calibration and decision of the split points to be used. After the parameters listed above are exchanged between the UE and network, either the UE or the network may run certain algorithms or calibrations in order to calculate or estimate the optimal split points to be used for the service. Such algorithms or calibrations typically use the parameters as described above, and in one example the parameters may be used to minimalize the end-to-end latency of the split inference service, taking into account the inference processing latencies at the different entities performing partial inference, as well as taking into account the delivery latencies of sending the respective data to the other entity via 5G (e.g. intermediate data uplink or downlink from the UE).
[0208] -- Configuration of dynamic split decisions.Depending on the AI media service at hand, including the nature of the AI service, this step may also decide the dynamic configuration of the service. For example, many of the parameters listed, especially those related to network QoS and device performance, will change dynamically throughout the duration of the service. In addition, for certain AI media services, the AI model parameters may also change, e.g. a change in AI model topology / structure / architecture, or an update in the weights and biases used for the trained model. In these cases, the UE and / or network may decide to renegotiate the split points for the split AI inference service.
[0209] --- In order to enable and initiate the renegotiation of the split points, the sending of feedback messages may be configured, to be sent between the UE network, as highlighted by steps 13 to 16, and described in more detail below.
[0210] --- The method of sending feedback messages are further described below in the descriptions of steps 13 to 16.
[0211] --- As part of the dynamic split decision configuration, this step may also define thresholds for the service, wherein when certain thresholds are exceeded, feedback messages triggering an update of the split decision configuration are exchanged between the UE and network.
[0212] --- Thresholds may be direct limits of certain parameters as listed in this invention, or may be a parameter which is calculated from a combination of the parameters listed, e.g. a latency threshold for the service, when may be caused by various factors such as network conditions (delivery related bandwidth, QoS), or by device performance status conditions (processing burden / load status)
[0213] As another embodiment of this invention, depending on the configuration of dynamic split configurations in step 9, feedback messages maybe shared between the UE and network. The purposes of such feedback messages are twofold:
[0214] 1) To report the status of an entity such that the other entity can monitor the status via parameters sent via the report.
[0215] And / or:
[0216] 2) To trigger and initiate the update of the split configuration, or the split points used for the split configuration
[0217] Feedback messages may be sent either:
[0218] i) Periodically, as periodic feedback messages - by defining a periodic continuous time period, feedback messages are sent at the defined time interval, for the purposes of reporting specific parameters configured for reporting. Upon receiving these periodic feedback messages, an entity may compare parameters, or use the parameters to calculate and compare to any defined thresholds for initiating the renegotiation / update of the split configuration or split points.
[0219] ii) Triggered, as trigger feedback messages - an entity (e.g. UE or network) may be configured to continuously monitor the parameters relevant to the service (as listed in this invention), and, through a defined threshold, if the related parameters exceed the defined threshold(s), it sends a trigger feedback message to the other entity, initiating the renegotiation / update of the split configuration or split points. In addition, if an entity has knowledge of the update dynamic nature of the AI model (e.g. if AI topology / weight / biases change throughout the timeline of the service), a trigger feedback message can also be send using the AI model update as a trigger for the update of the split configuration, since the AI model data used is changed during these moments.
[0220] To summarize:
[0221] Utilize feedback messages to report and monitor status of parameters, and / or trigger negotiation / update of split configuration and / or split points
[0222] 1. Periodic feedback message
[0223] - Continuous and periodic time period
[0224] - E.g. use case where media is long and continuous (CCTV)
[0225] 2. Trigger feedback message
[0226] - Threshold triggers from parameters
[0227] -- Latencies (E2E, processing, delivery)
[0228] -- Bandwidth
[0229] -- Computational load
[0230] - Split-point re-negotiation from AI model update
[0231] -- Update of AI topology
[0232] -- Update of weights and biases
[0233] As detailed above:
[0234] Step 13
[0235] - If configured to send periodic feedback messages (in step 9), these messages are exchanged between the UE and network at the period time interval defined.
[0236] Step 14
[0237] - On receiving the period feedback messages which contain reported parameters, the entity receiving the feedback message may compare the reported parameters with the threshold(s) agreed in step 9. If the threshold(s) are exceeding, then it may trigger an update of the split configuration and / or split point(s) direct, or it may send a trigger feedback message.
[0238] Step 15
[0239] - As mentioned, trigger feedback messages may be sent if certain conditions exceed thresholds within an entity. As such, a trigger feedback may contain the related parameters relevant to the threshold(s), and / or it may also contain information to initiate the update of the split configuration and / or split point(s).
[0240] Step 16
[0241] - On receiving the prompt to update the split configuration / split point, the two entities renegotiate to update the split configuration and / or split points accordingly.
[0242] Figure 16 shows another embodiment of this invention, detailing the a similar (be different in details) procedure as in figure 15, wherein the entity on the right is a network entity, which may be an edge server, application server, or other media related network entity.
[0243] STEP-1 : UE initiates a discovery mechanism to find (heterogeneous) devices having different capabilities. From the set of heterogeneous devices, a device with higher capabilities is selected based on network bandwidth / signal strength, load, etc.UE initiates a communication with theselectedEdge for Split Computation / Processing
[0244] STEP-2 :Edge shares Calibration Model (Parent Patent ID: MN-202305-026-1-KR0 - On UE capability negotiation for device AIML inferencing and split AIML inferencing ) for capability negotiationfor split Decision for load balancing.
[0245] STEP-3 :UE receives the current Edge capability in terms of same Calibration model execution time as well as Network bandwidth details for consideration ( based on the model transmission time, computation of AI / ML model on UE andcomputation of AI / ML model the selected edge)
[0246] STEP-4 :Computation / Processing capability ratio ( e.g. UE : Edge ) is calculated / considered in UE sidefor dynamic split ratio based on network bandwidth in real time.
[0247] STEP-5 :UE computes the Split Compute logic for the model (Parent Patent ID: US20220311678A1 -https: / patents.google.com / patent / US20220311678A1). Additional consideration done based on constraints of Model and Data availability as well as whether un-processed data can be shared or not ( e.g. due to privacy consideration, UE data may not be possible to be shared to Edge and might enforce minimum the layer-1 to be executed in UE side itself ).
[0248] STEP-6 : Client Register the threshold capability variation with the selected edge.
[0249] STEP-7 :Possibility of multi split consideration ( e.g. for a "N" layered model , layer 1 to "m1" executes in UE / layer "m1" to "m2" in Edge and finally layer "m2" to "N" in UE again ). This is again needed in case of privacy consideration where final output cannot be made available in Edge.
[0250] STEP-8 :Once the initial split plan is finalized ,the inference layer partitioned is executed among UE and edge based on Split Compute logic for the AI / ML model where sub-set of layers are executedto Edge.
[0251] STEP-9:Once split details & edge is finalized, client creates a Tensor metadata description message and initiate the connection with the edge server to the provided url & port (details are received in Notify_Edge_AI_Modified_Caps()).
[0252] STEP-10 : Edge server responds with TMD ANSWER
[0253] STEP-11 :Initiate the Split Computation and Data Sending process. In case , the data is just a single input based inference , based on the split plan , intermediate data and Split Model layers will be shared to Edge and Edge will generate the output and share back to UE.
[0254] Figure 17 details the continuation of the embodiment as shown in figure 16.
[0255] STEP-12 :During In situation where the data has a continuous processing requirement ( e.g. CCTV footage analysis throughout the day in real-time ) , periodic feedback message gets exchanged (at continuous / periodic time stamp)to evaluate the edge and UE device load. The method is same as Step-3 , where the calibration model gets re-evaluated in a periodic manner and details are exchanged back to UE.
[0256] STEP-13:Possibility also can be there for indication based re-evaluation of split logic ( instead of periodic ) to avoid additional calculation. For example , on recognizing heavy load in UE / Edge ( for some continuous threshold duration ) , indication can be exchanged between UE and Edge. Similar consideration can also be done during transmission of data , when there is a reduced / increased bandwidth availability ( again more than threshold ) , which can require a split re-evaluationfor the AI / ML model.
[0257] STEP-14 :In case the periodic evaluation results in a Computation / Processing capability ratio different than the original / previous one ( delta above the "capability threshold" ) or the Network bandwidth has changed considerably ( delta above "network fluctuation threshold" ) , re-evaluation of Split computation is done and new split points get identified.
[0258] Step-15: Upon receiving the "Notify_Edge_AI_Modified_Caps" from each edge, the client needs to update its "edge table" table and re-order the edge's with the best capable edge at the top. Re -evaluation of split points needs to be identified using the edge that is present at the top in edge table.
[0259] Step-16: Once the new split point is identified, Step-8 & 9 needs to be performed with the new edge or the current edge dynamically
[0260] STEP-17 :Split Re-negotiation will require sending of notification to Edge from UE. This notification will stop new data from being sent to Edge and instead keep in queue for temporary storage. Once the Edge processes the presently available data and sends back the output to UE with confirmation of data drain from Edge.
[0261] STEP-18:Once Data drain information is received in UE , queued data starts getting resend along with new split model ( delta if needed ) based on the recent calculation done and data sending resumes between UE and Edge.
[0262] STEP-19:Once the entire data processing ends or the use-case requirement finishes , the final indication is sent from UE to Server and Server stops processing data and sends closure response.
[0263] STEP-20 : If the edge suddenly wants to STOP providing the service, it can De-board the service by sending Notify_Edge_AI_Modified_Caps() with AI service status as "inactive" message
[0264] STEP-21 : Upon receiving De-board message, the client should remove its entry from the edges table, and recalculate the split point using the next best edge entry available in the edge table.
[0265] Figure 18 shows an example of a signalling message for step-1 in figure 16(UE initiates a discovery mechanism to find (heterogeneous) devices having different capabilities).
[0266] Figure 19 shows an example of a signalling message for step-3 in figure 16 (UE receives the current Edge capability in terms of same Calibration model execution time as well as Network bandwidth details for consideration).
[0267] Figure 20 shows an example of a signalling message for step-6 in figure 16 (Client Register the threshold capability variation with the selected edge).
[0268] Figure 21 shows an example of a signalling message for step-8 in figure 16 (Once the initial split plan is finalized,the inference layer partitioned is executed among UE and edge based on Split Compute logic for the AI / ML model where sub-set of layers are executedto Edge).
[0269] Figure 22 shows an example of a signalling message for step-9 in figure 16 (Once split details & edge is finalized, client creates a Tensor metadata description message and initiate the connection with the edge server to the provided url & port).
[0270] The methods according to various embodiments described in the claims or the specification of the disclosure may be implemented by hardware, software, or a combination of hardware and software.
[0271] When the methods are implemented by software, a computer-readable storage medium for storing one or more programs (software modules) may be provided. The one or more programs stored in the computer-readable storage medium may be configured for execution by one or more processors within the electronic device. The at least one program may include instructions that cause the electronic device to perform the methods according to various embodiments of the disclosure as defined by the appended claims and / or disclosed herein.
[0272] The programs (software modules or software) may be stored in non-volatile memories including a random access memory and a flash memory, a read only memory (ROM), an electrically erasable programmable read only memory (EEPROM), a magnetic disc storage device, a compact disc-ROM (CD-ROM), digital versatile discs (DVDs), or other type optical storage devices, or a magnetic cassette. Alternatively, any combination of some or all of them may form a memory in which the program is stored. Further, a plurality of such memories may be included in the electronic device.
[0273] In addition, the programs may be stored in an attachable storage device which may access the electronic device through communication networks such as the Internet, Intranet, Local Area Network (LAN), Wide LAN (WLAN), and Storage Area Network (SAN) or a combination thereof. Such a storage device may access the electronic device via an external port. Further, a separate storage device on the communication network may access a portable electronic device.
[0274] In the above-described detailed embodiments of the disclosure, an element included in the disclosure is expressed in the singular or the plural according to presented detailed embodiments. However, the singular form or plural form is selected appropriately to the presented situation for the convenience of description, and the disclosure is not limited by elements expressed in the singular or the plural. Therefore, either an element expressed in the plural may also include a single element or an element expressed in the singular may also include multiple elements.
[0275] Although specific embodiments have been described in the detailed description of the disclosure, it will be apparent that various modifications and changes may be made thereto without departing from the scope of the disclosure. Therefore, the scope of the disclosure should not be defined as being limited to the embodiments, but should be defined by the appended claims and equivalents thereof.
Claims
1.A method performed by a user equipment (UE) in a wireless communication system, the method comprising:receiving, from a first network entity, information regarding at least one artificial intelligence (AI) model for a service;identifying AI data inferencing capabilities for the AI model;requesting an AI split inference;negotiating a configuration for the service and a split point of the AI split inference;establishing an AI model data deliver pipeline for delivering the AI model;establishing an intermediate data deliver pipeline for delivering intermediated data;performing the split inference process using the AI model data deliver pipeline and the intermediate data deliver pipeline;transmitting a status report for service; andupdating the configuration based on the status report.2.The method of claim 1,wherein the information regarding at least one AI model includes a uniform resource locator (URL) to obtain a list the at least one AI model.3.The method of claim 1, further comprising:transmitting a feedback message based on a predefined time interval or a predefined threshold; andupdating the configuration and the split point based on the feedback message.4.The method of claim 1,wherein the feedback message includes at least one of information on a network capability parameter, information on a latency and information on an AI model parameter.5.The method of claim 1,wherein the split point of the AI split inference is determined based on a privacy police.6.A method performed by a first network entity in a wireless communication system, the method comprising:transmitting, to a user equipment (UE), information regarding at least one artificial intelligence (AI) model for a service;identifying AI data inferencing capabilities for the AI model;requesting an AI split inference;negotiating a configuration for the service and a split point of the AI split inference;establishing an AI model data deliver pipeline for delivering the AI model;establishing an intermediate data deliver pipeline for delivering intermediated data;performing the split inference process using the AI model data deliver pipeline and the intermediate data deliver pipeline;receiving a status report for service; andupdating the configuration based on the status report.7.The method of claim 6,wherein the information regarding at least one AI model includes a uniform resource locator (URL) to obtain a list the at least one AI model.8.The method of claim 6, further comprising:receiving a feedback message based on a predefined time interval or a predefined threshold; andupdating the configuration and the split point based on the feedback message.9.The method of claim 6,wherein the feedback message includes at least one of information on a network capability parameter, information on a latency and information on an AI model parameter.10.The method of claim 6,wherein the split point of the AI split inference is determined based on a privacy police.11.A user equipment (UE) in a wireless communication system, the UE comprising:at least one transceiver; andat least one processor operatively coupled with the at least one transceiver,wherein the at least one processor is configured to:receive, from a first network entity, information regarding at least one artificial intelligence (AI) model for a service,identify AI data inferencing capabilities for the AI model,request an AI split inference,negotiate a configuration for the service and a split point of the AI split inference,establish an AI model data deliver pipeline for delivering the AI model,establish an intermediate data deliver pipeline for delivering intermediated data,perform the split inference process using the AI model data deliver pipeline and the intermediate data deliver pipeline,transmit a status report for service, andupdate the configuration based on the status report.12.The UE of claim 11,wherein the information regarding at least one AI model includes a uniform resource locator (URL) to obtain a list the at least one AI model.13.The UE of claim 11, wherein the at least one processor is further configured to:transmit a feedback message based on a predefined time interval or a predefined threshold, andupdate the configuration and the split point based on the feedback message.14.The UE of claim 11,wherein the feedback message includes at least one of information on a network capability parameter, information on a latency and information on an AI model parameter.15.A first network entity in a wireless communication system, the first network entity comprising:at least one transceiver; andat least one processor operatively coupled with the at least one transceiver,wherein the at least one processor is configured to,transmit, to a user equipment (UE), information regarding at least one artificial intelligence (AI) model for a service,identify AI data inferencing capabilities for the AI model,request an AI split inference,negotiate configuration for the service and a split point of the AI split inference,establish an AI model data deliver pipeline for delivering the AI model,establish an intermediate data deliver pipeline for delivering intermediated data,perform the split inference process using the AI model data deliver pipeline and the intermediate data deliver pipeline,receive a status report for service, andupdate the configuration based on the status report.