Methods, architectures, apparatus, and systems for artificial intelligence model delivery in wireless networks
The communication system addresses the lack of AI model distribution in 5G by using diverse radio access technologies to deliver AI models to WTRUs, improving their functionality within the network.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-09
- Publication Date
- 2026-04-08
Smart Images

Figure 2026510551000001_ABST
Abstract
Description
Technical Field
[0004] , , , , ,
[0001] This disclosure relates to procedures, methods, architectures, devices, systems, devices, and computer program products for and / or directed to the distribution of artificial intelligence (AI) and / or machine learning (ML) models from a network to wireless transmit-receive units (WTRUs) within the network.
Background Art
[0002] Cross - reference to Related Applications This application claims the benefit of (i) European Patent Application No. 23315026.7 filed on February 10, 2023, and (ii) European Patent Application No. 24305126.5 filed on January 22, 2024, each of which is incorporated herein by reference.
[0003] The 5G system lacks an architecture for providing AI models. It would be beneficial to propose an architecture and related methods for the distribution of AI models from the network to WTRUs.
Prior Art Documents
Non - Patent Documents
[0004]
Non - Patent Document 1
Non - Patent Document 2
Non - Patent Document 3
[0005] [Figure 1A] This is a system diagram illustrating an exemplary communication system. [Figure 1B] Figure 1A is a system diagram showing an exemplary wireless transceiver unit (WTRU) that may be used in the communication system shown. [Figure 1C] Figure 1A is a system diagram showing exemplary radio access networks (RANs) and exemplary core networks (CNs) that may be used within the communication system shown. [Figure 1D] Figure 1A is a system diagram showing further exemplary RAN and further exemplary CN that may be used within the communication system shown. [Figure 2] This is an AI / ML (Artificial Intelligence / Machine Learning) model diagram showing examples of different AI / ML subset compositions based on various split points. [Figure 3] Block diagram shows an exemplary functional model distributed architecture according to one embodiment. [Figure 4] This is a block diagram of a 5G AI model distributed system for downlink communication of AI models. [Figure 5] This is a block diagram showing the components of an inference engine in a 5G AI model distributed system according to one embodiment. [Figure 6] This is a block diagram showing the components of an AI model session handler in a 5G AI model distributed system according to one embodiment. [Figure 7] This is a high-level signal flow diagram illustrating an example of progressive download for on-demand AI model content according to an embodiment. [Figure 8A] This is a signaling flow diagram illustrating the (e.g., high-level) procedure for progressive downloading an AI model in subset streaming / incremental loading mode according to an embodiment. [Figure 8B] This is a signaling flow diagram illustrating the (e.g., high-level) procedure for progressive downloading an AI model in subset streaming / incremental loading mode according to an embodiment. [Figure 9] This is a block diagram showing the components for processing a neural network (NN) into a neural network representation (NNR). [Figure 10] This is a block diagram of the components of an NNR bitstream. [Figure 11] A block diagram illustrating an exemplary overview of a protocol stack. [Figure 12] This is a system diagram illustrating an exemplary system that uses Dynamic Adaptive Streaming Over Hypertext Transfer Protocol (DASH) segmentation over a hypertext transfer protocol for AI model delivery. [Figure 13] This block diagram shows an example data structure for File Delivery Over Unidirectional transport (FLUTE). [Figure 14]This block diagram shows an exemplary protocol stack for Real-Time Object Delivery Over Unidirectional Transport (ROUTE) DASH. [Figure 15] This is a step diagram illustrating the exemplary procedure for WTRU102 to download an AI model from the network. [Modes for carrying out the invention]
[0006] A more detailed understanding can be obtained from the following detailed description, which is given as examples along with the drawings attached herein. The figures in the drawings, as well as the detailed description, are examples. Therefore, the figures (FIG.) and the detailed description should not be considered limiting, and other equally valid examples are possible and possible. Furthermore, similar reference numbers ("ref.") in the figures indicate similar elements.
[0007] The following detailed description includes numerous specific details to provide a thorough understanding of the embodiments and / or examples disclosed herein. However, it should be understood that such embodiments and examples may be carried out without some or all of the specific details described herein. In other examples, well-known methods, procedures, components, and circuits are not described in detail so as not to obscure the following description. Furthermore, embodiments and examples not specifically described herein may be carried out in place of, or in combination with, the embodiments and other examples explicitly, implicitly, and / or essentially described, disclosed, or otherwise provided herein (collectively, “provided”). Various embodiments of apparatus, systems, devices, etc., and / or any elements thereof that perform operations, processes, algorithms, functions, etc., and / or any part thereof are described and / or claimed herein, but it should be understood that any embodiment described and / or claimed herein assumes that any apparatus, systems, devices, etc., and / or any elements thereof are configured to perform any operation, process, algorithm, function, etc., and / or any part thereof.
[0008] Exemplary communication system The methods, apparatus, and systems provided herein are well suited to communications, including both wired and wireless networks. Outlines of various types of wireless devices and infrastructure are provided with respect to Figures 1A to 1D, and various elements of a network may utilize, operate, and be arranged in accordance with the methods, apparatus, and systems provided herein, as well as adapt and / or be configured for them.
[0009] Figure 1A is a system diagram showing an exemplary communication system 100 in which one or more disclosed embodiments may be implemented. The communication system 100 may be a multiple access system that provides content such as voice, data, video, messaging, broadcast, etc., to multiple wireless users. The communication system 100 may enable multiple wireless users to access such content through the sharing of system resources, including wireless bandwidth. For example, the communication system 100 may use one or more channel access methods such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single-carrier FDMA (SC-FDMA), zero-tail (ZT) unique word (UW) discrete Fourier transform (DFT) spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block filtered OFDM, filter bank multicarrier (FBMC), etc.
[0010] As shown in FIG. 1A, the communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, a radio access network (RAN) 104 / 113, a core network (CN) 106 / 115, a public switched telephone network (PSTN) 108, the Internet 110, and other networks 112, although it is understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a wireless environment. By way of example, the WTRUs 102a, 102b, 102c, 102d may each be referred to as a “station” and / or “STA” and may be configured to transmit and / or receive wireless signals, and may include, or be, a user equipment (UE), a mobile station, a fixed or mobile subscriber unit, a subscriber-based unit, a pager, a cellular telephone, a personal digital assistant (PDA), a smartphone, a laptop, a netbook, a personal computer, a wireless sensor, a hotspot or Mi-Fi device, an Internet of Things (IoT) device, a wristwatch or other wearable, a head-mounted display (HMD), a vehicle, a drone, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in the context of industrial and / or automated processing chains), home appliances, devices operating on commercial and / or industrial wireless networks, etc. Any of the WTRUs 102a, 102b, 102c, and 102d may be interchangeably referred to as a UE.
[0011] The communication system 100 may also include base stations 114a and / or base station 114b. Each of the base stations 114a and 114b may be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, and 102d, thereby facilitating access to one or more communication networks, such as CN 106 / 115, the Internet 110, and / or network 112. As an example, base stations 114a and 114b may be any of the following: Base Transceiver Station (BTS), Node-B (NB), eNode-B (eNB), Home Node-B (HNB), Home eNode-B (HeNB), gNode-B (gNB), NR Node-B (NR NB), site controller, access point (AP), wireless router, etc. Although base stations 114a and 114b are shown as single elements, it will be understood that base stations 114a and 114b may include any number of interconnected base stations and / or network elements.
[0012] Base station 114a may be part of RAN104 / 113, and RAN104 / 113 may also include other base stations and / or network elements (not shown) such as a base station controller (BSC), a radio network controller (RNC), a relay node, etc. Base station 114a and / or base station 114b may be configured to transmit and / or receive wireless signals on one or more carrier frequencies that may be referred to as a cell (not shown). These frequencies may be an authorized frequency band, an unauthorized frequency band, or a combination of an authorized frequency band and an unauthorized frequency band. A cell may provide coverage for wireless services to a specific geographic area that may be relatively fixed or may change over time. A cell may be further divided into cell sectors. For example, the cell associated with base station 114a may be divided into three sectors. Thus, in one embodiment, base station 114a may include three (i.e., one for each sector of the cell) transceivers. In one embodiment, base station 114a may use multiple-input multiple-output (MIMO) technology and may utilize multiple transceivers for each sector or any sector of the cell. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.
[0013] Base stations 114a, 114b can communicate with one or more of WTRUs 102a, 102b, 102c, 102d via air interface 116, and air interface 116 can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, millimeter wave, infrared (IR), ultraviolet (UV), visible light, etc.). Air interface 116 can be established using any suitable radio access technology (RAT).
[0014] More specifically, as described above, the communication system 100 may be a multiple access system and may use one or more channel access schemes such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, base stations 114a and WTRUs 102a, 102b, 102c in RAN 104 / 113 may implement radio technologies such as Universal Mobile Communications System (UMTS) and Terrestrial Radio Access (UTRA), which may establish an air interface 116 using broadband CDMA (WCDMA). WCDMA may include communication protocols such as High Speed Packet Access (HSPA) and / or Advanced HSPA (HSPA+). HSPA may include High Speed Downlink Packet Access (HSDPA) and / or High Speed Uplink Packet Access (HSUPA).
[0015] In one embodiment, base stations 114a and WTRUs 102a, 102b, and 102c can implement radio technologies such as Advanced UMTS Terrestrial Radio Access (E-UTRA), which can establish an air interface 116 using Long-Term Evolution (LTE), as well as / or LTE Advanced (LTE-A) and / or LTE Advanced Pro (LTE-A Pro).
[0016] In one embodiment, base station 114a and WTRU 102a, 102b, 102c can implement radio technologies such as NR radio access, which can establish an air interface 116 using New Radio (NR).
[0017] In one embodiment, base station 114a and WTRUs 102a, 102b, 102c can implement multiple radio access technologies. For example, base station 114a and WTRUs 102a, 102b, 102c can implement LTE radio access and NR radio access together, for example, using the dual connectivity (DC) principle. Thus, the air interface utilized by WTRUs 102a, 102b, 102c may be characterized by multiple types of radio access technologies and / or transmissions sent to / from multiple types of base stations (e.g., eNB and gNB).
[0018] In one embodiment, base stations 114a and WTRUs 102a, 102b, and 102c can implement wireless technologies such as IEEE 802.11 (i.e., Wireless Fidelity (Wi-Fi)), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Provisional Standard 2000 (IS-2000), Provisional Standard 95 (IS-95), Provisional Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rate for GSM Evolution (EDGE), GSM EDGE (GERAN), etc.
[0019] In Figure 1A, base station 114b may be, for example, a wireless router, Home Node-B, Home eNode-B, or access point, and may utilize any suitable RAT to facilitate wireless connectivity in local areas such as workplaces, homes, vehicles, campuses, industrial facilities, aerial corridors (for use by drones, for example), roads, etc. In one embodiment, base station 114b and WTRU 102c, 102d can implement radio technologies such as IEEE 802.11 to establish a wireless local area network (WLAN). In one embodiment, base station 114b and WTRU 102c, 102d can implement radio technologies such as IEEE 802.15 to establish a wireless personal area network (WPAN). In one embodiment, base stations 114b and WTRUs 102c, 102d can establish any small cell, picocell, or femtocell using a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.). As shown in Figure 1A, base station 114b may have a direct connection to the internet 110. Therefore, base station 114b may not need to access the internet 110 via CN 106 / 115.
[0020] RAN104 / 113 can communicate with CN106 / 115, which may be any type of network configured to provide voice, data, applications, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRU102a, 102b, 102c, and 102d. The data may have various Quality of Service (QoS) requirements, such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, and mobility requirements. CN106 / 115 may provide call control, billing services, mobile location-based services, prepaid calls, internet connectivity, video distribution, and / or perform high-level security functions such as user authentication. Although not shown in Figure 1A, it will be understood that RAN104 / 113 and / or CN106 / 115 may communicate directly or indirectly with other RANs using the same RAT as RAN104 / 113, or with other RATs, either different or otherwise. For example, in addition to connecting to RAN104 / 113 which may be using NR radio technology, CN106 / 115 can also communicate with another RAN (not shown) using one of the following technologies: GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or Wi-Fi radio technology.
[0021] CN106 / 115 can also act as a gateway for WTRU102a, 102b, 102c, and 102d to access PSTN108, the Internet 110, and / or other networks 112. PSTN108 may include a circuit-switched telephone network providing Plain Old Telephone Service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices using common communication protocols such as Transmit Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) in the TCP / IP Internet Protocol Suite. Network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, network 112 may include another CN connected to one or more RANs that can use the same RAT as RAN104 / 114 or a different RAT.
[0022] Some or all of the WTRUs 102a, 102b, 102c, and 102d in the communication system 100 can include multimode capability (for example, WTRUs 102a, 102b, 102c, and 102d can include multiple transceivers for communicating with different wireless networks via different wireless links). For example, WTRU 102c shown in Figure 1A may be configured to communicate with base station 114a, which can use cellular-based radio technology, and base station 114b, which can use IEEE 802 radio technology.
[0023] Figure 1B is a system diagram showing an exemplary WTRU 102. As shown in Figure 1B, the WTRU 102 may include, among other things, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power supply 134, a global positioning system (GPS) chipset 136, and / or other elements / peripherals 138. It will be understood that the WTRU 102 may include any subcombinations of the aforementioned elements while maintaining consistency with the embodiment.
[0024] The processor 118 could be a general-purpose processor, a dedicated processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. The processor 118 can perform signal coding, data processing, power control, input / output processing, and / or any other functionality that enables WTRU 102 to operate in a wireless environment. The processor 118 may be coupled to the transceiver 120, and the transceiver 120 may be coupled to the transmit / receive element 122. Although Figure 1B shows the processor 118 and the transceiver 120 as separate components, it will be understood that the processor 118 and the transceiver 120 may be integrated together, for example, in an electronic package or chip.
[0025] The transmitting / receiving element 122 may be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) via the air interface 116. For example, in one embodiment, the transmitting / receiving element 122 may be an antenna configured to transmit and / or receive RF signals. In one embodiment, the transmitting / receiving element 122 may be an emitter / detector configured to transmit and / or receive, for example, IR signals, UV signals, or visible light signals. In one embodiment, the transmitting / receiving element 122 may be configured to transmit and / or receive both RF signals and optical signals. It will be understood that the transmitting / receiving element 122 may be configured to transmit and / or receive any combination of wireless signals.
[0026] Although the transmit / receive element 122 is shown as a single element in Figure 1B, the WTRU 102 can include any number of transmit / receive elements 122. For example, the WTRU 102 can utilize MIMO technology. Thus, in one embodiment, the WTRU 102 can include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals via the air interface 116.
[0027] The transceiver 120 may be configured to modulate the signal transmitted by the transmit / receive element 122 and demodulate the signal received by the transmit / receive element 122. As described above, the WTRU 102 may have multimode capability. Therefore, the transceiver 120 may include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs such as NR and IEEE 802.11.
[0028] The processor 118 of the WTRU102 may be coupled to a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light-emitting diode (OLED) display unit), and may receive user input data from there. The processor 118 may also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. In addition, the processor 118 may access information from any type of suitable memory, such as non-removable memory 130 and / or removable memory 132, and store data therein. Non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. Removable memory 132 may include a subscriber identification module (SIM) card, a memory stick, a secure digital (SD) memory card, and the like. In other embodiments, the processor 118 can access information from memory that is not physically located on the WTRU 102, such as on a server or home computer (not shown), and store data in that memory.
[0029] The processor 118 can receive power from the power supply 134 and may be configured to distribute and / or control power to other components within the WTRU 102. The power supply 134 can be any suitable device for supplying power to the WTRU 102. For example, the power supply 134 may include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel-metal hydride (NiMH), lithium-ion (Li-ion), etc.), a solar cell, a fuel cell, etc.
[0030] The processor 118 may also be coupled to a GPS chipset 136 which can be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to, or instead of, the WTRU 102 may receive location information from base stations (e.g., base stations 114a, 114b) via the air interface 116 and / or determine its position based on the timing of signals received from two or more nearby base stations. It will be understood that the WTRU 102 may acquire location information by any suitable positioning method while maintaining consistency with the embodiments.
[0031] The processor 118 may be further coupled to other elements / peripherals 138, which may include one or more software and / or hardware modules / units that provide additional features, functionality, and / or wired or wireless connectivity. For example, elements / peripherals 138 may include an accelerometer, an electronic compass, a satellite transceiver, a digital camera (e.g., for photos and / or videos), a Universal Serial Bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth® module, a frequency modulation (FM) radio unit, a digital music player, a media player, a video game player module, an internet browser, a virtual reality and / or augmented reality (VR / AR) device, an activity tracker, and the like. Element / Peripheral 138 may include one or more sensors, which may be one or more of the following: gyroscope, accelerometer, Hall effect sensor, magnetometer, compass sensor, proximity sensor, temperature sensor, time sensor, geolocation sensor, altimeter, light sensor, touch sensor, barometer, gesture sensor, biometric sensor, and / or humidity sensor.
[0032] WTRU102 may include a full-duplex radio in which the transmission and reception of some or all of a signal associated with a particular subframe for both an uplink (for transmission) and a downlink (for reception) may be parallel and / or simultaneous. The full-duplex radio may include an interference management unit to reduce and / or substantially eliminate self-interference through signal processing either through hardware (e.g., chokes) or a processor (e.g., a separate processor (not shown) or processor 118). In one embodiment, WTRU102 may include a half-duplex radio for the transmission and reception of some or all of a signal associated with either an uplink (for transmission) or a downlink (for reception).
[0033] Figure 1C is a system diagram showing RAN104 and CN106 according to one embodiment. As described above, RAN104 can communicate with WTRU102a, 102b, and 102c via the air interface 116 using E-UTRA wireless technology. RAN104 can also communicate with CN106.
[0034] RAN104 may include eNode-B160a, 160b, and 160c, but it will be understood that RAN104 may include any number of eNode-B while maintaining consistency with the embodiment. Each eNode-B160a, 160b, and 160c may include one or more transceivers for communicating with WTRU102a, 102b, and 102c via the air interface 116. In one embodiment, eNode-B160a, 160b, and 160c can implement MIMO technology. Thus, eNode-B160a can, for example, use multiple antennas to transmit wireless signals to and receive wireless signals from WTRU102a.
[0035] Each of the eNode-B160a, 160b, and 160c may be associated with a specific cell (not shown) and may be configured to handle wireless resource management decisions, handover decisions, user scheduling on uplink (UL) and / or downlink (DL), etc. As shown in Figure 1C, the eNode-B160a, 160b, and 160c can communicate with each other via the X2 interface.
[0036] The CN106 shown in Figure 1C may include a Mobility Management Entity (MME) 162, a Serving Gateway (SGW) 164, and a Packet Data Network (PDN) Gateway (PGW) 166. Although each of the aforementioned elements is shown as part of CN106, it will be understood that any one of these elements may be owned and / or operated by an entity other than the CN operator.
[0037] The MME162 can be connected to each of the eNode-B160a, 160b, and 160c within RAN104 via the S1 interface and can function as a control node. For example, the MME162 can be responsible for authenticating users of WTRU102a, 102b, and 102c, activating / deactivating bearers, and selecting a specific serving gateway during the initial attachment of WTRU102a, 102b, and 102c. The MME162 can provide control plane functionality for switching between RAN104 and other RANs (not shown) using other radio technologies such as GSM and / or WCDMA.
[0038] The SGW164 can be connected to each of the eNode-B160a, 160b, and 160c within RAN104 via the S1 interface. The SGW164 can generally route and forward user data packets to and from WTRU102a, 102b, and 102c. The SGW164 can perform other functions such as anchoring the user plane during eNode-B handovers, triggering paging when DL data is available to WTRU102a, 102b, and 102c, and managing and remembering the status of WTRU102a, 102b, and 102c.
[0039] SGW164 may be connected to PGW166, which provides WTRU102a, 102b, and 102c with access to a packet-switched network such as the Internet 110, thereby facilitating communication between WTRU102a, 102b, and 102c and IP-enabled devices.
[0040] CN106 can facilitate communication with other networks. For example, CN106 can provide WTRU102a, 102b, and 102c with access to circuit-switched networks such as PSTN108, thereby facilitating communication between WTRU102a, 102b, and 102c and conventional land-line communication devices. For example, CN106 may include, or communicate with, an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between CN106 and PSTN108. In addition, CN106 can provide WTRU102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.
[0041] While the WTRU is shown as a wireless terminal in Figures 1A to 1D, in some typical embodiments, such a terminal is expected to be able to use a wired communication interface with a communication network (e.g., temporarily or permanently).
[0042] In a typical embodiment, the other network 112 may be a WLAN.
[0043] In Infrastructure Basic Service Set (BSS) mode, a WLAN may have access points (APs) for the BSS and one or more stations (STAs) associated with the APs. APs may have access to or interfaces with distributed systems (DSs) or other types of wired / wireless networks that carry traffic into and / or out of the BSS. Traffic originating outside the BSS and destined for the STAs may reach and be delivered to the STAs via the APs. Traffic originating from the STAs and destined for destinations outside the BSS may be sent to the APs for delivery to their respective destinations. Traffic between STAs within the BSS may be sent via APs; for example, a source STA can send traffic to an AP, which can then deliver the traffic to the destination STA. Traffic between STAs within the BSS may be considered and / or referred to as peer-to-peer traffic. Peer-to-peer traffic may be sent between a source STA and a destination STA (e.g., directly between them) using a Direct Link Setup (DLS). In some representative embodiments, the DLS may use 802.11e DLS or 802.11z Tunnel DLS (TDLS). A WLAN using Independent BSS (IBSS) mode may not have APs, and STAs within or using IBSS (e.g., all STAs) can communicate directly with each other. The IBSS communication mode may be referred to herein as the “ad hoc” communication mode.
[0044] When using the 802.11ac infrastructure operating mode or a similar operating mode, an AP may transmit beacons on a fixed channel, such as the primary channel. The primary channel may have a fixed width (e.g., a 20 MHz bandwidth) or a dynamically set width via signaling. The primary channel may be the operating channel of the BSS, which can be used by STAs to establish connections with the AP. In some typical embodiments, carrier-sensing multiple access and collision avoidance (CSMA / CA) may be implemented in an 802.11 system, for example. For CSMA / CA, STAs, including the AP (e.g., all STAs), can sense the primary channel. If a particular STA senses / detects and / or determines that the primary channel is busy, that STA may backoff. A single STA (e.g., only one station) may transmit at any given time within a given BSS.
[0045] A high-throughput (HT) STA can, for example, form a 40MHz wide channel for communication by using a 40MHz wide channel through a combination of a primary 20MHz channel and adjacent or non-adjacent 20MHz channels.
[0046] Ultra-high throughput (VHT) STAs may support 20MHz, 40MHz, 80MHz, and / or 160MHz wide channels. 40MHz and / or 80MHz channels can be formed by combining consecutive 20MHz channels. A 160MHz channel can be formed by combining eight consecutive 20MHz channels, or by combining two discontinuous 80MHz channels, which may be called an 80+80 configuration. For the 80+80 configuration, after channel encoding, the data may be passed through a segment parser that can split the data into two streams. Inverse fast Fourier transform (IFFT) processing and time-domain processing may be performed separately for each stream. The streams may be mapped onto two 80MHz channels, and the data may be transmitted by a transmitting STA. At the receiver of the receiving STA, the operation described above for the 80+80 configuration may be reversed, and the combined data may be sent to a medium access control (MAC) layer, entities, etc.
[0047] Sub-1GHz operating modes are supported by 802.11af and 802.11ah. Channel operating bandwidth and carrier are reduced in 802.11af and 802.11ah compared to those used in 802.11n and 802.11ac. 802.11af supports 5MHz, 10MHz, and 20MHz bandwidths in the TV white space (TVWS) frequency band, while 802.11ah supports 1MHz, 2MHz, 4MHz, 8MHz, and 16MHz bandwidths using non-TVWS frequency bands. According to a typical embodiment, 802.11ah may support meter-type control / machine-type communications (MTC), such as MTC devices in a macro coverage area. MTC devices may have limited capabilities, including support for some and / or limited bandwidths (e.g., support only for that). MTC devices may include batteries with battery life above a threshold (e.g., to maintain very long battery life).
[0048] WLAN systems that can support multiple channels and channel bandwidths, such as 802.11n, 802.11ac, 802.11af, and 802.11ah, include a channel that can be designated as the primary channel. The primary channel may have a bandwidth equal to the largest common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel may be set and / or limited by the STA that supports the mode operating at the smallest bandwidth among all STAs operating in the BSS. In the 802.11ah example, the primary channel may be 1 MHz wide for an STA (e.g., an MTC type device) that supports (e.g., only supports) the 1 MHz mode, even if other STAs in the AP and BSS support modes operating at 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidths. Carrier sensing and / or network allocation vector (NAV) settings may depend on the status of the primary channel. For example, if the primary channel is busy due to an STA (which only supports a mode operating at 1MHz) transmitting to the AP, the entire available frequency band may be considered busy even though a large portion of the frequency band remains idle and could be available.
[0049] In the United States, the available frequency band that can be used by 802.11ah is from 902 MHz to 928 MHz. In South Korea, the available frequency band is from 917.5 MHz to 923.5 MHz. In Japan, the available frequency band is from 916.5 MHz to 927.5 MHz. The total bandwidth available for 802.11ah is from 6 MHz to 26 MHz, depending on the country code.
[0050] Figure 1D is a system diagram showing RAN113 and CN115 according to one embodiment. As described above, RAN113 can communicate with WTRU102a, 102b, and 102c via the air interface 116 using NR radio technology. RAN113 can also communicate with CN115.
[0051] RAN113 may include gNB180a, 180b, and 180c, but it will be understood that RAN113 may include any number of gNBs while maintaining consistency with the embodiment. gNB180a, 180b, and 180c may each include one or more transceivers for communicating with WTRU102a, 102b, and 102c via the air interface 116. In one embodiment, gNB180a, 180b, and 180c may implement MIMO technology. For example, gNB180a and 180b can use beamforming to transmit signals to and / or receive signals from WTRU102a, 102b, and 102c. Thus, gNB180a can, for example, use multiple antennas to transmit wireless signals to and / or receive wireless signals from WTRU102a. In one embodiment, gNB180a, 180b, and 180c may implement carrier aggregation technology. For example, gNB180a may transmit multiple component carriers to WTRU102a (not shown). A subset of these component carriers may be on unlicensed frequency bands, while the remaining component carriers may be on licensed frequency bands. In one embodiment, gNB180a, 180b, and 180c may implement coordinated multipoint (CoMP) technology. For example, WTRU102a may receive coordinated transmissions from gNB180a and gNB180b (and / or gNB180c).
[0052] WTRU102a, 102b, and 102c can communicate with gNB180a, 180b, and 180c using transmits associated with scalable neurology. For example, OFDM symbol intervals and / or OFDM subcarrier intervals may vary for different transmits, different cells, and / or different portions of the wireless transmit frequency band. WTRU102a, 102b, and 102c can communicate with gNB180a, 180b, and 180c using subframes or transmit time intervals (TTIs) of varying or scalable lengths (e.g., including a changing number of OFDM symbols and / or a varying absolute time duration).
[0053] The gNB180a, 180b, and 180c can be configured to communicate with WTRU102a, 102b, and 102c in standalone and / or non-standalone configurations. In a standalone configuration, the WTRU102a, 102b, and 102c can communicate with the gNB180a, 180b, and 180c without accessing other RANs (e.g., eNode-B160a, 160b, and 160c). In a standalone configuration, the WTRU102a, 102b, and 102c can use one or more of the gNB180a, 180b, and 180c as mobility anchor points. In a standalone configuration, the WTRU102a, 102b, and 102c can communicate with the gNB180a, 180b, and 180c using signals in an unauthorized band. In a non-standalone configuration, WTRU102a, 102b, and 102c can communicate / connect with gNB180a, 180b, and 180c, while also communicating / connecting with other RANs such as eNode-B160a, 160b, and 160c. For example, WTRU102a, 102b, and 102c can implement DC principles for substantially simultaneous communication with one or more gNB180a, 180b, and 180c and one or more eNode-B160a, 160b, and 160c. In a non-standalone configuration, eNode-B160a, 160b, and 160c can act as mobility anchors for WTRU102a, 102b, and 102c, and gNB180a, 180b, and 180c can provide additional coverage and / or throughput to service WTRU102a, 102b, and 102c.
[0054] Each of the gNB180a, 180b, and 180c may be associated with a specific cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, support for network slicing, dual connectivity, interworking between NR and E-UTRA, routing of user plane data to user plane functions (UPF) 184a and 184b, and routing of control plane information to access and mobility management functions (AMF) 182a and 182b. As shown in Figure 1D, the gNB180a, 180b, and 180c may communicate with each other via the Xn interface.
[0055] The CN115 shown in Figure 1D may include at least one AMF182a, 182b, at least one UPF184a, 184b, at least one Session Management Function (SMF)183a, 183b, and at least one Data Network (DN)185a, 185b. Although each of the aforementioned elements is shown as part of the CN115, it will be understood that any of these elements may be owned and / or operated by entities other than the CN operator.
[0056] AMF182a and 182b can be connected to one or more of gNB180a, 180b, and 180c within RAN113 via the N2 interface and can function as control nodes. For example, AMF182a and 182b can perform roles such as authenticating users of WTRU102a, 102b, and 102c, supporting network slicing (e.g., handling different protocol data unit (PDU) sessions with different requirements), selecting specific SMF183a and 183b, managing registration areas, terminating NAS signaling, and mobility management. Network slicing can be used by AMF182a and 182b to customize CN support for WTRU102a, 102b, and 102c based on the type of service being used by WTRU102a, 102b, and 102c. For example, different network slices may be established for different use cases, such as services relying on ultra-reliable low latency (URLLC) access, services relying on enhanced massive mobile broadband (eMBB) access, services for MTC access, and / or similar. The AMF162 can provide control plane functionality for switching between RAN113 and other RANs (not shown), where other radio technologies, such as LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies like Wi-Fi, may be used.
[0057] SMF183a and 183b can be connected to AMF182a and 182b in CN115 via the N11 interface. SMF183a and 183b can also be connected to UPF184a and 184b in CN115 via the N4 interface. SMF183a and 183b can select and control UPF184a and 184b and configure the routing of traffic through them. SMF183a and 183b can perform other functions, such as managing and assigning UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, and providing downlink data notifications. PDU session types can be IP-based, non-IP-based, Ethernet-based, etc.
[0058] UPF184a and 184b may be connected via the N3 interface to one or more of gNB180a, 180b, and 180c in RAN113, which can provide WTRU102a, 102b, and 102c with access to a packet-switched network such as the Internet 110 to facilitate communication between WTRU102a, 102b, and 102c and IP-enabled devices. UPF184 and 184b can perform other functions such as routing and forwarding packets, enforcing user plane policies, supporting multi-homed PDU sessions, handling user plane QoS, buffering downlink packets, and providing mobility anchoring.
[0059] CN115 can facilitate communication with other networks. For example, CN115 may include, or can communicate with, an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between CN115 and PSTN108. In addition, CN115 can provide WTRU102a,102b,102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, WTRU102a,102b,102c may be connected to data networks (DN) 185a,185b through UPF184a,184b via an N3 interface to UPF184a,184b and an N6 interface between UPF184a,184b and DN185a,185b.
[0060] In terms of Figures 1A to 1D and the corresponding descriptions therein, one or more or all of the functions described herein with respect to WTRU 102a to d, base stations 114a to b, eNode-B 160a to c, MME 162, SGW 164, PGW 166, gNB 180a to c, AMF 182a to b, UPF 184a to b, SMF 183a to b, DN 185a to b, and / or any other elements / devices described herein may be performed by one or more emulation elements / devices (not shown). An emulation device may be one or more devices configured to emulate one or more or all of the functions described herein. For example, an emulation device may be used to test other devices and / or to simulate network functions and / or WTRU functions.
[0061] Emulation devices may be designed to implement one or more tests of other devices in a laboratory and / or operator network environment. For example, one or more emulation devices may perform one, more, or all functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices in a communication network. One or more emulation devices may perform one, more, or all functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. Emulation devices may be directly coupled to another device for test purposes and / or perform tests using over-the-air wireless communication.
[0062] One or more emulation devices can perform one or more functions, including all of them, while not implemented / deployed as part of a wired and / or wireless communication network. For example, an emulation device may be used in a test scenario in a test laboratory and / or an undeployed (e.g., test) wired and / or wireless communication network to implement testing of one or more components. One or more emulation devices may be test equipment. Wireless communication via direct RF coupling and / or RF circuitry (which may include, for example, one or more antennas) may be used by an emulation device to transmit and / or receive data.
[0063] Introduction Figure 2 is an AI / ML (Artificial Intelligence / Machine Learning) model diagram showing examples of different AI / ML subset compositions based on various split points. As shown in Figure 2, several compositions of the same AI / ML model M can be represented by AI / ML subsets (M0, M1), (M'0, M'1), or (M''0, M''1, M''2). The same AI / ML subset can be used in different compositions depending on the configuration of the model composition (e.g., M'0 and M''0).
[0064] In some representative embodiments, a set of subsets where some subsets within a set run on different network nodes (e.g., different WTRUs) may be referred to as model splitting and / or distributed inference. As can be seen in various compositions 202, an AI model may have multiple different candidate split points. For example, the lines separating different subsets in Figure 2 may represent delivery model partitions when each subset M0, M1, and M2 are sent and inferred (e.g., executed) on the same node (e.g., WTRU 102). A delivery model partition may define and / or refer to the boundaries of parts of a model (e.g., subsets) that can be sent and executed as units independently of other parts of the whole model (e.g., subsets).
[0065] Examples (a) and (b) in Figure 2 show, respectively, examples of AI / ML inference nodes running an AI / ML model M consisting of two subsets, M0 and M1. For example, a node (e.g., a network endpoint or WTRU endpoint) can run the AI / ML model subset M0 while downloading the other subset M1.
[0066] In some representative embodiments, a first subset (e.g., M0) may be requested (e.g., selected by WTRU102). The first subset may be received (e.g., by an inference engine run by WTRU102). The received first subset may be used to obtain (e.g., intermediate) inference results from the first subset, such as using a given media content (e.g., images, sequences of images, videos, audio, text, and / or video) as input to an inference performed using the first subset. For example, the inference engine of WTRU102 may run the first subset while a second subset (e.g., M1) is requested (e.g., by WTRU102) for delivery. The second subset may be received (e.g., by an inference engine run by WTRU102) and sent to the inference engine. The received second subset may be used to obtain (e.g., intermediate) inference results from an inference engine that performs inference using the second subset, for example, by using the inference results (e.g., the output from performing inference using the first subset) as input for an inference engine that performs inference using the second subset.
[0067] For example, an AI model file can be distinguished from a general-purpose software file. An AI model file may be segmented and composed of different sequentially executable subsets, where the output of a first inference (e.g., performed using a first subset) is used as input to a second inference (e.g., performed using a second subset). In other words, an AI model file may be decomposed into multiple sequential loadable and runnable subsets, where, for example, the output from an inference performed using each subset is provided as input to an inference performed using the next subset.
[0068] Examples (c) and (d) in Figure 2 show an AI / ML split model in which subsets M1 and M1' are run on the network while subsets M0 and M'0 are run on the WTRU. For example, these configurations can be called split AI / ML model inference.
[0069] As described herein, the terms AI model, ML model, and AI / ML model may be used interchangeably.
[0070] Currently, there is no detailed architecture for providing AI model downloads and / or streaming in 3GPP.
[0071] Therefore, the architecture, components, and methods for distributing AI models from the network to WTRU102 should preferably include the following:
[0072] • Representation for AI model data composition and distribution • 5G AI model distributed system for downlink communication • Steps for progressive download of a full AI model (e.g., on-demand) • Procedures for progressive download of AI models using streaming and / or incremental model loading • Steps to select whether to download a full AI model or an incremental AI model.
[0073] AI Model Data Composition Figure 3 is a block diagram showing an exemplary functional model distributed architecture 300 according to one embodiment. In Figure 3, the WTRU 102 and the network 302 (e.g., RAN 113 and core network 115) can communicate to exchange AI model data 304. For example, the AI model may be associated with the WTRU application 306 and / or the network application 308.
[0074] In some representative embodiments, the AI model delivery function 310 in the network can deliver AI model data 304 (e.g., AI / ML models) from the AI model repository 312 to the WTRU 102 via 5GS. For example, the AI model repository 312 can store multiple AI models and / or various compositions thereof received from the AI model builder 314. The AI model builder may include an encapsulation function 316 and / or a compression function 318.
[0075] In some representative embodiments, the AI model access function 320 within the WTRU 102 can receive AI model data 304 and feed it to the AI model inference engine 322. For example, the AI model inference engine 322 may include a decapsulation function 324 and / or a decompression function 326 to complement the encapsulation function 316 and / or compression function 318.
[0076] In some representative embodiments, the AI model data 304 may be divided into AI model subsets. The AI model subsets may be structured or unstructured. For example, a structured AI model (e.g., a subset) may include the structure of a neural network (NN) model, as well as the relevant data used for inference (e.g., a finite set of DNN (deep neural network) layers having the necessary DNN layer data). For example, an unstructured AI model may include fragments of model data that are divided (e.g., cut) into different data chunks that require aggregation to constitute a structured AI model (e.g., a subset).
[0077] In some representative embodiments, the AI model inference engine 322 may download a structured AI model subset, load it into memory 130 and / or 132, and run the subset (for example, to obtain intermediate results) before doing the same for any subsequent AI model subsets. For example, this could be called incremental model loading with progressive downloading of the AI model. For example, incremental model loading with progressive downloading may be applicable (for example, only) to a structured AI model subset.
[0078] In some representative embodiments, AI model data may be represented using different formats such as ONNX (Open Neural Network Exchange) and NNEF (Neural Network Exchange Format).
[0079] In some representative embodiments, AI model data may be compressed using different codecs, such as NNC (Neural Network Constructor). For example, WTRU102 and network 302 can use a common AI model data profile to provide interoperability.
[0080] In some representative embodiments, compressed AI model data or AI model data representations may be encapsulated using a container format such as the ISO-based media file format (ISOBMFF). For example, additional file format structures (e.g., boxes / atoms) may need to be defined for the encapsulation of AI model data. These boxes may define metadata that allows a parser to easily extract AI model data from the file. Additionally, if the container file carries model data applied or used at different points in time, additional track types and associated metadata boxes may be defined.
[0081] When downloading AI model data, the sending entity can use Multipurpose Internet Mail Extension (MIME) multipart messages to encapsulate different subsets of the AI model, signaling to the receiving end that the received model data has different logical parts. Additional top-level media types may be defined for AI model data.
[0082] Architectural Components for Distributed AI Models 5G AI Model Distributed System for Downlink (5GAImDSd) In some representative embodiments, a 5G AI model distributed system can be used for downlink communication of AI models. Figure 4 is a block diagram of a 5G AI model distributed system 400 for downlink communication of AI models.
[0083] In some representative embodiments, the WTRU 102 can run a 5GAImDSd application 402 (for example, an application for receiving inference results from an AI model) and / or a 5GAImDSd media client 404. For example, the 5GAImDSd media client 404 may include two (sub)functions, namely an AI model session handler 406 and an inference engine 408.
[0084] The AI model session handler 406 may be a (sub)function running on WTRU 102 that communicates with 5GAImDSd application functions (AF) 410 in the network to establish, control, and support the delivery of AI model sessions. The AI model session handler 406 may perform additional functions such as collecting and reporting consumption and QoE (Quality of Experience) metrics. The AI model session handler 406 may expose one or more APIs that can be used by the 5GAImDSd application 402.
[0085] The inference engine 408 may also be a (sub)function running on WTRU 102 that communicates with a 5GAImDSd application server (AS) 412 in the network to obtain model data, and can provide APIs to the 5GAImDSd application 402 for model data delivery and to the AI model session handler 406 for AI model session control.
[0086] The 5GAImDSd application 402 may be an external AI media application. The 5GAImDSd application 402 can control the 5GAImDSd AI media client 404, implement logic specific to the external application and / or content service provider, and / or enable the establishment of an AI model session. Although the 5GAImDSd application 402 may not be defined within the 5G Technical Specifications Group Services and System Aspects Working Group 4 (SA4) specifications, the application 402 or equivalent functionality can utilize the 5G AI media client 404 and network functions using the 5G model AI interface and APIs.
[0087] The 5GAImDSd application server (AS) 412 may be an application server that hosts 5G AI model functionality. In some typical embodiments, different implementations of the 5GAImDSd AS 412 may exist, including the distribution of 5GAImDSd AS functionality across different physical hosts, such as within a Content Delivery Network (CDN).
[0088] In some representative embodiments, service access information may refer to a set of parameters and addresses used (e.g., required) by the 5GAImDSd media client 404 to activate the reception of an AI model streaming session (e.g., downlink). For example, service access information may include one or more AI model data input points.
[0089] In some representative embodiments, service and content discovery may refer to functionality and / or procedures provided by a 5GAImSd application provider 414 to a 5GAImDSd recognition application 402 that enables an end user to discover available distributed service and content offerings and select a specific service or content item (e.g., a specific AI model) for access.
[0090] In some representative embodiments, service notification may refer to a procedure performed between the 5GAImDSd-aware application 402 and the 5GAImSd application provider 414 so that the 5GAImDSd-aware application 402 can obtain service access information. For example, the 5GAImDSd-aware application 402 can obtain service access information from the 5GAImSd application provider 414 (e.g., directly). For example, the 5GAImDSd-aware application 402 can obtain a reference and / or part of the service access information from the 5GAImSd application provider 414 (e.g., directly) and obtain other service access information (e.g., the remainder of the service access information) from the 5GAImDSd AF 410.
[0091] In some representative embodiments, 5GAImDSd AS412 may support and / or provide any of the following features: (i) ingesting AI / ML models from a 5GAImDSd application provider 414 (e.g., at reference point M2d); (ii) caching AI / ML model content (e.g., to reduce the need to repeatedly ingest the same content at reference point M2d); (iii) a framework (e.g., a general one) for AI model data preparation; (iv) domain name aliasing (e.g., at reference point M4d); (v) server certificate support (e.g., at reference point M4d); (vi) URL path rewriting (e.g., at reference point M4d); and / or (vii) URL signing (e.g., at reference point M4d).
[0092] The 5GAImDSd application provider 414 may be an external application or content-specific AI media functionality (e.g., AI model creation, encoding, and formatting) that distributes AI models to the 5GAImDSd recognition application 402 using the 5GAImDSd interface.
[0093] The 5GAImDSd AF410 may be an application function that provides various control functions to the AI model session handler 406 on the WTRU102 and / or to the 5GAImDSd application provider 414. It may relay and / or initiate requests for processing of different policies or billing functions (PCFs) 416, or interact with other network functions via the network exposure function (NEF) 418.
[0094] In some representative embodiments, the following interfaces may be defined for 5G downlink AI model distribution.
[0095] For example, interface M1d (e.g., 5GAImDSd provisioning API) may be an external API exposed by 5GAImDSd AF410, which allows 5GAImDSd application provider 414 to provision the use of the 5G AI model distributed system for downlink AI / ML model data and / or to receive feedback.
[0096] For example, an M2d interface (e.g., the 5GAImDSd Ingest API) could be an optional external API exposed by a 5GAImDSd AS412, used when a 5GAImDSd AS412 in a (e.g., trusted) DN is selected to host content for a delivery service.
[0097] For example, interface M3d could be an internal (e.g., non-3GPP designated) API used to exchange information for content hosted on 5GAImDSd AS412 within a (e.g., trusted) DN.
[0098] For example, the M4d interface (e.g., Model Distribution API) could be one or more APIs exposed to the inference engine 408 by 5GAImDSd AS412 to deliver AI model data content.
[0099] For example, interface M5d may be one or more APIs (e.g., model data session handling APIs) exposed to the AI / ML model session handler 406 by 5GAImDSd AF410 for model data session handling, control, reporting, and support. One or more security mechanisms, such as authorization and authentication, may be supported and / or included.
[0100] For example, interface M6d (e.g., WTRU AI media session handling API) may be one or more APIs exposed to the inference engine 408 by the AI / ML session handler 406 for client internal communication. Interface M6d may be exposed to the 5GAImDSd recognition application 402, which will be able to utilize 5GAImDSd functionality.
[0101] For example, interface M7d (e.g., WTRU inference engine API) may be one or more APIs exposed by the inference engine to the 5GAImDSd recognition application 402 and the AI model session handler 406 in order to utilize the inference engine 408.
[0102] For example, interface M8d (e.g., an application API) could be an application interface used for information exchange between a 5GAImDSd recognition application 402 and a 5GAImDSd application provider 414. For example, interface M8d could be used to provide service access information to the 5GAImDSd recognition application 414. As an API, interface M8d may be outside the 5G system and may not be specified by 5G Media Streaming (5GMS).
[0103] WTRU 5GAImDSd function In some representative embodiments, WTRU102 may include one or more (sub)functions that can be individually used and / or controlled by a 5GAImDSd recognition application 402 (for example, performed by WTRU102).
[0104] The 5GAImDSd recognition application 402 itself may include one or more (sub)functions not provided by the 5GAImDSd client 404 or WTRU 102. Examples include service and AI model discovery, notification, and social network integration. The 5GAImDSd recognition application 402 may also include functions equivalent to those provided by the 5GAImDSd media client 404, or it may use only a subset of the 5GAImDSd client functions. The 5GAImDSd recognition application 402 may operate based on user input, or it may receive remote control commands from the 5GAImDSd application provider 414, for example, via M8d.
[0105] Figure 5 is a block diagram showing the components of an inference engine in a 5G AI model distributed system according to one embodiment. Figure 5 shows the functional components of the inference engine 408 for accessing 5GMSd AS412. In some representative embodiments, one or more of the following components may be provided by the inference engine 408.
[0106] For example, the inference engine 408 may include an AI model access client 501 that accesses AI model content for AI model subset distribution, such as file-based AI model subsets or DASH-formatted AI model subsets.
[0107] For example, the inference engine 408 may include an inference framework library 503 that can extract AI model data (e.g., basic) such as neural network models, as well as relevant data used for inference.
[0108] For example, the inference engine 408 may include an AI model decompression function 505, which can extract (e.g., basic) AI model data when the AI model data is compressed (e.g., using a neural network representation (NNR)).
[0109] For example, the inference engine 408 may include a neural network hardware API 507, which can provide acceleration to the inference engine 408 using a supported hardware accelerator such as a graphics processing unit (GPU), a digital signal processor (DSP), a neural processing unit (NPU), or a central processing unit (CPU).
[0110] For example, the inference engine 408 may include a metric measurement and logging client 511 that can perform QoE metric measurement and logging according to a metric reporting configuration of provisioning data supplied to the 5GAImDSd AF410 by the 5GAImDSd application provider 414 and transferred to the inference engine 408 via the media session handler 406 by the 5GAImDSd AF410.
[0111] For example, the inference engine 408 may include an AI inference engine runtime 509 function that can feed and run AI model data (e.g., decapsulated and unpacked) received from the AI model access client 501.
[0112] In some representative embodiments, one or more of the following components may be provided as part of the AI model AS412.
[0113] For example, 5GAImDSd AS412 may include an AI model delivery server 521 that can deliver AI model content for AI model subset distribution, such as file-based AI model subsets or DASH-formatted AI model subsets.
[0114] For example, 5GAImDSd AS412 may include an inference framework library 523, which may encapsulate AI model data (e.g., basic) such as neural network models, as well as related data used for inference.
[0115] For example, 5GAImDSd AS412 may include an AI model compression function 525 that can compress (for example, basic) AI model data (for example, using NNR).
[0116] Figure 6 is a block diagram showing the components of an AI model session handler in a 5G AI model distributed system according to one embodiment. As shown in Figure 6, the AI model session handler 406 accesses the 5GAImDSd AF410. In some representative embodiments, the AI model session handler 406 may include one or more of the following components.
[0117] For example, the AI model session handler 406 may include one or more core functions 602 for realizing the "session" concept for media communication, and may optionally extend to multiple stateless sessions. The core functions 602 may interact with the network-based 5GAImDSd AF410.
[0118] For example, the AI model session handler 406 may include a metric collection and reporting function 604, which can collect QoE metric measurement logs from the inference engine 408 for the purpose of metric analysis or to enable potential transport optimization of AI model data distribution by network, and send metric reports to the 5GAImDSd AF410. An example may include model performance metrics including any (e.g., a combination) of (i) model precision, (ii) model fineness, (iii) model recovery, (iv) mean squared error, and / or (v) absolute error.
[0119] For example, the AI model session handler 406 may include network assistance and QoS functions 606 that coordinate the downlink distribution assistance provided by the network to the 5GAImDSd client 404 and the inference engine 408.
[0120] For example, the AI model session handler 406 may include a WTRU capability reporting function 608 that can monitor WTRU capability for the purpose of AI model selection on the network side and report the WTRU capability to the network monitoring function 612 of the 5GAImDSd AF410. Examples of WTRU capability information may include (i) available memory allocated to the AI / ML service, (ii) processing power available for WTRU inference (e.g., CPU, GPU, TPU, and / or NPU), (iii) maximum available energy consumption, (iv) computing performance capabilities (e.g., floating-point operations or flops per second), (v) current model inference latency, and / or (vi) WTRU location (e.g., a combination of these).
[0121] For example, the AI model session handler 406 may include a WTRU model selection function 610 that can select an AI / ML model for different tasks depending on the WTRU capabilities and network conditions for a WTRU-based selection mode. The WTRU model selection 610 can communicate with the network model selection entity 614 of the 5GAImDSd AF410, for example, when the model selection is shared between the WTRU 102 and the network.
[0122] In some representative embodiments, additional interfaces and / or APIs not shown in Figure 6 may reside within WTRU102, such as (i) an AI model control interface for configuring and interacting with different WTRU AI model functions, (ii) an AI model control interface for AI model session management, (iii) a control interface for collecting logged QoE metric measurements, (iv) a control interface for collecting logged content consumption measurements, (v) handling of AI model data samples to the inference engine 408, and / or (vi) handling of decoded and compressed AI model samples to an (e.g., trusted) AI model decoder (e.g., a combination thereof).
[0123] In some representative embodiments, the 5GAImDSd AF410 may include either (i) a WTRU capability monitoring function for monitoring WTRU capability on the network side, such as through communication with the WTRU capability reporting function 608, and / or (ii) a network model selection function that selects a model for different tasks depending on the monitored WTRU capability, network conditions, and server-side resources for distributing AI models to the WTRU 102, etc. (e.g., a combination thereof). The selection of (e.g., available) AI models may be called a network-based selection mode. When the selection is shared between the WTRU 102 and the network, the network model selection function may communicate with the WTRU model selection 610 function.
[0124] Progressive download of AI models Full AI model, full model progressive download In some representative embodiments, the WTRU102 can perform steps to establish a downlink session for streaming an AI model. For example, the streaming session may use the 3GP file format (e.g., for progressive download), 3GP Timed Text, or other (e.g., non-3GPP) formats.
[0125] Figure 7 is a high-level signal flow diagram illustrating an example of progressive download for on-demand AI model content according to an embodiment. For example, although not shown in Figure 7, it can be assumed that a 5GAImDSd application provider 414 provisions a 5G AI model distributed system for downlink and sets up content ingestion, and a 5GAImDSd recognition application 402 receives a service notification from the 5GAImDSd application provider 414.
[0126] In 702, the 5GAImDSd recognition application 402 can trigger service announcements and service and content discovery procedures. Service announcements may include either the entire service access information (e.g., details for AI model session handling via interface M5d, and details for AI model streaming access via interface M4d, etc.) or a reference to the service access information.
[0127] For example, the 5GAImDSd recognition application 402 may request an AI model in 702a by using the "Get AI Model Session Information" message. The request may indicate whether the application is requesting a full model, a structured model, and / or any full / structured model composition.
[0128] For example, application provider 414 may, in 702b, provide and transmit (e.g., in response to a request) a list of AI models (e.g., AI model session URLs) with additional metadata including information about the model type (e.g., full model type and / or structured model type) via 5GAImDSd AS412. The list may include different AI / ML compositions, including any full models and any structured models available for download. This can provide an alternative to WTRU102 for selecting models according to various WTRU capabilities or requirements.
[0129] For example, WTRU102 can select an AI / ML model based on any of the following: (i) evaluating its internal capabilities (e.g., memory and / or processing power) to process all or part of a structured model; (ii) obtaining intermediate results before continuing to download additional parts of the structured model; and / or (iii) evaluating inference latency to obtain early intermediate or final results from inferring part of the structured model or the full model itself.
[0130] For example, a list of mixed compositions of full models and structured models may include (i) full model #1, (ii) full model #2, (iii) structured model #3 composition (e.g., adapted for incremental loading) such as subset 1, subset 2, subset 3, and / or (iv) structured model #4 composition (e.g., adapted for incremental loading) such as subset 1', subset 2', subset 3'.
[0131] In some representative embodiments, the service and content discovery procedure in 702 may include (for example, only) a 5GAImDSd recognition application 402 and a 5GAImDSd application provider 412.
[0132] In 704, the full AI model may be selected, for example, from the list of candidate models obtained in 702. The selection may take into account any AI requirements regarding the achievable model performance for the capabilities that WTRU102 may or may want to assign to running the AI model.
[0133] In 706, the 5GAImDSd recognition application 402 can trigger the AI model session handler 406 to begin inference. An inference engine entry may be provided to the AI model session handler 406. The application 402 can select or assist the inference engine 408 in selecting an inference entry from among the available inference processes running within WTRU 102. A process may be an allocated TPU (Tensor Processing Unit), GPU (Graphical Processing Unit), CPU (Central Processing Unit) process, a software process, or a virtual machine instance running on WTRU 102.
[0134] In some embodiments, when the 5GAImDSd recognition application 402 receives only a reference to service access information in 702, the AI model session handler 406 may (optionally, e.g.) interact with the 5GAImDSd AF410 to obtain the entire service access information. In that information, the 5GAImDSd AF410 may provide the AI model session handler 406 with the AI model data encapsulation and / or compression format to be used (e.g., ONNX, NNEF, or NNC). In some embodiments, the 5GAImDSd AF410 may provide (e.g., transmit) information about the model data composition, whether or not the AI model subset is structured.
[0135] In 710, the AI model session handler 406 may provide and transmit information indicating the model selection (e.g., the URL associated with the selected AI / ML model) to the 5GAImDSd AF410. This may include information indicating the type of model selected (e.g., full or structured). If a structured model is selected, the AI model session handler 406 may provide (e.g., transmit) information about the parts of the structured model to download.
[0136] For example, depending on capabilities and / or requirements, WTRU102 may select structured model #3, which includes subsets 1 and 2, for initial download.
[0137] For example, WTRU102 can evaluate intermediate results (for example, based on inferences using subsets 1 and 2) before requesting the download of subset 3 (for example, if necessary).
[0138] In 710, the AI model session handler 406 can trigger the inference engine 408 to start the session.
[0139] At 712, the inference engine 406 can establish a transport session.
[0140] In 714, the inference engine 408 can send a request for progressive download of the selected AI model content to the 5GAImDSd AS412 in the network. For example, the request may include information indicating the selected AI model content (e.g., the URL associated with the selected AI / ML model).
[0141] In 716, the inference engine 408 can receive initial configuration information for progressive download of selected content from 5GAImDSd AS412. The initial configuration information may include configuration parameters for receiving AI model and / or digital rights management (DRM) information.
[0142] In 718, the inference engine 408 may configure its pipeline for loading AI model content for (for example, further) inference.
[0143] In 720, the inference engine 408 can notify the AI model session handler 406 by providing transport session information and (for example, some) AI model content-related information.
[0144] In 722, the inference engine 408 can (for example, optionally) obtain a license key and / or content key from the 5GAImDSd application provider 414 to decode the AI model data.
[0145] In 724, the inference engine 408 can receive AI model content. For example, the inference engine 408 can put the received AI model content into the rendering pipeline.
[0146] In 726, the inference engine 408 can receive (for example, continue to receive) AI model content. For example, the inference engine 408 can put each received part of the AI model content (for example, as they are received) into the rendering pipeline.
[0147] At 728, the inference engine 408 can receive the final portion of the AI model content. The entire AI model can be received by the inference engine 408.
[0148] In 730, the inference engine 408 can run (e.g., execute) the AI model to obtain one or more (e.g., final) results.
[0149] In step 732, the inference engine 408 can provide (e.g., send) the results obtained in step 730 to the application 402.
[0150] Model streaming or incremental model loading with progressive download of AI models In some representative embodiments, the WTRU102 can perform steps to establish a downlink streaming session for one or more subsets of AI models.
[0151] Figures 8A and 8B are signaling flow diagrams illustrating the (e.g., high-level) procedure for progressive download of AI models in subset streaming / incremental loading mode according to an embodiment. For example, although not shown in Figures 8A and 8B, it can be assumed that a 5GAImDSd application provider 414 provisions a 5G AI model distributed system for downlink and sets up content ingestion, and a 5GAImDSd recognition application 402 receives a service notification from the 5GAImDSd application provider 414.
[0152] In some representative embodiments, an AI model may consist of a set of AI model subsets. For example, the subsets may be organized in a linear sequence, such that the output of one subset serves as the input for the next subset.
[0153] In Figures 8A and 8B, the (e.g., selected) AI model subsets are structured. In Figure 7, the (e.g., selected) AI models are not structured. Because the AI model subsets are structured in Figures 8A and 8B, instead of waiting for the full AI model to be downloaded before performing inference using the full AI model (e.g., as in Figure 7), when the inference engine 408 receives and performs inference using the AI model subsets, it may begin to infer (e.g., perform) each AI model subset and send (e.g., immediately) their respective intermediate (e.g., partial) results to the 5GAImDSd recognition application 402.
[0154] In 802, the 5GAImDSd recognition application 402 can trigger service announcements and service and content discovery procedures. Service announcements may include either the entire service access information (e.g., details for AI model session handling via interface M5d, and details for AI model streaming access via interface M4d, etc.) or a reference to the service access information.
[0155] For example, the 5GAImDSd recognition application 402 can request an AI model by using the "Get AI Model Session Information" message in 802a. The request may indicate that the application is requesting a structured AI model.
[0156] For example, application provider 414 may, in 802b, provide and transmit (e.g., in response to a request) a list of AI models (e.g., AI model session URLs) with additional metadata including information about the model type (e.g., full model type and / or structured model type) via 5GAImDSd AS412. The list may include different AI / ML compositions, including any full models and any structured models available for download. This can provide an alternative to WTRU102 for selecting models according to various WTRU capabilities or requirements.
[0157] For example, WTRU102 can select an AI / ML model based on any of the following: (i) evaluating its internal capabilities (e.g., memory and / or processing power) to process all or part of a structured model; (ii) obtaining intermediate results before continuing to download additional parts of the structured model; and / or (iii) evaluating inference latency to obtain early intermediate or final results from inferring part of the structured model or the full model itself.
[0158] In some representative embodiments, the service and content discovery procedure in 802 may include (for example, only) a 5GAImDSd awareness application 402 and a 5GAImDSd application provider 412.
[0159] In 804, the structured AI model may be selected, for example, from the list of candidate models obtained in 802. The selection may take into account any AI requirements regarding the achievable model performance for the capability to which WTRU102 may or may want to be assigned to run the AI model.
[0160] In 806, the 5GAImDSd recognition application 402 can trigger the AI model session handler 406 to begin inference. An inference engine entry may be provided to the AI model session handler 406. Application 402 can select or assist the inference engine 408 in selecting an inference entry from among the available inference processes running within WTRU 102. A process may be an allocated TPU (Tensor Processing Unit), GPU (Graphical Processing Unit), CPU (Central Processing Unit) process, a software process, or a virtual machine instance running on WTRU 102.
[0161] In some embodiments, when the 5GAImDSd recognition application 402 receives only a reference to service access information in 802, the AI model session handler 406 may (optionally, e.g.) interact with the 5GAImDSd AF410 to obtain the entire service access information. In that information, the 5GAImDSd AF410 may provide the AI model session handler 406 with the AI model data encapsulation and / or compression format to be used (e.g., ONNX, NNEF, or NNC). In some embodiments, the 5GAImDSd AF410 may provide (e.g., transmit) information about the model data composition, whether or not the AI model subset is structured.
[0162] In 810, the AI model session handler 406 may provide and transmit information indicating model selection (e.g., a URL associated with the selected AI / ML model) to the 5GAImDSd AF410. This may include information indicating the type of model being selected (e.g., structured). The AI model session handler 406 may provide (e.g., transmit) information about the parts of the structured model to be downloaded.
[0163] For example, depending on capabilities and / or requirements, WTRU102 can select a structured model that includes a subset of all AI models for initial download.
[0164] For example, WTRU102 can evaluate intermediate results (for example, using subsets 1 and 2, based on inference) before requesting the download of any subsequent subsets (e.g., subset 3 if necessary).
[0165] In 810, the AI model session handler 406 can trigger the inference engine 408 to start a session.
[0166] At 812, the inference engine 406 can establish a transport session.
[0167] In 814, the inference engine 408 can send a request for progressive download of the selected AI model content to the 5GAImDSd AS412 in the network. For example, the request may include information indicating the selected AI model content (e.g., the URL associated with the selected AI / ML model).
[0168] In 816, the inference engine 408 can receive initial configuration information for progressive download of selected content from 5GAImDSd AS412. The initial configuration information may include configuration parameters for receiving AI model and / or digital rights management (DRM) information.
[0169] In 818, the inference engine 408 may configure its pipeline for loading AI model content for (for example, further) inference.
[0170] In 820, the inference engine 408 can notify the AI model session handler 406 by providing transport session information and (for example, some) AI model content-related information.
[0171] In 822, the inference engine 408 can (for example, optionally) obtain a license key and / or content key from the 5GAImDSd application provider 414 to decode the AI model data.
[0172] In 824, the inference engine 408 can download (e.g., receive) the first AI model subset and place the first AI model subset into the rendering pipeline.
[0173] In 826, the inference engine 408 can run (e.g., execute) a first AI model subset (e.g., even if the complete model has not yet been downloaded). Media content, such as images, sequences of images, videos, audio, text, and / or other data, may be provided as input to the first AI model subset.
[0174] In 828, the inference engine 408 can send intermediate results (e.g., from running a first AI model subset) to the 5GAImDSd recognition application 402. In some embodiments, the inference engine 408 may wait until the entire model (or any part thereof) has been run before sending any results to the 5GAImDSd recognition application 402.
[0175] In 830, the inference engine 408 can receive (for example, sequentially) (for example, continue receiving) AI model subsets. For example, the inference engine 408 can put each received AI model subset into the inference pipeline (for example, as they are received).
[0176] For example, in 832, the inference engine 408 can download (e.g., receive) the Nth AI model subset and place the Nth AI model subset into the rendering pipeline. In 834, the inference engine 408 can run (e.g., execute) the Nth AI model subset (e.g., even if the complete model has not yet been downloaded). For example, the Nth AI model subset can perform inference and use the output from a previous (e.g., the N-1th) AI model subset as input to generate intermediate results from the Nth AI model subset. In 836, the inference engine 408 can send the intermediate results (e.g., from running the Nth AI model subset) to the 5GAImDSd recognition application 402.
[0177] In 838, the inference engine 408 can receive (for example, sequentially) (for example, continue receiving) AI model subsets. For example, the inference engine 408 can put each received AI model subset into the rendering pipeline (for example, as they are received).
[0178] At step 840, the inference engine 408 can receive the last AI model subset. For example, the inference engine 408 can put each received AI model subset into the rendering pipeline (for example, when they are received). For example, at step 842, the inference engine 408 can run (e.g., execute) the last AI model subset, for example, by using the output from a previous (e.g., second-to-last) AI model subset.
[0179] In 844, the inference engine 408 can send the final result (e.g., from running the last AI model subset) to the 5GAImDSd recognition application 402. For example, the final result may correspond to the output of the last AI model subset, if the output of the previous AI model subset was provided as input to the last AI model subset.
[0180] AI model format In some representative embodiments, the AI model can use one of the following formats: Open Neural Network Exchange (ONNX), Neural Network Exchange (NNEF), or Neural Network Coding and Representation (NNR). In some representative embodiments, the AI model may undergo decomposition as follows:
[0181] For example, the ONNX format is built around protocol buffers, and the ONNX graph can be structured as a list of nodes forming an acyclic graph that describes the AI model. It can provide the metadata necessary for extra model completion. There is a large set of built-in operators that describe node behavior, including (i) mathematical operators such as Abs, (ii) DNN operators such as Conv and LSTM, (iii) activation operators such as Sigmoid and Relu, (iv) pooling operators such as MaxPool, and (v) other operators such as error calculation and data reformatting operators.
[0182] For example, the NNEF format allows for the encapsulation of both the structure of a neural network model (e.g., an AI model) and the associated data used for inference. An NNEF container may contain a textual file that describes the structure of the neural network, which is described through a computation graph, such as a directed graph containing data or computation nodes. Computation nodes may have attributes that describe the exact computations that need to be performed. Computation nodes may be configured together to generate more complex computations. An NNEF container may contain binary data files for each variable tensor. These files may be hierarchically structured into subfolders associated with the corresponding computations. Each tensor may have different representations, such as matching different quantized versions. An NNEF container may contain a quantization file that contains details about the quantization algorithm used to quantize the exported tensors.
[0183] For example, a neural network compiler (NNC) (e.g., form) can specify a compressed representation format for neural network data and processes for its decoding. An NNC consists of a toolbox that can be flexibly selected. In particular, an NNC defines data structures and syntactic elements to support features such as (i) packaging of NN data, (ii) signaling of metadata related to various methods of preprocessing for data reduction, (iii) compression of NN weights / tensor coefficients, and (iv) interoperability.
[0184] For example, different types of NN data can be packaged into neural network representation (NNR) units for access from the system or application layer. NNR parameter set units and NNR layer parameter set units can carry metadata and information related to the entire NN layer and individual NN layers, respectively. NNR topology units can contain information about the NN topology (e.g., connections between layers / tensors). Actual tensor data can be propagated in NNR quantized information and NNR compressed data units. NNR aggregation units can enable combinations of several related NNR units of different types.
[0185] For example, metadata related to various preprocessing methods for data reduction can be signaled. This may include parameters related to sparsification, pruning, low-rank decomposition, unification, batch norm folding, and local scaling.
[0186] For example, NN weight / tensor coefficient compression can be performed using quantization and entropy coding. The tensor / weight coefficients may be signaled as raw data or quantized in a different way. The quantized coefficients may be binarized and entropy coded using a context-adaptive arithmetic coder (e.g., DeepCABAC).
[0187] For example, interoperability with other exchanges (e.g., NNEF, ONNX) or native formats (e.g., PyTorch, TensorFlow) may be supported. NNC may allow embedding other forms of topology information into the NNR bitstream. NNR units representing encoded tensors / weights may be embedded in other forms of containers.
[0188] Figure 9 is a block diagram showing the components for processing an NN into an NNR. As shown in Figure 9, (for example, the original) NN 902 can be processed into an NNR bitstream 904. For example, NN data representing an NNR can be provided to a preprocessing and / or parameter reduction processing unit 906, a quantization processing unit 908, and / or an entropy coding processing unit 910. For example, the preprocessing and / or parameter reduction processing unit 906 may include any of the sparsification function 912, pruning function 914, local scaling function 916, LR decomposition function 918, unification function 920, and / or batch norm folding function 922. For example, the quantization processing unit 908 may receive NN data and / or the output from the preprocessing and / or parameter reduction processing unit 906 (for example, an NNR unit) as input, and the quantization processing may include any of the uniform function 924, codebook 926, and / or dependent function 928. For example, the entropy encoding processing unit 910 can receive NN data and / or the output from the quantization processing unit 908 (e.g., an NNR unit) as input and may include any of the binarization function 930, the context modeling function 932, and / or the arithmetic encoding function 934. The output from the preprocessing and / or parameter reduction processing unit 906, the quantization unit 908, and / or the entropy encoding processing unit 910 (e.g., an NNR unit) may form an NNR bitstream 904. In other examples, non-NN AI models may be processed into a bitstream.
[0189] Figure 10 is a block diagram showing the components of an NNR bitstream. As shown in Figure 10, the NNR bitstream 904 may comprise a plurality of NNR units 1002. For example, an NNR unit 1002 may contain information including an NNR unit size 1004, an NNR unit header 1006, and / or an NNR unit payload 1002.
[0190] 3GP file type In some representative embodiments, the 3GPP file format (3GP) can be used as an instance of an ISO-based media file format.
[0191] In some representative embodiments, the transmission of media content (e.g., an AI model) to a receiving terminal may use file download, streaming, or multimedia broadcast / multicast service (MBMS) download delivery. In the first and last cases, a self-contained file may be transmitted. In the second case, for example, real-time transport protocol (RTP) streaming, the content may be extracted from a file and streamed according to an open payload format. In this case, no trace of the file format remains in the transmitted content (e.g., via an air / wireless interface).
[0192] In segmented streaming on DASH, files can be divided into segments for transmission.
[0193] For example, an AI model might be a self-contained file download. An AI model composition can contain different AI subsets terminating at a specific neural network boundary.
[0194] Progressive Download Figure 11 is a block diagram illustrating an exemplary overview of a protocol stack that may be used for services such as those described herein. As shown in Figure 11, the protocol stack may include IP1102, TCP1104, HTTP1106, media presentation description1108, 3GP file format1110, and any of the video format, audio format, speech format, timed text format, and / or AI model format1112.
[0195] In some representative embodiments, 3GP files may be accessible using progressive download. In some representative embodiments, segments based on the 3GPP file format may be accessible via HTTP. For example, progressive download may provide partial transmission of 3GP files and / or segments (for example, using HTTP with the header "application / 3gpp-partial" in combination with an HTTP GET request).
[0196] In some representative embodiments, progressive downloading can be used to download AI models from the network to the WTRU102.
[0197] In some representative embodiments, partial transmission can be used for delivering (e.g., downloading) a subset of an AI model.
[0198] Dynamic Adaptive Streaming (DASH) via Hypertext Transfer Protocol Figure 12 is a system diagram illustrating an exemplary system using DASH segmentation for AI model delivery. As shown in Figure 12, content server 1202 can communicate with DASH client 1204 regarding transport protocols (e.g., performed by WTRU 102) and perform media presentation description (MDP) delivery. Content server 1202 may include an MDP unit 1206 and various AI models 1208 available for download. DASH client 1204 may comprise a control heuristic 1210, an MPD parser 1212, a segment parser 1214, a transport access client 1216, and one or more media players 1218.
[0199] In some typical embodiments, an AI model can be thought of as being similar to a file.
[0200] In some representative embodiments, DASH segmentation may be used to deliver an AI model consisting of a subset of AI model data. For example, segmentation may be independent of the AI model data composition as a bitstream of encapsulated, compressed, and / or serialized AI model data chunks.
[0201] For example, a segmentation representation can provide a description of a closed group of AI model data or model subsets that can be run by AI model inference. For instance, it may include a finite set of DNN layers that have the necessary DNN layer data.
[0202] One-way transport file delivery (FLUTE) for AI model access clients. Figure 13 is a block diagram illustrating an exemplary data structure for FLUTE. As shown in Figure 13, the transport unit 1300 may include a UDP header 1302, a default LCT header 1304, an LCT header extension 1306, an FEC payload ID 1308, and a FLUTE payload (e.g., encoding symbols) 1310. FLUTE provides file delivery over a one-way UDP-based transport. FLUTE can be used to optimize latency for file delivery. FLUTE can enable IP multicast according to a reliable multicast transport (RMT).
[0203] However, FLUTE adds a delivery size overhead to provide (e.g., additional) error correction techniques, known as forward error codes (FEC), which are used to detect and correct errors in the transmitted data.
[0204] Real-Time Transport Object Delivery Over Unidirectional Transport (ROUTE: Real-Time Transport Object Delivery Over Unidirectional Transport) DASH Figure 14 is a block diagram illustrating an exemplary protocol unit for ROUTE DASH. As shown in Figure 14, the protocol unit 1400 may include a DASH header 1402, a FLUTE header 1404, a UDP header 1406, and an IP multicast payload 1408. For example, ROUTE DASH provides DASH segmentation over the one-way ROUTE protocol. It can also enable IP multicast delivery of DASH segments (e.g., AI model subsets). An application server acting as a carousel multicast server can simultaneously provide a large set of WTRUs 102.
[0205] In some representative embodiments, a large set of WTRU102s may want to download and run AI models over a 5G link with limited network resources (e.g., congested locations and / or events). Route Dash metadata and signaling can be optimized to provide real-time delivery of AI models.
[0206] Typical procedure Figure 15 is a step diagram illustrating an exemplary procedure for WTRU102 to download AI models from the network. As shown in Figure 15, at 1502, WTRU102 can receive information from the application provider indicating a set of AI models and a set of addresses associated with the set of AI models. At 1504, WTRU102 can use the addresses associated with the AI models from the set of addresses to send a request for an AI model (e.g., from the set of AI models) to the network entity running the application server. At 1506, WTRU102 can receive information from the network entity indicating the AI model content corresponding to the AI models. At 1508, WTRU can use the received AI model content to obtain results (e.g., the full / final output of the inference).
[0207] In some representative embodiments, the WTRU102 can establish a session (e.g., transport) with a second network entity (e.g., one that performs application functions). For example, AI model content may be received from the second network entity via the established session.
[0208] In some representative embodiments, the AI model content may correspond to the full model of the AI model.
[0209] In some representative embodiments, WTRU102 can be configured with a full model to run an inference engine. The results in 1512 can be obtained from the configured inference engine.
[0210] In some representative embodiments, the WTRU102 can establish a session (e.g., transport) with a second network entity (e.g., one that performs application functions). For example, AI model content may be received from the second network entity via the established session.
[0211] In some representative embodiments, the AI model content may correspond to one or more subsets of the AI model.
[0212] In some representative embodiments, WTRU102 can configure an inference engine to be executed by WTRU102 using each subset of the AI model. For example, WTRU102 can obtain one or more intermediate results from the configured inference engine.
[0213] In some representative embodiments, one or more intermediate results (e.g., each of them) (e.g., from an inference engine) may be provided to an application executed by WTRU102.
[0214] In some representative embodiments, the WTRU102 can provide the final result (e.g., from an inference engine) to the application being executed by the WTRU102.
[0215] In some representative embodiments, WTRU102 can receive a selection of an AI model from a set of AI models.
[0216] In some representative embodiments, the WTRU102 can receive decryption information (e.g., a content key) associated with the AI model from the application provider. The WTRU102 can use the decryption information to decrypt information that indicates the AI model content.
[0217] In some representative embodiments, the AI model content may include or be composed of Dynamic Adaptive Streaming (DASH) segments transmitted via multiple hypertext transfer protocols. For example, WTRU102 can aggregate DASH segments to obtain an AI model.
[0218] In some representative embodiments, the AI model content may include or consist of one or more AI model files. For example, WTRU102 may receive any (e.g., each) AI model file as multiple sequentially executable subsets.
[0219] In some representative embodiments, WTRU102 can send information indicating an AI model (e.g., selected) from a set of AI models to a second network entity (e.g., one that performs an application function).
[0220] In some representative embodiments, the WTRU102 can receive information from a second network entity (e.g., one performing an application function) indicating the data encapsulation and / or compression format associated with the AI model.
[0221] In some representative embodiments, the WTRU102 can receive information representing the AI model content as a bitstream.
[0222] In some typical embodiments, the address received in 1502 may be a Uniform Resource Locator (URL) associated with a set of AI models.
[0223] In some representative embodiments, WTRU102 can perform procedures for the (e.g., progressive) download of (e.g., unstructured) AI models from the network. WTRU102 can run an inference engine (IE). For example, WTRU102 can select an AI model from a list of candidate AI models to download from the network. WTRU can trigger the IE to initiate the procedure for downloading the selected progressive AI model. The IE can establish a transport session with the network. The IE can send a request for the progressive download of the selected AI model to the network. The IE can receive a first portion of the selected progressive AI model from the network and place the first portion of the selected AI model into the rendering pipeline. The IE can receive a second portion of the selected AI model from the network and place the second portion of the selected AI model into the rendering pipeline. The IE can receive a final portion of the selected AI model from the network and place the final portion of the selected AI model into the rendering pipeline. After the final portion of the selected AI model is received and placed into the rendering pipeline, IE can run the selected AI model (for example, to obtain inference results).
[0224] In some representative embodiments, the WTRU102 can perform network-based service announcement and service and content discovery procedures to obtain at least one of service access information and AI model streaming access information.
[0225] In some representative embodiments, the WTRU 102 can select an AI model to download depending on the WTRU's ability to run the AI model and / or the functional requirements of the candidate AI model.
[0226] In some representative embodiments, the WTRU102 can receive initial configuration information for the selected AI model content in response to a request. For example, the initial configuration information may include configuration parameters for receiving the selected AI model.
[0227] In some representative embodiments, WTRU102 can perform procedures for the (e.g., progressive) download of (e.g., structured) AI models from the network. WTRU102 can run an inference engine (IE). For example, WTRU102 can select an AI model from a list of candidate AI models to download from the network. WTRU102 can trigger the IE to begin the procedure for downloading the selected AI model. The IE can establish a transport session with the network. The IE can send a request for the progressive download of the selected (e.g., structured) AI model. The IE can receive a first part of the selected AI model from the network and execute the first part of the selected AI model. The IE can receive a second part of the selected AI model from the network and execute the second part of the selected AI model. The IE can receive a final part of the selected AI model from the network and execute the final part of the selected AI model.
[0228] In some representative embodiments, the WTRU102 can perform network-based service announcement and service and content discovery procedures to obtain at least one of service access information and AI model streaming access information.
[0229] In some representative embodiments, the WTRU 102 can select an AI model depending on the WTRU's ability to run the AI model and the functional requirements of the candidate AI models.
[0230] In some representative embodiments, the WTRU102 can receive initial configuration information for the selected AI model content in response to a request. For example, the initial configuration information may include configuration parameters for receiving the selected AI model.
[0231] References The contents of the following references are incorporated herein by reference: (1) Non-Patent Document 1, (2) Non-Patent Document 2, (3) Non-Patent Document 3, and (4) Non-Patent Document 4.
[0232] conclusion While features and elements are provided above in specific combinations, those skilled in the art will understand that each feature or element may be used alone or in any combination with other features and elements. This disclosure is intended to be illustrative of various embodiments and should not be limited to the specific embodiments described in this application. As will be apparent to those skilled in the art, many modifications and variations can be made without departing from the spirit and scope of the invention. Elements, actions, or instructions used in the description of this application should not be construed as important or essential to the invention unless so expressly provided. In addition to those enumerated herein, functionally equivalent methods and apparatus within the scope of this disclosure will be apparent to those skilled in the art from the foregoing description. Such modifications and variations are intended to fall within the scope of the appended claims. This disclosure should be limited only by the entire scope equivalent to that which such claims are granted, together with the appended claims. It should be understood that this disclosure is not limited to any particular method or system.
[0233] The embodiments described above are explained for simplicity in terms of the terminology and structure of wirelessly communicable devices (e.g., radio wave emitters and receivers). However, the embodiments described are not limited to these systems and may be applied to other systems using other forms of electromagnetic waves or non-electromagnetic waves such as sound waves.
[0234] It should be understood that the terminology used herein is for the purpose of describing a particular embodiment and is not intended to be limiting. When used herein, the term “video” or the term “imagery” may mean a snapshot, a single image, and / or multiple images displayed on a time basis. As another example, when referred herein, the term “user equipment” and its abbreviation “UE,” the term “remote,” and / or the term “head-mounted display” or its abbreviation “HMD” may mean or include (i) a wireless transmit and / or receive unit (WTRU), (ii) any of several embodiments of a WTRU, (iii) a wirelessly and / or wired (e.g., tetherable) device configured using some or all of the structure and functionality of a WTRU, (iii) a wirelessly and / or wired device configured using less than all of the structure and functionality of a WTRU, or (iv) the same. Details of exemplary WTRUs that may represent any WTRU listed herein are provided herein with respect to Figures 1A to 1D. As another example, the various embodiments disclosed above and below in this specification are described as utilizing a head-mounted display. Devices other than head-mounted displays may be used, and those skilled in the art will recognize that some or all of the present disclosure and the various embodiments may be modified accordingly without undue experimentation. Examples of such other devices may include drones or other devices configured to stream information for providing an adapted reality experience.
[0235] In addition, the methods provided herein may be implemented in computer programs, software, or firmware embedded in computer-readable media for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted via wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital multipurpose disks (DVDs). Processors associated with software may be used to implement radio frequency transceivers for use in WTRUs, UEs, terminals, base stations, RNCs, or any host computer.
[0236] Modifications of the methods, apparatus, and systems provided above are possible without departing from the scope of the invention. In view of the wide range of applicable embodiments, it should be understood that the exemplary embodiments are merely examples and should not be taken as limiting the scope of the following claims. For example, embodiments provided herein include a handheld device which may include, or may be used with, any suitable voltage source, such as a battery, that provides any suitable voltage.
[0237] Furthermore, the embodiments provided above also illustrate other devices, including processing platforms, computing systems, controllers, and processors. These devices may include at least one central processing unit ("CPU") and memory. In accordance with the practice of those skilled in computer programming, references to acts and symbolic representations of actions or instructions may be performed by various CPUs and memories. Such acts and actions or instructions may be referred to as "executed," "computer-executed," or "CPU-executed."
[0238] Those skilled in the art will understand that actions and symbolically represented operations or instructions involve the manipulation of electrical signals by the CPU. The electrical system causes transformations or reductions resulting from the electrical signals and the preservation of data bits at memory locations within the memory system, thereby representing data bits that can reconfigure or otherwise alter the CPU's operation and other processing of the signals. The memory locations where the data bits are preserved are physical locations having specific electrical, magnetic, optical, or organic properties corresponding to or representing the data bits. It should be understood that the embodiments are not limited to the platforms or CPUs described above, and other platforms and CPUs may support the methods provided.
[0239] The data bits may also be maintained on a computer-readable medium, including magnetic disks, optical disks, and any other volatile (e.g., Random Access Memory (RAM)) or non-volatile (e.g., Read-On Memory (ROM)) mass storage systems readable by the CPU. The computer-readable medium may include collaborative or interconnected computer-readable media distributed among multiple interconnected processing systems, which may reside exclusively on a processing system or be local or remote to the processing system. It should be understood that the embodiments are not limited to the memory described above, and other platforms and memories may support the methods provided.
[0240] In exemplary embodiments, any of the operations, processes, etc., described herein may be implemented as computer-readable instructions stored on a computer-readable medium. These computer-readable instructions may be executed by processors in mobile units, network elements, and / or any other computing devices.
[0241] There is little distinction left between hardware and software implementations of a system configuration. The use of hardware or software is generally (but not always, in that in certain situations the choice between hardware and software can be important) a design choice representing a cost-efficiency trade-off. There may be various means (e.g., hardware, software, and / or firmware) by which the processes and / or systems and / or other technologies described herein can be delivered, and the preferred means may vary depending on the context in which the processes and / or systems and / or other technologies are deployed. For example, if the implementer determines that speed and accuracy are paramount, the implementer may choose primarily hardware and / or firmware means. If flexibility is paramount, the implementer may choose primarily software implementation. Alternatively, the implementer may choose any combination of hardware, software, and / or firmware.
[0242] The detailed description above illustrates various embodiments of devices and / or processes through the use of block diagrams, flowcharts, and / or examples. Those skilled in the art will understand that, insofar as such block diagrams, flowcharts, and / or examples include one or more functions and / or operations, each function and / or operation within such block diagrams, flowcharts, or examples can be implemented individually and / or collectively by a wide range of hardware, software, firmware, or virtually any combination thereof. In one embodiment, some parts of the subject matter described herein may be implemented via application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), and / or other integrated formats. However, those skilled in the art will recognize that some aspects of the embodiments disclosed herein can be uniformly implemented in integrated circuits, either as one or more computer programs running on one or more computers (e.g., as one or more programs running on one or more computer systems), as one or more programs running on one or more processors (e.g., as one or more programs running on one or more microprocessors), as firmware, or in virtually any combination thereof, and that designing circuits and / or writing code for software and / or firmware is well within the scope of the art of those skilled in the art in view of this disclosure. In addition, those skilled in the art will understand that the mechanisms of the subject matter described herein can be distributed as various forms of program products, and that the exemplary embodiments of the subject matter described herein apply regardless of the particular type of signal-carrying medium used to actually carry out the distribution.Examples of signal-carrying media include, but are not limited to, recordable media such as floppy disks, hard disk drives, CDs, DVDs, digital tapes, and computer memory, as well as transmittable media such as digital and / or analog communication media (e.g., fiber optic cables, waveguides, wired communication links, wireless communication links, etc.).
[0243] Those skilled in the art will recognize that it is common in the industry to describe devices and / or processes in the manner described herein and then, using engineering practices, integrate such described devices and / or processes into data processing systems. That is, at least some of the devices and / or processes described herein can be integrated into data processing systems through a reasonable amount of experimentation. Those skilled in the art will recognize that a typical data processing system may generally include one or more of the following: a system unit housing, a video display device, memory such as volatile and non-volatile memory, a processor such as a microprocessor and a digital signal processor, a computing entity such as an operating system, drivers, a graphical user interface and application program, one or more interaction devices such as a touchpad or screen, and / or control systems including feedback loops and control motors (e.g., feedback for sensing position and / or velocity, control motors for moving and / or adjusting components and / or quantities). A typical data processing system may be implemented using any suitable commercially available components, such as those typically found in data computing / communication and / or network computing / communication systems.
[0244] The subject matter described herein may include different components contained within or connected to other different components. Such illustrated architectures are merely examples, and it should be understood that many other architectures can be implemented to achieve the same functionality. Conceptually, any arrangement of components to achieve the same functionality is effectively “associated” in such a way that the desired functionality can be achieved. Thus, any two components combined herein to achieve a particular functionality, regardless of architecture or intermediate components, can be understood as “associated” with each other in such a way that the desired functionality can be achieved. Similarly, any two components thus associated can be considered “operably connected” or “operably coupled” with each other to achieve the desired functionality, and any two components that can be associated in such a way can also be considered “operably coupled” with each other to achieve the desired functionality. Specific examples of operable coupling include, but are not limited to, physically matable and / or physically interacting components, as well as / or wirelessly interactable and / or wirelessly interacting components, as well as / or logically interacting and / or logically interactable components.
[0245] With regard to the use of substantially any plural and / or singular terms herein, those skilled in the art can convert from plural to singular and / or singular to plural as appropriate to the context and / or use. Various singular / plural substitutions may be explicitly stated herein for clarity.
[0246] In general, it will be understood by those skilled in the art that terms used herein, particularly in the appended claims (e.g., the body of the appended claims), are generally intended as “open” terms (e.g., the term “including” should be interpreted as “including but not limited to,” the term “having” as “at least having,” and the term “includes” as “including but not limited to,” etc.). It will further be understood by those skilled in the art that if a particular number relating to the description of an introduced claim is intended, such intention is explicitly stated in the claim, and if such statement is not made, such intention does not exist. For example, if only one item is intended, the term “single” or similar wording may be used. For the sake of understanding, the following appended claims and / or descriptions herein may include the use of the introductory phrases “at least one” and “one or more” to introduce the description of a claim. However, the use of such phrases should not be interpreted as meaning that the introduction of a claim by the indefinite article "a" or "an" limits any particular claim containing such introduced claim content to only one embodiment containing such content (for example, "a" and / or "an" should be interpreted as meaning "at least one" or "one or more"). The same applies to the use of definite articles used to introduce claim content. In addition, even if a specific number relating to the introduced claim content is explicitly stated, a person skilled in the art will recognize that such a statement should be interpreted as meaning at least the number stated (for example, the mere statement "two statements" without other modifiers means at least two statements, or two or more statements).Furthermore, when a similar expression is used, such a configuration is generally intended to mean that a person skilled in the art will understand the expression (for example, "a system having at least one of A, B, and C" includes, but is not limited to, a system having only A, only B, only C, A and B together, A and C together, B and C together, and / or a system having A, B, and C together). It will be further understood by those skilled in the art that any virtually any disjunct word and / or disjunct phrase presenting two or more alternative terms should be understood as potentially including one of the terms, either or both of the terms in the specification, claims, or drawings. For example, the phrase “A or B” will be understood as including the possibilities of “A” or “B” or “A and B.” Furthermore, as used herein, the term “any of” followed by an enumeration of multiple items and / or categories of multiple items is intended to include, individually or in combination with other items and / or categories of items, “any of,” “any combination of,” “any multiple,” and / or “any combination of multiple.” Also, as used herein, the term “set” is intended to include any number of items, including zero. Furthermore, as used herein, the term “number” is intended to include any number, including zero. Also, as used herein, the term “multiple” is intended to be synonymous with “a plurality.”
[0247] In addition, if any feature or aspect of the present disclosure is described in terms of the Markush Group, a person skilled in the art will recognize that the present disclosure also describes any individual member or subgroup of a member of the Markush Group.
[0248] As will be understood by those skilled in the art, for any purpose, including written descriptions, all scopes disclosed herein include all possible subscopes and combinations thereof. Any named scope can be readily recognized as sufficiently explaining and enabling that the same scope may be divided into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each scope described herein can be readily divided into a lower third, a middle third, an upper third, etc. As will also be understood by those skilled in the art, all phrases such as “up to,” “at least,” “greater than,” and “less than” include the named number and refer to the scope that may subsequently be divided into subscopes as described above. Finally, as will be understood by those skilled in the art, a scope includes its individual members. Thus, for example, a group having 1 to 3 cells refers to a group having 1, 2, or 3 cells. Similarly, a group having 1 to 5 cells refers to a group having 1, 2, 3, 4, or 5 cells, and so on.
[0249] Furthermore, unless otherwise stated, the claims should not be interpreted as being limited to the order or elements provided. In addition, the use of the term “means for” in any claim is intended to exercise § 112(6) of the U.S. Patent Act, or the means-plus-function claim format, and any claim without the term “means for” is not intended to do so.
Claims
1. A wireless transceiver unit (WTRU), The application provider receives information indicating a set of artificial intelligence (AI) models and a set of addresses associated with the said set of AI models. A request regarding the AI model is sent to the network entity running the application server, using the address associated with the AI model from the set of addresses. Information indicating the AI model content corresponding to the AI model is received from the second network entity. The results are obtained using the AI model content received as described above. Processor, memory, and transceiver configured in such a way A WTRU characterized by having the following features.
2. The processor, memory, and transceiver are configured to establish a session with the network entity. The WTRU according to claim 1, characterized in that the AI model content is received from the network entity via the established session, and the AI model content corresponds to the full model of the AI model.
3. The processor, memory, and transceiver are configured, in the full model, to constitute an inference engine executed by the WTRU. The WTRU according to claim 2, characterized in that the above results are obtained from the configured inference engine.
4. The processor, memory, and transceiver are configured to establish a session with the network entity. The WTRU according to claim 1, wherein the AI model content is received from the network entity via the established session, and the AI model content corresponds to one or more subsets of the AI model.
5. The processor, memory, and transceiver are, Each subset of the AI model constitutes an inference engine executed by the WTRU, Obtain one or more intermediate results from the configured inference engine. The WTRU according to claim 4, characterized in that it is configured in such a way.
6. The WTRU according to claim 5, characterized in that the processor, memory, and transceiver are configured to provide each of the one or more intermediate results to an application executed by the WTRU.
7. The WTRU according to any one of claims 1 to 6, characterized in that the processor, memory, and transceiver are configured to provide the results to an application executed by the WTRU.
8. The WTRU according to any one of claims 1 to 7, characterized in that the processor, memory, and transceiver are configured to receive a selection of the AI model from the set of AI models.
9. The WTRU according to any one of claims 1 to 8, wherein the processor, memory, and transceiver are configured to receive decryption information associated with the AI model from the application provider and to decrypt the information indicating the content of the AI model using the decryption information.
10. The WTRU according to any one of claims 1 to 9, wherein the AI model content includes a plurality of dynamic adaptive streaming (DASH) segments via a hypertext transfer protocol, and the processor, memory, and transceiver are configured to aggregate the DASH segments to acquire the AI model.
11. The WTRU according to any one of claims 1 to 9, wherein the AI model content includes one or more AI model files, and the processor, memory, and transceiver are configured to receive each AI model file as a plurality of sequentially executable subsets.
12. The processor, memory, and transceiver are, Information indicating the AI model from the set of AI models is sent to another network entity that executes the application function. The other network entity receives information indicating the data encapsulation and / or compression format associated with the AI model. The other network entity receives the information indicating the AI model content corresponding to the AI model, based on the data encapsulation and / or compression format. The WTRU according to any one of claims 1 to 11, characterized in that it is configured as follows.
13. The WTRU according to any one of claims 1 to 12, characterized in that the processor, memory, and transceiver are configured to receive the information indicating the AI model content as a bitstream.
14. The WTRU according to any one of claims 1 to 13, characterized in that the address is a uniform resource locator (URL) associated with each of the sets of AI models.
15. A method implemented by a wireless transceiver unit (WTRU), Receiving information from an application provider indicating a set of artificial intelligence (AI) models and a set of addresses associated with the said set of AI models, A request regarding the AI model is sent to the network entity running the application server, using the address associated with the AI model from the set of addresses. Receiving information from the network entity indicating the AI model content corresponding to the AI model, based on the data encapsulation and / or compression format, Obtaining results using the AI model content received as described above. A method characterized by comprising:
16. To establish a session with the aforementioned network entity. Furthermore, The method according to 15, characterized in that the AI model content is received from the network entity via the established session, and the AI model content corresponds to the full model of the AI model.
17. The full model constitutes an inference engine executed by the WTRU. Furthermore, The method according to 16, characterized in that the above results are obtained from the configured inference engine.
18. To establish a session with the aforementioned network entity. Furthermore, The method according to 15, characterized in that the AI model content is received from the network entity via the established session, and the AI model content corresponds to one or more subsets of the AI model.
19. Each subset of the aforementioned AI model constitutes an inference engine executed by the WTRU, To obtain one or more intermediate results from the configured inference engine. The method according to 18, further comprising:
20. To provide the application executed by the WTRU with each of the one or more intermediate results. The method according to 19, further comprising:
21. To provide the results to the application executed by the WTRU The method according to any one of claims 15 to 20, further comprising:
22. Receiving a selection of an AI model from the aforementioned set of AI models. The method according to any one of claims 15 to 21, further comprising:
23. The application provider receives decryption information associated with the AI model from the application provider, and uses the decryption information to decrypt the information indicating the content of the AI model. The method according to any one of claims 15 to 22, further comprising:
24. The AI model is obtained by aggregating multiple Dynamic Adaptive Streaming (DASH) segments via a hypertext transfer protocol. The method according to any one of claims 15 to 23, further comprising:
25. The method according to any one of claims 15 to 23, characterized in that the AI model content includes one or more AI model files, and each AI model file is received as a plurality of sequentially executable subsets.
26. Sending information indicating the AI model from the set of AI models to another network entity that executes the application function, Receiving information from the other network entity indicating the data encapsulation and / or compression format associated with the AI model. Furthermore, The method according to any one of claims 15 to 25, characterized in that the information indicating the AI model content corresponding to the AI model is received based on the data encapsulation and / or compression format.
27. The method according to any one of claims 15 to 24, characterized in that the information indicating the AI model content is received as a bitstream.
28. The method according to any one of claims 1 to 15 to 25, characterized in that the address is a uniform resource locator (URL) associated with each of the sets of AI models.
Citation Information
Patent Citations
Optical deflector
JP1989002021A