Network-aware streaming adaptation

A network-aware streaming adaptation system using AI models to predict bandwidth and adjust video delivery policies addresses the instability of 360-degree video streaming over cellular networks, enhancing efficiency and playback quality.

WO2026040028A1PCT designated stage Publication Date: 2026-02-26ALCATEL LUCENT SHANGHAI BELL CO LTD +2
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/113789
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-21
Publication Date
2026-02-26

AI Technical Summary

Technical Problem

The challenge of streaming high-resolution 360-degree videos over unstable cellular networks, such as 5G, is exacerbated by fluctuations in signal strength and network congestion, leading to potential packet loss, visual artifacts, and degraded playback quality.

Method used

Implementing a network-aware streaming adaptation system that determines network throughput metrics using AI models to predict bandwidth fluctuations and adjust video delivery policies, including viewport-dependent delivery and quality control, to optimize video streaming based on real-time network conditions.

Benefits of technology

This approach enhances video delivery efficiency and resource utilization by dynamically adapting streaming quality and frame rate to match available network capacity, reducing latency and improving playback stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024113789_26022026_PF_FP_ABST
    Figure CN2024113789_26022026_PF_FP_ABST
Patent Text Reader

Abstract

Example embodiments of the present disclosure are directed to network-aware streaming adaptation. A method comprises: determining a characteristic metric of a wireless communication of the apparatus; determining a metric related to a network throughput of the apparatus available for video delivery, based on the determined characteristic metric of the wireless communication; determining a control policy for the video delivery, based on the determined metric related to the available network throughput; and delivering a video or partial video images of the video based on the control policy, using the wireless communication.
Need to check novelty before this filing date? Find Prior Art

Description

NETWORK-AWARE STREAMING ADAPTATION

[0001] FIELDS

[0002] Various example embodiments of the present disclosure generally relate to the field of wireless communication and in particular, to methods, devices, apparatuses and computer readable storage medium for network-aware streaming adaptation.BACKGROUND

[0003] To capture a panoramic video or images, for example full 360-degree images, panoramic cameras, for example, 360-degree cameras (also known as omnidirectional cameras) may utilize 2, 3, 4, or more camera sensors to simultaneously capture images from surroundings, and then stitch these images together to form a uniform panoramic, for example, 360-degree image and video. To stream the image or video over a network, a (built-in) communication module, e.g., for a fourth generation (4G) or the fifth generation (5G) cellular radio network may be used. Quality of video streaming, however, depends on connection stability and rate limitations of the provided communication network. It may depend on signal strength and signal stability of the wireless communication network, e.g., a cellular network. The signal strength may vary depending on locations, network congestion, and environmental factors, which may lead to potential fluctuations in connection quality.SUMMARY

[0004] In a first aspect of the present disclosure, there is provided a apparatus. The apparatus comprises at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: determine a characteristic metric of a wireless communication of the apparatus; determine a metric related to a network throughput of the apparatus available for video delivery, based on the determined characteristic metric of the wireless communication; determine a control policy for the video delivery, based on the determined metric related to the available network throughput; and deliver a video or partial video images of the video based on the control policy, using the wireless communication.

[0005] In a second aspect of the present disclosure, there is provided a method. The  method comprises: determining a characteristic metric of a wireless communication of the apparatus; determining a metric related to a network throughput of the apparatus available for video delivery, based on the determined characteristic metric of the wireless communication; determining a control policy for the video delivery, based on the determined metric related to the available network throughput; and delivering a video or partial video images of the video based on the control policy, using the wireless communication.

[0006] In a third aspect of the present disclosure, there is provided a apparatus. The apparatus comprises means for determining a characteristic metric of a wireless communication of the apparatus; means for determining a metric related to a network throughput of the apparatus available for video delivery, based on the determined characteristic metric of the wireless communication; means for determining a control policy for the video delivery, based on the determined metric related to the available network throughput; and means for delivering a video or partial video images of the video based on the control policy, using the wireless communication.

[0007] In a fourth aspect of the present disclosure, there is provided a computer readable medium. The computer readable medium comprises instructions stored thereon for causing an apparatus to perform at least the method according to the second aspect.

[0008] It is to be understood that the Summary section is not intended to identify key or essential features of embodiments of the present disclosure, nor is it intended to be used to limit the scope of the present disclosure. Other features of the present disclosure will become easily comprehensible through the following description.BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Some example embodiments will now be described with reference to the accompanying drawings, where:

[0010] FIG. 1A and FIG. 1B illustrate example block diagrams of a camera device according to some example embodiments;

[0011] FIG. 2 illustrates a flowchart of an example streaming adaptation method according to some example embodiments;

[0012] FIG. 3A to 3C illustrate example prediction processes of an Artificial Intelligence (AI) model according to some example embodiments;

[0013] FIG. 4 illustrates an example process of three-dimension (3D) sphere to two-dimension (2D) viewport picture projection according to some example embodiments;

[0014] FIG. 5A shows an example process of apply different quality of service (QoS) strategies depending on priority configuration according to some example embodiments;

[0015] FIG. 5B shows an example process of quality adaptation on real-time 360-degree stitching ranking according to some example embodiments;

[0016] FIG. 6A illustrates an example case of viewport-dependent delivery (VDD) according to some example embodiments;

[0017] FIG. 6B illustrates pre-defined viewports to cover the full 360-degree degree space according to some example embodiments;

[0018] FIG. 7A illustrates an example extension of a viewport image according to some example embodiments;

[0019] FIG. 7B illustrates an example quality adaptation process according to some example embodiments;

[0020] FIG. 8A illustrates a flow of an example strategy selection process according to some example embodiments;

[0021] FIG. 8B illustrates a flow of another example strategy selection process according to some example embodiments;

[0022] FIG. 9 illustrates an example of a one-byte header extension 900 according to some example embodiments;

[0023] FIG. 10 illustrates a simplified block diagram of a device that is suitable for implementing example embodiments of the present disclosure; and

[0024] FIG. 11 illustrates a block diagram of an example computer readable medium in accordance with some example embodiments of the present disclosure.

[0025] Throughout the drawings, the same or similar reference numerals represent the same or similar element.DETAILED DESCRIPTION

[0026] Principle of the present disclosure will now be described with reference to some  example embodiments. It is to be understood that these embodiments are described only for the purpose of illustration and help those skilled in the art to understand and implement the present disclosure, without suggesting any limitation as to the scope of the disclosure. Embodiments described herein can be implemented in various manners other than the ones described below.

[0027] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.

[0028] References in the present disclosure to “one embodiment, ” “an embodiment, ” “an example embodiment, ” and the like indicate that the embodiment described may comprise a particular feature, structure, or characteristic, but it is not necessary that every embodiment comprises the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is submitted that it is within the knowledge of one skilled in the art to affect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.

[0029] It is to be understood that although the terms “first, ” “second” and the like may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and similarly, a second element could be termed a first element, without departing from the scope of example embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.

[0030] As used herein, “at least one of the following: <a list of two or more elements>” and “at least one of <a list of two or more elements>” and similar wording, where the list of two or more elements are joined by “and” or “or” , mean at least any one of the elements, or at least any two or more of the elements, or at least all the elements.

[0031] As used herein, unless stated explicitly, performing a step “in response to A” does not indicate that the step is performed immediately after “A” occurs and one or more intervening steps may be included.

[0032] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments. As used herein, the singular forms “a” , “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” , “comprising” , “has” , “having” , “includes” and / or “including” , when used herein, specify the presence of stated features, elements, and / or components etc., but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof.

[0033] As used in this application, the term “circuitry” may refer to one or more or all of the following:

[0034] (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry) and

[0035] (b) combinations of hardware circuits and software, such as (as applicable) :

[0036] (i) a combination of analog and / or digital hardware circuit (s) with software / firmware and

[0037] (ii) any portions of hardware processor (s) with software (including digital signal processor (s) ) , software, and memory (ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions) and

[0038] (c) hardware circuit (s) and or processor (s) , such as a microprocessor (s) or a portion of a microprocessor (s) , that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.

[0039] This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit (IC) or processor integrated circuit (IC) for a mobile device or a similar integrated circuit (IC) in server, a cellular network device, or other computing or network device.

[0040] As used herein, the term “wireless communication network” refers to a network following any suitable communication standards and / or protocols, such as New Radio (NR) , Long Term Evolution (LTE) , LTE-Advanced (LTE-A) , Wideband Code Division Multiple Access (WCDMA) , High-Speed Packet Access (HSPA) , Narrow Band Internet of Things (NB-IoT) , WLAN Wireless Local Area Network (WLAN) or Wi-Fi, and so on, and any of their further generations, or any combination thereof. Furthermore, the communications between a terminal device and a network device in the communication network may be performed according to any suitable generation of telecommunication protocols, including, but not limited to, the first generation (1G) , the second generation (2G) , 2.5G, 2.75G, the third generation (3G) , the fourth generation (4G) , 4.5G, the fifth generation (5G) , the sixth generation (6G) , any further generation telecommunication protocols, and / or any other communication protocols either currently known or to be developed in the future. Embodiments of the present disclosure may be applied in various communication systems. Given the rapid development in communications, there will of course also be future type communication technologies and systems with which the present disclosure may be embodied. It should not be seen as limiting the scope of the present disclosure to only the aforementioned system.

[0041] As used herein, the term “network device” refers to a node in a communication network via which a terminal device accesses the network and receives services therefrom. The network device may refer to a base station (BS) or an access point (AP) , for example, a node B (NodeB or NB) , an evolved NodeB (eNodeB or eNB) , an NR NB (also referred to as a gNB) , a Remote Radio Unit (RRU) , a radio header (RH) , a remote radio head (RRH) , a relay, an Integrated Access and Backhaul (IAB) node, a low power node such as a femto, a pico, a non-terrestrial network (NTN) or non-ground network device such as a satellite network device, a low earth orbit (LEO) satellite and a geosynchronous earth orbit (GEO) satellite, an aircraft network device, and so forth, depending on the applied terminology and technology. In some example embodiments, radio access network (RAN) split architecture comprises a Centralized Unit (CU) and a Distributed Unit (DU) at an IAB donor node. An IAB node comprises a Mobile Terminal (IAB-MT) part that behaves like a UE toward the parent node, and a DU part of an IAB node behaves like a base station toward the next-hop IAB node.

[0042] The term “terminal device” refers to any end device that may be capable of wireless communication. By way of example rather than limitation, a terminal device may  also be referred to as a communication device, user equipment (UE) , a Subscriber Station (SS) , a Portable Subscriber Station, a Mobile Station (MS) , or an Access Terminal (AT) . The terminal device 100 in FIG. 1A may comprise, but not limited to, a camera device, a video camera device, a video surveillance camera, a MR or AR (Mixed Reality or Augmented Reality) camera device, a video conference device, a mobile phone, a cellular phone, a smart phone, voice over IP (VoIP) phones, wireless local loop phones, a tablet, a wearable terminal device, a personal digital assistant (PDA) , a mobile communication device, portable computers, desktop computer, laptop computer, image capture terminal devices such as digital cameras, gaming terminal devices, music storage and playback appliances, vehicle-mounted wireless terminal devices, a vehicle with one or more camera devices, wireless endpoints, mobile stations, laptop-embedded equipment (LEE) , laptop-mounted equipment (LME) , USB dongles, smart devices, wireless customer-premises equipment (CPE) , an Internet of Things (IoT) device, a watch or other wearable, a head-mounted display (HMD) , a television (TV) , a set-top box, a display, a vehicle, an infotainment unit, a drone, a medical device and applications (e.g., remote surgery) , an industrial device and applications (e.g., a robot and / or other wireless devices operating in an industrial and / or an automated processing chain contexts) , a consumer electronics device, a device operating on commercial and / or industrial wireless networks, and the like, or any combination thereof. The terminal device may also correspond to a Mobile Termination (MT) part of an IAB node (e.g., a relay node) . In the following description, the terms “terminal device” , “communication device” , “terminal” , “user equipment” and “UE” may be used interchangeably.

[0043] As described above, streaming of the image or video over a 4G / 5G / 6G cellular radio network may depend on signal strength and signal stability of the cellular network. However, the signal strength may vary depending on locations, network congestion, and environmental factors, which may lead to potential fluctuations in connection quality. This situation is getting worse for 360-degree cameras, as a resolution of a 360-degree video is much higher than videos generated by traditional two-dimension (2D) cameras. Nowadays, 360-degree cameras can produce 4K and even 8K live videos in an equirectangular projection (ERP) format, with video encoding in high bit rates up to 100 or 150Mbps.

[0044] The video resolution is expected to further increase to satisfy Extended Reality (XR) , such as Virtual Reality (VR) , graphical accuracy requirements, while keeping the  frames per second (fps) performance at a reasonable level (at least 30fps) . The real-time of 360-degree video streaming with a high resolution (12K, 16K, etc. ) requires solutions to deal with the network stability and rate limitation.

[0045] An open issue to be overcome by these solutions is to enable effective 360-degree video streaming over an unstable radio network with a built-in cellular communication module such as 5G integrated customer premise equipment (CPE) for 5G non-standalone (NSA)  / standalone (SA) . Compared with traditional videos, a difference is that adaptive video streaming techniques are required in 360-degree videos with different QoS criteria. Without knowledge of the network condition, the 360-degree video streaming may become much unstable and have a higher packet loss rate with a downgraded video playback quality such as visual artifacts, stuttering, and even long pausing (e.g., higher end-to-end (e2e) latency) .

[0046] Example embodiments of the present disclosure propose a solution for streaming adaptation based on determination (such as estimation, prediction or measurement ) of a network throughput available for video delivery. In this solution, based on a determined characteristic metric of a wireless communication of an apparatus, the apparatus determines a metric related to a network throughput available for video delivery. Based on the determined metric related to the available network throughput, the apparatus determines a control policy for the video delivery by means of the wireless communication.

[0047] This solution allows dynamic adaptation of video streaming based on the network condition. In this way, efficiency of the video delivery may be improved, and network resource utilization may be increased. The following description for the 360-degree image and video process and manipulation is also suitable for less than 360-degree image or video presentation.

[0048] FIG. 1A illustrates an example block diagram 100A of a terminal device 100, for example, a camera device 100, or any wireless communication device with one or more cameras, according to some example embodiments. The camera device 100 may be a 360-degree camera device, which comprises a mobile communication module 102, for example, having a 4G / 5G modem applied in a 4G / 5G network. In some example embodiments, the camera device 100 may act as a terminal device, such as a UE, with uplink traffic (also called an uplink (UL) device) . In some other example embodiments, the camera device 100 may act as a network device such as a base station with downlink  traffic.

[0049] The uplink device may be responsible for transmitting uplink data from the UE to a network. The uplink device can be any device capable of transmitting the uplink data, such as a smartphone, a modem, or an IoT device equipped with cellular connectivity. The uplink device communicates with a network infrastructure, which comprises one or more network devices such as base stations, evolved NodeBs (eNodeBs) , or Next-Generation NodeBs (gNodeBs) in 5G networks.

[0050] The camera device 100 further may comprise a plurality of camera sensors (also called camera lenses) 104-1, …, 104-N (individually or collectively referred to as camera sensor (s) 104) which may capture images from the surroundings. N is a positive integer. Some or all of these camera sensors 104-1, …, 104-N may focus on different viewing directions. The camera device 100 may obtain (e.g., block 106) and stitch (e.g., at a stitching module 108) the images from the camera sensors 104 into a video for delivery. It is to be understood that the camera sensors 104-1, …, 104-N are shown to be included in the camera device 100 only for the purpose of illustration without suggesting any limitation. In some example embodiments, one or more camera sensors may be deployed outside of the camera device 100 in the surroundings. The camera device 100 may exchange data with the camera sensors remotely.

[0051] Furthermore, the camera device 100 may comprise a network-aware quality video control application (or network-aware quality control application, also called a VCA) 110. In some example embodiments, the VCA 110 may perform streaming adaptation based on a network bandwidth fluctuation with a bandwidth (or throughput) prediction based on machine leaning.

[0052] Some example components of the VCA 110 are shown in FIG. 1B which shows another example block diagram 100B of the camera device 100 according to some example embodiments. As shown in FIG. 1B, a traffic adaption control loop may be organized by 3 modules: the wireless communication module 102, the VCA (model) 110, and a viewport control, graphics pipeline, video encoding and packetize and transmit module 112. The network related metrics e.g. a modulation and coding scheme (MCS) , a block error rate (BLER) , a buffer size, a received signal strength indicator (RSSI) , a reference signal receiving power (RSRP) , a reference signal receiving quality (RSRQ) , and / or the like measured by the 4G / 5G modem 102 are sent to the network-aware quality  control application (VCA) 110.

[0053] A viewport is a portion of a 360-degree video that a viewer can see. A viewport may be defined by its size and orientation or direction. Its size may be the same as a Field of View (FoV) that may be characterized by width and height in pixel or angular degrees. A change or control of the viewport implies a change of a viewing window in terms of the FoV and / or the direction. In the context of this invention, a viewport may represent a video created from original 360 video with the size of a FoV.

[0054] The VCA comprises functions: an AI network-aware function (also referred to as an AI network-aware model or AI model 114) and a quality control application function (also referred to as a Quality Control application 116) . In some example embodiments, the AI model 114 may be implemented using a long short-term memory (LSTM) model. The AI network-aware model 114 may perform uplink (UL) bandwidth prediction and propose the predicated values to the quality control application 116. The quality control application 116 generates a control policy e.g. quality of service (QoS) encoding parameters and / or field of view (FOV) control and QoS parameters, for packaging according to the UL bandwidth values. Then, the quality control application 116 may output the control policy to the viewport control, graphics pipeline, video encoding and packetize and transmit module 112.

[0055] The VCA 110 may allow the following adaptation aspects, for example comprising Adaptation Aspect 1, Adaptation Aspect 2, Adaptation Aspect 3 and Adaptation Aspect 4. In Adaptation Aspect 1, real-time network uplink (UL) bandwidth prediction is provided. This prediction (also called an on-device prediction) is based on inference from a trained Artificial Intelligence (AI) network-aware model (or an AI model) 114. The AI model 114 may use any machine learning (ML) algorithms or architecture. The AI model 114 can be trained offline on a server or online in the device 100, depending on real needs and specific 4G / 5G network providers. Some data or metrics measured by the mobile communication module (e.g., the 4G / 5G modem) 102 may be used to train the AI model 114 and perform the inference. Such data and / or metrics may comprise real-time UL properties such as physical resource blocks (PRBs) , modulation and coding schemes (MCSs) , and a transmission (or transmit) power level, as well as some UL metrics such as block error rates (BLERs) . An output of the AI model 114 may comprise UL bandwidth prediction values that may indicate a predicted network condition in terms of rate control parameters for video streaming, such as suggested bit rates in next future  periods. The bandwidth prediction may be performed once or in several times periodically.

[0056] By using the bandwidth prediction, a predicted throughput may be directed to a video encoding bit rate. Traditional network-aware adaptive bit rate (ABR) streaming (which may be performed in an application layer) is based on feedback collected from a receiver, and bit rate changes happens gradually and slowly over time based on heuristic algorithms. Thus, fast wireless network channel varies cannot be tracked, which potentially leads service degradation. The bandwidth prediction may improve dynamic and timely streaming adaptation and thus improve video delivery efficiency.

[0057] In Adaptation Aspect 2, a picture quality of a full 360-degree video is adapted. In some example embodiments, a full 360-degree (omnidirectional) video may be delivered with the quality control application 116 that may provide quality-based rate control. The AI model 114 may communicate with the built-in communication module 102 to predict near future bandwidth changes. The quality control application 116 may instruct a video encoder 120 in FIG. 1A to increase or decrease an output quality of the stitching module 108 (which may be implemented by a stitching processor unit) in FIG. 1A for stitching the images from the multiple camera lenses 104 to a full 360-degree frame (e.g., in an ERP format) . For example, the video encoder 120 may apply encoding compression control, for example, by adapting values of a video quantization parameter (QP) .

[0058] In some example embodiments, to control the quality of the full 360-degree video, the overall quality or bit rate may be changed uniformly, depending on the network situation or condition. In some other example embodiments, a region-based approach may be used to control the quality to one or more active the camera lenses 104 dynamically, based on primary or active viewing directions, e.g., field of views (FoV) or viewports, from one or more terminal devices. A 360-degree video image may be generated by stitching in real-time the image from the multiple camera sensors (lenses) 104, e.g. from front lenses (covering a front space) and back lenses (covering a back space) or 4 lenses covering a 360-degree space. The stitching module 108 (for example, a customized stitching IC) may keep one region in the 360-degree video picture (or frame) formed by the images from one or more lenses to have higher quality than the rest of the frame (off-region) formed by the images from other lenses, depending on the network bandwidth prediction from the VCA 110. In an example, the off-region (less important region) quality may be corresponding to the highest QP.

[0059] In Adaptation Aspect 3, a fixed viewport-dependent delivery (VDD) quality is adapted. Viewport-dependent delivery in 360-degree video streaming refers to a technique used to optimize the delivery of video content by prioritizing the portion of the video that is currently being requested by the client, rather than streaming the entire 360-degree video at once. The camera device 100 may output a viewport-only video that may have dynamic bit rates with respect to the network condition. Some fixed viewports with logic names such as front, left, right, back, up, and down may be partially preset or customizable by end users.

[0060] The bit rate changes may be a result of multiple techniques. In an example, dynamic viewport quality adaptation may be implemented by controlling an encoding bit rate min-max range from predicted bandwidth throughput. In another example, dynamic viewport margin quality adaptation may be applied . For example, the quality adaptation may be applied to an extra viewport margin only. The viewport part may keep an original bit rate, but the bit rate of the margin area may be lowered. In yet another example, dynamic viewport margin size adaptation may be applied. For example, a margin size may vary according to an available throughput. In some cases, the whole viewport region may be treated as “margin” with a low quality (e.g., using a high quantization value) .

[0061] In Adaptation Aspect 4, bandwidth-aware adaptation profiles are used for frame rate and quality-based adaptation strategies. The adaptation may be applied to the video output frame rates together with the video quality. Based on a bandwidth (or throughput) estimation threshold value or value sets, for example, corresponding to a network congestion severity level, both the frame rate and quality-based adaptations may be combined. The adaptation may be defined as different profiles. Each profile may define adaptation priorities as a sequence to change the quality and / or frame rate accordingly.

[0062] For example, some orders could be defined as follows: quality priority for VDD streaming and frame rate adaptation only after the available threshold exceeds a certain throughput threshold; and frame rate priority for 360-degree streaming and quality adaptation after the available threshold exceeds a certain throughput threshold. In some example embodiment, the adaptation profiles or preferences may be signaled from receivers of the video, e.g., terminal devices, to suit different use cases, which means different adaptation strategies may be chosen to apply for individual streaming sessions or viewports.

[0063] Video bandwidth adaptation may be based on changes either frame rate changes or quality changes, depending on specific requirements and constraints of use cases or application scenarios, as well as available adaptation mechanisms. Some example implementations for streaming adaptation will be described below with reference to FIGS. 2 to 9.

[0064] Now, referring to FIG. 2, a flowchart of an example streaming adaptation method 200 is illustrated. The method 200 may be implemented at an apparatus 100 such as the camera device 100 in FIG. 1A. At block 210, the apparatus 100 determines a characteristic metric of a wireless communication of the apparatus 100. The characteristic metric of the wireless communication may comprise any metric of the characteristics of the wireless communication. In an example, the metric may comprise network metrics measured by the 4G / 5G / 6G modem 102.

[0065] In some example embodiments, the characteristic metric of the wireless communication may be related to scheduling information for the wireless communication of the apparatus 100. The uplink (UL) scheduling information may comprise, but not limited, one or more of a transport block (TB) size, a modulation and coding scheme (MCS) , a modulation type, transmit power control, or a retransmission version.

[0066] Alternatively, or in addition, the characteristic metric of the wireless communication may be related to a quality of the wireless communication. In an example, the quality of the wireless communication may be indicated by a transmission result such as a positive acknowledgement (ACK) or a negative acknowledgement (NACK) . Alternatively, or in addition, the quality of the wireless communication may be indicated by one or more of a radio link control (RLC) status such as a protocol data unit (PDU) retransmission rate, a PDU NACK rate, a mean service data unit (SDU) throughput, a mean SDU latency, or a buffer occupancy. Alternatively, or in addition, the quality of the wireless communication may be indicated by one or more of a packet data convergence protocol (PDCP) status such as a SDU throughput, a SDU packet rate, or a lost PDU rate.

[0067] At block 220, based on the estimated characteristic metric of the wireless communication, the apparatus 100 determines a predicted metric related to a network throughput of the apparatus 100 available for video delivery, for example, for delivery of a 360-degree panoramic video. The predicted metric related to the network throughput of the apparatus 100 may comprise any metric related to the available network throughput  of the apparatus 100.

[0068] In some example embodiments, the predicted metric related to the available network throughput may comprise a predicted value of a maximum bit rate (referred to as a first maximum bit rate) available for the video delivery. Alternatively, or in addition, the predicted metric may comprise a ratio of the first maximum bit rate to a maximum bit rate (referred to as a second maximum bit rate) allocated to the apparatus 100. Alternatively, or in addition, the predicted metric may comprise a value of the available network throughput.

[0069] In some example embodiments, the predicted metric related to the available network throughput of the apparatus 100 may be determined using a machine learning model (for example, the AI model 114 in FIG. 1B) based on the characteristic metric of the wireless communication. As shown in FIG. 3A, in a process 300A, based on the input network related metrics, the AI network-aware model 114 may output a predicted bandwidth (e.g., a maximum bit rate) . In a 5G network, a UL bandwidth (e.g., the maximum bit rate) change as a radio channel condition varies. The AI Network-aware model 114 may forecast the channel condition and UL bandwidth in advance.

[0070] In some example embodiments, the apparatus 100 may determine predicted spectrum efficiency from the characteristic metric of the wireless communication and obtain a resource allocation scheme for the wireless communication. The resource allocation scheme may comprise a current resource allocation scheme and / or a predicted resource allocation scheme for the wireless communication. Based on the predicted spectrum efficiency and the resource allocation scheme, the apparatus 100 may determine the predicted metric related to the available network throughput.

[0071] In some example embodiments, different machine learning modules may be applied for the spectrum efficiency prediction and the bandwidth prediction. For example, the predicted spectrum efficiency may be determined using a machine learning model from the characteristic metric of the wireless communication, and the predicted metric related to the available network throughput is determined using a further machine learning module based on the predicted spectrum efficiency and the resource allocation scheme.

[0072] FIG. 3B show an example prediction process 300B according to some example embodiments. The process 300B may be a UL bandwidth prediction procedure implemented by the AI model 114. As shown in FIG. 3B, the process 300B comprises the  following steps. In step 1, a data input from the 4G / 5G / 6G modem 102 is processed to prepare and transform raw data of the one or more network metric e.g. MCSs, BLERs, RSSIs, RSSPs, RSSQs, or buffer status, etc. into a format that can be used for model training or make prediction.

[0073] In step 2, a prediction of UL spectrum efficiency is performed. The spectrum efficiency refers to bits that may be transmitted per Hz. In general, the UL spectrum efficiency of a UE becomes higher when the UE is close to a cell center and becomes lower when the UE is close to a cell edge. A Long Short-Term Memory (LSTM) network may be used for the spectrum efficiency prediction. The AI model 114 may be trained via large amount real data collected by the modem 102, or the data generated by simulations.

[0074] Step 3 is to apply one or more resource allocation algorithms, which calculate / estimate the resource (e.g., a physical resource block (PRB) number) that will be assigned by the network for the applications. In general, a network device such as gNB is responsible for scheduling and assign resource allocation to the application of a UE in UL, the UE or the modem has no gNB scheduling information beforehand. In this case, in an example, the apparatus 100 such as the UE 100 or model may use a current resource allocation scheme (e.g., the PRB number) when it performs the UL bandwidth prediction and assume the same amount resource will be assigned to the UE / modem by the gNB in near future. In another example, resource allocation may be calculated or estimated according to a time sequence of resources allocated in past several slots by using some algorithms.

[0075] In some example embodiments, the predicted metric related to the available network throughput of the apparatus 100 may be determined directly from the characteristic metric of the wireless communication. An example process in this regard will be described below with reference to FIG. 3C. A process 300C in FIG. 3C comprises the following steps. Step 1 in FIG. 3C is same as step 1 in FIG. 3B, which is to process the data input from the 4G / 5G / 6G modem 102, and transform the data into a format that can be used for model training or make prediction.

[0076] Step 2 in FIG. 3C is to predicate the UL bandwidth via the AI model 114 in a process different from the process 300B in FIG. 3B. In the process 300B, the UL bandwidth is predicted by 2 steps where in the first step, spectrum efficiency is predicated by the AI model 114; and in the second step, resource allocation is calculated / estimated  by the resource allocation algorithm. Then, the UL bandwidth output = spectrum efficiency × resource allocation. In process 300C, the AI model 114 may perform the UL bandwidth directly via the AI model 114. The AI model 114 may be trained considering both spectrum efficiency related data and resource allocation related data, both of which may be input from the 4G / 5G / 6G modem 102.

[0077] Still with reference to FIG. 2, at block 230, the apparatus 100 determines a control policy for the video delivery, based on the predicted metric related to the available network throughput. In some example embodiments, the apparatus 100 may determine a quality level of the video delivery based on the predicted metric related to the available network throughput; and then determine the control policy based on the quality ranking.

[0078] Table 1, called Quality Ranking Table, shows example mapping to convert the predicted network throughputs to a discrete quality measurement. In this example, the predicted linear throughputs are denoted by a ratio of the predicted bandwidth to the allocated or optimal bandwidth of the apparatus 100. This table may be used to control all the following quality adaptation in the apparatus 100 such as the camera device 100. It is to be understood that the conversion is one example, and different use cases and requirements from the apparatus 100 may have different weighting to affect the mapping from the predicted bandwidth (e.g., ratio) to the quality.

[0079] Table 1: Example of a Quality Ranking Table from predicted network throughput

[0080] In some example embodiments, the control policy may comprise an adaptation strategy for a streaming quality and / or a streaming rate of the video delivery. In some example embodiments, the adaptation strategy may be determined based on an adaptation priority for the streaming quality and / or the streaming rate. In some example embodiments, the apparatus 100 may receive an adaptation preference from a further apparatus 100 (e.g.,  the receiver of the video) to which the video is to be delivered. The apparatus 100 may determine the control policy based on the received adaptation preference.

[0081] At block 240, based on the control policy, the apparatus 100 delivers a video or partial video images of the video using the wireless communication. For example, a 360-degree camera may produce a video with a full-stitched frame. It may produce 4K (3840×1960 pixels) and even 8K (7680×3840 pixels) resolution ERP frames. In the context of a 360-degree video, "ERP" stands for "Equirectangular Projection. " Equirectangular projection is a mapping approach commonly used to represent a three-dimension (3D) spherical surface as a two-dimension (2D) image, suitable for displaying spherical content such as 360-degree videos or panoramas on flat screens. In an equirectangular projection, as shown in FIG. 4, a spherical surface 405 is projected onto a rectangle 410, where a horizontal axis represents an azimuth angle (which is horizontally rotated around the sphere) , and a vertical axis represents an altitude angle (which is vertically rotated around the sphere) . This results in an image where the entire 360-degree horizontal field of view and 180-degree vertical field of view are mapped onto a rectangular frame.

[0082] In some example embodiments, a quality of the video may be adapted based on the control policy. In some example embodiments, the video to be delivered may comprise a 360-degree panoramic video stitched from a plurality of image capturing devices such as cameras, lenses, and sensors. In an example, to adapt the quality of the video, a stitching quality and / or an encoding quality of the video may be adapted based on the control policy. In some example embodiments, the apparatus 100 may determine whether either or both of the stitching quality and the encoding quality of the video are to be adapted, based on an adaptation priority for at least one of the stitching quality or the encoding quality.

[0083] If the stitching quality of the video is adapted, the apparatus 100 may select and change image capturing devices (such as the camera sensors 104) that are participated in the stitching. For example, the apparatus 100 may select one or more image capturing devices from a plurality of image capturing devices based on the control policy, and then stitching one or more images from the one or more image capturing devices.

[0084] FIG. 5A shows an example process 500 of apply different QoS strategies depending on priority configuration according to some example embodiments. After starting the 360 live streaming at block 501, the process 500 estimates radio network  bandwidth (bitrate in uplink) in the next time period at block 502. As shown, if a stitching quality priority is selected as “QOS STRATEGY PREFERENCE” at block 503 (aleft flow 505) , the network bandwidth may be converted to quality ranking preset (which may be corresponding to a range of the stitching quality, e.g. from high to low, and a special single lens only, as shown in FIG. 5B) at block 506. Further, at block 507, the quality raking value may affect the output quality of the stitched 360-degree raw frame before video encoding at block 508.

[0085] Depending on the available uplink bandwidth, the VCA 110 may determine which stitching quality is used by the 360-degree image stitching component (e.g., the stitching module 108 FIG. 1A) . As shown in FIG. 5B, in a lowest ranking mode 511, the stitched 360-degree raw frame contains the input from a single sensor (e.g., a front senor 512) and the rest (which contains the inputs from other sensors, such as a left sensor 514, a back sensor 516, and a right sensor 518) of the 360-degree ERP picture is black.

[0086] It is to be noted that the quality change may be the following two cases: 1) a sensor (lens) is configurable to different quality output; and 2) the sensor output is fixed in terms of the quality. In case 2) , an extra imaging processing step may be needed to degrade the quality. Any techniques may be applied here. For instance, downsampling may be used to reduce the number of pixels in an image by discarding pixels or averaging pixel values.

[0087] If the QoS strategy is prioritized by the encoding quality (aright flow 509 in FIG. 5A) , at block 510, the process 500 applies a predicted target bitrate to the video encoder to increase or decrease the encoding bitrate and the update can happen immediately or scheduled to the next video keyframe time. For example, the VCA 110 may pass the estimated encoding bit rates from the bandwidth bit rate. If the estimated bit rate is not presenting directly the encoding bit rate, the VCA 110 may convert the network UL bandwidth bit rate to video encoding bit rate proportionally. Then, at block 520, the process 500 performs video packaging for network delivery.

[0088] In some example embodiments, the video or the partial video images of the video to be delivered may comprise a limited viewport video. In some example embodiments, one or more viewports of the video may be delivered based on the control policy. For example, a raw 8K frame (7680×3840) with 3 or 4 color channels (RGB or RGBA) may occupy a significant memory space, no matter which image format or color space it uses,  e.g. YUV planar or packed RGB format. A viewport-only output may thereafter be provided to produce one part of the full 360-degree space only. The viewport-only output means that the outputted video contains viewport frames focusing on one or more viewing directions, instead of 360-degree frames.

[0089] FIG. 6A shows an example case of viewport-dependent delivery (VDD) according to some example embodiments. A VDD mode is a commonly used delivery mode in 360-degree streaming. Both fixed-viewports and dynamic viewports may be allowed. Fixed viewports mean the number of viewport outputs and resolutions with respect to the field of view (FoV) size is pre-configured and rather fixed. Dynamic viewports mean that the viewport resolution and viewport directions are variable.

[0090] In a camera device 600, the viewport generation may be done by a Viewport Control Driver SW (VCDSW) 610. In an embodiment, the quality adaptation is to deal with the full 360-degree frames. After the uplink prediction is received from the Network-aware Quality Video Control Application (VCA) 110, different adaptation strategies may be applied based on the 2 different delivery modes: 1) original 360-degree video delivery and 2) viewport-dependent delivery adaptation. If mode 2) is applied, the delivered video may comprise partial video images and may also be referred to as a viewport video or a limited viewport video. In some example embodiments, the VCA 110 may maintain a bandwidth-to-quality ranking table (e.g., Table 1) for throttling the final video throughput within the expected bandwidth.

[0091] For instance, a 90° × 60° out of 360° × 180° viewport may save up to 12 times amount of data. A 360-degree camera may have a number of viewports predefined. FIG. 6B illustrates 6 pre-defined fixed viewports 620-1, …, 620-6 to cover the full 360-degree degree space according to some example embodiments. Each viewport may have the same fixed resolution. Moreover, each viewport has a bigger margin than 90 × 60. The dimension of the margin can be zero, which means that the viewport output is exactly as same as the viewport size in terms of its viewing width and height in an angular degree (between 0-360 degree horizontal and 0-90 vertical space) .

[0092] Table 2 shows stitched viewport-only resolution for 360-degree cameras with 2 resolutions of 8K and 16K.

[0093] Table 2

[0094] The viewports may need to be projected first to a unit sphere from the 2D plane and rotated around the sphere and re-projected back to the 2D plane. As shown in FIG. 4 which illustrates the sphere and plane projection with theta and phi degrees, with a given 3D orientation (x, y, z) in a 3D coordinate system, the x / y / z vec3 can be converted to the theta and phi degree and then to the 2D plane; and vice versa, from a 2D x / y offset in the 2D plane to a 3D position (x / y / z) in the sphere. This ensures that each viewport is at the center (0 degree longitude and 0 degree latitude, given the degree ranges of -180-180 degree horizontal and -90-90 degree vertical) .

[0095] In some example embodiments, a margin size of a viewport of the one or more viewports may be adapted based on the control policy. FIG. 7A shows an example extension of a viewport image according to some example embodiments. The viewport image contains extra pixels in all edges named “margin” . The margin size can be fixed or dynamic as well, depending on the available bandwidth.

[0096] In some example embodiments, qualities of pixels of a viewport of the one or more viewports are adapted in a pattern based on the control policy. FIG. 7B illustrates  another embodiment of the quality adaptation according to some example embodiments. The image quality is changed gradually, based on the bandwidth-to-quality ranking (e.g., Table 1) . The brightness in FIG. 7B represents different quality levels. The white means the original quality; and the dark means downgraded pixels. The downgraded quality can be in different patterns, like centered radial, or linear (high quality vertically) .

[0097] In some example embodiments, an encoding bit rate of the video may be adapted based on the control policy. The encoding bit rate of the video may be less than or equal to a maximum bit rate available for the video.

[0098] In addition to or instead of the quality of the video, a frame rate of the video may be adapted based on the control policy. For example, in addition to the adaptation based on the visual quality of individual video frames, another effective way to control the data rate is to change the video frame rate with or without the quality adaptation. In general, a smooth video playback requires 30 frames per second (FPS) . A low FPS video may result in stuttering video playback and bad user experience, without sacrificing any video quality. It is a trade-off when designing an adaptation strategy. Different use cases need different strategies.

[0099] FIG. 8A shows a flow of an example strategy selection process 800A according to some example embodiments. In this example, the adaptation may be fixed on a pre-defined or pre-selected strategy. After starting the 360 live streaming at block 801, the process 800A estimates radio network bandwidth (bitrate in uplink) in the next time period at block 802. If it is determined at block 803 that a smooth video is preferred, the process 800A proceeds to a left branch 805 where a predicted target bitrate is applied to the video quality-based adaptation, but the original output frame rate is maintained. If it is determined at block 803 that a high-quality video is preferred, the process 800A proceeds to a left branch 805 where the predicted target bitrate is applied to a video framerate so that the output bitrate can meet the expected bandwidth bitrate. Then, at block 806, the process 800A performs video packaging for network delivery.

[0100] FIG. 8B shows another an example strategy selection process 800B according to some example embodiments where strategy selection is more dynamic. It is based on the real-time bandwidth estimation and chose different strategies within a lifespan of a live delivery session. In this case, the VCA 110 in FIGS 1A and 1B may re-use the bandwidth-to-quality ranking table (e.g. Table 1) to plan the strategies. After starting the 360 live  streaming at block 811, the process 800B estimates radio network bandwidth (bitrate in uplink) in the next time period at block 812. At block 813, the process 800B determines the strategy based on the ranking table. The perceived quality adaptation varies over time and depends on the network congestion and available uplink bandwidth. When the network condition is poor and the quality ranking is low, the VCA 110 may choose the frame rate-adaptation at block 814 where a predicted target bitrate is applied to the video quality-based adaptation, but the original output frame rate is maintained. Likewise, the VCA 110 may switch to the quality-based adaptation, when the ranking is becoming better at block 815 where the predicted target bitrate is applied to a video framerate so that the output bitrate can meet the expected bandwidth bitrate. The process 800B also applies a hybrid strategy (block 820) in which both the frame rate adaptation and the quality adaptation may be applied when the network condition is extremely poor, e.g. the quality ranking is in the “low” level. Then, at block 821, the process 800B performs video packaging for network delivery.

[0101] Table 3 shows strategies based on a quality ranking table. It is to be noted that the number of quality ranking levels in Table 3 is only illustrative, but not limited. The levels may be classified into more fine-grained granularity, more than 3 levels, as illustrated in Table 3.

[0102] Table 3

[0103] In some example embodiments, the video may be delivered along with an indication whether the video comprises a viewport frame. In some example embodiments, the indication is carried in a header of the video, and / or a channel different from a channel for the delivering of the video.

[0104] In an embodiment, the output may switch from 360-degree video to viewport-only content. In this case, the output coded video may need to carry with content  identification metadata to indicate the content type, whether or not it is a full 360-degree ERP frame, or a viewport frame. The metadata may be coded using in-band approach like H264 / H265’s custom SEI NAL unit, or one-byte header in Real-time Transport Protocol (RTP) header extension, e.g., following RFC 8285. Alternatively, the metadata may be delivered in an out-of-band approach via control channels such as a Real-Time Control Protocol (RTCP) sender report or other channels to the receiver (s) .

[0105] FIG. 9 shows an example of a one-byte header extension 900 according to some example embodiments. The 4-bit “ID” is a local identifier of this element in the range 1 to 14 inclusive. The 4-bit “len” (length) is the number, minus one, of data bytes of this header extension element following the one-byte header. In the example header extension 900, the header extension defines one extension ( “length=1” ) 905, which length is 1 byte, to carry the content identification, e.g. 0 indicates a 360-degree video, 1 indicates a viewport video.

[0106] In some example embodiments, an apparatus (for example, the camera device 100 in FIG. 1A) may comprise means for performing any or any combination of the methods or processes 200, 300A, 300B, 300C, 500, 602, 800A or 800B. The means may be implemented in any suitable form. For example, the means may be implemented in a circuitry or least one processor and at least one software module with storing instructions, for example, computer program code, that is executed by the at least one circuitry or processor. The apparatus may be implemented as or included in the terminal device 100, for example, the camera device 100 in FIG. 1A.

[0107] In some example embodiments, the apparatus (for example, the camera device 100 in FIG. 1A) comprises means for determining a characteristic metric of a wireless communication of the apparatus; means for determining a predicted metric related to a network throughput of the apparatus available for video delivery, based on the estimated characteristic metric of the wireless communication; means for determining a control policy for the video delivery, based on the predicted metric related to the available network throughput; and means for delivering a video or partial video images of the video based on the control policy, using the wireless communication.

[0108] In some example embodiments, the characteristic metric of the wireless communication is related to at least one of: scheduling information for the wireless communication of the apparatus, or a quality of the wireless communication of the  apparatus.

[0109] In some example embodiments, the predicted metric related to the available network throughput comprises at least one of: a predicted value of a first maximum bit rate available for the video delivery, a ratio of the first maximum bit rate to a second maximum bit rate allocated to the apparatus, or a value of the available network throughput.

[0110] In some example embodiments, the predicted metric related to the available network throughput is determined using a machine learning (ML) model based on the characteristic metric of the wireless communication.

[0111] In some example embodiments, the means for determining a predicted metric comprises at least one of: means for determining predicted spectrum efficiency from the characteristic metric of the wireless communication, means for obtaining a resource allocation scheme for the wireless communication, or means for determining the predicted metric related to the available network throughput, based on the predicted spectrum efficiency and the resource allocation scheme.

[0112] In some example embodiments, the resource allocation scheme comprises at least one of a current resource allocation scheme or a predicted resource allocation scheme for the wireless communication.

[0113] In some example embodiments, the predicted spectrum efficiency is determined using a machine learning model from the characteristic metric of the wireless communication, and the predicted metric related to the available network throughput is determined using a further machine learning module based on the predicted spectrum efficiency and the resource allocation scheme.

[0114] In some example embodiments, the means for determining the control policy comprises: means for determining a quality level of the video delivery based on the predicted metric related to the available network throughput; and means for determining the control policy based on the quality level.

[0115] In some example embodiments, the control policy comprises an adaptation strategy for at least one of a streaming quality or a streaming rate of the video delivery.

[0116] In some example embodiments, the adaptation strategy is determined based on an adaptation priority for at least one of the streaming quality or the streaming rate.

[0117] In some example embodiments, the video is delivered to a further apparatus, and the apparatus further comprises: means for receiving an adaptation preference from the further apparatus, wherein the control policy is determined, at least partly, based on the received adaptation preference.

[0118] In some example embodiments, an encoding bit rate of the video is adapted based on the control policy.

[0119] In some example embodiments, the encoding bit rate of the video is less than or equal to a maximum bit rate available for the video.

[0120] In some example embodiments, the video to be delivered comprises a 360-degree panoramic video stitched from a plurality of image capturing devices. Further, in some other example embodiments, the video to be delivered comprises less than a 360-degree panoramic video stitched from a one or more of image capturing devices.

[0121] In some example embodiments, at least one of a stitching quality or an encoding quality of the video is adapted based on the control policy.

[0122] In some example embodiments, the apparatus further comprises: means for determining whether either or both of the stitching quality and the encoding quality of the video are to be adapted, based on an adaptation priority for at least one of the stitching quality or the encoding quality.

[0123] In some example embodiments, the stitching quality of the video is adapted, and the apparatus further comprises: means for selecting one or more image capturing devices from a plurality of image capturing devices based on the control policy; and means for stitching one or more images from the one or more image capturing devices.

[0124] In some example embodiments, the video or the partial video images of the video to be delivered comprises a limited viewport video.

[0125] In some example embodiments, one or more viewports of the video are delivered based on the control policy.

[0126] In some example embodiments, a margin size of a viewport of the one or more viewports is adapted based on the control policy.

[0127] In some example embodiments, qualities of pixels of a viewport of the one or more viewports are adapted in a pattern based on the control policy.

[0128] In some example embodiments, the video is delivered along with an indication whether the video comprises one or more viewport frames.

[0129] In some example embodiments, the indication, whether the video comprises one or more viewport frames, is carried in at least one of: a header of the video, or a channel different from a channel for the delivering of the video.

[0130] In some example embodiments, the video delivery comprises delivery of a 360-degree panoramic video. In some other example embodiments, the video delivery comprises delivery of a less than 360-degree panoramic video.

[0131] FIG. 10 is a simplified block diagram of a device 1000 that is suitable for implementing example embodiments of the present disclosure. The device 1000 may be provided to implement an apparatus, for example, the camera device 100 as shown in FIG. 1A. As shown in FIG. 10, the device 1000 comprises at least one or more processors 1010, one or more memories 1020 coupled to the one or more processors 1010, one or more communication modules 1040 coupled to the one or more processor 1010, one or more camera sensors 1050 coupled to the one or more processor 1010, and one or more input / output (I / O) means coupled to the one or more processor 1010.

[0132] The communication module 1040 is for wireless or wired bidirectional communications. The communication module 1040 has one or more communication interfaces to facilitate communication with one or more other modules or devices. The communication interfaces may represent any interface that is necessary for communication with other network elements. In some example embodiments, the communication module 1040 may comprise at least one antenna.

[0133] The processor 1010 may be of any type suitable to the local technical network and may comprise one or more of the following: general purpose computers, special purpose computers, microprocessors, central processing units (CPUs) , digital signal processors (DSPs) , graphical processing units (GPUs) , circuitries, or processors based on multicore processor architecture, as non-limiting examples, or any combination thereof. The device 1000 may have multiple processors, such as an application specific integrated circuit chip that is slaved in time to a clock which synchronizes the main processor.

[0134] The memory 1020 may comprise one or more non-volatile memories and one or more volatile memories. Examples of the non-volatile memories comprise, but are not  limited to, a Read Only Memory (ROM) 1024, an electrically programmable read only memory (EPROM) , a flash memory, a hard disk, a compact disc (CD) , a digital video disk (DVD) , an optical disk, a laser disk, or other magnetic storage and / or optical storage. Examples of the volatile memories comprise, but are not limited to, a random-access memory (RAM) 1022 or other volatile memories that will not last in the power-down duration.

[0135] One or more computer programs 1030 comprises computer executable instructions that are executed by the associated processor 1010. The instructions of the program 1030 may comprise instructions for performing operations / acts of example embodiments of the present disclosure. The program 1030 may be stored in the memory, e.g., the ROM 1024. The processor 1010 may perform any suitable actions and processing by loading the program 1030 into the RAM 1022.

[0136] The example embodiments of the present disclosure may be implemented by means of the program 1030 so that the device 1000 may perform any process of the disclosure as discussed with reference to FIG. 1A to FIG. 9. The example embodiments of the present disclosure may also be implemented by hardware or by a combination of software and hardware.

[0137] In some example embodiments, the program 1030 may be tangibly contained in a computer readable medium which may be comprised in the device 1000 (such as in the memory 1020) or other storage devices that are accessible by the device 1000. The device 1000 may load the program 1030 from the computer readable medium to the RAM 1022 for execution. In some example embodiments, the computer readable medium may comprise any types of non-transitory storage medium, such as ROM, EPROM, a flash memory, a hard disk, CD, DVD, and the like. The term “non-transitory, ” as used herein, is a limitation of the medium itself (i.e., tangible, not a signal) as opposed to a limitation on data storage persistency (e.g., RAM vs. ROM) .

[0138] FIG. 11 shows an example of the computer readable medium 1100 which may be in form of CD, DVD or other optical storage disk. The computer readable medium 1100 has the program 1030 stored thereon.

[0139] Generally, various embodiments of the present disclosure may be implemented in hardware or special purpose circuits, software, logic or any combination thereof. Some aspects may be implemented in hardware, and other aspects may be implemented in  firmware or software which may be executed by a controller, microprocessor or other computing device. Although various aspects of embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representations, it is to be understood that the block, apparatus, system, technique or method described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.

[0140] Some example embodiments of the present disclosure also provide at least one computer program product tangibly stored on a computer readable medium, such as a non-transitory computer readable medium. The computer program product comprises computer-executable instructions, such as those comprised in program modules, being executed in a device on a target physical or virtual processor, to carry out any of the methods as described above. Generally, program modules comprise routines, programs, libraries, objects, classes, components, data structures, or the like that perform particular tasks or implement particular abstract data types. The functionality of the program modules may be combined or split between program modules as desired in various embodiments. Machine-executable instructions for program modules may be executed within a local or distributed device. In a distributed device, program modules may be located in both local and remote storage media.

[0141] Program code for carrying out methods of the present disclosure may be written in any combination of one or more programming languages. The program code may be provided to a processor or controller of a general-purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program code, when executed by the processor or controller, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may execute entirely on a machine, partly on the machine, as a stand-alone software package, partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0142] In the context of the present disclosure, the computer program code or related data may be carried by any suitable carrier to enable the device, apparatus or processor to perform various processes and operations as described above. Examples of the carrier comprise a signal, computer readable medium, and the like.

[0143] The computer readable medium may be a computer readable signal medium or  a computer readable storage medium. A computer readable medium may comprise but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer readable storage medium would comprise an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random-access memory (RAM) , a read-only memory (ROM) , an erasable programmable read-only memory (EPROM or Flash memory) , an optical fiber, a portable compact disc read-only memory (CD-ROM) , an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0144] Further, although operations are depicted in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Likewise, although several specific implementation details are contained in the above discussions, these should not be construed as limitations on the scope of the present disclosure, but rather as descriptions of features that may be specific to particular embodiments. Unless explicitly stated, certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, unless explicitly stated, various features that are described in the context of a single embodiment may also be implemented in a plurality of embodiments separately or in any suitable sub-combination.

[0145] Although the present disclosure has been described in languages specific to structural features and / or methodological acts, it is to be understood that the present disclosure defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

Claims

1.A apparatus comprising:at least one processor; andat least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to:determine a characteristic metric of a wireless communication of the apparatus;determine a metric related to a network throughput of the apparatus available for video delivery, based on the determined characteristic metric of the wireless communication;determine a control policy for the video delivery, based on the determined metric related to the available network throughput; anddeliver a video or partial video images of the video based on the control policy, using the wireless communication.2.The apparatus of claim 1, wherein the characteristic metric of the wireless communication is related to at least one of:scheduling information for the wireless communication of the apparatus, ora quality of the wireless communication of the apparatus.3.The apparatus of claim 1 or 2, wherein the determined metric related to the available network throughput comprises at least one of:a predicted value of a first maximum bit rate available for the video delivery,a ratio of the first maximum bit rate to a second maximum bit rate allocated to the apparatus, ora value of the available network throughput.4.The apparatus of any of claims 1 to 3, wherein the determined metric related to the available network throughput is determined using a machine learning model based on  the characteristic metric of the wireless communication.5.The apparatus of any of claims 1 to 3, wherein the instructions that, when executed by the at least one processor, further cause the apparatus to:determine a spectrum efficiency from the characteristic metric of the wireless communication;obtain a resource allocation scheme for the wireless communication; anddetermine the predicted metric related to the available network throughput, based on the determined spectrum efficiency and the obtained resource allocation scheme.6.The apparatus of claim 5, wherein the resource allocation scheme comprises at least one of a current resource allocation scheme or a predicted resource allocation scheme for the wireless communication.7.The apparatus of claim 5 or 6, wherein the determined spectrum efficiency is determined using a machine learning model based on the characteristic metric of the wireless communication, and the determined metric related to the available network throughput is determined using a further machine learning module based on the predicted spectrum efficiency and the resource allocation scheme.8.The apparatus of any of claims 1 to 7, wherein the instructions that, when executed by the at least one processor, further cause the apparatus to:determine a quality level of the video delivery based on the determined metric related to the available network throughput; anddetermine the control policy based on the quality level.9.The apparatus of any of claims 1 to 8, wherein the control policy comprises an adaptation strategy for at least one of a streaming quality or a streaming rate of the video delivery.10.The apparatus of claim 9, wherein the adaptation strategy is determined based on an adaptation priority for at least one of the streaming quality or the streaming rate.11.The apparatus of claim 9 or 10, wherein the video is delivered to a further apparatus, and wherein the instructions that, when executed by the at least one processor, further cause the apparatus to:receive an adaptation preference from the further apparatus,wherein the control policy is determined based on the received adaptation preference.12.The apparatus of any of claims 1 to 11, wherein an encoding bit rate of the video is adapted based on the control policy.13.The apparatus of claim 12, wherein the encoding bit rate of the video is less than or equal to a maximum bit rate available for the video.14.The apparatus of any of claims 1 to 13, wherein the video to be delivered comprises a 360-degree panoramic video stitched from a plurality of image capturing devices.15.The apparatus of claim 14, wherein at least one of a stitching quality or an encoding quality of the video is adapted based on the control policy.16.The apparatus of claim 15, wherein the instructions that, when executed by the at least one processor, further cause the apparatus to:determine whether either or both of the stitching quality and the encoding quality of the video are to be adapted, based on an adaptation priority for at least one of the stitching quality or the encoding quality.17.The apparatus of claim 15 or 16, when the stitching quality of the video is adapted, and then the instructions that, when executed by the at least one processor, further cause the apparatus to:select one or more image capturing devices from a plurality of image capturing devices based on the control policy; andstitch one or more images from the one or more image capturing devices.18.The apparatus of any of claims 1 to 17, wherein the video or the partial video images of the video to be delivered comprises a limited viewport video.19.The apparatus of claim 18, wherein one or more viewports of the video are delivered based on the control policy.20.The apparatus of claim 19, wherein a margin size of a viewport of the one or more viewports is adapted based on the control policy.21.The apparatus of claim 19 or 20, wherein qualities of pixels of the viewport of the one or more viewports are adapted in a pattern based on the control policy.22.The apparatus of any of claims 1 to 21, wherein the video is delivered along with an indication whether the video includes a viewport frame.23.The apparatus of claim 22, wherein the indication is carried in at least one of:a header of the video, ora channel different from a channel for the delivering of the video.24.The apparatus of any of claims 1 to 23, wherein the video delivery comprises delivery of a 360-degree panoramic video.25.A method comprising:determining a characteristic metric of a wireless communication of the apparatus;determining a metric related to a network throughput of the apparatus available for video delivery, based on the determined characteristic metric of the wireless communication;determining a control policy for the video delivery, based on the determined metric related to the available network throughput; anddelivering a video or partial video images of the video based on the control policy, using the wireless communication.26.An apparatus comprising:means for determining a characteristic metric of a wireless communication of the apparatus;means for determining a metric related to a network throughput of the apparatus available for video delivery, based on the determined characteristic metric of the wireless communication;means for determining a control policy for the video delivery, based on the determined metric related to the available network throughput; andmeans for delivering a video or partial video images of the video based on the control policy, using the wireless communication.27.A computer readable medium comprising instructions stored thereon for causing an apparatus at least to perform the method of claim 25.

Citation Information

Patent Citations

  • Adaptive streaming media control method and system, computer equipment and application

    CN112953922A

  • Method of dynamic adaptive streaming for 360-degree videos

    US20190281318A1

  • Transport controlled video coding

    US20210218954A1