Network-aware streaming adaptation
A network-aware streaming adaptation system using AI for bandwidth prediction and quality control optimizes 360-degree video and audio delivery over unstable cellular networks, enhancing playback quality and resource efficiency.
Patent Information
- Application Number
- PCT/CN2024/133970
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-21
- Filing Date
- 2024-11-22
- Publication Date
- 2026-02-26
AI Technical Summary
The challenge of effectively streaming high-resolution 360-degree videos and spatial audio over unstable cellular networks, such as 5G, is exacerbated by fluctuations in signal strength and network congestion, leading to poor playback quality and high packet loss rates.
Implementing a network-aware streaming adaptation system that determines network throughput metrics to dynamically adjust video and audio delivery policies, using AI models for bandwidth prediction and quality control to optimize delivery based on real-time network conditions.
Enhances video and audio streaming quality by adapting to network fluctuations, reducing artifacts and latency, and improving network resource utilization.
Smart Images

Figure CN2024133970_26022026_PF_FP_ABST
Abstract
Description
NETWORK-AWARE STREAMING ADAPTATIONCROSS REFERENCE
[0001] The present disclosure claims priority to patent cooperation treaty (PCT) Patent Application No. PCT / CN2024 / 113789 filed on August 21, 2024 and entitled “NETWORK-AWARE STREAMING ADAPTATION” , the entirety of which is incorporated herein by reference. FIELDS
[0002] Various example embodiments of the present disclosure generally relate to the field of wireless communication and in particular, to methods, devices, apparatuses and computer readable storage medium for network-aware streaming adaptation.BACKGROUND
[0003] To capture a panoramic video or images, for example full 360-degree images, panoramic cameras, for example, 360-degree cameras (also known as omnidirectional cameras) may utilize 2, 3, 4, or more camera sensors to simultaneously capture images from surroundings, and then stitch these images together to form a uniform panoramic, for example, 360-degree image and video. In addition, the panoramic cameras, for example, the 360-degree cameras, may further comprise a plurality of microphone sensors to capture a spatial audio or a stereo audio from the surroundings. Some or all of these microphone sensors may focus on different spatial directions. To stream the image or video or the audio over a network, a (built-in) communication module, e.g., for a fourth generation (4G) or the fifth generation (5G) cellular radio network may be used. Quality of video streaming and audio streaming, however, depends on connection stability and rate limitations of the provided communication network. It may depend on signal strength and signal stability of the wireless communication network, e.g., a cellular network. The signal strength may vary depending on locations, network congestion, and environmental factors, which may lead to potential fluctuations in connection quality.SUMMARY
[0004] In a first aspect of the present disclosure, there is provided an apparatus. The apparatus comprises at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: obtain a metric related to a network throughput of the apparatus available for at least one of video delivery or audio delivery; determine a control policy for at least one of the video delivery or the audio delivery, based on the metric related to the available network throughput; and deliver, based on the control policy, at least one of: a video or partial video images of the video, or a spatial audio or a stereo audio.
[0005] In a second aspect of the present disclosure, there is provided a method. The method comprises: obtaining a metric related to a network throughput of the apparatus available for at least one of video delivery or audio delivery; determining a control policy for at least one of the video delivery or the audio delivery, based on the metric related to the available network throughput; and delivering, based on the control policy, at least one of:a video or partial video images of the video, or a spatial audio or a stereo audio.
[0006] In a third aspect of the present disclosure, there is provided an apparatus. The apparatus comprises means for obtaining a metric related to a network throughput of the apparatus available for at least one of video delivery or audio delivery; means for determining a control policy for at least one of the video delivery or the audio delivery, based on the metric related to the available network throughput; and means for delivering, based on the control policy, at least one of: a video or partial video images of the video, or a spatial audio or a stereo audio.
[0007] In a fourth aspect of the present disclosure, there is provided a computer readable medium. The computer readable medium comprises instructions stored thereon for causing an apparatus to perform at least the method according to the second aspect when executed by at least one processor.
[0008] It is to be understood that the Summary section is not intended to identify key or essential features of embodiments of the present disclosure, nor is it intended to be used to limit the scope of the present disclosure. Other features of the present disclosure will become easily comprehensible through the following description.BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Some example embodiments will now be described with reference to the accompanying drawings, where:
[0010] FIG. 1A to 1D illustrate example block diagrams of a camera device according to some example embodiments;
[0011] FIG. 2 illustrates a flowchart of an example streaming adaptation method according to some example embodiments;
[0012] FIG. 3A to 3C illustrate example prediction processes of an Artificial Intelligence (AI) model according to some example embodiments;
[0013] FIG. 4 illustrates an example process of three-dimension (3D) sphere to two-dimension (2D) viewport picture projection according to some example embodiments;
[0014] FIG. 5A shows an example process of apply different quality of service (QoS) strategies depending on priority configuration according to some example embodiments;
[0015] FIG. 5B shows an example process of quality adaptation on real-time 360-degree stitching ranking according to some example embodiments;
[0016] FIG. 6A illustrates an example case of viewport-dependent delivery (VDD) according to some example embodiments;
[0017] FIG. 6B illustrates pre-defined viewports to cover the full 360-degree degree space according to some example embodiments;
[0018] FIG. 7A illustrates an example extension of a viewport image according to some example embodiments;
[0019] FIG. 7B illustrates an example quality adaptation process according to some example embodiments;
[0020] FIG. 8A illustrates a flow of an example strategy selection process according to some example embodiments;
[0021] FIG. 8B illustrates a flow of another example strategy selection process according to some example embodiments;
[0022] FIG. 9 illustrates an example of a one-byte header extension 900 according to some example embodiments;
[0023] FIG. 10 illustrates a flowchart of an example streaming adaptation method according to some example embodiments;
[0024] FIG. 11 shows an example process of apply different quality of service (QoS) strategies depending on priority configuration according to some example embodiments;
[0025] FIG. 12A and 12B shows an example of viewport quality adaptation based on the audio source according to some example embodiments;
[0026] FIG. 13 illustrates a simplified block diagram of a device that is suitable for implementing example embodiments of the present disclosure; and
[0027] FIG. 14 illustrates a block diagram of an example computer readable medium in accordance with some example embodiments of the present disclosure.
[0028] Throughout the drawings, the same or similar reference numerals represent the same or similar element.DETAILED DESCRIPTION
[0029] Principle of the present disclosure will now be described with reference to some example embodiments. It is to be understood that these embodiments are described only for the purpose of illustration and help those skilled in the art to understand and implement the present disclosure, without suggesting any limitation as to the scope of the disclosure. Embodiments described herein can be implemented in various manners other than the ones described below.
[0030] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.
[0031] References in the present disclosure to “one embodiment, ” “an embodiment, ” “an example embodiment, ” and the like indicate that the embodiment described may comprise a particular feature, structure, or characteristic, but it is not necessary that every embodiment comprises the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is submitted that it is within the knowledge of one skilled in the art to affect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.
[0032] It is to be understood that although the terms “first, ” “second” and the like may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and similarly, a second element could be termed a first element, without departing from the scope of example embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.
[0033] As used herein, “at least one of the following: <a list of two or more elements>” and “at least one of <a list of two or more elements>” and similar wording, where the list of two or more elements are joined by “and” or “or” , mean at least any one of the elements, or at least any two or more of the elements, or at least all the elements.
[0034] As used herein, unless stated explicitly, performing a step “in response to A” does not indicate that the step is performed immediately after “A” occurs and one or more intervening steps may be included.
[0035] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments. As used herein, the singular forms “a” , “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” , “comprising” , “has” , “having” , “includes” and / or “including” , when used herein, specify the presence of stated features, elements, and / or components etc., but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof.
[0036] As used in this application, the term “circuitry” may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry) and (b) combinations of hardware circuits and software, such as (as applicable) : (i) a combination of analog and / or digital hardware circuit (s) with software / firmware and (ii) any portions of hardware processor (s) with software (including digital signal processor (s) ) , software, and memory (ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions) and (c) hardware circuit (s) and or processor (s) , such as a microprocessor (s) or a portion of a microprocessor (s) , that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.
[0037] This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit (IC) or processor integrated circuit (IC) for a mobile device or a similar integrated circuit (IC) in server, a cellular network device, or other computing or network device.
[0038] As used herein, the term “wireless communication network” refers to a network following any suitable communication standards and / or protocols, such as New Radio (NR) , Long Term Evolution (LTE) , LTE-Advanced (LTE-A) , Wideband Code Division Multiple Access (WCDMA) , High-Speed Packet Access (HSPA) , Narrow Band Internet of Things (NB-IoT) , WLAN Wireless Local Area Network (WLAN) or Wi-Fi, and so on, and any of their further generations, or any combination thereof. Furthermore, the communications between a terminal device and a network device in the communication network may be performed according to any suitable generation of telecommunication protocols, including, but not limited to, the first generation (1G) , the second generation (2G) , 2.5G, 2.75G, the third generation (3G) , the fourth generation (4G) , 4.5G, the fifth generation (5G) , the sixth generation (6G) , any further generation telecommunication protocols, and / or any other communication protocols either currently known or to be developed in the future. Embodiments of the present disclosure may be applied in various communication systems. Given the rapid development in communications, there will of course also be future type communication technologies and systems with which the present disclosure may be embodied. It should not be seen as limiting the scope of the present disclosure to only the aforementioned system.
[0039] As used herein, the term “network device” refers to a node in a communication network via which a terminal device accesses the network and receives services therefrom. The network device may refer to a base station (BS) or an access point (AP) , for example, a node B (Node or NB) , an evolved NodeB (eNodeB or eNB) , an NR NB (also referred to as a gNB) , a Remote Radio Unit (RRU) , a radio header (RH) , a remote radio head (RRH) , a relay, an Integrated Access and Backhaul (IAB) node, a low power node such as a femto, a pico, a non-terrestrial network (NTN) or non-ground network device such as a satellite network device, a low earth orbit (LEO) satellite and a gosynchronous earth orbit (GEO) satellite, an aircraft network device, and so forth, depending on the applied terminology and technology. In some example embodiments, radio access network (RAN) split architecture comprises a Centralized Unit (CU) and a Distributed Unit (DU) at an IAB donor node. An IAB node comprises a Mobile Terminal (IAB-MT) part that behaves like a UE toward the parent node, and a DU part of an IAB node behaves like a base station toward the next-hop IAB node.
[0040] The term “terminal device” refers to any end device that may be capable of wireless communication. By way of example rather than limitation, a terminal device may also be referred to as a communication device, user equipment (UE) , a Subscriber Station (SS) , a Portable Subscriber Station, a Mobile Station (MS) , or an Access Terminal (AT) . The terminal device 100 in FIG. 1A may comprise, but not limited to, a camera device, a video camera device, a video surveillance camera, a MR or AR (Mixed Reality or Augmented Reality) camera device, a video conference device, a mobile phone, a cellular phone, a smart phone, voice over IP (VoIP) phones, wireless local loop phones, a tablet, a wearable terminal device, a personal digital assistant (PDA) , a mobile communication device, portable computers, desktop computer, laptop computer, image capture terminal devices such as digital cameras, gaming terminal devices, music storage and playback appliances, vehicle-mounted wireless terminal devices, a vehicle with one or more camera devices, wireless endpoints, mobile stations, laptop-embedded equipment (LEE) , laptop-mounted equipment (LME) , USB dongles, smart devices, wireless customer-premises equipment (CPE) , an Internet of Things (IoT) device, a watch or other wearable, a head-mounted display (HMD) , a television (TV) , a set-top box, a display, a vehicle, an infotainment unit, a drone, a medical device and applications (e.g., remote surgery) , an industrial device and applications (e.g., a robot and / or other wireless devices operating in an industrial and / or an automated processing chain contexts) , a consumer electronics device, a device operating on commercial and / or industrial wireless networks, and the like, or any combination thereof. The terminal device may also correspond to a Mobile Termination (MT) part of an IAB node (e.g., a relay node) . In the following description, the terms “terminal device” , “communication device” , “terminal” , “user equipment” and “UE” may be used interchangeably.
[0041] As described above, streaming of the image or video over a 4G / 5G / 6G cellular radio network may depend on signal strength and signal stability of the cellular network. However, the signal strength may vary depending on locations, network congestion, and environmental factors, which may lead to potential fluctuations in connection quality. This situation is getting worse for 360-degree cameras, as a resolution of a 360-degree video is much higher than videos generated by traditional two-dimension (2D) cameras. Nowadays, 360-degree cameras can produce 4K and even 8K live videos in an equirectangular projection (ERP) format, with video encoding in high bit rates up to 100 or 150Mbps.
[0042] The video resolution is expected to further increase to satisfy Extended Reality (XR) , such as Virtual Reality (VR) , graphical accuracy requirements, while keeping the frames per second (fps) performance at a reasonable level (at least 30fps) . The real-time of 360-degree video streaming with a high resolution (12K, 16K, etc. ) requires solutions to deal with the network stability and rate limitation.
[0043] For increasing the video resolution, an open issue to be overcome by these solutions is to enable effective 360-degree video streaming over an unstable radio network with a built-in cellular communication module such as 5G integrated customer premise equipment (CPE) for 5G non-standalone (NSA) / standalone (SA) . Compared with traditional videos, a difference is that adaptive video streaming techniques are required in 360-degree videos with different QoS criteria. Without knowledge of the network condition, the 360-degree video streaming may become much unstable and have a higher packet loss rate with a downgraded video playback quality such as visual artifacts, stuttering, and even long pausing (e.g., higher end-to-end (e2e) latency) .
[0044] Example embodiments of the present disclosure propose a solution#1 for streaming adaptation based on determination (such as estimation, prediction or measurement ) of a network throughput available for video delivery. In this solution#1, based on a determined characteristic metric of a wireless communication of an apparatus, the apparatus determines a metric related to a network throughput available for video delivery. Based on the determined metric related to the available network throughput, the apparatus determines a control policy for the video delivery by means of the wireless communication.
[0045] This solution#1 allows dynamic adaptation of video streaming based on the network condition. In this way, efficiency of the video delivery may be improved, and network resource utilization may be increased. The following description for the 360-degree image and video process and manipulation is also suitable for less than 360-degree image or video presentation.
[0046] In addition, this situation is also getting worse for the 360-degree cameras, when considering audio collection and delivery. In some use cases, the audio delivery is very important or even more important than the video delivery for a remote user, for example, concert live, dim mine underground, etc. A high-resolution audio may have a sampling frequency at 48 kHz or even 192 kHz, 24 bits or 32 bits per sample and 4 audio channels. That means the high-resolution audio may reach 192 k × 32 bits × 4 = 24.576 Mbits / s at most for a raw audio data which is a big amount that may not be ignored in the video and audio delivery.
[0047] When the 360-degree video plus a spatial audio resolution is expected to grow further to satisfy augmented reality (AR) and the graphical accuracy requirements and keep the fps performance at the reasonable level (at least 30 fps) , the real-time of a high resolution (12K, 16K, etc. ) 360-degree video streaming and a spatial audio streaming requires novel solutions to deal with the network stability and bandwidth limitation.
[0048] For increasing the resolution of the 360-degree video plus the spatial audio, a key problem is to make the effective 360-degree video and the spatial audio combined streaming over the unstable radio network with a built-in cellular communication module (e.g. integrated 5G CPE for 5G NSA / SA etc. ) . The main difference is the adaptive video and audio combined streaming techniques in the 360-degree video with different QoS criteria. Without the network-aware technique, the 360-degree video and the spatial audio combined streaming would become much unstable and high packet loss rate with downgraded video / audio playback quality such as visual / acoustical artifacts, stuttering, and even long pausing (extremely high e2e latency) .
[0049] Example embodiments of the present disclosure also propose a solution#2 for streaming adaptation based on determination (such as estimation, prediction or measurement) of the network throughput available for at least one of the video delivery or the audio delivery. The target is to take advantage of the on-device predicted bandwidth and provide adaptive (full or partial) video and (spatial or stereo) audio output. In this solution#2, the apparatus obtains a metric related to a network throughput of the apparatus available for at least one of the video delivery or the audio delivery. Based on the metric related to the available network throughput, the apparatus determines a control policy for at least one of the video delivery or the audio delivery.
[0050] In this way, this solution#2 may produce constant-high quality video / audio combined streams and increased full network bandwidth utilization rates. The following description for the 360-degree image and video process and manipulation is also suitable for less than 360-degree image or video presentation. The following description for a spatial audio process and manipulation is also suitable for a stereo audio presentation.
[0051] Some example processes will be described with reference to FIGS. 1A to 1D of a terminal device 100 according to some example embodiments. FIGS. 1A and FIGS. 1B shows example block diagrams of the terminal device with one or more cameras in the solution#1, while FIGS. 1C and FIGS. 1D shows example block diagrams of the terminal device with one or more cameras and spatial audio in the solution#2.
[0052] FIG. 1A illustrates an example block diagram 100A of a terminal device 100, for example, a camera device 100, or any wireless communication device with one or more cameras, according to some example embodiments. The camera device 100 may be a 360-degree camera device, which comprises a mobile communication module 102, for example, having a 4G / 5G modem applied in a 4G / 5G network. In some example embodiments, the camera device 100 may act as a terminal device, such as a UE, with uplink traffic (also called an uplink (UL) device) . In some other example embodiments, the camera device 100 may act as a network device such as a base station with downlink traffic.
[0053] The uplink device may be responsible for transmitting uplink data from the UE to a network. The uplink device may be any device capable of transmitting the uplink data, such as a smartphone, a modem, or an IoT device equipped with cellular connectivity. The uplink device communicates with a network infrastructure, which comprises one or more network devices such as base stations, evolved NodeBs (eNodeBs) , or Next-Generation NodeBs (gNodeBs) in 5G networks.
[0054] The camera device 100 further may comprise a plurality of camera sensors (also called camera lenses) 104-1, …, 104-N (individually or collectively referred to as camera sensor (s) 104) which may capture images from the surroundings. N is a positive integer. Some or all of these camera sensors 104-1, …, 104-N may focus on different viewing directions. The camera device 100 may obtain (e.g., block 106) and stitch (e.g., at a stitching module 108) the images from the camera sensors 104 into a video for delivery. It is to be understood that the camera sensors 104-1, …, 104-N are shown to be included in the camera device 100 only for the purpose of illustration without suggesting any limitation. In some example embodiments, one or more camera sensors may be deployed outside of the camera device 100 in the surroundings. The camera device 100 may exchange data with the camera sensors remotely.
[0055] Furthermore, the camera device 100 may comprise a network-aware quality video control application (or network-aware quality control application, also called a VCA) 110. In some example embodiments, the VCA 110 may perform streaming adaptation based on a network bandwidth fluctuation with a bandwidth (or throughput) prediction based on machine leaning.
[0056] Some example components of the VCA 110 are shown in FIG. 1B which shows another example block diagram 100B of the camera device 100 according to some example embodiments. As shown in FIG. 1B, a traffic adaptation control loop may be organized by 3 modules: the wireless communication module 102, the VCA (model) 110, and a viewport control, graphics pipeline, video encoding and packetize and transmit module 112. The network related metrics e.g. a modulation and coding scheme (MCS) , a block error rate (BLER) , a buffer size, a received signal strength indicator (RSSI) , a reference signal receiving power (RSRP) , a reference signal receiving quality (RSRQ) , and / or the like measured by the 4G / 5G modem 102 are sent to the network-aware quality control application (VCA) 110.
[0057] A viewport is a portion of a 360-degree video that a viewer can see. A viewport may be defined by its size and orientation or direction. Its size may be the same as a Field of View (FoV) that may be characterized by width and height in pixel or angular degrees. A change or control of the viewport implies a change of a viewing window in terms of the FoV and / or the direction. In the context of this embodiment, a viewport may represent a video created from original 360 video with the size of a FoV.
[0058] The VCA comprises functions: an AI network-aware function (also referred to as an AI network-aware model or AI model 114) and a quality control application function (also referred to as a Quality Control application 116) . In some example embodiments, the AI model 114 may be implemented using, for example, a long short-term memory (LSTM) model. The AI network-aware model 114 may perform uplink (UL) bandwidth prediction and propose the predicated values to the quality control application 116. The quality control application 116 generates a control policy e.g. quality of service (QoS) encoding parameters and / or field of view (FOV) control and QoS parameters, for packaging according to the UL bandwidth values. Then, the quality control application 116 may output the control policy to the viewport control, graphics pipeline, video encoding and packetize and transmit module 112.
[0059] The VCA 110 may allow the following adaptation aspects, for example comprising Adaptation Aspect 1, Adaptation Aspect 2, Adaptation Aspect 3 and Adaptation Aspect 4. In Adaptation Aspect 1, real-time network uplink (UL) bandwidth prediction is provided. This prediction (also called an on-device prediction) is based on inference from a trained Artificial Intelligence (AI) network-aware model (or an AI model) 114. The AI model 114 may use any machine learning (ML) algorithms or architecture. The AI model 114 can be trained offline on a server or online in the device 100, depending on real needs and specific 4G / 5G network providers. Some data or metrics measured by the mobile communication module (e.g., the 4G / 5G modem) 102 may be used to train the AI model 114 and perform the inference. Such data and / or metrics may comprise real-time UL properties such as physical resource blocks (PRBs) , modulation and coding schemes (MCSs) , and a transmission (or transmit) power level, as well as some UL metrics such as block error rates (BLERs) . An output of the AI model 114 may comprise UL bandwidth prediction values that may indicate a predicted network condition in terms of rate control parameters for video streaming, such as suggested bit rates in next future periods. The bandwidth prediction may be performed once or in several times periodically.
[0060] By using the bandwidth prediction, a predicted throughput may be directed to a video encoding bit rate. Traditional network-aware adaptive bit rate (ABR) streaming (which may be performed in an application layer) is based on feedback collected from a receiver, and bit rate changes happens gradually and slowly over time based on heuristic algorithms. Thus, fast wireless network channel varies cannot be tracked, which potentially leads service degradation. The bandwidth prediction may improve dynamic and timely streaming adaptation and thus improve video delivery efficiency.
[0061] In Adaptation Aspect 2, a picture quality of a full 360-degree video is adapted. In some example embodiments, a full 360-degree (omnidirectional) video may be delivered with the quality control application 116 that may provide quality-based rate control. The AI model 114 may communicate with the built-in communication module 102 to predict near future bandwidth changes. The quality control application 116 may instruct a video encoder 120 in FIG. 1A to increase or decrease an output quality of the stitching module 108 (which may be implemented by a stitching processor unit) in FIG. 1A for stitching the images from the multiple camera lenses 104 to a full 360-degree frame (e.g., in an ERP format) . For example, the video encoder 120 may apply encoding compression control, for example, by adapting values of a video quantization parameter (QP) .
[0062] In some example embodiments, to control the quality of the full 360-degree video, the overall quality or bit rate may be changed uniformly, depending on the network situation or condition. In some other example embodiments, a region-based approach may be used to control the quality to one or more active the camera lenses 104 dynamically, based on primary or active viewing directions, e.g., field of views (FoV) or viewports, from one or more terminal devices. A 360-degree video image may be generated by stitching in real-time the image from the multiple camera sensors (lenses) 104, e.g. from front lenses (covering a front space) and back lenses (covering a back space) or 4 lenses covering a 360-degree space. The stitching module 108 (for example, a customized stitching IC) may keep one region in the 360-degree video picture (or frame) formed by the images from one or more lenses to have higher quality than the rest of the frame (off-region) formed by the images from other lenses, depending on the network bandwidth prediction from the VCA 110. In an example, the off-region (less important region) quality may be corresponding to the highest QP.
[0063] In Adaptation Aspect 3, a fixed viewport-dependent delivery (VDD) quality is adapted. Viewport-dependent delivery in 360-degree video streaming refers to a technique used to optimize the delivery of video content by prioritizing the portion of the video that is currently being requested by the client, rather than streaming the entire 360-degree video at once. The camera device 100 may output a viewport-only video that may have dynamic bit rates with respect to the network condition. Some fixed viewports with logic names such as front, left, right, back, up, and down may be partially preset or customizable by end users.
[0064] The bit rate changes may be a result of multiple techniques. In an example, dynamic viewport quality adaptation may be implemented by controlling an encoding bit rate min-max range from predicted bandwidth throughput. In another example, dynamic viewport margin quality adaptation may be applied. For example, the quality adaptation may be applied to an extra viewport margin only. The viewport part may keep an original bit rate, but the bit rate of the margin area may be lowered. In yet another example, dynamic viewport margin size adaptation may be applied. For example, a margin size may vary according to an available throughput. In some cases, the whole viewport region may be treated as “margin” with a low quality (e.g., using a high quantization value) .
[0065] In Adaptation Aspect 4, bandwidth-aware adaptation profiles are used for frame rate and quality-based adaptation strategies. The adaptation may be applied to the video output frame rates together with the video quality. Based on a bandwidth (or throughput) estimation threshold value or value sets, for example, corresponding to a network congestion severity level, both the frame rate and quality-based adaptations may be combined. The adaptation may be defined as different profiles. Each profile may define adaptation priorities as a sequence to change the quality and / or frame rate accordingly.
[0066] For example, some orders could be defined as follows: quality priority for VDD streaming and frame rate adaptation only after the available threshold exceeds a certain throughput threshold; and frame rate priority for 360-degree streaming and quality adaptation after the available threshold exceeds a certain throughput threshold. In some example embodiment, the adaptation profiles or preferences may be signaled from receivers of the video, e.g., terminal devices, to suit different use cases, which means different adaptation strategies may be chosen to apply for individual streaming sessions or viewports.
[0067] FIG. 1C illustrates an example block diagram 100C of the terminal device 100, for example, the camera device 100, or any wireless communication device with one or more cameras and spatial audio, according to some example embodiments. As in shown in FIG. 1C, there are two separate processing components for the video and the audio. In other words, other components for the audio are deployed in the block diagram 100C compared with the block diagram 100A.
[0068] As is shown in FIG. 1C, beyond the camera sensors 104-1, …, 104-N, the camera device 100 may further comprise a plurality of microphone sensors 114-1, …, 114-N (individually or collectively referred to as microphone sensor (s) 114) which may capture audio from the surroundings. N is a positive integer. Some or all of these microphone sensors 114-1, …, 114-N may focus on different spatial directions. The camera device 100 may obtain (e.g., block 116) and mix (e.g., at a mixing module 118) the audio from the microphone sensors 114 into an audio for delivery. It is to be understood that the microphone sensors 114-1, …, 114-N are shown to be included in the camera device 100 only for the purpose of illustration without suggesting any limitation. In some example embodiments, one or more microphone sensors may be deployed outside of the camera device 100 in the surroundings. The camera device 100 may exchange data with the microphone sensors remotely.
[0069] As is shown in FIG. 1C, 360-degree image renderer 108 (also referred as video renderer) may use a viewport renderer to convert 360 stitched projected 360 video frame (equirectangular or cube map, etc. ) or even raw fisheye inputs from multiple camera lenses into a 2D frame with the size of a predefined or dynamic FoV. The cropping may not as simple as image cropping. In the context of 3D, the ERP frame may need to be re-projected by rotating according to the viewing direction (e.g. rotation vector with yaw / pitch / roll) . For example, the following 3D rotation matrix represents a rotation whose yaw, pitch, and roll angles are α, β and γ. Normal cropping based on the FoV width and height in pixel may be done only after the cropped area in the original ERP image is re-projected to the center of the new ERP image, for example,
[0070] As is shown in FIG. 1C, spatial audio renderer 118 may use the binaural renderer to convert the spatial audio channels to binaural audio (typically stereo channels) . The spatial audio renderer 118 may process the sound using head-related transfer functions (HRTFs) , which simulate how sound interacts with the human head and ears to create the perception of direction and distance.
[0071] As is shown in FIG. 1C, furthermore, the camera device 100 may comprise a network-aware quality control (QC) module 140 (also referred as QC module 140) for the quality control of at least one of the video and the audio. The functions of the QC module 140 in FIG. 1C are similar to those of the VAC 110 in FIG. 1A, while the QC module 140 may also adapt the quality of the audio. In some example embodiments, the QC module 140 may perform an adaptation strategy based on a network bandwidth fluctuation with a bandwidth (or throughput) prediction.
[0072] Some example components of the QC module 140 are shown in FIG. 1D which shows another example block diagram 100D of the camera device 100 according to some example embodiments. It is noted that the components for the video quality control shown in FIG. 1B may be also included in FIG. 1D if the video delivery is also needed. As shown in FIG. 1D, a traffic adaptation control loop may be organized by 3 modules: the wireless communication module 102, the QC module 140, and an audio pipeline, audio encoding and packetize and transmit module 112. For example, the network related metrics e.g. a modulation and coding scheme (MCS) , a block error rate (BLER) , a buffer size, a received signal strength indicator (RSSI) , a reference signal receiving power (RSRP) , a reference signal receiving quality (RSRQ) , and / or the like measured by the 4G / 5G modem 102 are sent to the QC module 140.
[0073] In some embodiments, as is shown in FIG. 1D, the QC module 140 also includes 2 functions: AI network-aware model 113 and QC application 115. The AI network-aware model 113 may perform the UL bandwidth prediction and propose the predicated values to the QC application 115. The QC application 115 may generate the control policy e.g. QoS encoding parameters&FOV control &QoS parameters to packaging according to the UL Bandwidth values, then output the control policy to the audio pipeline, audio encoding and packetize and transmit module 112.
[0074] It is to be understood all operations and / or features related to the video in the block diagram 100A as described above with reference to FIG. 1A are likewise applicable to the block diagram 100C shown in the FIG. 1C and have similar effects. It is also to be understood all operations and / or features related to the video in the block diagram 100B as described above with reference to FIG. 1B are likewise applicable to the block diagram 100D shown in the FIG. 1D and have similar effects.
[0075] In the solution#1, video bandwidth adaptation may be based on changes either frame rate changes or quality changes, depending on specific requirements and constraints of use cases or application scenarios, as well as available adaptation mechanisms. Some example implementations for streaming adaptation will be described below with reference to FIGS. 2 to 9.
[0076] Now, referring to FIG. 2, a flowchart of an example streaming adaptation method 200 is illustrated. The method 200 may be implemented at an apparatus 100 such as the camera device 100 in FIG. 1A. At block 210, the apparatus 100 determines a characteristic metric of a wireless communication of the apparatus 100. The characteristic metric of the wireless communication may comprise any metric of the characteristics of the wireless communication. In an example, the metric may comprise network metrics measured by the 4G / 5G / 6G modem 102.
[0077] In some example embodiments, the characteristic metric of the wireless communication may be related to scheduling information for the wireless communication of the apparatus 100. The uplink (UL) scheduling information may comprise, but not limited, one or more of a transport block (TB) size, a modulation and coding scheme (MCS) , a modulation type, transmit power control, or a retransmission version.
[0078] Alternatively, or in addition, the characteristic metric of the wireless communication may be related to a quality of the wireless communication. In an example, the quality of the wireless communication may be indicated by a transmission result such as a positive acknowledgement (ACK) or a negative acknowledgement (NACK) . Alternatively, or in addition, the quality of the wireless communication may be indicated by one or more of a radio link control (RLC) status such as a protocol data unit (PDU) retransmission rate, a PDU NACK rate, a mean service data unit (SDU) throughput, a mean SDU latency, or a buffer occupancy. Alternatively, or in addition, the quality of the wireless communication may be indicated by one or more of a packet data convergence protocol (PDCP) status such as a SDU throughput, a SDU packet rate, or a lost PDU rate.
[0079] At block 220, based on the estimated characteristic metric of the wireless communication, the apparatus 100 determines a predicted metric related to a network throughput of the apparatus 100 available for video delivery, for example, for delivery of a 360-degree panoramic video. The predicted metric related to the network throughput of the apparatus 100 may comprise any metric related to the available network throughput of the apparatus 100.
[0080] In some example embodiments, the predicted metric related to the available network throughput may comprise a predicted value of a maximum bit rate (referred to as a first maximum bit rate) available for the video delivery. Alternatively, or in addition, the predicted metric may comprise a ratio of the first maximum bit rate to a maximum bit rate (referred to as a second maximum bit rate) allocated to the apparatus 100. Alternatively, or in addition, the predicted metric may comprise a value of the available network throughput.
[0081] In some example embodiments, the predicted metric related to the available network throughput of the apparatus 100 may be determined using a machine learning model (for example, the AI model 114 in FIG. 1B) based on the characteristic metric of the wireless communication. As shown in FIG. 3A, in a process 300A, based on the input network related metrics, the AI network-aware model 114 may output a predicted bandwidth (e.g., a maximum bit rate) . In a 5G network, a UL bandwidth (e.g., the maximum bit rate) change as a radio channel condition varies. The AI Network-aware model 114 may forecast the channel condition and UL bandwidth in advance.
[0082] In some example embodiments, the apparatus 100 may determine predicted spectrum efficiency from the characteristic metric of the wireless communication and obtain a resource allocation scheme for the wireless communication. The resource allocation scheme may comprise a current resource allocation scheme and / or a predicted resource allocation scheme for the wireless communication. Based on the predicted spectrum efficiency and the resource allocation scheme, the apparatus 100 may determine the predicted metric related to the available network throughput.
[0083] In some example embodiments, different machine learning modules may be applied for the spectrum efficiency prediction and the bandwidth prediction. For example, the predicted spectrum efficiency may be determined using a machine learning model from the characteristic metric of the wireless communication, and the predicted metric related to the available network throughput is determined using a further machine learning module based on the predicted spectrum efficiency and the resource allocation scheme.
[0084] FIG. 3B show an example prediction process 300B according to some example embodiments. The process 300B may be a UL bandwidth prediction procedure implemented by the AI model 114. As shown in FIG. 3B, the process 300B comprises the following steps. In step 1, a data input from the 4G / 5G / 6G modem 102 is processed to prepare and transform raw data of the one or more network metric e.g. MCSs, BLERs, RSSIs, RSSPs, RSSQs, or buffer status, etc. into a format that can be used for model training or make prediction.
[0085] In step 2, a prediction of UL spectrum efficiency is performed. The spectrum efficiency refers to bits that may be transmitted per Hz. In general, the UL spectrum efficiency of a UE becomes higher when the UE is close to a cell center and becomes lower when the UE is close to a cell edge. A Long Short-Term Memory (LSTM) network may be used for the spectrum efficiency prediction. The AI model 114 may be trained via large amount real data collected by the modem 102, or the data generated by simulations.
[0086] Step 3 is to apply one or more resource allocation algorithms, which calculate / estimate the resource (e.g., a physical resource block (PRB) number) that will be assigned by the network for the applications. In general, a network device such as gNB is responsible for scheduling and assign resource allocation to the application of a UE in UL, the UE or the modem has no gNB scheduling information beforehand. In this case, in an example, the apparatus 100 such as the UE 100 or model may use a current resource allocation scheme (e.g., the PRB number) when it performs the UL bandwidth prediction and assume the same amount resource will be assigned to the UE / modem by the gNB in near future. In another example, resource allocation may be calculated or estimated according to a time sequence of resources allocated in past several slots by using some algorithms.
[0087] In some example embodiments, the predicted metric related to the available network throughput of the apparatus 100 may be determined directly from the characteristic metric of the wireless communication. An example process in this regard will be described below with reference to FIG. 3C. A process 300C in FIG. 3C comprises the following steps. Step 1 in FIG. 3C is same as step 1 in FIG. 3B, which is to process the data input from the 4G / 5G / 6G modem 102, and transform the data into a format that can be used for model training or make prediction.
[0088] Step 2 in FIG. 3C is to predicate the UL bandwidth via the AI model 114 in a process different from the process 300B in FIG. 3B. In the process 300B, the UL bandwidth is predicted by 2 steps where in the first step, spectrum efficiency is predicated by the AI model 114; and in the second step, resource allocation is calculated / estimated by the resource allocation algorithm. Then, the UL bandwidth output = spectrum efficiency × resource allocation. In process 300C, the AI model 114 may perform the UL bandwidth directly via the AI model 114. The AI model 114 may be trained considering both spectrum efficiency related data and resource allocation related data, both of which may be input from the 4G / 5G / 6G modem 102.
[0089] Still with reference to FIG. 2, at block 230, the apparatus 100 determines a control policy for the video delivery, based on the predicted metric related to the available network throughput. In some example embodiments, the apparatus 100 may determine a quality level of the video delivery based on the predicted metric related to the available network throughput; and then determine the control policy based on the quality ranking.
[0090] Table 1, called Quality Ranking Table, shows example mapping to convert the predicted network throughputs to a discrete quality measurement. In this example, the predicted linear throughputs are denoted by a ratio of the predicted bandwidth to the allocated or optimal bandwidth of the apparatus 100. This table may be used to control all the following quality adaptation in the apparatus 100 such as the camera device 100. It is to be understood that the conversion is one example, and different use cases and requirements from the apparatus 100 may have different weighting to affect the mapping from the predicted bandwidth (e.g., ratio) to the quality. Table 1: Example of a Quality Ranking Table from predicted network throughput
[0091] In some example embodiments, the control policy may comprise an adaptation strategy for a streaming quality and / or a streaming rate of the video delivery. In some example embodiments, the adaptation strategy may be determined based on an adaptation priority for the streaming quality and / or the streaming rate. In some example embodiments, the apparatus 100 may receive an adaptation preference from a further apparatus 100 (e.g., the receiver of the video) to which the video is to be delivered. The apparatus 100 may determine the control policy based on the received adaptation preference.
[0092] At block 240, based on the control policy, the apparatus 100 delivers a video or partial video images of the video using the wireless communication. For example, a 360-degree camera may produce a video with a full-stitched frame. It may produce 4K (3840×1960 pixels) and even 8K (7680×3840 pixels) resolution ERP frames. In the context of a 360-degree video, "ERP" stands for "Equirectangular Projection. " Equirectangular projection is a mapping approach used to represent a three-dimension (3D) spherical surface as a two-dimension (2D) image, suitable for displaying spherical content such as 360-degree videos or panoramas on flat screens. In an equirectangular projection, as shown in FIG. 4, a spherical surface 405 is projected onto a rectangle 410, where a horizontal axis represents an azimuth angle (which is horizontally rotated around the sphere) , and a vertical axis represents an altitude angle (which is vertically rotated around the sphere) . This results in an image where the entire 360-degree horizontal field of view and 180-degree vertical field of view are mapped onto a rectangular frame.
[0093] In some example embodiments, a quality of the video may be adapted based on the control policy. In some example embodiments, the video to be delivered may comprise a 360-degree panoramic video stitched from a plurality of image capturing devices such as cameras, lenses, and sensors. In an example, to adapt the quality of the video, a stitching quality and / or an encoding quality of the video may be adapted based on the control policy. In some example embodiments, the apparatus 100 may determine whether either or both of the stitching quality and the encoding quality of the video are to be adapted, based on an adaptation priority for at least one of the stitching quality or the encoding quality.
[0094] If the stitching quality of the video is adapted, the apparatus 100 may select and change image capturing devices (such as the camera sensors 104) that are participated in the stitching. For example, the apparatus 100 may select one or more image capturing devices from a plurality of image capturing devices based on the control policy, and then stitching one or more images from the one or more image capturing devices.
[0095] FIG. 5A shows an example process 500 of apply different QoS strategies depending on priority configuration according to some example embodiments. After starting the 360 live streaming at block 501, the process 500 estimates radio network bandwidth (bitrate in uplink) in the next time period at block 502. As shown, if a stitching quality priority is selected as “QOS STRATEGY PREFERENCE” at block 503 (a left flow 505) , the network bandwidth may be converted to quality ranking preset (which may be corresponding to a range of the stitching quality, e.g. from high to low, and a special single lens only, as shown in FIG. 5B) at block 506. Further, at block 507, the quality raking value may affect the output quality of the stitched 360-degree raw frame before video encoding at block 508.
[0096] Depending on the available uplink bandwidth, the VCA 110 may determine which stitching quality is used by the 360-degree image stitching component (e.g., the stitching module 108 FIG. 1A) . As shown in FIG. 5B, in a lowest ranking mode 511, the stitched 360-degree raw frame contains the input from a single sensor (e.g., a front senor 512) and the rest (which contains the inputs from other sensors, such as a left sensor 514, a back sensor 516, and a right sensor 518) of the 360-degree ERP picture is black.
[0097] It is to be noted that the quality change may be the following two cases: 1) a sensor (lens) is configurable to different quality output; and 2) the sensor output is fixed in terms of the quality. In case 2) , an extra imaging processing step may be needed to degrade the quality. Any techniques may be applied here. For instance, downsampling may be used to reduce the number of pixels in an image by discarding pixels or averaging pixel values.
[0098] If the QoS strategy is prioritized by the encoding quality (a right flow 509 in FIG. 5A) , at block 510, the process 500 applies a predicted target bitrate to the video encoder to increase or decrease the encoding bitrate and the update can happen immediately or scheduled to the next video keyframe time. For example, the VCA 110 may pass the estimated encoding bit rates from the bandwidth bit rate. If the estimated bit rate is not presenting directly the encoding bit rate, the VCA 110 may convert the network UL bandwidth bit rate to video encoding bit rate proportionally. Then, at block 520, the process 500 performs video packaging for network delivery.
[0099] In some example embodiments, the video or the partial video images of the video to be delivered may comprise a limited viewport video. In some example embodiments, one or more viewports of the video may be delivered based on the control policy. For example, a raw 8K frame (7680×3840) with 3 or 4 color channels (RGB or RGBA) may occupy a significant memory space, no matter which image format or color space it uses, e.g. YUV planar or packed RGB format. A viewport-only output may thereafter be provided to produce one part of the full 360-degree space only. The viewport-only output means that the outputted video contains viewport frames focusing on one or more viewing directions, instead of 360-degree frames.
[0100] FIG. 6A shows an example case of viewport-dependent delivery (VDD) according to some example embodiments. A VDD mode is used delivery mode in 360-degree streaming. Both fixed-viewports and dynamic viewports may be allowed. Fixed viewports mean the number of viewport outputs and resolutions with respect to the field of view (FoV) size is pre-configured and rather fixed. Dynamic viewports mean that the viewport resolution and viewport directions are variable.
[0101] In a camera device 600, the viewport generation may be done by a Viewport Control Driver SW (VCDSW) 610. In an embodiment, the quality adaptation is to deal with the full 360-degree frames. After the uplink prediction is received from the Network-aware Quality Video Control Application (VCA) 110, different adaptation strategies may be applied based on the 2 different delivery modes: 1) original 360-degree video delivery and 2) viewport-dependent delivery adaptation. If mode 2) is applied, the delivered video may comprise partial video images and may also be referred to as a viewport video or a limited viewport video. In some example embodiments, the VCA 110 may maintain a bandwidth-to-quality ranking table (e.g., Table 1) for throttling the final video throughput within the expected bandwidth.
[0102] For instance, a 90° × 60° out of 360° × 180° viewport may save up to 12 times amount of data. A 360-degree camera may have a number of viewports predefined. FIG. 6B illustrates 6 pre-defined fixed viewports 620-1, …, 620-6 to cover the full 360-degree degree space according to some example embodiments. Each viewport may have the same fixed resolution. Moreover, each viewport has a bigger margin than 90 × 60. The dimension of the margin can be zero, which means that the viewport output is as same as the viewport size in terms of its viewing width and height in an angular degree (between 0-360 degree horizontal and 0-90 vertical space) .
[0103] Table 2 shows stitched viewport-only resolution for 360-degree cameras with 2 resolutions of 8K and 16K. Table 2
[0104] The viewports may need to be projected first to a unit sphere from the 2D plane and rotated around the sphere and re-projected back to the 2D plane. As shown in FIG. 4 which illustrates the sphere and plane projection with theta and phi degrees, with a given 3D orientation (x, y, z) in a 3D coordinate system, the x / y / z vec3 can be converted to the theta and phi degree and then to the 2D plane; and vice versa, from a 2D x / y offset in the 2D plane to a 3D position (x / y / z) in the sphere. This ensures that each viewport is at the center (0 degree longitude and 0 degree latitude, given the degree ranges of -180-180 degree horizontal and -90-90 degree vertical) .
[0105] In some example embodiments, a margin size of a viewport of the one or more viewports may be adapted based on the control policy. FIG. 7A shows an example extension of a viewport image according to some example embodiments. The viewport image contains extra pixels in all edges named “margin” . The margin size can be fixed or dynamic as well, depending on the available bandwidth.
[0106] In some example embodiments, qualities of pixels of a viewport of the one or more viewports are adapted in a pattern based on the control policy. FIG. 7B illustrates another embodiment of the quality adaptation according to some example embodiments. The image quality is changed gradually, based on the bandwidth-to-quality ranking (e.g., Table 1) . The brightness in FIG. 7B represents different quality levels. The white means the original quality; and the dark means downgraded pixels. The downgraded quality can be in different patterns, like centered radial, or linear (high quality vertically) .
[0107] In some example embodiments, an encoding bit rate of the video may be adapted based on the control policy. The encoding bit rate of the video may be less than or equal to a maximum bit rate available for the video.
[0108] In addition to or instead of the quality of the video, a frame rate of the video may be adapted based on the control policy. For example, in addition to the adaptation based on the visual quality of individual video frames, another effective way to control the data rate is to change the video frame rate with or without the quality adaptation. In general, a smooth video playback requires 30 frames per second (FPS) . A low FPS video may result in stuttering video playback and bad user experience, without sacrificing any video quality. It is a trade-off when designing an adaptation strategy. Different use cases need different strategies.
[0109] In some embodiments, a view depth of the video may be adapted based on the control policy. For example, an operation of digit zooming on may be used to get more deeper depth of a field if a viewport of the camera is fixed in a specific direction. When the viewport for the camera is fixed in the specific direction, there is no need to stitch images from other viewports. Then the view depth can be changed. Normally when the view depth is zoomed on which means the deeper field is in the view, the image’s size may reduce since less content is included in the view then the needed network bandwidth is reduced. This kind of adjustment doesn’ t reduce the image’s quality.
[0110] In some embodiments, the control policy is determined based on a motion quantity between a first frame and a subsequent second frame of the video, and in accordance with a determination the motion quantity between the first and second frames is greater than a second threshold, the apparatus 100 determines to deliver the second frame. The basic idea is to check the difference between two or more frames and to see whether the difference is bigger than a threshold. Then the apparatus 100 may enlarge the motion threshold in encoder part and only transmit the frame that the motion quantity is bigger than the threshold when the available bandwidth reduced. In this way, a motion detector may be used to detect the quantity of the motion in video and the apparatus 100 only transmits the frame that the motion quantity is bigger than the threshold. The motion may be detected in the video by several methods, such as temporal differencing, gaussian mixture models, codebook algorithm, Self-organizing background subtraction, etc.
[0111] FIG. 8A shows a flow of an example strategy selection process 800A according to some example embodiments. In this example, the adaptation may be fixed on a pre-defined or pre-selected strategy. After starting the 360 live streaming at block 801, the process 800A estimates radio network bandwidth (bitrate in uplink) in the next time period at block 802. If it is determined at block 803 that a smooth video is preferred, the process 800A proceeds to a left branch 805 where a predicted target bitrate is applied to the video quality-based adaptation, but the original output frame rate is maintained. If it is determined at block 803 that a high-quality video is preferred, the process 800A proceeds to a left branch 805 where the predicted target bitrate is applied to a video framerate so that the output bitrate can meet the expected bandwidth bitrate. Then, at block 806, the process 800A performs video packaging for network delivery.
[0112] FIG. 8B shows another an example strategy selection process 800B according to some example embodiments where strategy selection is more dynamic. It is based on the real-time bandwidth estimation and chose different strategies within a lifespan of a live delivery session. In this case, the VCA 110 in FIGS 1A and 1B may re-use the bandwidth-to-quality ranking table (e.g. Table 1) to plan the strategies. After starting the 360 live streaming at block 811, the process 800B estimates radio network bandwidth (bitrate in uplink) in the next time period at block 812. At block 813, the process 800B determines the strategy based on the ranking table. The perceived quality adaptation varies over time and depends on the network congestion and available uplink bandwidth. When the network condition is poor and the quality ranking is low, the VCA 110 may choose the frame rate-adaptation at block 814 where a predicted target bitrate is applied to the video quality-based adaptation, but the original output frame rate is maintained. Likewise, the VCA 110 may switch to the quality-based adaptation, when the ranking is becoming better at block 815 where the predicted target bitrate is applied to a video framerate so that the output bitrate can meet the expected bandwidth bitrate. The process 800B also applies a hybrid strategy (block 820) in which both the frame rate adaptation and the quality adaptation may be applied when the network condition is extremely poor, e.g. the quality ranking is in the “low” level. Then, at block 821, the process 800B performs video packaging for network delivery.
[0113] Table 3 shows strategies based on a quality ranking table. It is to be noted that the number of quality ranking levels in Table 3 is only illustrative, but not limited. The levels may be classified into more fine-grained granularity, more than 3 levels, as illustrated in Table 3. Table 3
[0114] In some example embodiments, the video may be delivered along with an indication whether the video comprises a viewport frame. In some example embodiments, the indication is carried in a header of the video, and / or a channel different from a channel for the delivering of the video.
[0115] In an embodiment, the output may switch from 360-degree video to viewport-only content. In this case, the output coded video may need to carry with content identification metadata to indicate the content type, whether or not it is a full 360-degree ERP frame, or a viewport frame. The metadata may be coded using in-band approach like H264 / H265’s custom SEI NAL unit, or one-byte header in Real-time Transport Protocol (RTP) header extension, e.g., following RFC 8285. Alternatively, the metadata may be delivered in an out-of-band approach via control channels such as a Real-Time Control Protocol (RTCP) sender report or other channels to the receiver (s) .
[0116] FIG. 9 shows an example of a one-byte header extension 900 according to some example embodiments. The 4-bit “ID” is a local identifier of this element in the range 1 to 14 inclusive. The 4-bit “len” (length) is the number, minus one, of data bytes of this header extension element following the one-byte header. In the example header extension 900, the header extension defines one extension ( “length=1” ) 905, which length is 1 byte, to carry the content identification, e.g. 0 indicates a 360-degree video, 1 indicates a viewport video.
[0117] In the solution#2, at least one of the video and the audio bandwidth adaptation may be based on the control policy for at least one of the video delivery or the audio delivery depending on specific requirements and constraints of use cases or application scenarios, as well as available adaptation mechanisms. Some example implementations for streaming adaptation will be described below with reference to FIGS. 10 to 12.
[0118] Now, referring to FIG. 10, a flowchart of an example streaming adaptation method 1000 is illustrated. The method 1000 may be implemented at an apparatus 100 such as the camera device 100 in FIG. 1C.
[0119] At block 1020, the apparatus 100 obtains a metric related to a network throughput of the apparatus available for at least one of video delivery or audio delivery, for example, for delivery of a 360-degree panoramic video with a spatial audio or a stereo audio. The predicted metric related to the network throughput of the apparatus 100 may comprise any metric related to the available network throughput of the apparatus 100.
[0120] In some example embodiments, the predicted metric related to the available network throughput may comprises at least one of: a predicted value of a first maximum bit rate available for the at least one of the video delivery or the audio delivery, a predicted value of a first maximum bandwidth for the at least one of the video delivery or the audio delivery, a ratio of the first maximum bit rate to a second maximum bit rate allocated to the apparatus, a ratio of the first maximum bandwidth to a second maximum bandwidth allocated to the apparatus, a value of an available bitrate, a value of an available bandwidth, or a value of the available network throughput.
[0121] In some example embodiments, the predicted metric related to the available network throughput of the apparatus 100 may be determined using a machine learning model (for example, the AI model 114 in FIG. 1B or the AI model 113 in FIG. 1D) based on a characteristic metric of the wireless communication. In another embodiment, the prediction may come from other external network modules or network terminals, or any 3rd-party services.
[0122] For example, the AI model may be trained offline or online, depending on the real needs and specific 4G / 5G network providers. As the camera device 110 has a built-in communication module, some key data / metric measured by the communication module may be used to train the AI model and perform the inference, such as real-time uplink (UL) properties like PRB (Physical Resource Block) , MCS, and transmit power level, etc., plus some UL metrics such as BLER. The output of the bandwidth prediction unit is the estimated network condition in terms of video streaming rate control parameters such as suggested bitrates in the next prediction timing window (prediction period in million seconds) . The prediction may encompass multiple estimated bandwidths across various future time windows and each time window may be fixed in duration or extended: ( {Pt0, Pt1, Pt2 ... Ptn} .
[0123] It is to be understood all operations and / or features related to the prediction by the AI model in the procedure 320 as described above with reference to FIG. 3A to 3C are likewise applicable to the AI model in the procedure 1020 and have similar effects. For simplification, the details for the AI model will not be repeated.
[0124] Still with reference to FIG. 10, at block 1030, the apparatus 100 determines a control policy for at least one of the video delivery or the audio delivery, based on the metric related to the available network throughput.
[0125] In some embodiments, the control policy indicates a limited content delivery for a plurality of video images of the video and multi-channel spatial audios.
[0126] Since there are two separate processing components for the video and the audio as shown in FIG. 1C, it may define the priority for the video and the audio in bandwidth adaptation for different use cases. In some use cases for 360 cameras, the video may be more important than the audio. But there are still some use cases that the audio may be more important than the video, for example, classical concert, speech live, etc. and also in some cases, the audio and the video may have similar priorities, for example, in traditional broadcasting scenario.
[0127] Then, there may be a setting for the video and the audio priority in the 360-degree camera based on different use cases the camera may apply. The setting may be pre-configured and dynamically changed from a remote control. The priority may have more accurate representation such as using percentage or ratio of the bandwidth. What’s more, since the 360-degree camera with spatial audio covers the whole space from nearly any directions, the priority may also be set per directions. For example, suppose the directions are front, rear, left, right, upside and underside, the video’s priority may set to high in the front direction, the audio’s priority may set to high in the left and right direction, and set similar priority for audio and video in other directions.
[0128] Upon receiving the prediction in the next bandwidth prediction window, there are several adaptation strategies for the video / audio combined streams. The strategies may be preconfigured and adjusted dynamically based on different use cases. For example, the strategies comprise at least one of an equal adaptation aspect 1, an unbalanced adaptation aspect 2 and an unbalanced adaptation aspect 3. Table 4 shows an example of the video and the audio adaptation strategy for the video and the audio priority in the 360-degree camera. Table 4
[0129] For example, as shown in Table 4, in the equal adaptation aspect 1, the audio and the video have a similar priority such as both being high priority or both being low priority. In other words, this strategy is to set similar or equal priority for the video and the audio. This means when the QC module 140 gets the instruction to increase or decrease the output quality for the video and the audio because of the bandwidth change, the QC module 140 may consider increasing or decreasing the quality for the video and the audio in the similar QoS level.
[0130] For example, as shown in Table 4, in the unbalanced adaptation aspect 2, the video is set a higher priority while the audio is set a lower priority. The QC module 140 communicates with the built-in communication module which predicts the near future bandwidth changes and instructs the QC module 140 to increase or decrease the output quality for the video and the audio. The QC module 140 may first consider increasing or decreasing the quality of the audio when the priority for video is higher than audio to adapt to the bandwidth change.
[0131] In some examples, the control policy to adapt the quality of the audio comprises an adaptation strategy for at least one of the followings: a sampling frequency of the audio to be delivered, a number of quantization bits of the audio, a compression scheme of the audio, a number of sound channels of the audio, or a spatial direction of the sound channels. For example, the sound channels may be decreased, e.g. from 4 (e.g. ambisonics format) to 2 (stereo format) , or increased from 2 (stereo format) to 4 (e.g. ambisonics format) . The sampling frequency for the audio may be increased or decreased in all or part of the sound channels which may depend on the 360-degree camera active video region (current viewport) .
[0132] For example, as shown in Table 4, in the unbalanced adaptation aspect 3, the video is set a lower priority while the audio is set a higher priority. The QC module 140 will first consider to increasing or decreasing the quality of the video when the priority for audio is higher than video to adapt the bandwidth change.
[0133] It is to be understood all operations and / or features related to the adaptation for the quality of the video as described above shown in FIG. 5 to 8 are likewise applicable to the adaptation for the quality of the video in the adaptation aspect 1 to 3. For simplification, the details will not be repeated.
[0134] In some example embodiments, the apparatus 100 may determine at least one quality level of the at least one of the video delivery or the audio delivery based on the metric related to the available network throughput, where the control policy is determined at least based on the at least one quality level. In some other example embodiments, the at least one quality level comprises a first quality level of the video delivery and the apparatus 100 determines a first priority level for the video delivery and a second priority level for the audio delivery, where the control policy is determined based on the first and second priority levels and the first and second quality levels.
[0135] TABLES. 5 to 7 introduces a mapping technique to convert the predicted linear throughputs to a discrete quality measurement or quality level called the Quality Ranking Table.
[0136] Table 5 shows an example mapping to convert the predicted network throughputs to a quality level in the equal adaptation aspect 1. Predicted Bandwidth in Table 5 is the whole bandwidth predicted and split to video and audio based on their current ratio. In this example, the predicted linear throughputs are denoted by a ratio of the predicted bandwidth to the allocated or optimal bandwidth of the apparatus 100. This table may be used to control all the following quality adaptation in the apparatus 100 such as the camera device 100. It is to be understood that the conversion is one example, and different use cases and requirements from the apparatus 100 may have different weighting to affect the mapping from the predicted bandwidth (e.g., ratio) to the quality. Table 5
[0137] Table 6 shows an example mapping to convert the predicted network throughputs to a discrete quality measurement in the unequal adaptation aspect. Predicted available bandwidth in Table 6 is the whole bandwidth predicted in the future subtracting the minimum bandwidth for the low priority bandwidth adaptation (the video or the audio) . In this case, the bandwidth after subtracting may be used for high priority bandwidth adaptation (video or audio) . The minimum bandwidth is a preconfigured parameter with calculation, for example, it may be min (128kbps, bandwidth / 10) for the audio and min (1Mbps, bandwidth / 10) for the video, and of course, it may be zero. For example, the whole bandwidth predicted in the future is 5 Mbps, the minimum bandwidth for the low priority bandwidth adaptation (the audio) is 128 Kbps. Then the predicted available bandwidth is 5M –0.128M = 4.872 Mbps in the future. Table 6
[0138] Table 7 shows an example mapping to convert the predicted network throughputs to a discrete quality measurement in the unequal adaptation aspect. Predicted remained bandwidth in Table 7 is the whole bandwidth for the low priority bandwidth adaptation after the allocation of the high priority bandwidth adaptation (the video or the audio) . For example, suppose the current bandwidth is 50 Mbps, and the predicted bandwidth may reduce to 5 Mbps. The video is set to have higher priority than the audio. Then the maximum bandwidth allocated to the video will be 5M –0.128M = 4.872 Mbps in the future. If the video only needs to use 4 Mbps, the audio will then get the remaining bandwidth which is 1 Mbps. Table 7
[0139] In some embodiments, the control policy comprises an adaptation strategy for allocating a resource to the video delivery and the audio delivery. For example, in the equal adaptation aspect 1, in accordance with a determination that a difference between the first and second priority levels is less than or equal to a first threshold, the apparatus 100 determines that the resource is allocated based on Table 5 to the video delivery and the audio delivery according to a ratio of a first part of the resource for the video delivery to a second part of the resource for the audio delivery. For example, in the unequal adaptation aspect 2, in accordance with a determination that a difference of the first priority level to the second priority level is greater than a first threshold, the apparatus 100 determines that allocation of the resource to the video delivery is prioritized over allocation of the resource to the audio delivery based on Table 6 and 7. For example, in the unequal adaptation aspect 3, in accordance with a determination that a difference of the second priority level to the first priority level is greater than a first threshold, the apparatus 100 determines that allocation of the resource to the audio delivery is prioritized over allocation of the resource to video delivery based on Table 6 and 7.
[0140] Reference is now made to FIG. 11, which illustrates an example signaling flow 1100 according to some example embodiments. The signaling flow 1100 shows an example signaling flow involves the different QoS strategies the QC module 140 may apply depending on the priority configuration for the video and the audio.
[0141] As is shown in FIG. 11, after 360 live video &audio streaming is started (1101) and radio network bandwidth (bitrate in uplink) in the next future period (millisecond) is estimated (1102) , when the priority for the audio and the video is similar (mid flow (adaptation aspect 1) in FIG. 11) , the QC module 140 may consider increasing or decreasing (1104 and 1105) the quality for video and audio in a similar ratio to meet the available network bandwidth based on Table 5.
[0142] In this case, the video and the audio has similar priority, then when the available network bandwidth changed, the QC module 140 may calculate the changing amount for the video and the audio part based on a changing amount of the bandwidth and a current video / audio ratio and try to keep the current ratio for the video and audio. For example, the current video / audio ratio of the bandwidth usage is 9: 1 and the available bandwidth changed from 50 M to 10 M, then the QC module 140 will reduce both video and audio quality proportionally to adapt this change and keep their ratio which means the video will get about 9 M bandwidth and the audio will get 1 M bandwidth. The QC module 140 may choose the proper adaptation method mentioned in the two sections based on the use case and the reduced amount the method may bring for video and audio separately.
[0143] When the high priority for the video and low priority for the audio is selected (1103) (left flow (adaptation aspect 2) in FIG. 11) , the QC module 140 may first consider meeting (1106) the video needs of the network available bandwidth based on Table 6 and 7 which means to decrease the audio quality first when the available network bandwidth reduced and to increase the video quality first when the available network bandwidth augmented. Then, the QC module 140 may allocate (1107) remain bandwidth into the audio delivery.
[0144] In this setting, as mentioned above, several factors impact the audio quality which are, for example, at least one of the following: the sampling frequency, quantization bit number, sound channel, or encoding format. When not count in compression, the audio code rate may be obtained based on the sampling frequency, the number of the quantization bit and the number of the sound channel. For example, possible value for the sampling frequency (kHz) may be 192, 96, 50, 48, 44.1, 22.05, 11.025, 8; possible value for the number of the quantization bit (bit) may be 32, 24, 16, 8; possible value for the number of the sound channel may be 4, 2, 1.
[0145] The sampling frequency, the change of the number of the quantization bit and compression method may be done in encoder part. The change of the number of the sound channel may be done in raw data collection part. The number of sound channels may also relate to the video viewport. For example, if the video is collected in a specific view port, it may only need to collect the sound channel in the direction of the video viewport. The spatial audio mixing component (Mixer) is responsible for spatial audio rendering when viewport mode is instructed by the QC module 140. The Mixer takes the viewport direction as input and generates spatial (2-channel) audio raw output with respect to the spatial direction related to the viewport direction. The direction data may be represented as a 3 elements vector that contains the yaw, the pitch, and the roll rotation rotated about three orthogonal axes, known as the rotation matrix. The 2D viewport is rotated and cropped from a 360-panorama video frame. Likewise, a 2-channel raw audio sample is re-rendered from the mixer from a 4-channel audio source, e.g. 4-channel ambisonics audio source.
[0146] From the above, the audio rate may decrease to 128 kbps (MP3 level) when the audio quality may be sacrificed. This means the audio only occupying very little network bandwidth and main bandwidth may allocate to video. Also in this setting, when the available network bandwidth is augmented, the QC module 140 may first consider increasing the video quality, for example, increase the video frame rate, increase the quality of less important region, etc.
[0147] When the high priority for the audio and low priority for the video is selected (1103) (right flow (adaptation aspect 3) in FIG. 11) , the QC module 140 will first consider meeting (1108) the audio needs of the network available bandwidth based on Table 6 and 7 which means to decrease the video quality first when the available network bandwidth reduced and to increase the audio quality first when the available network bandwidth augmented. Then, the QC module 140 may allocate (1109) remain bandwidth into the video delivery.
[0148] There can be several embodiments to control the quality of the video and it is to be understood all operations and / or features related to control the quality of the video are likewise applicable to the example signaling flow 1100 and have similar effects. For simplification, the details will not be repeated.
[0149] As is shown in FIG. 11, after allocating the resource for the video and the audio, the quality of the video and the audio is adapted (1110 and 1111) based on the TABLES 5 to 7. For example, the apparatus 100 passes (1110) the ranking value to the 360-image stitching component and lows the quality of generated 360 or viewpoint-only raw frames according to the ranking. The apparatus 100 passes (1111) the ranking value to audio mixing component and lows the quality of generated spatial or stereo audio raw frames according to the ranking. Finally, the video and the audio packages are delivered (1112) .
[0150] In some other embodiments, a viewport of the one or more viewports is related to an orientation of a sound audio source. In other words, the viewport quality adaptation is based on the audio source, rather than the current viewing (viewport) direction. Reference is now made to FIG. 12A and 12B, which illustrates an example embodiment 1200 according to some example embodiments.
[0151] As is shown in FIG. 12A, the direction of the viewport of Lens 4 is following the direction of the dominant audio source of Mic 4 via sound tracking techniques identifying audio sources around the device and allowing to apply audio zoom, for example, when recording your soundtrack. The audio source tracking follows the sound automatically even when the audio source changes direction or is a bit far away.
[0152] As is shown in FIG. 12B, actual viewport-only (Lens 4) output may cover the direction of the Mic 4 and cropped from the rotated 360 equirectangular frame. The rotation is needed to make sure the viewport in the center degree (0, 0 in the unit of –180-180 degree horizontal and –90 to 90 vertical) . In this embodiment, the viewport video and rendered 2-channel audio are synchronized in real-time and follow the quality adaptation strategy table.
[0153] Still with reference to FIG. 10, at block 1040, the apparatus 100 delivers, based on the control policy, at least one of: a video or partial video images of the video, or a spatial audio or a stereo audio.
[0154] In some embodiments, at block 1010, the apparatus 100 determines a characteristic metric of a wireless communication of the apparatus 100 and the metric related to the network throughput of the apparatus is determined at least based on the characteristic metric. The characteristic metric of the wireless communication may comprise any metric of the characteristics of the wireless communication. In an example, the metric may comprise network metrics measured by the 4G / 5G / 6G modem.
[0155] It is to be understood all operations and / or features related to the characteristic metric as described above with reference to block 210 are likewise applicable to the characteristic metric in block 1010 and have similar effects. For simplification, the details will not be repeated.
[0156] In addition, it is to be understood all operations and / or features related to the adaptation for the quality control of the video as described above with reference to FIG. 5 to 8 are likewise applicable to the adaptation for the quality control of the audio and have similar effects.
[0157] For example, the adaptation strategy for at least one of a streaming quality or a streaming rate of the video delivery is likewise applicable to the adaptation strategy for at least one of a streaming quality or a streaming rate of the audio delivery. For example, at least one encoding bit rate for the video delivery adapted based on the control policy is likewise applicable to the at least one encoding bit rate for the audio delivery.
[0158] In some example embodiments, an apparatus (for example, the camera device 100 in FIG. 1A to 1D) may comprise means for performing any or any combination of the methods or processes 200, 300A, 300B, 300C, 500, 800A, 800B, 1000 or 1100. The means may be implemented in any suitable form. For example, the means may be implemented in a circuitry or least one processor and at least one software module with storing instructions, for example, computer program code, that is executed by the at least one circuitry or processor. The apparatus may be implemented as or included in the terminal device 100, for example, the camera device 100 in FIG. 1A.
[0159] In some example embodiments, the apparatus (for example, the camera device 100 in FIG. 1C) comprises means for obtaining a metric related to a network throughput of the apparatus available for at least one of video delivery or audio delivery; means for determining a control policy for at least one of the video delivery or the audio delivery, based on the metric related to the available network throughput; and means for delivering, based on the control policy, at least one of a video or partial video images of the video, or a spatial audio or a stereo audio.
[0160] In some example embodiments, the apparatus further comprises: means for determining a characteristic metric of a wireless communication of the apparatus, where the metric related to the network throughput of the apparatus is determined at least based on the characteristic metric.
[0161] In some example embodiments, the characteristic metric of the wireless communication is related to at least one of: scheduling information for the wireless communication of the apparatus, or a quality of the wireless communication of the apparatus.
[0162] In some example embodiments, the metric related to the available network throughput is determined using a machine learning model based on the characteristic metric of the wireless communication.
[0163] In some example embodiments, the metric related to the available network throughput is determined based on a spectrum efficiency determined from the characteristic metric of the wireless communication and a resource allocation scheme for the wireless communication.
[0164] In some example embodiments, the resource allocation scheme comprises at least one of a current resource allocation scheme or a predicted resource allocation scheme for the wireless communication.
[0165] In some example embodiments, the spectrum efficiency is determined using a machine learning model based on the characteristic metric of the wireless communication, and the metric related to the available network throughput is determined using a further machine learning module based on the spectrum efficiency and the resource allocation scheme.
[0166] In some example embodiments, the metric related to the available network throughput comprises at least one of: a predicted value of a first maximum bit rate available for the at least one of the video delivery or the audio delivery, a predicted value of a first maximum bandwidth for the at least one of the video delivery or the audio delivery, a ratio of the first maximum bit rate to a second maximum bit rate allocated to the apparatus, a ratio of the first maximum bandwidth to a second maximum bandwidth allocated to the apparatus, a value of an available bitrate, a value of an available bandwidth, or a value of the available network throughput.
[0167] In some example embodiments, the control policy comprises an adaptation strategy for at least one of a streaming quality or a streaming rate of at least one of the video delivery or the audio delivery.
[0168] In some example embodiments, the adaptation strategy is determined based on an adaptation priority for at least one of the streaming quality or the streaming rate.
[0169] In some example embodiments, at least one of the video delivery or the audio delivery is directed to a further apparatus, and the means for determining a control policy for at least one of the video delivery or the audio delivery comprises: means for receiving the adaptation strategy from the further apparatus, where the control policy is determined based on the received adaptation strategy.
[0170] In some example embodiments, the means for determining a control policy for at least one of the video delivery or the audio delivery comprises: means for determining at least one quality level of the at least one of the video delivery or the audio delivery based on the metric related to the available network throughput, where the control policy is determined at least based on the at least one quality level.
[0171] In some example embodiments, the at least one quality level comprises a first quality level of the video delivery and a second quality level of the audio delivery, and the means for determining at least one quality level of the at least one of the video delivery or the audio delivery comprises: means for determining the first priority level for the video delivery and the second priority level for the audio delivery, where the control policy is determined based on the first and second priority levels and the first and second quality levels.
[0172] In some example embodiments, the control policy comprises an adaptation strategy for allocating a resource to the video delivery and the audio delivery.
[0173] In some example embodiments, the means for determining a control policy for at least one of the video delivery or the audio delivery comprises: means for, in accordance with a determination that a difference between the first and second priority levels is less than or equal to a first threshold, determining that the resource is allocated to the video delivery and the audio delivery according to a ratio of a first part of the resource for the video delivery to a second part of the resource for the audio delivery.
[0174] In some example embodiments, the means for determining a control policy for at least one of the video delivery or the audio delivery comprises: means for, in accordance with a determination that a difference of the first priority level to the second priority level is greater than a first threshold, determining that allocation of the resource to the video delivery is prioritized over allocation of the resource to the audio delivery.
[0175] In some example embodiments, the means for determining a control policy for at least one of the video delivery or the audio delivery comprises: means for, in accordance with a determination that a difference of the second priority level to the first priority level is greater than a first threshold, determining that allocation of the resource to the audio delivery is prioritized over allocation of the resource to video delivery.
[0176] In some example embodiments, at least one encoding bit rate for the at least one of the video delivery or the audio delivery is adapted based on the control policy.
[0177] In some example embodiments, the at least one encoding bit rate for the at least one of the video delivery or the audio delivery is less than or equal to a maximum bit rate available for the at least one of the video delivery or the audio delivery.
[0178] In some example embodiments, the video to be delivered comprises a 360-degree panoramic video stitched from a plurality of images captured by a plurality of image capturing devices.
[0179] In some example embodiments, at least one of a stitching quality or an encoding quality of the video is adapted based on the control policy.
[0180] In some example embodiments, the apparatus further comprises: means for determining whether either or both of the stitching quality and the encoding quality of the video are to be adapted, based on an adaptation priority for at least one of the stitching quality or the encoding quality.
[0181] In some example embodiments, the apparatus further comprises: means for selecting one or more image capturing devices from the plurality of image capturing devices based on the control policy; and stitch one or more images from the one or more image capturing devices.
[0182] In some example embodiments, a quality level of a first image captured by a first image capturing device of the plurality of image capturing devices is different from a quality level of a second image captured by a second image capturing device of the plurality of image capturing devices.
[0183] In some example embodiments, the video or the partial video images of the video comprises a limited viewport video.
[0184] In some example embodiments, one or more viewports of the video are delivered based on the control policy.
[0185] In some example embodiments, a margin size of a viewport of the one or more viewports is adapted based on the control policy.
[0186] In some example embodiments, qualities of pixels of a viewport of the one or more viewports are adapted in a pattern based on the control policy.
[0187] In some example embodiments, at least one of a view direction or a view depth of a viewport of the one or more viewports is adapted based on the control policy.
[0188] In some example embodiments, a viewport of the one or more viewports is related to an orientation of an audio source.
[0189] In some example embodiments, the control policy is determined based on a motion quantity between a first frame and a subsequent second frame of the video, and the means for determining a control policy for at least one of the video delivery or the audio delivery comprises: means for, in accordance with a determination the motion quantity between the first and the second frames is greater than a second threshold, determining to deliver the second frame.
[0190] In some example embodiments, the video is delivered along with an indication whether the video includes a viewport frame.
[0191] In some example embodiments, the indication is carried in at least one of: a header of the video, or a channel different from a channel for the delivering of the video.
[0192] In some example embodiments, the video delivery comprises delivery of a 360-degree panoramic video.
[0193] In some example embodiments, the control policy comprises an adaptation strategy for at least one of: a sampling frequency of the audio to be delivered, a number of quantization bits of the audio, a compression scheme of the audio, a number of sound channels of the audio, or a spatial direction of the sound channels.
[0194] In some example embodiments, the video or the partial video images of the video to be delivered comprises a limited viewport video, and where at least one of the number or the spatial direction of the sound channels is related to a viewport of the limited viewport video.
[0195] In some example embodiments, the control policy indicates a limited content delivery for a plurality of video images of the video and multi-channel spatial audios.
[0196] FIG. 13 is a simplified block diagram of a device 1300 that is suitable for implementing example embodiments of the present disclosure. The device 1300 may be provided to implement an apparatus, for example, the camera device 100 as shown in FIG. 1A or FIG. 1C. As shown in FIG. 13, the device 1300 comprises at least one or more processors 1310, one or more memories 1320 coupled to the one or more processors 1310, one or more communication modules 1340 (such as one or more wireless or wired communication modules) coupled to the one or more processor 1310, one or more camera sensors 1350 coupled to the one or more processor 1010, one or more microphone sensors 1360 coupled to the one or more processor 1310, and one or more input / output (I / O) means coupled to the one or more processor 1310, where the one or more processors 1310 and the one or more memories 1320 storing instructions that, when executed by the one or more processors 1310, cause the device 1300 at least to perform functions as described in FIGS. 1A –12.
[0197] The communication module 1340 is for wireless or wired bidirectional communications. The communication module 1340 has one or more communication interfaces to facilitate communication with one or more other modules or devices. The communication interfaces may represent any interface that is necessary for communication with other network elements. In some example embodiments, the communication module 1340 may comprise at least one antenna.
[0198] The processor 1310 may be of any type suitable to the local technical network and may comprise one or more of the following: general purpose computers, special purpose computers, microprocessors, central processing units (CPUs) , digital signal processors (DSPs) , graphical processing units (GPUs) , circuitries, or processors based on multicore processor architecture, as non-limiting examples, or any combination thereof. The device 1300 may have multiple processors, such as an application specific integrated circuit chip that is slaved in time to a clock which synchronizes the main processor.
[0199] The memory 1320 may comprise one or more non-volatile memories and one or more volatile memories. Examples of the non-volatile memories comprise, but are not limited to, a Read Only Memory (ROM) 1324, an electrically programmable read only memory (EPROM) , a flash memory, a hard disk, a compact disc (CD) , a digital video disk (DVD) , an optical disk, a laser disk, or other magnetic storage and / or optical storage. Examples of the volatile memories comprise, but are not limited to, a random-access memory (RAM) 1322 or other volatile memories that will not last in the power-down duration.
[0200] One or more computer programs 1330 comprises computer executable instructions that are executed by the associated processor 1310. The instructions of the program 1330 may comprise instructions for performing operations / acts of example embodiments of the present disclosure. The program 1330 may be stored in the memory, e.g., the ROM 1324. The processor 1310 may perform any suitable actions and processing by loading the program 1330 into the RAM 1322.
[0201] The example embodiments of the present disclosure may be implemented by means of the program 1330 so that the device 1300 may perform any process of the disclosure as discussed with reference to FIG. 1A to FIG. 12. The example embodiments of the present disclosure may also be implemented by hardware or by a combination of software and hardware.
[0202] In some example embodiments, the program 1330 may be tangibly contained in a computer readable medium which may be comprised in the device 1300 (such as in the memory 1320) or other storage devices that are accessible by the device 1300. The device 1300 may load the program 1330 from the computer readable medium to the RAM 1322 for execution. In some example embodiments, the computer readable medium may comprise any types of non-transitory storage medium, such as ROM, EPROM, a flash memory, a hard disk, CD, DVD, and the like. The term “non-transitory, ” as used herein, is a limitation of the medium itself (i.e., tangible, not a signal) as opposed to a limitation on data storage persistency (e.g., RAM vs. ROM) .
[0203] FIG. 14 shows an example of the computer readable medium 1400 which may be in form of CD, DVD or other optical storage disk. The computer readable medium 1400 has the program 1330 stored thereon.
[0204] Generally, various embodiments of the present disclosure may be implemented in hardware or special purpose circuits, software, logic or any combination thereof. Some aspects may be implemented in hardware, and other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device. Although various aspects of embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representations, it is to be understood that the block, apparatus, system, technique or method described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
[0205] Some example embodiments of the present disclosure also provide at least one computer program product tangibly stored on a computer readable medium, such as a non-transitory computer readable medium. The computer program product includes computer-executable instructions, such as those included in program modules, being executed in a device on a target physical or virtual processor, to carry out any of the methods as described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, or the like that perform particular tasks or implement particular abstract data types. The functionality of the program modules may be combined or split between program modules as desired in various embodiments. Machine-executable instructions for program modules may be executed within a local or distributed device. In a distributed device, program modules may be located in both local and remote storage media.
[0206] Program code for carrying out methods of the present disclosure may be written in any combination of one or more programming languages. The program code may be provided to a processor or controller of a general-purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program code, when executed by the processor or controller, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may execute entirely on a machine, partly on the machine, as a stand-alone software package, partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0207] In the context of the present disclosure, the computer program code or related data may be carried by any suitable carrier to enable the device, apparatus or processor to perform various processes and operations as described above. Examples of the carrier include a signal, computer readable medium, and the like.
[0208] The computer readable medium may be a computer readable signal medium or a computer readable storage medium. A computer readable medium may include but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random-access memory (RAM) , a read-only memory (ROM) , an erasable programmable read-only memory (EPROM or Flash memory) , an optical fiber, a portable compact disc read-only memory (CD-ROM) , an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0209] Further, although operations are depicted in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Likewise, although several specific implementation details are contained in the above discussions, these should not be construed as limitations on the scope of the present disclosure, but rather as descriptions of features that may be specific to particular embodiments. Unless explicitly stated, certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, unless explicitly stated, various features that are described in the context of a single embodiment may also be implemented in a plurality of embodiments separately or in any suitable sub-combination.
[0210] Although the present disclosure has been described in languages specific to structural features and / or methodological acts, it is to be understood that the present disclosure defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
An apparatus comprising:at least one processor; andat least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to:obtain a metric related to a network throughput of the apparatus available for at least one of video delivery or audio delivery;determine a control policy for at least one of the video delivery or the audio delivery, based on the metric related to the available network throughput; anddeliver, based on the control policy, at least one of:a video or partial video images of the video, ora spatial audio or a stereo audio.The apparatus of claim 1, wherein the instructions that, when executed by the at least one processor, further cause the apparatus to:determine a characteristic metric of a wireless communication of the apparatus,wherein the metric related to the network throughput of the apparatus is determined at least based on the characteristic metric.The apparatus of claim 2, wherein the characteristic metric of the wireless communication is related to at least one of:scheduling information for the wireless communication of the apparatus, ora quality of the wireless communication of the apparatus.The apparatus of any of claims 2 or 3, wherein the metric related to the available network throughput is determined using a machine learning model based on the characteristic metric of the wireless communication.The apparatus of any of claims 2 to 4, wherein the metric related to the available network throughput is determined based on:a spectrum efficiency determined from the characteristic metric of the wireless communication, anda resource allocation scheme for the wireless communication.The apparatus of claim 5, wherein the resource allocation scheme comprises at least one of a current resource allocation scheme or a predicted resource allocation scheme for the wireless communication.The apparatus of claim 5 or 6, wherein the spectrum efficiency is determined using a machine learning model based on the characteristic metric of the wireless communication, and the metric related to the available network throughput is determined using a further machine learning module based on the spectrum efficiency and the resource allocation scheme.The apparatus of claim 1 to 7, wherein the metric related to the available network throughput comprises at least one of:a predicted value of a first maximum bit rate available for the at least one of the video delivery or the audio delivery,a predicted value of a first maximum bandwidth for the at least one of the video delivery or the audio delivery,a ratio of the first maximum bit rate to a second maximum bit rate allocated to the apparatus,a ratio of the first maximum bandwidth to a second maximum bandwidth allocated to the apparatus,a value of an available bitrate,a value of an available bandwidth, ora value of the available network throughput.The apparatus of any of claims 1 to 8, wherein the control policy comprises an adaptation strategy for at least one of a streaming quality or a streaming rate of at least one of the video delivery or the audio delivery.The apparatus of claim 9, wherein the adaptation strategy is determined based on an adaptation priority for at least one of the streaming quality or the streaming rate.The apparatus of claim 9 or 10, wherein at least one of the video delivery or the audio delivery is directed to a further apparatus, and wherein the instructions that, when executed by the at least one processor, further cause the apparatus to:receive the adaptation strategy from the further apparatus,wherein the control policy is determined based on the received adaptation strategy.The apparatus of any of claims 1 to 11, wherein the instructions that, when executed by the at least one processor, further cause the apparatus to:determine at least one quality level of the at least one of the video delivery or the audio delivery based on the metric related to the available network throughput, wherein the control policy is determined at least based on the at least one quality level.The apparatus of claim 12, wherein the at least one quality level comprises a first quality level of the video delivery and a second quality level of the audio delivery, and wherein the instructions that, when executed by the at least one processor, further cause the apparatus to:determine the first priority level for the video delivery and the second priority level for the audio delivery, wherein the control policy is determined based on the first and second priority levels and the first and second quality levels.The apparatus of claim 13, wherein the control policy comprises an adaptation strategy for allocating a resource to the video delivery and the audio delivery.The apparatus of claim 14, wherein the instructions that, when executed by the at least one processor, further cause the apparatus to:in accordance with a determination that a difference between the first and second priority levels is less than or equal to a first threshold, determine that the resource is allocated to the video delivery and the audio delivery according to a ratio of a first part of the resource for the video delivery to a second part of the resource for the audio delivery.The apparatus of claim 14, wherein the instructions that, when executed by the at least one processor, further cause the apparatus to:in accordance with a determination that a difference of the first priority level to the second priority level is greater than a first threshold, determine that allocation of the resource to the video delivery is prioritized over allocation of the resource to the audio delivery.The apparatus of claim 14, wherein the instructions that, when executed by the at least one processor, further cause the apparatus to:in accordance with a determination that a difference of the second priority level to the first priority level is greater than a first threshold, determine that allocation of the resource to the audio delivery is prioritized over allocation of the resource to the video delivery.The apparatus of any of claims 1 to 17, wherein at least one encoding bit rate for the at least one of the video delivery or the audio delivery is adapted based on the control policy.The apparatus of claim 18, wherein the at least one encoding bit rate for the at least one of the video delivery or the audio delivery is less than or equal to a maximum bit rate available for the at least one of the video delivery or the audio delivery.The apparatus of any of claims 1 to 19, wherein the video to be delivered comprises a 360-degree panoramic video stitched from a plurality of images captured by a plurality of image capturing devices.The apparatus of claim 20, wherein at least one of a stitching quality or an encoding quality of the video is adapted based on the control policy.The apparatus of claim 21, wherein the instructions that, when executed by the at least one processor, further cause the apparatus to:determine whether either or both of the stitching quality and the encoding quality of the video are to be adapted, based on an adaptation priority for at least one of the stitching quality or the encoding quality.The apparatus of claim 21 or 22, when the stitching quality of the video is adapted, and wherein the instructions that, when executed by the at least one processor, further cause the apparatus to:select one or more image capturing devices from the plurality of image capturing devices based on the control policy; andstitch one or more images from the one or more image capturing devices.The apparatus of claim 23, wherein a quality level of a first image captured by a first image capturing device of the plurality of image capturing devices is different from a quality level of a second image captured by a second image capturing device of the plurality of image capturing devices.The apparatus of any of claims 1 to 24, wherein the video or the partial video images of the video comprises a limited viewport video.The apparatus of claim 25, wherein one or more viewports of the video are delivered based on the control policy.The apparatus of claim 26, wherein a margin size of a viewport of the one or more viewports is adapted based on the control policy.The apparatus of claim 26 or 27, wherein qualities of pixels of a viewport of the one or more viewports are adapted in a pattern based on the control policy.The apparatus of claim 26 to 28, wherein at least one of a view direction or a view depth of a viewport of the one or more viewports is adapted based on the control policy.The apparatus of claim 26 to 29, wherein a viewport of the one or more viewports is related to an orientation of an audio source.The apparatus of claim 20 to 30, wherein the control policy is determined based on a motion quantity between a first frame and a subsequent second frame of the video, and the instructions that, when executed by the at least one processor, further cause the apparatus to:in accordance with a determination the motion quantity between the first and the second frames is greater than a second threshold, determine to deliver the second frame.The apparatus of any of claims 1 to 31, wherein the video is delivered along with an indication whether the video includes a viewport frame.The apparatus of claim 32, wherein the indication is carried in at least one of:a header of the video, ora channel different from a channel for the delivering of the video.The apparatus of any of claims 1 to 33, wherein the video delivery comprises delivery of a 360-degree panoramic video.The apparatus of any of claims 1 to 34, wherein the control policy comprises the adaptation strategy for at least one of: a sampling frequency of the audio to be delivered, a number of quantization bits of the audio, a compression scheme of the audio, a number of sound channels of the audio, or a spatial direction of the sound channels.The apparatus of any of claims 1 to 35, wherein the video or the partial video images of the video to be delivered comprises a limited viewport video, and wherein at least one of the number or the spatial direction of the sound channels is related to a viewport of the limited viewport video.The apparatus of any of claims 1 to 36, wherein the control policy indicates a limited content delivery for a plurality of video images of the video and multi-channel spatial audios.A method comprising:obtaining a metric related to a network throughput of the apparatus available for at least one of video delivery or audio delivery.determining a control policy for at least one of the video delivery or the audio delivery, based on the metric related to the available network throughput. anddelivering, based on the control policy, at least one of:a video or partial video images of the video, ora spatial audio or a stereo audio.An apparatus comprising:means for obtaining a metric related to a network throughput of the apparatus available for at least one of video delivery or audio delivery;means for determining a control policy for at least one of the video delivery or the audio delivery, based on the metric related to the available network throughput; andmeans for delivering, based on the control policy, at least one of:a video or partial video images of the video, ora spatial audio or a stereo audio.A computer readable medium comprising instructions stored thereon for causing an apparatus at least to perform the method of claim 38.
Citation Information
Patent Citations
Adaptive streaming media control method and system, computer equipment and application
CN112953922A
Method of dynamic adaptive streaming for 360-degree videos
US20190281318A1
Transport controlled video coding
US20210218954A1