Method and apparatus for sub-picture adaptive resolution change

Adaptive resolution change methods in video coding address inefficiencies by dynamically adjusting resolution, enhancing coding efficiency and resilience in communication systems.

JP2026016694APending Publication Date: 2026-02-03INTERDIGITAL VC HOLDINGS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025184474
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-06-25
Filing Date
2025-10-31
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Existing video coding technologies struggle to efficiently adapt to changes in resolution, leading to suboptimal performance in communication systems.

Method used

Implementing methods and apparatus for adaptive resolution change (ARC) in picture and video coding, utilizing sub-picture characteristic information and inter-layer prediction to manage transitions between low and high-resolution representations.

Benefits of technology

Enhances coding efficiency and resilience by dynamically adjusting resolution based on viewport changes, improving user experience and reducing bandwidth requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026016694000001_ABST
    Figure 2026016694000001_ABST
Patent Text Reader

Abstract

Methods and apparatus related to picture and video coding in a communication system are provided.SOLUTION: The method includes determining one or more layers associated with a parameter set, generating a syntax element including an indication of whether the one or more layers associated with the parameter set are independently coded, and generating a message including the syntax element. The indication is an all layer independent flag indicating that each of the one or more layers associated with the VPS is independently coded.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to picture and / or video coding in communication systems. [Background technology]

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to and the benefit of U.S. Provisional Patent Application No. 62 / 816,686, filed in the U.S. Patent and Trademark Office on March 11, 2019, and U.S. Provisional Patent Application No. 62 / 866,528, filed in the U.S. Patent and Trademark Office on June 25, 2019, the entire contents of each of which are incorporated herein by reference as if fully set forth in their entirety below and for all applicable purposes. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] JCTVC-F158, “Resolution switching for coding efficiency and resilience”, July 2011 [Non-patent document 2] JVET-M0135, “On adaptive resolution change for VVC”, January 2019 [Non-patent document 3] JVET-M0259, “Use cases and proposed design choices or adaptive resolution changing”, January 2019 [Non-patent document 4] JVET-M0261, “AHG12:On grouping of tiles”, January 2019 Summary of the Invention

[0004] FIELD OF THE INVENTION The embodiments disclosed herein generally relate to methods and apparatus for picture and / or video coding in communication systems. [Brief explanation of the drawings]

[0005] A more detailed understanding may be had from the following detailed description, given by way of example in conjunction with the drawings attached hereto. The figures in such drawings, like the detailed description, are examples. Accordingly, the figures and detailed description should not be understood as limiting, as other equally effective examples are possible and likely to be so. Moreover, like reference numerals within the figures indicate like elements. [Figure 1A] FIG. 1 is a system diagram illustrating an example communication system in which one or more disclosed embodiments may be implemented. [Figure 1B] 1B is a system diagram illustrating an exemplary wireless transmit / receive unit (WTRU) that may be used within the communications system shown in FIG. 1A, according to one embodiment. [Figure 1C] 1B is a system diagram illustrating an example radio access network (RAN) and an example core network (CN) that may be used within the communication system shown in FIG. 1A, according to one embodiment. [Figure 1D] FIG. 1B is a system diagram illustrating a further exemplary RAN and a further exemplary CN that may be used within the communication system shown in FIG. 1A, according to one embodiment. [Figure 2] FIG. 1 is a block diagram illustrating an example of a video streaming architecture, according to one or more embodiments. [Figure 3] FIG. 1 is a diagram of an example of adaptive resolution change (ARC), according to one or more embodiments. [Figure 4] FIG. 1 is a diagram of an example of viewport adaptive streaming, according to one or more embodiments. [Figure 5] FIG. 1 is a diagram of an example of viewport switching, according to one or more embodiments. [Figure 6] FIG. 1 illustrates an example of viewport switching using ARC, according to one or more embodiments. [Figure 7] FIG. 1 is a diagram of an example of ARC implemented based on sub-picture characteristic information provided by a supplemental enhancement information (SEI) message, according to one or more embodiments. [Figure 8] FIG. 1 illustrates an example of an ARC transition between a low-resolution representation and a high-resolution representation, according to one or more embodiments. [Figure 9] FIG. 10 is a diagram of an example encoding procedure for determining one or more ARC transition points, according to one or more embodiments. [Figure 10] FIG. 1 is a diagram of an example of ARC based on inter-layer prediction, according to one or more embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0006] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the embodiments and / or examples disclosed herein. However, it should be understood that such embodiments and examples may be practiced without some or all of the specific details set forth herein. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to obscure the following description. Furthermore, embodiments and examples not specifically described herein may be practiced in place of, or in combination with, these embodiments and other examples described, disclosed, or otherwise explicitly, implicitly, and / or inherently provided (collectively, "provided") herein. Although various embodiments are described and / or claimed herein in which apparatuses, systems, devices, etc., and / or any elements thereof, perform operations, processes, algorithms, functions, etc., and / or any portions thereof, it should be understood that any embodiment described and / or claimed herein may be configured with any apparatus, system, device, etc., and / or any elements thereof, performing any operation, process, algorithm, function, etc., and / or any portion thereof.

[0007] In this application, the terms "reconstructed" and "decoded" are used interchangeably, the terms "pixel" and "sample" are used interchangeably, and the terms "image," "picture," and "frame" may be used interchangeably. Typically, but not necessarily, the term "reconstructed" is used on the encoder side, while "decoded" is used on the decoder side.

[0008] Various methods are described herein, each of which includes one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for the proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. Furthermore, terms such as “first,” “second,” and the like may be used in various embodiments to modify elements, components, step actions, etc., e.g., “first decode” and “second decode.” The use of such terms does not imply an order to the modified actions unless specifically required. Consequently, in this example, the first decode need not be performed before the second decode, but may occur before, during, or within a period overlapping with the second decode.

[0009] Representative communication networks The methods, apparatus, and systems provided herein are well suited for communications involving both wired and wireless networks. Wired networks are well known. With reference to Figures 1A-1D, an overview of various types of wireless devices and infrastructure is provided, in which various elements of the network may utilize, perform, be arranged and / or adapted and / or configured in accordance with the methods, apparatus, and systems provided herein.

[0010] 1A illustrates an exemplary communication system 100 in which one or more disclosed embodiments may be implemented. The communication system 100 may be a multiple-access system that provides content, such as voice, data, video, messaging, broadcasts, etc., to multiple wireless users. The communication system 100 may enable the multiple wireless users to access such content through the sharing of system resources, including wireless bandwidth. For example, the communication system 100 may employ one or more channel access schemes, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single-carrier FDMA (SC-FDMA), zero-tailed unique word DFT-spread OFDM (ZT-UW-DFT-S-OFDM), unique word OFDM (UW-OFDM), resource block filtered OFDM, filter bank multicarrier (FBMC), etc.

[0011] 1A, communications system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RANs 104 / 113, CNs 106 / 115, public switched telephone network (PSTN) 108, the Internet 110, and other networks 112, although it will be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of WTRUs 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a wireless environment. For example, the WTRUs 102a, 102b, 102c, 102d, all of which may be referred to as “stations” and / or “STAs,” may be configured to transmit and / or receive wireless signals and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, notebooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, IoT devices, watches or other wearables, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in the context of industrial and / or automated processing chains), consumer electronic devices, devices operating on commercial and / or industrial wireless networks, etc. The WTRUs 102a, 102b, 102c, and 102d may all be referred to interchangeably as UEs.

[0012] The communications system 100 may also include a base station 114a and / or a base station 114b. Each of the base stations 114a, 114b may be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communications networks, such as the CN 106 / 115, the Internet 110, and / or other networks 112. For example, the base stations 114a, 114b may be a base transceiver station (BTS), a Node B, an eNodeB, a Home Node B, a Home eNodeB, a gNB, a New Radio (NR) Node B, a site controller, an access point (AP), a wireless router, etc. Although the base stations 114a, 114b are each depicted as a single element, it will be understood that the base stations 114a, 114b may include any number of interconnected base stations and / or network elements.

[0013] The base station 114a may be part of the RAN 104 / 113, which may also include other base stations and / or network elements (not shown), such as a base station controller (BSC), a radio network controller (RNC), relay nodes, etc. The base station 114a and / or base station 114b may be configured to transmit and / or receive wireless signals on one or more carrier frequencies, which may be referred to as a cell (not shown). These frequencies may be in the licensed spectrum, the unlicensed spectrum, or a combination of the licensed and unlicensed spectrum. A cell may provide coverage for wireless services in a particular geographic area, which may be relatively fixed or which may change over time. A cell may be further divided into cell sectors. For example, the cell associated with the base station 114a may be divided into three sectors. Thus, in one embodiment, the base station 114a may include three transceivers, e.g., one transceiver for each sector of the cell. In one embodiment, the base station 114a may employ multiple-input multiple-output (MIMO) technology and may utilize multiple transceivers per sector of the cell. For example, beamforming may be used to transmit and / or receive signals in desired spatial directions.

[0014] The base stations 114a, 114b may communicate with one or more of the WTRUs 102a, 102b, 102c, 102d over the air interface 116, which may be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 116 may be established using any suitable radio access technology (RAT).

[0015] More particularly, as noted above, the communication system 100 may be a multiple-access system and may employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, the base stations 114a and WTRUs 102a, 102b, 102c in the RAN 104 / 113 may implement a radio technology such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which may establish the air interface 115 / 116 / 117 using Wideband CDMA (WCDMA). WCDMA may include communication protocols such as High Speed ​​Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High Speed ​​Downlink (DL) Packet Access (HSDPA) and / or High Speed ​​UL Packet Access (HSUPA).

[0016] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which may establish the air interface 116 using Long Term Evolution (LTE) and / or LTE Advanced (LTE-A) and / or LTE Advanced Pro (LTE-A Pro).

[0017] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as New Radio (NR) radio access, which may establish the air interface 116 using NR.

[0018] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement multiple radio access technologies. For example, the base station 114a and the WTRUs 102a, 102b, 102c may jointly implement LTE and NR radio access, e.g., using a dual connectivity (DC) principle. Thus, the air interface utilized by the WTRUs 102a, 102b, 102c may be characterized by multiple types of radio access technologies and / or transmissions sent to and from multiple types of base stations (e.g., eNBs and gNBs).

[0019] In other embodiments, the base station 114a and the WTRUs 102a, 102b, 102c may implement a wireless technology such as IEEE 802.11 (i.e., Wireless Fidelity (WiFi)), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile communications (GSM), Enhanced Data Rates for GSM Evolution (EDGE), GSM EDGE (GERAN), or the like.

[0020] 1A may be, for example, a wireless router, a Home NodeB, a Home eNodeB, or an access point and may utilize any suitable RAT to facilitate wireless connectivity in a local area, such as a workplace, a home, a vehicle, a premises, an industrial facility, an air corridor (e.g., for use by drones), a road, etc. In one embodiment, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In one embodiment, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 114b and the WTRUs 102c, 102d may utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a picocell or femtocell. 1A, the base station 114b may have a direct connection to the Internet 110. Therefore, the base station 114b may not need to access the Internet 110 via the CN 106 / 115.

[0021] The RAN 104 / 113 may be in communication with the CN 106 / 115, which may be any type of network configured to provide voice, data, application, and / or VoIP services to one or more of the WTRUs 102a, 102b, 102c, 102d. The data may have varying quality of service (QoS) requirements, such as different throughput requirements, latency requirements, error resilience requirements, reliability requirements, data throughput requirements, mobility requirements, etc. The CN 106 / 115 may provide call control, billing services, mobile location services, prepaid calling, Internet connectivity, video distribution, etc., and / or perform high-level security functions such as user authentication. Although not shown in FIG. 1A , it will be understood that the RAN 104 / 113 and / or the CN 106 / 115 may be in direct or indirect communication with other RANs employing the same RAT as the RAN 104 / 113 or a different RAT. For example, in addition to being connected to the RAN 104 / 113, which may utilize NR radio technology, the CN 106 / 115 may also be in communication with another RAN (not shown) that employs GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.

[0022] The CN 106 / 115 may also serve as a gateway for the WTRUs 102a, 102b, 102c, 102d to access the PSTN 108, the Internet 110, and / or other networks 112. The PSTN 108 may include a circuit-switched telephone network providing plain old telephone service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices that use common communication protocols, such as TCP, UDP, and / or IP in the TCP / IP Internet protocol suite. The network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, the network 112 may include another CN connected to one or more RANs, which may employ the same RAT as the RAN 104 / 113 or a different RAT.

[0023] Some or all of the WTRUs 102a, 102b, 102c, 102d in the communications system 100 may include multi-mode capabilities (e.g., the WTRUs 102a, 102b, 102c, 102d may include multiple transceivers for communicating with different wireless networks over different wireless links). For example, the WTRU 102c shown in FIG. 1A may be configured to communicate with a base station 114a, which may employ a cellular-based wireless technology, and with a base station 114b, which may employ an IEEE 802.2 wireless technology.

[0024] 1B is a system diagram illustrating an example WTRU 102. As shown in FIG. 1B, the WTRU 102 may include, among other things, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power source 134, a GPS chipset 136, and / or other peripherals 138. It will be understood that the WTRU 102 may include any subcombination of the foregoing elements while remaining consistent with an embodiment.

[0025] The processor 118 may be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, other types of integrated circuits (ICs), a state machine, etc. The processor 118 may perform signal coding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 102 to operate in a wireless environment. The processor 118 may be coupled to the transceiver 120, which may be coupled to the transmit / receive element 122. While FIG. 1B depicts the processor 118 and the transceiver 120 as separate components, it will be understood that the processor 118 and the transceiver 120 may be integrated together in an electronic package or chip.

[0026] The transmit / receive element 122 may be configured to transmit and receive signals to and from a base station (e.g., base station 114a) over the air interface 116. For example, in one embodiment, the transmit / receive element 122 may be an antenna configured to transmit and / or receive RF signals. In one embodiment, the transmit / receive element 122 may be an emitter / detector configured to transmit and / or receive IR, UV, or visible light signals, for example. In yet another embodiment, the transmit / receive element 122 may be configured to transmit and / or receive both RF and light signals. It will be understood that the transmit / receive element 122 may be configured to transmit and / or receive any combination of wireless signals.

[0027] 1B depicts the transmit / receive element 122 as a single element, the WTRU 102 may include any number of transmit / receive elements 122. More particularly, the WTRU 102 may employ MIMO technology. Thus, in one embodiment, the WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 116.

[0028] The transceiver 120 may be configured to modulate signals to be transmitted by the transmit / receive element 122 and to demodulate signals received by the transmit / receive element 122. As mentioned above, the WTRU 102 may have multi-mode capabilities. Thus, the transceiver 120 may include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs, such as, for example, NR and IEEE 802.11.

[0029] The processor 118 of the WTRU 102 may be coupled to and may receive user input data from a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light emitting diode (OLED) display unit). The processor 118 may also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. Furthermore, the processor 118 may access information and store data in any type of suitable memory, such as non-removable memory 130 and / or removable memory 132. The non-removable memory 130 may include RAM, ROM, a hard disk, or any other type of memory storage device. The removable memory 132 may include a SIM card, a memory stick, a secure digital (SD) memory card, etc. In other embodiments, the processor 118 may access information and store data in memory that is not physically located on the WTRU 102, such as on a server or home computer (not shown).

[0030] The processor 118 may receive power from the power source 134 and may be configured to distribute and / or control power to other components in the WTRU 102. The power source 134 may be any suitable device for powering the WTRU 102. For example, the power source 134 may include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel-metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.

[0031] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to, or instead of, information from the GPS chipset 136, the WTRU 102 may receive location information from base stations (e.g., base stations 114a, 114b) over the air interface 116 and / or determine its location based on the timing of signals received from two or more nearby base stations. It will be understood that the WTRU 102 may obtain location information by way of any suitable location-determination method while remaining consistent with an embodiment.

[0032] The processor 118 may also be coupled to other peripherals 138, which may include one or more software and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, the peripherals 138 may include an accelerometer, an e-compass, a satellite transceiver, a digital camera (for photos and / or videos), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth module, a frequency modulation (FM) radio unit, a digital music player, a media player, a video game player module, an internet browser, a virtual reality and / or augmented reality (VR / AR) device, an activity tracker, etc. The peripherals 138 may include one or more sensors, which may be one or more of a gyroscope, an accelerometer, a Hall effect sensor, a magnetometer, an orientation sensor, a proximity sensor, a temperature sensor, a time sensor, a geolocation sensor, an altimeter, a light sensor, a touch sensor, a magnetometer, a barometer, a gesture sensor, a biometric sensor, and / or a humidity sensor.

[0033] The WTRU 102 may include a full-duplex radio where transmission and reception of some or all of the signals (e.g., associated with a particular subframe for both the UL (e.g., for transmission) and the downlink (e.g., for reception)) may be parallel and / or simultaneous. The full-duplex radio may include an interference management unit 139 to reduce and / or substantially eliminate self-interference through either hardware (e.g., chokes) or signal processing via a processor (e.g., a separate processor (not shown) or the processor 118). In one embodiment, the WTRU 102 may include a half-duplex radio for transmission and reception of some or all of the signals (e.g., associated with a particular subframe for either the UL (e.g., for transmission) or the downlink (e.g., for reception)).

[0034] 1C is a system diagram illustrating the RAN 104 and the CN 106, according to one embodiment. As noted above, the RAN 104 may employ E-UTRA radio technology to communicate with the WTRUs 102a, 102b, 102c over the air interface 116. The RAN 104 may also be in communication with the CN 106.

[0035] The RAN 104 may include eNodeBs 160a, 160b, and 160c, although it will be understood that the RAN 104 may include any number of eNodeBs while remaining consistent with an embodiment. The eNodeBs 160a, 160b, and 160c may each include one or more transceivers for communicating with the WTRUs 102a, 102b, and 102c over the air interface 116. In one embodiment, the eNodeBs 160a, 160b, and 160c may implement MIMO technology. Thus, the eNodeB 160a, for example, may use multiple antennas to transmit wireless signals to and / or receive wireless signals from the WTRU 102a.

[0036] Each of the eNodeBs 160a, 160b, 160c may be associated with a particular cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, scheduling of users in the UL and / or DL, etc. As shown in FIG. 1C, the eNodeBs 160a, 160b, 160c may communicate with each other via an X2 interface.

[0037] 1C may include a mobility management entity (MME) 162, a serving gateway (SGW) 164, and a packet data network (PDN) gateway (or PGW) 166. While each of the above elements is shown as part of the CN 106, it will be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.

[0038] The MME 162 may be connected to each of the eNodeBs 160a, 160b, 160c in the RAN 104 via an S1 interface and may act as a control node. For example, the MME 162 may be responsible for authenticating users of the WTRUs 102a, 102b, 102c, activating / deactivating bearers, selecting a particular serving gateway during initial attach of the WTRUs 102a, 102b, 102c, etc. The MME 162 may provide a control plane function for switching between the RAN 104 and other RANs (not shown) that employ other radio technologies such as GSM and / or WCDMA.

[0039] The SGW 164 may be connected to each of the eNodeBs 160a, 160b, 160c in the RAN 104 via an S1 interface. The SGW 164 may generally route and forward user data packets to and from the WTRUs 102a, 102b, 102c. The SGW 164 may perform other functions such as anchoring the user plane during handovers between eNodeBs, triggering paging when DL data is available for the WTRUs 102a, 102b, 102c, and managing and storing the context of the WTRUs 102a, 102b, 102c.

[0040] The SGW 164 may be connected to a PGW 166 that may provide the WTRUs 102a, 102b, 102c with access to packet-switched networks, such as the Internet 110, to facilitate communications between the WTRUs 102a, 102b, 102c and IP-enabled devices.

[0041] The CN 106 may facilitate communication with other networks. For example, the CN 106 may provide the WTRUs 102a, 102b, 102c with access to circuit-switched networks, such as the PSTN 108, to facilitate communication between the WTRUs 102a, 102b, 102c and traditional fixed communication devices. For example, the CN 106 may include or communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between the CN 106 and the PSTN 108. Additionally, the CN 106 may provide the WTRUs 102a, 102b, 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.

[0042] Although the WTRU is depicted in FIGS. 1A-1D as a wireless terminal, it is contemplated that in some representative embodiments such a terminal may use a wired communication interface with the communication network (e.g., temporarily or permanently).

[0043] In an exemplary embodiment, the other network 112 may be a WLAN.

[0044] A WLAN in infrastructure basic service set (BSS) mode has an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP may have access or an interface to a distribution system (DS) or another type of wired / wireless network that carries traffic to and / or from the BSS. Originating traffic to a STA may arrive through the AP and be sent to the STA. Traffic originating from a STA to a destination outside the BSS may be sent to the AP for delivery to the respective destination. Traffic between STAs within a BSS may be sent through the AP; for example, a source STA may send traffic to the AP, which then delivers the traffic to the destination STA. Traffic between STAs within a BSS may be considered or referred to as peer-to-peer traffic. Peer-to-peer traffic may be sent (e.g., directly) between a source STA and a destination STA using direct link setup (DLS). In certain representative embodiments, the DLS may use 802.11e DLS or 802.11z tunneled DLS (TDLS). A WLAN using an Independent BSS (IBSS) mode may not have an AP, and STAs within or using the IBSS (e.g., all STAs) may communicate directly with each other. The IBSS communication mode is sometimes referred to herein as an "ad hoc" communication mode.

[0045] When using 802.11ac infrastructure mode operation or a similar mode of operation, an AP may transmit beacons on a fixed channel, such as a primary channel. The primary channel may be a fixed width (e.g., a wide 20 MHz bandwidth) or a width dynamically set via signaling. The primary channel may be the operating channel of the BSS and may be used by STAs to establish a connection with the AP. In certain representative embodiments, carrier sense multiple access with collision avoidance (CSMA / CA) may be implemented, for example, in an 802.11 system. With CSMA / CA, STAs (e.g., all STAs), including the AP, can sense the primary channel. If the primary channel is sensed / detected by a particular STA and / or determined to be busy, the particular STA can back off. One STA (e.g., only one station) may transmit on a particular BSS at any time.

[0046] High-throughput (HT) STAs may use 40 MHz wide channels for communication, for example, by combining a primary 20 MHz channel with adjacent or non-adjacent 20 MHz channels to form the 40 MHz wide channel.

[0047] A very high throughput (VHT) STA can support channels of 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz width. A 40 MHz and / or 80 MHz channel can be formed by combining contiguous 20 MHz channels. A 160 MHz channel can be formed by combining eight contiguous 20 MHz channels or two non-contiguous 80 MHz channels, sometimes referred to as an 80+80 configuration. For the 80+80 configuration, data can be passed through a segment parser after channel encoding, which can split the data into two streams. Inverse fast Fourier transform (IFFT) processing and time-domain processing can be performed separately on each stream. The streams can be mapped onto two 80 MHz channels, and the data can be transmitted by the transmitting STA. At the receiver of the receiving STA, the above operations for the 80+80 configuration can be reversed, and the combined data can be transmitted to the medium access control (MAC).

[0048] Sub-1 GHz operating modes are supported by 802.11af and 802.11ah. In 802.11af and 802.11ah, the channel operating bandwidths and carriers are reduced relative to those used in 802.11n and 802.11ac. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV White Space (TVWS) spectrum, while 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. According to a representative embodiment, 802.11ah can support meter-type control / machine-type communications, such as MTC devices in macro coverage areas. MTC devices may have limited functionality, including specific features, such as support for (e.g., support only) specific and / or limited bandwidths. MTC devices may include batteries with above-threshold battery life (e.g., maintain very long battery life).

[0049] A WLAN system may support multiple channels and channel bandwidths, such as 802.11n, 802.11ac, 802.11af, and 802.11ah, and the WLAN system includes a channel that may be designated as a primary channel. The primary channel may have a bandwidth equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel may be configured and / or limited by the STA from among all STAs operating in the BSS that support the smallest bandwidth operating mode. In an 802.11ah example, the primary channel may be 1 MHz wide for a STA (e.g., an MTC-type device) that supports (e.g., only supports) the 1 MHz mode, even if the AP and other STAs in the BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operating modes. Carrier sensing and / or network allocation vector (NAV) configuration may depend on the status of the primary channel. For example, if the primary channel is busy because a STA (that only supports a 1 MHz mode of operation) is transmitting to the AP, the entire available frequency band may be considered busy, even though most of the frequency band may remain idle and available for use.

[0050] In the United States, the available frequency bands available for 802.11ah are 902MHz to 928MHz. In South Korea, the available frequency bands are 917.5MHz to 923.5MHz. In Japan, the available frequency bands are 916.5MHz to 927.5MHz. The total available bandwidth for 802.11ah is 6MHz to 26MHz depending on the country code.

[0051] 1D is a system diagram illustrating the RAN 113 and the CN 115, according to one embodiment. As mentioned above, the RAN 113 may employ NR radio technology to communicate with the WTRUs 102a, 102b, 102c over the air interface 116. The RAN 113 may also be in communication with the CN 115.

[0052] While the RAN 113 may include gNBs 180a, 180b, and 180c, it will be understood that the RAN 113 may include any number of gNBs while remaining consistent with an embodiment. The gNBs 180a, 180b, and 180c may each include one or more transceivers for communicating with the WTRUs 102a, 102b, and 102c over the air interface 116. In one embodiment, the gNBs 180a, 180b, and 180c may implement MIMO technology. For example, the gNBs 180a, 180b may utilize beamforming to transmit and / or receive signals to and from the gNBs 180a, 180b, and 180c. Thus, for example, the gNB 180a may use multiple antennas to transmit and / or receive wireless signals to and from the WTRU 102a. In one embodiment, the gNBs 180a, 180b, and 180c may implement carrier aggregation technology. For example, the gNB 180a may transmit multiple component carriers to the WTRU 102a (not shown). A subset of these component carriers may be on the unlicensed spectrum, and the remaining component carriers may be on the licensed spectrum. In one embodiment, the gNBs 180a, 180b, and 180c may implement coordinated multipoint (CoMP) technology. For example, the WTRU 102a may receive coordinated transmissions from the gNBs 180a and 180b (and / or 180c).

[0053] The WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using transmissions associated with scalable numerology. For example, the OFDM symbol spacing and / or OFDM subcarrier spacing may vary for different transmissions, different cells, and / or different portions of the wireless transmission spectrum. The WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using subframes or transmission time intervals (TTIs) of different or scalable lengths (e.g., different lengths of absolute time including and / or lasting different numbers of OFDM symbols).

[0054] The gNBs 180a, 180b, 180c may be configured to communicate with the WTRUs 102a, 102b, 102c in a standalone configuration and / or a non-standalone configuration. In a standalone configuration, the WTRUs 102a, 102b, 102c can communicate with the gNBs 180a, 180b, 180c without accessing another RAN (e.g., eNodeBs 160a, 160b, 160c, etc.). In a standalone configuration, the WTRUs 102a, 102b, 102c can utilize one or more gNBs 180a, 180b, 180c as mobility anchor points. In a standalone configuration, the WTRUs 102a, 102b, 102c can communicate with the gNBs 180a, 180b, 180c using signals in unlicensed bands. In a non-standalone configuration, the WTRUs 102a, 102b, 102c may communicate / connect with a gNB 180a, 180b, 180c while also communicating / connecting with another RAN, such as an eNodeB 160a, 160b, 160c. For example, the WTRUs 102a, 102b, 102c may implement the DC principle to communicate with one or more gNBs 180a, 180b, 180c and one or more eNodeBs 160a, 160b, 160c at approximately the same time. In a non-standalone configuration, the eNodeBs 160a, 160b, 160c may act as mobility anchors for the WTRUs 102a, 102b, 102c, and the gNBs 180a, 180b, 180c may provide additional coverage and / or throughput for serving the WTRUs 102a, 102b, 102c.

[0055] Each of the gNBs 180a, 180b, 180c may be associated with a particular cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, scheduling of users in the UL and / or DL, support for network slicing, dual connectivity, interworking between NR and E-UTRA, routing of user plane data to User Plane Functions (UPFs) 184a, 184b, routing of control plane information to Access and Mobility Management Functions (AMFs) 182a, 182b, etc. As shown in FIG. 1D , the gNBs 180a, 180b, 180c may communicate with each other via an Xn interface.

[0056] 1D may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and possibly a Data Network (DN) 185a, 185b. While each of the above elements is shown as part of the CN 115, it will be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.

[0057] The AMF 182a, 182b may be connected to one or more gNBs 180a, 180b, 180c in the RAN 113 via an N2 interface and may function as a control node. For example, the AMF 182a, 182b may be responsible for authenticating users of the WTRUs 102a, 102b, 102c, supporting network slicing (e.g., handling different PDU sessions with different requirements), selecting a particular SMF 183a, 183b, managing registration areas, terminating NAS signaling, mobility management, etc. Network slicing may be used by the AMF 182a, 182b to customize the CN support of the WTRUs 102a, 102b, 102c based on the type of service being utilized by the WTRUs 102a, 102b, 102c. Different network slices may be established for different use cases, for example, services relying on Ultra-Reliable Low-Latency (URLLC) access, services relying on enhanced Massive Mobile Broadband (eMBB) access, services for Machine Type Communications (MTC) access, etc. The AMF 182 may provide a control plane function for switching between the RAN 113 and other RANs (not shown) that use other radio technologies, such as LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies, such as WiFi.

[0058] The SMFs 183a, 183b may be connected to the AMFs 182a, 182b in the CN 115 via an N11 interface. The SMFs 183a, 183b may also be connected to the UPFs 184a, 184b in the CN 115 via an N4 interface. The SMFs 183a, 183b may select and control the UPFs 184a, 184b and configure the routing of traffic through the UPFs 184a, 184b. The SMFs 183a, 183b may perform other functions such as managing and assigning IP addresses for WTRUs or UEs, managing PDU sessions, controlling policy enforcement and QoS, providing downlink data notification, etc. The PDU session type may be IP-based, non-IP-based, Ethernet-based, etc.

[0059] The UPFs 184a, 184b may be connected to one or more gNBs 180a, 180b, 180c in the RAN 113 via an N3 interface, which provides the WTRUs 102a, 102b, 102c with access to packet-switched networks such as the Internet 110 and facilitates communication between the WTRUs 102a, 102b, 102c and IP-enabled devices. The UPFs 184, 184b may perform other functions such as routing and forwarding packets, enforcing user plane policy, supporting multi-homed PDU sessions, handling user plane QoS, buffering downlink packets, providing a mobility anchor, etc.

[0060] The CN 115 may facilitate communication with other networks. For example, the CN 115 may include or communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between the CN 115 and the PSTN 108. Additionally, the CN 115 may provide the WTRUs 102a, 102b, 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, the WTRUs 102a, 102b, 102c may be connected to the local data networks (DNs) 185a, 185b through the UPFs 184a, 184b via an N3 interface to the UPFs 184a, 184b and an N6 interface between the UPFs 184a, 184b and the DNs 185a, 185b.

[0061] 1A-1D and the corresponding description thereof, one or more or all of the functions described herein for one or more of the WTRUs 102a-d, base stations 114a-b, eNodeBs 160a-c, MME 162, SGW 164, PGW 166, gNBs 180a-c, AMFs 182a-b, UPFs 184a-b, SMFs 183a-b, DNs 185a-b, and / or any other devices described herein may be performed by one or more emulation devices (not shown). The emulation devices may be one or more devices configured to emulate one or more or all of the functions described herein. For example, the emulation devices may be used to test other devices and / or to simulate network and / or WTRU functions.

[0062] The emulation device can be designed to implement one or more tests of other devices in a lab environment and / or an operator network environment. For example, one or more emulation devices can be fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to perform one or more, or all, functions for testing other devices in the communication network. One or more emulation devices can be temporarily implemented / deployed as part of a wired and / or wireless communication network to perform one or more, or all, functions. The emulation device can be directly coupled to another device for testing purposes and / or can perform testing using wireless communication.

[0063] The one or more emulation devices may also perform one or more functions, including all functions, without being implemented / deployed as part of a wired and / or wireless communication network. For example, the emulation devices may be utilized in test scenarios in a test lab and / or in test scenarios in non-deployed (e.g., test) wired and / or wireless communication networks to implement testing of one or more components. The one or more emulation devices may be test equipment. Direct RF coupling and / or wireless communication via RF circuitry (which may, for example, include one or more antennas) may be used by the emulation devices to transmit and / or receive data.

[0064] Representative architectures / frameworks Video coding systems may be used to compress digital video signals, which may reduce storage requirements and / or transmission bandwidth of the video signal. Video coding systems may include block-based systems, wavelet-based systems, and / or object-based systems. Block-based video coding systems may be based on, use, follow, comply with, etc., one or more standards such as MPEG-1 / 2 / 4 Part 2, H.264 / MPEG-4 Part 10 AVC, VC-1, High Efficiency Video Coding (HEVC), and / or Versatile Video Coding (VVC). Block-based video coding systems may include block-based hybrid video coding frameworks.

[0065] 2 is a block diagram illustrating an example of a video streaming architecture 200. In this example, the server 202 may be composed of one or more video encoders (e.g., encoders 204, 206, and 208), each of which may generate a video bitstream at a different resolution, frame rate, or bitrate. A middlebox 210 may be used or configured. In one example, the middlebox 210 may be a media aware network element (MANE). The middlebox 210 may generate, forward, identify, or parse higher-level syntax of an input video bitstream, extract sub-bitstreams from one input video bitstream, and / or output the extracted sub-bitstreams to a client or decoder 212. The middlebox 210 may also extract multiple sub-bitstreams from multiple input video bitstreams and combine them to form a new output video bitstream for delivery to the client or decoder 212.

[0066] In various embodiments, one or more encoders (e.g., encoders 204, 206, and / or 208), middlebox 210, and / or decoder 212 may be implemented in a device having a processor communicatively coupled to a memory. The memory may include instructions executable by the processor, including instructions for performing any of the various embodiments (e.g., representative procedures) disclosed herein. In various embodiments, the device may be configured as and / or configured with various elements of a wireless transmit / receive unit (WTRU). Exemplary details of a WTRU and its elements are provided herein in FIGS. 1A-1D and the accompanying disclosure.

[0067] Various methods and aspects described herein may be used to modify modules, such as the intra-prediction module, entropy coding module, and / or decoding module, of one or more video encoders (e.g., encoders 204, 206, and 208) and decoder 212 shown in FIG. 2. Furthermore, the aspects are not limited to VVC or HEVC, but may be applied to other standards and recommendations, whether pre-existing or developed in the future, and to extensions of any such standards and recommendations (including VVC and HEVC). Unless otherwise indicated or technically precluded, the aspects described herein may be used individually or in combination.

[0068] Various numerical values ​​are used in this application. The particular values ​​are for illustrative purposes and the described aspects are not limited to these particular values.

[0069] Typical Procedure for Adaptive Resolution Change Various schemes using AVC and / or HEVC may not have the ability to change resolution without introducing Intra Random Access Point (IRAP) pictures. IRAP pictures coded with reasonable quality generally have much larger frame sizes (e.g., more bits used to code the frame) than non-IRAP pictures. Furthermore, IRAP pictures are more complex to decode. Adaptive resolution change (ARC) may refer to any scheme or function where spatial resolution can be changed in non-IRAP pictures.

[0070] 3 shows an example of an ARC mechanism 300. Referring to FIG. 3, high-resolution picture frame No. 3 may be coded with inter-prediction from reference picture frame No. 2 having the same resolution, and low-resolution picture frame No. 5 may be coded with inter-prediction from reference picture frame No. 4 having the same resolution. During ARC, high-resolution picture frame No. 3 may be reconstructed by reference picture frame No. 2, which is upscaled from low-resolution picture frame No. 2, for motion compensation, and low-resolution picture frame No. 5 may be reconstructed by reference picture frame No. 4, which is downscaled from high-resolution picture frame No. 4, for motion compensation. As a result, resolution switching may occur in or be implemented for one or more non-IRAP frames.

[0071] In various implementations, multi-party video conferencing may benefit from processing picture and / or video (“picture / video”) frames using, for example, ARC, where one or more or all participants (i.e., their pictures / videos) are displayed individually on a shared screen, with the active speaker (i.e., their picture / video) displayed at a larger video size than the remaining participants. If the active speaker changes frequently, ARC may be used (e.g., may be required to be used) to efficiently achieve frequent and / or unpredictable resolution changes by swapping in new active speakers and swapping out old ones. Current adaptive video streaming techniques typically change the video representation bit rate or resolution after an IRAP picture to match changing network bandwidth.

[0072] ARC may improve adaptive streaming performance by eliminating the need to send one or more high (or large) frame-size IRAP pictures. ARC may reduce streaming start latency because applications typically buffer up to a certain number of decoded pictures and / or a range of decoding times before displaying, e.g., in view of smaller-sized pictures. Under current motion constrained tile set (MCTS)-based viewport-adaptive 360-degree video streaming, subpictures representing the viewport are typically delivered using a high resolution, and subpictures representing other areas (e.g., areas outside the user's field of view) are typically delivered using a lower resolution. When the viewport changes, the corresponding resolution of the subpictures is changed accordingly, and the user experience is affected by the high-quality viewport switching latency.

[0073] FIG. 4 illustrates an example of a viewport adaptive streaming mechanism 400. A viewport subpicture (e.g., a near view) is extracted from a high-resolution 360-degree video (Representation No. 1), and subpictures representing other areas are extracted from a low-resolution 360-degree video (Representation No. 2). The extracted subpictures may be combined into a single representation (e.g., as shown by the sequence of frames at the bottom right of FIG. 4). The resulting composed or merged viewport-adapted video is delivered to a user (or client) so that the user can experience a high-quality viewport with reduced delivery bandwidth. When a user changes the viewport from the near view to the right view, a subpicture of the high-resolution right view is extracted onto the IRAP picture, forming a new video frame of the composed or merged video that incorporates the high-resolution right view subpicture and the low-resolution near subpicture. As a result, the length of the IRAP distance may affect high-quality viewport switching latency, and high-bitrate IRAP pictures may also increase network and / or processing load. ARC may enable faster viewport switching and may support different IRAP distances for different (e.g., 360-degree) video representations.

[0074] By rescaling reference pictures (e.g., using the techniques proposed in Non-Patent Document 1 and / or Non-Patent Document 2), pictures / video frames can be predicted across resolutions. A picture resolution index (PRI) can be signaled in a picture parameter set (PPS) to indicate that a slice (e.g., associated with a picture) is from that picture and has the resolution indicated by the index. It may be desirable to allow a sub-picture derived from a random access (e.g., IRAP) picture and another sub-picture derived from a non-random access (e.g., non-IRAP) picture to be merged into the same coded picture that complies with Generic Video Coding (VVC) (e.g., using the scheme proposed in Non-Patent Document 3).

[0075] Typical steps for switching viewports In the case of viewport adaptive streaming (e.g., associated with 360-degree video), a sub-bitstream corresponding to each subpicture may be extracted from its original bitstream and / or representation, and multiple sub-bitstreams may be merged to form a new bitstream. The original bitstream and / or new bitstream may be, for example, HEVC, VVC, and / or similar type bitstreams. In current viewport adaptive streaming, sub-bitstream merging is performed in the compressed domain, which may introduce some issues. For example, viewport switching is performed only when all included subpictures are instantaneous decoding refresh (IDR) pictures to ensure that (temporally) following non-IRAP pictures have the correct reference pictures and / or reference subpictures. However, the use of IDR pictures may introduce latency issues.

[0076] Figure 5 shows a viewport switching mechanism 500 in which the current viewport is switched to a new viewport. Referring to Figure 5, the right view subpicture is switched from a lower resolution to a higher resolution, and the foreground view subpicture is switched from a higher resolution to a lower resolution to match the new viewport. The dashed lines represent temporal inter-prediction. In the current viewport switching scheme, switching can only occur when the higher resolution right view subpicture and the lower resolution foreground view subpicture are both IDR pictures, because subsequent subpictures of the same view are inter-predicted from the same resolution subpicture. Because inter-prediction continues for subpictures with the same resolution, other subpictures, such as the top view, back view, left view, and bottom view, do not need to be coded as IRAP pictures.

[0077] According to the methods and / or techniques provided herein, ARC may be implemented in conjunction with adaptive viewport switching (and / or adaptive viewport streaming), for example, so that subpictures can be predicted at different resolutions from subpictures of the same view. FIG. 6 shows an example of a viewport switching mechanism 600 using ARC. Referring to FIG. 6, a subpicture of a high-resolution right view may be predicted from a subpicture of a previous low-resolution right view, and a subpicture of a low-resolution foreground view may be predicted from a subpicture of a previous high-resolution foreground view. By using the previous subpicture, the subpicture to be predicted can be predicted without introducing an IRAP subpicture. Both switching latency and transport bitrate can be reduced by not introducing an IRAP subpicture.

[0078] Currently, VVC does not specify the decoding process for such subpicture extraction and repositioning schemes, including when the same view subpicture may be packed into different positions (e.g., positions that change from time to time) within a picture. The motion vector of each subpicture may result in an offset from coordinates in the decoded subpicture to coordinates in the reference subpicture, and these coordinates shall be consistent (or at least may be assumed to be consistent) across pictures with the same resolution.

[0079]

[0013] Still referring to Figure 6, the reference subpicture may be scaled up or down to match the resolution of the current subpicture, but the coordinates of the reference subpicture and the current subpicture within the picture may be different. According to the methods and / or techniques provided herein, current decoding processes and associated signaling may be modified to achieve and / or implement the viewport adaptation techniques disclosed herein. The methods and / or techniques provided herein address shortcomings in current signaling associated with decoding, including the lack of signaling or metadata for identifying sets or groups of subpictures of multiple representations that map to the same two-dimensional (2D) or three-dimensional (3D) content region, for example, at the elementary bitstream level or system level.

[0080] Several supplemental enhancement information (SEI) messages specify rectangular regions and 360-degree video information in HEVC. The pan-scan rectangle SEI message specifies the coordinates of one or more rectangular areas relative to the conformance cropping window specified by the active SPS. The equirectangular projection and cube-map projection SEI messages provide information to enable remapping of color samples of a projected picture onto a spherical coordinate space (e.g., addressed using spherical coordinates) to support 360-degree video pictures. The region-wise packing SEI message provides information to enable remapping of color samples of a cropped decoded picture onto a projected picture, as well as information about the location and size of guard bands, if any. However, all of these SEI messages are designed for a single representation (or layer) and, as a result, do not address the relationship between sub-pictures across multiple representations with different resolutions.

[0081] Representative Procedure for Subpicture-Based Applications For some sub-picture-based applications, each extracted sub-picture may be assigned to a different position within the new picture. As used herein, a sub-picture characteristics SEI message refers to an SEI message containing sub-picture characteristic information that may indicate (and / or define characteristics of or indicate defined characteristics of) one or more sub-pictures or tile groups across multiple layers or representations associated with the same source content region. In one embodiment, the sub-picture characteristic information may include one or more additional indicators for indicating one or more recommended ARC switching points. Alternatively, the sub-picture characteristics SEI message or a similar type of SEI message may include an indicator for indicating one or more recommended ARC switching points. In one embodiment, the sub-picture characteristic information may include one or more actual ARC switching points, for example, to achieve better reconstructed picture quality or to apply some constraints. Alternatively, the sub-picture characteristics SEI message or a similar type of SEI message may include one or more actual ARC switching points.

[0082] Subpicture Characteristics SEI Message In one embodiment, video content may be encoded into multiple coded versions or representations (or layers). Each representation may be coded at a different resolution and / or quality. In the case of 360-degree video, for example, each representation may be in a different projection and / or per-region packing format. An original content region may be mapped to a different portion of the representation, i.e., a sub-picture. The sub-picture may be rotated, scaled, or projected differently in different representations. The position and size of a sub-picture corresponding to the same content region may also vary across different representations. A middlebox or client may fetch one or more sub-pictures across the representations. The middlebox or client may use multiple fetched sub-pictures to form a new picture for viewport-dependent streaming applications. The new picture may be a combination of different sub-pictures at different quality levels and / or resolutions. The formation of a new picture may be performed, for example, to meet the streaming needs of a (e.g., 360-degree) video client.

[0083] In one embodiment, a content producer may generate multiple representations. All representations associated with the same content may be arranged in a multi-layer structure. In one embodiment, each representation may be a layer. Each layer may be coded independently or may depend on other layers. A layer ID may be used to identify a particular coded video representation. Each layer may have multiple sub-pictures. Each sub-picture may be identified by a unique sub-picture ID or tile group ID. From the projection and per-region packing SEI messages associated with each representation, it is possible to derive the correspondence between the sub-pictures and the original content region (e.g., the region 360-degree content sphere). However, doing so would require the middlebox or client to parse multiple SEI messages from each layer, and this derivation process may increase the workload of the middlebox or client. A single SEI message or parameter set that represents subpicture characteristics (e.g., correspondence between subpictures across multiple layers and mapping of subpictures within each layer to regions on a (e.g., 360-degree) video sphere) may simplify the mapping between subpictures and corresponding original content regions and facilitate subpicture-based applications.

[0084] In one embodiment, the subpicture partition layout may be signaled within the PPS. The subpicture resolution may be explicitly signaled (e.g., in the PPS or tile group header). Alternatively, the subpicture resolution may be derived from the tile group layout and overall picture resolution. To map one (e.g., each) subpicture to the corresponding spherical space, the subpicture characteristics SEI message may include and / or provide information such as a layer ID, a tile group ID, the coordinates of the subpicture, and its mapping to the spherical coordinate space. The SEI message may list any or all subpictures available for viewport adaptive streaming and subpicture region-by-region packing. The decoder may identify corresponding reference subpictures that may or may not be collocated with the current subpicture based on such an SEI message. Depending on the resolution of the reference subpicture, the decoder may scale the subpicture for ARC and align coordinates between the current subpicture and the reference subpicture.

[0085] Table 1 provides an example of a Subpicture Characteristics SEI message syntax structure. Table 1 lists the spherical area (viewport) and the number of subpictures that cover the same spherical area. Table 1 also provides the coordinates of each repositioned subpicture relative to the conformance cropping window specified by the active sequence parameter set (SPS).

[0086] [Table 1]

[0087] In Table 1, num_source_content_regions_minus1 plus 1 may specify the number of source content regions specified by the SEI message.

[0088] In Table 1, source_content_region_position may specify the position of the ith source content region. For 2D source content, it may be the top-left position of the region in 2D coordinates. For 360-degree video content, it may be the sphere-center azimuth and tilt position.

[0089] In Table 1, source_content_region_size may specify the size of the ith source content region. For 2D source content, it may be the width and height of the region. For 360-degree video content, it may be the tilt angle relative to the global coordinate axes, the azimuth angle range and elevation angle range of the spherical region passing through the center point of the spherical region in degrees.

[0090] In Table 1, num_subpics_minus1[i] plus 1 may specify the number of subpictures associated with the i-th source content region.

[0091] In Table 1, layer_id[i][j] may specify the layer identifier to which the jth subpicture associated with the ith source content region belongs.

[0092] In Table 1, subpic_coordinate[i][j] may specify the coordinate of the jth subpicture associated with the ith source content region in the picture. It can be the top-left position of the subpicture or the center position of the subpicture.

[0093] In Table 1, subpic_id[i][j] may specify the subpicture ID of the jth subpicture associated with the ith source content region. The subpicture ID and / or tile group ID may be used to identify the subpicture.

[0094] In Table 1, subpic_width[i][j] and subpic_height[i][j] may specify the resolution of the jth subpicture associated with the ith source content region.

[0095] In various embodiments, characteristics such as bit depth, color subsampling, coding profile, and / or coding level for each subpicture or group of subpictures may be included in the subpicture characteristics SEI message.

[0096] In various embodiments, the subpicture characteristics SEI message may indicate the number of subpictures associated with the same source content region among multiple representations or layers. The decoder may determine corresponding reference subpictures available in the previous picture and derive the reference subpicture position and size from the active PPS to implement ARC. The disclosed SEI message may explicitly signal the position and size of each subpicture, for example, to simplify derivation so that the decoder does not have to parse the parameter set or SEI message of each representation or layer. The disclosed syntax elements of the subpicture characteristics SEI message may also be implemented by a cross-layer parameter set, such as a video parameter set (VPS) or a decoder parameter set (DPS).

[0097] In various embodiments, a middlebox may rely on the disclosed SEI message. For example, the middlebox may extract a subpicture that matches the viewport based on the SEI message. The middlebox may form an ARC picture using the extracted subpicture, for example, to reduce frame size (e.g., the number of bits required to code the frame). A client may rely on the proposed SEI message. For example, the client may identify an ARC subpicture and an associated reference subpicture based on the SEI message. The client may align coordinates between the ARC subpicture and the reference subpicture for a proper motion compensation process.

[0098] FIG. 7 shows an example of a subpicture-based ARC mechanism 700. Referring to FIG. 7, each subpicture in a cube map projection format may be coded into two resolutions. The subpicture that matches the viewport may be extracted from the high-resolution representation. The remaining subpictures may be extracted from the low-resolution representation. In various embodiments, the subpicture characteristics SEI message may indicate subpictures associated with the same source content region. For example, tile group No. 0 and tile group No. 6 both cover the left side, and tile group No. 1 and tile group No. 7 both cover the near side, but at different resolutions and in different representations.

[0099] When the extractor extracts a non-IRAP high-resolution subpicture to match a viewport change, the extractor may signal an ARC occurrence in a subpicture characteristics SEI message. Because subpicture No. 2 (right) and subpicture No. 7 (foreground) were not available in the previous picture, the decoder may find that the right subpicture and the background subpicture are both ARC subpictures. To reconstruct the ARC subpicture, the decoder may parse the SEI message and identify subpictures associated with the same region as tile group No. 2 and tile group No. 7. For example, tile group No. 1 may be associated with the same content region as tile group No. 7 and is available in the previous decoded picture. Tile group No. 8 may be associated with the same content region as tile group No. 2 and is available in the previous decoded picture. Based on the subpicture positions and sizes derived from the PPS, the decoder may scale down decoded subpicture No. 1 and scale up decoded subpicture No. 8 for ARC motion compensation. The motion vector of each sample in subpicture No. 2 may (e.g., shall) be shifted by an offset between the coordinates of subpicture No. 2 and the coordinates of subpicture No. 8, for example, as marked (dMVx, dMVy) in Figure 7. The offset may be in units of 1 / 16 sample intervals relative to the luma sampling grid, for example. The same motion vector shift may apply to each sample in tile group No. 7.

[0100] In one embodiment, all subpictures associated with the same source content region may be assigned to a subpicture group, and each subpicture group may have its own unique ID. The subpicture group ID and its characteristics, such as region position and size, may be carried in the tile group header or in the PPS tile group layout syntax field.

[0101] In one embodiment, the SEI characteristics message may signal characteristics of the current subpicture and its associated reference subpictures during ARC, including identifier, coordinates, subpicture size, bit depth, chroma subsampling, projection format, per-region packing, etc. Table 2 provides an example of an ARC subpicture characteristics SEI message.

[0102] [Table 2]

[0103] In Table 2, num_arc_subpics_minus1 plus 1 may specify the number of subpictures associated with the adaptive resolution change.

[0104] In Table 2, arc_subpic_id[i] and reference_subpic_id[i] may specify the identifiers of the i-th ARC subpicture and associated reference subpicture.

[0105] In Table 2, arc_subpic_coordinate[i] and arc_subpic_size[i] may specify the coordinates (eg, center or top-left sample position) and size of the i-th ARC subpicture.

[0106] In Table 2, reference_subpic_coordinate[i] and reference_subpic_size[i] may specify the coordinates (eg, center or top-left sample position) and size of the reference subpicture for the i-th ARC subpicture.

[0107] In various embodiments, the ARC subpicture SEI message may include information indicating the format of the ARC subpicture and associated reference subpictures, including bit depth, color subsampling, coding profile, and / or coding level.

[0108] In various embodiments, the ARC subpicture SEI message may include information indicating the coded subpicture after ARC and the corresponding reference subpicture before ARC. Based on such an indication, the client can scale or transform (e.g., mirror, scale, or rotate) the corresponding reference picture for motion compensation. This message may be (e.g., typically) generated by a middlebox or extractor that extracts and repositions the subpicture bitstream. This message may be attached (e.g., shall be attached) to the ARC picture. Based on such a message, the client or decoder may identify the ARC subpicture and the associated reference subpicture and align coordinates between the two subpictures for motion compensation.

[0109] Recommended ARC Transition SEI Messages In various embodiments, ARC may use scaled, reconstructed pictures as reference pictures, for example, to avoid high-bitrate IDR or IRAP pictures and / or to maintain acceptable decoded picture quality. There may be multiple versions of coded video of the same content. ARC transition performance, including reconstructed picture quality and error propagation, may depend on the scaling filter design, reference picture, and temporal layer.

[0110] 8 shows a temporal scalability mechanism 800 in which ARC transition on POC No. 4 may not cause error propagation. Depending on the correlation of subpictures between different versions and coding structures, the encoder may find the best ARC transition point. For example, the encoder may simulate ARC within each IRAP interval or group of pictures (GOP) interval to determine the transition point for ARC.

[0111] 9 shows an example for determining an optimal ARC transition point. In addition to traditional temporal inter-prediction and motion compensation, the encoder may simulate reconstructing a picture by referencing scaled pictures from different representations, calculate PSNR, and mark the picture that achieves the highest peak signal-to-noise ratio (PSNR) as the recommended ARC transition point.

[0112] In one embodiment, to determine the ARC transition point (e.g., POC No. 3), mechanism 900 may periodically perform inter-layer prediction, as specified in SHVC, between the two representations. This additional inter-layer coded data may be delivered to the end user when ARC is implemented.

[0113] 10 shows an example mechanism 1000 in which high-resolution representation pictures are periodically predicted from scaled low-resolution IRAP pictures. These pictures are marked as ARC transition pictures. The transition from low-resolution to high-resolution representations can only occur at these ARC transition points, while the transition from high-resolution to low-resolution representations can occur at any low-resolution IRAP picture because the size of the low-resolution IRAP picture is much smaller than the high-resolution IRAP picture size. When implementing ARC, low-resolution IRAP pictures can be carried in the bitstream as inter-layer reference pictures for high-resolution pictures.

[0114] 10, bitstream No. 1 may be a low-resolution coded stream temporally predicted with periodic IRAP pictures, and bitstream No. 2 may be a high-resolution coded stream temporally predicted with ARC pictures inter-layer predicted from the low-resolution IRAP pictures. The ARC transition bitstream may include a low-resolution reference picture (POC No. 4), an inter-layer predicted ARC picture (POC No. 4), and a subsequent temporally predicted high-resolution picture.

[0115] In various embodiments, the recommended ARC transition SEI message may be included in the sub-picture characteristics SEI message or used or transmitted separately. The recommended ARC transition SEI message may include a POC value for the high-resolution representation, where the POC value is that of a reference picture of a corresponding lower layer with a representation ID such as a layer ID or tile group ID, scaling filter coefficients, and a prediction method (e.g., temporal prediction or inter-layer prediction).

[0116] Table 3 provides an example syntax for the recommended ARC transition SEI message. This message identifies the recommended ARC subpicture by its layer ID, POC value, and subpicture ID. This message also indicates multiple subpictures recommended to switch from. The scaling filter may support customized scaling filters to improve ARC quality.

[0117] [Table 3]

[0118] In Table 3, arc_layer_id may specify the identifier of the layer with which the ARC subpicture is associated.

[0119] In Table 3, arc_poc_lsb may specify the POC LSB value of the ARC subpicture.

[0120] In Table 3, arc_subpic_id may specify the identifier of the ARC subpicture.

[0121] In Table 3, num_ref_subpic_minus1 plus 1 may specify the recommended number of subpictures to switch from during ARC.

[0122] In Table 3, reference_subpic_id[i] may specify the i-th recommended subpicture to switch from during ARC.

[0123] In Table 3, scaling_filter can be a structure containing the recommended scaling filter coefficients.

[0124] In various embodiments, the SEI message may provide multiple recommended ARC transition points and prioritize these points. A priority indicator may be signaled in the SEI message to indicate the priority of the sub-picture for which ARC is to be implemented. A higher priority indicates that better ARC performance may be achieved for the associated sub-picture.

[0125] ARC Toggle Indicator To indicate the occurrence of ARC to the decoder, an indicator may be used and signaled in an SEI message, PPS, or sub-picture-related parameter set. This indicator may carry parameters such as layer ID and POC number. Because some pictures after the ARC switch point will use scaled pictures as references, motion compensation errors may reduce the performance of those coding tools that rely on accurate temporal information and reference samples. For example, bidirectional optical flow (BDOF) requires deriving a delta motion vector for each 4x4 sub-block based on a temporal prediction signal using the difference between two temporal predictions from forward and backward predictions, as well as spatial gradients. Both the temporal difference and spatial gradients will change after the spatial scaling of the reference picture after the ARC switch point, and the derivation results will change significantly. Decoder-side motion vector refinement (DMVR) derives a delta motion vector for each sub-block (e.g., 16x16) in a coding unit using two temporal predictions from two reference pictures. The derivation results will change if the temporal reference picture is scaled. Decoder-side derivation related coding techniques such as BDOF and DMVR shall be disabled in the ARC picture and all subsequent pictures preceding the next IRAP picture in decoding order. In various embodiments, these coding techniques may be disabled based on an indicator signaled in the ARC-related SEI message, or they may be adaptively disabled based on the scaling ratio. If the scaling ratio is close to 1, they may be enabled; otherwise, if the scaling ratio is far from 1, they may be disabled. Disabling these coding tools may further reduce decoder workload and power consumption without affecting coding performance.

[0126] Temporal motion vector prediction (TMVP) and sub-block temporal motion vector prediction (SbTMVP) may also be affected by motion vector scaling due to resolution changes, and motion vector error may be propagated within a picture by motion vector prediction. Motion vector error has a significant impact on picture quality due to motion compensation. In various embodiments, when encoding pictures such as interpictures after an ARC switch point that may be affected by reference picture scaling and motion vector scaling, the encoder may disable these decoder-side derived related coding techniques, TMVP and SbTMVP, so that picture quality is not too affected after the ARC switch.

[0127] In various embodiments, an indication or flag (e.g., a constraint flag) may be used to indicate whether one or more coding tools (e.g., a predetermined coding tool) will be disabled at an ARC transition point. In one example, a decoder may detect that a constraint flag is set in an SPS, indicating that one or more coding tool constraints apply to the entire coded video sequence (CVS). For example, layer-based spatial scalability may allow each high-resolution enhancement picture to be predicted from a lower-resolution (e.g., base layer) picture, and one or more coding tool constraints apply to each enhancement layer (e.g., a layer other than the base layer, which is the first layer in the bitstream) frame. In the case of a single-layer sequence, the encoder may set a constraint flag for some or predetermined frames for ARC transition. The decoder may detect the flag at the slice level and perform ARC on the associated frame (e.g., the frame containing the slice).

[0128] Table 4 provides an example syntax for an SPS, including the constraint flag sps_arc_constraint_flag that is signaled within the SPS.

[0129] [Table 4]

[0130] In Table 4, sps_arc_constraint_flag is a constraint flag. When the value of sps_arc_constraint_flag is equal to 1, it indicates or specifies that for the CVS associated with the SPS, the coding tools TMVP, DMVR, and / or BDOF shall be disabled (e.g., slice_temporal_mvp_enabled_flag, bdofFlag, and / or dmvrFlag having values ​​equal to 0). When sps_arc_constraint_flag has a value equal to 0, the flag indicates or specifies that no constraints are imposed.

[0131] Table 5 provides an example syntax for a slice header. The constraint flag slice_arc_constraint_flag is signaled in the slice header.

[0132] [Table 5]

[0133] In Table 5, slice_arc_constraint_flag is a constraint flag. When the value of slice_arc_constraint_flag is equal to 1, it indicates or specifies that for the associated slice, the coding tools TMVP, DMVR, and / or BDOF shall be disabled (e.g., slice_temporal_mvp_enabled_flag, bdofFlag, and / or dmvrFlag of the associated slice having values ​​equal to 0). When slice_arc_constraint_flag has a value equal to 0, the flag indicates or specifies that no constraint is imposed. When slice_arc_constraint_flag is not present, it is inferred to have a value equal to 1.

[0134] In various embodiments, a coding tool constraint (e.g., signaled using one or more constraint flags described above) may indicate (or provide reference to) a known or fixed set of tools that will be disabled after an ARC transition. For example, a coding tool constraint may disable BDOF and DMVR after an ARC transition and / or until the next IRAP picture. In another embodiment, additional signaling (e.g., one or more additional flags) may be provided within the syntax to adaptively specify which coding tools are disabled when a coding tool constraint is active. For example, a first flag may specify whether the coding tool constraint will disable BDOF, a second flag may specify whether the coding tool constraint will disable DMVR, and one or more additional flags may specify whether the coding tool constraint will disable other tools.

[0135] In various embodiments, Table 6 provides an example syntax for a VPS in which the all-layer-independent flag is used to indicate that each layer is coded independently and, as a result, dependencies between layers do not need to be specified.

[0136] [Table 6]

[0137] In Table 6, vps_all_layers_independent_flag equal to 1 may specify that each layer (specified by the VPS) is an independent layer, and vps_all_layers_independent_flag equal to 0 may specify that one or more layers (specified by the VPS) may not be independent layers. When vps_all_layers_independent_flag is set to 1, layer dependency signaling (e.g., direct_dependency_flag) will not need to be signaled.

[0138] In various embodiments, methods and apparatuses are provided for picture and / or video coding in a communication system. For example, the method may include selectively including sub-picture characteristic information in an SEI message for use with adaptive switching of viewports and generating / transmitting the SEI message. The method may also include identifying a set of sub-pictures associated with the viewport and identifying sub-picture characteristic information associated with the identified set of sub-pictures. In one example, selectively including the sub-picture characteristic information in the SEI message may include generating an SEI message that includes the identified sub-picture characteristic information.

[0139] In various embodiments, the subpicture characteristic information indicates and / or includes one or more layer identification information (IDs), one or more tile group IDs, coordinates of each subpicture, the position and format of each subpicture in the set of subpictures, mapping information for each subpicture to be mapped onto the spherical coordinate space of the picture, bit depth, color subsampling information, an encoding profile, and an encoding level for one or more subpictures.

[0140] In various embodiments, the method may include generating syntax to indicate sub-picture characteristic information in an SEI message.

[0141] In various embodiments, the set of sub-pictures may be a set of tile groups associated with a picture.

[0142] In various embodiments, the method may include selectively including subpicture characteristic information in an SEI message for use with ARC in connection with adaptive switching of viewports, and generating / transmitting the SEI message. In one example, the subpicture characteristic information may include information indicating which of a low-resolution subpicture and a high-resolution subpicture to which ARC may be applied for adaptive switching of viewports. In one example, the low-resolution subpicture and the high-resolution subpicture are associated with the same source content region used for the viewport.

[0143] In various embodiments, the SEI message may be transmitted after the frame containing the first IRAP picture and before the next frame containing the second IRAP picture.

[0144] In various embodiments, any of the methods discussed herein may include identifying a set of subpictures associated with a picture available for ARC, the set of subpictures including a low-resolution subpicture and a high-resolution subpicture, and may also include identifying subpicture characteristic information associated with the identified set of subpictures. In one example, selectively including the subpicture characteristic information in an SEI message may include generating an SEI message including the identified subpicture characteristic information.

[0145] In various embodiments, a method may include receiving an SEI message including sub-picture characteristic information for use with adaptive switching of viewports, and performing ARC based on the sub-picture characteristic information.

[0146] In various embodiments, the subpicture characteristic information may include information indicating which of the low-resolution and high-resolution subpictures ARC may be applied to for adaptive switching of viewports.

[0147] In various embodiments, ARC may be implemented, at least in part, by matching a presented low-resolution subpicture to a high-resolution subpicture and / or by matching a presented high-resolution subpicture to a low-resolution subpicture.

[0148] In various embodiments, performing ARC may include performing ARC in conjunction with adaptive switching of viewports based on subpicture characteristic information.

[0149] In various embodiments, ARC can be implemented without introducing IRAP picture frames.

[0150] In various embodiments, ARC may be implemented at least in part by matching presented low-resolution sub-pictures to high-resolution sub-pictures, and matching presented high-resolution sub-pictures to low-resolution sub-pictures.

[0151] In various embodiments, any of the methods discussed herein may include decoding the received SEI message.

[0152] In various embodiments, implementing ARC may include scaling up the low-resolution sub-picture based on the sub-picture characteristic information.

[0153] In various embodiments, implementing ARC may include scaling down the high resolution sub-picture based on the sub-picture characteristic information.

[0154] In various embodiments, performing ARC may include aligning coordinates between a set of subpictures and a corresponding set of reference subpictures based on subpicture characteristic information.

[0155] In various embodiments, the steps of performing ARC may include determining the position and format of the subpicture based on the subpicture characteristic information, determining the resolution of the corresponding reference subpicture, scaling the resolution of the subpicture based on the determined position and format and the resolution of the reference subpicture, and mapping the scaled subpicture to corresponding coordinates of the picture.

[0156] In various embodiments, a sub-picture may include a group of tiles.

[0157] In various embodiments, any of the methods discussed herein may include determining, based on the subpicture characteristic information, whether a corresponding reference subpicture is collocated with the subpicture.

[0158] In various embodiments, the SEI message may include syntax with sub-picture characteristic information.

[0159] In various embodiments, the syntax may be carried by a cross-layer parameter set.

[0160] In various embodiments, the subpicture characteristic information may indicate the coordinates of each repositioned subpicture relative to the conformance cropping window.

[0161] In various embodiments, any of the methods discussed herein may include determining corresponding reference subpictures available within one or more previous pictures, and deriving the location and format of the corresponding reference subpictures based on the subpicture characteristic information.

[0162] In various embodiments, the subpicture characteristic information may indicate characteristics of a set of subpictures that includes one or more low-resolution subpictures and one or more high-resolution subpictures.

[0163] In various embodiments, the method may include identifying an ARC switch point at which one or more pictures after the ARC switch point use a scaled picture as a reference, and sending an indicator indicating the ARC switch point.

[0164] In various embodiments, the indicator is transmitted in one of the SEI message, the picture parameter set (PPS), and the sub-picture related parameter set.

[0165] In various embodiments, the indicator includes a parameter, which includes one of a layer ID and a picture order count (POC) value.

[0166] In various embodiments, the method may include receiving an indicator indicating an ARC switch point and identifying the ARC switch point and a scaled picture. In various embodiments, one or more pictures received after the ARC switch point may use the scaled picture as a reference.

[0167] In various embodiments, the method may include identifying a set of sub-pictures associated with a picture that are usable or recommended for implementing ARC, and generating / sending an SEI message indicating one or more parameters of the set of sub-pictures.

[0168] In various embodiments, the one or more parameters discussed above may include a POC value for the high-resolution representation of the picture, a POC value of a corresponding lower-resolution reference picture having a representation ID, one or more scaling filter coefficients, and any of the prediction methods.

[0169] In various embodiments, the representation ID may be a layer ID or a tile group ID.

[0170] In various embodiments, prediction methods may include either temporal prediction or inter-layer prediction.

[0171] In various embodiments, the syntax may indicate information for one subpicture of a set of subpictures, where the information includes any of a layer ID, a POC value, and a subpicture ID associated with the subpicture.

[0172] In various embodiments, the method may include receiving an SEI message associated with a picture, identifying a set of subpictures recommended for implementing ARC based on one or more parameters indicated in the SEI message, and selecting one or more subpictures from the set of subpictures for implementing ARC.

[0173] In various embodiments, the SEI message may include a priority indicator that indicates the priority of each subpicture in the set of subpictures.

[0174] In various embodiments, the priorities discussed herein may be higher priorities indicating better ARC performance for the respective sub-pictures.

[0175] In various embodiments, an apparatus may comprise one or more processors, encoders, decoders, transmitters, receivers, and / or memories that implement or perform any of the methods discussed herein.

[0176] In various embodiments, the middlebox may comprise one or more processors, encoders, decoders, transmitters, receivers, and / or memories that implement or perform any of the methods discussed herein.

[0177] Each of the following references is incorporated herein by reference: [1] Non-Patent Document 1; [2] Non-Patent Document 2; [3] Non-Patent Document 3; [4] Non-Patent Document 4; and [5] U.S. Provisional Patent Application No. 62 / 775,130.

[0178] Conclusion While features and elements are described above in particular combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with the other features and elements. Furthermore, the methods described herein may be implemented in a computer program, software, or firmware embodied in a computer-readable medium for execution by a computer or processor. Examples of non-transitory computer-readable storage media include, but are not limited to, read-only memory (ROM), random-access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs). A processor in association with software may be used to implement a radio frequency transceiver for use in the WTRU 102, a UE, a terminal, a base station, an RNC, or any host computer.

[0179] Additionally, in the above embodiments, processing platforms, computing systems, controllers, and other devices including processors are described. These devices may include at least one central processing unit (CPU) and memory. In accordance with the practices of those skilled in the art of computer programming, references to acts and symbolic representations of operations or instructions may be performed by various CPUs and memories. Such acts and operations or instructions may be referred to as being "executed," "executed by a computer," or "executed by a CPU."

[0180] Those skilled in the art will understand that the acts and symbolically represented operations or instructions include the manipulation of electrical signals by a CPU. The electrical system represents the data bits, which may cause the transformation or reduction of the resulting electrical signals and the retention of the data bits in memory locations within a memory system, thereby reconfiguring or altering the operation of the CPU and other processing of the signals. The memory locations in which the data bits are maintained are physical locations with specific electrical, magnetic, optical, or organic properties that correspond to or represent the data bits. It should be understood that exemplary embodiments are not limited to the above platforms or CPUs, and that other platforms and CPUs may support the provided methods.

[0181] Data bits may also be maintained on computer-readable media, including magnetic disks, optical disks, and other volatile (e.g., random access memory (RAM)) or non-volatile (e.g., read-only memory (ROM)) CPU-readable mass storage systems. The computer-readable media may include cooperative or interconnected computer-readable media that reside exclusively on a processing system or that are distributed among multiple interconnected processing systems, which may be local or remote to the processing system. Representative embodiments are not limited to the above memories, and it will be understood that other platforms and memories may support the described methods.

[0182] In an exemplary embodiment, any of the operations, processes, etc. described herein may be implemented as computer-readable instructions stored on a computer-readable medium, which may be executed by a processor of a mobile unit, a network element, and / or any other computing device.

[0183] There is little difference between hardware and software implementations of system aspects. The use of hardware or software is generally (though not always, the choice between hardware and software may be significant in certain contexts) a design choice that represents a trade-off between cost and efficiency. There may be various mediums (e.g., hardware, software, and / or firmware) in which the processes and / or systems and / or other techniques described herein may be affected, and the preferred medium may vary depending on the context in which the processes and / or systems and / or other techniques are deployed. For example, if an implementer determines that speed and accuracy are paramount, the implementer may select a primarily hardware and / or firmware medium. If flexibility is paramount, the implementer may select a primarily software implementation. Alternatively, the implementer may select some combination of hardware, software, and / or firmware.

[0184] The foregoing detailed description has illustrated various embodiments of devices and / or processes through the use of block diagrams, flowcharts, and / or examples. To the extent that such block diagrams, flowcharts, and / or examples include one or more functions and / or operations, it will be understood by those skilled in the art that each function and / or operation within such block diagrams, flowcharts, or examples, individually and / or collectively, can be implemented by a wide range of hardware, software, firmware, or virtually any combination thereof. Suitable processors include, for example, general-purpose processors, special-purpose processors, conventional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors associated with a DSP core, controllers, microcontrollers, application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), field-programmable gate array (FPGA) circuits, other types of integrated circuits (ICs), and / or state machines.

[0185] Although features and elements have been described above in particular combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with the other features and elements. The present disclosure should not be limited in terms of the specific embodiments described herein, which are intended as illustrations of various aspects. As will be apparent to those skilled in the art, many modifications and variations can be made without departing from its spirit and scope. No element, act, or instruction used in the description of the present application should be construed as critical or essential to the invention unless explicitly provided as such. Functionally equivalent methods and apparatuses within the scope of the present disclosure, in addition to those enumerated herein, will be apparent to those skilled in the art from the foregoing description. Such modifications and variations are intended to be included within the scope of the appended claims. The present disclosure should be limited only by the terms of the appended claims, and the full scope of equivalents to which such claims are entitled. It is understood that this disclosure is not limited to any particular method or system.

[0186] It should also be understood that the terminology used herein is for the purpose of describing particular aspects only and is not intended to be limiting. As used herein, the terms “station” and its abbreviation “STA,” “user equipment” and its abbreviation “UE,” when referred to herein, mean (i) a wireless transmit / receive unit (WTRU), as described below; (ii) any of several embodiments of a WTRU, as described below; (iii) a wireless-enabled and / or wired-enabled (e.g., tetherable) device configured with some or all of the structure and functionality of a WTRU, as described below; (iv) a wireless-enabled and / or wired-enabled device configured with less than all of the structure and functionality of a WTRU, as described below; or (v) the like. Details of an exemplary WTRU, which can represent (or be interchangeable with) any UE or mobile device listed herein, are provided below with respect to FIGS. 1A-1D.

[0187] In certain representative embodiments, portions of the subject matter described herein may be implemented via application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), and / or other integrated formats. However, those skilled in the art will understand that some aspects of the embodiments disclosed herein may be implemented, in whole or in part, within an integrated circuit as one or more computer programs running on one or more computers (e.g., multiple programs running on one or more computer systems), as one or more programs running on one or more processors (e.g., one or more programs running on one or more microprocessors), as firmware, or virtually any combination thereof, and that designing circuitry and / or writing code for software and / or firmware is within the skill of those skilled in the art in light of this disclosure. Furthermore, those skilled in the art will understand that the mechanisms of the subject matter described herein can be distributed as program products in various forms, and that exemplary embodiments of the subject matter described herein apply regardless of the particular type of signal-bearing medium used to actually perform the distribution. Examples of signal bearing media include, but are not limited to, recordable-type media such as floppy disks, hard disk drives, CDs, DVDs, digital tape, computer memory, and transmission-type media such as digital and / or analog communications media (e.g., fiber optic cables, wave guides, wired communications links, wireless communications links, etc.).

[0188] The subject matter described herein may depict different components contained within or connected to different other components. It should be understood that such depicted architectures are merely examples, and that in fact many other architectures that achieve the same functionality may be implemented. In a conceptual sense, an arrangement of components to achieve the same functionality is effectively “associated” such that the desired functionality is achieved. Thus, any two components herein that are combined to achieve a particular function may be considered to be “associated” with each other such that the desired functionality is achieved, regardless of the architecture or intermediate components. Similarly, any two components so associated may also be considered to be “operably connected” or “operably coupled” to each other to achieve the desired functionality, and any two components that may be so associated may also be considered to be “operably couplable” with each other to achieve the desired functionality. Specific examples of operably couplable components include, but are not limited to, physically coupled and / or physically interacting components, wirelessly interacting and / or wirelessly interacting components, and / or logically interacting and / or logically interacting components.

[0189] With respect to the use of virtually any plural and / or singular term herein, those skilled in the art can translate from plural to singular and / or from singular to plural as appropriate to the context and / or application. For clarity, various singular / plural permutations may be expressly set forth herein.

[0190] In general, terms used herein, and particularly in the appended claims (e.g., the appended claims), are generally intended as "open" terms (e.g., the term "comprising" should be interpreted as "including but not limited to," the term "having" should be interpreted as "having at least," the term "including" should be interpreted as "including but not limited to," etc.). Where a specific number of introduced claim recitations are intended, such intention will be explicitly set forth in the claim; in the absence of such a recitation, it will further be understood by those skilled in the art that no such intention exists. For example, where only one item is intended, the term "single" or similar language may be used. However, the use of such phrases should not be construed as meaning that introducing a claim recitation with the indefinite article "a" or "an" limits a particular claim containing such introduced claim recitation to embodiments containing only one such recitation. The same claim may include the introductory phrase "one or more" or "at least one" and an indefinite article such as "a" or "an" (e.g., "a" and / or "an" should be interpreted as "at least one" or "one or more"). This also applies to the use of definite articles used to introduce claim recitations. Moreover, even when a specific number of introduced claims is explicitly recited, those skilled in the art will recognize that such recitation should be interpreted to mean at least the recited number (e.g., the literal recitation of "two recitations" without other modifiers means at least two recitations, or more than two recitations).

[0191] Furthermore, when a convention similar to "at least one of A, B, C, etc." is used, generally such a configuration is intended in the sense that one of ordinary skill in the art would understand the convention (e.g., "a system having at least one of A, B, and C" includes, but is not limited to, systems of A only, B only, C only, A and B together, A and C together, B and C together, and / or A, B, C together). When a convention similar to "at least one of A, B, or C, etc." is used, generally such a configuration is intended in the sense that one of ordinary skill in the art would understand the convention (e.g., "a system having at least one of A, B, or C" includes, but is not limited to, systems of A only, B only, C only, A and B together, A and C together, B and C together, and / or A, B, C together). It will be further understood by those skilled in the art that virtually all disjunctive words and / or phrases presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to contemplate the possibility of including one of the terms, either term, or both terms. For example, the phrase "A or B" would be understood to include the possibilities of "A" or "B" or "A and B." Furthermore, as used herein, the term "any of" following a list of multiple items and / or multiple categories of items is intended to include "any of," "any combination," "any plurality," and / or "any combination of multiple" of the items and / or categories of items, individually or in combination with other items and / or items from other categories. Furthermore, as used herein, the term "set" or "group" is intended to include any number of items, including zero. Furthermore, as used herein, the term "number" is intended to include any number, including zero.

[0192] Furthermore, when features or aspects of the disclosure are described in terms of a Markush group, those skilled in the art will recognize that the disclosure is also thereby described in terms of any individual members or subgroups of members of the Markush group.

[0193] As will be understood by those skilled in the art, for all purposes, including with respect to providing a written description, all ranges disclosed herein also encompass any and all possible subranges and combinations of subranges. Any listed range can be readily recognized as sufficiently descriptive so that the same range can be divided into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each range discussed herein can be readily broken down into lower, middle, and upper thirds, etc. As will be understood by those skilled in the art, all language, such as "up to," "at least," "greater than," and "less than," refers to a range that is inclusive of the recited numbers and can then be broken down into subranges as described above. Finally, as will be understood by those skilled in the art, a range includes individual members. Thus, for example, a group having 1 to 3 cells refers to a group having 1, 2, or 3 cells. Similarly, a group having 1 to 5 cells refers to a group having 1, 2, 3, 4, or 5 cells, etc.

[0194] Furthermore, the claims should not be construed as limited to the order or elements provided unless stated to that effect. Furthermore, the use of the term "means for" in a claim is intended to invoke 35 U.S.C. 112, paragraph 6 or means-plus-function claim format, and a claim without the term "means for" is not so intended.

[0195] A processor in association with software may be used to implement a radio transmit / receive unit (WTRU), user equipment (UE), terminal, base station, mobility management entity (MME) or evolved packet core (EPC), or radio frequency transceiver for use in any host computer. The WTRU may be used in combination with modules implemented in hardware and / or software, including a software-defined radio (SDR), as well as a camera, a video camera module, a videophone, a speakerphone, a vibration device, a speaker, a microphone, a television walkie-talkie, a hands-free headset, a keyboard, a Bluetooth module, a frequency modulation (FM) radio unit, a near-field communications (NFC) module, a liquid crystal display (LCD) display unit, an organic light-emitting diode (OLED) display unit, a digital music player, a media player, a video game player module, an internet browser, and / or a wireless local area network (WLAN) or ultra-wideband (UWB) module.

[0196] Although the present invention has been described with respect to a communications system, it is contemplated that the system may be implemented in software on a microprocessor / general purpose computer (not shown). In particular embodiments, one or more of the functions of the various components may be implemented in software controlling a general purpose computer.

[0197] Although the invention has been illustrated and described herein with reference to specific embodiments, it is not intended to limit the invention to the details shown. Rather, various changes in detail may be made within the scope and range of equivalents of the claims without departing from the invention.

[0198] Throughout this disclosure, those skilled in the art will understand that certain exemplary embodiments may be used in the alternative or in combination with other exemplary embodiments.

[0199] While features and elements are described above in particular combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with the other features and elements. Furthermore, the methods described herein may be implemented in a computer program, software, or firmware embodied in a computer-readable medium for execution by a computer or processor. Examples of non-transitory computer-readable storage media include, but are not limited to, read-only memory (ROM), random-access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs). A processor in conjunction with software may be used to implement a radio frequency transceiver for use in a WTRU, UE, terminal, base station, RNC, or any host computer.

[0200] Additionally, in the above embodiments, processing platforms, computing systems, controllers, and other devices including processors are described. These devices may include at least one central processing unit (CPU) and memory. In accordance with the practices of those skilled in the art of computer programming, references to acts and symbolic representations of operations or instructions may be performed by various CPUs and memories. Such acts and operations or instructions may be referred to as being "executed," "executed by a computer," or "executed by a CPU."

[0201] Those skilled in the art will understand that the acts and symbolically represented operations or instructions include the manipulation of electrical signals by the CPU. The electrical system represents the data bits, which may cause the transformation or reduction of the resulting electrical signals and the maintenance of the data bits in memory locations within the memory system, thereby reconfiguring or altering the operation of the CPU and other processing of the signals. The memory locations in which the data bits are maintained are physical locations with particular electrical, magnetic, optical, or organic properties that correspond to or represent the data bits.

[0202] Data bits may also be maintained on computer-readable media, including magnetic disks, optical disks, and other volatile (e.g., random access memory (RAM)) or non-volatile (e.g., read-only memory (ROM)) CPU-readable mass storage systems. The computer-readable media may include cooperative or interconnected computer-readable media that reside exclusively on a processing system or that are distributed among multiple interconnected processing systems, which may be local or remote to the processing system. Representative embodiments are not limited to the above memories, and it will be understood that other platforms and memories may support the described methods.

[0203] Suitable processors include, for example, general purpose processors, special purpose processors, conventional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors associated with a DSP core, controllers, microcontrollers, application specific integrated circuits (ASICs), application specific standard products (ASSPs), field programmable gate array (FPGA) circuits, other types of integrated circuits (ICs), and / or state machines.

[0204] Although the present invention has been described with respect to a communications system, it is contemplated that the system may be implemented in software on a microprocessor / general purpose computer (not shown). In particular embodiments, one or more of the functions of the various components may be implemented in software controlling a general purpose computer.

[0205] Although the invention has been illustrated and described herein with reference to specific embodiments, it is not intended to limit the invention to the details shown. Rather, various changes in detail may be made within the scope and range of equivalents of the claims without departing from the invention.

Claims

1. 1. A method for delivering a video stream, comprising: obtaining a plurality of video representations of the video content; obtaining subpicture property information associating subpictures representing the same region of the video content across the plurality of video representations; generating the video stream based on the sub-picture property information, wherein sub-pictures representing the same temporal region of the video content are extracted from one or more of the video representations, said generating including: at a switch point, switching from extracting a first sub-picture from a first video representation to extracting a second sub-picture from a second video representation, the second sub-picture representing the same region of the video content as a non-Intra Random Access Point (IRAP) sub-picture; delivering said video stream to a decoder in a bitstream; and A method for providing

2. The method of claim 1 , further comprising encoding a syntax element indicating the switching point into the bitstream.

3. updating the sub-picture property information, including adding a syntax element indicating the switching point; encoding the updated sub-picture property information into the bitstream; The method of claim 1 further comprising:

4. For each region of the video content, the sub-picture property information identifies a sub-picture from a plurality of video representations representing the region, and for each of these sub-pictures, the information includes: a subpicture ID that identifies the subpicture; a layer ID identifying a video representation containing said sub-picture; the coordinates of the subpicture; The method of claim 1 , comprising:

5. information identifying the first sub-picture and the second sub-picture, a subpicture ID identifying the first subpicture and the second subpicture; coordinates of the first sub-picture and the second sub-picture; the size of the first sub-picture and the second sub-picture; The method of claim 1 , further comprising encoding information including:

6. The method of claim 1 , wherein the first video representation represents the video content at a first resolution level and the second video representation represents the video content at a second resolution level.

7. The method of claim 1 , wherein the switching is performed when the second subpicture matches a viewport of the video content.

8. For a sub-picture, the sub-picture property information comprises: The method of claim 1 , further comprising mapping information used to map the sub-pictures into the space of the video content.

9. 2. The method of claim 1, wherein the subpicture property information further indicates one or more recommended switching points associated with each picture of the video content, each recommended switching point recommending switching from a subpicture of one of the plurality of video representations to a corresponding subpicture of another one of the plurality of video representations.

10. The method of claim 1 , further comprising encoding a supplemental enhancement information (SEI) message containing the sub-picture property information into the bitstream.

11. The method of claim 1 , wherein each of the plurality of video representations represents the video content according to a respective set of coding parameters, including a resolution level, a quality level, a projection format, a region-by-region packing, or a combination thereof.

12. A device for distributing a video stream, comprising: at least one processor; a memory storing instructions that, when executed by said at least one processor, cause said device to perform the method of any one of claims 1 to 11; A device comprising:

13. 1. A method for decoding a video stream, comprising: The video stream is generated from multiple video representations of video content, and sub-pictures representing the same temporal region of the video content are extracted from one or more of the video representations to form the video stream, and at a switching point, a first sub-picture extracted from a first video representation is followed by a second sub-picture extracted from a second video representation, the second sub-picture representing the same region of the video content as a non-Intra Random Access Point (IRAP) sub-picture, and the method comprises: obtaining subpicture property information associating subpictures representing the same region of the video content across the plurality of video representations; decoding the video stream from a bitstream, the decoding including decoding the second subpicture by referencing the first subpicture identified based on the subpicture property information; A method for providing

14. For each region of the video content, the sub-picture property information identifies a sub-picture from a plurality of video representations representing the region, and for each of these sub-pictures, the information includes: a subpicture ID that identifies the subpicture; a layer ID identifying a video representation containing said sub-picture; the coordinates of the subpicture; 14. The method of claim 13, comprising:

15. The decoding step comprises:

14. The method of claim 13, comprising decoding, from the bitstream, a syntax element indicating the switching point.

16. The method of claim 13 , further comprising scaling the referenced first subpicture to match the scale of the second subpicture.

17. The method of claim 13 , further comprising decoding, from the bitstream, a supplemental enhancement information (SEI) message that includes the sub-picture property information.

18. The method of claim 13 , wherein the second sub-picture coincides with a viewport of the video content.

19. 14. The method of claim 13, wherein the subpicture property information further indicates one or more recommended switch points associated with each picture of the video content, each recommended switch point recommending switching from a subpicture of one of the plurality of video representations to a corresponding subpicture of another one of the plurality of video representations.

20. 1. An apparatus for decoding a video stream, comprising: at least one processor; a memory storing instructions that, when executed by said at least one processor, cause said device to perform the method of any one of claims 13 to 19; A device comprising: