Sub-image bitstream extraction and repositioning
By encoding and decoding 360-degree video as a bitstream with sub-image level information and offsets, the method optimizes video delivery for immersive experiences with reduced complexity and latency.
Patent Information
- Application Number
- JP2024027813
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-05-31
- Filing Date
- 2024-02-27
- Publication Date
- 2025-11-12
- Estimated Expiration
- 2040-03-11
AI Technical Summary
360-degree video processing and delivery face challenges due to large video sizes hindering high-quality delivery, necessitating efficient video coding and decoding methods to achieve immersive user experiences with low latency.
The method involves encoding and decoding video as a bitstream with sub-images, signaling level information and position offsets for each sub-image, allowing independent decoding and repositioning of sub-images to optimize video delivery.
This approach enables efficient decoding and repositioning of sub-images, facilitating high-quality 360-degree video delivery with reduced complexity and latency.
Smart Images

Figure 0007769024000010 
Figure 0007769024000011 
Figure 0007769024000012
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a nonprovisional application of and claims the benefit under 35 U.S.C. § 119(e) of U.S. Provisional Patent Application No. 62 / 816,703, entitled "Sub-Picture Bitstream Extraction and Reposition," filed March 11, 2019, and U.S. Provisional Patent Application No. 62 / 855,446, entitled "Sub-Picture Bitstream Extraction and Reposition," filed May 31, 2019, both of which are incorporated herein by reference in their entireties. [Background technology]
[0002] 360-degree video is a rapidly growing new format in the media industry. It is made possible by the increasing availability of VR devices and can provide viewers with a profoundly new sense of presence. Compared to traditional rectilinear video (2D or 3D), 360-degree video poses a new and challenging set of engineering challenges for video processing and delivery. High-quality video and very low latency are necessary to achieve a comfortable and immersive user experience, but large video sizes can hinder the delivery of high-quality 360-degree video.
[0003] Video coding standards specify a syntax to be followed to convey video and related information in a bitstream. In some cases, it may be desirable to use only a specific subset of the available syntax, for example, to reduce complexity. Different subsets of the overall bitstream syntax are referred to as different "profiles." Even when using a particular profile, there is a large variation in the memory and processing capabilities of video encoder and decoder devices. While various videos may conform to the syntax specified in a particular profile, those various videos may require large variations in encoder and decoder performance. The required performance may be strongly correlated with certain values signaled in the bitstream, such as the size of the decoded image.
[0004] To address this issue, some video coding standards specify "levels" within each profile. A "level" is a predefined set of constraints imposed on the values that may be taken by syntax elements and variables signaled in the bitstream. Some of these constraints impose limits on individual values, while others impose limits on arithmetic combinations of values. For example, a particular level might impose a limit on the picture width multiplied by the picture height multiplied by the number of pictures decoded per second.
[0005] In some standards, levels are specified with "layers." Generally, levels specified for lower layers are more restrictive than levels specified for higher layers. Layers serve as categories of level constraints imposed on values signaled in a bitstream. Because level constraints are nested within layers, a decoder that can decode a bitstream with a particular layer and level is expected to be able to decode all bitstreams that conform to the same layer, layers below that level, or layers at any level below that.
[0006] In some video coding standards, profile, tier, and level information is signaled in syntactic structures such as the "profile_tier_level()" structure. For example, in HEVC, the "profile_tier_level()" structure contains a "general_level_idc" element, which indicates the level to which the coded video sequence in the bitstream conforms. Summary of the Invention
[0007] The embodiments described herein include methods used in the video encoding and decoding (collectively "coding") and rewriting processes in the bitstream.
[0008] In some embodiments, a method includes encoding a video including at least one image including a plurality of sub-images in a bitstream, and signaling, in the bitstream, level information for each of the respective sub-images, the level information indicating, for each sub-image, a set of predefined constraints on values of syntax elements of the respective sub-image.
[0009] Some embodiments further include signaling one or more of the layers or profiles for each sub-image.
[0010] In some embodiments, at least one of the sub-images is a layered sub-image coded in the bitstream using multiple layers, and level information is signaled in the bitstream for each of the layers.
[0011] In some embodiments, each of the sub-images is associated with a layer, and each sub-image within a layer is coded independently of other sub-images within the same layer.
[0012] In some embodiments, the method further includes signaling at least one output sub-image set in the bitstream, the output sub-image set identifying at least a subset of the plurality of sub-images and including level information for each of the sub-images in the subset.
[0013] In some embodiments, the method further includes signaling at least one output sub-image set in the bitstream, the output sub-image set identifying at least a subset of the plurality of sub-images and including position offset information for each of the sub-images in the subset.
[0014] In some embodiments, the method further includes signaling at least one output sub-image set in the bitstream, the output sub-image set identifying at least a subset of the plurality of sub-images and including size information for each of the sub-images in the subset.
[0015] In some embodiments, the level information of the sub-image is signaled in the profile_tier_level() data structure.
[0016] In some embodiments, the method includes decoding level information for each of a plurality of respective sub-images from the bitstream, the level information indicating, for each sub-image, a set of predefined constraints on values of syntax elements of the respective sub-image, and decoding the plurality of sub-images from the bitstream in accordance with the level information.
[0017] In some embodiments, the method further includes selecting an output sub-image set of sub-images based at least in part on the level information, and decoding the plurality of sub-images includes decoding the selected output sub-image set.
[0018] In some embodiments, the method further includes decoding, for at least one of the sub-images, information indicative of the layer of the respective sub-image.
[0019] In some embodiments, the method further includes decoding, for at least one of the sub-images, information indicative of a profile for the respective sub-image.
[0020] In some embodiments, at least one of the sub-images is a layered sub-image encoded in the bitstream using multiple layers, and the method further includes decoding level information from the bitstream for at least one layer.
[0021] In some embodiments, each of the sub-images is associated with a layer, and at least one sub-image within a layer is decoded independently from other sub-images within the same layer.
[0022] Some embodiments further include decoding at least one output sub-image set from the bitstream, the output sub-image set identifying at least a subset of the plurality of sub-images and including level information for each of the sub-images in the subset.
[0023] Some embodiments further include constructing at least one output frame from the decoded sub-images.
[0024] Some embodiments further include decoding at least one output sub-image set from the bitstream, the output sub-image set identifying at least a subset of the plurality of sub-images and including positional offset information for each of the sub-images in the subset, and the output frame being constructed based on the positional offset information.
[0025] Some embodiments further include decoding at least one output sub-image set from the bitstream, the output sub-image set identifying at least a subset of the plurality of sub-images and including size information for each of the sub-images in the subset, and the output frame being constructed based on the size information.
[0026] In some embodiments, the level information of the sub-image is decoded in the profile_tier_level() data structure.
[0027] In some embodiments, the signal comprises information encoding a video including at least one image including a plurality of sub-images and level information for each of the respective sub-images, the level information indicating, for each sub-image, a set of predefined constraints on values of syntax elements of the respective sub-image. The signal may be stored on a computer-readable medium. The computer-readable medium may be a non-transitory medium.
[0028] In additional embodiments, an encoder, decoder, and bitstream rewriting / extraction system are provided for performing the methods described herein.
[0029] Some embodiments include a processor configured to perform any of the methods described herein. In some such embodiments, a computer-readable medium (e.g., a non-transitory medium) is provided that stores instructions that operate to perform any of the methods described herein.
[0030] Some embodiments include a computer-readable medium (eg, a non-transitory medium) that stores video encoded using one or more of the methods disclosed herein.
[0031] One or more of the present embodiments also provide a computer-readable storage medium having stored thereon instructions for encoding or decoding video data according to any of the methods described above. The present embodiments also provide a computer-readable storage medium having stored thereon a bitstream generated according to the methods described above. The present embodiments also provide a method and apparatus for transmitting a bitstream generated according to the methods described above. The present embodiments also provide a computer program product including instructions for performing any of the methods described. [Brief explanation of the drawings]
[0032] [Figure 1A] FIG. 1 is a system diagram illustrating an example communication system in which one or more disclosed embodiments may be implemented. [Figure 1B] 1B is a system diagram illustrating an exemplary wireless transmit / receive unit (WTRU) that may be used within the communication system shown in FIG. 1A, according to one embodiment. [Figure 1C] FIG. 1 is a functional block diagram of a system used in some embodiments described herein. [Figure 2A] FIG. 1 is a functional block diagram of a block-based video encoder such as the encoder used in VVC. [Figure 2B] FIG. 1 is a functional block diagram of a block-based video decoder such as the decoder used in VVC. [Figure 3] FIG. 1 is a diagram of an example architecture of a two-layer scalable video encoder. [Figure 4] FIG. 1 is a diagram of an example architecture of a two-layer scalable video decoder. [Figure 5] FIG. 1 illustrates an example of a two-view video coding structure. [Figure 6] FIG. 10 is a diagram illustrating an example of an inter-layer prediction structure. [Figure 7] FIG. 1 illustrates an example of a coded bitstream structure. [Figure 8] FIG. 1 illustrates an example of a communication system. [Figure 9] FIG. 1 illustrates an example of adaptive streaming of a 360 video viewport. [Figure 10] FIG. 10 illustrates an example of a skipped region in an output image. [Figure 11] FIG. 2 is a diagram illustrating an example of a layer structure. [Figure 12] FIG. 10 illustrates the activation order of parameter sets. [Figure 13] FIG. 10 is a diagram illustrating an example of a sub-DPB structure. [Figure 14] FIG. 10 shows an example of POC derivation for sub-image extraction and repositioning. [Figure 15] FIG. 10 is a diagram illustrating an example of a hierarchical parameter set structure for a sub-image. [Figure 16] FIG. 1 illustrates a layer structure for multiple media types. [Figure 17] 1 is a flowchart of a method performed in some embodiments.
[0033] Examples of networks for implementation 1A illustrates an example communication system 100 in which one or more disclosed embodiments may be implemented. The communication system 100 may be a multiple-access system that provides content, such as voice, data, video, messaging, and broadcasts, to multiple wireless users. The communication system 100 may enable the multiple wireless users to access such content through the sharing of system resources, including wireless bandwidth. For example, the communication system 100 may employ one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single-carrier FDMA (SC-FDMA), zero-tailed unique word DFT-spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block filtering OFDM, filter bank multicarrier (FBMC), etc.
[0034] 1A, communications system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RAN 104, CN 106, public switched telephone network (PSTN) 108, Internet 110, and other networks 112, although it will be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of WTRUs 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a wireless environment. By way of example, the WTRUs 102a, 102b, 102c, 102d (any of which may be referred to as a “station” and / or “STA”) may be configured to transmit and / or receive wireless signals and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, mobile phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearables, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in the context of industrial and / or automated processing chains), consumer electronics devices, devices operating in commercial and / or industrial wireless networks, etc. Any of the WTRUs 102a, 102b, 102c, and 102d may be referred to interchangeably as a UE.
[0035] The communications system 100 may also include a base station 114a and / or a base station 114b. Each of the base stations 114a, 114b may be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communications networks, such as the CN 106, the Internet 110, and / or other networks 112. By way of example, the base stations 114a, 114b may be a base transceiver station (BTS), a Node-B, an eNode-B, a Home Node-B, a Home eNode-B, a gNB, an NR Node-B, a site controller, an access point (AP), a wireless router, etc. Although the base stations 114a, 114b are each shown as a single element, it will be understood that the base stations 114a, 114b may include any number of interconnected base stations and / or network elements.
[0036] The base station 114a may be part of the RAN 104, which may also include other base stations and / or network elements (not shown), such as a base station controller (BSC), a radio network controller (RNC), relay nodes, etc. The base station 114a and / or base station 114b may be configured to transmit and / or receive wireless signals on one or more carrier frequencies, which may be referred to as a cell (not shown). These frequencies may be licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide wireless service coverage for a particular geographic area, which may be relatively fixed or may change over time. A cell may be further divided into cell sectors. For example, the cell associated with the base station 114a may be divided into three sectors. Thus, in one embodiment, the base station 114a may include three transceivers, i.e., one for each sector of the cell. In one embodiment, the base station 114a may employ multiple-input multiple-output (MIMO) technology and utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in desired spatial directions.
[0037] The base stations 114a, 114b may communicate with one or more WTRUs 102a, 102b, 102c, 102d over an air interface 116, which may be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 116 may be established using any suitable radio access technology (RAT).
[0038] More specifically, as noted above, the communications system 100 may be a multiple-access system and may employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, the base station 114a and the WTRUs 102a, 102b, 102c of the RAN 104 may implement a radio technology, such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which may establish the air interface 116 using Wideband CDMA (WCDMA). WCDMA may include communication protocols such as High Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High Speed Downlink (DL) Packet Access (HSDPA) and / or High Speed UL Packet Access (HSUPA).
[0039] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which may establish the air interface 116 using Long Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-Advanced Pro (LTE-A Pro).
[0040] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as NR radio access, which may establish the air interface 116 using New Radio (NR).
[0041] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement multiple radio access technologies. For example, the base station 114a and the WTRUs 102a, 102b, 102c may jointly implement LTE radio access and NR radio access, e.g., using a dual connectivity (DC) principle. Thus, the air interface utilized by the WTRUs 102a, 102b, 102c may be characterized by transmissions sent to and from multiple types of radio access technologies and / or multiple types of base stations (e.g., eNBs and gNBs).
[0042] In other embodiments, the base station 114a and the WTRUs 102a, 102b, 102c may implement a wireless technology such as IEEE 802.11 (i.e., Wireless Fidelity (WiFi)), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rates for GSM Evolution (EDGE), GSM EDGE (GERAN), or the like.
[0043] 1A may be, for example, a wireless router, a Home Node-B, a Home eNode-B, or an access point and may utilize any suitable RAT to facilitate wireless connectivity in a local area, such as a business, a home, a vehicle, a campus, an industrial facility, an air corridor (e.g., for use by drones), a road, etc. In one embodiment, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In one embodiment, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 114b and the WTRUs 102c, 102d may utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a picocell or femtocell. 1A, the base station 114b may have a direct connection to the Internet 110. Therefore, the base station 114b may not need to access the Internet 110 through the CN 106.
[0044] The RAN 104 may communicate with the CN 106, which may be any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more WTRUs 102a, 102b, 102c, 102d. The data may have various quality of service (QoS) requirements, such as different throughput requirements, delay requirements, error tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, etc. The CN 106 may provide call control, billing services, mobile location-based services, prepaid calling, Internet connectivity, video distribution, etc., and / or perform high-level security functions such as user authentication. Although not shown in FIG. 1A , it will be understood that the RAN 104 and / or CN 106 may communicate directly or indirectly with other RANs that use the same RAT as the RAN 104 or a different RAT. For example, in addition to being connected to the RAN 104, which may utilize NR radio technology, the CN 106 may also communicate with another RAN (not shown) that uses GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.
[0045] The CN 106 may also serve as a gateway for the WTRUs 102a, 102b, 102c, 102d to access the PSTN 108, the Internet 110, and / or other networks 112. The PSTN 108 may include a circuit-switched telephone network providing plain old telephone service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices that use common communication protocols such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) of the TCP / IP Internet protocol suite. The networks 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, the network 112 may include another CN connected to one or more RANs, which may use the same RAT as the RAN 104 or a different RAT.
[0046] Some or all of the WTRUs 102a, 102b, 102c, 102d in the communications system 100 may include multi-mode capabilities (e.g., the WTRUs 102a, 102b, 102c, 102d may include multiple transceivers for communicating with different wireless networks over different wireless links.) For example, the WTRU 102c shown in FIG. 1A may be configured to communicate with the base station 114a, which may employ a cellular-based wireless technology, and with the base station 114b, which may employ IEEE 802.2 wireless technology.
[0047] 1B is a system diagram illustrating an example WTRU 102. As shown in FIG. 1B, the WTRU 102 may include, among other things, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power source 134, a global positioning system (GPS) chipset 136, and / or other peripherals 138. It will be understood that the WTRU 102 may include any subcombination of the foregoing elements while remaining consistent with an embodiment.
[0048] The processor 118 may be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, other types of integrated circuits (ICs), a state machine, etc. The processor 118 may perform signal coding, data processing, power control, input / output processing, and / or any other function that enables the WTRU 102 to operate in a wireless environment. The processor 118 may be coupled to the transceiver 120, which may be coupled to the transmit / receive element 122. While FIG. 1B depicts the processor 118 and the transceiver 120 as separate components, it will be understood that the processor 118 and the transceiver 120 may be integrated together in an electronic package or chip.
[0049] The transmit / receive element 122 may be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) over the air interface 116. For example, in one embodiment, the transmit / receive element 122 may be an antenna configured to transmit and / or receive RF signals. In one embodiment, the transmit / receive element 122 may be an emitter / detector configured to transmit and / or receive IR, UV, or visible light signals, for example. In yet another embodiment, the transmit / receive element 122 may be configured to transmit and / or receive both RF and light signals. It will be understood that the transmit / receive element 122 may be configured to transmit and / or receive any combination of wireless signals.
[0050] 1B as a single element, the WTRU 102 may include any number of transmit / receive elements 122. More specifically, the WTRU 102 may employ MIMO technology. Thus, in one embodiment, the WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 116.
[0051] The transceiver 120 may be configured to modulate signals transmitted by the transmit / receive element 122 and demodulate signals received by the transmit / receive element 122. As mentioned above, the WTRU 102 may have multi-mode capabilities. Thus, the transceiver 120 may include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs, such as, for example, NR and IEEE 802.11.
[0052] The processor 118 of the WTRU 102 may be coupled to and may receive user input data from a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light emitting diode (OLED) display unit). The processor 118 may also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. Furthermore, the processor 118 may access information from and store data in any type of suitable memory, such as non-removable memory 130 and / or removable memory 132. The non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 may include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, etc. In other embodiments, the processor 118 may access information from and store data in memory that is not physically located on the WTRU 102, such as a server or home computer (not shown).
[0053] The processor 118 may receive power from the power source 134 and may be configured to distribute and / or control the power to other components within the WTRU 102. The power source 134 may be any suitable device for providing power to the WTRU 102. For example, the power source 134 may include one or more dry batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel-metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.
[0054] The processor 118 may also be coupled to the GPS chipset 136, which may be adapted to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to, or instead of, information from the GPS chipset 136, the WTRU 102 may receive location information from a base station (e.g., base stations 114a, 114b) over the air interface 116 and / or may determine its location based on the timing of signals received from two or more nearby base stations. It will be appreciated that the WTRU 102 may obtain location information by way of any suitable location-determination method while remaining consistent with an embodiment.
[0055] The processor 118 may further be coupled to other peripherals 138, which may include one or more software and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, the peripherals 138 may include an accelerometer, an e-compass, a satellite transceiver, a digital camera (for photos and / or videos), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth module, a frequency modulation (FM) radio unit, a digital music player, a media player, a video game player module, an internet browser, a virtual reality and / or augmented reality (VR / AR) device, an activity tracker, etc. The peripherals 138 may include one or more sensors, which may be one or more of a gyroscope, an accelerometer, a Hall effect sensor, a magnetometer, a direction sensor, a proximity sensor, a temperature sensor, a time sensor, a geolocation sensor, an altimeter, a light sensor, a touch sensor, a magnetometer, a barometer, a gesture sensor, a biometric sensor, and / or a humidity sensor.
[0056] The WTRU 102 may include a full-duplex radio for transmitting and receiving some or all of the signals (e.g., associated with a particular subframe on both the UL (e.g., for transmission) and downlink (e.g., for reception)) in parallel and / or simultaneously. The full-duplex radio may include an interference management unit for reducing and / or substantially eliminating self-interference (e.g., choking) either through hardware or signal processing via a processor (e.g., via a separate processor (not shown) or processor 118). In one embodiment, the WTRU 102 may include a half-duplex radio for transmitting and receiving some or all of the signals (e.g., associated with a particular subframe on either the UL (e.g., for transmission) or downlink (e.g., for reception)).
[0057] Although the WTRU is depicted in FIGS. 1A-1B as a wireless terminal, it is contemplated that in certain representative embodiments such a terminal may use a wired communication interface (e.g., temporarily or permanently) with the communication network.
[0058] In a representative embodiment, the other network 112 may be a WLAN.
[0059] 1A-1B and in view of the corresponding description, one or more or all of the functions described herein may be performed by one or more emulation devices (not shown). The emulation devices may be one or more devices configured to emulate one or more or all of the functions described herein. For example, the emulation devices may be used to test other devices and / or to simulate network and / or WTRU functions.
[0060] The emulation devices may be designed to implement one or more tests of other devices in a lab environment and / or an operator network environment. For example, one or more emulation devices may perform one or more or all functions while fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices in the communication network. One or more emulation devices may perform one or more or all functions while temporarily implemented / deployed as part of a wired and / or wireless communication network. The emulation devices may be directly coupled to another device for testing purposes and / or may perform testing using wireless communication over a wireless network.
[0061] One or more emulation devices may perform one or more functions, inclusive, while not being implemented / deployed as part of a wired and / or wireless communication network. For example, the emulation devices may be utilized in test scenarios in a test lab and / or in an undeployed (e.g., test) wired and / or wireless communication network to implement testing of one or more components. One or more emulation devices may be test equipment. Direct RF coupling and / or wireless communication via RF circuitry (which may, for example, include one or more antennas) may be used by the emulation devices to transmit and / or receive data.
[0062] Exemplary system. The embodiments described herein are not limited to being implemented on a WTRU. Such embodiments may be practiced using other systems, such as the system of FIG. 1C. FIG. 1C illustrates a block diagram of an example system in which various aspects and embodiments may be implemented. System 2000 may be embodied as a device including various components described below and configured to perform one or more of the aspects described herein. Examples of such devices include, but are not limited to, various electronic devices, such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television sets, personal video recording systems, connected home appliances, and servers. The elements of system 2000, alone or in combination, may be embodied in a single integrated circuit (IC), multiple ICs, and / or separate components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 2000 are distributed across multiple ICs and / or separate components. In various embodiments, system 2000 is communicatively coupled to one or more other systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various embodiments, the system 2000 is configured to implement one or more of the aspects described in this document.
[0063] The system 2000 includes at least one processor 2010 configured to execute loaded instructions, for example, to implement various aspects described herein. The processor 2010 may include embedded memory, input / output interfaces, and various other circuits, as known in the art. The system 2000 includes at least one memory 2020 (e.g., a volatile memory device and / or a nonvolatile memory device). The system 2000 includes a storage device 2040, which may include nonvolatile and / or volatile memory, including, but not limited to, electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash, magnetic disk drives, and / or optical disk drives. The storage device 2040 may include, by way of non-limiting example, an internal storage device, an attached storage device (including removable and non-removable storage devices), and / or a network-accessible storage device.
[0064] The system 2000 includes an encoder / decoder module 2030 configured to process data to provide, for example, encoded or decoded video, which may include its own processor and memory. The encoder / decoder module 2030 represents a module or modules that may be included in a device that performs encoding and / or decoding functions. As is known, a device may include one or both of an encoding and a decoding module. Furthermore, the encoder / decoder module 2030 may be implemented as a separate element of the system 2000 or may be incorporated within the processor 2010 as a combination of hardware and software, as is known to those skilled in the art.
[0065] Program code loaded into the processor 2010 or the encoder / decoder 2030 to perform various aspects described herein may be stored in the storage device 2040 and subsequently loaded into the memory 2020 for execution by the processor 2010. According to various embodiments, one or more of the processor 2010, the memory 2020, the storage device 2040, and the encoder / decoder module 2030 may store one or more of various items during execution of the processes described herein. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from processing of equations, expressions, operations, and computational logic.
[0066] In some embodiments, internal memory of the processor 2010 and / or encoder / decoder module 2030 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, external memory of the processing device (e.g., the processing device may be either the processor 2010 or the encoder / decoder module 2030) is used for one or more of these functions. The external memory may be memory 2020 and / or storage device 2040, and may be, for example, dynamic volatile memory and / or non-volatile flash memory. In some embodiments, external non-volatile flash memory is used, for example, to store the television's operating system. In at least one embodiment, a high-speed external dynamic volatile memory such as RAM is used as working memory for video coding and decoding operations such as MPEG-2 (MPEG refers to the Moving Picture Experts Group, MPEG-2 is also referred to as ISO / IEC 13818, 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, an emerging standard developed by JVET, i.e., the Joint Video Experts Team).
[0067] Inputs to the elements of system 2000 may be provided through various input devices, as shown in block 2130. Such input devices include, but are not limited to, (i) a radio frequency (RF) section that receives, for example, RF signals transmitted over the air by a broadcast station, (ii) a component (COMP) input terminal (or set of COMP input terminals), (iii) a universal serial bus (USB) input terminal, and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Other examples not shown in FIG. 1C include composite video.
[0068] In various embodiments, the input devices of block 2130 have associated respective input processing elements as known in the art. For example, the RF section may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal or band-limiting a signal to a frequency band), (ii) downconverting the selected signal, (iii) band-limiting again to a narrower frequency band to select a signal frequency band, which in certain embodiments may be referred to as a channel (for example), (iv) demodulating the downconverted, band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF section of various embodiments includes one or more elements that perform these functions, e.g., a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a downconverter, a demodulator, an error corrector, and a demultiplexer. The RF section may include, for example, a tuner that performs a variety of these functions, including downconverting a received signal to a lower frequency (e.g., an intermediate frequency or a frequency near baseband) or to baseband. In one set-top box embodiment, the RF section and its associated input processing elements perform frequency selection by receiving, filtering, downconverting, and filtering again to the desired frequency band an RF signal transmitted over a wired (e.g., cable) medium. In various embodiments, the order of the above (and other) elements is rearranged, some of these elements are removed, and / or other elements that perform similar or different functions are added. Adding elements can include inserting elements between existing elements, such as inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.
[0069] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 2000 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of the input processing, e.g., Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or within processor 2010, as desired. Similarly, aspects of the USB or HDMI interface processing may be implemented, as desired, within a separate interface IC or within processor 2010. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 2010 and an encoder / decoder 2030, which operates in combination with memory and storage elements to process the data stream as desired for display on an output device.
[0070] The various elements of the system 2000 may be provided within an integrated housing in which the various elements are interconnected and capable of transmitting data therebetween using a suitable connection arrangement 2140, e.g., an internal bus as is known in the art, including an Inter-IC (I2C) bus, wiring, and printed circuit boards.
[0071] The system 2000 includes a communication interface 2050 that enables communication with other devices over a communication channel 2060. The communication interface 2050 may include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 2060. The communication interface 2050 may include, but is not limited to, a modem or a network card, and the communication channel 2060 may be implemented in a wired and / or wireless medium, for example.
[0072] In various embodiments, data is streamed or otherwise provided to system 2000 using a wireless network such as a Wi-Fi network, e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signal in these embodiments is received via communication channel 2060 and communication interface 2050 adapted for Wi-Fi communication. Communication channel 2060 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, enabling streaming applications and other over-the-top communications. Other embodiments provide streamed data to system 2000 using a set-top box that delivers data via an HDMI connection in input block 2130. Still other embodiments provide streamed data to system 2000 using an RF connection in input block 2130. As noted above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as a cellular network or a Bluetooth network.
[0073] The system 2000 can provide output signals to various output devices, including a display 2100, speakers 2110, and other peripheral devices 2120. The display 2100 of various embodiments includes, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 2100 may be for a television, a tablet, a laptop, a mobile phone, or other device. The display 2100 may also be integrated with other components (e.g., as in a smartphone) or may be separate (e.g., an external monitor for a laptop). The other peripheral devices 2120, in various example embodiments, include one or more of a standalone digital video disc (or digital versatile disc) (both terms DVR), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 2120 that provide functionality based on the output of the system 2000. For example, a disc player performs the function of playing the output of the system 2000.
[0074] In various embodiments, control signals are communicated between system 2000 and display 2100, speakers 2110, or other peripheral devices 2120 using signaling protocols such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable inter-device control with or without user intervention. Output devices may be communicatively coupled to system 2000 via dedicated connections through respective interfaces 2070, 2080, and 2090. Alternatively, output devices may be connected to system 2000 using communication channel 2060 via communication interface 2050. Display 2100 and speakers 2110 may be integrated into a single unit with other components of system 2000 in an electronic device such as a television. In various embodiments, display interface 2070 includes a display driver, such as a timing controller (T Con) chip.
[0075] The display 2100 and speakers 2110 may alternatively be separate from one or more of the other components, for example, if the RF portion of the input 2130 is part of a separate set-top box. In various embodiments in which the display 2100 and speakers 2110 are external components, the output signal may be provided via a dedicated output connection including, for example, an HDMI port, a USB port, or a COMP output.
[0076] The embodiments may be executed by the processor 2010, or by computer software implemented by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. The memory 2020 may be of any type appropriate to the technical environment, and may be implemented using any suitable data storage technology, such as, by way of non-limiting examples, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. The processor 2010 may be of any type appropriate to the technical environment, and may include, by way of non-limiting examples, one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture. DETAILED DESCRIPTION OF THE INVENTION
[0077] Block-based video coding. Like HEVC, VVC is built on a block-based hybrid video coding framework. Figure 2A provides a block diagram of a general block-based hybrid video coding system. An input video signal 103 is processed block by block. HEVC uses extended block sizes (called "coding units" or CUs) to efficiently compress high-resolution (1080p and above) video signals. In HEVC, CUs can be up to 64x64 pixels. CUs can be further divided into prediction units or PUs, to which individual prediction methods are applied. For each input video block (MB or CU), spatial prediction (161) and / or temporal prediction (163) may be performed. Spatial prediction (or "intra-prediction") predicts the current video block using pixels from already coded neighboring blocks within the same video image / slice. Spatial prediction reduces spatial redundancy inherent in video signals. Temporal prediction (also referred to as "inter-prediction" or "motion-compensated prediction") predicts the current video block using pixels from already coded video images. Temporal prediction reduces temporal redundancy inherent in video signals. The temporal prediction signal for a given video block is typically signaled by one or more motion vectors, which indicate the amount and direction of motion between the current block and its reference block. Also, if multiple reference pictures are supported (as in recent video coding standards such as H.264 / AVC or HEVC), a reference picture index is additionally transmitted for each video block, and the reference index is used to identify which reference picture in the reference picture store (165) the temporal prediction signal comes from. After spatial and / or temporal prediction, the encoder's mode decision block (181) selects the best prediction mode, for example, based on a rate-distortion optimization method. The prediction block is then subtracted (117) from the current video block, and the prediction residual is decorrelated using a transform (105) and quantization (107) to achieve the target bitrate.The quantized residual coefficients are inverse quantized (111) and inverse transformed (113) to form a reconstructed residual, which is then added to the prediction block (127) to form a reconstructed video block. Further in-loop filtering, such as a deblocking filter and an adaptive loop filter, may be applied to the reconstructed video block (167) before it is placed in a reference picture store (165) and used to code future video blocks. To form the output video bitstream 121, the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to an entropy coding unit (109) for further compression and packing to form the bitstream.
[0078] FIG. 2B provides a block diagram of a block-based video decoder. A video bitstream 202 is first unpacked and entropy decoded in an entropy decoding unit 208. Coding mode and prediction information is sent to either a spatial prediction unit 260 (if intra-coded) or a temporal prediction unit 262 (if inter-coded) to form a prediction block. Residual transform coefficients are sent to an inverse quantization unit 210 and an inverse transform unit 212 to reconstruct the residual block. The prediction block and the residual block are then summed at 226. The reconstructed block may further pass in-loop filtering before being stored in a reference picture store 264. The reconstructed video in the reference picture store is then sent to drive a display device and used to predict future video blocks.
[0079] In modern video codecs, bidirectional motion compensation prediction (MCP) is known for its high efficiency in removing temporal redundancy by exploiting the temporal correlation between images and is widely adopted in most state-of-the-art video codecs. However, a bi-prediction signal is formed by simply combining two uni-prediction signals using a weight value equal to 0.5. This is not necessarily optimal for combining uni-prediction signals, especially when illumination changes rapidly from one reference image to another. Therefore, some prediction techniques aim to compensate for illumination fluctuations over time by applying global or local weights and offset values to each of the sample values of the reference images.
[0080] Scalable video coding. A single-layer video encoder may receive a single video sequence input and generate a single compressed bitstream that is transmitted to a single-layer decoder. Video codecs may be designed for digital video services (e.g., transmission of TV signals via satellite, cable, and terrestrial transmission channels). In video-centric applications deployed in heterogeneous environments, multi-layer video coding techniques may be developed as extensions to video coding standards to enable various applications. For example, multi-layer video coding techniques, such as scalable video coding and / or multi-view video coding, may be designed to process multiple video layers, where each layer may be decoded to reconstruct a video signal of a particular spatial resolution, temporal resolution, fidelity, and / or view. While single-layer encoders and decoders are described with reference to FIGS. 2A and 2B, the concepts described herein may utilize multi-layer encoders and / or decoders, for example, for multi-view and / or scalable coding techniques.
[0081] Scalable video coding may improve the quality of experience for video applications running on devices with various capabilities over heterogeneous networks. Scalable video coding may encode a signal once at the highest representation (e.g., temporal resolution, spatial resolution, quality, etc.) but allow decoding from a subset of the video stream, depending on the specific rate and representation required for a particular application running on a client device. Scalable video coding may save bandwidth and / or storage compared to non-scalable solutions. International video standards, such as MPEG-2 Video, H.263, MPEG4 Visual, H.264, etc., may include tools and / or profiles that support modes of scalability.
[0082] Table 1 shows examples of various types of scalability and the corresponding standards that may support them. Bit depth scalability and / or chroma format scalability may be associated, for example, with video formats that may be primarily used in professional video applications (e.g., higher than 8-bit video and higher chroma sampling formats than YUV4:2:0). Aspect ratio scalability may also be provided. [Table 1]
[0083] Scalable video coding may use a base layer bitstream to provide a first level of video quality associated with a first set of video parameters. Scalable video coding may use one or more enhancement layer bitstreams to provide one or more levels of high quality associated with one or more sets of enhancement parameters. The set of video parameters may include one or more of spatial resolution, frame rate, reconstructed video quality (e.g., in formats such as SNR, PSNR, VOM, visual quality), 3D capabilities (e.g., two or more views), luma and chroma bit depth, chroma format, and underlying single-layer coding standard. For example, as shown in Table 1, various types of scalability may be used for various use cases. A scalable coding architecture may provide a common structure that can be configured to support one or more scalabilities (e.g., the scalabilities listed in Table 1). The scalable coding architecture may be flexible and support various scalabilities with minimal configuration effort. The scalable coding architecture may include at least one preferred mode of operation that may not require modifications to block-level operations, so that coding logic (e.g., encoding and / or decoding logic) can be maximized for reuse within the scalable coding system. For example, a scalable coding architecture based on a picture-level inter-layer processing and management unit may be provided, where inter-layer prediction may be performed at the picture level.
[0084] Figure 3 is a diagram of an example architecture of a two-layer scalable video encoder. A video encoder 900 may receive video (e.g., an enhancement layer video input). The enhancement layer video may be downsampled using a downsampler 902 to create a lower-level video input (e.g., a base layer video input). The enhancement layer video input and the base layer video input may correspond to each other through the downsampling process, achieving spatial scalability. A base layer encoder 904 (e.g., an HEVC encoder in this example) may encode the base layer video input block-by-block and generate a base layer bitstream. Figure 2A is a diagram of an example block-based single-layer video encoder that may be used as the base layer encoder of Figure 3.
[0085] In an enhancement layer, an enhancement layer (EL) encoder 906 may receive an EL input video input that may be of higher spatial resolution (e.g., and / or higher values of other video parameters) than the base layer video input. The EL encoder 906 may generate an EL bitstream in a manner substantially similar to the base layer video encoder 904, e.g., using spatial and / or temporal prediction to achieve compression. Inter-layer prediction (ILP) may be available in the EL encoder 906 to improve its coding performance. Unlike spatial and temporal prediction, which may derive a prediction signal based on the coded video signal of the current enhancement layer, inter-layer prediction may derive a prediction signal based on the coded video signal from the base layer (e.g., when there are two or more layers in a scalable system and / or other lower layers). A scalable system may use at least two forms of inter-layer prediction: picture-level ILP and block-level ILP. Here, we consider picture-level ILP and block-level ILP. The bitstream multiplexer 908 may combine the base layer and enhancement layer bitstreams together to generate a scalable bitstream.
[0086] Figure 4 is a diagram of an example architecture of a two-layer scalable video decoder. The two-layer scalable video decoder architecture of Figure 4 may correspond to the scalable encoder of Figure 3. A video decoder 1000 may, for example, receive a scalable bitstream from a scalable encoder (e.g., scalable encoder 900). A demultiplexer 1002 may separate the scalable bitstream into a base layer bitstream and an enhancement layer bitstream. A base layer decoder 1004 may decode the base layer bitstream and reconstruct the base layer video. Figure 2B is a diagram of an example block-based single-layer video decoder that may be used as the base layer decoder of Figure 4.
[0087] The enhancement layer decoder 1006 may decode the enhancement layer bitstream. The EL decoder 1006 may decode the EL bitstream in a manner substantially similar to the base layer video decoder 1004. The enhancement layer decoder may decode using information from the current layer and / or information from one or more dependent layers (e.g., the base layer). For example, such information from one or more dependent layers may undergo inter-layer processing, which may be achieved when picture-level ILP and / or block-level ILP are used. Although not shown, additional ILP information may be multiplexed together with the base and enhancement layer bitstreams in MUX 908. The ILP information may be demultiplexed by DEMUX 1002.
[0088] 5 is a diagram illustrating an example of a two-view video coding structure. As generally indicated at 1100, FIG. 5 illustrates an example of temporal and inter-dimensional / layer prediction for two-view video coding. In addition to typical temporal prediction, inter-layer prediction (e.g., illustrated by dashed lines) may be used to improve compression efficiency by exploring correlations between multiple video layers. In this example, inter-layer prediction may be performed between two views.
[0089] Inter-layer prediction may be used in HEVC scalable coding extensions, for example, to explore strong correlations between multiple layers and / or to improve scalable coding efficiency.
[0090] 6 is a diagram illustrating an exemplary inter-layer prediction structure that may be considered, for example, for an HEVC scalable coding system. As generally shown at 1200, enhancement layer prediction relies on motion-compensated prediction from the reconstructed base layer signal (e.g., after upsampling if the spatial resolution between the two layers is different), the current enhancement layer, and / or averaging the base layer reconstruction signal with a temporal prediction signal. A complete reconstruction of the lower layer image may be performed. Similar concepts may be utilized for HEVC scalable coding with more than two layers.
[0091] Coded bitstream structure. Figure 7 illustrates an example of a coded bitstream structure. A coded bitstream 1300 consists of several NAL (Network Abstraction Layer) units 1301. NAL units may contain coded sample data, such as coded slices 1306, or high-level syntactic metadata, such as parameter set data, slice header data 1305, or supplemental enhancement information data 1307 (which may be referred to as SEI messages). A parameter set is a high-level syntactic structure that contains essential syntactic elements that may apply to multiple bitstream layers (e.g., video parameter set 1302 (VPS)), or to a coded video sequence within a single layer (e.g., sequence parameter set 1303 (SPS)), or to several coded pictures within a coded video sequence (e.g., picture parameter set 1304 (PPS)). Parameter sets can be transmitted together with the coded pictures of the video bitstream or via other means (e.g., out-of-band transmission using a reliable channel, hard coding, etc.). The slice header 1305 is also a high-level syntactic structure that may be relatively small or contain some picture-related information that is only relevant to a particular slice or picture type. The SEI message 1307 carries information that may not be needed by the decoding process but may be used for various other purposes, such as picture output timing or display, and loss detection and concealment.
[0092] Communication devices and systems. FIG. 8 illustrates an example of a communication system. The communication system 1400 may include an encoder 1402, a communication network 1404, and a decoder 1406. The encoder 1402 may communicate with the network 1404 via a connection 1408, which may be a wired or wireless connection. The encoder 1402 may be similar to the block-based video encoder of FIG. 2A. The encoder 1402 may include a single-layer codec (e.g., FIG. 2A) or a multi-layer codec. The decoder 1406 may communicate with the network 1404 via a connection 1410, which may be a wired or wireless connection. The decoder 1406 may be similar to the block-based video decoder of FIG. 2B. The decoder 1406 may include a single-layer codec (e.g., FIG. 2B) or a multi-layer codec.
[0093] The encoder 1402 and / or decoder 1406 may be incorporated into a wide variety of wired communication devices and / or wireless transmit / receive units (WTRUs), such as, but not limited to, digital televisions, wireless broadcast systems, network elements / terminals, servers such as content or web servers (e.g., HyperText Transfer Protocol (HTTP) servers), personal digital assistants (PDAs), laptop or desktop computers, tablet computers, digital cameras, digital recording devices, video game devices, video game consoles, cellular or satellite wireless telephones, digital media players, and the like.
[0094] The communication network 1404 may be any suitable type of communication network. The communication system 1404 may be a multi-access system providing content, such as voice, data, video, messaging, broadcasts, etc., to multiple wireless users. The communication system 1404 may enable multiple wireless users to access such content through the sharing of system resources, including wireless bandwidth. For example, the communication system 1404 may employ one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single-carrier FDMA (SC-FDMA), etc. The communication network 1404 may include multiple connected communication networks. The communication network 1404 may include the Internet and / or one or more private commercial networks, such as cellular networks, WiFi hotspots, Internet Service Provider (ISP) networks, etc.
[0095] Sub image. A sub-picture is a picture that represents a spatial subset of the original video content, which has been divided into spatial subsets by content creators before video encoding. A sub-picture bitstream is a coded version of one or more representations that contain sub-pictures. (The terms sub-picture and sub-picture bitstream may be used interchangeably in this context.)
[0096] Subimages can be used in omnidirectional video for region-of-interest (ROI) applications or viewport-adaptive streaming. Figure 9 shows an example of viewport-adaptive streaming. In this example, subimages can represent surfaces, for example, in a cube-map projection format. Content is encoded at two spatial resolutions. In both resolutions, a 3x2 subimage grid is used, with each subimage coded independently of the other subimages. Each coded subimage sequence is stored as a subimage bitstream, resulting in 12 subimage bitstreams available for extraction. Depending on the user's viewing orientation, various combinations of high-resolution and low-resolution subimages can be extracted and repackaged for delivery to the user's 360 video streaming client. For example, if the user's orientation closely matches the content of the frontal subimage, a high-resolution frontal subimage (e.g., a frontal view) can be extracted along with various other low-resolution subimages (e.g., left, right, top, back, and bottom views). The extracted subimages are then repositioned to form an output bitstream, providing the user with a high-resolution viewport at an overall reduced bitrate.
[0097] Sub-bitstream. The sub-bitstream extraction process is specified in HEVC as a process whereby NAL units in a bitstream that do not belong to a target set are removed from the bitstream, by outputting sub-pictures consisting of NAL units in the bitstream that belong to the target set, as determined by the highest target TemporalID and the target layer identifier list. The inputs to the sub-bitstream extraction process are the bitstream, the highest target TemporalID value, and the target layer identifier list, and the output of such a process is a sub-bitstream.
[0098] In the sub-image extraction and repositioning process, the sub-bitstreams are not only extracted from one bitstream but also repositioned into another bitstream to form the output bitstream.
[0099] Problems addressed in some embodiments. Sub-picture-related proposals were revisited at the 13th JVET meeting to enable flexible tiling and independently decodable rectangular regions. JVET-M0261, "On Tile Grouping," January 2019, proposed that a sub-picture may reference its own PPS with sub-picture size signaling, and furthermore, that a sub-picture may be treated like a picture in the decoding process, e.g., by using padding to treat the sub-picture boundary as the picture boundary. JVET-M0388, "On Merging MCTS for Viewport-Dependent Streaming," January 2019, proposed that sub-pictures of the same coded picture may have different NAL unit type values to accommodate different IRAP (Intra Random Access Point) distances for each representation. Sub-pictures may also be treated as motion-constrained tile sets (MCTS), as specified in HEVC using SEI messages.
[0100] Disclosed herein are systems and methods for decoder buffer management, picture order count (POC) value synchronization, use of sub-picture parameter sets to facilitate sub-picture-based extraction and repositioning processes, and other operations.
[0101] As described in JVET-N0826, sub-image-based coding can be implemented in VVC. An image can be divided into sub-images, each of which can reference its own PPS with its own tile division. The location and size of each sub-image are indicated in the SPS. The SPS also specifies one or more output sub-image sets. Each output sub-image set can contain multiple sub-images to form an output image with a particular resolution, profile, tier, and level. However, such a system's output sub-image sets only apply sub-images that reference the same SPS. In contrast, it may be desirable to use viewport-dependent streaming to create different sub-images from images of different resolutions that reference different SPSs. Furthermore, if only one layout configuration of the sub-images is signaled in the SPS, the SPS may not be shared across multiple layers.
[0102] For accessing and delivering immersive media, it is desirable to use a system decoder model generalized to new applications. New media applications may consist of multiple components, and rendering of media data operates by decoding all or a subset of the component data. Each component data may be encoded by a different media codec, and multiple scalable versions (e.g., spatial coding resolution, coding quality, temporal coding rate) of the same component content may be available for adaptive access and delivery. For example, video-based point cloud compression (VPCC) projects point cloud data into geometry, texture, occupancy map, and patch components. Each video component data may be encoded by an AVC, HEVC, or VVC encoder, and the point cloud data can be reconstructed by combining all or partially decoded video component data with timed metadata. 3DoF+ visual coding provides multiple base views and additional view data along with metadata to facilitate client-side view synthesis and rendering. These multi-stream scenarios are traditionally handled by system specifications such as file formats and streaming protocols, which address access, delivery, and presentation synchronization. Video coding standards to handle multi-stream scenarios may require new decoding models and NAL unit designs.
[0103] The concept of layers without inter-layer prediction has been adopted as a starting point for supporting immersive media access and delivery in VVC, as the layer structure can directly support multiple streams. For sub-image scenarios, each sub-image can be represented as an independent image of a specific layer, and a sub-image can represent a patch or group of patches of VPCC data. The output image can be a composite image of multiple sub-images from different layers. However, appropriate signaling to indicate the content and spatial correlation between sub-images from different layers has not been developed until now.
[0104] Overview of some embodiments. The exemplary systems and methods described herein employ a high-level syntax design that supports the sub-image extraction and repositioning process. An input video may be coded into multiple representations, and each representation may be represented as a layer. A layer image may be divided into multiple sub-images. Each sub-image may have its own tile division, resolution, color format, and bit depth. Each sub-image is coded independently from other sub-images in the same layer but may be inter-predicted from corresponding sub-images in dependent layers. Each sub-image may reference a sub-image parameter set, in which properties of the sub-image, such as resolution and coordinates, are signaled. Each sub-image parameter set may reference a PPS, in which the resolution of the entire image is signaled.
[0105] The POC values of each sub-picture NAL unit within an associated picture are preferably consistent, although NAL unit types may vary across access units. A POC reset method is used to ensure that ID and non-ID RNAL units share the same POC value.
[0106] A DPB is divided into multiple sub-DPBs, and each sub-DPB is associated with a sub-picture. The maximum sub-DPB size and sorted picture number can be signaled for each sub-picture during session negotiation.
[0107] The output sub-image set is used to indicate the sub-images to be extracted and repositioned for the output image. The sub-image extraction process removes all NAL units whose sub-image identifier or tile group ID is not included in the output sub-image set, and removes all NAL units whose time ID is greater than the target time ID.
[0108] After repositioning the subimages for the output image, each subimage parameter set may reference a new PPS associated with the output image. A POC value for each subimage may be derived based on the POC anchor image of the new output sequence. Constraints are proposed to enable adaptive resolution change (ARC) and ensure that the corresponding reference image is available in the DPB. During ARC, the reference image of the previous subimage may be scaled and transformed to match the ARC subimage being switched to. The scaled and transformed reference image may be placed in the subDPB associated with the new subimage, and the subDPB associated with the previous subimage may be released. The size of each subimage in the output image may change, and the size of the output image may also change. The maximum output image resolution, profile, and level may be signaled in the output subimage set or in an output parameter set associated with the subimage extraction and repositioning process.
[0109] The original video content may be encoded into multiple versions or representations with different resolutions, depths, or color formats. Each of these representations may be packed into a multi-layer structure. Each bitstream may be coded independently or may be inter-layer predicted from other layers. Each representation has its own layer ID and temporal ID. Because each sub-image has only one tile group, the tile group ID may be used as an identifier (e.g., a unique identifier) for the sub-image. As further described in the following subsections, described herein are output sub-image sets, sub-image parameter sets, and sub-DPB management.
[0110] Output of sub-image sets. As described in ISO / IEC DIS23008-2:2018(E), "High Efficiency Video Coding," Scalable HEVC (SHVC) specifies layer sets to identify the set of layers represented in a bitstream created from another bitstream by the operation of a sub-bitstream extraction process.
[0111] In some embodiments, for the sub-image extraction process, an output sub-image set is used to further identify sub-images across multiple layers or representations to be included in the output bitstream. The output sub-image set may be carried in a parameter set for layered coding or session negotiation, such as a video parameter set (VPS), a sequence parameter set (SPS), or a decoder parameter set (DPS). The output sub-image set indicates the number of sub-images included in the set and the tile group ID of each sub-image. Each sub-image may be associated with a layer ID and may be inter-layer predicted from another sub-image of another dependent layer. The output sub-image set identifies the sub-image extraction operation point along with the target temporal layer ID. The middlebox or client may derive the output sub-bitstream by removing all NAL units with layer IDs and sub-image tile group IDs that are not between the values included in the output sub-image set and by removing NAL units with temporal IDs greater than the target temporal layer ID.
[0112] In some embodiments, the output sub-image set includes parameters indicating one or more of output image size, color format, bit depth, and layout of sub-images within the output image for bitstream packing, output image reconstruction, and rendering. Multiple layouts may be provided for the purposes of bitstream packing and output image reconstruction. The sub-image layout may indicate the position and size of each sub-image within the output image. The sub-image layout may indicate per-region transformation types, such as mirroring, flipping, rotating, and scaling of sub-images for output image reconstruction and rendering. In some embodiments, the sub-images are packed in the bitstream at a low resolution but are reconstructed and rendered at an upscaled high resolution based on output sub-image set signaling. [Table 2]
[0113] Table 2 shows an example of a proposed output sub-image set syntax structure. Each output sub-image set (OSPS) in this example specifies the output frame resolution, the number of sub-images to be output, and the layer ID, sub-image ID, position, and size of each output sub-image that makes up the output frame. In the example of Table 2, the output sub-image set signals the profile, tier, and / or level of each sub-image in the set. In some embodiments, this information is signaled in the profile_tier_level() data structure for each sub-image. The element sub_pic_max_tId[i][j] specifies the maximum time ID of the sub-images involved in the extraction process. The transform syntax element specifies the type of transform for the particular sub-image that makes up the output image. Each OSPS may indicate the profile, tier, and level to which it conforms.
[0114] In another embodiment, the output sub-image layout specified by x_offset, y_offset may be optional for recommended per-region packing and rendering, and the client may configure and render the output image in any output layout format. The output image may contain certain skipped regions that are not filled with sub-images, and the client may decide how to fill and render these skipped regions. Figure 10 shows an example of an output image with skipped regions (i.e., skipped regions #0 and #1).
[0115] A parameter set, such as a VPS or SPS, may specify multiple sub-images across layers, with each sub-image having its own unique ID. Sub-images referencing the same VPS or SPS may originate from the same content but are coded into different versions. The coded versions may reference a particular spatial resolution, temporal frame rate, color space, depth, or components. All sub-images of the same coded version may reference the same SPS across layers.
[0116] If an SPS can be shared by multiple layers, the SPS may signal all sub-picture configurations associated with the image of each layer, and the parameter set associated with the PPS or an image consisting of multiple sub-pictures may reference an index into such a sub-picture configuration list. Table 3 shows the proposed SPS syntax structure, where num_sub_pic_cfgs_minus1 plus1 specifies the number of sub-picture configurations available, and each sub-picture configuration may consist of multiple sub-pictures, each with its own position and size. The index pps_sub_pic_cfg_idx specified in Table 4 is an index into the sub-picture configuration list of the SPS, and the corresponding sub-picture layout is applied to the image associated with the PPS. [Table 3] [Table 4]
[0117] In another embodiment, all sub-image configurations may be signaled to the VPS, and the parameter set associated with the SPS, PPS, or image consisting of multiple sub-images may reference an index into such a sub-image configuration list.
[0118] In another embodiment, it may be proposed to override the SPS sub-image configuration in the picture-level parameter set or header of the sub-image whose properties are changed during the coded video sequence (CVS). The override syntax element may include an override flag, the number of updated sub-images, and the updated sub-image configuration when the override flag is set. Table 5 is an example of an SPS syntax structure for overriding the sub-image position and size specified in the SPS or VPS. [Table 5]
[0119] In another embodiment, if the sub-picture configuration associated with a layer referencing an SPS differs from the default sub-picture configuration signaled in the VPS or DPS, the default sub-picture configuration associated with each layer may be signaled in the VPS or DPS, and the SPS may override the sub-picture configuration. The sps_sub_pic_cfg_override_flag may be indicated in the SPS to specify the presence of a sub-picture configuration syntax element in the SPS.
[0120] Figure 11 shows an example layer structure. Sub-images #0, #1, and #2 represent multiple regions of the first source content. Sub-images #5 and #6 represent regions of the second source content. Sub-image #3 is an enhanced (e.g., higher resolution) version of sub-image #0, sub-image #4 is an enhanced version of sub-image #1, and sub-image #7 is an enhanced version of sub-image #6. Sub-image #3 can be predicted from sub-image #0, and sub-image #4 can be predicted from sub-image #1. Sub-image #7 is coded independently. A total of five layers are available (layers 0 through 4), and each layer contains one version of the content. A layer may contain an entire image or only one or more sub-images. Each sub-image may reference its own PPS. All layers associated with the same source may reference the same SPS or VPS. One advantage of sharing the same SPS across layers is the guarantee of coding configurations such as CTU size, bit depth, and chroma format. In some embodiments, a constraint is proposed that each sub-image referencing the same SPS has a unique sub-image ID.
[0121] Table 6 provides examples of sub-image correspondence and dependency indices that indicate sub-image relationships between layers. If the flag sub_pic_corresponding_flag[i][j] is equal to 1, then the identifier corresponding_sub_pic_id[i][j] is provided to specify the sub-image that corresponds to the jth sub-image in the ith layer. Although both corresponding sub-images may cover the same area of the original content, the resolution, quality, and transformation of the two sub-images may be different.
[0122] If the flag sub_pic_dependent_flag[i][j] is equal to 1, the identifier dependent_sub_pic_id[i][j] specifies the ID of the subimage from which the jth subimage of the ith layer is predicted. In FIG. 11, subimage #0 is a dependent subimage and corresponding subimage of subimage #3. Subimage #1 is a dependent subimage and corresponding subimage of subimage #4. Subimage #6 is not a dependent subimage of subimage #7, but is a corresponding subimage of subimage #7. The base layer image may carry all regions of the source content, and the enhancement layer may carry one or more subimage regions. The relative alignment of each enhancement layer subimage in the source content can be inferred from the base layer subimage layout. If there are multiple components of content in the layer structure, the layer of the corresponding subimage may carry region alignment information of the source content. [Table 6]
[0123] In another embodiment, a list of sub-image correspondence groups may be specified in a parameter set. Sub-images covering the same content region may share the same index in the sub-image correspondence group list. Coordination relationships between multiple regions may be signaled individually, as shown in Table 7. [Table 7]
[0124] The value num_regions_minus1 specifies one less than the total number of regions covered by the independently coded regions (sub-images). The values nominal_pic_width and nominal_pic_height specify the nominal image resolution. The index corresponding_sub_pic_group_idx[i] specifies the index into the corresponding sub-picture group list; the identified corresponding sub-image covers the i-th region of the image. The offset values region_x_offset[i] and region_y_offset[i] specify the location of the i-th region, and nominal_region_width[i] and nominal_region_height[i] specify the nominal size of the i-th region.
[0125] Sub-image parameter set. In some embodiments, the parameter set, sub-image parameter set, is used to indicate one or more sub-image parameters such as tile division, coordinates and size of the sub-image, and dependent sub-image layers.
[0126] The subpicture coordinates may indicate the subpicture's location within the picture. The dependent subpicture layers indicate the layers to which the current subpicture may be predicted. The subpicture parameter set may also include DPB management signaling, such as a reference picture list and the maximum DPB buffer size required for each subpicture. Each subpicture may reference its own subpicture parameters, set by a subpicture parameter set ID. Figure 12 shows the order of sequence parameter sets (SPSs), picture parameter sets (PPSs), and subpicture parameter sets (sPPSs), as well as their activation. A subpicture parameter set becomes active when referenced by a tile group, and a PPS becomes active when referenced by an sPPS or a tile group. A subpicture parameter set becomes available to the decoding process before its activation, and the NAL unit containing the sPPS may have a NAL unit layer ID equal to 0. A tile group may also reference both a PPS and an sPPS, with the PPSID and sPPSID signaled in the tile group header. The advantage of including syntax elements in the sPPS is that it avoids the redundant overhead signaled in each tile group header and simplifies the process of rewriting the sub-bitstreams.
[0127] In some embodiments, during sub-image extraction, those NAL units containing sub-picture parameter sets that are not referenced by tile groups of sub-images included in the sub-image set are removed.
[0128] DPB management of sub-images. The decoded picture buffer (DPB) holds decoded pictures for references, output reordering, or output delays specified for a virtual reference decoder. Some embodiments utilize independent coding of each sub-picture by employing a DPB structure that operates at the sub-picture level. In some embodiments, each sub-picture shares the same reference picture list with other sub-pictures within the picture. In some embodiments, each sub-picture may have its own reference picture list to improve coding performance, and in such embodiments, the corresponding reference picture list may be signaled in the sub-picture parameter set.
[0129] JCTVC-O0217, "Sub-DPB Based DPB Operation", October 2013, proposed two modes: (i) a layer-specific sub-DPB mode in which a separate sub-DPB is assigned for each layer, and (ii) a resolution-specific sub-DPB operation mode in which all images with the same spatial resolution, color format, and bit depth share the same sub-DPB.
[0130] In some embodiments, a DPB is divided into multiple sub-DPBs, and each sub-DPB is managed independently for each sub-image. In sub-image-specific sub-DPB modes, decoded sub-images can be inserted, marked, and deleted independently of other sub-images. In some embodiments, the maximum sub-DPB size, maximum number of reordered images, and maximum latency increase are signaled in the PPS or SEI message of each sub-image for session negotiation. This allows a middlebox or client to derive the maximum DPB size used for sub-image repositioning. The PPS may be an appropriate parameter set for carrying sub-image-related properties across multiple sub-images.
[0131] 13 shows an example of a DPB division into multiple layer-based sub-DPBs. Each layer-based sub-DPB can be further divided into multiple sub-picture-based sub-DPBs. Each sub-picture can be coded differently within the corresponding sub-DPB.
[0132] The method described herein of using sub-DPBs for sub-images may simplify the decoding process when the region-wise packing method rotates or flips the sub-images or packs the sub-images to different positions within the image. Because each sub-image is coded independently, a sub-image may find a reference sub-image inside a particular sub-DPB based on a picture order count (POC), a time ID, and a tile group ID, regardless of the sub-image's coordinates within the image.
[0133] For sub-picture switching, such as adaptive resolution change (ARC), an SEI message or external means may indicate the first sub-picture identifier before the ARC switch and the second sub-picture identifier after the ARC switch. To provide a consistent decoding process, constraints may be applied to the ARC sub-pictures. For example, the second sub-picture after the switch may have the same temporal sub-layer structure and coding structure as the first sub-picture before the ARC switch. Such constraints ensure that the reference images for the second sub-picture can be derived from the reference images for the first sub-picture available in the DPB (e.g., in one of the sub-DPBs). The POC values of the sub-pictures in the sub-picture sequence before and after the ARC switch are preferably aligned. For example, if an ARC operation switches from sub-picture #A to sub-picture #B, sub-pictures #A and #B may be coded at different resolutions, color formats, or bit depths. Sub-DPB #A is allocated for sub-picture #A, and sub-DPB #B is allocated for sub-picture #B. During ARC, the client may increase or decrease the buffer size to match the size of sub-DPB #B. The reference images of sub-DPB #A are scaled or transformed to match the properties of sub-image #B, including the new resolution, color format, and / or bit depth. These scaled or transformed reference images may then be assigned to sub-DPB #B, and sub-DPB #A may be freed.
[0134] Derivation of POC values. HEVC and VVC operate to reset the POC value to zero for Instantaneous Decoding Refresh (IDR) pictures and PicOrderCntMsb to zero for Intra Random Access Point (IRAP) pictures with NoRaslOutputFlag equal to 1. If the IRAP distances of different representations are different, the output picture from the sub-picture extraction and repositioning process may consist of sub-picture NAL units of different types, and the derived POC values of related sub-pictures or tile groups may not align.
[0135] In some embodiments, to align POC values across sub-images within an image, the POCLSB values signaled in the tile group header are rewritten. The length of the tile_group_pic_order_cnt_lsb syntax element is log2_max_pic_order_cnt_lsb_minus4+4 bits, and the log2_max_pic_order_cnt_lsb_minus4 syntax element is signaled in the SPS of the associated representation. In some embodiments, a constraint is proposed requiring that the SPSs associated with sub-images across representations share the same value of log2_max_pic_order_cnt_lsb_minus4, simplifying the syntax rewriting process of the tile_group_pic_order_cnt_lsb element. In other embodiments, log2_max_pic_order_cnt_lsb_minus4 may be explicitly signaled for each sub-image in the PPS or sub-image parameter set.
[0136] If one or more NAL unit types are IDR, the POC value of the associated sub-image is zero and may not be the same as the POC values of other non-IDR sub-images in the same image. In some embodiments, a POC reset scheme is proposed to reset the tile_group_pic_order_cnt_lsb values of all sub-images to zero if the image contains at least one IDR sub-image.
[0137] Since PicOrderCntMsb may be inconsistent between the input sub-picture bitstream and the output repositioned bitstream, long-term reference pictures in the decoding order of ARC pictures and pictures following the ARC picture may not be allowed.
[0138] In some embodiments, a POC reset flag is carried with each sub-image or sub-image parameter set, so that POC derivation can be performed by external means, without regard to previous images.
[0139] Figure 14 shows an example of POC value reset for an image created from three sub-images. Sub-image #0 has an IDR interval of 8, and sub-images #1 and #2 have an IDR interval of 4. The newly formed image POC value is reset to 0 if at least one IDRNAL unit is included in the access unit or if the POC reset flag is set externally.
[0140] The output parameter set. A number of coding configuration parameters and coding enablement flags may be specified in the SPS, indicating syntax element lengths, coding unit sizes, and tool configurations that apply to the entire coded video sequence (CVS). For example, log2_ctu_size_minus2 defines the CTU size, log2_min_luma_coding_block_size_minus2 defines the minimum luma coding block size, and sps_sao_enabled_flag determines whether a sample adaptive offset process is applied to the reconstructed image. Each representation may use its own coding parameter settings, and each representation CVS may reference a different SPS. After subimage extraction and repositioning, a single output CVS is formed, referencing one SPS. The values of these SPS configuration parameters may be applied to all subimages from multiple representations. One way to align them is to require all subimages included in the output subimage set to reference the same SPS or share the same parameter values. However, coding performance may be affected if high-resolution and low-resolution representations are coded with the same configuration. One alternative embodiment is to explicitly signal these configuration parameters or coding enable flags individually for each sub-image included in an output sub-image set to PPS or SPS, and each sub-image may reference its corresponding coding configuration parameters using a sub-image identifier.
[0141] In another embodiment, each sub-image may be treated as a single image, which may reference its own PPS, each PPS may reference an SPS specified in HEVC or VVC, and multiple SPSs may reference a DPS covering all potential coding parameters or maximum coding capabilities for the entire decoded sequence. In some embodiments, characteristics of the composite output image, such as output image resolution and sub-image layout, are signaled as syntax elements in a PPS or SPS. In some embodiments, characteristics of the composite output image, such as output image resolution and sub-image layout, are signaled in a separate parameter set, e.g., an output parameter set (OPS). The OPS may be used to indicate properties of the output image for rendering and presentation, which may include the size of the output image and the layout of repositioned sub-images. The OPS may be referenced by a PPS or a sub-image parameter set. Figure 15 shows an example of the relationship between parameter sets in which the same sub-image may be associated with multiple output image resolutions and layouts.
[0142] The process of extracting and repositioning sub-images. HEVC specifies a sub-bitstream extraction process that extracts sub-bitstreams from an input layer bitstream by removing all NAL units whose TemporalID is greater than tIdTarget and all NAL units whose nuh_layer_id is not equal to lidTarget.
[0143] A new media application may operate by extracting multiple sub-picture streams from different layers of a layer bitstream and merging the extracted sub-bitstreams in a specific order to form a new adapted bitstream. Here, we propose a process for extracting and repositioning sub-bitstreams. The inputs to this process are a bitstream and a target sub-picture set, subPicSetTarget. The output of this process is a bitstream.
[0144] To achieve bitstream conformance for an input bitstream, the following conditions may be imposed: where any output sub-bitstream is the output of a process specified in the bitstream, all nuh_layer_id values associated with subPicSetTarget specified in the active VPS, lidTarget, are equal to any value in the range of 0 to 126, and all highest temporal ID values associated with subPicSetTarget specified in the active VPS, tIdTarget, are equal to any value in the range of 0 to 6, and a sub-picture ID value sIdtarget equal to the sub_pic_id associated with subPicSetTarget specified in the active VPS as input, is a conforming bitstream that satisfies the following conditions: The output sub-bitstream contains at least one VCNLAL unit with sub_pic_id equal to sIdTarget, TemporalID equal to tIdTarget, and nuh_layer_id equal to lidTarget.
[0145] The extracted sub-bitstream may be derived in a manner that includes (i) removing all NAL units with a TemporalID greater than tIdTarget, and (ii) removing all NAL units with a nuh_layer_id not equal to lidTarget.
[0146] The repositioned bitstream may be derived in a manner that includes merging collocated access units of extracted sub-bitstreams in the order specified in subPicSetTarget. The access units of the extracted sub-bitstreams represent frames of the corresponding subpictures. The collocated access units of multiple subpictures may share the same timestamp, such as a picture ordinal number.
[0147] The order of the NAL units of different sub-pictures is either signaled in the output sub-picture set or inferred from the sub-picture layout indicated in the output sub-picture set. Each access unit of an output picture may consist of multiple groups of NAL units of the sub-pictures in the order specified in the output sub-picture set.
[0148] A layered architecture for immersive media access and delivery. Various media types (e.g., video data, metadata), components (e.g., geometry, texture, attributes, depth, tiles), and coded versions (resolution, frame rate, bit depth, color space, codec) may be referred to as various representation data in various layers. Specific layer combinations may be output to form an output bitstream to support an application. Clients may access the reconstructed media data and present it in full or partial representations. Figure 16 shows an example in which 360-degree scalable video, PCC data, and 3DoF+ data are multiplexed into a layer bitstream. Different layers may be in different formats and encoded by different media encoders.
[0149] Several syntax elements may be specified in a cross-layer media parameter set to support immersive media access, delivery, and rendering, as follows: An example syntax structure is shown in Table 8.
[0150] In some embodiments, the layer availability flag is used to specify whether the representation data associated with a particular layer is available within the bitstream or is provided by external means outside the specification. For example, mps_layer_available_flag[i] specified in Table 8 indicates whether the ith layer is available in the layer bitstream (mps_layer_available_flag[i] equals 1) or is provided by external means (mps_layer_available_flag[i] equals 0).
[0151] In some embodiments, the layer presentation flag is used to specify whether the representation data of the associated layer is intended to be output separately. For example, a layer containing geometry video data associated with a point cloud object may not be output, decoded, and rendered independently. For example, mps_layer_output_flag[i] in Table 8 indicates whether the ith layer can be decoded and output separately.
[0152] The mapping table may map each layer representation data to a specific media type, a specific media component, and / or a subset of representation data. For example, the layer representation data may represent a specific tile group of point cloud geometry video data, and such mapping may be derived from the layer ID and sub-image ID. For example, the mps_media_type specified in Table 8 indicates the media type or codec type included in the layer structure. The index mps_media_type_idx[i] specifies an index into a list of mps_media_type syntax structures used to map to a specific media or component type.
[0153] In some embodiments, an output media set is used to specify some layer representation data and / or specific layer representation data subsets with temporal IDs, sub-picture IDs, or slice IDs to form an output bitstream representing a whole or partial media presentation. An output media set may also indicate the output representation data rate, maximum resolution and codec profile, and supported layers and levels. The syntax element mps_num_output_set_minus1 specified in Table 8 indicates the number of output media sets. The element mps_media_type_idx[i] specifies the media type of the ith output set, and num_sub_layers[i] specifies the number of sub-pictures or sub-components (e.g., VPCC geometry layers) included in the ith output set. [Table 8]
[0154] Within each layer representation data, a sub-layer dataset is proposed for some embodiments to indicate the matching sub-layer bitstream. The set can identify the sub-layer output data using NAL unit type, temporal ID, and sub-picture ID. The sub-layer dataset can also indicate the sub-layer data contained in the layer data using a byte offset or byte number. The sub-layer dataset ID can be used in the output media set to extract and reposition the output media representation data.
[0155] In the embodiment shown in Table 9, num_sub_layer_minus1 specifies one less than the total number of sub-layers for a particular layer data. The identifier for the i-th sub-layer is specified by sub_layer_id[i]. The element sub_layer_entry_count_minus1 plus1 specifies the number of sub-layer entries available in the layer data. The data length of the i-th entry is indicated by entry_byte_length[i]. The element sub_layer_idx[i] specifies the index of the sub-layer set associated with the i-th entry. [Table 9]
[0156] In some embodiments, a client or middlebox may extract partial media representation data based on the output mediaset. The client may apply a specific media codec to each layer or sub-layer representation to reconstruct complete or partial media data based on the proposed media parameter set and sub-layer data set. If only part of the media data can be reconstructed, a space mapping between the layer or sub-layer data and the target 3D presentation space can be used.
[0157] For example, a multi-layer representation may include both 360-degree video and point cloud objects. A group of layers may represent a specific VPCC object, and another group of layers may represent 360-degree video. Layer data associated with a specific VPCC object may represent VPCC components (e.g., attributes), and sub-layer data may represent independently decodable regions of the component, such as VPCC geometry slices, or component dimensions, such as VPCC geometry layers or attribute types. If a client partially renders a VPCC object against a 360-degree video background, the client may not have access to all of the 360-degree video and VPCC data. The client may access one slice of each component associated with the VPCC object based on one output media set, and the client may access specific viewport data of the 360-degree video based on another output media set. The client can reconstruct the partial VPCC object and 360-degree viewport by decoding, composing, and rendering the two output media sets.
[0158] Syntax design overview. The exemplary systems and methods described herein employ a high-level syntax design that supports the sub-image extraction and repositioning process. An input video may be coded into multiple representations, each of which may be represented as a layer. A layer image may be divided into multiple sub-images. Each sub-image may have its own tile division, resolution, color format, and bit depth. Each sub-image is coded independently of other sub-images in the same layer but may be inter-predicted from corresponding sub-images in dependent layers. Each sub-image may reference a sub-image parameter set, which signals the properties of the sub-image. The sub-image properties may include information such as the resolution of each sub-image and coordinates indicating the position of each sub-image within the output image. Each sub-image parameter set may reference a PPS, which signals the resolution of the entire image.
[0159] The POC values of each sub-picture NAL unit within an associated picture are preferably consistent, although NAL unit types may vary across access units. A POC reset method is used to ensure that ID and non-ID RNAL units share the same POC value.
[0160] A DPB is divided into multiple sub-DPBs, and each sub-DPB is associated with a sub-picture. The maximum sub-DPB size and reordered picture number can be signaled for each sub-picture for session negotiation.
[0161] The output sub-image set may be used to indicate the sub-images to be extracted and repositioned for the output image. The sub-image extraction process removes all NAL units whose sub-image identifier or tile group ID is not included in the output sub-image set, and removes all NAL units whose time ID is greater than the target time ID.
[0162] After repositioning the subimages in the output image, each subimage parameter set may reference a new PPS associated with the output image. A POC value for each subimage may be derived based on the POC anchor image of the new output sequence. Constraints are proposed to enable ARC and ensure the corresponding reference image is available in the DPB. During ARC, the reference image of the previous subimage may be scaled and transformed to match the ARC subimage. The scaled and transformed reference image is placed in the subDPB associated with the new subimage, and the subDPB associated with the previous subimage is released. The size of each subimage in the output image may change, and the size of the output image may also change. The maximum output image resolution, profile, and level may be signaled in the output subimage set or in the output parameter set associated with the subimage extraction and repositioning process.
[0163] Example Systems and Methods. 17, a method performed in some embodiments includes encoding 1702 a video including at least one image including multiple sub-images in a bitstream. The sub-images may be encoded using constraints determined for the sub-images 1704. Level information for each of the respective sub-images is signaled 1706 in the bitstream, where the level information indicates, for each sub-image, a set of predefined constraints on the values of the respective sub-image's syntax elements.
[0164] In some embodiments, the method includes decoding level information for each of a plurality of respective sub-images from the bitstream, the level information indicating, for each sub-image, a set of predefined constraints on values of syntax elements of the respective sub-image, and decoding the plurality of sub-images from the bitstream in accordance with the level information.
[0165] In some embodiments, a signal is provided, the signal comprising information encoding a video including at least one image including a plurality of sub-images and level information for each of the respective sub-images, the level information indicating, for each sub-image, a set of predefined constraints on values of syntax elements of the respective sub-image. The signal may be stored on a computer-readable medium. The computer-readable medium may be a non-transitory medium.
[0166] In some embodiments, the apparatus comprises one or more processors configured to perform an encoding method such as that shown in FIG.
[0167] In some embodiments, an apparatus comprises one or more processors configured to perform a method including decoding level information for each of a plurality of respective sub-images from a bitstream, the level information indicating, for each sub-image, a set of predefined constraints on values of syntax elements of the respective sub-image; and decoding the plurality of sub-images from the bitstream in accordance with the level information.
[0168] In some embodiments, an apparatus described herein includes at least one of: (i) an antenna configured to receive a signal, the signal including data representing an image; (ii) a band limiter configured to limit the received signal to a frequency band including the data representing the image; or (iii) a display configured to display the image. The device may be, for example, a television, a mobile phone, a tablet, a set-top box, or a middlebox.
[0169] In some embodiments, the apparatus includes an access unit configured to access data including a plurality of sub-images and level information for each of the sub-images. The apparatus may further include a transmitter configured to transmit the data.
[0170] In some embodiments, the method includes accessing data including a plurality of sub-images and level information for each of the sub-images. The method may further include transmitting the data including the plurality of sub-images and level information for each of the sub-images.
[0171] In some embodiments, a computer readable medium and computer program product are provided that include a plurality of sub-images and level information for each of the sub-images.
[0172] In some embodiments, the computer-readable medium includes a plurality of sub-images and level information for each of the sub-images.
[0173] In some embodiments, a computer-readable medium includes instructions that cause one or more processors to encode, in a bitstream, a video including at least one image including a plurality of sub-images, and signal, in the bitstream, level information for each of the respective sub-images, the level information indicating, for each sub-image, a set of predefined constraints on values of syntax elements of the respective sub-image.
[0174] In some embodiments, a computer-readable medium includes instructions that cause one or more processors to decode, from a bitstream, level information for each of a plurality of respective sub-images; and decode the plurality of sub-images from the bitstream according to the level information, wherein the level information indicates, for each sub-image, a set of predefined constraints on values of syntax elements of the respective sub-image.
[0175] In some embodiments, a computer program product includes instructions that, when executed by one or more processors, cause the one or more processors to: encode a video including at least one image including a plurality of sub-images; and signal, in the bitstream, level information for each of the respective sub-images, wherein the level information indicates, for each sub-image, a set of predefined constraints on values of syntax elements of the respective sub-image.
[0176] In some embodiments, the computer program product includes instructions that, when executed by one or more processors, cause the one or more processors to decode level information for each of a plurality of respective sub-images from the bitstream; and decode the plurality of sub-images from the bitstream in accordance with the level information, wherein the level information indicates, for each sub-image, a set of predefined constraints on values of syntax elements of the respective sub-image.
[0177] Additional embodiments. In some embodiments, a video bitstream rewriting method includes receiving an input bitstream including a plurality of NAL units, each NAL unit having a layer ID and a sub-image tile group ID; selecting a temporal ID and an output sub-image set, the output sub-image set identifying at least one layer ID and at least one tile group ID; performing a rewriting process on the input bitstream to generate a sub-bitstream; and the rewriting process includes removing from the input bitstream (i) NAL units having a layer ID that is not identified in the output sub-image set, (ii) NAL units having units with a tile group ID that is not identified in the output sub-image set, and (iii) NAL units having a temporal ID greater than the selected temporal ID.
[0178] In some embodiments, the input bitstream further comprises at least one sub-picture parameter set.
[0179] In some embodiments, the sub-image parameter set includes information indicating one or more of the tile division, coordinates of the sub-image within the image, size of the sub-image, and dependent sub-image layers.
[0180] In some embodiments, the sub-picture parameter set includes decoded picture buffer management signaling.
[0181] In some embodiments, the decoded picture buffer management signaling includes one or more of a reference picture list and a maximum decoded picture buffer (DPB) buffer size for each sub-picture.
[0182] In some embodiments, the sub-picture parameter set includes an identifier of a picture parameter set (PPS).
[0183] In some embodiments, the rewriting process further includes removing from the input bitstream (iv) NAL units containing sub-picture parameter sets that are not referenced by tile groups of sub-pictures included in the output sub-picture set.
[0184] In some embodiments, a video decoding method includes receiving a bitstream of video including a plurality of sub-images, wherein the bitstream includes, for at least one of the sub-images, DPB information that divides the DPB into a plurality of sub-DPBs based on a maximum sub-DPB size, a maximum number of reordered images, and a maximum latency increase, each sub-DPB being associated with a corresponding sub-image and including DPB information indicating at least one of decoding each of the sub-images using the corresponding sub-DPB.
[0185] In some embodiments, the video includes multiple layers, and each sub-DPB is associated with a corresponding layer and a corresponding sub-image.
[0186] In some embodiments, the DPB information is included in the PPS in the bitstream.
[0187] In some embodiments, a method includes receiving a bitstream of video, the video including a plurality of images, each image including a plurality of sub-images, and in response to determining that at least one of the sub-images in the corresponding image is an instantaneous decoding refresh (IDR) image, setting a picture order count (POC) value of the corresponding image to zero.
[0188] In some embodiments, a method includes receiving a bitstream of video, the bitstream encoding a plurality of sub-images, the bitstream further including an output parameter set (OPS), the OPS indicating the positions of the sub-images in an output image, decoding the sub-images, and further constructing the output image by positioning the decoded sub-images according to the OPS.
[0189] In some embodiments, a method includes receiving video including an input image; dividing the input image into a plurality of sub-images; encoding each of the sub-images in at least two layers using scalable coding, wherein each of the sub-images is encoded independently of the other sub-images; and encoding a sub-image parameter set for each sub-image, wherein the sub-image parameter set indicates layer dependency for inter-layer prediction of the respective sub-image.
[0190] In some embodiments, each sub-image corresponds to a tile group, and the tile group header of each respective tile group references the corresponding sub-image parameter set.
[0191] In some embodiments, the sub-picture parameter set refers to a picture parameter set (PPS).
[0192] In some embodiments, each sub-image parameter set identifies the resolution of the corresponding sub-image.
[0193] In some embodiments, each sub-image parameter set identifies the location of the corresponding sub-image within the output image.
[0194] In some embodiments, a video bitstream rewriting method includes receiving an input bitstream including a plurality of NAL units, each NAL unit having a layer ID and a sub-picture ID; and receiving an output sub-picture parameter set, the output sub-picture parameter set specifying, for each of a plurality of output sub-picture sets, a layer ID and a sub-picture ID for each sub-picture within the respective output sub-picture; selecting (i) a temporal ID and (ii) an output sub-picture set identified in the output sub-picture parameter set; and performing a rewriting process on the input bitstream to generate a sub-bitstream, the rewriting process including removing from the input bitstream (i) NAL units of sub-pictures that are not in the selected output sub-picture set, as indicated by the layer ID and sub-picture ID, and (ii) NAL units that have a temporal ID greater than the selected temporal ID.
[0195] In some embodiments, the output sub-image parameter set further specifies, for each output sub-image set, a sub-image offset position of each sub-image within the respective output sub-image set.
[0196] In some embodiments, the output sub-image parameter set further specifies, for each output sub-image set, a sub-image width and height for each sub-image in the respective output sub-image set.
[0197] In some embodiments, the output sub-image parameter set further specifies, for each output sub-image set, a width and height of the respective output sub-image set.
[0198] In some embodiments, the subimage ID is a tile group ID.
[0199] In some embodiments, a video decoding method includes receiving an input bitstream including a plurality of sub-images, each sub-image having a respective sub-image ID; receiving an output sub-image parameter set, the output sub-image parameter set specifying, for a plurality of output sub-image sets including at least one selected output sub-image set, a sub-image ID for each sub-image in a selected output sub-image set; decoding each of the sub-images in the selected output sub-image set; and arranging the decoded sub-images into an output frame.
[0200] In some embodiments, the output sub-image parameter set further specifies, for the selected output sub-image set, a sub-image offset position for each sub-image within the respective output sub-image set, and constructing the decoded sub-images includes positioning each of the decoded sub-images at a respective offset position.
[0201] In some embodiments, the output sub-image parameter set further specifies, for the selected output sub-image set, a sub-image width and height for each sub-image in the respective output sub-image set, and constructing the decoded sub-images includes scaling each of the decoded sub-images to the respective width and height.
[0202] In some embodiments, the output sub-image parameter set further specifies, for a selected output sub-image set, a width and height for each output sub-image set.
[0203] In some embodiments, the subimage ID is a tile group ID.
[0204] In some embodiments, a video bitstream rewriting method includes receiving an input bitstream including a plurality of NAL units, each NAL unit having a layer ID and a sub-picture ID; receiving a picture parameter set, the picture parameter set specifying, for each of a plurality of sub-picture configurations, a sub-picture ID in the respective sub-picture configuration; selecting an output sub-picture set identified by (i) a temporal ID and (ii) the output sub-picture parameter set; and performing a rewriting process on the input bitstream to generate a sub-bitstream, the rewriting process including removing from the input bitstream (i) NAL units of sub-pictures that are not in the selected sub-picture configuration, as indicated by the sub-picture ID, and (ii) NAL units that have a temporal ID greater than the selected temporal ID.
[0205] In some embodiments, a video decoding method includes receiving an input bitstream including a plurality of sub-images, each sub-image having a respective sub-image ID; receiving a sequence parameter set, the sequence parameter set specifying, for each of a plurality of sub-image configurations including a selected sub-image configuration, a sub-image ID for each sub-image in the respective sub-image configuration; decoding each of the sub-images in the selected output sub-image set; and composing the decoded sub-images into an output frame.
[0206] In some embodiments, the sequence parameter set further specifies, for the selected sub-image configuration, a sub-image offset position for each sub-image in the selected sub-image configuration, and composing the decoded sub-images includes positioning each of the decoded sub-images at a respective offset position.
[0207] In some embodiments, the sequence parameter set further specifies, for the selected sub-image configuration, a sub-image width and height for each sub-image in the selected sub-image configuration, and composing the decoded sub-images includes scaling each of the decoded sub-images to the respective width and height.
[0208] Some embodiments further include receiving an image parameter set including a sub-image configuration index, wherein the selected sub-image configuration is selected based on the sub-image configuration index.
[0209] In some embodiments, a video decoding method includes receiving an input bitstream including a plurality of sub-images, each sub-image having a respective sub-image ID; receiving a picture parameter set, the picture parameter set including a sub-image configuration override flag; in response to determining that the sub-image configuration override flag is set, determining a sub-image output configuration conveyed in the picture parameter set, the sub-image output configuration including an ID of each sub-image in the output configuration; decoding each sub-image in the output configuration; and composing the decoded sub-images into an output frame.
[0210] In some embodiments, a video decoding method includes receiving an input bitstream including a plurality of sub-images, each sub-image having a respective sub-image ID; receiving a video parameter set, the video parameter set indicating, for each sub-image, whether the sub-image is dependent on another sub-image; and decoding the input bitstream according to the video parameter set.
[0211] In some embodiments, the video parameter set further indicates, for each sub-image that is indicated to be dependent on another sub-image, the sub-image ID of the sub-image on which it depends.
[0212] In some embodiments, the video parameter set further provides, for each sub-image that is not indicated as being dependent on another sub-image, a flag indicating whether the sub-image corresponds to another sub-image.
[0213] In some embodiments, the video parameter set further indicates, for each sub-image that is indicated to correspond to another sub-image, the sub-image ID of the corresponding sub-image.
[0214] In some embodiments, a video decoding method includes receiving an input bitstream including a plurality of sub-images, each sub-image having a respective sub-image ID; receiving a parameter set identifying a plurality of sub-image groups, each group having an index; for each of a plurality of regions of an output frame, receiving a parameter set identifying an index of a sub-image group corresponding to the respective region; for each of the regions, decoding at least one of the sub-images in the sub-image group corresponding to the region; and constructing an output frame from the decoded sub-images.
[0215] In some embodiments, a media decoding method includes receiving a bitstream including a plurality of layers and sub-layers, each sub-layer having a respective sub-layer ID; receiving a media parameter set, the media parameter set indicating, for each of the plurality of layers, whether the layer is available in the bitstream; and decoding the bitstream according to the media parameter set.
[0216] In some embodiments, a media decoding method includes receiving a bitstream including a plurality of layers and sub-layers, each sub-layer having a respective sub-layer ID; receiving a media parameter set, the media parameter set indicating, for each of the plurality of layers, whether the layer can be decoded and output independently; and decoding the bitstream in accordance with the media parameter set.
[0217] In some embodiments, a media decoding method includes receiving a bitstream including a plurality of layers and sub-layers, each sub-layer having a respective sub-layer ID; receiving a media parameter set, the media parameter set indicating a media type for each of the plurality of layers; and decoding the bitstream according to the media parameter set.
[0218] In some embodiments, a media decoding method includes receiving a bitstream including multiple layers and sub-layers, each sub-layer having a respective sub-layer ID; receiving a media parameter set, the media parameter set indicating multiple output sets for the sub-layer including at least one selected output set; and decoding the sub-layer of the selected output set.
[0219] In some embodiments, the media parameter set further indicates a media type for each of the output sets.
[0220] In some embodiments, the media parameter set further indicates the layer ID of each output set.
[0221] In some embodiments, a bitstream extraction method includes receiving a bitstream including a plurality of layers and sub-layers, each sub-layer having a respective sub-layer ID; receiving a sub-layer parameter set, the sub-layer parameter set indicating an entry byte length of each sub-layer; and extracting at least a partial media representation according to the sub-layer parameter set.
[0222] In some embodiments, a system is provided that includes a processor and a non-transitory computer-readable medium storing instructions that operate to perform any of the methods described herein.
[0223] In some embodiments, a non-transitory computer-readable storage medium is provided to store a video bitstream generated using any of the methods described herein.
[0224] This disclosure describes a wide variety of aspects, including tools, features, embodiments, models, approaches, and the like. Many of these aspects are described specifically and may be described in a definitive manner to at least indicate their individual characteristics. However, this is for clarity of description and does not limit the disclosure or scope of these aspects. In fact, all of the different aspects can be combined and interchanged to provide additional aspects. Furthermore, these aspects can also be combined and interchanged with aspects described in previous applications.
[0225] Aspects described and contemplated in this application can be implemented in many different formats. While some embodiments are specifically illustrated, other embodiments are contemplated, and discussion of a particular embodiment is not intended to limit the breadth of implementations. At least one of these aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a generated or encoded bitstream. These and other aspects can be implemented as a method, an apparatus, a computer-readable storage medium having stored thereon instructions for encoding or decoding video data according to any of the described methods, and / or a computer-readable storage medium having stored thereon a bitstream generated according to any of the described methods.
[0226] In this application, the terms "reconstructed" and "decoded" may be used interchangeably, the terms "pixel" and "sample" may be used interchangeably, and the terms "image," "picture," and "frame" may be used interchangeably.
[0227] Various methods are described herein, each of which includes one or more steps or acts for achieving the described method. Unless a specific order of steps or acts is required for the proper operation of the method, the order and / or use of specific steps and / or acts may be varied or combined. Furthermore, terms such as "first," "second," and the like may be used in various embodiments to vary elements, components, steps, operations, etc., e.g., "first decoding" and "second decoding." The use of such terms does not imply a varied order of operations unless specifically required. Thus, in this example, the first decoding need not be performed before the second decoding, but could occur, for example, before, during, or during an overlapping period with the second decoding.
[0228] For example, various numerical values may be used in this disclosure. The specific values are for illustrative purposes, and the described aspects are not limited to these specific values.
[0229] The embodiments described herein may be performed by computer software implemented by a processor or other hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. The processor may be of any type appropriate to the technical environment, including, by way of non-limiting example, one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.
[0230] Various implementations involve decoding. As used herein, "decoding" can encompass, for example, all or part of the processes performed on a received encoded sequence to generate a final output suitable for display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes also or alternatively include processes performed by decoders in various implementations described herein, such as extracting an image from a tiled (packed) image, determining an upsample filter to use, then upsampling the image, and flipping the image back to its intended orientation.
[0231] As a further example, in one embodiment, "decoding" refers only to entropy decoding, in another embodiment, "decoding" refers only to differential decoding, and in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to refer specifically to a subset of operations or to the broader decoding process generally will be clear based on the context of the particular description and will be well understood by one of ordinary skill in the art.
[0232] Various implementations involve encoding. Similar to the above discussion regarding "decoding," "encoding," as used herein, can encompass all or some of the processes performed on an input video sequence to, for example, generate an encoded bitstream. In various embodiments, such processes include one or more processes typically performed by an encoder, such as segmentation, differential encoding, transform, quantization, and entropy encoding. In various embodiments, such processes also or alternatively include processes performed by the encoders of the various implementations described herein.
[0233] As a further example, in one embodiment, "encoding" refers only to entropy encoding, in another embodiment, "encoding" refers only to differential encoding, and in another embodiment, "encoding" refers to a combination of differential and entropy encoding. Whether the phrase "encoding process" is intended to refer specifically to a subset of operations or to the broader encoding process in general will be clear based on the context of the particular description and will be well understood by one of ordinary skill in the art.
[0234] Where a figure is presented as a flow diagram, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, where a figure is presented as a block diagram, it should be understood that it also provides a flow diagram of the corresponding method / process.
[0235] Various embodiments refer to rate-distortion optimization. In particular, during the encoding process, a balance or trade-off between rate and distortion is typically considered, often given computational complexity constraints. Rate-distortion optimization is typically formulated to minimize a rate-distortion function, which is a weighted sum of rate and distortion. There are various approaches to solving the rate-distortion optimization problem. For example, an approach may be based on extensive testing of all encoding options, including all modes or coding parameter values considered, with a thorough evaluation of the coding cost and associated distortion of the reconstructed signal after coding and decoding. In particular, faster approaches may be used to reduce encoding complexity by calculating approximate distortion based on a prediction or prediction residual signal rather than the reconstructed signal. These two approaches may also be used in combination, such as by using approximate distortion for only some of the possible encoding options and full distortion for others. Other approaches evaluate only a subset of the possible encoding options. More generally, many approaches employ any of a variety of techniques for performing the optimization, but the optimization does not necessarily involve a thorough evaluation of both the coding cost and the associated distortion.
[0236] The implementations and aspects described herein may be implemented in, for example, a method or process, an apparatus, a software program, a data stream, or a signal. Even if discussed in the context of only a single implementation (e.g., discussed only as a method), the implementation of the discussed features may also be implemented in other forms (e.g., an apparatus or a program). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. The methods may be implemented in, for example, a processor, which refers generally to processing devices including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, for example, computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end users.
[0237] References to "one embodiment" or "one embodiment," or "one implementation" or "one implementation," as well as other variations thereof, mean that a particular feature, structure, characteristic, etc. described in connection with an embodiment is included in at least one embodiment. Thus, appearances of the phrases "in one embodiment" or "in one embodiment" or "in one implementation" or "in one implementation," as well as any other variations, in various places throughout this application are not necessarily all referring to the same embodiment.
[0238] Additionally, this disclosure may refer to "determining" various portions of information. Determining information may include, for example, one or more of evaluating information, calculating information, predicting information, or retrieving information from memory.
[0239] Additionally, the application may refer to "accessing" various portions of information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or evaluating information.
[0240] Additionally, the application may refer to "receiving" various portions of information. Receiving, like "accessing," is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from a memory). Furthermore, "receiving" typically involves in some manner, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or evaluating information.
[0241] For example, in the case of "A / B," "A and / or B," and "at least one of A and B," it should be understood that the use of any of the following " / ," "and / or," and "at least one of" is intended to encompass the selection of only the first-listed alternative (A), or the selection of only the second-listed alternative (B), or the selection of both alternatives (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C," such phraseology is intended to encompass the selection of only the first-listed alternative (A), or the selection of only the second-listed alternative (B), or the selection of only the third-listed alternative (C), or the selection of only the first and second-listed alternatives (A and B), or the selection of only the first and third-listed alternatives (A and C), or the selection of only the second and third-listed alternatives (B and C), or the selection of all three alternatives (A, B, and C). This may be extended to a number of listed items.
[0242] Also, as used herein, the word "signaling" refers, among other things, to instructing a corresponding decoder. For example, in certain embodiments, an encoder signals a specific one of multiple parameters for refinement. In this manner, in embodiments, the same parameters are used at both the encoder and decoder sides. Thus, for example, an encoder can transmit a specific parameter to a decoder (explicit signaling), so that the decoder can use the same specific parameter. Conversely, if the decoder already has a specific parameter as well as other parameters, signaling can be used without transmission (implicit signaling) to allow the decoder to easily recognize and select the specific parameter. By avoiding the transmission of any actual function, bit savings are realized in various embodiments. It should be appreciated that signaling can be achieved in various ways. For example, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder in various embodiments. Although the above refers to the verb form of the word "signaling," the word "signaling" can also be used as a noun herein.
[0243] Implementations can generate various signals formatted to carry information that can be stored or transmitted, for example. The information can include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal can be formatted to carry a bitstream of the described embodiments. Such a signal can be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.
[0244] Several embodiments are described. Features of these embodiments can be provided alone or in any combination across various claim categories and types. Furthermore, embodiments can include one or more of the following features, devices, or aspects, alone or in any combination across various claim categories and types. Insertion of signaling syntax elements that allow a decoder or middlebox to identify the profile, layer, and / or level of a sub-image. A bitstream or signal containing one or more of the listed syntax elements, or variations thereof. A bitstream or signal including syntax conveying information generated according to any of the described embodiments. Creating and / or transmitting and / or receiving and / or decoding a bitstream or signal that includes one or more of the described syntax elements, or variations thereof. · Created and / or transmitted and / or received and / or decoded according to any of the described embodiments. · A method, process, apparatus, instruction storage medium, data storage medium, or signal according to any of the described embodiments. A television, set-top box, mobile phone, tablet, or other electronic device operable to decode syntax elements indicating a profile, stage, and / or level of a sub-image according to any of the described embodiments. A television, set-top box, mobile phone, tablet, or other electronic device operable to decode syntax elements indicating the profile, tier, and / or level of a sub-image in accordance with any of the described embodiments, and displaying the resulting image (e.g., using a monitor, screen, or other type of display). A television, set-top box, mobile phone, tablet, or other electronic device that selects a channel (e.g., using a tuner) to receive a signal containing the encoded image and decodes syntax elements indicating the profile, tier, and / or level of the sub-image according to any of the described embodiments. A television, set-top box, mobile phone, tablet, or other electronic device that receives a signal containing an encoded image wirelessly (e.g., using an antenna) and decodes syntax elements indicating the profile, tier, and / or level of a sub-image according to any described embodiment.
[0245] It should be noted that various hardware elements of one or more of the described embodiments are referred to as "modules" that perform (i.e., perform, implement, etc.) various functions described herein with respect to the respective modules. As used herein, a module includes hardware deemed suitable by one of ordinary skill in the relevant art for a given implementation (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs), one or more memory devices). It should be noted that each described module may also include executable instructions to perform one or more functions described as being performed by the respective module, and these instructions may take the form of hardware (i.e., hardwired) instructions, firmware instructions, software instructions, etc., or may be stored on any suitable non-transitory computer-readable medium or media, commonly referred to as RAM, ROM, etc.
[0246] Although features and elements are described above in particular combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with the other features and elements. Furthermore, the methods described herein may be implemented in a computer program, software, or firmware embodied in a computer-readable medium for execution by a computer or processor. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random-access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, optical media such as CD-ROM disks, and digital versatile disks (DVDs). A processor in association with software may be used to implement a radio frequency transceiver for use in a WTRU, UE, terminal, base station, RNC, or any host computer.
Claims
1. 1. A video decoding method comprising: Obtaining a bitstream representing a video having a plurality of images, at least one of the plurality of images including a plurality of sub-images, at least one of the sub-images being a layered sub-image coded using a plurality of layers, each layered sub-image being associated with a respective layer-based and sub-image-based sub-DPB; obtaining, for each of the plurality of sub-images, information from the bitstream indicating a size of a sub-image-based sub-DPB corresponding to the respective sub-image, the sub-image being a layered sub-image coded using a plurality of layers; decoding each of the plurality of sub-images using a respective sub-DPB having the indicated size; A method comprising:
2. 10. The method of claim 1, A method in which each sub-image shares a reference image list with other sub-images within the image.
3. 3. The method of claim 2, A method wherein the reference image list is signaled with respective sub-image parameter sets.
4. 10. The method of claim 1, A method in which information indicating the size of the sub-image-based sub-DPB is obtained from a data structure signaled in a PPS or SEI message, the data structure including the maximum size of each layer-based and sub-image-based associated with each layer of the sub-image.
5. 10. The method of claim 1, A method in which information indicating the size of the sub-image-based sub-DPB is obtained from a data structure signaled in a PPS or SEI message, the data structure including the width and height of each layer-based and sub-image-based sub-DPB associated with each layer of the sub-image.
6. 10. The method of claim 1, A method in which information indicating the size of a sub-image-based sub-DPB is obtained from a data structure signaled in a PPS or SEI message, and the data structure indicates the level of each of the multiple sub-images, including a level of each layer of the layered sub-image and a level indicating values of syntax elements of the respective sub-image or a set of predefined constraints for each layer of the layered sub-image.
7. 10. The method of claim 1, Dividing a decoded picture buffer (DPB) into a plurality of layer-based sub-DPBs; dividing each layer-based sub-DPB into respective layer-based and sub-image-based sub-DPBs; A method comprising:
8. 1. A video decoding device, comprising: Obtaining a bitstream representing a video having a plurality of images, at least one of the plurality of images including a plurality of sub-images, at least one of the sub-images being a layered sub-image coded using a plurality of layers, each layered sub-image being associated with a respective layer-based and sub-image-based sub-DPB; obtaining, for each of the plurality of sub-images, information from the bitstream indicating a size of a sub-image-based sub-DPB corresponding to the respective sub-image, the sub-image being a layered sub-image coded using a plurality of layers; decoding each of the plurality of sub-images using a respective sub-DPB having the indicated size; a video decoding device comprising one or more processors configured to perform at least
9. 9. The apparatus of claim 8, Each sub-image shares a reference image list with other sub-images within the image.
10. 10. The apparatus of claim 9, The reference image list is signaled with respective sub-image parameter sets.
11. 9. The apparatus of claim 8, The device, wherein information indicating the size of the sub-image-based sub-DPB is obtained from a data structure signaled in a PPS or SEI message, the data structure including the maximum size of each layer-based and sub-image-based associated with each layer of the sub-image.
12. 9. The apparatus of claim 8, The device, wherein information indicating the size of the sub-image-based sub-DPB is obtained from a data structure signaled in a PPS or SEI message, the data structure including the width and height of each layer-based and sub-image-based sub-DPB associated with each layer of the sub-image.
13. 9. The apparatus of claim 8, The device, wherein information indicating the size of the sub-image-based sub-DPB is obtained from a data structure signaled in a PPS or SEI message, the data structure indicating the level of each layer of the layered sub-image and the level of each of the plurality of sub-images including a level indicating the value of a syntax element of the respective sub-image or a set of predefined constraints for each layer of the layered sub-image.
14. 9. The apparatus of claim 8, Dividing a decoded picture buffer (DPB) into a plurality of layer-based sub-DPBs; dividing each layer-based sub-DPB into respective layer-based and sub-image-based sub-DPBs; The apparatus is further configured to perform the following:
15. 1. A video encoding method comprising: encoding a video having a plurality of images into a bitstream, wherein at least one of the plurality of images includes a plurality of sub-images, at least one of the sub-images being a layered sub-image coded using a plurality of layers, each of the plurality of layered sub-images being coded using a respective layer-based and sub-image-based sub-DPB; For each of the plurality of sub-images, encoding into the bitstream information indicating a size of a sub-image-based sub-DPB corresponding to each sub-image, the sub-image being a layered sub-image encoded using a plurality of layers; A method comprising:
16. 16. The method of claim 15, A method in which each sub-image shares a reference image list with other sub-images within the image.
17. 17. The method of claim 16, A method wherein the reference image list is signaled with respective sub-image parameter sets.
18. 16. The method of claim 15, A method in which information indicating the size of the sub-image-based sub-DPB is obtained from a data structure signaled in a PPS or SEI message, the data structure including the maximum size of each layer-based and sub-image-based associated with each layer of the sub-image.
19. 16. The method of claim 15, A method in which information indicating the size of the sub-image-based sub-DPB is obtained from a data structure signaled in a PPS or SEI message, the data structure including the width and height of each layer-based and sub-image-based sub-DPB associated with each layer of the sub-image.
20. 16. The method of claim 15, A method in which information indicating the size of a sub-image-based sub-DPB is obtained from a data structure signaled in a PPS or SEI message, and the data structure indicates the level of each of the multiple sub-images, including a level of each layer of the layered sub-image and a level indicating values of syntax elements of the respective sub-image or a set of predefined constraints for each layer of the layered sub-image.
21. 16. The method of claim 15, Dividing a decoded picture buffer (DPB) into a plurality of layer-based sub-DPBs; dividing each layer-based sub-DPB into respective layer-based and sub-image-based sub-DPBs; A method comprising:
22. 1. A video encoding device, comprising: encoding a video having a plurality of images into a bitstream, wherein at least one of the plurality of images includes a plurality of sub-images, at least one of the sub-images being a layered sub-image coded using a plurality of layers, each of the plurality of layered sub-images being coded using a respective layer-based and sub-image-based sub-DPB; For each of the plurality of sub-images, encoding information into the bitstream indicating a size of a sub-image-based sub-DPB corresponding to the respective sub-image, the sub-image being a layered sub-image encoded using a plurality of layers; 1. A video encoding device comprising: one or more processors configured to perform at least
23. 23. The apparatus of claim 22, Each sub-image shares a reference image list with other sub-images within the image.
24. 24. The apparatus of claim 23, The reference image list is signaled with respective sub-image parameter sets.
25. 23. The apparatus of claim 22, The device, wherein information indicating the size of the sub-image-based sub-DPB is obtained from a data structure signaled in a PPS or SEI message, the data structure including the maximum size of each layer-based and sub-image-based associated with each layer of the sub-image.
26. 23. The apparatus of claim 22, The device, wherein information indicating the size of the sub-image-based sub-DPB is obtained from a data structure signaled in a PPS or SEI message, the data structure including the width and height of each layer-based and sub-image-based sub-DPB associated with each layer of the sub-image.
27. 23. The apparatus of claim 22, The device, wherein information indicating the size of the sub-image-based sub-DPB is obtained from a data structure signaled in a PPS or SEI message, the data structure indicating the level of each layer of the layered sub-image and the level of each of the plurality of sub-images including a level indicating the value of a syntax element of the respective sub-image or a set of predefined constraints for each layer of the layered sub-image.
28. 23. The apparatus of claim 22, Dividing a decoded picture buffer (DPB) into a plurality of layer-based sub-DPBs; dividing each layer-based sub-DPB into respective layer-based and sub-image-based sub-DPBs; The apparatus is further configured to perform the following:
Citation Information
Patent Citations
Method and apparatus for managing buffer for encoding and decoding multi-layer video
CN106105210A
Decoded picture buffer operation for video coding
JP2016528804A
Signaling and Derivation of Coded Picture Buffer Parameters
JP2017510100A