Temporal prediction of block partition parameters

By introducing multi-type tree depth value configuration and partition mode selection into video encoding devices, and combining CABAC and temporally adjacent block partition parameter prediction, the problem of low block partition parameter prediction efficiency in video encoding systems is solved, thereby improving the compression performance of video data.

CN121444444APending Publication Date: 2026-01-30INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480044521.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-05-11
Filing Date
2024-05-10
Publication Date
2026-01-30

AI Technical Summary

Technical Problem

Existing video coding systems suffer from inefficiency and insufficient accuracy in temporal prediction of block partitioning parameters, especially lacking effective prediction methods for configuring multi-type tree depth values ​​and selecting partitioning modes.

Method used

By introducing configuration of multiple types of tree depth values ​​and partition mode selection in video encoding devices, combined with context-adaptive binary arithmetic coding (CABAC), partition parameters are predicted and refined based on the candidate prediction list of partition parameters of temporally adjacent blocks, and the depth parameters are adjusted by the splitting types of quadtree splitting, ternary splitting and binary splitting.

Benefits of technology

It improves the efficiency and accuracy of block partitioning parameter prediction in video coding systems, enhances the compression performance of video data, and reduces the demand for storage and transmission bandwidth.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121444444A_ABST
    Figure CN121444444A_ABST
Patent Text Reader

Abstract

Systems, methods, and instrumentalities are disclosed for enhancing temporal prediction of block partition parameters. A video decoding apparatus may obtain, for a video block, a partition parameter candidate prediction list including a plurality of partition modes. A device may receive an index indicating a partition mode selected from a plurality of partition modes. The device may apply the selected partition mode to the video block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications This application claims the benefit of European Provisional Patent Application No. 23315193.5, filed on May 11, 2023, the contents of which are hereby incorporated by reference. Background Technology

[0002] Video coding systems can be used to compress digital video signals, for example, by reducing the storage and / or transmission bandwidth required for such signals. Video coding systems can include, for example, block-based, wavelet-based, and / or object-based systems. Summary of the Invention

[0003] Systems, methods, and means for enhancing the temporal prediction of block partitioning parameters are disclosed. A video decoding device can obtain a list of candidate partitioning parameter predictions, including multiple partitioning modes, for a video block. The device can receive an index indicating the partitioning mode to be selected from the multiple partitioning modes. The device can apply the selected partitioning mode to the video block.

[0004] The device can identify multiple temporally adjacent blocks within a video block. A candidate prediction list for partitioning parameters can be obtained based on these temporally adjacent blocks.

[0005] The first of multiple partition modes can be associated with a bypass partition. The device can determine that the selected partition mode includes the first partition mode. The device can apply the first partition mode to a video block. The second of multiple partition modes can be associated with an enabled partition. Enabling a partition can include configuring a multi-type tree depth value. The device can determine that the selected partition mode includes the second partition mode. The device can apply the second partition mode to a video block based on the multi-type tree depth value.

[0006] The device can sort multiple partition patterns based on the prediction parameters of video blocks. The device can determine the selected partition pattern based on the sorted multiple partition patterns. The index indicating the partition pattern selected from the multiple partition patterns can be encoded at the coding tree unit level. The index indicating the partition pattern selected from the multiple partition patterns can be encoded based on context-adaptive binary arithmetic coding (CABAC).

[0007] A video encoding device can determine a candidate prediction list of partitioning parameters for a video block, including multiple partitioning modes. The device can select a partitioning mode from the multiple partitioning modes. The device can apply the selected partitioning mode to the video block. The device can include an index in the video data indicating the partitioning mode selected from the multiple partitioning modes.

[0008] The device can identify multiple temporally adjacent blocks within a video block. A candidate prediction list for partitioning parameters can be obtained based on these temporally adjacent blocks.

[0009] The first of multiple partition modes can be associated with a bypass partition. The device can determine that the selected partition mode includes the first partition mode. The device can apply the first partition mode to a video block. The second of multiple partition modes can be associated with an enabled partition. Enabling a partition can include configuring a multi-type tree depth value. The device can determine that the selected partition mode includes the second partition mode. The device can apply the second partition mode to a video block based on the multi-type tree depth value.

[0010] The device can sort multiple partition patterns based on the prediction parameters of video blocks. The device can determine the selected partition pattern based on the sorted multiple partition patterns. The index indicating the partition pattern selected from the multiple partition patterns can be encoded at the coding tree unit level. The index indicating the partition pattern selected from the multiple partition patterns can be encoded based on context-adaptive binary arithmetic coding (CABAC).

[0011] Systems, methods, and means for enhancing the temporal prediction of block partitioning parameters are disclosed. In an example, a device, such as a video decoding apparatus, can predict block partitioning parameters based on temporal partitioning information from a reference image. The apparatus can obtain partitioning parameter residuals. The apparatus can refine the predicted partitioning parameters based on the partitioning parameter residuals. The partitioning parameter residuals can be obtained based on an instruction indicating the use of the residuals to refine the partitioning parameter prediction. Partitioning parameters may include split partitioning information and / or maximum multitree depth information. Partitioning parameters can be predicted using a depth parameter associated with the block size. The depth parameter can be adjusted based on splitting types including quadtree splits, ternary splits, and / or binary splits.

[0012] Systems, methods, and means for enhancing the temporal prediction of block partitioning parameters are disclosed. In examples, a device, such as a video encoding apparatus, can predict block partitioning parameters based on temporal partitioning information from a reference image. The apparatus can obtain partitioning parameter residuals. The apparatus can include an indication of the partitioning parameter residuals in the video data. Based on the obtained partitioning parameter residuals, the apparatus can include an indication in the video data that uses the residuals to refine partitioning parameter prediction is enabled. Partitioning parameters can include split partitioning information and / or maximum multitree depth information. Partitioning parameters can be predicted using a depth parameter associated with the block size. The depth parameter can be adjusted based on a splitting type including quadtree splits, ternary splits, and / or binary splits.

[0013] The systems, methods, and means described herein may relate to a decoder. In examples, the systems, methods, and means described herein may relate to an encoder. In examples, the systems, methods, and means described herein may relate to a signal (e.g., from an encoder and / or received by a decoder). A computer-readable medium may include instructions for causing one or more processors to perform the methods described herein. A computer program product may include instructions that, when executed by one or more processors, cause one or more processors to perform the methods described herein. Attached Figure Description

[0014] Figure 1A This is a system diagram illustrating an example communication system in which one or more of the disclosed embodiments may be implemented.

[0015] Figure 1B It is shown that, according to the embodiment, it is possible to Figure 1A The diagram shows a system diagram of an example wireless transmit / receive unit (WTRU) used in the communication system shown.

[0016] Figure 1C It is shown that, according to the embodiment, it is possible to Figure 1A The system diagram shows an example radio access network (RAN) and an example core network (CN) used in the communication system shown.

[0017] Figure 1D It is shown that, according to the embodiment, it is possible to Figure 1A The system diagram shows another example RAN and another example CN used in the communication system shown.

[0018] Figure 2 An example video encoder is shown.

[0019] Figure 3 An example video decoder is shown.

[0020] Figure 4 Examples of systems in which various aspects and examples can be implemented are shown.

[0021] Figure 5 The diagram illustrates various tree splitting patterns.

[0022] Figure 6 The signaling mechanism for partition splitting information in a quadtree with a nested multi-type tree-coded tree structure is illustrated.

[0023] Figure 7 An example of a quadtree with a nested multi-type tree coding block structure is shown. Detailed Implementation

[0024] A more detailed understanding can be obtained through the following description, which is given with reference to the accompanying drawings and examples.

[0025] Figure 1A This diagram illustrates an example communication system 100, in which one or more of the disclosed embodiments can be implemented. Communication system 100 can be a multiple access system providing content such as voice, data, video, messaging, and broadcasting to multiple wireless users. Communication system 100 enables multiple wireless users to access such content by sharing system resources, including wireless bandwidth. For example, communication system 100 can employ one or more channel access methods, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal FDMA (OFDMA), Single Carrier FDMA (SC-FDMA), Zero-Tail Unique Word DFT Spread Spectrum OFDM (ZT-UW-DFT-S-OFDM), Unique Word OFDM (UW-OFDM), Resource Block Filtered OFDM, Filter Bank Multicarrier (FBMC), etc.

[0026] like Figure 1A As shown, the communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RAN 104 / 113, CN 106 / 115, Public Switched Telephone Network (PSTN) 108, Internet 110, and other networks 112. However, it should be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, and 102d may be any type of device configured to operate and / or communicate in a wireless environment. For example, WTRUs 102a, 102b, 102c, and 102d (any of which can be referred to as a "station" and / or "STA") can be configured to transmit and / or receive wireless signals and can include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearable devices, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated processing chain environments), consumer electronics devices, devices operating on commercial and / or industrial wireless networks, etc. Any of WTRUs 102a, 102b, 102c, and 102d can be interchangeably referred to as a UE.

[0027] The communication system 100 may also include base station 114a and / or base station 114b. Each of base stations 114a and 114b may be any type of device configured to wirelessly interface with at least one of WTRUs 102a, 102b, 102c, and 102d to facilitate access to one or more communication networks, such as CN 106 / 115, the Internet 110, and / or other networks 112. For example, base stations 114a and 114b may be base transceiver stations (BTS), node B, eNode-B, home node B, home eNode-B, NB, NR node B, site controller, access point (AP), or wireless router. Although base stations 114a and 114b are depicted as single elements, it should be understood that base stations 114a and 114b may include any number of interconnected base stations and / or network elements.

[0028] Base station 114a may be part of RAN 104, and RAN 104 / 113 may also include other base stations and / or network elements (not shown), such as base station controllers (BSCs), radio network controllers (RNCs), relay nodes, etc. Base station 114a and / or base station 114b may be configured to transmit and / or receive radio signals on one or more carrier frequencies, which may be referred to as cells (not shown). These frequencies may be in licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide coverage of a specific geographic area, which may be relatively fixed or may change over time. A cell may also be divided into cell sectors. For example, the cell associated with base station 114a may be divided into three sectors. Thus, in one embodiment, base station 114a may include three transceivers, i.e., one transceiver per sector of the cell. In one embodiment, base station 114a may employ multiple-input multiple-output (MIMO) technology and may utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.

[0029] Base stations 114a and 114b can communicate with one or more of WTRUs 102a, 102b, 102c, and 102d via air interface 116. Air interface 116 can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). Any suitable radio access technology (RAT) can be used to establish air interface 116.

[0030] More specifically, as described above, the communication system 100 can be a multiple access system and can employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, base stations 114a and WTRUs 102a, 102b, and 102c in RAN 104 / 113 can implement wireless technologies such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which can establish air interfaces 115 / 116 / 117 using Wideband CDMA (WCDMA). WCDMA may include communication protocols such as High-Speed ​​Packet Access (HSPA) and / or evolved HSPA (HSPA+). HSPA may include High-Speed ​​Downlink (DL) Packet Access (HSDPA) and / or High-Speed ​​UL Packet Access (HSUPA).

[0031] In one embodiment, base station 114a and WTRUs 102a, 102b, 102c can implement wireless technologies such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which can use Long Term Evolution (LTE) and / or LTE-A Advanced (LTE-A) and / or LTE-A Pro Advanced (LTE-A Pro) to establish air interface 116.

[0032] In one embodiment, base station 114a and WTRUs 102a, 102b, 102c can implement radio technologies such as NR wireless access, which can establish an air interface 116 using a new radio (NR).

[0033] In this embodiment, base station 114a and WTRUs 102a, 102b, and 102c can implement various radio access technologies. For example, base station 114a and WTRUs 102a, 102b, and 102c can jointly implement LTE radio access and NR radio access, for example, using the dual connectivity (DC) principle. Therefore, the air interface utilized by WTRUs 102a, 102b, and 102c can be characterized by various types of radio access technologies and / or transmissions sent to / from various types of base stations (e.g., eNBs and gNBs).

[0034] In other embodiments, base station 114a and WTRUs 102a, 102b, 102c can implement radio technologies such as IEEE 802.11 (i.e., Wi-Fi), IEEE 802.16 (i.e., WiMAX), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Provisional Standard 2000 (IS-2000), Provisional Standard 95 (IS-95), Provisional Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rate GSM Evolution (EDGE), GSMEDGE (GERAN), etc.

[0035] For example, Figure 1A Base station 114b can be a wireless router, home node B, home eNode-B, or access point, and can utilize any suitable RAT to facilitate wireless connectivity in a local area, such as commercial locations, homes, vehicles, campuses, industrial facilities, air corridors (e.g., for drone use), roads, etc. In one embodiment, base station 114b and WTRUs 102c, 102d can implement radio technologies such as IEEE 802.11 to establish a wireless local area network (WLAN). In one embodiment, base station 114b and WTRUs 102c, 102d can implement radio technologies such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, base station 114b and WTRUs 102c, 102d can utilize cellular-based RATs (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish picocells or femtocells. Figure 1A As shown, base station 114b can have a direct connection to Internet 110. Therefore, it is not required that base station 114b access Internet 110 via CN 106 / 115.

[0036] RAN 104 / 113 can communicate with CN 106 / 115, which can be any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRUs 102a, 102b, 102c, and 102d. Data can have different Quality of Service (QoS) requirements, such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, etc. CN 106 / 115 can provide call control, billing services, location-based services, prepaid calling, internet connectivity, video distribution, and / or perform advanced security functions such as user authentication. Although in Figure 1AAlthough not shown, it should be understood that RAN104 / 113 and / or CN 106 / 115 can communicate directly or indirectly with other RANs using the same RAT as RAN 104 / 113 or a different RAT. For example, in addition to connecting to RAN 104 / 113, which may utilize NR radio technology, CN 106 / 115 can also communicate with another RAN (not shown) using GSM, UMTS, CDMA 2000, WiMAX, E-UTRA, or WiFi radio technology.

[0037] CN 106 / 115 can also serve as a gateway for WTRU 102a, 102b, 102c, 102d to access PSTN 108, the Internet 110, and / or other networks 112. PSTN 108 may include a circuit-switched telephone network providing Common Old-Style Telephone Service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices using common communication protocols such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) from the TCP / IP Internet Protocol suite. Network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, network 112 may include another CN connected to one or more RANs, which may use the same RAT as RAN 104 / 113 or a different RAT.

[0038] Some or all of the WTRUs 102a, 102b, 102c, and 102d in communication system 100 may include multimodal capabilities (e.g., WTRUs 102a, 102b, 102c, and 102d may include multiple transceivers for communicating with different wireless networks via different wireless links). For example, Figure 1A The WTRU 102c shown can be configured to communicate with base station 114a, which can use cellular-based radio technology, and to communicate with base station 114b, which can use IEEE 802 radio technology.

[0039] Figure 1B This shows a system diagram of an example WTRU 102. (See diagram below.) Figure 1B As shown, among other things, WTRU 102 may include a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power supply 134, a Global Positioning System (GPS) chipset 136, and / or other peripherals 138. It should be understood that WTRU 102 may include any sub-combination of the foregoing elements while remaining consistent with the embodiments.

[0040] Processor 118 may be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. Processor 118 may perform signal encoding, data processing, power control, input / output processing, and / or any other functions that enable WTRU 102 to operate in a wireless environment. Processor 118 may be coupled to transceiver 120, and transceiver 120 may be coupled to transmitting / receiving element 122. Although Figure 1B The processor 118 and transceiver 120 are depicted as separate components, but it should be understood that the processor 118 and transceiver 120 may be integrated together in an electronic package or chip.

[0041] The transmitting / receiving element 122 can be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) via the air interface 116. For example, in one embodiment, the transmitting / receiving element 122 can be an antenna configured to transmit and / or receive RF signals. In one embodiment, the transmitting / receiving element 122 can be, for example, a transmitter / detector configured to transmit and / or receive IR, UV, or visible light signals. In yet another embodiment, the transmitting / receiving element 122 can be configured to transmit and / or receive both RF and optical signals. It should be understood that the transmitting / receiving element 122 can be configured to transmit and / or receive any combination of wireless signals.

[0042] Although the transmitting / receiving element 122 is in Figure 1B While depicted as a single element, WTRU 102 may include any number of transmit / receive elements 122. More specifically, WTRU 102 may employ MIMO technology. Thus, in one embodiment, WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals via air interface 116.

[0043] Transceiver 120 can be configured to modulate signals transmitted by transmit / receive element 122 and demodulate signals received by transmit / receive element 122. As described above, WTRU 102 can have multimode capability. Therefore, transceiver 120 can include multiple transceivers to enable WTRU 102 to communicate via multiple RATs, such as NR and IEEE 802.11.

[0044] The processor 118 of WTRU 102 can be coupled to a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) unit or an organic light-emitting diode (OLED) display unit) and can receive user input data therefrom. The processor 118 can also output user data to the speaker / microphone 124, keypad 126, and / or display / touchpad 128. Furthermore, the processor 118 can access and store information from any type of suitable memory, such as non-removable memory 130 and / or removable memory 132. Non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. Removable memory 132 may include a user identity module (SIM) card, a memory stick, a secure digital storage (SD) card, etc. In other embodiments, the processor 118 can access and store information from memory that is not physically located on WTRU 102 (such as on a server or home computer (not shown)).

[0045] The processor 118 may receive power from the power supply 134 and may be configured to distribute and / or control power to other components in the WTRU 102. The power supply 134 may be any suitable device for powering the WTRU 102. For example, the power supply 134 may include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.

[0046] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) about the current location of the WTRU 102. In addition to, or instead of, information from the GPS chipset 136, the WTRU 102 may receive location information from base stations (e.g., base stations 114a, 114b) via the air interface 116, and / or determine its location based on the timing of signals received from two or more nearby base stations. It should be understood that the WTRU 102 may acquire location information using any suitable location determination method while remaining consistent with the embodiments.

[0047] The processor 118 may also be coupled to other peripherals 138, which may include one or more software and / or hardware modules providing additional features, functions, and / or wired or wireless connectivity. For example, peripherals 138 may include accelerometers, electronic compasses, satellite transceivers, digital cameras (for photos and / or videos), Universal Serial Bus (USB) ports, vibration devices, television transceivers, hands-free headsets, Bluetooth® modules, FM radio units, digital music players, media players, video game player modules, internet browsers, virtual reality and / or augmented reality (VR / AR) devices, activity trackers, etc. Peripherals 138 may include one or more sensors, which may be one or more of the following: gyroscopes, accelerometers, Hall effect sensors, magnetometers, orientation sensors, proximity sensors, temperature sensors, time sensors, geolocation sensors; altimeters, light sensors, touch sensors, magnetometers, barometers, attitude sensors, biometric sensors, and / or humidity sensors.

[0048] WTRU 102 may include a full-duplex radio for which the transmission and reception of some or all signals (e.g., associated with a specific subframe of both UL (e.g., for transmission) and downlink (e.g., for reception)) may be concurrent and / or simultaneous. The full-duplex radio may include an interference management unit to reduce and / or substantially eliminate self-interference through hardware (e.g., chokes) or through signal processing by a processor (e.g., a separate processor (not shown) or through processor 118). In one embodiment, WTRU 102 may include a half-duplex radio for which the transmission and reception of some or all signals (e.g., associated with a specific subframe of either UL (e.g., for transmission) or downlink (e.g., for reception)) may be concurrent and / or simultaneous.

[0049] Figure 1C This is a system diagram illustrating RAN 104 and CN 106 according to an embodiment. As described above, RAN 104 can communicate with WTRUs 102a, 102b, and 102c via air interface 116 using E-UTRA radio technology. RAN 104 can also communicate with CN 106.

[0050] RAN 104 may include eNode-B 160a, 160b, and 160c; however, it should be understood that RAN 104 may include any number of eNode-Bs while remaining consistent with the embodiments. Each eNode-B 160a, 160b, and 160c may include one or more transceivers for communicating with WTRU 102a, 102b, and 102c via air interface 116. In one embodiment, eNode-B 160a, 160b, and 160c may implement MIMO technology. Therefore, for example, eNode-B 160a may use multiple antennas to transmit radio signals to and / or receive radio signals from WTRU 102a.

[0051] Each of the eNode-B 160a, 160b, and 160c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, etc. Figure 1C As shown, eNode-B 160a, 160b, and 160c can communicate with each other via the X2 interface.

[0052] Figure 1C The CN 106 shown may include a Mobility Management Entity (MME) 162, a Serving Gateway (SGW) 164, and a Packet Data Network (PDN) Gateway (or PGW) 166. While each of the foregoing elements is described as part of CN 106, it should be understood that any of these elements may be owned and / or operated by an entity other than a CN operator.

[0053] The MME 162 can connect to each of the eNode-Bs 162a, 162b, and 162c in RAN 104 via the S1 interface and can act as a control node. For example, the MME 162 can be responsible for authenticating users of WTRUs 102a, 102b, and 102c, bearer activation / deactivation, selecting a specific serving gateway during the initial attachment of WTRUs 102a, 102b, and 102c, etc. The MME 162 can provide control plane functions for handover between RAN 104 and other RANs (not shown) employing other radio technologies such as GSM and / or WCDMA.

[0054] The SGW 164 can connect to each of the eNode-Bs 160a, 160b, and 160c in RAN 104 via the S1 interface. The SGW 164 can typically route and forward user data packets to / from WTRUs 102a, 102b, and 102c. The SGW 164 can perform other functions such as anchoring the user plane during inter-eNode-B handover, triggering paging when DL data is available for WTRUs 102a, 102b, and 102c, and managing and storing the context of WTRUs 102a, 102b, and 102c.

[0055] SGW 164 can connect to PGW 166, which can provide WTRU 102a, 102b, 102c with access to packet-switched networks such as Internet 110, so as to facilitate communication between WTRU 102a, 102b, 102c and IP-enabled devices.

[0056] CN 106 can facilitate communication with other networks. For example, CN 106 can provide WTRU 102a, 102b, 102c with access to a circuit-switched network such as PSTN 108, facilitating communication between WTRU 102a, 102b, 102c and traditional landline communication equipment. For example, CN 106 may include, or be able to communicate with, an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between CN 106 and PSTN 108. Furthermore, CN 106 can provide WTRU 102a, 102b, 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.

[0057] Despite WTRU in Figure 1A-1D While described as a wireless terminal, it is conceivable that in some representative embodiments, such a terminal may use (e.g., temporarily or permanently) a wired communication interface with a communication network.

[0058] In a representative embodiment, another network 112 may be a WLAN.

[0059] In Infrastructure Basic Services Set (BSS) mode, a WLAN may have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP may have access or an interface to a distributed system (DS) or another type of wired / wireless network that transmits traffic to and / or out of the BSS. Traffic originating outside the BSS destined for a STA can be delivered to the AP. Traffic from a STA to a destination outside the BSS can be sent to the AP for delivery to the appropriate destination. For example, traffic between STAs within the BSS can be transmitted via the AP, where the source STA can send traffic to the AP, and the AP can deliver traffic to the destination STA. Traffic between STAs within the BSS can be considered and / or referred to as peer-to-peer traffic. Peer-to-peer traffic can be transmitted between source and destination STAs (e.g., directly between them) using Direct Link Establishment (DLS). In some representative embodiments, the DLS may use 802.11e DLS or 802.11z Tunneled DLS (TDLS). A WLAN using Standalone BSS (IBSS) mode may not have an access point (AP), and STAs within the IBSS or using the IBSS (e.g., all STAs) can communicate directly with each other. The IBSS communication mode is sometimes referred to here as an "ad-hoc" communication mode.

[0060] When operating in 802.11ac infrastructure mode or a similar mode, the AP can transmit beacons on a fixed channel, such as the primary channel. The primary channel can be of a fixed width (e.g., a wide bandwidth of 20 MHz) or dynamically set via signaling. The primary channel can be the operating channel of the BSS and can be used by the STA to establish a connection with the AP. In some representative embodiments, Carrier Sense Multiple Access with Collision Avoidance (CSMA / CA) can be implemented, for example in an 802.11 system. For CSMA / CA, the STAs including the AP (e.g., each STA) can sense the primary channel. If a particular STA senses / detects and / or determines that the primary channel is busy, that particular STA can back off. A single STA (e.g., only one station) can transmit at any given time within a given BSS.

[0061] High-throughput (HT) STAs can communicate using a 40 MHz wide channel, for example, by combining a primary 20 MHz channel with adjacent or non-adjacent 20 MHz channels.

[0062] Very High Throughput (VHT) STAs can support 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz wide channels. 40 MHz and / or 80 MHz channels can be formed by combining consecutive 20 MHz channels. A 160 MHz channel can be formed by combining eight consecutive 20 MHz channels, or by combining two non-consecutive 80 MHz channels, which can be referred to as an 80+80 configuration. For the 80+80 configuration, after channel coding, the data can be divided into two streams by a segment parser. Each stream can be processed separately using Inverse Fast Fourier Transform (IFFT) and time-domain processing. The streams can be mapped onto the two 80 MHz channels, and the data can be transmitted by the transmitting STA. At the receiver of the receiving STA, the operation of the 80+80 configuration can be reversed, and the combined data can be sent to the Media Access Control (MAC).

[0063] 802.11af and 802.11ah support sub-1 GHz operating modes. The channel operating bandwidth and carrier are reduced in 802.11af and 802.11ah compared to those used in 802.11n and 802.11ac. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV whitespace (TVWS) spectrum, while 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. According to representative embodiments, 802.11ah can support metering-type control / machine-type communications, such as MTC devices in macro coverage areas. MTC devices may have certain capabilities, such as limited capabilities, including support (e.g., only support) certain and / or limited bandwidths. MTC devices may include batteries with a battery life exceeding a threshold (e.g., to maintain very long battery life).

[0064] WLAN systems that can support multiple channels and channel bandwidths (such as 802.11n, 802.11ac, 802.11af, and 802.11ah) include a channel that can be designated as the primary channel. The primary channel can have a bandwidth equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel can be set and / or limited by the STAs operating in the BSS that support the minimum bandwidth operating mode. In the example of 802.11ah, for STAs that support (e.g., only support) the 1 MHz mode (e.g., MTC type devices), the primary channel can be 1 MHz wide, even if the AP and other STAs in the BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operating modes. Carrier Sense and / or Network Allocation Vector (NAV) settings may depend on the status of the primary channel. If the primary channel is busy transmitting to the AP, for example due to an STA (only supporting the 1 MHz operating mode), the entire available band can be considered busy, even if most of the band remains idle and can be available.

[0065] In the United States, the available frequency band for 802.11ah is from 902 MHz to 928 MHz. In South Korea, the available frequency band is from 917.5 MHz to 923.5 MHz. In Japan, the available frequency band is from 916.5 MHz to 927.5 MHz. The total available bandwidth for 802.11ah is 6 MHz to 26 MHz, depending on the country code.

[0066] Figure 1D This is a system diagram illustrating RAN 113 and CN 115 according to an embodiment. As described above, RAN 113 can communicate with WTRUs 102a, 102b, and 102c via air interface 116 using NR radio technology. RAN 113 can also communicate with CN 115.

[0067] RAN 113 may include gNBs 180a, 180b, and 180c; however, it should be understood that RAN 113 may include any number of gNBs while remaining consistent with the embodiments. gNBs 180a, 180b, and 180c may each include one or more transceivers for communicating with WTRUs 102a, 102b, and 102c via air interface 116. In one embodiment, gNBs 180a, 180b, and 180c may implement MIMO technology. For example, gNBs 180a and 108b may utilize beamforming to transmit signals to and / or receive signals from gNBs 180a, 180b, and 180c. Therefore, for example, gNB 180a may use multiple antennas to transmit radio signals to and / or receive radio signals from WTRU 102a. In one embodiment, gNBs 180a, 180b, and 180c may implement carrier aggregation technology. For example, gNB 180a can transmit multiple component carriers (not shown) to WTRU 102a. A subset of these component carriers may be on unlicensed spectrum, while the remaining component carriers may be on licensed spectrum. In one embodiment, gNBs 180a, 180b, and 180c may implement Coordinated Multipoint (CoMP) technology. For example, WTRU 102a may receive coordinated transmissions from gNBs 180a and 180b (and / or gNB 180c).

[0068] WTRUs 102a, 102b, and 102c can communicate with gNB180a, 180b, and 180c using transmissions associated with scalable digitization. For example, the OFDM symbol spacing and / or OFDM subcarrier spacing can differ for different transmissions, different cells, and / or different portions of the radio transmission spectrum. WTRUs 102a, 102b, and 102c can communicate with gNB180a, 180b, and 180c using subframes or transmission time intervals (TTIs) of various or scalable lengths (e.g., containing a variable number of OFDM symbols and / or a continuously variable absolute time).

[0069] gNB180a, 180b, and 180c can be configured to communicate with WTRU 102a, 102b, and 102c in standalone and / or non-standalone configurations. In standalone configuration, WTRU 102a, 102b, and 102c can communicate with gNB180a, 180b, and 180c without accessing other RANs (e.g., eNode-B 160a, 160b, and 160c). In standalone configuration, WTRU 102a, 102b, and 102c can utilize one or more of gNB180a, 180b, and 180c as mobility anchors. In standalone configuration, WTRU 102a, 102b, and 102c can communicate with gNB180a, 180b, and 180c using signals in unlicensed frequency bands. In a non-standalone configuration, WTRUs 102a, 102b, and 102c can communicate / connect with gNBs 180a, 180b, and 180c, while also communicating / connecting with another RAN such as eNode-Bs 160a, 160b, and 160c. For example, WTRUs 102a, 102b, and 102c can implement DC principles to communicate substantially simultaneously with one or more gNBs 180a, 180b, and 180c, as well as one or more eNode-Bs 160a, 160b, and 160c. In a non-standalone configuration, eNode-Bs 160a, 160b, and 160c can act as mobility anchors for WTRUs 102a, 102b, and 102c, and gNBs 180a, 180b, and 180c can provide additional coverage and / or throughput for serving WTRUs 102a, 102b, and 102c.

[0070] Each of gNB180a, 180b, and 180c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, network slicing support, dual connectivity, interoperability between NR and E-UTRA, routing user plane data to User Plane Functions (UPF) 184a and 184b, and routing control plane information to Access and Mobility Management Functions (AMF) 182a and 182b, etc. Figure 1D As shown, gNB180a, 180b, and 180c can communicate with each other via the Xn interface.

[0071] Figure 1DThe CN 115 shown may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and may include a Data Network (DN) 185a, 185b. While each of the foregoing elements is described as part of the CN 115, it should be understood that any of these elements may be owned and / or operated by an entity other than a CN operator.

[0072] AMF 182a and 182b can connect to one or more gNBs 180a, 180b, and 180c in RAN 113 via the N2 interface and can act as control nodes. For example, AMF 182a and 182b can be responsible for authenticating users of WTRU 102a, 102b, and 102c, supporting network slicing (e.g., handling different PDU sessions with different requirements), selecting specific SMF 183a and 183b, managing registration areas, terminating NAS signaling, mobility management, and so on. AMF 182a and 182b can use network slicing to customize CN support for WTRU 102a, 102b, and 102c based on the service types used by WTRU 102a, 102b, and 102c. For example, different network slices can be established for different use cases, such as services relying on Ultra Reliable Low Latency (URLLC) access, services relying on Enhanced Massive Mobile Broadband (eMBB) access, and services for Machine-Type Communication (MTC) access. AMF 162 can provide control plane functions for handover between RAN 113 and other RANs (not shown) employing other radio technologies such as LTE, LTE-A, LTE-A Pro and / or non-3GPP access technologies such as WiFi.

[0073] SMFs 183a and 183b can connect to AMFs 182a and 182b in CN 115 via the N11 interface. SMFs 183a and 183b can also connect to UPFs 184a and 184b in CN 115 via the N4 interface. SMFs 183a and 183b can select and control UPFs 184a and 184b, and configure the routing of services through UPFs 184a and 184b. SMFs 183a and 183b can perform other functions, such as managing and allocating UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, and providing downlink data notifications. PDU session types can be IP-based, non-IP-based, Ethernet-based, etc.

[0074] UPF 184a and 184b can be connected to one or more gNB180a, 180b, and 180c in RAN 113 via the N3 interface. This interface provides WTRU 102a, 102b, and 102c with access to packet-switched networks (such as Internet 110) to facilitate communication between WTRU 102a, 102b, 102c and IP-enabled devices. UPF 184 and 184b can perform other functions such as routing and forwarding packets, enforcing user plane policies, supporting multi-destination PDU sessions, handling user plane QoS, buffering downlink packets, and providing mobility anchoring.

[0075] CN 115 can facilitate communication with other networks. For example, CN 115 may include, or be able to communicate with, an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between CN 115 and PSTN 108. Furthermore, CN 115 can provide WTRUs 102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, WTRUs 102a, 102b, and 102c may be connected to the local data network (DN) 185a and 185b via the N3 interface to UPFs 184a and 184b and the N6 interface between UPFs 184a and 184b and DNs 185a and 185b.

[0076] Given Figure 1A-1D as well as Figure 1A-1D The corresponding descriptions herein refer to one or more of the following functions: WTRU 102a-d, Base Station 114a-b, eNode-B 160a-c, MME 162, SGW 164, PGW166, gNB 180 ac, AMF 182 ab, UPF 184a-b, SMF 183 ab, DN185a-b, and / or any other device described herein. These functions can be performed by one or more emulation devices (not shown). An emulation device can be one or more devices configured to emulate the functions described herein. For example, an emulation device can be used to test other devices and / or simulate network and / or WTRU functions.

[0077] Simulation devices can be designed to perform tests on one or more other devices in laboratory and / or carrier network environments. For example, one or more simulation devices can perform one or more or all of their functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices within the communication network. One or more simulation devices can perform one or more or all of their functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. For testing purposes, simulation devices can be directly coupled to another device and / or can use over-the-air wireless communication to perform tests.

[0078] One or more simulation devices may perform one or more functions, including all functions, rather than being implemented / deployed as part of a wired and / or wireless communication network. For example, simulation devices may be used to test test scenarios in laboratory and / or non-deployment (e.g., testing) wired and / or wireless communication networks to implement the testing of one or more components. One or more simulation devices may be test devices. Simulation devices may transmit and / or receive data using direct RF coupling and / or wireless communication via RF circuitry (e.g., which may include one or more antennas).

[0079] This application describes multiple aspects, including tools, features, examples, models, methods, etc. Many of these aspects are described in detail, and often in a manner that may sound restrictive, at least to illustrate individual characteristics. However, this is for the purpose of clarity and does not limit the application or scope of those aspects. In fact, all the different aspects can be combined and interchanged to provide other aspects. Furthermore, aspects can also be combined and interchanged with aspects described in earlier applications.

[0080] The aspects described and envisioned in this application can be implemented in many different forms. Figures 5-7 Examples can be provided, but other examples are also expected. Figures 5-7 The discussion does not limit the scope of implementation. At least one aspect generally relates to video encoding and decoding, and at least one other aspect generally relates to the transmission of generated or encoded bitstreams. These and other aspects can be implemented as methods, apparatus, computer-readable storage media having instructions thereon stored thereon for encoding or decoding video data according to any of the described methods, and / or computer-readable storage media having bitstreams generated according to any of the described methods stored thereon.

[0081] In this application, the terms “reconstruction” and “decoding” are used interchangeably, the terms “pixel” and “sample” are used interchangeably, and the terms “image”, “picture” and “frame” are used interchangeably.

[0082] This document describes various methods, and each method includes one or more steps or actions for implementing the method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions can be modified or combined. Additionally, terms such as "first," "second," etc., can be used in various examples to modify elements, components, steps, operations, etc., such as, for example, "first decoding" and "second decoding." Unless specifically required, the use of such terms does not imply a reordering of the modified operations. Thus, in this example, the first decoding does not need to be performed before the second decoding and can occur, for example, before, during, or in a time period overlapping with the second decoding.

[0083] The various methods and other aspects described in this application can be used to modify, for example... Figure 2 and Figure 3 The illustrated video encoder 200 and decoder 300 modules are, for example, decoding modules. Furthermore, the subject matter disclosed herein can be applied to, for example, any type, format, or version of video encoding, whether described in standards or recommendations, whether pre-existing or future-developed, and to extensions applied to any such standards and recommendations. Unless otherwise stated or technically excluded, the aspects described in this application may be used individually or in combination.

[0084] Various numerical values ​​are used in the examples described in this application, such as those referenced in the examples shown and discussed relative to Figures 9 and 10, Equation (5), Table 1, the number of intra-frame modes, the downsampling rate, the number of adjustment values, etc. These and other specific values ​​are for illustrative purposes only, and the aspects described are not limited to these specific values.

[0085] Figure 2 This is a diagram illustrating an example video encoder. Variations of the example encoder 200 are envisioned, but for clarity, encoder 200 is described below without describing all anticipated variations.

[0086] Before being encoded, the video sequence may undergo pre-coding processing (201), such as applying color transformations to the input color image (e.g., a conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing remapping of the input image components to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization with one of the color components). Metadata may be associated with the pre-processing and attached to the bitstream.

[0087] In encoder 200, the image is encoded by encoder elements as described below. The image to be encoded is partitioned (202) and processed in units, for example, coding units (CUs). Each unit is encoded using, for example, intra-frame or inter-frame modes. When a unit is encoded in intra-frame mode, it performs intra-frame prediction (260). In inter-frame mode, motion estimation (275) and compensation (270) are performed. The encoder determines (205) which of the intra-frame or inter-frame modes to use to encode the unit and indicates the intra-frame / inter-frame decision by, for example, a prediction mode flag. For example, the prediction residual is calculated by subtracting (210) the prediction block from the original image block.

[0088] The predicted residual is then transformed (225) and quantized (230). The quantized transform coefficients, along with the motion vector and other syntax elements (e.g., image partitioning information), are entropy encoded (245) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass the transform and quantization, i.e., the residual is directly encoded without applying the transform or quantization process.

[0089] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inversely transformed (250) to decode the prediction residual. The decoded prediction residual and the prediction block are combined (255) to reconstruct the image block. A loop filter (265) is applied to the reconstructed image to perform, for example, deblocking / SAO (Sample Adaptive Offset) / ALF (Adaptive Loop Filter) filtering, thereby reducing coding artifacts. The filtered image is stored in a reference image buffer (280).

[0090] Figure 3 This is a diagram illustrating an example video decoder. In the example decoder 300, the bitstream is decoded by decoder elements, as described below. The video decoder 300 typically performs the same operations as... Figure 2 The encoding process described herein is the opposite of the decoding process. Encoder 200 typically also performs video decoding as part of the encoded video data.

[0091] Specifically, the decoder's input includes a video bitstream, which can be generated by the video encoder 200. The bitstream is first entropy decoded (330) to obtain transform coefficients, prediction modes, motion vectors, and other encoded information. Picture partitioning information indicates how the picture is partitioned. Therefore, the decoder can partition (335) the picture based on the decoded picture partitioning information. The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residuals. The decoded prediction residuals and prediction blocks are combined (355) to reconstruct the image blocks. The prediction blocks can be obtained (370) from intra-frame prediction (360) or motion-compensated prediction (i.e., inter-frame prediction) (375). A loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (380). In the example, for a given picture, the contents of the reference picture buffer 380 on the decoder 300 side can be the same as the contents of the reference picture buffer 280 on the encoder 200 side (e.g., for the same picture).

[0092] The decoded image may undergo further post-decoding processing (385), such as inverse color transformation (e.g., a conversion from YCbCr4:2:0 to RGB 4:4:4) or inverse remapping, which is the inverse of the remapping process performed in the pre-encoding process (201). The post-decoding process may use metadata derived in the pre-encoding process and signaled in the bitstream. In the example, the decoded image (e.g., after applying an in-loop filter (365) and / or after post-decoding processing (385), if post-decoding processing is used) may be sent to a display device for presentation to the user.

[0093] Figure 4 This is a block diagram of an example system in which various aspects and examples described herein can be implemented. System 400 can be embodied as a device including the various components described below and configured to perform one or more aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 400 can be embodied individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one example, the processing and encoder / decoder elements of system 400 are distributed across multiple ICs and / or discrete components. In various examples, system 400 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various examples, system 400 is configured to implement one or more aspects described in this document.

[0094] System 400 includes at least one processor 410 configured to execute instructions loaded therein for implementing various aspects, such as those described in this application. Processor 410 may include embedded memory, input / output interfaces, and various other circuit systems as known in the art. System 400 includes at least one memory 420 (e.g., a volatile memory device and / or a non-volatile memory device). System 400 includes a storage device 440, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 440 may include internal storage devices, attached storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices.

[0095] System 400 includes an encoder / decoder module 430 configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 430 may include its own processor and memory. The encoder / decoder module 430 represents one or more modules that may be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both encoding and decoding modules. Additionally, the encoder / decoder module 430 may be implemented as a separate element of system 400, or it may be incorporated within processor 410 as a combination of hardware and software as known to those skilled in the art.

[0096] Program code to be loaded onto processor 410 or encoder / decoder 430 to execute the various aspects described in this document may be stored in storage device 440 and subsequently loaded onto memory 420 for execution by processor 410. According to various examples, one or more of processor 410, memory 420, storage device 440, and encoder / decoder module 430 may store one or more various items during the execution of the processes described in this document. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.

[0097] In the example, the memory within processor 410 and / or encoder / decoder module 430 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other examples, external memory (e.g., the processing device may be processor 410 or encoder / decoder module 430) is used for one or more of these functions. External memory may be memory 420 and / or storage device 440, such as dynamically volatile memory and / or non-volatile flash memory. In several examples, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one example, fast external dynamically volatile memory such as RAM is used as working memory for video encoding and decoding operations.

[0098] As indicated in box 445, inputs to the components of system 400 can be provided through various input devices. Such input devices include, but are not limited to, (i) a radio frequency (RF) section that receives, for example, RF signals transmitted over the air by a broadcaster, (ii) component (COMP) input terminals (or sets of COMP input terminals), (iii) universal serial bus (USB) input terminals, and / or (iv) high-definition multimedia interface (HDMI) input terminals. Figure 4 Other examples not shown include composite video.

[0099] In various examples, the input device of block 405 has associated corresponding input processing elements, as known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or limiting a signal band to a band), (ii) down-converting the selected signal, (iii) further band-limiting to a narrower band to select, for example, a signal band that may be referred to as a channel in some examples, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and / or (vi) demultiplexing to select a desired data packet stream. The RF section of various examples includes one or more elements performing these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include tuners performing various functions among these functions, such as down-converting a received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or baseband. In one set-top box example, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to the desired frequency band. Various examples rearrange the order of the aforementioned (and other) components, remove some of these components, and / or add other components that perform similar or different functions. Adding components may include inserting components between existing components, such as, for example, inserting amplifiers and analog-to-digital converters. In various examples, the RF section includes an antenna.

[0100] USB and / or HDMI terminals may include corresponding interface processors for connecting system 400 to other electronic devices across USB and / or HDMI connections. It should be understood that various aspects of input processing—such as Reed-Solomon error correction—may be implemented as needed, for example, within a separate input processing IC or within processor 410. Similarly, as needed, various aspects of USB or HDMI interface processing may be implemented within a separate interface IC or within processor 410. The demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, processor 410 and encoder / decoder 430, which operate in combination with memory and storage elements to process the data streams as needed for presentation on the output device.

[0101] Various components of system 400 can be provided within an integrated housing, in which the various components can be interconnected and transmit data therebetween using suitable connection means 425, such as internal buses as known in the art, including internal IC (I2C) buses, wiring and printed circuit boards.

[0102] System 400 includes a communication interface 450 that enables communication with other devices via a communication channel 460. The communication interface 450 may include, but is not limited to, a transceiver configured to send and receive data via the communication channel 460. The communication interface 450 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 460 may be implemented within, for example, wired and / or wireless media.

[0103] In various examples, wireless networks, such as Wi-Fi networks (e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers)), are used to stream or otherwise provide data to system 400. In these examples, the Wi-Fi signal is received via a communication channel 460 and a communication interface 450 adapted for Wi-Fi communication. The communication channel 460 in these examples is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other examples use a set-top box to provide streaming data to system 400, delivering data via an HDMI connection to input block 445. Still other examples use an RF connection to input block 445 to provide streaming data to system 400. As indicated above, various examples provide data in a non-streaming manner. Additionally, various examples use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth® networks.

[0104] System 400 can provide output signals to various output devices, including a display 475, a speaker 485, and other peripheral devices 495. Various examples of the display 475 include one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a flexible display, and / or a foldable display. The display 475 can be used in televisions, tablet devices, laptop computers, telephones (mobile phones), or other devices. The display 475 can also be integrated with other components (e.g., as in a smartphone) or standalone (e.g., an external monitor for a laptop computer). In various examples, other peripheral devices 495 include one or more of a standalone digital video disc (or digital universal disc) (DVD, both terms), a disc player, a stereo system, and / or a lighting system. Various examples utilize one or more peripheral devices 495 that provide functionality based on the output of system 400. For example, a disc player performs the function of playing the output of system 400.

[0105] In various examples, signaling is used to communicate control signals between system 400 and display 475, speaker 485, or other peripheral devices 495. This signaling may include AV.Link, Consumer Electronics Control (CEC), or other communication protocols enabling device-to-device control with or without user intervention. Output devices may be communicatively coupled to system 400 via dedicated connections through corresponding interfaces 470, 480, and 490. Alternatively, output devices may be connected to system 400 via communication interface 450 using communication channel 460. Display 475 and speaker 485 may be integrated into a single unit with other components of system 400 in electronic devices, such as, for example, a television set. In various examples, display interface 470 includes a display driver, such as, for example, a timing controller (TCon) chip.

[0106] For example, if the RF portion of input 445 is part of a separate set-top box, then display 475 and speaker 485 can alternatively be separated from one or more other components. In various examples where display 475 and speaker 485 are external components, the output signal can be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output. Examples can be implemented by computer software implemented by processor 410, or by hardware, or by a combination of hardware and software. As a non-limiting example, examples can be implemented by one or more integrated circuits. Memory 420 can have any type suitable for the technical environment and can be implemented using any suitable data storage technology, such as, as a non-limiting example, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. As a non-limiting example, processor 410 can have any type suitable for the technical environment and can comprise one or more of microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.

[0107] Various implementations involve decoding. As used herein, "decoding" can include, for example, performing all or part of a process on a received encoded sequence to produce a final output suitable for display. In various examples, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various examples, such processes also or alternatively include processes performed by a decoder of the various implementations described herein, such as, for a video block, obtaining a candidate prediction list of partitioning parameters comprising multiple partitioning modes; receiving an index indicating the partitioning mode selected from the multiple partitioning modes; and applying the selected partitioning mode to the video block.

[0108] As another example, in one example, "decoding" refers only to entropy decoding; in another example, "decoding" refers only to differential decoding; and in yet another example, "decoding" refers to a combination of entropy decoding and differential decoding. It will be clear, and is considered well understood by those skilled in the art, whether the phrase "decoding process" is intended to refer specifically to a subset of operations or to a broader decoding process, based on the specific context of the description.

[0109] Various implementations involve encoding. In a manner similar to the discussion above regarding “decoding,” “encoding,” as used herein, can include, for example, all or part of a process performed on an input video sequence to produce an encoded bitstream. In various examples, such a process includes one or more processes typically performed by an encoder, such as partitioning, differential coding, transform, quantization, and entropy coding. In various examples, such processes also or alternatively include processes performed by an encoder of the various implementations described herein, such as determining a candidate prediction list of partitioning parameters comprising multiple partitioning modes for a video block; selecting a partitioning mode from the multiple partitioning modes; applying the selected partitioning mode to the video block; and including in the video data an index indicating the partitioning mode selected from the multiple partitioning modes.

[0110] As further examples, in one example, "encoding" refers only to entropy coding; in another, "encoding" refers only to differential coding; and in yet another, "encoding" refers to a combination of differential and entropy coding. It will be clear, and considered well understood by those skilled in the art, whether the phrase "encoding process" is intended to refer specifically to a subset of operations or to a broader encoding process, based on the specific context of the description. Note that the syntactic elements used herein, such as the encoding syntax indicating the temporal prediction of splitting parameters, the usage flag for temporal prediction of splitting parameters at the image level, the constraint flag indicating partitioning limitations, etc., are descriptive terms. Therefore, they do not preclude the use of other syntactic element names or functions.

[0111] When a diagram is presented as a flowchart, it should be understood that it also provides a block diagram of the corresponding device. Similarly, when a diagram is presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / process.

[0112] The implementations and aspects described herein can be implemented, for example, in methods or processes, apparatuses, software programs, data streams, or signals. Even if discussed only in the context of a single implementation (e.g., discussed only as a method), implementations of the discussed features can be implemented in other forms (e.g., apparatuses or programs). Apparatuses can be implemented, for example, in suitable hardware, software, and firmware. The method can be implemented, for example, in a processor, which generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as, for example, computers, cellular phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate information communication between end users.

[0113] References to “an example” or “an instance” or “an implementation” or “an implementation”, as well as their other variations, imply that the specific features, structures, characteristics, etc., described in connection with that example are included in at least one example. Therefore, the appearance of the phrase “in an example” or “in one example” or “in one implementation” or “in one implementation”, and any other variations appearing in various places throughout the application, do not necessarily all refer to the same example.

[0114] Additionally, this application may refer to "determining" various information pieces. Determining information may include, for example, one or more of estimation information, calculation information, prediction information, or information retrieved from memory. Obtaining may include receiving, retrieving, constructing, generating, and / or determining.

[0115] Furthermore, this application may refer to "accessing" various information chips. Accessing information may include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information, or one or more of these.

[0116] Additionally, this application may refer to "receiving" various pieces of information. Like "accessing," receiving is intended to be a broad term. Receiving information may include, for example, accessing information or retrieving information (e.g., from memory) one or more of the following: Furthermore, "receiving" is generally referred to in one way or another during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0117] It will be understood that, for example, in the cases of “A / B,” “A and / or B,” and “at least one of A and B,” any of the following uses of “ / ,” “and / or,” and “at least one of…” are intended to include selecting only the first listed option (A), or only the second listed option (B), or both options (A and B). As another example, in the cases of “A, B, and / or C” and “at least one of A, B, and C,” such wording is intended to include selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A, B, and C). As will be apparent to those skilled in the art and related fields, this can be extended to a large number of listed items.

[0118] Furthermore, as used herein, the word "signal" specifically refers, among other things, to instructing the corresponding decoder to do something. Encoder signals can include, for example, signals that instruct the encoder to select one or more adjustment values, whether implicitly or explicitly determined. Thus, in the example, the same parameters are used on both the encoder and decoder sides. Therefore, for example, the encoder can transmit (explicitly signal) a specific parameter to the decoder so that the decoder can use the same specific parameter. Conversely, if the decoder already has the specific parameter as well as other parameters, signaling can be used without transmission (implicit signaling) to simply allow the decoder to know and select the specific parameter. Bit savings are achieved in various examples by avoiding the transmission of any actual functionality. It will be understood that signaling can be implemented in many ways. For example, in various examples, one or more syntax elements, flags, etc., are used to signal information to the corresponding decoder. Although the verb form of the word "signal" has been used above, the word "signal" can also be used as a noun in this article.

[0119] As will be apparent to those skilled in the art, the implementation can generate various signals, which are formatted to carry information, for example, that can be stored or transmitted. This information may include, for example, instructions for performing a method, or data generated by one of the described implementations. For example, the signal may be formatted to carry a bitstream of the described example. Such a signal may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or baseband signals. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is known, the signal can be transmitted via a variety of different wired or wireless links. The signal may be stored on, or accessed or received from, a processor-readable medium.

[0120] Numerous examples are described herein. Features of the examples may be provided individually or in any combination across various claim classes and types. Furthermore, examples may include one or more of the features, devices, or aspects described herein, individually or in any combination across various claim classes and types. For example, the features described herein may be implemented in a bitstream or signal including information generated as described herein. This information may allow a decoder to decode the bitstream, an encoder according to any of the described embodiments, a bitstream, and / or a decoder. For example, the features described herein may be implemented by creating and / or transmitting and / or receiving and / or decoding a bitstream or signal. For example, the features described herein may be implemented as a method, process, apparatus, medium storing instructions, medium storing data, or signal. For example, the features described herein may be implemented by a television, set-top box, mobile phone, tablet computer, or other electronic device performing decoding. The television, set-top box, mobile phone, tablet computer, or other electronic device may display (e.g., using a monitor, screen, or other type of display) a resulting image (e.g., an image reconstructed from the residual of a video bitstream). The television, set-top box, mobile phone, tablet computer, or other electronic device may receive a signal including an encoded image and perform decoding.

[0121] The features described in this article can be associated with video coding. Examples can be related to the temporal prediction of block partitioning parameters in video coding. By predicting partitioning parameters in time, it is possible to achieve coding gain and runtime reduction. Split parameters can be predicted to impose constraints on the partitioning of the current block or a set of blocks within a region. The examples described in this article can influence the temporal prediction of partitioning parameters.

[0122] The features described in this paper can be associated with block partitioning. A CTU can be split into CUs using a quadtree structure represented as a coding tree to accommodate local characteristics. A decision can be made at the leaf CU level whether to use inter-picture (e.g., temporal) or intra-picture (e.g., spatial) prediction to encode a picture region. Depending on the PU splitting type, a leaf CU can be split into one, two, or four PUs. Within a PU, the (e.g., the same) prediction process can be applied, and relevant information can be transferred to the decoder based on the PU. After obtaining residual blocks by applying prediction processes based on the PU splitting type, the leaf CUs can be partitioned into transform units (TUs) according to a quadtree structure (e.g., a coding tree similar to a CU). The partitioning concept can include CUs, PUs, and TUs.

[0123] In video processing, quadtrees with nested multi-type trees using binary and ternary splitting structures can replace the concept of multiple prediction unit types. For example, a quadtree with nested multi-type trees can remove the separation of CU, PU, ​​and TU concepts, except for CUs with a maximum transform size. Quadtrees with nested multi-type tree partitions can provide flexibility in the shape of CU partitions. In the coding tree structure, CUs can have square or rectangular shapes. Coding tree units (CTUs) can be partitioned by quadtree (e.g., quadtree) structures. Quadtree leaf nodes can be partitioned by multi-type tree structures.

[0124] Figure 5 This illustrates various tree splitting patterns. For example... Figure 5 As shown, there are four possible split types in a multi-type tree structure: vertical binary split (SPLIT_BT_VER), horizontal binary split (SPLIT_BT_HOR), vertical ternary split (SPLIT_TT_VER), and horizontal ternary split (SPLIT_TT_HOR). Multi-type tree leaf nodes can be called coding units (CUs), and unless the CU is too large for the maximum transform length, the split can be used for prediction and transform processing (e.g., without any further partitioning). In a quadtree with a nested multi-type tree coding block structure, CUs, PUs, and TUs can have the same block size. Exceptions may occur when the maximum supported transform length is less than the width or height of the color component of the CU.

[0125] Figure 6The signaling mechanism for partition splitting information in a quadtree with a nested multi-type tree coding tree structure is illustrated. A coding tree unit (CTU) can be considered the root of a quadtree and can be partitioned by the quadtree structure. Quadtree leaf nodes can be partitioned by the multi-type tree structure when the size of the resulting partition is equal to or greater than a given minimum partition size. In a quadtree with a nested multi-type tree coding tree structure, for a CU node, a first indicator (split_cu_indication) can be signaled to indicate whether the node is to be further partitioned. If the current CU node is a quadtree CU node, a second indicator (split_qt_indication) can be signaled to indicate whether it is a quadtree (QT) or multi-type tree (MTT) partitioning mode. When partitioning a node using the MTT partitioning mode, a third indicator (mtt_split_cu_vertical_indication) can be signaled to indicate the split direction, and a fourth indicator (mtt_split_cu_binary_indication) can be signaled to indicate whether the split is a binary or ternary split. Based on the values ​​of mtt_split_cu_vertical_indication and mtt_split_cu_binary_indication, the multi-type tree splitting mode (MttSplitMode) of the CU can be derived, as shown in Table 1.

[0126] Table 1 – MttSplitMode Derivation Based on Multi-Type Tree Syntax Elements MttSplitMode mtt_split_cu_vertical_indication mtt_split_cu_binary_indication SPLIT_TT_HOR 0 0 SPLIT_BT_HOR 0 1 SPLIT_TT_VER 1 0 SPLIT_BT_VER 1 1

[0127] Figure 7 The diagram illustrates a CTU divided into multiple CUs, featuring a quadtree and nested multi-type tree coding block structure. Bold block edges represent quadtree partitions, and the remaining edges represent multi-type tree partitions. The quadtree with nested multi-type tree partitions provides a Content Adaptive Coding (CABAC) tree structure composed of CUs. The size of a CU can be as large as the CTU or as small as 4×4 in units of luma samples. For the 4:2:0 chroma format, the maximum chroma CB size can be 64×64, and the minimum chroma CB size can include 16 chroma samples.

[0128] The maximum supported luminance transformation size is 64×64, and the maximum supported chrominance transformation size is 32×32. When the width or height of the CB is greater than the maximum transformation width or height, the CB can be implicitly split in the horizontal and / or vertical directions to meet the transformation size limit in that direction.

[0129] Figure 7An example of a quadtree with a nested multi-type tree coding block structure is shown. The following parameters can be defined for a quadtree with a nested multi-type tree coding tree scheme. One or more of the following parameters can be specified by the SPS syntax element and can be refined by the picture header syntax element: CTU size: the size of the root node of the quadtree; MinQTSize: the minimum allowed size of the leaf node of the quadtree; MaxBtSize: the maximum allowed size of the root node of the binary tree; MaxTtSize: the maximum allowed size of the root node of the ternary tree; MaxMttDepth: the maximum allowed level depth for splitting the multi-type tree from the leaves of the quadtree; and / or MinCbSize: the minimum allowed size of the coding block node.

[0130] In an example of a quadtree with a nested multi-type tree coding tree structure, the CTU size can be set to 128×128 luma samples, with two corresponding 64×64 blocks of 4:2:0 chroma samples. As an example, MinQTSize can be equal to 8×8, MaxBtSize can be set to 128×128, MaxTtSize can be set to 64×64, MinCbsize (e.g., for width and height) can be set to 4×4, and MaxMttDepth can be set to 4. Quadtree partitioning can be applied to the CTU to generate quadtree leaf nodes. The size of the quadtree leaf node can be between 8×8 (e.g., MinQTSize) and the CTU size. If the size of the leaf QT node is greater than MaxBtSize and MaxTtSize, the leaf QT node cannot be further split using binary or ternary splitting patterns. If the size of the leaf QT node is less than or equal to MaxBtSize and MaxTtSize, the leaf quadtree node can be further partitioned by multi-type trees. A quaternion leaf node can be the root node of a polytree, and it can have a polytree depth (mttDepth) of 0. Further splits are not considered when the polytree depth reaches MaxMttDepth (e.g., 4). Further horizontal splits are not considered when the width of a polytree node equals MinCbsize. Further vertical splits are not considered when the height of a polytree node equals MinCbsize.

[0131] The coding tree scheme can support the ability for luma and chroma to have separate block tree structures. For P and B slices, the luma and chroma CTBs in the CTU can share the same coding tree structure. For I slices, luma and chroma can have separate block tree structures. When applying the separate block tree mode, the luma CTB can be partitioned into CUs by the coding tree structure, and the chroma CTB can be partitioned into chroma CUs by another coding tree structure. CUs in I slices can include coding blocks for the luma components or coding blocks for both chroma components, and CUs in P or B slices can include coding blocks for all three color components (e.g., unless the video is monochrome).

[0132] The QT, BT, and TT partitions of a block or a set of blocks can be modified based on temporal partitioning parameters in the local neighborhood. This method can predict the permissible splits and maximum depth of a given block based on (e.g., previous) coded frames in a sequence. Two examples are provided below.

[0133] Whether a split partition is allowed can be determined. For example, for a block, the allowance of a split partition can be predicted based on the minimum QT split and the average QT split obtained from a time region (e.g., a set of juxtaposed CTUs or juxtaposed images of CTUs). A QT split is allowed when the current QT depth is less than the minimum QT depth minus 1. A split is not allowed when the current QT depth is less than the average QT depth minus 1, and both QT splits and TT splits are allowed. BT splits are allowed if TT is selected in the parent node.

[0134] The MTT depth parameter can be changed adaptively. For example, the maximum multi-tree depth of a block can be predicted. If the temporal maximum multi-tree depth is lower than the current block's maximum multi-tree depth, it can be decreased (e.g., the MTT depth parameter) (e.g., unless the current QP is higher than or equal to the reference frame's QP). If the temporal maximum multi-tree depth is higher than the current block's maximum multi-tree depth, and if the current depth equals the current QT depth, the maximum multi-tree depth can be increased.

[0135] Sub-block temporal motion vector prediction (SbTMVP) can use temporal correlations to predict the motion of sub-blocks from previous blocks. Context-adaptive binary arithmetic coding (CABAC) initialization can depend on (e.g., become dependent on) previous inter-frame slices.

[0136] Local splitting parameters (e.g., coding tree parameters such as maximum multi-type tree depth, QT depth, or the allowance for splitting partitions) can be derived from the local temporal neighborhood. In a slow-moving scene, partitions in two consecutive frames may appear similar. Parameters associated with a partition can be predicted based on previous frames. In the example, when there is fast motion in the scene, the splitting parameters between predicted frames can be disabled or deactivated.

[0137] Time-based prediction can be used to predict partition parameters. The predicted partition parameters can be exported.

[0138] Juxtaposed images can be used as reference images for temporal prediction of partition-related parameters. For example, juxtaposed images can be identified as follows.

[0139] At the beginning of decoding a P or B slice, after decoding the slice header, and once a list of reference images has been calculated for the current slice, the juxtaposed image can be identified as the reference image whose index in the list of reference images L0 or L1, identified by the image header syntax element ph_collocated_from_L0_indication, is equal to the parsed image header syntax element ph_collocated_ref_idx. If these two syntax elements do not exist in the image header, the juxtaposed image for the current slice can be identified using the slice header syntax elements sh_collocated_from_l0_indication and sh_collocated_ref_idx.

[0140] When Temporal Motion Vector Prediction (TMVP) is enabled, juxtaposed images can be defined. For example, TMVP can be enabled when ph_temporal_MVP_enabled_indication equals 1. As described in this article, images from juxtaposed images can be used.

[0141] When the current image is an Intra-Random Access Point (IRAP), prediction may not occur. When the current image is not an IRAP, prediction can be made based on one of the following: prediction based on the previous image in the encoding order; prediction based on the most recent preceding image in the encoding order with a time ID equal to or less than the current image; prediction based on the most recent preceding image in the encoding order, which has a time ID equal to or less than the current image and the same resolution as the current image; prediction based on the image with the highest POC value among the images preceding the current image in the encoding order; prediction based on the image with the highest POC value among the images preceding the current image in the encoding order and with a time ID equal to or less than the current image; prediction based on the image with the highest POC value among the images preceding the current image in the encoding order and with a time ID equal to or less than the current image and the same resolution as the current image; prediction based on a juxtaposed image (or the previous image in the encoding order if no such image is defined); and / or prediction based on an image referenced by an index explicitly encoded in the image header. The index can point to an image with a time ID equal to or less than the current image and a resolution equal to the current image's resolution.

[0142] Prediction can be disabled based on one or more of the following: the current image does not contain an I-slice, but the image used for prediction contains an I-slice (e.g., only an I-slice); the current image does not contain an I-slice, but the image used for prediction contains one or more I-slices; the slice type associated with the current CTU is different from the slice type associated with the juxtaposed CTU in the image used for prediction; and / or the current image has a lower time ID than the image used for prediction.

[0143] Based on any of the following examples, parameters can be applied to the temporal prediction of coding tree parameters (e.g., partitioning parameters). If the image used to predict the coding tree parameters of the current image is encoded in a dual-tree mode, and the current image is in a non-dual-tree mode, then the luma coding tree parameters of the referenced image can be used for prediction (e.g., instead of the chroma coding tree parameters). If adaptive resolution coding is used in the video bitstream under consideration, the temporal prediction of partitioning parameters may occur between images encoded at the same resolution level. If adaptive resolution coding is used in the video bitstream under consideration, the temporal prediction of partitioning parameters may occur between images encoded at full resolution.

[0144] The predicted partitioning parameters can be refined (e.g., corrected) by encoding the residual data. In the example, predicted parameters derived from temporal information (e.g., split partitioning information and / or maximum multi-tree depth) can be corrected based on the residuals in the video data. For a given block or set of blocks, partitioning parameters can be predicted based on temporal partitioning information. Once the predicted parameters are derived, partial or full partitioning searches can be performed at the encoder. The encoded syntax can be the residuals (e.g., the difference between the partitioning information computed by the encoder and the predicted partitioning information) of the predicted partitioning information.

[0145] The decoder can perform predictions of partition parameters, decode the residuals of these parameters, and use the decoded residuals to refine the predictions.

[0146] In the example, flags (e.g., at the image level) can be encoded to indicate whether to use residual syntax to correct the predicted partitioning information or to rely on unrefined prediction parameters.

[0147] Multiple candidate prediction lists can be used. In the example, multiple types of partition parameter predictions (e.g., multiple modes) can be exported at the encoder and / or decoder and stored in the candidate list. The index can be encoded by region (e.g., at the CTU level) to indicate which mode is suitable for that region.

[0148] Mode 0 can correspond to, for example, disabling splitting; Mode 1 can correspond to, for example, enabling splitting, with the maximum multi-tree depth set to 4; Mode 2 can correspond to, for example, enabling splitting, with the maximum multi-tree depth set to 3; Mode 3 can correspond to, for example, enabling splitting, with the maximum multi-tree depth set to 2.

[0149] Patterns can be sorted based on predicted parameters. Patterns most similar to those predicted parameters can be placed in the list first. The index encoding can be based on CABAC encoding, using the predicted parameters from the CABAC context model.

[0150] The features described herein can be associated with a CABAC context. In the example, the predicted partitioning parameters can be used to constrain the split decision. In the example, the predicted partitioning parameters may not be used to constrain the split decision. These parameters can be used to determine which CABAC contexts are used to encode the split decision, for example, which contexts are used to encode syntax elements and flags (e.g., split_qt_flag, split_cu_flag, mtt_split_cu_vertical_flag, and / or mtt_split_cu_binary_flag).

[0151] If both values ​​of the flag are allowed according to the predicted partitioning parameters, the first set of CABAC contexts can be used to encode the flag. If the first value of the flag is allowed according to the predicted partitioning parameters, the second set of CABAC contexts can be used to encode the flag. If the second value of the flag is allowed according to the predicted partitioning parameters, the third set of CABAC contexts can be used to encode the flag. If neither value of the flag is allowed according to the predicted partitioning parameters, the fourth set of CABAC contexts can be used.

[0152] The number of CABAC context sets can be reduced. For example, if both values ​​of a flag are either disallowed or both are allowed based on the predicted partitioning parameters, then a set of CABAC contexts can be used to encode the flag. If one of the possible binary values ​​of a flag is allowed based on the predicted partitioning parameters, then a second set of CABAC contexts can be used to encode the flag. The flag can be interpreted as follows: a first value (e.g., 0) can be interpreted as a split flag value allowed based on the predicted partitioning parameters, while a second value (e.g., Example 1) can be interpreted as a split flag value disallowed based on the predicted partitioning parameters.

[0153] Partitioning can be predicted based on the average block size in the time reference region. The coding tree depth (e.g., quadtree depth or multi-type tree depth) can be used as a partitioning parameter for prediction along the time axis. A block characteristic that might be associated with one coded image from another is the block size (e.g., not the coding tree itself). Block size can be associated with the spatial activity of the compressed images signaled by the image, which can be correlated between images. Reaching the considered block size during the partitioning process can be achieved through various series of splitting patterns from the coding tree root to the coding tree leaves, resulting in different depth values ​​associated with the leaves. Multi-type tree depth can be related to or unrelated to the size of the coded blocks. For example, a horizontal and vertical ternary splitting pattern can divide a block into three sub-blocks of different sizes. The area of ​​the middle sub-block (e.g., the number of samples) can be twice the area of ​​the two other sub-blocks.

[0154] Temporal propagation of block size parameters can be performed. The average block area of ​​coded blocks in a given region of a reference image can be calculated. Partitions can be determined (e.g., predicted) based on the average block size.

[0155] For a given coding tree node, further splitting of the current node is permitted if the current node's surface area is greater than half the area of ​​the propagation block. Splitting of the current tree node is not permitted if the current coding tree node's surface area is less than half the size of the propagation block.

[0156] In some examples, partitioning based on the average block size can be performed on the encoder side. In other examples, partitioning based on the average block size can be performed synchronously on both the encoder and decoder sides.

[0157] The partitioning process can be driven by parameters that reflect the temporal correlation between two images in terms of block partitioning, thereby producing compression efficiency.

[0158] Partitions can be predicted based on a temporal depth parameter representing the block size. In the example, partitioning parameters that propagate from one image to another can be included in the depth parameter. The depth parameter can be different from quadtree depth and multi-type tree depth, and can represent the average, minimum, or maximum block size.

[0159] The depth parameter can be determined as follows. At the root level of the encoding tree, set depth = 0. When a depth value is specified... parentDepth When a coding tree node is divided into multiple sub-coding tree nodes, the sub-coding tree can be assigned the following depth values, as shown in (1): (1) Using the depth parameters described here, a bijection can exist between the depth value and the block area.

[0160] Temporal propagation of the depth parameter can be applied, and the propagation of the depth parameter can be used to determine whether split blocks are allowed or not. The maximum and minimum values ​​of the depth parameter can be obtained in the spatial region of the reference image. For the current tree node in the current image, the depth parameter described here can be calculated.

[0161] Splitting rules based on time-propagation depth parameters can be applied. For example, if the current node depth is at least the maximum time depth plus 1, splitting the current tree node is not allowed. If the current node depth is less than the minimum time depth minus 1, splitting the current tree node is forced. If the current node depth is less than the maximum time depth plus 1, or if the current node depth is at least the minimum time depth minus 1, splitting the current tree node is allowed. Splitting parameters can be selected on the encoder side, and the splitting decision can be signaled in the bitstream.

[0162] In some examples, during rate-distortion optimized coding tree decisions, allowing / disallowing block splitting based on propagation depth parameters can be implemented on the encoder side. In some examples, allowing / disallowing block splitting based on propagation depth parameters can be implemented on both the encoder and decoder sides. Certain splitting operations can be allowed or disallowed based on one or more rules on the encoder and decoder sides. During partitioning, the depth parameters described herein can be similar to the cbSubDiv parameter.

[0163] The cbSubDiv parameter can be propagated in time and used to drive the partitioning process on the encoder and / or decoder side. The same time propagation and splitting rules as described herein can be applied (e.g., replacing the depth parameter described herein with the cbSubDiv variable).

[0164] While the examples provided herein assume that media content is streamed to a display device, there are no specific limitations on the type of display device that can benefit from the example technologies described herein. For example, a display device could be a television, projector, mobile phone, tablet, etc. Furthermore, the example technologies described herein can be applied not only to streaming use cases but also to teleconferencing setups. Additionally, the decoder and display described herein can be separate devices or parts of the same device. For example, a set-top box can decode the input video stream and (e.g., subsequently) provide the decoded stream to a display device (e.g., via HDMI), and information about viewing conditions (e.g., viewing distance) can be transmitted from the display device to the set-top box (e.g., via HDMI).

[0165] Although the features and elements have been described above in specific combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with other features and elements. Furthermore, the methods described herein can be implemented in a computer program, software, or firmware incorporated in a computer-readable medium for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted via wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROMs and digital multifunction discs (DVDs). A processor associated with the software can be used to implement a radio frequency transceiver used in a WTRU, UE, terminal, base station, RNC, or any host computer.

Claims

1. A video decoding device comprising: a processor configured to: obtain, for a video block, a partition parameter candidate prediction list comprising a plurality of partition modes; receive an index indicating a partition mode selected from the plurality of partition modes; and apply the selected partition mode to the video block.

2. The device of claim 1, wherein the processor is further configured to: identify a plurality of temporally neighboring blocks of the video block, wherein the partition parameter candidate prediction list is obtained based on the plurality of temporally neighboring blocks of the video block.

3. The device of any of claims 1 or 2, wherein a first mode of the plurality of partition modes is associated with a bypass partition, and wherein the processor is further configured to: determine that the selected partition mode comprises the first partition mode; and apply the first partition mode to the video block.

4. The device of any of claims 1 or 2, wherein a second partition mode of the plurality of partition modes is associated with an enabled partition, and wherein the enabled partition comprises configuring a multi-type tree depth value, and wherein the processor is further configured to: determine that the selected partition mode comprises the second partition mode; and apply the second partition mode to the video block based on the multi-type tree depth value.

5. The device of any of claims 1-4, wherein the processor is further configured to: order the plurality of partition modes based on prediction parameters of the video block; and determine the selected partition mode based on the ordered plurality of partition modes.

6. The device of any of claims 1-5, wherein the index indicating the partition mode selected from the plurality of partition modes is encoded at a coding tree unit level.

7. The device of any of claims 1-6, wherein the index indicating the partition mode selected from the plurality of partition modes is encoded based on context adaptive binary arithmetic coding (CABAC).

8. A video encoding device comprising: a processor configured to: determine, for a video block, a partition parameter candidate prediction list comprising a plurality of partition modes; select a partition mode from the plurality of partition modes; apply the selected partition mode to the video block; and include, in video data, an index indicating the partition mode selected from the plurality of partition modes.

9. The device of claim 8, wherein the processor is further configured to: identify a plurality of temporally neighboring blocks of the video block, wherein the partition parameter candidate prediction list is determined based on the plurality of temporally neighboring blocks of the video block.

10. The device of any of claims 8 or 9, wherein a first mode of the plurality of partition modes is associated with a bypass partition, and wherein the processor is further configured to: determine that the selected partition mode comprises the first partition mode; and apply the first partition mode to the video block.

11. The device of any of claims 8 or 9, wherein a second partition mode of the plurality of partition modes is associated with an enabled partition, and wherein the enabled partition comprises configuring a multi-type tree depth value, and wherein the processor is further configured to: determine that the selected partition mode comprises the second partition mode; and apply the second partition mode to the video block based on the multi-type tree depth value.

12. The device of any of claims 8 to 11, wherein the processor is further configured to: rank the plurality of partition modes based on prediction parameters of the video block; and determine the selected partition mode based on the ranked plurality of partition modes.

13. The device of any of claims 8 to 12, wherein an index indicating the selected partition mode from the plurality of partition modes is encoded at a coding tree unit level.

14. The device of any of claims 8 to 13, wherein the index indicating the selected partition mode from the plurality of partition modes is encoded based on context adaptive binary arithmetic coding (CABAC).

15. A method for a video decoder, the method comprising: obtaining, for a video block, a partition parameter candidate prediction list including a plurality of partition modes; receiving an index indicating a selected partition mode from the plurality of partition modes; and applying the selected partition mode to the video block.

16. The method of claim 15, wherein the method further comprises: identifying a plurality of temporally neighboring blocks of the video block, wherein the partition parameter candidate prediction list is obtained based on the plurality of temporally neighboring blocks of the video block.

17. The method of any of claims 15 or 16, wherein a first mode of the plurality of partition modes is associated with a bypass partition, and wherein the method further comprises: determining that the selected partition mode includes the first partition mode; and applying the first partition mode to the video block.

18. The method of any of claims 15 or 16, wherein a second partition mode of the plurality of partition modes is associated with an enabled partition, and wherein the enabled partition includes configuring a multi-type tree depth value, and wherein the method further comprises: determining that the selected partition mode includes the second partition mode; and applying the second partition mode to the video block based on the multi-type tree depth value.

19. The method of any of claims 15 to 18, wherein the method further comprises: ranking the plurality of partition modes based on prediction parameters of the video block; and determining the selected partition mode based on the ranked plurality of partition modes.

20. The method of any of claims 15 to 19, wherein an index indicating the selected partition mode from the plurality of partition modes is encoded at a coding tree unit level.

21. The method of any of claims 15 to 20, wherein the index indicating the selected partition mode from the plurality of partition modes is encoded based on context adaptive binary arithmetic coding (CABAC).

22. A method for a video encoder, the method comprising: determining, for a video block, a partition parameter candidate prediction list including a plurality of partition modes; selecting a partition mode from the plurality of partition modes; applying the selected partition mode to the video block; and including, in video data, an index indicating the selected partition mode from the plurality of partition modes.

23. The method of claim 22, wherein the method further comprises: identifying a plurality of temporally neighboring blocks of the video block, wherein the partition parameter candidate prediction list is determined based on the plurality of temporally neighboring blocks of the video block. ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ 24. The method of any of claims 22 or 23, wherein a first mode of the plurality of partition modes is associated with a bypass partition, and wherein the method further comprises: determining that the selected partition mode comprises the first partition mode; and applying the first partition mode to the video block.

25. The method of any of claims 22 or 23, wherein a second mode of the plurality of partition modes is associated with an enabled partition, and wherein the enabled partition comprises configuring a multi-type tree depth value, and wherein the method further comprises: determining that the selected partition mode comprises the second partition mode; and applying the second partition mode to the video block based on the multi-type tree depth value.

26. The method of any of claims 22 to 25, wherein the method further comprises: ordering the plurality of partition modes based on a prediction parameter of the video block; and determining the selected partition mode based on the ordered plurality of partition modes.

27. The method of any of claims 22 to 26, wherein an index indicating the selected partition mode from the plurality of partition modes is encoded at a coding tree unit level.

28. The device of any of claims 22 to 27, wherein the index indicating the selected partition mode from the plurality of partition modes is encoded based on context adaptive binary arithmetic coding (CABAC).

29. A computer program product stored on a non-transitory computer readable medium and comprising program code instructions for implementing the steps of the method of any of claims 15 to 28 when executed by a processor.

30. A computer program comprising program code instructions for implementing the steps of the method of any of claims 15 to 28 when executed by a processor.

31. Video data comprising information representative of an encoded block encoded according to one of the methods of any of claims 22 to 28.