Low frequency non-separable transform and use of non-dct2 primary transform

CN122514950APending Publication Date: 2026-08-04INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INTERDIGITAL CE PATENT HOLDINGS SAS
Filing Date
2025-01-03
Publication Date
2026-08-04

Smart Images

  • Figure CN122514950A_ABST
    Figure CN122514950A_ABST
Patent Text Reader

Abstract

Systems, methods, and means for video encoding / decoding are disclosed for the use of Low Frequency Inseparable Transform (LFNST) and Non-Discrete Cosine Transform 2 (DCT2) master transforms. In the examples, a video decoding device can determine whether to disable the Inseparable Master Transform (NSPT) for the current block. If NSPT is disabled for the current block, the video decoding device can determine to use LFNST for the current block. If LFNST is used, the video decoding device can decode the current block based on Multiple Transform Selection (MTS). In the examples, a video encoding device can disable NSPT for the current block. A video encoding device can enable LFNST for the current block. A video encoding device can encode the current block based on MTS.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications This application claims the benefit of European Provisional Patent Application No. 24305013.5, filed on 5 January 2024, the contents of which are incorporated herein by reference. Background Technology

[0002] Video codec systems can be used to compress digital video signals, for example, to reduce the storage and / or transmission bandwidth required for such signals. Video codec systems can include, for example, block-based, wavelet-based, and / or object-based systems. Summary of the Invention

[0003] Systems, methods, and means for video encoding and decoding are disclosed, such as video coding and / or video decoding using the Low Frequency Inseparable Transform (LFNST) and Non-Discrete Cosine Transform 2 (DCT2) master transform.

[0004] In the example, the video decoding device (e.g., the decoder) may include a processor configured to perform one or more of the following.

[0005] Video decoding devices can determine whether the Non-Separable Main Transform (NSPT) is disabled for a codec unit. For example, the device can obtain an NSPT enable indication from video data (e.g., a bitstream). The NSPT enable indication can be configured to indicate whether NSPT is enabled for a codec unit. Determining whether NSPT is enabled for a codec unit can be based on the NSPT enable indication. The NSPT enable indication can be an NSPT enable indication flag.

[0006] Based on the determination that NSPT is disabled for the codec unit, the video decoding device can determine that LFNST is enabled for the codec unit. For example, the device can obtain an LFNST enable indication from the video data. The LFNST enable indication can be configured to indicate whether LFNST is enabled for the codec unit. Determining that LFNST is enabled for the codec unit can be based on the LFNST enable indication. The LFNST enable indication can be an LFNST enable indication flag.

[0007] Based on the determination that LFNST is enabled for the codec unit, the video decoding device can decode the codec unit based on Multiple Transform Selection (MTS).

[0008] The video decoding device can determine that the LFNST index is not 0. Based on determining that the LFNST index is not 0 and that NSPT is disabled for the codec unit, the device can obtain the MTS index associated with the MTS in the video data. The MTS index can be or can include and / or indicate an MTS index indicator. The MTS index indicator can be configured to indicate the use of a transform pair. The video decoding device can perform an inverse transform based on the transform pair. The transform pair can be or can include at least one of Discrete Cosine Transform Type 2 (DCT2) and DCT2 transform pair or DCT5 and DCT5 transform pair.

[0009] In the example, the video encoding device (e.g., an encoder) may include a processor configured to perform one or more of the following.

[0010] Video encoding devices can determine whether to disable the Non-Separable Main Transform (NSPT) for codec units. For example, the device can include an NSPT enable indicator in the video data (e.g., a bitstream). The NSPT enable indicator can be configured to indicate whether NSPT is enabled for codec units. The NSPT enable indicator can be an NSPT enable indicator flag.

[0011] Based on the determination that NSPT is disabled for the codec unit, the video encoding device can determine whether LFNST is enabled for the codec unit. For example, the device can include an LFNST enable indication in the video data. The LFNST enable indication can be configured to indicate whether LFNST is enabled for the codec unit. The LFNST enable indication can be an LFNST enable indication flag.

[0012] Based on the determination that LFNST is enabled for the codec unit, the encoding device can encode the codec unit based on MTS.

[0013] The video encoding device can determine that the LFNST index is not 0. Based on determining that the LFNST index is not 0 and that NSPT is disabled for the codec unit, the device can include an MTS index associated with the MTS in the video data. The MTS index can be, or can include, and / or indicate an MTS index indicator. The MTS index indicator can be configured to indicate the use of a transform pair. The transform pair can be, or can include, at least one of a DCT2 and DCT2 transform pair or a DCT5 and DCT5 transform pair.

[0014] The video decoding device can determine whether the Inseparable Main Transform (NSPT) is disabled for the current block. The current block can be associated with an intra-frame codec block. If NSPT is disabled for the current block, the video decoding device can determine whether to use LFNST for the current block. The video decoding device can determine whether to use LFNST for the current block based on obtaining at least one of the following: Codec Unit (CU) LFNST indication, Transform Unit (TU) LFNST indication, Sequence Parameter Set (SPS) LFNST indication, Slice Header LFNST indication, and Picture Header LFNST indication. If LFNST is used, the video decoding device can decode the current block based on Multiple Transform Selection (MTS).

[0015] In the example, an MTS can be associated with an MTS index indicator. The MTS index indicator can indicate the use of a transform pair associated with the MTS. The transform pair can be or may include at least one of DCT2 and DCT2 transform pairs or DCT5 and DCT5 transform pairs. In the example, an MTS candidate set can be associated with transform pairs of the MTS. The MTS candidate set can be or may include at least one of DCT2 and DCT2 transform pairs, DCT5 and DCT5 transform pairs, DCT2 and DCT5 transform pairs, or DCT5 and DCT2 transform pairs.

[0016] In the example, the video decoding device can determine that the LFNST index is not 0. If the LFNST index is not 0, the video decoding device can obtain the MTS index associated with the MTS in the video data.

[0017] In the example, the video decoding device can determine an MTS candidate set associated with the MTS based on the sum of the transform coefficient amplitudes. If the sum of the transform coefficient amplitudes is less than or equal to a first threshold, the video decoding device can determine that the MTS candidate set is a DCT2 and DCT2 transform pair. If the sum of the transform coefficient amplitudes is greater than the first threshold and less than or equal to a second threshold, the video decoding device can determine that the MTS candidate set is at least one of a DCT2 and DCT2 transform pair or a DCT5 and DCT5 transform pair. If the sum of the transform coefficient amplitudes is greater than the second threshold, the video decoding device can determine that the MTS candidate set is at least one of a DCT2 and DCT2 transform pair, a DCT5 and DCT5 transform pair, a DCT2 and DCT5 transform pair, or a DCT5 and DCT2 transform pair.

[0018] In the example, the video decoding device can obtain the MTS based on at least one of intra-frame mode parity check or decoder-side intra-frame mode derivation (DIMD).

[0019] The video encoding device can disable NSPT for the current block. The current block can be associated with an intra-frame codec block. The video encoding device can enable LFNST for the current block. The video encoding device can include CU LFNST indication, TU LFNST indication, SPS LFNST indication, slice header LFNST indication, and picture header LFNST indication in the video data. The video encoding device can encode the current block based on MTS.

[0020] In the example, an MTS can be associated with an MTS index indicator. The MTS index indicator can indicate the use of a transform pair associated with the MTS. The transform pair can be or may include at least one of DCT2 and DCT2 transform pairs or DCT5 and DCT5 transform pairs. In the example, an MTS candidate set can be associated with transform pairs of the MTS. The MTS candidate set can be or may include at least one of DCT2 and DCT2 transform pairs, DCT5 and DCT5 transform pairs, DCT2 and DCT5 transform pairs, or DCT5 and DCT2 transform pairs.

[0021] In the example, the video encoding device can determine that the LFNST index is not 0. If the LFNST index is not 0, the video encoding device can include the MTS index associated with the MTS in the video data.

[0022] In the example, the video coding device can determine the MTS candidate set associated with the MTS based on the sum of the transform coefficient amplitudes. If the sum of the transform coefficient amplitudes is less than or equal to a first threshold, the video coding device can determine the MTS candidate set as a DCT2 and DCT2 transform pair. If the sum of the transform coefficient amplitudes is greater than the first threshold and less than or equal to a second threshold, the video coding device can determine the MTS candidate set as at least one of a DCT2 and DCT2 transform pair or a DCT5 and DCT5 transform pair. If the sum of the transform coefficient amplitudes is greater than the second threshold, the video coding device can determine the MTS candidate set as at least one of a DCT2 and DCT2 transform pair, a DCT5 and DCT5 transform pair, a DCT2 and DCT5 transform pair, or a DCT5 and DCT2 transform pair.

[0023] In the example, the video encoding device can obtain the MTS based on at least one of intra-frame mode parity check or DIMD. Attached Figure Description

[0024] Figure 1A This is a system diagram illustrating an example communication system in which one or more of the disclosed embodiments may be implemented.

[0025] Figure 1B To illustrate, according to an embodiment, it is possible to Figure 1AThe diagram shows a system diagram of an example wireless transmit / receive unit (WTRU) used in a communication system.

[0026] Figure 1C To illustrate, according to an embodiment, it is possible to Figure 1A The diagram illustrates a system diagram of an example radio access network (RAN) and an example core network (CN) used within a communication system.

[0027] Figure 1D To illustrate, according to an embodiment, it is possible to Figure 1A The diagram shows another example RAN and another example CN used in the communication system.

[0028] Figure 2 The illustration shows a sample video encoder.

[0029] Figure 3 The illustration shows an example video decoder.

[0030] Figure 4 The illustration shows an example of a system that can implement various aspects and examples.

[0031] Figure 5 An example of the region of interest (ROI) for the low-frequency non-separable transform (LFNST)16 is illustrated.

[0032] Figure 6 The diagram illustrates an example of the ROI for LFNST8.

[0033] Figure 7 The illustration shows an example of the Non-Separable Master Transform (NSPT) applied to a block (e.g., a small block), and an example of the LFNST applied to the rest.

[0034] Figure 8A The illustration shows an example of the lowest frequency basis function of the Discrete Cosine Transform V (DCT-V) and DCT-II. Figure 8B Examples of the lowest frequency basis functions for DCT-IV and DCT-VIII are illustrated. Figure 8C An example of the lowest frequency basis function of the Discrete Sine Transform IV (DST-IV) and DST-VII is illustrated. Figure 8D Examples of the lowest frequency basis functions for DST-I and DST-II are illustrated.

[0035] Figure 9 The illustration shows an example image group organization in low-latency B codec. Detailed Implementation

[0036] A more detailed understanding can be obtained from the following description, which is given in illustrative form in conjunction with the accompanying drawings.

[0037] Figure 1A The diagram illustrates an example communication system 100 in which one or more of the disclosed embodiments may be implemented. The communication system 100 may be a multiple access system that provides content such as voice, data, video, messaging, broadcasting, etc., to multiple wireless users. The communication system 100 enables multiple wireless users to access such content by sharing system resources, including wireless bandwidth. For example, the communication system 100 may employ one or more channel access methods, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal FDMA (OFDMA), Single Carrier FDMA (SC-FDMA), Zero-Tail Unique Word DFT Spread Spectrum OFDM (ZT UW DTS-s OFDM), Unique Word OFDM (UW-OFDM), Resource Block Filtered OFDM, Filter Bank Multicarrier (FBMC), etc.

[0038] like Figure 1A As shown, the communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RAN 104 / 113, CN 106 / 115, public switched telephone network (PSTN) 108, Internet 110, and other networks 112. However, it will be appreciated that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a wireless environment. For example, any of the WTRUs 102a, 102b, 102c, and 102d, which may be referred to as a “station” and / or “STA”, may be configured to transmit and / or receive wireless signals and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearable devices, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated processing chain scenarios), consumer electronics devices, devices operating on commercial and / or industrial wireless networks, etc. Any of the WTRUs 102a, 102b, 102c, and 102d may be interchangeably referred to as a UE.

[0039] The communication system 100 may also include base station 114a and / or base station 114b. Each of base stations 114a and 114b may be any type of device configured to wirelessly interface with at least one of WTRUs 102a, 102b, 102c, and 102d to facilitate access to one or more communication networks such as CN 106 / 115, the Internet 110, and / or other networks 112. For example, base stations 114a and 114b may be a basic transceiver station (BTS), Node-B, eNode B, home NodeB, home eNode B, gNB, NR NodeB, site controller, access point (AP), wireless router, etc. Although each of base stations 114a and 114b is depicted as a single element, it will be appreciated that base stations 114a and 114b may include any number of interconnected base stations and / or network elements.

[0040] Base station 114a may be part of RAN 104 / 113, which may also include other base stations and / or network elements (not shown), such as base station controllers (BSCs), radio network controllers (RNCs), relay nodes, etc. Base station 114a and / or base station 114b may be configured to transmit and / or receive radio signals on one or more carrier frequencies, which may be referred to as cells (not shown). These frequencies may be in licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide coverage of a specific geographic area, which may be relatively fixed or may change over time. A cell may be further divided into cell sectors. For example, the cell associated with base station 114a may be divided into three sectors. Therefore, in one embodiment, base station 114a may include three transceivers, i.e., one transceiver for each sector of the cell. In embodiments, base station 114a may employ multiple-input multiple-output (MIMO) technology and may utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.

[0041] Base stations 114a and 114b can communicate with one or more of WTRUs 102a, 102b, 102c, and 102d via air interface 116. The air interface can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). Air interface 116 can be established using any suitable radio access technology (RAT).

[0042] More specifically, as noted above, communication system 100 can be a multiple access system and can employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, base station 114a in RAN 104 / 113 and WTRUs 102a, 102b, 102c can implement radio technologies such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which can use Wideband CDMA (WCDMA) to establish air interfaces 115 / 116 / 117. WCDMA can include communication protocols such as High-Speed ​​Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA can include High-Speed ​​Downlink (DL) Packet Access (HSDPA) and / or High-Speed ​​UL Packet Access (HSUPA).

[0043] In the embodiment, base station 114a and WTRUs 102a, 102b, 102c can implement radio technologies such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which can establish air interface 116 using Long Term Evolution (LTE) and / or LTE Advanced (LTE-A) and / or LTE Advanced Pro (LTE-A Pro).

[0044] In the embodiment, base station 114a and WTRUs 102a, 102b, 102c can implement radio technologies such as NR radio access, which can establish air interface 116 using new radio (NR).

[0045] In the embodiments, base station 114a and WTRUs 102a, 102b, and 102c can implement multiple radio access technologies. For example, base station 114a and WTRUs 102a, 102b, and 102c can, for example, use the dual connectivity (DC) principle to implement both LTE and NR radio access together. Therefore, the air interface utilized by WTRUs 102a, 102b, and 102c can be characterized by multiple types of radio access technologies and / or transmissions sent to / from multiple types of base stations (e.g., eNBs and gNBs).

[0046] In other embodiments, base station 114a and WTRUs 102a, 102b, 102c can implement radio technologies such as IEEE 802.11 (i.e., Wi-Fi), IEEE 802.16 (i.e., Global Microwave Access Interoperability (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Provisional Standard 2000 (IS-2000), Provisional Standard 95 (IS-95), Provisional Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced GSM Evolution Data Rate (EDGE), GSM EDGE (GERAN), etc.

[0047] Figure 1A Base station 114b can be, for example, a wireless router, a home Node B, a home eNode B, or an access point, and can utilize any suitable RAT to facilitate wireless connectivity in a local area, such as a business premises, residence, vehicle, campus, industrial facility, air corridor (e.g., for drone use), road, etc. In one embodiment, base station 114b and WTRUs 102c, 102d can implement radio technologies such as IEEE 802.11 to establish a wireless local area network (WLAN). In another embodiment, base station 114b and WTRUs 102c, 102d can implement radio technologies such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, base station 114b and WTRUs 102c, 102d can utilize cellular-based RATs (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a picocell or femtocell. Figure 1A As shown, base station 114b can have a direct connection to Internet 110. Therefore, it is not required that base station 114b access Internet 110 via CN 106 / 115.

[0048] RAN 104 / 113 can communicate with CN 106 / 115, which can be any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRU 102a, 102b, 102c, and 102d. Data can have varying Quality of Service (QoS) requirements, such as different throughput requirements, latency requirements, fault tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, etc. CN 106 / 115 can provide call control, billing services, location-based services, prepaid calling, internet connectivity, video distribution, etc., and / or implement advanced security functions such as user authentication. Although not explicitly stated... Figure 1AAs shown, but will be understood, RAN 104 / 113 and / or CN 106 / 115 can communicate directly or indirectly with other RANs that use the same RAT as RAN 104 / 113 or a different RAT. For example, in addition to connecting to RAN 104 / 113, which may be using NR radio technology, CN 106 / 115 can also communicate with another RAN (not shown) using GSM, UMTS, CDMA 2000, WiMAX, E-UTRA, or WiFi radio technology.

[0049] CN 106 / 115 can also serve as a gateway for WTRU 102a, 102b, 102c, 102d to access PSTN 108, the Internet 110, and / or other networks 112. PSTN 108 may include a circuit-switched telephone network providing Common Old-Style Telephone Service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices using common communication protocols such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) from the TCP / IP Internet Protocol suite. Network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, network 112 may include another CN connected to one or more RANs that may employ the same RAT as or a different RAT than RAN 104 / 113.

[0050] Some or all of the WTRUs 102a, 102b, 102c, and 102d in communication system 100 may include multi-mode capabilities (e.g., WTRUs 102a, 102b, 102c, and 102d may include multiple transceivers for communicating with different wireless networks via different wireless links). For example, Figure 1A The WTRU 102c shown can be configured to communicate with base station 114a, which may employ cellular-based radio technology, and with base station 114b, which may employ IEEE 802 radio technology.

[0051] Figure 1B The following diagram illustrates the system of example WTRU 102. Figure 1B As shown, among other things, WTRU 102 may include a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power supply 134, a Global Positioning System (GPS) chipset 136, and / or other peripheral devices 138. It will be appreciated that WTRU 102 may include any sub-combination of the foregoing elements while remaining consistent with the embodiments.

[0052] Processor 118 may be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. Processor 118 may perform signal encoding / decoding, data processing, power control, input / output processing, and / or any other functions that enable WTRU 102 to operate in a wireless environment. Processor 118 may be coupled to transceiver 120, and transceiver 120 may be coupled to transmitting / receiving element 122. Although... Figure 1B The processor 118 and transceiver 120 are depicted as separate components, but it will be understood that the processor 118 and transceiver 120 can be integrated together into an electronic package or chip.

[0053] Transmitting / receiving element 122 can be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) via air interface 116. For example, in one embodiment, transmitting / receiving element 122 can be an antenna configured to transmit and / or receive RF signals. In another embodiment, transmitting / receiving element 122 can be a transmitter / receiver configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, transmitting / receiving element 122 can be configured to transmit and / or receive both RF signals and optical signals. It will be appreciated that transmitting / receiving element 122 can be configured to transmit and / or receive any combination of wireless signals.

[0054] Although the transmitting / receiving element 122 is in Figure 1B While depicted as a single element, WTRU 102 may include any number of transmit / receive elements 122. More specifically, WTRU 102 may employ MIMO technology. Thus, in one embodiment, WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals via air interface 116.

[0055] Transceiver 120 can be configured to modulate signals to be transmitted by transmitting / receiving element 122 and demodulate signals to be received by transmitting / receiving element 122. As noted above, WTRU 102 can have multi-mode capability. Therefore, transceiver 120 can include multiple transceivers for enabling WTRU 102 to communicate via various RATs such as NR and IEEE 802.11, for example.

[0056] The processor 118 of WTRU 102 can be coupled to a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) unit or an organic light-emitting diode (OLED) display unit), and can receive user input data from these devices. The processor 118 can also output user data to the speaker / microphone 124, keypad 126, and / or display / touchpad 128. Furthermore, the processor 118 can access information from any type of suitable memory and store the data in memory such as non-removable memory 130 and / or removable memory 132. Non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. Removable memory 132 may include a subscriber identity module (SIM) card, memory stick, secure digital storage (SD) card, etc. In other embodiments, the processor 118 can access information from memory that is not physically located on WTRU 102 (such as on a server or home computer (not shown)) and store the data in that memory.

[0057] The processor 118 can receive power from the power supply 134 and can be configured to distribute and / or control power to other components in the WTRU 102. The power supply 134 can be any suitable device for powering the WTRU 102. For example, the power supply 134 may include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cell units, fuel cell units, etc.

[0058] The processor 118 may also be coupled to the GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to, or instead of, information from the GPS chipset 136, the WTRU 102 may receive location information from base stations (e.g., base stations 114a, 114b) via the air interface 116, and / or determine its location based on the timing of signals received from two or more nearby base stations. It will be understood that the WTRU 102 may acquire location information using any suitable location determination method, while remaining consistent with the embodiments.

[0059] The processor 118 may be further coupled to other peripheral devices 138, which may include one or more software and / or hardware modules providing additional features, functions, and / or wired or wireless connectivity. For example, peripheral devices 138 may include accelerometers, electronic compasses, satellite transceivers, digital cameras (for photos and / or video), Universal Serial Bus (USB) ports, vibration devices, television transceivers, hands-free headsets, Bluetooth® modules, FM radio units, digital music players, media players, video game player modules, internet browsers, virtual reality and / or augmented reality (VR / AR) devices, activity trackers, etc. Peripheral devices 138 may include one or more sensors, which may be one or more of the following: gyroscopes, accelerometers, Hall effect sensors, magnetometers, orientation sensors, proximity sensors, temperature sensors, time sensors; geolocation sensors; altimeters, light sensors, touch sensors, magnetometers, barometers, gesture sensors, biometric sensors, and / or humidity sensors.

[0060] WTRU 102 may include a full-duplex radio, for which the transmission and reception of some or all signals (e.g., associated with a specific subframe of both UL (e.g., for transmission) and downlink (e.g., for reception)) may be concurrent and / or simultaneous. The full-duplex radio may include an interference management unit to reduce and / or substantially eliminate self-interference via hardware (e.g., a choke) or via signal processing (e.g., a separate processor (not shown) or via processor 118). In embodiments, WTRU 102 may include a half-duplex radio, for which the transmission and reception of some or all signals (e.g., associated with a specific subframe of either UL (e.g., for transmission) or downlink (e.g., for reception) may be concurrent and / or simultaneous.

[0061] Figure 1C The diagram illustrates a system diagram of RAN 104 and CN 106 according to an embodiment. As noted above, RAN 104 may employ E-UTRA radio technology to communicate with WTRUs 102a, 102b, and 102c via air interface 116. RAN 104 may also communicate with CN 106.

[0062] RAN 104 may include eNode-B 160a, 160b, 160c, but will be appreciated that RAN 104 may include any number of eNode-Bs while remaining consistent with the embodiments. Each of eNode-B 160a, 160b, 160c may include one or more transceivers for communicating with WTRU 102a, 102b, 102c via air interface 116. In one embodiment, eNode-B 160a, 160b, 160c may implement MIMO technology. Thus, for example, eNode-B 160a may use multiple antennas to transmit radio signals to and / or receive radio signals from WTRU 102a.

[0063] Each of the eNode-B 160a, 160b, and 160c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, etc. Figure 1C As shown, eNode-B 160a, 160b, and 160c can communicate with each other via the X2 interface.

[0064] Figure 1C The CN 106 shown may include a Mobility Management Entity (MME) 162, a Serving Gateway (SGW) 164, and a Packet Data Network (PDN) Gateway (or PGW) 166. While each of the foregoing elements is depicted as part of the CN 106, it will be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.

[0065] The MME 162 can connect to each of the eNode-Bs 160a, 160b, and 160c in RAN 104 via the S1 interface and can be used as a control node. For example, the MME 162 can be responsible for authenticating users of WTRUs 102a, 102b, and 102c, bearer activation / deactivation, selecting a specific serving gateway during the initial attachment of WTRUs 102a and 102c, etc. The MME 162 can provide control plane functions for handover between RAN 104 and other RANs (not shown) employing other radio technologies such as GSM and / or WCDMA.

[0066] The SGW 164 can connect to each eNode B 160a, 160b, or 160c in RAN 104 via the S1 interface. The SGW 164 can typically route user data packets to or forward user data packets from WTRUs 102a, 102b, or 102c. The SGW 164 can also perform other functions, such as anchoring the user plane during inter-eNode B handover, triggering paging when DL data is available to WTRUs 102a, 102b, or 102c, and managing and storing the context of WTRUs 102a, 102b, or 102c.

[0067] The SGW 164 can connect to the PGW 166, which can provide WTRU 102a, 102b, and 102c with access to packet-switched networks such as Internet 110, facilitating communication between WTRU 102a, 102b, 102c and IP-enabled devices.

[0068] CN 106 can facilitate communication with other networks. For example, CN 106 can provide WTRU 102a, 102b, and 102c with access to circuit-switched networks such as PSTN 108 to facilitate communication between WTRU 102a, 102b, and 102c and traditional landline communication equipment. For example, CN 106 may include, or be able to communicate with, an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) serving as an interface between CN 106 and PSTN 108. Furthermore, CN 106 can provide WTRU 102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.

[0069] Despite WTRU in Figures 1A-1D While described as a wireless terminal, it is conceivable that in some representative embodiments, such a terminal may (e.g., temporarily or permanently) use a wired communication interface with a communication network.

[0070] In a representative embodiment, the other network 112 may be a WLAN.

[0071] A WLAN in Infrastructure Basic Services Set (BSS) mode may have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP may have access or an interface to a distribution system (DS) or another type of wired / wireless network that introduces and / or leads traffic into and / or out of the BSS. Traffic originating outside the BSS and destined for a STA can be delivered to the AP. Traffic originating from a STA and destined for a destination outside the BSS can be sent to the AP for delivery to its respective destination. Traffic between STAs within the BSS can be sent, for example, through the AP, where a source STA can send traffic to the AP, and the AP can deliver the traffic to the destination STA. Traffic between STAs within the BSS can be considered and / or referred to as peer-to-peer traffic. Peer-to-peer traffic can be sent between the source STA and the destination STA (e.g., directly between them) using a direct link setup (DLS). In some representative embodiments, the DLS may use 802.11e DLS or 802.11z tunneled DLS (TDLS). A WLAN using the Standalone BSS (IBSS) mode may not have an access point (AP), and STAs within the IBSS or using the IBSS (e.g., all STAs) can communicate directly with each other. The IBSS communication mode may sometimes be referred to as the "ad-hoc" communication mode in this document.

[0072] When operating in 802.11ac infrastructure mode or a similar mode, the AP can transmit beacons on a fixed channel, such as the primary channel. The primary channel can be of fixed width (e.g., a 20 MHz bandwidth) or dynamically set via signaling. The primary channel can be the operating channel of the BSS and can be used by the STA to establish a connection with the AP. In some representative embodiments, such as in an 802.11 system, Carrier Sense Multiple Access (CSMA / CA) with collision avoidance can be implemented. For CSMA / CA, each STA (e.g., every STA), including the AP, can sense the primary channel. If a particular STA senses / detects and / or determines that the primary channel is busy, that particular STA can exit. A single STA (e.g., only one station) can transmit at any given time within a given BSS.

[0073] High-throughput (HT) STAs can communicate using a 40 MHz wide channel, for example, by combining a primary 20 MHz channel with adjacent or non-adjacent 20 MHz channels to form a 40 MHz wide channel.

[0074] Very High Throughput (VHT) STAs can support channels with widths of 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz. A 40 MHz and / or 80 MHz channel can be formed by combining consecutive 20 MHz channels. A 160 MHz channel can be formed by combining eight consecutive 20 MHz channels, or by combining two non-consecutive 80 MHz channels (this can be referred to as an 80+80 configuration). For the 80+80 configuration, after channel coding, data is transmitted through a segmented parser that divides the data into two streams. Inverse Fast Fourier Transform (IFFT) processing and time-domain processing can be performed on each stream separately. The streams can be mapped onto the two 80 MHz channels, and the data can be transmitted by the transmitting STA. At the receiver of the receiving STA, the operation of the 80+80 configuration can be reversed, and the combined data can be sent to the Media Access Control (MAC).

[0075] 802.11af and 802.11ah support operating modes below 1 GHz. The channel operating bandwidth and carrier in 802.11af and 802.11ah are reduced compared to those used in 802.11n and 802.11ac. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV whitespace (TVWS) spectrum, while 802.11ah uses non-TVWS spectrum to support 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths. According to a representative embodiment, 802.11ah can support instrument-type control / machine-type communication, such as MTC devices in macro coverage areas. MTC devices may have certain capabilities, such as limited capabilities, including support (e.g., only support) certain and / or limited bandwidths. MTC devices may include batteries with a battery life exceeding a threshold (e.g., to maintain a very long battery life).

[0076] WLAN systems that can support multiple channels and channel bandwidths, such as 802.11n, 802.11ac, 802.11af, and 802.11ah, include channels that can be designated as the primary channel. The primary channel can have a bandwidth equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel can be set and / or limited by the STAs supporting the minimum bandwidth operating mode from all STAs operating in the BSS. In the example of 802.11ah, for STAs supporting (e.g., only supporting) the 1 MHz mode (e.g., MTC type devices), the primary channel can still be 1 MHz wide even if the AP and other STAs in the BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operating modes. Carrier sensing and / or Network Allocation Vector (NAV) settings may depend on the status of the primary channel. If the primary channel is busy, for example, due to STAs (which only support the 1 MHz operating mode) transmitting to the AP, the entire available band may be considered busy even if most of the band remains idle and may be available.

[0077] In the United States, the available frequency band for 802.11ah is 902 MHz to 928 MHz. In South Korea, the available frequency band is from 917.5 MHz to 923.5 MHz. In Japan, the available frequency band is from 916.5 MHz to 927.5 MHz. Depending on the country code, the total bandwidth available for 802.11ah ranges from 6 MHz to 26 MHz.

[0078] Figure 1D The diagram illustrates a system diagram of RAN 113 and CN 115 according to an embodiment. As noted above, RAN 113 can use NR radio technology to communicate with WTRUs 102a, 102b, and 102c via air interface 116. RAN 113 can also communicate with CN 115.

[0079] RAN 113 may include gNBs 180a, 180b, and 180c; however, it will be understood that RAN 113 may include any number of gNBs while remaining consistent with the embodiments. Each of gNBs 180a, 180b, and 180c may include one or more transceivers for communicating with WTRUs 102a, 102b, and 102c via air interface 116. In one embodiment, gNBs 180a, 180b, and 180c may implement MIMO technology. For example, gNBs 180a and 180b may utilize beamforming to transmit signals to and / or receive signals from gNBs 180a, 180b, and 180c. Thus, for example, gNB 180a may use multiple antennas to transmit radio signals to and / or receive radio signals from WTRU 102a. In an embodiment, gNBs 180a, 180b, and 180c may implement carrier aggregation technology. For example, gNB 180a can transmit multiple component carriers to WTRU 102a (not shown). A subset of these component carriers may be located on unlicensed spectrum, while the remaining component carriers may be located on licensed spectrum. In embodiments, gNBs 180a, 180b, and 180c can implement Coordinated Multipoint (CoMP) technology. For example, WTRU 102a can receive coordinated transmissions from gNBs 180a and 180b (and / or gNB 180c).

[0080] WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using transmissions associated with Scalable Digital Numerology (SDN). For example, OFDM symbol spacing and / or OFDM subcarrier spacing can vary for different transmissions, different cells, and / or different portions of the radio transmission spectrum. WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using subframes of various or scalable lengths or Transmission Time Intervals (TTIs) (e.g., containing different numbers of OFDM symbols and / or absolute times of varying durations).

[0081] gNBs 180a, 180b, and 180c can be configured to communicate with WTRUs 102a, 102b, and 102c in standalone and / or non-standalone configurations. In standalone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c without also accessing other RANs (e.g., eNode-Bs 160a, 160b, and 160c). In standalone configuration, WTRUs 102a, 102b, and 102c can utilize one or more of gNBs 180a, 180b, and 180c as mobility anchors. In standalone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using signals in unlicensed frequency bands. In a non-standalone configuration, WTRUs 102a, 102b, and 102c can communicate / connect with gNBs 180a, 180b, and 180c, and also with another RAN such as eNode-Bs 160a, 160b, and 160c. For example, WTRUs 102a, 102b, and 102c can implement DC principles to communicate essentially simultaneously with one or more gNBs 180a, 180b, and 180c and one or more eNode-Bs 160a, 160b, and 160c. In a non-standalone configuration, eNode-Bs 160a, 160b, and 160c can act as mobility anchors for WTRUs 102a, 102b, and 102c, and gNBs 180a, 180b, and 180c can provide additional coverage and / or throughput to serve WTRUs 102a, 102b, and 102c.

[0082] Each of gNBs 180a, 180b, and 180c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, network slicing support, dual connectivity, interoperability between NR and E-UTRA, routing of user plane data to User Plane Functions (UPF) 184a and 184b, and routing of control plane information to Access and Mobility Management Functions (AMF) 182a and 182b, etc. Figure 1D As shown, gNB 180a, 180b, and 180c can communicate with each other via the Xn interface.

[0083] Figure 1DThe CN 115 shown may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and possibly a Data Network (DN) 185a, 185b. While each of the foregoing elements is depicted as part of the CN 115, it will be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.

[0084] AMF 182a and 182b can connect to one or more of gNBs 180a, 180b, and 180c in RAN 113 via the N2 interface and can be used as control nodes. For example, AMF 182a and 182b can be responsible for authenticating users of WTRU 102a, 102b, and 102c, supporting network slicing (e.g., handling different PDU sessions with different requirements), selecting specific SMF183a and 183b, managing registration areas, terminating NAS signaling, mobility management, etc. AMF 182a and 182b can use network slicing to customize CN support for WTRU 102a, 102b, and 102c based on the service types being utilized by WTRU 102a, 102b, and 102c. For example, different network slices can be established for different use cases, such as services relying on Ultra Reliable Low Latency (URLLC) access, services relying on Enhanced Massive Mobile Broadband (eMBB) access, and services for Machine Type Communication (MTC) access. AMF 162 can provide control plane functions for handover between RAN 113 and other RANs (not shown) employing other radio technologies such as LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies such as WiFi.

[0085] SMFs 183a and 183b can connect to AMFs 182a and 182b in CN 115 via the N11 interface. SMFs 183a and 183b can also connect to UPFs 184a and 184b in CN 115 via the N4 interface. SMFs 183a and 183b can select and control UPFs 184a and 184b, and configure the routing of traffic through UPFs 184a and 184b. SMFs 183a and 183b can perform other functions, such as managing and allocating UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, and providing downlink data notifications. PDU session types can be IP-based, non-IP-based, or Ethernet-based.

[0086] UPF 184a and 184b can connect to one or more of gNB 180a, 180b, and 180c in RAN 113 via the N3 interface. The N3 interface can provide WTRU 102a, 102b, and 102c with access to packet-switched networks such as Internet 110, facilitating communication between WTRU 102a, 102b, and 102c and IP-enabled devices. UPF 184 and 184b can perform other functions such as routing and forwarding packets, enforcing user plane policies, supporting multi-homed PDU sessions, handling user plane QoS, buffering downlink packets, and providing mobility anchoring.

[0087] CN 115 can facilitate communication with other networks. For example, CN 115 may include, or be able to communicate with, an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) serving as an interface between CN 115 and PSTN 108. Furthermore, CN 115 can provide WTRUs 102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, WTRUs 102a, 102b, and 102c may be connected to local data networks (DNs) 185a and 185b via the N3 interface to UPFs 184a and 184b and the N6 interface between UPFs 184a and 184b and DNs 185a and 185b.

[0088] Given Figures 1A-1D as well as Figures 1A-1D The corresponding descriptions herein indicate that one or more of the functions described herein in relation to one or more of the following can be implemented by one or more emulation devices (not shown): WTRU 102a-102d, base station 114a-114b, eNode-b 160a-160c, MME 162, SGW 164, PGW 166, gNB 180a-180c, AMF 182a-182b, UPF 184a-184b, SMF 183a-183b, DN 185a-185b, and / or any other device(s) described herein. An emulation device can be one or more devices configured to emulate one or more of the functions described herein. For example, an emulation device can be used to test other devices and / or simulate network and / or WTRU functions.

[0089] Simulation devices can be designed to perform one or more tests on other devices in laboratory and / or carrier network environments. For example, one or more simulation devices can perform one or more functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices within the communication network. One or more simulation devices can perform one or more functions while being temporarily implemented / deployed as part of a wired or wireless communication network. Simulation devices can be directly coupled to another device for testing purposes and / or can be used for testing via over-the-air wireless communication.

[0090] One or more simulation devices can perform one or more functions without being implemented / deployed as part of a wired and / or wireless communication network. For example, simulation devices can be used in test scenarios within test laboratories and / or undeployed (e.g., testing) wired and / or wireless communication networks to perform testing of one or more components. One or more simulation devices can be test rigs. Simulation devices can transmit and / or receive data using direct RF coupling and / or wireless communication via RF circuitry (e.g., which may include one or more antennas).

[0091] This application describes a wide variety of aspects, including tools, features, examples, models, methods, etc. Many of these aspects are described in a targeted manner and, at least to show individual characteristics, are generally described in a way that may sound restrictive. However, this is for clarity and does not limit the application or scope of those aspects. In fact, all the different aspects can be combined and interchanged to provide other aspects. Furthermore, these aspects can also be combined and interchanged with aspects described in earlier applications.

[0092] The aspects described and envisioned in this application can be implemented in many different forms. Figures 5-9 Some examples can be provided, but other examples are also envisioned. Figures 5-9 The discussion does not limit the breadth of implementations. At least one of these aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to the transmission of the generated or encoded bitstream. These and other aspects can be implemented as methods, apparatus, computer-readable storage media having instructions stored thereon for encoding or decoding video data according to any of the described methods, and / or computer-readable storage media having bitstreams generated according to any of the described methods stored thereon.

[0093] In this application, the terms “reconstruction” and “decoding” are used interchangeably, as are the terms “pixel” and “sample”, and the terms “image”, “picture” and “frame” are used interchangeably.

[0094] This document describes various methods, each of which includes one or more steps or actions to achieve the described method. Unless a specific order of steps or actions is required for the method to function correctly, the order and / or use of specific steps and / or actions can be modified or combined. Furthermore, terms such as "first," "second," etc., may be used in various examples to modify elements, components, steps, operations, etc., such as "first decoding" and "second decoding," for example. Unless specifically required, the use of such terms does not imply a reordering of the modified operations. Therefore, in this example, the first decoding does not need to be performed before the second decoding and can occur, for example, before, during, or within a time period overlapping with the second decoding.

[0095] The various methods and other aspects described in this application can be used to modify modules (e.g., decoding modules) of the video encoder 200 and video decoder 300, such as... Figure 2 and Figure 3 As shown. Furthermore, the subject matter disclosed herein can be applied to, for example, any type, format, or version of video codecs, whether described in pre-existing or future-developed standards or recommendations, and any extensions to such standards and recommendations. Unless otherwise indicated or technically excluded, the aspects described in this application may be used alone or in combination.

[0096] Various numerical values, such as 0, 1, 3, 7, etc., are used in the examples described in this application. These and other specific values ​​are used for illustrative purposes only, and the aspects described are not necessarily limited to these specific values.

[0097] Figure 2 A diagram illustrating an example video encoder is provided. Variations of the example encoder 200 are envisioned, but for clarity, encoder 200 is described below without depicting all anticipated variations.

[0098] Before being encoded, the video sequence can undergo pre-coding (201), for example, applying color transformations to the input color image (e.g., converting from RGB 4:4:4 to YCbCr 4:2:0), or remapping the input image components to obtain a more flexible signal distribution for compression (e.g., using histogram equalization of one of the color components). Metadata can be associated with pre-processing and attached to the bitstream.

[0099] In encoder 200, the image is encoded by encoder elements as described below. The image to be encoded is partitioned (202) and processed in units such as encoding / decoding units (CUs). Each unit is encoded using, for example, an intra-frame or inter-frame mode. When a unit is encoded intra-frame, it performs intra-frame prediction (260). In inter-frame mode, motion estimation (275) and compensation (270) are performed. The encoder determines (205) which of the intra-frame or inter-frame modes to use to encode the unit and indicates the intra / inter-frame determination by, for example, a prediction mode flag. For example, the prediction residual is calculated by subtracting (210) the predicted block from the original image block.

[0100] The predicted residual is then transformed (225) and quantized (230). The quantized transform coefficients, along with the motion vector and other syntax elements, are entropy encoded / decoded (245) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can also bypass both the transform and quantization, for example, the residual can be directly encoded / decoded without applying either the transform or quantization process.

[0101] The encoder decodes the encoded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inversely transformed (250) to decode the prediction residual. The decoded prediction residual and the predicted block are combined (255) to reconstruct the image block. An in-loop filter (265) is applied to the reconstructed image to perform, for example, deblocking / SAO (sample adaptive offset) filtering to reduce coding artifacts. The filtered image is stored in a reference image buffer (280).

[0102] Figure 3 This diagram illustrates an example of a video decoder. In the example decoder 300, as described below, the bitstream is decoded by decoder elements. The video decoder 300 is typically implemented as follows... Figure 2 The encoder 200 is the inverse of the decoder pass described in the document. The encoder 200 typically also performs video decoding as part of the encoding of the video data.

[0103] Specifically, the decoder's input includes a video bitstream, which can be generated by the video encoder 200. The bitstream is first entropy-decoded (330) to obtain transform coefficients, motion vectors, and other encoding / decoding information. Picture partitioning information indicates how the picture should be partitioned. Therefore, the decoder can partition the picture according to the decoded picture partitioning information (335). The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residual. The decoded prediction residual and the predicted block are combined (355) to reconstruct the image block. The predicted block can be obtained from intra-frame prediction (360) or motion-compensated prediction (e.g., inter-frame prediction) (375) (370). An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (380).

[0104] The decoded image can also undergo post-decoding processing (385), such as inverse color transformation (e.g., from YCbCr 4:2:0 to RGB 4:4:4) or inverse remapping of the remapping process performed in pre-encoding processing (201). Post-decoding processing can use metadata derived in pre-encoding processing and signaled in the bitstream. In the example, the decoded image (e.g., after applying an in-loop filter (365) and / or after post-decoding processing (385) if post-decoding processing is used) can be sent to a display device for rendering to the user.

[0105] Example context-adaptive binary arithmetic codecs (e.g., encoders and / or decoders) with dynamic model switching and / or parameterization can be used.

[0106] Figure 4 The diagram illustrates examples of systems in which the various aspects and examples described herein may be implemented. System 400 may be embodied as a device including the various components described below and configured to implement one or more aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptops, smartphones, tablets, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 400 may be embodied individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one example, the processing and encoder / decoder elements of system 400 are distributed across multiple ICs and / or discrete components. In various examples, system 400 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various examples, system 400 is configured to implement one or more aspects described in this document.

[0107] System 400 includes at least one processor 410 configured to execute instructions loaded thereon to implement various aspects, such as those described in this document. Processor 410 may include embedded memory, input / output interfaces, and various other circuitry known in the art. System 400 includes at least one memory 420 (e.g., a volatile memory device and / or a non-volatile memory device). System 400 includes a storage device 440, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 440 may include internal storage devices, attached storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices.

[0108] System 400 includes an encoder / decoder module 430 configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 430 may include its own processor and memory. The encoder / decoder module 430 represents one or more modules that can be included in a device to perform encoding and / or decoding functions. It is well known that a device may include one or both encoding and decoding modules. Additionally, as those skilled in the art will appreciate, the encoder / decoder module 430 may be implemented as a separate element of system 400, or may be incorporated into processor 410 as a combination of hardware and software.

[0109] Program code to be loaded onto processor 410 or encoder / decoder 430 to implement the various aspects described in this document may be stored in storage device 440 and subsequently loaded onto memory 420 for execution by processor 410. According to various examples, one or more of processor 410, memory 420, storage device 440, and encoder / decoder module 430 may store one or more various items during the execution of the processes described in this document. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.

[0110] In some examples, the memory within processor 410 and / or encoder / decoder module 430 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other examples, external memory (e.g., the processing device could be processor 410 or encoder / decoder module 430) is used for one or more of these functions. External memory could be memory 420 and / or storage device 440, such as volatile memory and / or non-volatile flash memory. In several examples, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one example, fast external volatile memory (such as RAM) is used as working memory for video encoding and decoding operations.

[0111] Inputs to the components of system 400 may be provided by various input devices as indicated in box 445. Such input devices include, but are not limited to: (i) a radio frequency (RF) section that receives, for example, RF signals transmitted over the air by a broadcaster; (ii) a component (COMP) input terminal (or a set of COMP input terminals); (iii) a universal serial bus (USB) input terminal; and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Figure 4 Other examples not shown include composite video.

[0112] In various examples, the input device of block 445 has associated corresponding input processing elements known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or limiting a signal band to a certain band); (ii) down-converting the selected signal; (iii) band-limiting it again to a narrower band to select, for example, a signal band that may be referred to as a channel in some examples; (iv) demodulating the down-converted and band-limited signal; (v) performing error correction; and / or (vi) demultiplexing to select a desired data packet stream. The RF section of various examples includes one or more elements for performing these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include tuners that perform various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or baseband. In one set-top box example, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to the desired frequency band. Various examples rearrange the order of the above (and other) components, remove some of these components, and / or add other components that perform similar or different functions. Adding components may include inserting components between existing components, such as, for example, inserting amplifiers and analog-to-digital converters. In various examples, the RF section includes an antenna.

[0113] USB and / or HDMI terminals may include corresponding interface processors for connecting system 400 to other electronic devices across USB and / or HDMI connections. It is to be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented as needed, for example, within a separate input processing IC or within processor 410. Similarly, aspects of USB or HDMI interface processing may be implemented as needed within a separate interface IC or within processor 410. Demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, processor 410 and encoder / decoder 430 operating in conjunction with memory and storage elements, to process the data streams as needed for presentation on an output device.

[0114] Various components of system 400 can be provided within an integrated housing. Within the integrated housing, the various components can be interconnected using a suitable connection arrangement 425 and data can be transmitted therebetween, such as internal buses known in the art, including inter-IC (I2C) buses, wiring, and printed circuit boards.

[0115] System 400 includes a communication interface 450 that enables communication with other devices via a communication channel 460. The communication interface 450 may include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 460. The communication interface 450 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 460 may be implemented, for example, within a wired and / or wireless medium.

[0116] In various examples, data is streamed or otherwise provided to system 400 using a wireless network such as Wi-Fi (e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers)). In these examples, the Wi-Fi signal is received via a communication channel 460 and a communication interface 450 suitable for Wi-Fi communication. The communication channel 460 in these examples is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other examples use a set-top box to provide streaming data to system 400, delivering data via an HDMI connection in input box 445. Still other examples use an RF connection in input box 445 to provide streaming data to system 400. As shown above, various examples provide data in a non-streaming manner. Additionally, various examples use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth® networks.

[0117] System 400 can provide output signals to various output devices, including displays 475, speakers 485, and other peripheral devices 495. Displays 475 in various examples include one or more of, for example, touchscreen displays, organic light-emitting diode (OLED) displays, curved displays, and / or foldable displays. Displays 475 can be used in televisions, tablets, laptops, mobile phones, or other devices. Displays 475 can also be integrated with other components (e.g., in smartphones) or standalone (e.g., as an external monitor for a laptop). In various examples, other peripheral devices 495 include one or more of standalone digital video discs (or digital multifunction discs) (DVDs, for both terms), disc players, stereo systems, and / or lighting systems. Various examples use one or more peripheral devices 495 that provide functionality based on the output of system 400. For example, a disc player performs the function of playing the output of system 400.

[0118] In various examples, signaling such as AV links, Consumer Electronics Control (CEC), or other communication protocols that enable device-to-device control with or without user intervention is used to transmit control signals between system 400 and display 475, speaker 485, or other peripheral devices 495. Output devices can be communicatively coupled to system 400 via dedicated connections through their respective interfaces 470, 480, and 490. Alternatively, output devices can be connected to system 400 via communication interface 450 using communication channel 460. Display 475 and speaker 485 can be integrated into a single unit with other components of system 400 in electronic devices such as, for example, televisions. In various examples, display interface 470 includes display drivers, such as, for example, timing controller (TCon) chips.

[0119] Alternatively, the display 475 and speaker 485 can be separated from one or more other components, for example, if the RF section of input 445 is part of a separate set-top box. In various examples where the display 475 and speaker 485 are external components, the output signal can be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.

[0120] The example can be implemented by computer software implemented by processor 410, or by hardware, or by a combination of hardware and software. As a non-limiting example, the example can be implemented by one or more integrated circuits. Memory 420 can be of any type suitable for the technical environment and can be implemented using any suitable data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples. Processor 410 can be of any type suitable for the technical environment and, as a non-limiting example, can encompass one or more of microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.

[0121] Various implementations involve decoding. As used in this application, "decoding" can encompass all or part of a process, such as performing a received encoded sequence to produce a final output suitable for display. In various examples, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various examples, such a process also includes, or alternatively includes, processes performed by decoders of the various implementations described in this application, such as: determining whether to disable the Non-Separable Master Transform (NSPT) for the current block; determining to use the Low-Frequency Non-Separable Transform (LFNST) for the current block based on the determination that NSPT is disabled for the current block; and decoding the current block based on the determination that LFNST is used, using Multiple Transform Selection (MTS), etc.

[0122] As further examples, in one example, "decoding" refers only to entropy decoding; in another example, "decoding" refers only to differential decoding; and in yet another example, "decoding" refers to a combination of entropy decoding and differential decoding. It will be clear, and those skilled in the art, whether the phrase "decoding processing" is intended to specifically refer to a subset of operations or generally to broader decoding processing, based on the specific context of the description.

[0123] Various implementations involve encoding. Similar to the discussion of "decoding" above, the term "encoding" as used in this application can encompass all or part of the process performed on an input video sequence to generate an encoded bitstream. In various examples, such a process includes one or more processes typically performed by an encoder, such as partitioning, differential coding, transform, quantization, and entropy coding. In various examples, such a process also includes, or alternatively includes, the processes performed by encoders of the various implementations described in this application, such as: disabling the Non-Separable Master Transform (NSPT) for the current block; enabling the Low-Frequency Non-Separable Transform (LFNST) for the current block; and encoding the current block based on Multiple Transform Selection (MTS), etc.

[0124] As further examples, in one example, "encoding" refers only to entropy encoding; in another, "encoding" refers only to differential encoding; and in yet another, "encoding" refers to a combination of differential and entropy encoding. It will be clear, and those skilled in the art, whether the phrase "encoding processing" is intended to specifically refer to a subset of operations or generally to broader encoding processing, based on the specific context of the description.

[0125] Note that all grammatical elements used in this article are descriptive terms. Therefore, they do not preclude the use of other grammatical element names.

[0126] When a diagram is presented as a flowchart, it should be understood that it also provides a block diagram of the corresponding device. Similarly, when a diagram is presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / process.

[0127] The implementations and aspects described herein can be implemented, for example, in methods or processes, apparatuses, software programs, data streams, or signals. Even if discussed only in the context of a single implementation (e.g., discussed only as a method), the features under discussion can be implemented in other forms (e.g., apparatuses or programs). Apparatuses can be implemented, for example, in appropriate hardware, software, and firmware. Methods can be implemented, for example, in a processor, which generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as, for example, computers, mobile phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate information communication between end users.

[0128] The reference to "an example" or "example" or "an implementation" or "implementation" and its variations means that the specific features, structures, characteristics, etc., described in connection with the example are included in at least one example. Therefore, the phrases "in an example" or "in the example" or "in an implementation" or "in the implementation" and any other variations appearing throughout this application do not necessarily all refer to the same example.

[0129] Additionally, this application may involve "determining" various pieces of information. Determining information may include, for example, one or more of estimated information, calculated information, predicted information, or information retrieved from memory. Acquiring may include receiving, retrieving, constructing, generating, and / or determining.

[0130] Furthermore, this application may involve "accessing" various information fragments. Accessing information may include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information, or one or more of these.

[0131] Additionally, this application may relate to "receiving" various pieces of information. Like "access," receiving is intended to be a broad term. Receiving information may include, for example, accessing information or retrieving information (e.g., from memory) from one or more sources. Furthermore, "receiving" typically relates in one or more ways during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0132] To be clear, for example, in the cases of “A / B,” “A and / or B,” and “at least one of A and B,” the use of any of the following “ / ,” “and / or,” and “at least one of…” is intended to cover selecting only the first listed option (A), or only the second listed option (B), or both options (A and B). As another example, in the cases of “A, B, and / or C” and “at least one of A, B, and C,” this wording is intended to cover selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A, B, and C). As will be apparent to those skilled in the art and related fields, this can be extended to as many items as are listed.

[0133] Furthermore, as used herein, the term "signal" refers, among other things, to something indicated to the corresponding decoder. Encoder signals may include, for example, sh_cabac_init_flag. In this way, in the examples, the same parameters are used on both the encoder and decoder sides. Thus, for example, the encoder can transmit (explicit signaling) specific parameters to the decoder so that the decoder can use the same specific parameters. Conversely, if the decoder already has specific parameters as well as other parameters, then signaling can then be used without transmission (implicit signaling) to simply allow the decoder to know and select specific parameters. In various examples, bit savings are achieved by avoiding the transmission of any actual functionality. It should be understood that signaling can be done in a variety of ways. For example, in various examples, one or more syntax elements, flags, etc., are used to signal information to the corresponding decoder. Although the foregoing refers to the verb form of the term "signal," the term "signal" can also be used as a noun in this text.

[0134] It will be apparent to those skilled in the art that implementations can generate a wide variety of signals formatted to carry, for example, information that can be stored or transmitted. The information may include, for example, instructions for implementing a method or data generated by one of the described implementations. For example, a signal may be formatted to carry a bit stream as described in the example. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of a spectrum) or a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. It is well known that signals can be transmitted via a wide variety of wired or wireless links. Signals may be stored on, accessed from, or received from a processor-readable medium.

[0135] This document describes numerous examples. Features of the examples may be provided individually or in any combination across various claim classes and types. Furthermore, examples may include one or more features, devices, or aspects described herein, individually or in any combination, across various claim classes and types. For example, features described herein may be implemented in a bitstream or signal including information generated as described herein. This information may allow a decoder to decode the bitstream, encoder, bitstream, and / or decoder according to any of the described embodiments. For example, features described herein may be implemented by creating and / or transmitting and / or receiving and / or decoding a bitstream or signal. For example, features described herein may be implemented by a method, process, apparatus, medium storing instructions, medium storing data, or signal. For example, features described herein may be implemented by a TV, set-top box, mobile phone, tablet computer, or other electronic device that performs decoding. The TV, set-top box, mobile phone, tablet computer, or other electronic device may display (e.g., using a monitor, screen, or other type of display) the obtained image (e.g., an image reconstructed from a residual of a video bitstream). The TV, set-top box, mobile phone, tablet computer, or other electronic device may receive and perform decoding of a signal including an encoded image.

[0136] Forward Low-Frequency Inseparable Transform (LFNST) can be applied. For example, forward LFNST can be applied after performing Discrete Cosine Transform 2 (DCT2) on one or more intra-frame codec blocks. On the decoder side, an inverse transform can be applied, for example, before the inverse DCT2 transform. The transform stages described herein can be associated with codec gain (e.g., high codec gain).

[0137] One or more transform stages (e.g., a single transform stage) may be added, for example, named the Inseparable Master Transform (NSPT). The added transform stages may replace the DCT2-LFNST stage, for example, in cases where an inseparable transform (e.g., a single inseparable transform) is used. The replacement of the DCT2-LFNST stage described herein may be permitted (e.g., only permitted) for small blocks, for example, because large blocks of inseparable transforms may require enormous memory and / or computational complexity.

[0138] Multiple transform selection (MTS) can be used. For MTS, Discrete Sine Transform 7 (DST7) and DST8 (e.g., DST7 and DST8 only) transform kernels can be utilized. DST7 and DST8 transform kernels can be used for intra-frame and inter-frame encoding and decoding.

[0139] Additional master transforms, including DCT5, DST4, DST1, and the identity transform (IDT), can be employed. Furthermore, the MTS set can be based on the size of the transform unit (TU) and / or intra-frame mode information. For blocks predicted via Intra-Template Matching Prediction (IntraTMP), a decoder-side intra-frame mode derivation (DIMD) process can be used on the predicted blocks, for example, to derive the intra-frame mode used for transform selection. For example, the horizontal and vertical gradients of the predicted samples can be computed, for example, to construct a gradient histogram (HoG). Intra-frame prediction modes with histogram amplitude values ​​(e.g., maximum histogram amplitude values) can be used for the MTS transform set.

[0140] Sixteen different TU sizes can be considered. For each TU size, one or more (e.g., five) different categories can be considered, for example, based on intra-frame mode information. For each category, one, four, and / or six different transform pairs can be considered. The number of intra-frame MTS candidates can be selected (e.g., adaptively selected between 1, 4, and 6 MTS candidates). For example, the number of intra-frame MTS candidates can be adaptively selected based on the sum of the absolute values ​​of the transform coefficients. The sum can be compared with one or more thresholds (e.g., two fixed thresholds). The sum after comparison can be used to determine the total number of allowed MTS candidates: One candidate: sum <= th0 Four candidates: th0 < sum <= th1 Six candidates: total > th1.

[0141] A total of 80 different categories can be considered. Some different categories may share similar (e.g., completely identical) transformation sets. The final lookup table (LUT) can have 58 entries (e.g., fewer than 80 unique entries).

[0142] For angular patterns, joint symmetry on the TU shape and intra-prediction can be considered. An intra-prediction pattern i (i>34) with a TU shape A×B can be mapped to the same category as a pattern j=(68–i) with a TU shape B×A. For transform pairs, the order of the horizontal and vertical transform kernels can be interchanged. For example, a 16×4 block with pattern 18 (e.g., horizontal prediction) and / or a 4×16 block with pattern 50 (e.g., vertical prediction) can be mapped to the same category. The vertical and horizontal transform kernels can be interchanged. For wide-angle intra-prediction patterns, an angular pattern (e.g., the closest regular angular pattern) can be used for transform set determination. For example, pattern 2 can be used for one or more (e.g., all) patterns between -2 and -14. Similarly, pattern 66 can be used for patterns 67 through 80.

[0143] Implicit MTS can be used. For example, implicit MTS can specify a transform selection mode, such as eliminating the need for MTS index signaling in the video data (e.g., the codec bitstream). If implicit MTS is enabled, the transform type for a given transform unit can be deduced from the block shape. Implicit MTS modes can be useful for encoders (e.g., fast encoder implementations). For example, the primary transform selection for rate distortion optimization can be skipped (e.g., it may not be necessary).

[0144] In implicit MTS mode, the master transform can be derived as follows: trTypeHor = (4<= tuWidth<= 16) DST7 : DCT2 trTypeVer = (4<= tuHeight<= 16) DST7 : DCT2.

[0145] For example, in a given direction (e.g., horizontal and / or vertical), the 1D transformation used could be DST7. The TU width could be between 4 and 16 (inclusive). And, for example, otherwise, the 1D transformation used could be DCT2.

[0146] The use of implicit MTS (e.g., compared to explicit MTS) can provide signaling at the sequence level (in the Sequence Parameter Set (SPS)). Explicit MTS, as described herein, may refer to the case where the MTS index is encoded or decoded within the video data (e.g., a bitstream).

[0147] Inter-frame MTS optimization can be used. For the MTS of inter-frame codec CUs, four candidates can be used: {(DST7,DST7), (DST7, DCT8), (DCT8, DST7), (DCT8, DCT8)}, which can be used for one or more (e.g., each) CUs. For higher resolution sequences (e.g., width > 1080), the maximum CU size used for inter-frame MTS can be set to 32 (e.g., for CUs with width <= 32 and height <= 32, use inter-frame MTS). For the remaining sequences (e.g., lower resolution), the maximum CU size used for inter-frame MTS can be set to 16. For 4-point, 8-point, and / or 16-point transforms, one or more adaptive multi-transform cores (e.g., DST-7 and / or DCT-8) can be replaced with separable Karhunen-Loève transforms (KLT).

[0148] Low-frequency non-separable transform (LFNST) (e.g., a secondary transform that may be applied after one or more primary transforms have been applied) and non-separable primary transform (NSPT) can be performed using the video codecs described in this paper. LFNSTs can include several LFNST sets (S) and candidates (C), which extend to S=35 and C=3. For a given intra-frame mode (predModeIntra), the LFNST set (lfnstTrSetIdx) can be derived according to the following formula: For predModeIntra < 2, lfnstTrSetIdx equals 2. lfnstTrSetIdx = predModeIntra, for predModeIntra in [0,34] lfnstTrSetIdx = 68 – predModeIntra, for predModeIntra in [35,66].

[0149] Three different kernels (LFNST4, LFNST8, and LFNST16) can be defined to indicate the LFNST kernel set, and they can be applied to 4×N / N×4 (N≥4), 8×N / N×8 (N≥8), and M×N (M, N≥16), respectively. The kernel dimension can be specified as follows: (LFSNT4, LFNST8*, LFNST16*) = (16x16, 32x64, 32x96).

[0150] Forward LFNST can be applied to the low-frequency region in the upper left corner, which can be called the region of interest (ROI). If LFNST is applied, the principal transform coefficients present in regions other than the ROI may be set to zero.

[0151] Figure 5 The diagram illustrates an example of the ROI for LFNST16. LFNST 16 can comprise six 4x4 sub-blocks, which can be scanned consecutively in a zigzag pattern. Since the input sample count is 96, the transformation matrix used for the forward LFNST16 can be R×96. R can be chosen as 32. Accordingly, 32 coefficients (two 4x4 sub-blocks) can be generated from the forward LFNST16, which can be scanned sequentially as follows... Figure 5 The coefficients shown are placed in the scanning order.

[0152] Figure 6 The diagram illustrates an example of an ROI for LFNST8. The forward LFNST8 matrix can be Rx64. R can be chosen as 32. The generated coefficients can be positioned in the same way as LFNST16.

[0153] Table 1 below shows the mapping from intra-frame prediction modes to these sets. Table 1: Mapping of intra-frame prediction modes to LFNST set indices.

[0154] LFNST and non-DCT2 master transforms can be used (e.g., mutually exclusive).

[0155] LFNST can be used (e.g., possibly only for) blocks that are encoded and decoded using DCT2-DCT2 as transform pairs (e.g., the main transform pairs in some video codec tools).

[0156] In an exemplary decoding process, the LFNST index can be decoded (e.g., first decoded from video data such as a bitstream). If the LFNST index of the CU under consideration is non-zero, the MTS index can be skipped (e.g., decoding can be omitted), and DCT2-DCT2 can be inferred as the transform pair (e.g., the master transform pair) of the block under consideration (e.g., corresponding to an MTS index equal to 0).

[0157] One or more conditions (e.g., additional conditions) for an MTS index to be decoded that is not inferred to be 0 may include one or more of the following: a transform skip indication (e.g., a flag) that decodes one or more transform coefficients may be equal to false; the position of the last significant decoded coefficient in the scan order may be at least 1; the codec block (e.g., codec unit) may not be in intra-segment sub-partition (ISP) mode or sub-block transform (SBT) mode; and / or MTS may be enabled relative to the CU codec mode (e.g., intra-frame mode and / or inter-frame mode) and / or CU size.

[0158] Figure 7 An example NSPT is illustrated, which replaces a two-stage transformation (DCT2-LFNST) with an inseparable transformation (e.g., a single inseparable transformation). This may (e.g., may only) be allowed for small blocks.

[0159] One or more (e.g., all) NSPTs can include 35 sets and 3 candidates (e.g., similar to the current LFNST). The kernel of an NSPT can have at least one of the following shapes: NSPT4x4: 16x16; NSPT4x8 / NSPT8x4: 32x20; NSPT8x8: 64x32; NSPT4x16 / NSPT16x4: 64x24; or NSPT8x16 / NSPT16x8: 128x40. Therefore, 12, 32, 40, and 88 coefficients can be zeroed using NSPT4x8 / NSPT8x4, NSPT8x8, NSPT4x16 / NSPT16x4, and NSPT8x16 / NSPT16x8, respectively.

[0160] In video encoding and decoding, compression efficiency can be improved, for example, for the transform portion. For instance, if NSPT is not used for a given codec unit, the possibility of jointly using LFNST and non-DCT2 main transforms can be introduced.

[0161] In the example, LFNST can be enabled for (e.g., only for) MTS index 0. Enabling LFNST for MTS index 0 can reduce complexity (e.g., without sacrificing any compression efficiency).

[0162] As described herein, a set of master transforms (e.g., supporting more than DCT2, DST7, and / or DCT8 master transforms) may be available. For example, DCT5, DST4, DST1, IDT, and / or other transforms described herein may share one or more features with DCT2. For example, the low-frequency basis functions of DCT2 may be flat. The low-frequency basis functions of DCT5 may also be flat. LFNST and DCT5 can work synergistically with each other, as is the case with LFNST and DCT2.

[0163] One or more other discrete cosine transforms (e.g., DCT1 and / or DCT6) may also have flat low-frequency basis functions. One or more transforms (e.g., DCT1 and / or DCT6) may also be combined with LFNST to form the master transform.

[0164] Figure 8A The illustration shows examples of the lowest frequency basis functions for DCT-V and DCT-II. Figure 8B Examples of the lowest frequency basis functions for DCT-IV and DCT-VIII are illustrated. Figure 8C Examples of the lowest frequency basis functions for DST-IV and DST-VII are illustrated. Figure 8D Examples of the lowest frequency basis functions for DST-I and DST-II are illustrated.

[0165] In the example, the MTS index and / or LFNST index can be signaled. For instance, the MTS index and / or LFNST index can be explicitly signaled.

[0166] If LFNST is used for a given codec unit, transform pairs other than DCT2-DCT2 (e.g., master transform pairs) can be allowed, even if NSPT is not allowed. For example, if the LFNST index is not 0, the MTS index can be not 0, and the NSPT phase can be disabled (e.g., not allowed) for the codec unit under consideration.

[0167] If the LFNST index is not 0 and NSPT is not allowed for the block under consideration, the MTS index can be signaled in the video data (e.g., in the bitstream).

[0168] In the examples, if LFNST is used for the codec unit, the MTS transform used for the intra-codec unit can be a transform (e.g., any transform used in the example video codec tool). In some examples, LFNST may be allowed (e.g., only allowed) for intra-blocks.

[0169] Decoding the MTS index can be done in a manner similar to (e.g., exactly the same as) decoding the MTX index in an intra-frame MTS extension (e.g., the intra-frame MTS extension in some video codecs), for example, regardless of the value of the LFNST index if NSPT is not allowed for the block under consideration.

[0170] In the example, the MTS index may include indications, such as flags. Indicators included in the MTS index may indicate the use of (DCT2,DCT2) or (DCT5,DCT5) transform pairs.

[0171] A dedicated MTS can be used for one or more non-zero LFNST index cases.

[0172] The master transform set (e.g., a specific set) can be used for intra-CUs using LFNST. For example, a reduced set of MTS can be used compared to intra-MTS (e.g., the intra-MTS of some video codec tools).

[0173] In the example, the reduced MTS set can be, or may include, transform pairs (e.g., (DCT5, DCT5) transform pairs) as a supplement and / or alternative to the (DCT2, DCT2) transform pair. For example, the MTS for LFNST adaptation can be the following: {(DCT2, DCT2), (DCT5, DCT5)}.

[0174] In the example, the MTS set dedicated to LFNST opening for one or more blocks can be one of the following: {(DCT2,DCT2),(DCT5,DCT5),(DCT2,DCT5),(DCT5,DCT2)}.

[0175] In the example, the set of transform pairs (e.g., the set of primary transform pairs) for a block using LFNST can be adapted to one or more properties of the transform coefficients included in the transform unit under consideration. For example, intra-frame MTS can be based on the sum of the transform coefficient magnitudes. sum The number of MTS candidates can be based on the following values: sum : One candidate, if sum≤th0 , corresponding to the MTS set {(DCT2,DCT2)} Two candidates, if th0 <sum≤th1 This corresponds to the MTS set {(DCT2,DCT2),(DCT5,DCT5)}. Four candidates, if th1 <sum , corresponding to the MTS set {(DCT2,DCT2),(DCT5,DCT5),(DCT2,DCT5),(DCT5,DCT2)}.

[0176] In the example, one or more extended master transformation sets may include one or more other transformations, such as DST1, DST7, DCT8, separable KLT, etc.

[0177] In the example for IntraMTS, for a TU with LFNST, the allowed set of transform pairs can also be based on the TU size and / or the intra-prediction mode associated with the TU under consideration. If the TU belongs to an intra-CU encoded and decoded in an intra-TMP mode (e.g., intra-template matching involves motion compensation from a reference block), the intra-mode associated with the current block can be derived using DIMD intra-mode derivation. For example, one or more specific transform classes different from those used in the intra-MTS (e.g., a regular intra-MTS) can be employed to map parameters (e.g., TU size, intra-mode) to a candidate set of MTS transform pairs. The candidate set can include one or more candidates, depending on the codec transform coefficients contained in the TU, as described herein.

[0178] In the examples, the MTS set used in conjunction with non-zero LFNST indices can be, or can include, one or more other master transforms, such as DCT1 and DCT6. The MTS set described herein can also have flat low-frequency basis function properties, such as DCT2 and DCT5.

[0179] The MTS set dedicated to the LFNST case can be signaled via intra-frame prediction mode (e.g., implicitly).

[0180] In the example, MTS index signaling can be skipped. For instance, for a CU that encodes and decodes with LFNST enabled, the MTS index can be omitted.

[0181] In the example, the transform pair can be inferred from a video decoding device, such as a decoder, based on one or more parameters associated with the CU under consideration.

[0182] Such parameters can be included in or can be found in the intra-prediction mode of the CU under consideration. For intra-CUs encoded and decoded using intra-TMP, the intra-mode can be derived using DIMD on the IntraTMP prediction module.

[0183] In the example, intra-frame mode parity can be used to indicate the transform pair between two candidates, such as (DCT2,DCT2) and (DCT5,DCT5).

[0184] In the example, the modulus value intramode%nbMTSCandidates (e.g., the Euclidean remainder of the intramode divided by the number of candidate transform pairs) can represent the index of the transform pairs used for the current block. The number of candidates and the candidate set can be designed based on transform unit size, intramode, encoder / decoder transform coefficient magnitude, etc., in the same manner as described herein (e.g., using a dedicated MTS for one or more non-zero LFNST indexes).

[0185] In the example, signaling can be sent from the video data (e.g., bitstream) to indicate CU-level and / or TU-level indicators (e.g., CU-level flags and / or TU-level flags). For example, if LFNST is enabled for the CU under consideration, signaling can be sent to indicate the CU-level flag and / or TU-level flag. Indicators (e.g., flags) can generally indicate whether implicit MTS derivation is used for the block under consideration. If implicit MTS derivation is not used for the block under consideration, a DCT2-DCT2 transform can be used (e.g., the default DCT2-DCT2 transform). If implicit MTS derivation is used for the block under consideration, a transform pair different from DCT2-DCT2 can be derived for the block under consideration.

[0186] The examples described herein can also be limited to implicit MTS as described herein. For example, if an implicit MTS is specified according to an indication (such as an advanced flag), the choice of the primary transform can alternate between DCT2 and DCT5 when using LFNST / NSPT. For example, DCT2 (e.g., for horizontal and / or vertical directions) can be considered for even-numbered intra-frame modes (e.g., and / or virtual intra-frame modes), and DCT5 can be considered for odd-numbered modes. Virtual intra-frame modes can correspond to (e.g., can typically correspond to) intra-frame modes derived from prediction blocks. For example, if prediction blocks are available in IntraTMP mode (e.g., relative to reconstructed samples located in the template region surrounding the current block), virtual intra-frame modes can correspond to intra-frame modes derived from prediction blocks. The considerations described herein can improve encoding / decoding flexibility without adding any additional overhead and / or complexity.

[0187] For one or more CU parameters, such as block size or codec mode (e.g., inter-frame codec mode), LFNST or NSPT may not be allowed. If LFNST and / or NSPT are not allowed, implicit derivation of the primary transform type described herein can be used to switch between DCT2 and DCT5 usage, for example. For example, the switch between DCT2 and DCT5 usage can be based on the intra-frame mode associated with the CU. Switching between DCT2 and DCT5 usage based on the intra-frame mode associated with the CU can improve compression efficiency. For example, the increased diversity of primary transform types introduced can improve compression efficiency compared to codec designs that support DCT2 for the CU (e.g., DCT2 only).

[0188] In the example, if LFNST and MTS are not used for the current block (e.g., MTS indicators (such as the MTS flag) are zero and LFNST indicators (such as the LFNST flag) are zero), the methods described herein can be used to diversify the selection of horizontal and vertical transforms. DCT2 can be used in both directions (e.g., by default). The alternation between DCT2 and DCT5 can be configured based on the current intra-frame mode and / or the current virtual intra-frame mode. That is, for even-numbered modes, DCT2 can be used in one or more (e.g., two) directions. For odd-numbered modes, DCT5 can be used. To further expand the diversity, the modulus of the intra-frame mode (e.g., and / or the virtual intra-frame mode) can be used to alternate between DCT2-DCT2, DCT2-DCT5, DCT5-DCT2, and / or DCT5-DCT5. The first transform can be used in the vertical direction and the second transform can be used in the horizontal direction.

[0189] One or more examples described herein may also include the possibility of implicitly signaling and / or using DCT1 and / or DCT6. Implicit signaling and / or using DCT1 and / or DCT6 may also have flat low-frequency basis function properties, such as DCT2 and DCT5.

[0190] Advanced signaling and / or restrictions can be configured based on the slice type.

[0191] In the example, for one or more CUs with LFNST enabled, enabling non-DCT2 master transform can (e.g., can only) be used for intra-slices.

[0192] In the example, for one or more CUs with LFNST enabled, the enabling of the non-DCT2 master transform can be used based on advanced indicators (e.g., advanced flags). For example, advanced indicators can be signaled in the image header and / or slice header.

[0193] In the examples, the advanced indicators described herein (e.g., advanced flags) can be signaled at the SPS level (e.g., SPS level indicators, such as SPS level flags). SPS level indicators can be overridden at the picture level and / or slice level. Some picture headers and / or slice headers may have dedicated indicators (e.g., dedicated flags).

[0194] One or more instances of inter-frame codec units can be configured.

[0195] One or more examples can be configured (e.g., only for intra-codec unit configuration) because LFNST can (e.g., can only) be used in intra-codec units.

[0196] The use of LFNST can be configured for inter-frame codec units. Such enabling can potentially result in compression efficiency gains.

[0197] In the example, the use of non-DCT2 transform can be configured for inter-frame codec units that use LFNST.

[0198] In video codec tools, the inter-frame MTS transform mode can enable the use of transforms such as DST7 and DCT8 (e.g., in addition to DCT2). As described in this article, separable KLT transforms for small blocks can be enabled.

[0199] In the examples, DCT5 can be configured to be enabled for inter-frame CUs in the LFNST active case, just as in the case of intra-frame codec units. For example, one or more examples described herein can be applied to inter-frame CUs as described herein. This may involve the derivation of the intra-frame codec mode for the inter-frame CU under consideration. For example, the derivation can be performed by applying an intra-frame mode derivation similar to DIMD to the inter-frame prediction block, such as one or more intra-frame TMP examples.

[0200] In the examples, transforms such as DST7 and DCT8 can also be used in conjunction with LFNST for (e.g., some) inter-frame codec units. Separable KLTs can also be used in conjunction with LFNST for (e.g., some) inter-frame codec units. Separable KLTs can be learned offline.

[0201] Inter-frame CUs can be in merge mode. If an inter-frame CU is in merge mode, motion information can be derived from merge candidates. Merge candidates can be previously encoded or decoded inter-frame CUs. Previously encoded or decoded CUs can serve as candidates, for example, for predicting motion data for the current inter-frame CU. Inter-frame CUs can be assigned merge indices. Merge indices can identify merge candidates for the CUs under consideration, for example, in a list of merge candidates constructed by video encoding devices (e.g., encoders) and / or video decoding devices (e.g., decoders).

[0202] In the example, if inter-frame merged CUs are used and LFNST is enabled, the use of the non-DCT2 master transform can be derived from the merge candidate. That is, if the selected merge candidate uses non-DCT2, the current CU can also be (e.g., it can be derived from the merge candidate).

[0203] Unmerged inter-frame CUs can be encoded and decoded, for example, in Advanced Motion Vector Prediction (AMVP) mode. In AMVP mode, motion data can be encoded and decoded, for example, in the form of motion vector predictor indices (e.g., indicating MV predictors in two possible nodes) and / or motion vector differences for those predictors. In the example, if LFNST is enabled, the use of non-(DCT2,DCT2) transform pairs can be hidden in the motion vector difference (MVD). For example, indications (such as flags) can be hidden in parity information of one or more MVD component values. For example, the binary element of an XOR operation equal to one or more x and y components of the MVD can be used to signal the use of non-(DCT2,DCT2).

[0204] Video encoding and decoding can be limited using picture and / or slice time depth, such as Low Latency-B (LDB) codec configuration and / or Low Latency-P (LDP) codec configuration.

[0205] In LDB codec configuration and / or LDP codec configuration, the following can be used: Figure 9 The image group structure is shown in the figure. Figure 9 The illustration shows the organization of example image groups in LDB encoding / decoding. For example... Figure 9 As illustrated, one or more consecutive images can be assigned a quantization parameter (QP). The assigned QP can vary from image to image, for example, making the rate distributions between different images extremely unequal.

[0206] For example, the QP (Queries Per Pixel) of an image can be set based on the arrangement of virtual hierarchies between images. Virtual hierarchies can assign depth values ​​to images, such as at the encoder level. The QP assigned to a given image can be based on the image's depth values. For example, as... Figure 9 The image shown is of the following type. B 3 You can specify a B-image with a depth value of 3. Figure 9 The diagram illustrates the QP assigned to the image depth, for example, relative to the sequence-level QP parameter.

[0207] Depth values ​​may not be signaled in video data (e.g., bitstream). For example, depth values ​​may be used on the encoder side (e.g., may be used only on the encoder side), such as in a GOP structure for organizing sequences using low-latency encoding and decoding.

[0208] In the example, MTS usage for the LFNST case can be activated / deactivated based on the temporal depth of the image and / or slice under consideration.

[0209] LFNST+ non-DCT2 usage can be restricted based on one or more surrounding blocks.

[0210] In non-inter-frame slices, the use of LFNST+non-DCT2 (also known as LFNST+MTS) can be limited to one or more CUs (e.g., certain CUs), for example, based on one or more characteristics of the surrounding CUs. One or more of the following options can be considered.

[0211] In the example, if one or more surrounding CUs (e.g., if most of the surrounding CUs) are intra-coded, then LFNST+MTS may be allowed (e.g., only allowed). In cases where intra-coding is used (e.g., overused), LFNST+MTS may be restricted to using a certain proportion of blocks (e.g., a small proportion of blocks).

[0212] In the example, LFNST+MTS may be permitted (e.g., only permitted if one or more surrounding CUs are inter-frame encoded) if one or more surrounding CUs are inter-frame encoded. LFNST+MTS scenarios may be limited if the current block is difficult to encode and / or if LFNST is expected to provide codec gain (e.g., significant codec gain).

[0213] In one or more examples (e.g., the two examples described in this article), advanced syntax elements can be used to indicate that LFNST+MTS is restricted based on the surrounding blocks.

[0214] Based on one or more examples described herein, methods, processes, apparatus, TVs, set-top boxes, mobile phones, tablets, media storing instructions, media storing data, syntax elements, bit streams, or signals can be used for encoding, decoding, storing, displaying, transmitting, and / or receiving data, or for one or more of these purposes. Although features and elements have been described above in specific combinations, those skilled in the art will appreciate that each feature or element can be used alone or in any combination with other features and elements. Furthermore, the methods described herein can be implemented in a computer program, software, or firmware incorporated in a computer-readable medium for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted via wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor storage devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROMs and digital multifunction discs (DVDs). A processor associated with the software can be used to implement a radio frequency transceiver for use in a WTRU, UE, terminal, base station, RNC, or any host computer.

Claims

1. An apparatus for video decoding, the apparatus comprising: The processor is configured as follows: Determine that the Non-Separable Main Transform (NSPT) is disabled for the codec unit; Based on the determination that NSPT is disabled for the codec unit, and the determination that Low Frequency Non-Separable Transform (LFNST) is enabled for the codec unit; and Based on the determination that LFNST is enabled for the codec unit, the codec unit is decoded based on Multiple Transform Selection (MTS).

2. The apparatus according to claim 1, wherein, The processor is configured as follows: Obtain the NSPT enable indicator from the video data, wherein the NSPT enable indicator is configured to indicate whether NSPT is enabled for the codec unit, and determine whether NSPT is enabled for the codec unit based on the NSPT enable indicator.

3. The apparatus according to claim 1 or claim 2, wherein, The processor is configured as follows: The LFNST enable indication is obtained from the video data, wherein the LFNST enable indication is configured to indicate whether LFNST is enabled for the codec unit, and the determination of whether LFNST is enabled for the codec unit is based on the LFNST enable indication.

4. The apparatus according to any one of claims 1-3, wherein, The processor is configured as follows: Ensure the LFNST index is not 0; and Based on determining that the LFNST index is not 0 and based on determining that NSPT is disabled for the codec unit, the MTS index associated with the MTS in the video data is obtained, wherein the MTS index includes the MTS index indicator.

5. The apparatus according to claim 4, wherein, The MTS index is configured to indicate the use of transform pairs, and the processor is further configured to perform an inverse transform based on the transform pairs.

6. The apparatus according to claim 5, wherein, The transform pair includes at least one of Discrete Cosine Transform 2 (DCT2) and DCT2 transform pair or DCT5 and DCT5 transform pair.

7. A method for video decoding, the method comprising: Determine that the Non-Separable Main Transform (NSPT) is disabled for the codec unit; Based on the determination that NSPT is disabled for the codec unit, and the determination that Low Frequency Non-Separable Transform (LFNST) is enabled for the codec unit; and Based on the determination that LFNST is enabled for the codec unit, the codec unit is decoded based on Multiple Transform Selection (MTS).

8. The method according to claim 7, wherein, The methods include: Obtain the NSPT enable indicator from the video data, wherein the NSPT enable indicator is configured to indicate whether NSPT is enabled for the codec unit, and determine whether NSPT is enabled for the codec unit based on the NSPT enable indicator.

9. The method according to claim 7 or claim 8, wherein, The methods include: The LFNST enable indication is obtained from the video data, wherein the LFNST enable indication is configured to indicate whether LFNST is enabled for the codec unit, and the determination of whether LFNST is enabled for the codec unit is based on the LFNST enable indication.

10. The method according to any one of claims 7-9, wherein, The methods include: Ensure the LFNST index is not 0; and Based on determining that the LFNST index is not 0 and based on determining that NSPT is disabled for the codec unit, the MTS index associated with the MTS in the video data is obtained, wherein the MTS index includes the MTS index indicator.

11. The method according to claim 10, wherein, The MTS index is configured to indicate the use of a transform pair, and the method further includes performing an inverse transform based on the transform pair.

12. The method according to claim 11, wherein, The transform pair includes at least one of Discrete Cosine Transform 2 (DCT2) and DCT2 transform pair or DCT5 and DCT5 transform pair.

13. An apparatus for video encoding, the apparatus comprising: The processor is configured as follows: Determine whether to disable the Inseparable Main Transform (NSPT) for the codec unit; Based on determining that NSPT is disabled for the codec unit, it is determined whether Low Frequency Inseparable Transform (LFNST) is enabled for the codec unit; and Based on the determination that LFNST is enabled for the codec unit, the codec unit is encoded based on Multiple Transform Selection (MTS).

14. The apparatus according to claim 13, wherein, The processor is configured as follows: The video data includes at least one of an NSPT enable indicator or an LFNST enable indicator, wherein the NSPT enable indicator is configured to indicate whether NSPT is enabled for the codec unit, and wherein the LFNST enable indicator is configured to indicate whether LFNST is enabled for the codec unit.

15. A method for a video encoding device, the method comprising: Determine whether to disable the Inseparable Main Transform (NSPT) for the codec unit; Based on the determination that NSPT is disabled for the codec unit, it is determined whether to enable Low Frequency Non-Separable Transform (LFNST) for the codec unit. Based on the determination that LFNST is enabled for the codec unit, the codec unit is encoded based on Multiple Transform Selection (MTS).

16. The method according to claim 15, wherein, The methods include: The video data includes at least one of an NSPT enable indicator or an LFNST enable indicator, wherein the NSPT enable indicator is configured to indicate whether NSPT is enabled for the codec unit, and wherein the LFNST enable indicator is configured to indicate whether LFNST is enabled for the codec unit.

17. A computer-readable medium comprising instructions for video decoding, the instructions causing one or more processors to perform the method according to any one of claims 7-12.

18. A computer-readable medium comprising instructions for video encoding, the instructions causing one or more processors to perform the method according to any one of claims 15-16.

19. Video data comprising information representing encoded output generated by one or more methods according to any one of claims 15-16.