Selecting decoder-side intra mode derivation (DIMD) merge mode based on template filtering usage
By introducing enable indicators for template filtering mode and DIMD merging mode in the video coding system, the problem of poor filter mode selection in existing video coding systems is solved, thereby improving coding efficiency and quality and optimizing the compression effect of video signals.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INTERDIGITAL CE PATENT HOLDINGS SAS
- Filing Date
- 2024-09-19
- Publication Date
- 2026-05-01
AI Technical Summary
Existing video coding systems lack effective filtering mode selection methods when choosing the intra-frame mode derivation and merging mode on the decoder side, resulting in poor coding efficiency and quality.
A template filtering mode-based enable indication and merging mode selection method is adopted. By analyzing the filtering consistency of adjacent blocks, it is determined whether to enable the template filtering mode and DIMD merging mode, thereby improving the encoding efficiency of the encoder and decoder.
It improves the efficiency and quality of video encoding, reduces data transmission and storage requirements, and optimizes the compression effect of video signals.
Smart Images

Figure CN121970316A_ABST
Abstract
Description
Cross-references to applications related to template filtering for selecting decoder-side intra-frame mode derivation (DIMD) merging modes.
[0001] This application claims the benefit of European Provisional Patent Application No. 23306681.0, filed on 2 October 2023, the contents of which are incorporated herein by reference. Background Technology
[0002] Video coding systems can be used to compress digital video signals, for example, to reduce the storage and / or transmission bandwidth required for such signals. Video coding systems can include, for example, block-based, wavelet-based, and / or object-based systems. Summary of the Invention
[0003] Systems, methods, and tools for selecting decoder-side intra-frame mode derivation (DIMD) merging modes based on enabled template filtering modes are disclosed. Exemplary devices for video encoding (e.g., video encoders) and / or exemplary devices for video decoding (e.g., video decoders) may have processors configured to perform one or more of the actions described herein.
[0004] In the example, a decoding device (such as a video decoder) can determine that a template filtering mode is enabled for a block. Based on this determination, the device can decode the block based on a DIMD merging mode. In the example, the device can obtain indications in the video data (e.g., a video bitstream), such as a template filtering mode enabled indication. This indication can be configured to indicate whether a template filtering mode is enabled for a block. Based on the obtained indication, the video decoder can determine whether a template filtering mode is enabled for that block.
[0005] In the example, a device (e.g., a video decoder) can determine that one or more neighboring blocks use the same filter. For example, one or more neighboring blocks (e.g., a majority of neighboring blocks among one or more neighboring blocks) can be associated with the calculation of the MHog of that block. The device can determine whether to enable template filtering mode for that block based on the fact that one or more neighboring blocks use the same filter.
[0006] In the example, the device (e.g., a video decoder) can determine that at least one DIMD neighboring block uses the same filtering as the current DIMD block. Based on the determination that at least one DIMD neighboring block uses the same filtering, the device can allow selection of a DIMD merging mode for decoding the block. For example, the device can decode the block using at least one of a template filtering mode and a DIMD merging mode.
[0007] In the example, a video encoding device (such as a video encoder) can determine that a block is associated with a template filtering mode. Based on this determination, the device can encode the block based on a DIMD merging mode. The device can include an indication (such as a template filtering mode enable indication) in the video data (e.g., a video bitstream). The template filtering mode enable indication can be configured to indicate that a template filtering mode is enabled.
[0008] In the example, a device (e.g., a video encoder) can determine that one or more neighboring blocks use the same filter. One or more neighboring blocks (e.g., a majority of neighboring blocks) can be associated with a computed merged gradient histogram (MHog) of that block. Based on the determination that one or more neighboring blocks use the same filter, the device can determine that the block is associated with a template filtering pattern.
[0009] In the example, a device (e.g., a video encoder) can determine that at least one DIMD neighboring block uses the same filtering as the current DIMD block. Based on the determination that at least one DIMD neighboring block uses the same filtering, the device can allow selection of a DIMD merging mode for encoding the block. For example, the device can encode the block based on at least one of a template filtering mode and a DIMD merging mode. The device can include an indication, such as a DIMD merging mode enable indication, in the video data (e.g., a video bitstream). The DIMD merging mode enable indication can be configured to indicate whether DIMD merging mode is enabled. Attached Figure Description
[0010] Figure 1A is a system diagram illustrating an exemplary communication system in which one or more of the disclosed embodiments may be implemented.
[0011] Figure 1B is a system diagram illustrating an exemplary wireless transmit / receive unit (WTRU) that can be used within the communication system shown in Figure 1A according to an embodiment.
[0012] Figure 1C is a system diagram illustrating an exemplary radio access network (RAN) and an exemplary core network (CN) that can be used within the communication system shown in Figure 1A according to an embodiment.
[0013] Figure 1D is a system diagram illustrating yet another exemplary RAN and yet another exemplary CN that can be used within the communication system shown in Figure 1A according to an embodiment.
[0014] Figure 2 shows an exemplary video encoder.
[0015] Figure 3 shows an exemplary video decoder.
[0016] Figure 4 shows an example of a system in which various aspects and examples can be implemented.
[0017] Figure 5 shows an exemplary L-shaped template around the current block (e.g., coding unit (CU)).
[0018] Figure 6 illustrates an exemplary technique for deriving prediction blocks.
[0019] Figure 7 illustrates an exemplary decoder-side intra-frame mode derivation (DIMD) process.
[0020] Figure 8 shows exemplary reference regions for DIMD, DIMD_T, and DIMD_L.
[0021] Figure 9 illustrates an exemplary DIMD process, for example, based on the merged gradient histogram (MHoG).
[0022] Figure 10 shows a set of exemplary samples that have been filtered. Detailed Implementation
[0023] A more detailed understanding can be obtained through the following description, which is given by way of example in conjunction with the accompanying drawings.
[0024] Figure 1A is a diagram illustrating an example communication system 100 in which one or more of the disclosed embodiments may be implemented. The communication system 100 may be a multiple access system providing content such as voice, data, video, messaging, and broadcasting to multiple wireless users. The communication system 100 enables multiple wireless users to access such content by sharing system resources (including wireless broadband). For example, the communication system 100 may employ one or more channel access methods, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal FDMA (OFDMA), Single Carrier FDMA (SC-FDMA), Zero-Tail Unique Word DFT Extended OFDM (ZT UW DTS-s OFDM), Unique Word OFDM (UW-OFDM), Resource Block Filtered OFDM, Filter Bank Multicarrier (FBMC), etc.
[0025] As shown in Figure 1A, the communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RAN 104 / 113, CN 106 / 115, public switched telephone network (PSTN) 108, Internet 110, and other networks 112. However, it will be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, and 102d may be any type of device configured to operate and / or communicate in a wireless environment. For example, WTRUs 102a, 102b, 102c, and 102d (any of which may be referred to as a “station” and / or “STA”) may be configured to transmit and / or receive wireless signals and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearable devices, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in the context of industrial and / or automated processing chains), consumer electronics devices, devices operating on commercial and / or industrial wireless networks, etc. Any of WTRUs 102a, 102b, 102c, and 102d may be interchangeably referred to as a UE.
[0026] The communication system 100 may also include base station 114a and / or base station 114b. Each of base stations 114a and 114b can be any type of device configured to wirelessly connect to at least one of WTRUs 102a, 102b, 102c, and 102d to facilitate access to one or more communication networks, such as CN 106 / 115, the Internet 110, and / or other networks 112. For example, base stations 114a and 114b can be base transceiver stations (BTS), Node-B, eNode B, home Node B, main eNode B, gNB, NR Node B, site controllers, access points (APs), wireless routers, etc. Although base stations 114a and 114b are each depicted as a single element, it will be understood that base stations 114a and 114b can include any number of interconnected base stations and / or network elements.
[0027] Base station 114a may be part of RAN 104 / 113, which may also include other base stations and / or network elements (not shown), such as base station controllers (BSCs), radio network controllers (RNCs), relay nodes, etc. Base station 114a and / or base station 114b may be configured to transmit and / or receive radio signals on one or more carrier frequencies, which may be referred to as cells (not shown). These frequencies may be in licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide coverage for a specific geographic area that may be relatively fixed or may change over time. A cell may also be divided into cell sectors. For example, the cell associated with base station 114a may be divided into three sectors. Therefore, in one embodiment, base station 114a may include three transceivers, i.e., one transceiver per sector of the cell. In embodiments, base station 114a may employ multiple-input multiple-output (MIMO) technology and may utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.
[0028] Base stations 114a and 114b can communicate with one or more of WTRUs 102a, 102b, 102c, and 102d via air interface 116, which can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). Any suitable radio access technology (RAT) can be used to establish air interface 116.
[0029] More specifically, as described above, the communication system 100 can be a multiple access system and can employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, base station 114a in RAN 104 / 113 and WTRUs 102a, 102b, 102c can implement radio technologies such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which can use Wideband CDMA (WCDMA) to establish the air interface 116. WCDMA can include communication protocols such as High-Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA can include High-Speed Downlink (DL) Packet Access (HSDPA) and / or High-Speed UL Packet Access (HSUPA).
[0030] In the embodiment, base station 114a and WTRUs 102a, 102b, 102c can implement radio technologies such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which can use Long Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-Advanced Pro (LTE-A Pro) to establish air interface 116.
[0031] In the embodiment, base station 114a and WTRUs 102a, 102b, 102c can implement radio technologies such as NR radio access, which can use New Radio (NR) to establish air interface 116.
[0032] In the embodiments, base station 114a and WTRUs 102a, 102b, and 102c can implement multiple radio access technologies. For example, base station 114a and WTRUs 102a, 102b, and 102c can, for instance, use the dual connectivity (DC) principle to jointly implement LTE radio access and NR radio access. Therefore, the air interface used by WTRUs 102a, 102b, and 102c can be characterized by multiple types of radio access technologies and / or transmissions sent to / from multiple types of base stations (e.g., eNBs and gNBs).
[0033] In other embodiments, base station 114a and WTRUs 102a, 102b, 102c can implement radio technologies such as IEEE 802.11 (i.e., Wi-Fi), IEEE 802.16 (i.e., WiMAX), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Provisional Standard 2000 (IS-2000), Provisional Standard 95 (IS-95), Provisional Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rate GSM Evolution (EDGE), GSMEDGE (GERAN), etc.
[0034] Base station 114b in Figure 1A can be, for example, a wireless router, a home NodeB, a home eNodeB, or an access point, and can utilize any suitable RAT to facilitate wireless connectivity in localized areas such as commercial locations, homes, vehicles, campuses, industrial facilities, air corridors (e.g., for use by drones), roads, etc. In one embodiment, base station 114b and WTRUs 102c, 102d can implement radio technologies such as IEEE 802.11 to establish a wireless local area network (WLAN). In another embodiment, base station 114b and WTRUs 102c, 102d can implement radio technologies such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, base station 114b and WTRUs 102c, 102d can utilize cellular-based RATs (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-a, LTE-a Pro, NR, etc.) to establish a pico or femtocell. As shown in Figure 1A, base station 114b can be directly connected to the Internet 110. Therefore, base station 114b does not need to access Internet 110 via CN 106 / 115.
[0035] RAN 104 / 113 can communicate with CN 106 / 115, which can be any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRUs 102a, 102b, 102c, and 102d. Data can have different Quality of Service (QoS) requirements, such as different throughput requirements, latency requirements, fault tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, etc. CN 106 / 115 can provide call control, billing services, location-based services, prepaid calling, internet connectivity, video distribution, etc., and / or perform advanced security functions (such as user authentication). Although not shown in Figure 1A, it will be understood that RAN 104 / 113 and / or CN 106 / 115 can communicate directly or indirectly with other RANs using the same RAT as RAN 104 / 113 or a different RAT. For example, in addition to connecting to RAN 104 / 113, which may be using NR radio technology, CN 106 / 115 can also communicate with another RAN (not shown) using GSM, UMTS, CDMA 2000, WiMAX, E-UTRA, or WiFi radio technology.
[0036] CN 106 / 115 can also serve as a gateway for WTRU 102a, 102b, 102c, 102d to access PSTN 108, the Internet 110, and / or other networks 112. PSTN 108 may include a circuit-switched telephone network providing Common Old-Style Telephone Service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices using common communication protocols such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) from the TCP / IP Internet Protocol suite. Network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, network 112 may include another CN connected to one or more RANs, which may use the same RAT as RAN 104 / 113 or a different RAT.
[0037] Some or all of the WTRUs 102a, 102b, 102c, and 102d in the communication system 100 may include multi-mode capabilities (e.g., WTRUs 102a, 102b, 102c, and 102d may include multiple transceivers for communicating with different wireless networks via different wireless links). For example, the WTRU 102c shown in Figure 1A may be configured to communicate with a base station 114a that may employ cellular-based radio technology and with a base station 114b that may employ IEEE 802 radio technology.
[0038] Figure 1B is a system diagram illustrating an example WTRU 102. As shown in Figure 1B, WTRU 102 may include a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power supply 134, a Global Positioning System (GPS) chipset 136, and / or other peripheral devices 138, etc. It will be understood that, while remaining consistent with the embodiments, WTRU 102 may include any sub-combination of the foregoing elements.
[0039] Processor 118 may be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. Processor 118 may perform signal encoding, data processing, power control, input / output processing, and / or any other functions that enable WTRU 102 to operate in a wireless environment. Processor 118 may be coupled to transceiver 120, which may be coupled to transmitting / receiving element 122. Although Figure 1B depicts processor 118 and transceiver 120 as separate components, it will be understood that processor 118 and transceiver 120 may be integrated together in an electronic package or on a chip.
[0040] Transmitting / receiving element 122 can be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) via air interface 116. For example, in one embodiment, transmitting / receiving element 122 can be an antenna configured to transmit and / or receive RF signals. In another embodiment, transmitting / receiving element 122 can be a transmitter / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, transmitting / receiving element 122 can be configured to transmit and / or receive both RF signals and optical signals. It will be understood that transmitting / receiving element 122 can be configured to transmit and / or receive any combination of wireless signals.
[0041] Although the transmit / receive element 122 is depicted as a single element in FIG. 1B, the WTRU 102 may include any number of transmit / receive elements 122. More specifically, the WTRU 102 may employ MIMO technology. Thus, in one embodiment, the WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals via the air interface 116.
[0042] Transceiver 120 can be configured to modulate signals transmitted by transmitting / receiving element 122 and demodulate signals received by transmitting / receiving element 122. As described above, WTRU 102 can have multi-mode capability. Therefore, transceiver 120 can include multiple transceivers for enabling WTRU 102 to communicate via various RATs (e.g., such as NR and IEEE 802.11).
[0043] The processor 118 of WTRU 102 can be coupled to and receive user input data from: a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) unit or an organic light-emitting diode (OLED) display unit). The processor 118 can also output user data to the speaker / microphone 124, keypad 126, and / or display / touchpad 128. Additionally, the processor 118 can access information and store data from any suitable type of memory, such as non-removable memory 130 and / or removable memory 132. Non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. Removable memory 132 may include a subscriber identity module (SIM) card, memory stick, secure digital storage (SD) card, etc. In other embodiments, the processor 118 can access information and store data from memory that is not physically located on WTRU 102 (such as on a server or home computer (not shown)).
[0044] The processor 118 may receive power from the power supply 134 and may be configured to distribute power to other components in the WTRU 102 and / or control power to those other components. The power supply 134 may be any suitable device for powering the WTRU 102. For example, the power supply 134 may include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.
[0045] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) about the current location of the WTRU 102. In addition to or instead of information from the GPS chipset 136, the WTRU 102 may receive location information from base stations (e.g., base stations 114a, 114b) via air interface 116 and / or determine its location based on the timing of signals received from two or more nearby base stations. It will be understood that, while remaining consistent with the embodiments, the WTRU 102 may acquire location information using any suitable location determination method.
[0046] The processor 118 may also be coupled to other peripheral devices 138, which may include one or more software and / or hardware modules that provide additional features, functions, and / or wired or wireless connectivity. For example, peripheral devices 138 may include accelerometers, electronic compasses, satellite transceivers, digital cameras (for photos and / or videos), Universal Serial Bus (USB) ports, vibration devices, television transceivers, hands-free headsets, Bluetooth® modules, FM radio units, digital music players, media players, video game player modules, internet browsers, virtual reality and / or augmented reality (VR / AR) devices, activity trackers, etc. Peripheral devices 138 may include one or more sensors, which may be one or more of the following: gyroscopes, accelerometers, Hall effect sensors, magnetometers, orientation sensors, proximity sensors, temperature sensors, time sensors; geolocation sensors; altimeters, light sensors, touch sensors, magnetometers, barometers, gesture sensors, biometric sensors, and / or humidity sensors.
[0047] WTRU 102 may include a full-duplex radio, for which the transmission and reception of some or all of the signals (e.g., signals associated with a specific subframe of both UL (e.g., for transmission) and downlink (e.g., for reception)) may be concurrent and / or simultaneous. The full-duplex radio may include an interference management unit to reduce and / or substantially eliminate self-interference through signal processing via hardware (e.g., a choke) or via a processor (e.g., a separate processor (not shown) or via processor 118). In embodiments, WTRU 102 may include a half-duplex radio, for which the transmission and reception of some or all of the signals (e.g., signals associated with a specific subframe of both UL (e.g., for transmission) and downlink (e.g., for reception)) are separate.
[0048] Figure 1C is a system diagram illustrating RAN 104 and CN 106 according to an embodiment. As described above, RAN 104 can employ E-UTRA radio technology to communicate with WTRUs 102a, 102b, and 102c via air interface 116. RAN 104 can also communicate with CN 106.
[0049] RAN 104 may include eNode-Bs 160a, 160b, and 160c, but it will be understood that RAN 104 may include any number of eNode-Bs while remaining consistent with the embodiments. eNode-Bs 160a, 160b, and 160c may each include one or more transceivers for communicating with WTRUs 102a, 102b, and 102c via air interface 116. In one embodiment, eNode-Bs 160a, 160b, and 160c may implement MIMO technology. Therefore, for example, eNode-B 160a may use multiple antennas to transmit radio signals to and / or receive radio signals from WTRU 102a.
[0050] Each of the eNode-B 160a, 160b, and 160c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, etc. As shown in Figure 1C, the eNode-B 160a, 160b, and 160c can communicate with each other via the X2 interface.
[0051] The CN 106 shown in Figure 1C may include a Mobility Management Entity (MME) 162, a Serving Gateway (SGW) 164, and a Packet Data Network (PDN) Gateway (or PGW) 166. While each of the foregoing elements is described as part of the CN 106, it will be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.
[0052] The MME 162 can connect to each of the eNode-Bs 160a, 160b, and 160c in RAN 104 via the S1 interface and can act as a control node. For example, the MME 162 can be responsible for authenticating users of WTRUs 102a, 102b, and 102c, activating / deactivating bearers, selecting a specific serving gateway during the initial attachment of WTRUs 102a, 102b, and 102c, etc. The MME 162 can provide control plane functions for handover between RAN 104 and other RANs (not shown) employing other radio technologies such as GSM and / or WCDMA.
[0053] The SGW 164 can connect to each of the eNode Bs 160a, 160b, and 160c in RAN 104 via the S1 interface. The SGW 164 can typically route and forward user data packets to or from WTRUs 102a, 102b, and 102c. The SGW 164 can perform other functions, such as anchoring the user plane during eNode-B handover, triggering paging when DL data is available to WTRUs 102a, 102b, and 102c, and managing and storing the context of WTRUs 102a, 102b, and 102c.
[0054] SGW 164 can be connected to PGW 166, which can provide WTRU 102a, 102b, 102c with access to packet-switched networks (such as Internet 110) to facilitate communication between WTRU 102a, 102b, 102c and IP-enabled devices.
[0055] CN 106 can facilitate communication with other networks. For example, CN 106 can provide WTRUs 102a, 102b, and 102c with access to circuit-switched networks (such as PSTN 108) to facilitate communication between WTRUs 102a, 102b, and 102c and traditional terrestrial line communication equipment. For example, CN 106 may include an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between CN 106 and PSTN 108, or can communicate with it. Additionally, CN 106 can provide WTRUs 102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.
[0056] Although the WTRU is depicted as a wireless terminal in Figures 1A through 1D, it is conceivable that in some representative embodiments, such a terminal may (e.g., temporarily or permanently) use a wired communication interface with a communication network.
[0057] In a representative embodiment, the other network 112 may be a WLAN.
[0058] A WLAN in Infrastructure Basic Services Set (BSS) mode may have an Access Point (AP) for the BSS and one or more Stations (STAs) associated with the AP. The AP may access or interface with a Distribution System (DS) or another type of wired / wireless network that carries traffic entering and / or leaving the BSS. Traffic originating outside the BSS destined for a STA can be delivered to the AP. Traffic originating from a STA destined for a destination outside the BSS can be sent to the AP for delivery to the appropriate destination. Traffic between STAs within the BSS can be sent via the AP, for example, where a source STA can send traffic to the AP, and the AP can deliver the traffic to the destination STA. Traffic between STAs within the BSS can be considered and / or referred to as peer-to-peer traffic. Peer-to-peer traffic can be sent between a source STA and a destination STA using a Direct Link Setup (DLS) (e.g., directly between them). In some representative embodiments, the DLS may use 802.11e DLS or 802.11z Tunneled DLS (TDLS). A WLAN using the Standalone BSS (IBSS) mode may not have an access point (AP), and STAs within the IBSS or using the IBSS (e.g., all STAs) can communicate directly with each other. The IBSS communication mode may sometimes be referred to as the "ad-hoc" communication mode in this article.
[0059] When operating in 802.11ac infrastructure mode or a similar mode, the AP can transmit beacons on a fixed channel, such as the primary channel. The primary channel can be of fixed width (e.g., a bandwidth of 20 MHz) or dynamically set via signaling. The primary channel can be the operating channel of the BSS and can be used by the STA to establish a connection with the AP. In some representative embodiments, Carrier Sense Multiple Access (CSMA / CA) with collision avoidance can be implemented, for example, in an 802.11 system. For CSMA / CA, the AP STA (e.g., each STA) can sense the primary channel. If a particular STA senses / detects that the primary signal is busy and / or determines that the primary signal is busy, that particular STA can back off. In a given BSS, at any given time, only one STA (e.g., only one station) can transmit.
[0060] High-throughput (HT) STAs can communicate using a 40 MHz wide channel, for example, by combining a primary 20 MHz channel with adjacent or non-adjacent 20 MHz channels.
[0061] Very High Throughput (VHT) STAs can support channels with widths of 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz. 40 MHz and / or 80 MHz channels can be formed by combining consecutive 20 MHz channels. A 160 MHz channel can be formed by combining eight consecutive 20 MHz channels, or by combining two non-consecutive 80 MHz channels, which can be referred to as an 80+80 configuration. In the 80+80 configuration, data, after channel coding, can be passed through a fragment resolver that splits the data into two streams. Inverse Fast Fourier Transform (IFFT) processing and time-domain processing can be performed on each stream separately. The streams can be mapped onto the two 80 MHz channels, and the data can be transmitted by the transmitting STA. At the receiver of the receiving STA, the above operations for the 80+80 configuration can be reversed, and the combined data can be sent to the Media Access Control (MAC).
[0062] 802.11af and 802.11ah support operating modes below 1 GHz. The channel operating bandwidth and carrier in 802.11af and 802.11ah are reduced compared to those used in 802.11n and 802.11ac. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV Blank (TVWS) spectrum, and 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. According to a representative embodiment, 802.11ah may support instrument-type control / machine-type communication (MTC), such as MTC devices in macro coverage areas. MTC devices may have certain capabilities, including, for example, limited capabilities to support (e.g., only support) certain and / or limited bandwidths. MTC devices may include batteries with a battery life exceeding a threshold (e.g., to maintain a very long battery life).
[0063] WLAN systems that can support multiple channels and channel bandwidths (such as 802.11n, 802.11ac, 802.11af, and 802.11ah) include a channel that can be designated as the primary channel. The primary channel can have a bandwidth equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel can be set and / or limited by the STAs operating in the BSS that support the minimum bandwidth operating mode. In the 802.11ah example, for STAs that support (e.g., only support) the 1 MHz mode (e.g., MTC type devices), the primary channel can be 1 MHz wide even if the AP and other STAs in the BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operating modes. Carrier Sense and / or Network Assignment Vector (NAV) settings can depend on the status of the primary channel. If the primary channel is busy, for example, because an STA (which only supports the 1 MHz operating mode) is transmitting to the AP, the entire available band may be considered busy even if most of the band remains idle and potentially available.
[0064] In the United States, the available frequency band for 802.11ah is 902 MHz to 928 MHz. In South Korea, the available frequency band is 917.5 MHz to 923.5 MHz. In Japan, the available frequency band is 916.5 MHz to 927.5 MHz. The total available bandwidth for 802.11ah is 6 MHz to 26 MHz, depending on the country code.
[0065] Figure 1D is a system diagram illustrating RAN 113 and CN 115 according to an embodiment. As described above, RAN 113 may employ NR radio technology to communicate with WTRUs 102a, 102b, and 102c via air interface 116. RAN 113 may also communicate with CN 115.
[0066] RAN 113 may include gNBs 180a, 180b, and 180c, but it will be understood that RAN 113 may include any number of gNBs while remaining consistent with the embodiments. gNBs 180a, 180b, and 180c may each include one or more transceivers for communicating with WTRUs 102a, 102b, and 102c via air interface 116. In one embodiment, gNBs 180a, 180b, and 180c may implement MIMO technology. For example, gNBs 180a and 180b may utilize beamforming to transmit signals to and / or receive signals from gNBs 180a, 180b, and 180c. Thus, for example, gNB 180a may use multiple antennas to transmit radio signals to and / or receive radio signals from WTRU 102a. In an embodiment, gNBs 180a, 180b, and 180c may implement carrier aggregation technology. For example, gNB 180a can transmit multiple component carriers to WTRU 102a (not shown). A subset of these component carriers may be located on unlicensed spectrum, while the remaining component carriers may be located on licensed spectrum. In embodiments, gNBs 180a, 180b, and 180c can implement Coordinated Multipoint (CoMP) technology. For example, WTRU 102a can receive coordinated transmissions from gNBs 180a and 180b (and / or gNB 180c).
[0067] WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using transmissions associated with a scalable set of parameters. For example, the OFDM symbol spacing and / or OFDM subcarrier spacing can be varied for different transmissions, different cells, and / or different portions of the radio transmission spectrum. WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using subframes or transmission time intervals (TTIs) of various lengths or scalable lengths (e.g., including different numbers of OFDM symbols and / or absolute times of varying durations).
[0068] gNBs 180a, 180b, and 180c can be configured to communicate with WTRUs 102a, 102b, and 102c in standalone and / or non-standalone configurations. In standalone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c without accessing other RANs (e.g., eNodeBs 160a, 160b, and 160c). In standalone configuration, WTRUs 102a, 102b, and 102c can use one or more of gNBs 180a, 180b, and 180c as mobile anchors. In standalone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using signals in unlicensed frequency bands. In a non-standalone configuration, WTRUs 102a, 102b, and 102c can communicate / connect with gNBs 180a, 180b, and 180c while also communicating / connecting with another RAN (such as eNode-Bs 160a, 160b, and 160c). For example, WTRUs 102a, 102b, and 102c can implement DC principles to communicate substantially simultaneously with one or more gNBs 180a, 180b, and 180c and one or more eNode-Bs 160a, 160b, and 160c. In a non-standalone configuration, eNode-Bs 160a, 160b, and 160c can act as mobile anchors for WTRUs 102a, 102b, and 102c, and gNBs 180a, 180b, and 180c can provide additional coverage and / or throughput to serve WTRUs 102a, 102b, and 102c.
[0069] Each of gNBs 180a, 180b, and 180c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in the uplink (UL) and / or downlink (DL), support for network slicing, dual connectivity, interoperability between NR and E-UTRA, routing of user plane data to User Plane Functions (UPF) 184a and 184b, routing of control plane information to Access and Mobility Management Functions (AMF) 182a and 182b, etc. As shown in Figure 1D, gNBs 180a, 180b, and 180c can communicate with each other via the Xn interface.
[0070] The CN 115 shown in Figure 1D may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and possibly a Data Network (DN) 185a, 185b. While each of the foregoing elements is described as part of the CN 115, it will be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.
[0071] AMF 182a and 182b can connect to one or more of the gNBs 180a, 180b, and 180c in RAN 113 via the N2 interface and can act as control nodes. For example, AMF 182a and 182b can be responsible for authenticating users of WTRU 102a, 102b, and 102c, supporting network slicing (e.g., handling different PDU sessions with different requirements), selecting specific SMFs 183a and 183b, managing registration areas, terminating NAS signaling, mobility management, etc. AMF 182a and 182b can use network slicing to customize CN support for WTRU 102a, 102b, and 102c based on the service types being used by WTRU 102a, 102b, and 102c. For example, different network slices can be established for different use cases, such as services relying on Ultra Reliable Low Latency (URLLC) access, services relying on Enhanced Massive Mobile Broadband (eMBB) access, and services for Machine Type Communication (MTC) access. AMF 162 can provide control plane functions for handover between RAN 113 and other RANs (not shown) that employ other radio technologies (such as LTE, LTE-A, LTE-A Pro) and / or non-3GPP access technologies (such as WiFi).
[0072] SMFs 183a and 183b can connect to AMFs 182a and 182b in CN 115 via the N11 interface. SMFs 183a and 183b can also connect to UPFs 184a and 184b in CN 115 via the N4 interface. SMFs 183a and 183b can select and control UPFs 184a and 184b, and configure them to route traffic through UPFs 184a and 184b. SMFs 183a and 183b can perform other functions, such as managing and allocating UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, and providing downlink data notifications. PDU session types can be IP-based, non-IP-based, or Ethernet-based.
[0073] UPF 184a and 184b can connect via the N3 interface to one or more of the gNBs 180a, 180b, and 180c in RAN 113. These gNBs can provide WTRU 102a, 102b, and 102c with access to packet-switched networks (such as the Internet 110) to facilitate communication between WTRU 102a, 102b, and 102c and IP-enabled devices. UPF 184a and 184b can perform other functions such as routing and forwarding packets, enforcing user plane policies, supporting multi-homed PDU sessions, handling user plane QoS, buffering downlink packets, and providing mobility anchoring.
[0074] CN 115 can facilitate communication with other networks. For example, CN 115 may include or be able to communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between CN 115 and PSTN 108. Additionally, CN 115 can provide WTRUs 102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, WTRUs 102a, 102b, and 102c can be connected to DN 185a and 185b via UPF 184a and 184b through the N3 interface to UPF 184a and 184b and the N6 interface between UPF 184a and 184b and local data networks (DNs) 185a and 185b.
[0075] Based on the corresponding descriptions in Figures 1A-1D, one or more emulation devices (not shown) can perform one or more or all of the functions described herein with respect to one or more of the following: WTRU 102a-102d, base stations 114a-114b, eNode-B 160a-160c, MME 162, SGW 164, PGW 166, gNB 180a-180c, AMF 182a-182b, UPF 184a-184b, SMF 183a-183b, DN 185a-185b, and / or any other devices described herein. An emulation device can be one or more devices configured to emulate one or more or all of the functions described herein. For example, an emulation device can be used to test other devices and / or simulate network and / or WTRU functions.
[0076] Simulation devices can be designed to perform one or more tests on other devices in laboratory and / or carrier network environments. For example, one or more simulation devices may perform one or more functions when fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices within the communication network. One or more simulation devices may perform one or more functions when temporarily implemented / deployed as part of a wired and / or wireless communication network. Simulation devices may be directly coupled to another device for testing purposes and / or may use over-the-air wireless communication to perform tests.
[0077] One or more simulation devices may perform one or more functions without being implemented / deployed as part of a wired and / or wireless communication network. For example, a simulation device may be used to test scenarios in a laboratory and / or an undeployed (e.g., tested) wired and / or wireless communication network to perform testing of one or more components. One or more simulation devices may be test devices. Simulation devices may transmit and / or receive data using direct RF coupling and / or wireless communication via an RF circuit system (e.g., which may include one or more antennas).
[0078] This application describes various aspects, including tools, features, examples, models, methods, etc. Many of these aspects are described in a specific manner, and are generally described in a way that may sound restrictive, at least to illustrate the individual features. However, this is for the purpose of clarity of description and does not limit the application or scope of these aspects. In fact, all the different aspects can be combined and interchanged to provide other aspects. Furthermore, these aspects can also be combined and interchanged with those described in previous documents.
[0079] The aspects described and contemplated in this application can be implemented in many different forms. Figures 5 through 11 described herein provide some examples, but other examples are contemplated. The discussion of Figures 5 through 11 does not limit the breadth of implementations. At least one aspect generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting the generated or encoded bitstream. These and other aspects can be implemented as methods, apparatus, computer-readable storage media having instructions thereon stored thereon for encoding or decoding video data according to any of the methods, and / or computer-readable storage media having bitstreams generated according to any of the methods stored thereon.
[0080] In this application, the terms “reconstruction” and “decoding” are used interchangeably, the terms “pixel” and “sample” are used interchangeably, and the terms “image”, “picture” and “frame” are used interchangeably.
[0081] This document describes various methods, each of which includes one or more steps or actions to implement the described method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions can be modified or combined. Additionally, in various examples, terms such as "first," "second," etc., may be used to modify elements, components, steps, operations, etc., such as, for example, "first decoding" and "second decoding." Unless specifically required, the use of such terms does not imply a sequence of operations. Therefore, in this example, the first decoding does not need to be performed before the second decoding, but can be performed, for example, before, during, or in a time period overlapping with the second decoding.
[0082] The various methods and other aspects described in this application can be used to modify modules of the video encoder 200 and decoder 300 shown in Figures 2 and 3, such as decoding modules. Furthermore, the subject matter disclosed herein can be applied to, for example, any type, format, or version of video coding, whether described in standards or recommendations, whether pre-existing or future-developed, and any extensions to such standards and recommendations. Unless otherwise stated or technically excluded, these aspects described in this application may be used alone or in combination.
[0083] Various numerical values, such as 1, 2, 4, 7, 8, 16, 32, 64, etc., are used in the examples described in this application. These and other specific values are for illustrative purposes only, and the aspects described are not limited to these specific values.
[0084] Figure 2 is a diagram illustrating an exemplary video encoder 200. Variations of the exemplary encoder 200 are contemplated, but for clarity, the encoder 200 is described below without describing all anticipated variations.
[0085] Before being encoded, the video sequence may undergo pre-coding processing 201 (e.g., applying a color transformation to the input color image (e.g., a conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing a remapping of the input image components) to obtain a signal distribution more suited to compression (e.g., using histogram equalization of one of the color components). Metadata (e.g., which may include film grain parameters determined by the pre-processing described herein) may be associated with the pre-processing and appended to the bitstream.
[0086] In encoder 200, the frame is encoded by encoder elements as described below. The frame to be encoded is partitioned (202) and processed in units such as coding units (CUs). Each unit is encoded using, for example, an intra-frame or inter-frame mode. When a unit is encoded in intra-frame mode, intra-frame prediction (260) is performed. In inter-frame mode, motion estimation (275) and compensation (270) are performed. The encoder determines (205) which mode, intra-frame or inter-frame, to use to encode the unit, and indicates the intra-frame / inter-frame decision by, for example, a prediction mode flag. For example, the prediction residual is calculated by subtracting (210) the prediction block from the original image block.
[0087] The predicted residual is then transformed (225) and quantized (230). The quantized transform coefficients, motion vectors, and other syntax elements are entropy encoded (245) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., directly encode the residual without applying either the transform or quantization process.
[0088] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residual. The decoded prediction residual and the prediction block are combined (255) to reconstruct the image block. An in-loop filter (265) is applied to the reconstructed image to perform, for example, deblocking / SAO (sample adaptive offset) filtering, thereby reducing coding artifacts. The filtered image is stored in a reference image buffer (280).
[0089] Figure 3 is a diagram illustrating an exemplary video decoder. In the exemplary decoder 300, the bitstream is decoded by decoder elements as described below. The video decoder 300 typically performs a decoding process that is the inverse of the encoding process described in Figure 2. The encoder 200 typically also performs video decoding as part of the encoded video data.
[0090] Specifically, the decoder's input includes a video bitstream, which can be generated by the video encoder 200. First, entropy decoding (330) is performed on the bitstream to obtain transform coefficients, motion vectors, and other encoded information. Picture partitioning information indicates how the picture should be partitioned. Therefore, the decoder can partition (335) the picture based on the decoded picture partitioning information. The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residuals. The decoded prediction residuals and prediction blocks are combined (355) to reconstruct the image blocks. Prediction blocks (370) can be obtained from intra-frame prediction (360) or motion-compensated prediction (i.e., inter-frame prediction) (375). An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (380).
[0091] The decoded image can also undergo post-decoding processing (385), such as inverse color transformation (e.g., a conversion from YCbCr 4:2:0 to RGB 4:4:4) or inverse remapping, which is the inverse of the remapping process performed in the pre-encoding process (201). Post-decoding processing can use metadata derived in the pre-encoding process and signaled in the bitstream. In the example, the decoded image (e.g., after applying an in-loop filter (365) and / or, in the case of post-decoding processing, after post-decoding processing (385)) can be sent to a display device for presentation to the user.
[0092] Figure 4 is a diagram illustrating an example of a system in which the various aspects and examples described herein can be implemented. System 400 can be implemented as a device including the various components described below and configured to perform one or more aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 400 can be implemented individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one example, the processing elements and encoder / decoder elements of system 400 are distributed across multiple ICs and / or discrete components. In various examples, system 400 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various examples, system 400 is configured to implement one or more aspects described in this document.
[0093] System 400 includes at least one processor 410 configured to execute instructions loaded thereon to implement various aspects described herein, such as those described. Processor 410 may include embedded memory, input / output interfaces, and various other circuitry known in the art. System 400 includes at least one memory 420 (e.g., a volatile memory device and / or a non-volatile memory device). System 400 includes a storage device 440 that may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 440 may include internal storage devices, attached storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices.
[0094] System 400 includes an encoder / decoder module 430 configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 430 may include its own processor and memory. The encoder / decoder module 430 represents a module that can be included in a device to perform encoding and / or decoding functions. It is well known that a device can include one or both of the encoding and decoding modules. Alternatively, the encoder / decoder module 430 may be implemented as a separate element of system 400, or it may be incorporated into processor 410 as a combination of hardware and software known to those skilled in the art.
[0095] Program code to be loaded onto processor 410 or encoder / decoder 430 to execute the various aspects described in this document may be stored in storage device 440 and subsequently loaded onto memory 420 for execution by processor 410. Depending on various examples, one or more of processor 410, memory 420, storage device 440, and encoder / decoder module 430 may store one or more of various items during the execution of the processes described in this document. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from equations, formulas, operations, and operational logic processing.
[0096] In some examples, the memory within processor 410 and / or encoder / decoder module 430 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other examples, external memory (e.g., the processing device could be processor 410 or encoder / decoder module 430) is used for one or more of these functions. External memory could be memory 420 and / or storage device 440, such as volatile memory and / or non-volatile flash memory. In several examples, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one example, fast external volatile memory such as RAM is used as working memory for video encoding and decoding operations.
[0097] Inputs to the components of system 400 can be provided through various input devices, as indicated in box 445. Such input devices include, but are not limited to, (i) a radio frequency (RF) section that receives, for example, RF signals transmitted over the air by a broadcaster, (ii) a component (COMP) input terminal (or a set of COMP input terminals), (iii) a universal serial bus (USB) input terminal, and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Other examples not shown in Figure 4 include composite video.
[0098] In various examples, the input device of block 445 has associated input processing elements known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or limiting the signal band to a band), (ii) down-converting the selected signal, (iii) further band-limiting to a narrower band to select (e.g.,) a signal band that may be referred to as a channel in some examples), (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and / or (vi) demultiplexing to select the desired data packet stream. The RF section of various examples includes one or more elements performing these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include tuners performing various functions among these functions, including, for example, down-converting a received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or baseband. In one set-top box example, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and filtering again to the desired frequency band. Various examples rearrange the above (and other) components, remove some of them, and / or add other components that perform similar or different functions. Adding components may include inserting components between existing components, such as inserting amplifiers and analog-to-digital converters. In various examples, the RF section includes an antenna.
[0099] USB and / or HDMI terminals may include corresponding interface processors for connecting system 400 to other electronic devices via USB and / or HDMI connections. It should be understood that aspects of input processing (e.g., Reed-Solomon error correction) may be implemented, for example, within a separate input processing IC or within processor 410 as needed. Similarly, aspects of USB or HDMI interface processing may be implemented, as needed, within a separate interface IC or within processor 410. The demodulated, error-corrected, and demultiplexed stream is provided to various processing elements (including, for example, processor 410 and encoder / decoder 430), which operate in conjunction with memory and storage elements to process the data stream as needed for presentation on the output device.
[0100] Various components of system 400 can be housed within an integrated housing. Within the integrated housing, the various components can be interconnected and transmit data between them using suitable connection devices 425 (e.g., internal buses known in the art, including inter-IC (I2C) buses, wiring, and printed circuit boards).
[0101] System 400 includes a communication interface 450 that enables communication with other devices via a communication channel 460. The communication interface 450 may include, but is not limited to, a transceiver configured to send and receive data via the communication channel 460. The communication interface 450 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 460 may be implemented, for example, within a wired and / or wireless medium.
[0102] In various examples, wireless networks, such as Wi-Fi networks (e.g., IEEE 802.11, where IEEE refers to the Institute of Electrical and Electronics Engineers), are used to stream or otherwise provide data to system 400. In these examples, the Wi-Fi signal is received via a communication channel 460 and a communication interface 450 adapted for Wi-Fi communication. The communication channel 460 in these examples is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other examples use a set-top box to provide streaming data to system 400, delivering data via an HDMI connection to input block 445. Still other examples use an RF connection to input block 445 to provide streaming data to system 400. As indicated above, various examples provide data in non-streaming modes. Additionally, various examples use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth® networks.
[0103] System 400 can provide output signals to various output devices, including display 475, speaker 485, and other peripheral devices 495. Various examples of display 475 include one or more of, for example, touchscreen displays, organic light-emitting diode (OLED) displays, curved displays, and / or foldable displays. Display 475 can be used in televisions, tablets, laptops, cellular phones (mobile phones), or other devices. Display 475 can also be integrated with other components (e.g., in a smartphone) or separate (e.g., an external monitor for a laptop computer). In various examples, other peripheral devices 495 include one or more of stand-alone digital video discs (or digital multifunction discs) (DVDs, for both terms), disc players, stereo systems, and / or lighting systems. Various examples use one or more peripheral devices 495 that provide functionality based on the output of system 400. For example, a disc player performs the function of playing the output of system 400.
[0104] In various examples, signaling (such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable device-to-device control with or without user intervention) is used to transmit control signals between system 400 and display 475, speaker 485, or other peripheral devices 495. Output devices may be communicatively coupled to system 400 via dedicated connections through corresponding interfaces 470, 480, and 490. Alternatively, output devices may be connected to system 400 via communication interface 450 using communication channel 460. Display 475 and speaker 485 may be integrated into a single unit along with another component of system 400 in an electronic device, such as a television. In various examples, display interface 470 includes a display driver, such as a timing controller (TCon) chip.
[0105] For example, if the RF input section 445 is part of a separate set-top box, the display 475 and speaker 485 can alternatively be separate from one or more of the other components. In various examples where the display 475 and speaker 485 are external components, the output signal can be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.
[0106] The example can be executed by processor 410 or by computer software implemented by hardware or a combination of hardware and software. As a non-limiting example, the example can be implemented by one or more integrated circuits. As a non-limiting example, memory 420 can be of any type suitable for the technical environment and can be implemented using any suitable data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. As a non-limiting example, processor 410 can be of any type suitable for the technical environment and can encompass one or more microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.
[0107] Various implementations involve decoding. As used herein, "decoding" can encompass all or part of a process performed, for example, on a received encoded sequence to produce a final output suitable for display. In various examples, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various examples, such a process may also include, or alternatively may include, processes performed by the decoder of the various implementations described herein.
[0108] As another example, in one example, "decoding" refers only to entropy decoding; in another example, "decoding" refers only to differential decoding; and in yet another example, "decoding" refers to a combination of entropy decoding and differential decoding. Based on the specific context of the description, it will be clear whether the phrase "decoding process" is intended to specifically refer to a subset of operations or to refer to the broader decoding process, and it is believed that those skilled in the art will understand this well.
[0109] Various implementations involve encoding. Similar to the discussion above regarding "decoding," the "encoding" used in this application can include, for example, all or part of a process performed on an input video sequence to produce an encoded bitstream. In various examples, such processes include one or more processes typically performed by an encoder, such as partitioning, differential coding, transform, quantization, and entropy coding. In various examples, such processes can also include, or alternatively may include, processes performed by an encoder of the various implementations described in this application.
[0110] As another example, in one example, "encoding" refers only to entropy encoding; in another example, "encoding" refers only to differential encoding; and in yet another example, "encoding" refers to a combination of entropy encoding and differential encoding. It will be clear from the specific context of the description whether the phrase "encoding process" is intended to refer specifically to a subset of operations or to a broader encoding process, and it is believed that those skilled in the art will understand this well.
[0111] It should be noted that the names of grammatical elements used in this article are descriptive terms only. Therefore, the use of other grammatical element names is not excluded.
[0112] When a diagram is presented as a flowchart, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when a diagram is presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / process.
[0113] The implementations and aspects described herein can be implemented in, for example, methods or processes, apparatuses, software programs, data streams, or signals. Even if discussed only in the context of a single implementation (e.g., discussed only as a method), the implementation of the features in question can be implemented in other forms (e.g., apparatuses or programs). Apparatuses can be implemented, for example, with appropriate hardware, software, and firmware. Methods can be implemented in, for example, a processor, which generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as computers, cellular phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate information communication between end users.
[0114] The reference to “an example” or “an example” or “an implementation” or “an implementation”, and their variations, means that the specific features, structures, characteristics, etc., described in connection with the example are included in at least one example. Therefore, the phrases “in an example” or “in the example” or “in an implementation” or “in the implementation”, and any other variations, appearing throughout this application, do not necessarily refer to the same example.
[0115] Additionally, this application may relate to "determining" various types of information. Determining information may include one or more of the following: for example, estimation information, calculation information, prediction information, and information retrieved from memory. Obtaining may include receiving, retrieving, constructing, generating, and / or determining.
[0116] Furthermore, this application may involve "accessing" various types of information. Accessing information may include one or more of the following: for example, receiving information, retrieving information (e.g., retrieving information from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, and estimating information.
[0117] Additionally, this application may relate to "receiving" various types of information. Like "accessing," receiving is a broad term. Receiving information may include one or more of the following: for example, accessing information and retrieving information (e.g., retrieving information from memory). Furthermore, "receiving" is generally referred to in one or more ways during operation, such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0118] It should be understood that the use of any of the following “ / ”, “and / or”, and “…” (e.g., in the cases of “A / B”, “A and / or B”, and “at least one of A and B”) is intended to cover selecting only the first listed option (A), or only the second listed option (B), or selecting both options (A and B). As yet another example, in the cases of “A, B, and / or C” and “at least one of A, B, and C”, this wording is intended to include selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or selecting all three options (A, B, and C). As will be apparent to those skilled in the art and related fields, this can be extended to a large number of listed items.
[0119] Furthermore, as used herein, the term "signaling" specifically refers to instructing the corresponding decoder to provide certain information. Encoder signals may include, for example, the number of intensity intervals, the number of model values, particle parameters, particle identifiers, scaling factors, etc. In this way, in the examples, the same parameters are used on both the encoder and decoder sides. Thus, for example, the encoder can send (explicit signaling) specific parameters to the decoder so that the decoder can use the same specific parameters. Conversely, if the decoder already has specific parameters as well as other parameters, signaling can be used without sending (implicit signaling) to simply allow the decoder to know and select specific parameters. Bit savings are achieved in various examples by avoiding the transmission of any actual functionality. It should be understood that signaling can be implemented in various ways. For example, in various examples, one or more syntax elements, flags, etc., are used to signal information to the corresponding decoder. Although the verb form of the word "signal" was mentioned above, the word "signal" can also be used as a noun in this article.
[0120] As will be apparent to those skilled in the art, implementations can generate various signals, which are formatted to carry, for example, information that can be stored or transmitted. The information may include, for example, instructions for performing a method, or data generated by one of the described implementations. For example, a signal may be formatted to carry a bitstream of the described example. Such a signal may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or baseband signals. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. It is well known that signals can be transmitted via a variety of different wired or wireless links. Signals may be stored on, accessed from, or received from, a processor-readable medium.
[0121] This document describes numerous examples. Features of the examples may be provided individually or in any combination across various claim classes and types. Furthermore, examples may include one or more of the features, devices, or aspects described herein, individually or in any combination across various claim classes and types. For example, the features described herein may be implemented in a bitstream or signal that includes information generated as described herein. This information may allow a decoder to decode the bitstream, and the encoder, bitstream, and / or decoder may be any of the embodiments described. For example, the features described herein may be implemented by creating and / or transmitting and / or receiving and / or decoding a bitstream or signal. For example, the features described herein may be implemented by a method, process, apparatus, medium storing instructions, medium storing data, or signal. For example, the features described herein may be implemented by a TV, set-top box, cellular phone, tablet computer, or other electronic device performing decoding. The TV, set-top box, cellular phone, tablet computer, or other electronic device may (e.g., using a monitor, screen, or other type of display) display the resulting image (e.g., an image reconstructed from the residual of a video bitstream). The TV, set-top box, cellular phone, tablet computer, or other electronic device may receive a signal including the encoded image and perform decoding.
[0122] This article describes the features associated with decoder-side intra-frame mode derivation (DIMD).
[0123] If DIMD is applied, one or more (e.g., up to five) intra-frame modes can be derived from one or more reconstructed neighboring samples (e.g., by analyzing the directionality of the content around the current block). One or more predictors (e.g., five predictors) can be combined with planar mode predictors (e.g., with weights derived from the gradient histogram (HoG)). The HoG can be computed on an L-shaped template formed by the reconstructed samples (e.g., an L-shaped template with three sample widths / heights), as shown in Figure 5. A Sobel filter can be used to obtain the template (e.g., for samples within the gray area in Figure 5, accumulating the magnitudes of all gradients in a given direction). The direction with the highest accumulated magnitude can be selected as the primary DIMD mode and the secondary DIMD mode. Predictors obtained using DIMD modes can be blended to form the final DIMD prediction. Uniform or spatial blending can be used (e.g., where DIMD predictors are combined with planar predictors, e.g., using weights that depend on the relative magnitudes of the modes in the gradient histogram).
[0124] The same lookup table (LUT)-based integerization scheme used in cross-component linear models (CCLM) can be used to perform division in weight derivation. The division in orientation calculation can be described in (4).
[0125]
[0126] The division operation in orientation computation can be computed using LUT-based schemes in (5), (6), (7) and / or (8):
[0127] The following can be used to describe the (e.g., possible) values of DivSigTable in (9).
[0128]
[0129] For size For example, if either the upper histogram magnitude or the left histogram magnitude is twice that of the other, the weights of the derivation patterns (e.g., each of the five derivation patterns) can be modified. In this case, the weights may be position-dependent. The weights can be calculated. For example, if the upper histogram is twice the size of the left histogram, the weights can be calculated according to equation (8):
[0130] For example, if the left histogram is twice the size of the upper histogram, the weights can be calculated using equation (9):
[0131] Refer to equations (8) and (9). It can be the unmodified uniform weights of DIMD, and This can be predefined (e.g., set to 10). The weight of the plane is fixed at 21 / 64 (~1 / 3). The remaining 43 / 64 (~2 / 3) of the weight can be shared between HoG IPMs (e.g., two HoG IPMs), for example, proportional to the amplitude of their HoG bars, as shown in Figure 6.
[0132] Figure 6 illustrates an exemplary technique for deriving prediction blocks.
[0133] Deduced intra-frame patterns can be included in a master list of most probable intra-frame patterns (MPMs). For example, the DIMD procedure can be performed before constructing the MPM list. The master deduced intra-frame pattern of a DIMD block can be stored along with the block. The master deduced intra-frame pattern can be used to construct the MPM list for adjacent blocks.
[0134] The regions of adjacent reconstructed samples (e.g., for calculating gradient histograms) can be based on (e.g., depending on) the availability of one or more reconstructed samples. The region of the decoding reference sample for the current W×H luminance coded block (CB) can extend towards the upper right (e.g., up to W additional columns, if available). A CB can be a subset of coded units (CUs) associated with (e.g., related to) a given component. For example, a CU can be and / or can include (e.g., up to) 3 CBs, such as one CU corresponding to one component (Y, Cb, and / or Cr components). A luminance CB can be and / or can correspond to a set of luminance coded samples in a given CU. The region of the decoding reference sample for the current W×H luminance CB can extend towards the lower left (e.g., up to H additional rows, if available).
[0135] Figure 7 illustrates an exemplary DIMD procedure 700 (e.g., sometimes referred to as a conventional DIMD procedure). As shown, at 710, adjacent reconstruction reference samples (e.g., as reference regions) can be obtained and / or selected. At 720, oG can be obtained and / or derived. At 730, intra-frame modes (e.g., up to five intra-frame modes) can be obtained and / or derived. At 740, one or more DIMD fusion weights can be obtained and / or derived. Predictions can then be established.
[0136] This paper describes the features associated with the adaptive reference region DIMD.
[0137] One or more additional (e.g., two additional) DIMD patterns (e.g., DIMD_T and DIMD_L) can allow the selection of one or more distinct neighborhoods as reference regions. In DIMD_T mode, the top-left, top-up, and / or top-right neighborhoods can be used as reference regions. In DIMD_L mode, the top-left, left-side, and / or bottom-left neighborhoods can be used as reference regions. In both DIMD_T and DIMD_L modes, the number of rows in the reference region can be four. The original DIMD pattern is called DIMD_TL.
[0138] Figure 8 shows example reference areas for DIMD (e.g., DIMD_TL), DIMD_T, and DIMD_L modes.
[0139] In an encoding unit (CU) (e.g., each CU), an indication such as a flag (e.g., cu_dimd_mode) can be signaled (e.g., after the DIMD enable flag cu_dimd_flag). For example, an indication such as a flag can be signaled to determine which DIMD mode to use. Exemplary binarization of cu_dimd_mode can be summarized in Table 1.
[0140] Table 1: Binarization of cu_dimd_mode
[0141] This article can describe the characteristics associated with DIMD merging. Figure 9 illustrates an exemplary DIMD process 900, for example, based on the merge gradient histogram (MHoG).
[0142] If DIMD merging is used, DIMD information extracted from one or more neighboring blocks can be used, for example, to compute the intra-prediction of the current block. The MHoG of the current block can be computed (e.g., based on the HoG of neighboring blocks). One or more neighboring blocks encoded using DIMD and / or DIMD merging can be considered (e.g., only neighboring blocks encoded using DIMD and / or DIMD merging) (e.g., as shown at 930).
[0143] If a DIMD or DIMD-merged neighboring block (e.g., a single DIMD or DIMM-merged neighboring block) is available, its gradient histogram can be used to form the MHoG of the current block. If more than one DIMD and / or DIMD-merged neighboring block is available, one or more corresponding histograms can be combined (e.g., by interval amplitude averaging) to obtain and / or derive the MHoG (as shown at 940).
[0144] MHoG can be used to compute intra-frame prediction modes and weights (e.g., as in conventional DIMD). Orientation modes and / or weights corresponding to the N (e.g., N=5) highest amplitudes in the MHoG can be obtained and / or selected (e.g., as shown at 950). The corresponding predictors can be mixed (e.g., as in conventional DIMD), as shown at 960.
[0145] DIMD merging may be available (e.g., available only) if the current block has at least one neighboring block encoded using DIMD and / or DIMD merge (as shown at 920). In this case, the use of DIMD merging can be signaled (e.g., using a CU-level flag and / or indication encoded in CABAC). A CABAC context may be included to support the encoding of the DIMD merge flag and / or indication. In terms of signaling, DIMD merging can be considered a sub-mode of DIMD. For example, if DIMDflag=1 and if there are neighboring CUs encoded using DIMD and / or DIMD merge, the DIMD merge flag and / or indication can be signaled (e.g., signaled only).
[0146] DIMD derivation can be configured with a filter template.
[0147] The template filtering method for DIMD mode can be configured. For example, the template can be filtered using a 3×3 filter operator before obtaining the gradient histogram. If the gradient is computed, one or more different gradient operators (e.g., 3×2 and / or 2×3 gradient operators) can be used (e.g., instead of 3×3 filtering and / or other than 3×3 filtering).
[0148] In the example, the template can be filtered using a 3×3 filter operator, as depicted in (10).
[0149]
[0150] In the example, a 3×3 Sobel gradient operator as shown in (11) can be used, for example, to determine the HoG in the template.
[0151]
[0152] In the example, in addition to and / or replacing the 3×3 Sobel gradient operator, the 3×2 and 2×3 gradient operators depicted in (12) and (13) can be used for the gradient histogram derivation of the left template and the top template, respectively.
[0153]
[0154]
[0155] Figure 10 shows a set of exemplary samples of filtering using the 3×2 and 2×3 gradient operators depicted in (12) and (13) described herein. In the examples, a 3×3 window can be used to filter circles filled with stripes or dots. Circles filled with dots can be used to obtain and / or derive gradient histograms using 2×3 or 3×2 windows. The location circled in bold may be the center of the convolution. Indicators such as flags can be signaled to specify whether the template is filtered.
[0156] The DIMD methods described herein (e.g., the DIMD merging method shown in Figure 7 and the template filtering for the DIMD methods depicted in (10)) can be combined. For example, the DIMD methods described herein can be combined to improve efficiency (e.g., benefiting from improved coding efficiency of the DIMD merging method without increasing encoder complexity). The combination of DIMD methods described herein can provide a gain in compression performance and / or may also add one or more options regarding how to perform DIMD predictions for a given block in rate-distortion optimization on the encoder side. The combined use of DIMD methods as described herein may further increase encoder complexity (e.g., this is undesirable). Since the DIMD methods (e.g., each of the DIMD methods) provide two options for the encoder, the combination of DIMD methods described herein can provide four options, for example, in the rate-distortion optimization (RDO) process.
[0157] This document may describe one or more exemplary methods for combining the DIMD merging method described herein (e.g., shown in Figure 7) with a template filtering method used for the DIMD merging method described herein (e.g., depicted in (10)). One or more examples may be configured, and the one or more examples may provide one or more variant examples as described herein.
[0158] In the examples, the template filtering and DIMD merging modes described herein can be configured to be mutually exclusive. For example, the template filtering method described herein (e.g., depicted in (10)) may be permitted (e.g., may be permitted only in intra-blocks that are not encoded in the DIMD merging mode).
[0159] For the syntactic parsing and / or decoding process, if the indication used for DIMD merging (e.g., a signaling flag) is true, the indication (e.g., a flag) indicating the use of template filtering described herein (e.g., depicted in (10)) can be skipped in the encoded bitstream (e.g., video data) and / or the indication (e.g., the flag) can be inferred as false by the decoder.
[0160] In addition and / or alternatively, if an indication that uses template filtering (e.g., a signaling flag) can be skipped in the encoded bitstream (e.g., video data) (e.g., no signaling), an indication that uses DIMD merging mode (e.g., a flag) can be skipped, and / or that indication (e.g., a flag) can be inferred as false by the decoder.
[0161] The encoder can be selected and / or obtained (e.g., chosen) among DIMD merging, DIMD (e.g., regular DIMD), and template filtering. The selection of the encoder can be configured (e.g., limited to) 3 options, for example, compared to 4 options in the case of combining methods as described herein (e.g., easily combined).
[0162] In the example, the DIMD merging pattern may employ one or more neighboring blocks that have not adopted the template filtering described herein (e.g., depicted in (10)), for example, to compute the MHog of the current block.
[0163] Normative rules regarding permitted DIMD merge modes may be applied. For example, if at least one adjacent block utilizes a DIMD mode without template filtering and / or a DIMD merge mode, then the DIMD merge mode may be permitted (e.g., only permitted).
[0164] It is possible to configure the system to use DIMD merging and template filtering in combination.
[0165] Signals (e.g., a single flag) can be sent to indicate the combined use of DIMD merging mode and template filtering.
[0166] For example, if the signaling indication (e.g., the signaling flag) is true, then DIMD merge mode and template filtering (e.g., both DIMD merge mode and template filtering) can be used.
[0167] If the signaling indication (e.g., the signaling flag) is false, then DIMD merge mode and template filtering can be skipped (e.g., DIMD merge mode and template filtering can be omitted) (e.g., no DIMD merge mode and template filtering).
[0168] In the example, if the indication (e.g., a flag) is true and there are no adjacent blocks in the DIMD mode, template filtering can be applied, for example, not using the DIMD merging method.
[0169] The system can be configured to use template filtering in DIMD merging mode.
[0170] In the example, when using the DIMD merge mode, template filtering can be applied (e.g., systematically applied) within the current DIMD block.
[0171] For example, the encoder can select and / or obtain (e.g., choose) between DIMD merging, DIMD (e.g., regular DIMD), and template filtering. The selection of the encoder can be configured (e.g., limited to) 3 options, for example, compared to 4 options when DIMD merging and filtering methods are combined (e.g., easily combined as described herein).
[0172] In the example, a specification configuration can be applied. For instance, neighboring blocks that employ template filtering can be used (e.g., only neighboring blocks can be used when calculating the MHoG of the current block).
[0173] In the example, a canonical configuration that allows DIMD merge modes (e.g., a second canonical configuration) can be applied. For example, if at least one adjacent block uses DIMD mode as well as template filtering and / or DIMD merge mode, then DIMD merge mode can be allowed (e.g., only allowed).
[0174] Template filtering can be configured based on the DIMD merging mode.
[0175] One or more template samples from the second row (e.g., the circular sample with filled dots shown in Figure 10) and / or the third row (e.g., away from the current block boundary) that are far from the current block boundary can be subjected to the template filtering process described herein.
[0176] For example, if the current intra-block uses DIMD merging mode, filtering can be performed using template samples from the same row.
[0177] In the example, the second row template sample, which is far from the current block boundary, can be filtered (e.g., only the second row template sample). In the example, one or more adjacent DIMD blocks (e.g., only adjacent DIMD blocks) can be configured (e.g., consider) in the current block where DIMD merging and template filtering are active.
[0178] In the example, the template sample of the third row, which is far from the current block boundary, can be filtered (e.g., only the template sample of the third row). In the example, one or more adjacent DIMD blocks (e.g., only adjacent DIMD blocks) can be configured (e.g., consider) in the current block where DIMD merging and template filtering are active.
[0179] Adaptively select DIMD merging of adjacent blocks based on template filtering.
[0180] For example, in DIMD merge mode, the template filtering used can be the same as that used for most neighboring blocks in the MHog computation of the current block (e.g., it can always be the same).
[0181] As described herein, the encoder can select and / or obtain (e.g., choose) between DIMD merging, DIMD (e.g., regular DIMD), and template filtering. The selection of the encoder can be configured (e.g., limited to) 3 options, for example, compared to 4 options in the case of combining methods as described herein (e.g., easily combining).
[0182] It is permissible (e.g., conditionally permissible) for DIMD merge mode to adaptively select the template filtering state for DIMD merge neighboring blocks based on template filtering usage.
[0183] In DIMD merge mode, one or more neighboring blocks used for MHoG computation (e.g., only neighboring blocks) can use the same filtering mode (e.g., on or off) as the current block.
[0184] The use of template filtering can be signaled (e.g., using indicators and / or flags). If there is at least one adjacent DIMD block with the same template filtering use indicator (such as a flag) as the current DIMD block, a DIMD merging mode can be allowed (e.g., canonically permitted).
[0185] Although the features and elements have been described above in specific combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with other features and elements. Furthermore, the methods described herein can be implemented in a computer program, software, or firmware incorporated into a computer-readable medium for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted via a wired or wireless connection) and computer-readable storage media. Examples of non-transitory computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM discs and digital multifunction disks (DVDs). The processor associated with the software can be used to implement a radio frequency transceiver for a WTRU, UE, terminal, base station, RNC, or any host computer.
Claims
1. An apparatus for video decoding, the apparatus comprising: A processor configured to: determine that a template filtering mode is enabled for a block; And based on determining that the template filtering mode is enabled for the block, the block is decoded based on the decoder-side intra-frame mode derivation (DIMD) merging mode.
2. The apparatus according to claim 1, wherein, The processor is configured to: obtain a template filtering enable indication in video data, wherein the template filtering enable indication is configured to indicate whether the template filtering mode is enabled for the block, and to determine that enabling the template filtering mode for the block is based on the template filtering enable indication.
3. The apparatus according to claim 1 or 2, wherein, Determining whether to enable the template filtering mode includes the processor being further configured to: determine that multiple adjacent blocks use the same filter, wherein the multiple adjacent blocks are associated with computing a merged gradient histogram (MHog) of the blocks.
4. The apparatus according to any one of claims 1 to 3, wherein, The processor is configured to: determine that at least one DIMD neighboring block uses the same filter as the current DIMD block; and allow selection of the DIMD merging mode to decode the block based on the determination that the at least one DIMD neighboring block uses the same filter.
5. A method for video decoding, the method comprising: Determine whether to enable template filtering mode for the block; And based on determining that the template filtering mode is enabled, the block is decoded based on the decoder-side intra-frame mode derivation (DIMD) merging mode.
6. The method according to claim 5, wherein, The method includes: obtaining a template filtering mode enable indication in video data, wherein the template filtering mode enable indication is configured to indicate whether the template filtering mode is enabled for the block, and determining that enabling the template filtering mode for the block is based on the template filtering mode enable indication.
7. The method according to claim 5 or claim 6, wherein, Determining whether to enable the template filtering mode further includes: determining that multiple adjacent blocks use the same filter, wherein the multiple adjacent blocks are associated with the calculation of a merged gradient histogram (MHog) of the blocks.
8. The method according to any one of claims 5 to 7, wherein, The method includes: determining that at least one DIMD neighboring block uses the same filter as the current DIMD block; and allowing the selection of the DIMD merging mode to decode the block based on determining that the at least one DIMD neighboring block uses the same filter.
9. An apparatus for video encoding, the apparatus comprising: A processor configured to: determine a block associated with a template filtering mode; And based on determining that the block is associated with the template filtering mode, the block is encoded using a decoder-side intra-frame mode derivation (DIMD) merging mode.
10. The apparatus according to claim 9, wherein, The processor is configured to include a template filtering mode enable indication in the video data, wherein the template filtering mode enable indication is configured to indicate that the template filtering mode is enabled.
11. The apparatus according to claim 9 or claim 10, wherein, Determining that the block is associated with the template filtering mode includes the processor being further configured to: determine that a plurality of adjacent blocks use the same filtering, wherein the plurality of adjacent blocks are associated with computing a merged gradient histogram (MHog) of the block.
12. The apparatus according to any one of claims 9 to 11, wherein, The processor is configured to: determine that at least one DIMD neighboring block uses the same filter as the current DIMD block; and allow selection of the DIMD merging mode to encode the block based on the determination that the at least one DIMD neighboring block uses the same filter.
13. A method for video encoding, the method comprising: The block is associated with the template filtering mode; And based on determining that the block is associated with the template filtering mode, the block is encoded using a decoder-side intra-frame mode derivation (DIMD) merging mode.
14. The method according to claim 13, wherein, The method includes: including a template filtering mode enable indication in video data, wherein the template filtering mode enable indication is configured to indicate that the template filtering mode is enabled.
15. The method according to claim 13 or claim 14, wherein, Determining whether to enable the template filtering mode further includes: determining that multiple adjacent blocks use the same filter, wherein the multiple adjacent blocks are associated with the calculation of a merged gradient histogram (MHog) of the blocks.
16. The method according to any one of claims 13 to 15, wherein, The method includes: determining that at least one DIMD neighboring block uses the same filter as the current DIMD block; and allowing the selection of the DIMD merging mode to encode the block based on the determination that the at least one DIMD neighboring block uses the same filter.
17. A computer-readable medium comprising instructions for video decoding, the instructions causing one or more processors to perform the method according to any one of claims 5 to 8.
18. Video data comprising information representing blocks encoded based on the method according to any one of claims 5 to 8.
19. A computer-readable medium comprising instructions for video encoding, the instructions causing one or more processors to perform the method according to any one of claims 13 to 16.
20. Video data comprising information representing blocks encoded based on the method according to any one of claims 13 to 16.