Extended angle prediction mode with decoder-side refinement

By identifying and refining intra prediction modes with different resolutions in the video decoding system and using gradient histograms to select the best prediction mode, the encoding efficiency and decoding quality problems of the video decoding system during resolution conversion are solved, and more efficient video signal compression is achieved.

CN120283398APending Publication Date: 2025-07-08INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380072234.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-11
Filing Date
2023-10-11
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

When existing video decoding systems process intra prediction modes of different resolutions, it is difficult to effectively refine the prediction mode to improve encoding efficiency and decoding quality.

Method used

By identifying intra prediction modes associated with different resolutions, the prediction mode on the template is evaluated, the prediction error is calculated, and the refinement intra prediction mode is selected based on the gradient histogram to achieve decoding of the video block.

Benefits of technology

It improves the encoding efficiency and decoding quality of the video decoding system at different resolutions, reduces prediction errors, and improves the compression effect of the video signal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120283398A_ABST
    Figure CN120283398A_ABST
Patent Text Reader

Abstract

Systems, methods, and instrumentalities are disclosed for refining extended angle prediction modes. A device (e.g., a video decoding device) may receive, for a video block, an indication of a first intra prediction mode associated with a first resolution. Based on the first intra prediction mode, the device may identify a second intra prediction mode associated with a second resolution. The device may evaluate a first intra prediction mode and a second intra prediction mode on a template of a video block. The device may select a refined intra prediction mode for a video block in a first intra prediction mode and a second intra prediction mode. The device may decode the video block based on the refined intra prediction mode.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the benefit of European Provisional Patent Application No. 22306532.7, filed on October 11, 2022, the content of which is incorporated herein by reference. Background of the Invention

[0003] Video coding systems can be used to compress digital video signals, for example, to reduce the storage and / or transmission bandwidth required for such signals. Video coding systems can include, for example, block - based, wavelet - based, and / or object - based systems. Summary of the Invention

[0004] Systems, methods, and means for refining an extended - angle prediction mode are disclosed. A device (e.g., a video decoding device) can determine a first intra - prediction mode associated with a first resolution for a video block. For example, the first intra - prediction mode associated with the first resolution can be determined based on an intra - prediction mode indication associated with the video block in video data (e.g., a video bitstream). Based on the first intra - prediction mode, the device can identify a second intra - prediction mode associated with a second resolution. The second resolution can be higher than the first resolution. The device can evaluate the first intra - prediction mode and the second intra - prediction mode on a template of the video block. The device can select a refined intra - prediction mode for the video block from the first intra - prediction mode and the second intra - prediction mode. The device can decode the video block based on the refined intra - prediction mode.

[0005] For example, the device can identify a third intra - prediction mode associated with a second (e.g., higher) resolution. The device can calculate corresponding predictions of the reconstructed template of the video block based on the three intra - prediction modes. The device can select a refined intra - prediction mode for the video block from the intra - prediction at the first resolution and the two intra - prediction modes at the second resolution based on their corresponding prediction errors. The refined intra - prediction mode can be selected based on the determination that the prediction error is the lowest among the prediction errors when using the refined intra - prediction mode to predict the template of the video block.

[0006] The device can generate a gradient histogram on a template of the video block. The histogram can include directions associated with the first intra - prediction mode, the second intra - prediction mode, and the third intra - prediction mode. The device can select a refined intra - prediction mode from the first intra - prediction mode, the second intra - prediction mode, and the third intra - prediction mode based on their associated histogram magnitude values. The refined intra - prediction mode can be selected based on the determination that the histogram magnitude value corresponding to the refined intra - prediction mode is the highest among the histogram magnitude values.

[0007] A device (e.g., a video encoding device) may determine a first intra prediction mode associated with a first resolution for a video block. For example, the first intra prediction mode associated with the first resolution may be determined based on an intra prediction mode indication associated with the video block in video data (e.g., a video bitstream). Based on the first intra prediction mode, the device may identify a second intra prediction mode associated with a second resolution. The second resolution may be higher than the first resolution. The device may evaluate the first intra prediction mode and the second intra prediction mode on a template of the video block. The device may select a refined intra prediction mode for the video block from the first intra prediction mode and the second intra prediction mode. The device may decode the video block based on the refined intra prediction mode.

[0008] For example, the device may identify a third intra prediction mode associated with a second (e.g., higher resolution). The device may calculate corresponding predictions of the reconstructed template of the video block based on three intra prediction modes. The device may select a refined intra prediction mode for the video block from the intra prediction at the first resolution and the two intra prediction modes at the second resolution based on their corresponding prediction errors. The refined intra prediction mode may be selected based on the determination that it has the lowest prediction error among the prediction errors when predicting the template of the video block.

[0009] The device may generate a gradient histogram on a template of the video block. The histogram may include directions associated with the first intra prediction mode, the second intra prediction mode, and the third intra prediction mode. The device may select a refined intra mode from the first intra prediction mode, the second intra prediction mode, and the third intra prediction mode based on their associated histogram magnitude values. The refined intra prediction mode may be selected based on the determination that the histogram magnitude value corresponding to the refined intra prediction mode is the highest among the histogram magnitude values.

[0010] The systems, methods, and means described herein may relate to a decoder. In some examples, the systems, methods, and means described herein may relate to an encoder. In some examples, the systems, methods, and means described herein may relate to a signal (e.g., from an encoder and / or received by a decoder). A computer-readable medium may include instructions for causing one or more processors to perform the methods described herein. A computer program product may include instructions that, when executed by one or more processors, may cause one or more processors to perform the methods described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1A is a system diagram illustrating an example communication system in which one or more of the disclosed embodiments may be implemented.

[0012] Figure 1B illustrates, according to an embodiment, that Figure 1ASystem diagram of an exemplary wireless transmit / receive unit (WTRU) used within the illustrated communication system.

[0013] Figure 1C Illustrates an example radio access network (RAN) and an example core network (CN) that may be used within the Figure 1A illustrated communication system, according to an embodiment.

[0014] Figure 1D Illustrates another example RAN and another example CN that may be used within the Figure 1A illustrated communication system, according to an embodiment.

[0015] Figure 2 Illustrates an example video encoder.

[0016] Figure 3 Illustrates an example video decoder.

[0017] Figure 4 Illustrates an example of a system in which various aspects and examples may be implemented.

[0018] Figure 5 Illustrates examples of prediction modes and prediction directions.

[0019] Figure 6 Shows an example of a template of current luminance and decoded reference samples of the template.

[0020] Figure 7 Shows neighboring reconstructed samples for (e.g.) decoder-side intra mode derivation (DIMD) chrominance modes.

[0021] Figure 8 Illustrates an example block diagram for refinement. Detailed Description

[0022] A more detailed understanding may be obtained from the following description given by way of example in conjunction with the accompanying drawings.

[0023] Figure 1AFIG. is an illustration of an example communication system 100 in which one or more of the disclosed embodiments may be implemented. The communication system 100 may be a multiple access system that provides content such as voice, data, video, messaging, broadcast, etc. to a plurality of wireless users. The communication system 100 may enable the plurality of wireless users to access such content through the sharing of system resources, including wireless bandwidth. For example, the communication system 100 may employ one or more channel access methods such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single carrier FDMA (SC-FDMA), zero-tail unique word DFT spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block filtered OFDM, filter bank multicarrier (FBMC), etc.

[0024] As Figure 1A shown, the communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, a RAN 104 / 113, a CN 106 / 115, a public switched telephone network (PSTN) 108, the Internet 110, and other networks 112, but it should be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a wireless environment. By way of example, the WTRUs 102a, 102b, 102c, 102d (any of which may be referred to as a "station" and / or "STA") may be configured to transmit and / or receive wireless signals and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular telephones, personal digital assistants (PDA), smartphones, laptop computers, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearable devices, head-mounted displays (HMD), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in an industrial and / or automated processing chain environment), consumer electronic devices, devices operating on commercial and / or industrial wireless networks, etc. Any of the WTRUs 102a, 102b, 102c, and 102d may be interchangeably referred to as a UE.

[0025] The communication system 100 may further include base station 114a and / or base station 114b. Each of base stations 114a, 114b may be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communication networks such as CN 106 / 115, Internet 110, and / or other networks 112. As an example, base stations 114a, 114b may be transceiver base stations (BTSs), Node Bs, evolved Node Bs, Home Node Bs, Home evolved Node Bs, gNBs, NR Node Bs, site controllers, access points (APs), wireless routers, etc. Although base stations 114a, 114b are each depicted as a single element, it should be understood that base stations 114a, 114b may include any number of interconnected base stations and / or network elements.

[0026] Base station 114a may be part of RAN 104 / 113, which may also include other base stations and / or network elements (not shown) such as base station controllers (BSCs), radio network controllers (RNCs), relay nodes, etc. Base station 114a and / or base station 114b may be configured to transmit and / or receive wireless signals on one or more carrier frequencies, which may be referred to as cells (not shown). These frequencies may be in licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide coverage of wireless services to a particular geographic area, which may be relatively fixed or may change over time. A cell may be further divided into cell sectors. For example, the cell associated with base station 114a may be divided into three sectors. Thus, in one embodiment, base station 114a may include three transceivers, i.e., one transceiver for each sector of the cell. In an embodiment, base station 114a may employ multiple-input multiple-output (MIMO) technology and may utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.

[0027] Base stations 114a, 114b may communicate with one or more of the WTRUs 102a, 102b, 102c, 102d via air interface 116, which may be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, millimeter wave, infrared (IR), ultraviolet (UV), visible light, etc.). Any suitable radio access technology (RAT) may be used to establish air interface 116.

[0028] More specifically, as noted above, the communication system 100 can be a multi-access system and can employ one or more channel access schemes such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, the base stations 114a and the WTRUs 102a, 102b, 102c in the RAN 104 / 113 can implement a radio technology such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which can use Wideband CDMA (WCDMA) to establish the air interfaces 115 / 116 / 117. WCDMA can include communication protocols such as High-Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA can include High-Speed Downlink (DL) Packet Access (HSDPA) and / or High-Speed UL Packet Access (HSUPA).

[0029] In an embodiment, the base stations 114a and the WTRUs 102a, 102b, 102c can implement a radio technology such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which can use Long-Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-A Pro to establish the air interface 116.

[0030] In an embodiment, the base stations 114a and the WTRUs 102a, 102b, 102c can implement a radio technology such as NR radio access, which can use New Radio (NR) to establish the air interface 116.

[0031] In an embodiment, the base stations 114a and the WTRUs 102a, 102b, 102c can implement multiple radio access technologies. For example, the base stations 114a and the WTRUs 102a, 102b, 102c can implement LTE radio access and NR radio access together using, for example, the Dual Connectivity (DC) principle. Thus, the air interfaces utilized by the WTRUs 102a, 102b, 102c can be characterized by multiple types of radio access technologies and / or transmissions to / from multiple types of base stations (e.g., eNBs and gNBs).

[0032] In other embodiments, the base station 114a and the WTRUs 102a, 102b, 102c may implement radio technologies such as IEEE 802.11 (i.e., Wireless Fidelity (WiFi)), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rates for GSM Evolution (EDGE), GSM EDGE (GERAN), etc.

[0033] Figure 1A The base station 114b in Figure 1A may be, for example, a wireless router, a Home Node B, a Home eNode B, or an access point, and may utilize any suitable RAT to facilitate wireless connectivity in a local area such as a business location, a home, a vehicle, a campus, an industrial facility, an air corridor (e.g., for drones), a road, etc. In one embodiment, the base station 114b and the WTRUs 102c, 102d may implement a radio technology (such as IEEE 802.11) to establish a Wireless Local Area Network (WLAN). In an embodiment, the base station 114b and the WTRUs 102c, 102d may implement a radio technology (such as IEEE 802.15) to establish a Wireless Personal Area Network (WPAN). In yet another embodiment, the base station 114b and the WTRUs 102c, 102d may utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a pico cell or a femto cell. As Figure 1A shown, the base station 114b may have a direct connection to the Internet 110. Thus, the base station 114b may not need to access the Internet 110 via the CN 106 / 115.

[0034] The RAN 104 / 113 may communicate with the CN 106 / 115, which may be any type of network configured to provide voice, data, applications, and / or Voice over Internet Protocol (VoIP) services to one or more of the WTRUs 102a, 102b, 102c, 102d. The data may have different Quality of Service (QoS) requirements, such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, etc. The CN 106 / 115 may provide call control, billing services, location-based services for mobile devices, prepaid calling, Internet connectivity, video distribution, etc., and / or perform advanced security functions such as user authentication. Although not shown in Figure 1Ais shown, but it should be understood that RAN 104 / 113 and / or CN 106 / 115 may communicate directly or indirectly with other RANs that employ the same RAT as RAN 104 / 113 or a different RAT. For example, in addition to being connected to RAN 104 / 113 that may utilize NR radio technology, CN 106 / 115 may also communicate with another RAN (not shown) that employs GSM, UMTS, CDMA 2000, WiMAX, E-UTRA, or WiFi radio technology.

[0035] CN 106 / 115 may also serve as a gateway for WTRUs 102a, 102b, 102c, 102d to access the PSTN 108, the Internet 110, and / or other networks 112. The PSTN 108 may include a circuit-switched telephone network that provides plain old telephone service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices that use common communication protocols such as the Transmission Control Protocol (TCP), the User Datagram Protocol (UDP), and / or the Internet Protocol (IP) in the TCP / IP Internet protocol suite. The network 112 may include wired communication networks and / or wireless communication networks that are owned and / or operated by other service providers. For example, the network 112 may include another CN that is connected to one or more RANs, and the one or more RANs may employ the same RAT as RAN 104 / 113 or a different RAT.

[0036] Some or all of the WTRUs 102a, 102b, 102c, 102d in the communication system 100 may include multimode capabilities (e.g., WTRUs 102a, 102b, 102c, 102d may include multiple transceivers for communicating with different wireless networks over different wireless links). For example, Figure 1A the illustrated WTRU 102c may be configured to communicate with a base station 114a that may employ a cellular-based radio technology and with a base station 114b that may employ IEEE 802 radio technology.

[0037] Figure 1B is a system diagram of an exemplary WTRU 102. As Figure 1B shown, the WTRU 102 may include a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, a non-removable memory 130, a removable memory 132, a power supply 134, a Global Positioning System (GPS) chipset 136, and / or other peripheral devices 138, among other things. It should be understood that the WTRU 102 may include any sub-combination of the foregoing elements while remaining consistent with the embodiments.

[0038] The processor 118 can be a general-purpose processor, a dedicated processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. The processor 118 can perform signal decoding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 102 to operate in a wireless environment. The processor 118 can be coupled to a transceiver 120, which can be coupled to a transmit / receive element 122. Although Figure 1B the processor 118 and the transceiver 120 are depicted as separate components, it should be understood that the processor 118 and the transceiver 120 can be integrated together in an electronic package or chip.

[0039] The transmit / receive element 122 can be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) via an air interface 116. For example, in one embodiment, the transmit / receive element 122 can be an antenna configured to transmit and / or receive RF signals. In an embodiment, the transmit / receive element 122 can be a transmitter / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, the transmit / receive element 122 can be configured to transmit and / or receive both RF signals and optical signals. It should be understood that the transmit / receive element 122 can be configured to transmit and / or receive any combination of wireless signals.

[0040] Although the transmit / receive element 122 is depicted as a single element in Figure 1B the WTRU 102 can include any number of transmit / receive elements 122. More specifically, the WTRU 102 can employ MIMO technology. Thus, in one embodiment, the WTRU 102 can include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals via the air interface 116.

[0041] The transceiver 120 can be configured to modulate signals to be transmitted by the transmit / receive element 122 and demodulate signals received by the transmit / receive element 122. As noted above, the WTRU 102 can have multi-mode capabilities. For example, thus, the transceiver 120 can include multiple transceivers for enabling the WTRU 102 to communicate via multiple RATs (such as NR and IEEE 802.11).

[0042] The processor 118 of the WTRU 102 may be coupled to the speaker / microphone 124, keypad 126, and / or the display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light emitting diode (OLED) display unit) and may receive user input data therefrom. The processor 118 may also output user data to the speaker / microphone 124, keypad 126, and / or the display / touchpad 128. In addition, the processor 118 may access information from any type of suitable memory (such as non-removable memory 130 and / or removable memory 132) and store data in any type of suitable memory. The non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 may include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, etc. In other embodiments, the processor 118 may access information from a memory that is not physically located on the WTRU 102 (such as on a server or a home computer (not shown)) and store data in that memory.

[0043] The processor 118 may receive power from a power supply 134 and may be configured to distribute and / or control power to other components in the WTRU 102. The power supply 134 may be any suitable device for powering the WTRU 102. For example, the power supply 134 may include one or more dry battery packs (e.g., nickel cadmium (NiCd), nickel zinc (NiZn), nickel metal hydride (NiMH), lithium ion (Li-ion), etc.), a solar cell, a fuel cell, etc.

[0044] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to or instead of the information from the GPS chipset 136, the WTRU 102 may receive location information via the air interface 116 from a base station (e.g., base stations 114a, 114b) and / or determine its location based on the timing of signals received from two or more nearby base stations. It should be understood that the WTRU 102 may obtain location information by any suitable location determination method while remaining consistent with the embodiments.

[0045] The processor 118 may also be coupled to other peripheral devices 138, which may include one or more software modules and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, the peripheral devices 138 may include an accelerometer, an electronic compass, a satellite transceiver, a digital camera (for photos and / or videos), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands-free headset, Modules, FM radio units, digital music players, media players, video game player modules, Internet browsers, virtual reality and / or augmented reality (VR / AR) devices, activity trackers, etc. Peripheral device 138 may include one or more sensors, which may be one or more of the following: gyroscopes, accelerometers, Hall effect sensors, magnetometers, orientation sensors, proximity sensors, temperature sensors, time sensors; geographic location sensors; altimeters, light sensors, touch sensors, magnetometers, barometers, gesture sensors, biometric sensors, and / or humidity sensors.

[0046] WTRU 102 may include a full-duplex radio, for which the transmission and reception of some or all signals (e.g., associated with a particular subframe for both UL (e.g., for transmission) and downlink (e.g., for reception)) may be concurrent and / or simultaneous. The full-duplex radio may include an interference management unit 139 for reducing and / or substantially eliminating self-interference via signal processing performed by hardware (e.g., chokes) or via a processor (e.g., a separate processor (not shown) or via processor 118). In an embodiment, WRTU 102 may include a half-duplex radio, for which the transmission and reception of some or all signals (e.g., associated with a particular subframe for UL (e.g., for transmission) or downlink (e.g., for reception)).

[0047] Figure 1C Is a system diagram illustrating RAN 104 and CN 106 according to an embodiment. As noted above, RAN 104 may communicate with WTRU 102a, 102b, 102c via air interface 116 using E-UTRA radio technology. RAN 104 may also communicate with CN106.

[0048] RAN 104 may include evolved Node Bs 160a, 160b, 160c, but it should be understood that RAN 104 may include any number of evolved Node Bs while remaining consistent with the embodiment. Each of evolved Node Bs 160a, 160b, 160c may include one or more transceivers for communicating with WTRU 102a, 102b, 102c via air interface 116. In one embodiment, evolved Node Bs 160a, 160b, 160c may implement MIMO technology. Thus, evolved Node B 160a, for example, may use multiple antennas to transmit wireless signals to and / or receive wireless signals from WTRU 102a.

[0049] Each of evolved Node Bs 160a, 160b, 160c may be associated with a specific cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, scheduling of users in UL and / or DL, etc. As Figure 1C shown, the evolved Node Bs 160a, 160b, 160c may communicate with each other via the X2 interface.

[0050] Figure 1C The CN 106 shown may include a Mobility Management Entity (MME) 162, a Serving Gateway (SGW) 164, and a Packet Data Network (PDN) Gateway (or PGW) 166. Although each of the foregoing elements is depicted as part of the CN 106, it should be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.

[0051] The MME 162 may be connected to each of the evolved Node Bs 162a, 162b, 162c in the RAN 104 via the S1 interface and may act as a control node. For example, the MME 162 may be responsible for authenticating users of the WTRUs 102a, 102b, 102c, bearer activation / deactivation, selecting a specific serving gateway during the initial attachment of the WTRUs 102a, 102b, 102c, etc. The MME 162 may provide control plane functions for exchanges between the RAN 104 and other RANs (not shown) employing other radio technologies such as GSM and / or WCDMA.

[0052] The SGW 164 may be connected to each of the evolved Node Bs 160a, 160b, 160c in the RAN 104 via the S1 interface. The SGW 164 may generally route and forward user data packets to / from the WTRUs 102a, 102b, 102c. The SGW 164 may perform other functions such as anchoring the user plane during handover between evolved Node Bs, triggering paging when DL data is available for the WTRUs 102a, 102b, 102c, managing and storing the context of the WTRUs 102a, 102b, 102c, etc.

[0053] The SGW 164 may be connected to the PGW 166, which may provide the WTRUs 102a, 102b, 102c with access to a packet switched network such as the Internet 110 to facilitate communication between the WTRUs 102a, 102b, 102c and IP-enabled devices.

[0054] CN 106 can facilitate communication with other networks. For example, CN 106 can provide the WTRUs 102a, 102b, 102c with access to a circuit-switched network (such as the PSTN 108) to facilitate communication between the WTRUs 102a, 102b, 102c and traditional landline communication devices. For example, CN 106 can include an IP gateway (such as an IP Multimedia Subsystem (IMS) server) that acts as an interface between CN 106 and the PSTN 108 or can communicate with the IP gateway. Additionally, CN 106 can provide the WTRUs 102a, 102b, 102c with access to other networks 112, which can include other wired and / or wireless networks owned and / or operated by other service providers.

[0055] Although the WTRU is described as a wireless terminal in Figures 1A to 1D it is contemplated that in some representative embodiments, such a terminal may (e.g., temporarily or permanently) use a wired communication interface to a communication network.

[0056] In a representative embodiment, the other network 112 can be a WLAN.

[0057] A WLAN in infrastructure basic service set (BSS) mode can have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP can have access to or an interface to a distribution system (DS) or another type of wired / wireless network that carries traffic to and / or from the BSS. Traffic originating outside the BSS and destined for an STA can reach the STA through the AP and can be delivered to the STA. Traffic originating from an STA and destined for a destination outside the BSS can be delivered to the AP for delivery to the corresponding destination. Traffic between STAs within the BSS can be delivered through the AP. For example, where the source STA can deliver traffic to the AP and the AP can deliver the traffic to the destination STA. Traffic between STAs within the BSS can be considered and / or referred to as peer-to-peer traffic. Peer-to-peer traffic can be delivered (e.g., directly between them) between the source STA and the destination STA using direct link setup (DLS). In some representative embodiments, the DLS can use 802.11e DLS or 802.11z tunnel DLS (TDLS). A WLAN using independent BSS (IBSS) mode may not have an AP, and the STAs within the IBSS or using the IBSS (e.g., all STAs in the IBSS) can communicate directly with each other. The IBSS communication mode may sometimes be referred to as an "ad hoc" communication mode in this document.

[0058] When using the 802.11ac infrastructure operation mode or a similar operation mode, the AP may send beacons on a fixed channel, such as the primary channel. The primary channel may be of a fixed width (e.g., 20 MHz bandwidth) or a width dynamically set via signaling. The primary channel may be the operation channel of the BSS and may be used by the STA to establish a connection with the AP. In some representative embodiments, for example, Carrier Sense Multiple Access / Collision Avoidance (CSMA / CA) may be implemented in the 802.11 system. For CSMA / CA, the STA (e.g., each STA) (including the AP) may listen to the primary channel. If the primary channel is listened to / detected and / or determined to be busy by a particular STA, the particular STA may back off. One STA (e.g., only one station) may transmit at any given time in a given BSS.

[0059] High Throughput (HT) STAs may communicate using 40 MHz wide channels, e.g., by combining the primary 20 MHz channel with an adjacent or non - adjacent 20 MHz channel to form a 40 MHz wide channel.

[0060] Very High Throughput (VHT) STAs may support 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz wide channels. The 40 MHz channel and / or 80 MHz channel may be formed by combining consecutive 20 MHz channels. The 160 MHz channel may be formed by combining eight consecutive 20 MHz channels, or by combining two non - consecutive 80 MHz channels (which may be referred to as an 80 + 80 configuration). For the 80 + 80 configuration, after channel coding, the data may pass through a segment parser that may divide the data into two streams. The Inverse Fast Fourier Transform (IFFT) processing and time - domain processing may be performed separately on each stream. These streams may be mapped to two 80 MHz channels, and the data may be transmitted by the transmitting STA. At the receiver of the receiving STA, the operations for the 80 + 80 configuration described above may be reversed, and the combined data may be delivered to the Medium Access Control (MAC).

[0061] 802.11af and 802.11ah support operation modes below 1 GHz. Compared to those used in 802.11n and 802.11ac, the channel operation bandwidth and carriers are reduced in 802.11af and 802.11ah. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV white space (TVWS) spectrum, and 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. According to a representative embodiment, 802.11ah may support meter type control / machine type communication, such as MTC devices in a macro coverage area. MTC devices may have certain capabilities, such as limited capabilities, including supporting (e.g., only supporting) certain bandwidths and / or limited bandwidths. MTC devices may include a battery with a battery life higher than a threshold (e.g., to maintain a very long battery life).

[0062] A WLAN system that can support multiple channels and channel bandwidths (such as 802.11n, 802.11ac, 802.11af, and 802.11ah) includes a channel that can be designated as a primary channel. The primary channel may have a bandwidth equal to the maximum common operation bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel may be set and / or restricted by an STA (which supports the minimum bandwidth operation mode) from all STAs operating in the BSS. In an example of 802.11ah, for an STA that supports (e.g., only supports) the 1 MHz mode (e.g., an MTC type device), the primary channel may be 1 MHz wide, even if the AP and other STAs in the BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operation modes. Carrier sensing and / or network allocation vector (NAV) settings may depend on the state of the primary channel. If the primary channel is busy, for example, because an STA (only supporting the 1 MHz operation mode) is sending to the AP, the entire available frequency band may be considered busy even if most of the frequency bands remain idle and may be available.

[0063] In the United States, the available frequency band for 802.11ah is 902 MHz to 928 MHz. In Korea, the available frequency band is 917.5 MHz to 923.5 MHz. In Japan, the available frequency band is 916.5 MHz to 927.5 MHz. The total available bandwidth for 802.11ah is 6 MHz to 26 MHz, depending on the country code.

[0064] Figure 1D FIG. is a system diagram illustrating RAN 113 and CN 115 according to an embodiment. As noted above, RAN 113 may employ NR radio technology to communicate with WTRU 102a, 102b, 102c via air interface 116. RAN 113 may also communicate with CN 115.

[0065] RAN 113 may include gNBs 180a, 180b, 180c, but it should be understood that while remaining consistent with the embodiments, RAN 113 may include any number of gNBs. Each of gNBs 180a, 180b, 180c may include one or more transceivers for communicating with WTRUs 102a, 102b, 102c via air interface 116. In one embodiment, gNBs 180a, 180b, 180c may implement MIMO technology. For example, gNBs 180a, 108b may utilize beamforming to transmit signals to and / or receive signals from gNBs 180a, 180b, 180c. Thus, gNB 180a, for example, may use multiple antennas to transmit wireless signals to WTRU 102a and / or receive wireless signals from that WTRU. In an embodiment, gNBs 180a, 180b, 180c may implement carrier aggregation technology. For example, gNB 180a may transmit multiple component carriers to WTRU 102a (not shown). A subset of these component carriers may be on unlicensed spectrum, while the remaining component carriers may be on licensed spectrum. In an embodiment, gNBs 180a, 180b, 180c may implement coordinated multi-point (CoMP) technology. For example, WTRU 102a may receive coordinated transmissions from gNB 180a and gNB 180b (and / or gNB 180c).

[0066] WTRUs 102a, 102b, 102c may communicate with gNBs 180a, 180b, 180c using transmissions associated with parameter sets that are capable of being extended. For example, the OFDM symbol interval and / or the OFDM subcarrier interval may vary for different transmissions, different cells, and / or different parts of the radio transmission spectrum. WTRUs 102a, 102b, 102c may communicate with gNBs 180a, 180b, 180c using subframes or transmission time intervals (TTIs) of various or extendable lengths (e.g., containing different numbers of OFDM symbols and / or having an absolute time length that continuously varies).

[0067] gNBs 180a, 180b, 180c can be configured to communicate with WTRUs 102a, 102b, 102c in a stand-alone configuration and / or a non-stand-alone configuration. In the stand-alone configuration, WTRUs 102a, 102b, 102c can communicate with gNBs 180a, 180b, 180c without accessing other RANs (e.g., such as evolved Node Bs 160a, 160b, 160c). In the stand-alone configuration, WTRUs 102a, 102b, 102c can use one or more of gNBs 180a, 180b, 180c as a mobility anchor. In the stand-alone configuration, WTRUs 102a, 102b, 102c can communicate with gNBs 180a, 180b, 180c using signals in an unlicensed band. In the non-stand-alone configuration, WTRUs 102a, 102b, 102c can communicate / connect with gNBs 180a, 180b, 180c while also communicating / connecting with another RAN (such as evolved Node Bs 160a, 160b, 160c). For example, WTRUs 102a, 102b, 102c can implement the DC principle to communicate with one or more of gNBs 180a, 180b, 180c and one or more of evolved Node Bs 160a, 160b, 160c substantially simultaneously. In the non-stand-alone configuration, evolved Node Bs 160a, 160b, 160c can act as the mobility anchor for WTRUs 102a, 102b, 102c, and gNBs 180a, 180b, 180c can provide additional coverage and / or throughput for serving WTRUs 102a, 102b, 102c.

[0068] Each of gNBs 180a, 180b, 180c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, scheduling of users in UL and / or DL, support for network slicing, dual connectivity, interworking between NR and E-UTRA, routing of user plane data towards user plane functions (UPFs) 184a, 184b, routing of control plane information towards access and mobility management functions (AMFs) 182a, 182b, etc. As Figure 1D shown, gNBs 180a, 180b, 180c can communicate with each other via the Xn interface.

[0069] Figure 1DThe illustrated CN 115 may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and possibly Data Networks (DN) 185a, 185b. Although each of the foregoing elements is depicted as part of CN 115, it should be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.

[0070] The AMF 182a, 182b may be connected to one or more of the gNBs 180a, 180b, 180c in the RAN 113 via the N2 interface and may act as a control node. For example, the AMF 182a, 182b may be responsible for authenticating users of the WTRUs 102a, 102b, 102c, support for network slicing (e.g., handling of different PDU sessions with different requirements), selection of a particular SMF 183a, 183b, management of the registration area, termination of NAS signaling, mobility management, etc. The AMF 182a, 182b may use network slicing in order to customize CN support for the WTRUs 102a, 102b, 102c based on the type of service utilized by the WTRUs 102a, 102b, 102c. For example, different network slices may be established for different use cases such as services that rely on ultra-reliable low-latency (URLLC) access, services that rely on enhanced mobile broadband (eMBB) access, services for machine type communication (MTC) access, etc. The AMF 162 may provide control plane functions for exchange between the RAN 113 and other RANs (not shown) that employ other radio technologies such as LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies such as WiFi.

[0071] The SMF 183a, 183b may be connected to the AMF 182a, 182b in the CN 115 via the N11 interface. The SMF 183a, 183b may also be connected to the UPF 184a, 184b in the CN 115 via the N4 interface. The SMF 183a, 183b may select and control the UPF 184a, 184b and configure the traffic routing through the UPF 184a, 184b. The SMF 183a, 183b may perform other functions such as management and allocation of UE IP addresses, management of PDU sessions, control of policy enforcement and QoS, provision of downlink data notification, etc. The PDU session type may be IP-based, non-IP-based, Ethernet-based, etc.

[0072] UPF 184a and 184b can be connected to one or more of gNBs 180a, 180b, and 180c in RAN 113 via the N3 interface. The one or more gNBs can provide access to a packet-switched network (such as the Internet 110) to WTRU 102a, 102b, and 102c to facilitate communication between WTRU 102a, 102b, and 102c and IP-enabled devices. UPF 184 and 184b can perform other functions such as routing and forwarding packets, implementing user plane policies, supporting multi-homed PDU sessions, handling user plane QoS, buffering downlink packets, providing mobility anchoring, etc.

[0073] CN 115 can facilitate communication with other networks. For example, CN 115 can include an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between CN 115 and the PSTN 108 or can communicate with the IP gateway. In addition, CN 115 can provide access to other networks 112 to WTRU 102a, 102b, and 102c. The other networks can include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, WTRU 102a, 102b, and 102c can be connected to DN 185a and 185b via UPF 184a and 184b through the N3 interface to UPF 184a and 184b and the N6 interface between UPF 184a and 184b and the local data network (DN) 185a and 185b.

[0074] In view of Figures 1A to 1D and Figures 1A to 1D In view of the corresponding descriptions of, one or more or all of the functions described herein for one or more of the following may be performed by one or more emulation devices (not shown): WTRU 102a - 102d, base stations 114a - 114b, evolved Node Bs 160a - 160c, MME 162, SGW 164, PGW 166, gNBs 180a - 180c, AMF 182a - 182b, UPF 184a - 184b, SMF 183a - 183b, DN 185a - 185b, and / or any other device described herein. The emulation device(s) can be one or more devices configured to mimic one or more or all of the functions described herein. For example, the emulation device(s) can be used to test other devices and / or simulate network and / or WTRU functions.

[0075] The simulation device can be designed to implement one or more tests of other devices in a laboratory environment and / or in an operator network environment. For example, one or more simulation devices can perform one or more functions or all functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices within the communication network. One or more simulation devices can perform one or more functions or all functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. The simulation device can be directly coupled to another device for testing purposes and / or can perform tests using over-the-air wireless communication.

[0076] One or more simulation devices can perform one or more (including all) functions without being implemented / deployed as part of a wired and / or wireless communication network. For example, the simulation device can be used in a test laboratory and / or in a test scenario in a non-deployed (e.g., test) wired and / or wireless communication network to implement tests of one or more components. One or more simulation devices can be test equipment. Direct RF coupling and / or wireless communication via an RF circuit system (e.g., which can include one or more antennas) can be used by the simulation device to transmit and / or receive data.

[0077] This application describes multiple aspects, including tools, features, examples, models, methods, etc. Many of these aspects are described in detail and are typically described in a way that may sound restrictive, at least to illustrate individual features. However, this is for the purpose of clarity of description and does not limit the application or scope of those aspects. In fact, all different aspects can be combined and interchanged to provide further aspects. Additionally, these aspects can also be combined and interchanged with aspects described in earlier applications.

[0078] The aspects described and contemplated in this application can be implemented in many different forms. Examples are described herein Figures 5 to 8 but other examples can be contemplated. Figures 5 to 8 The discussion does not limit the breadth of the implementation. At least one of the aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to generating a bitstream, storing a bitstream, and / or transmitting the generated or encoded bitstream. These and other aspects can be implemented as methods, devices, computer-readable storage media storing instructions for encoding or decoding video data according to any of the described methods, and / or computer-readable storage media storing a bitstream generated according to any of the described methods. As used herein, the bitstream may or may not be sent.

[0079] In this application, the terms "reconstructed" and "decoded" may be used interchangeably, the terms "pixel" and "sample" may be used interchangeably, and the terms "image", "picture", and "frame" may be used interchangeably.

[0080] This document describes various methods, and each method includes one or more steps or actions for implementing the described method. Unless the correct operation of the method requires steps or actions in a specific order, the order and / or use of specific steps and / or actions may be modified or combined. Additionally, terms such as "first", "second", etc. may be used in various examples to modify elements, components, steps, operations, etc., such as for example "first decoding" and "second decoding". Unless specifically required, the use of these terms does not imply an ordering of the modified operations. Thus, in this example, the first decoding does not need to be performed before the second decoding, and may occur, for example, before, during, or in a time period overlapping with the second decoding.

[0081] The various methods and other aspects described in this application can be used to modify modules of a video encoder 200 and a decoder 300 as shown in Figure 2 and Figure 3 e.g., the decoding module. Additionally, the subject matter disclosed herein can be applied to (for example) any type, format, or version of video coding, whether described in a standard or recommendation, whether pre - existing or future - developed, and to extensions of any such standard and recommendation. Unless otherwise stated or technically precluded, the aspects described in this application can be used alone or in combination.

[0082] Various numerical values are used in the examples described in this application, such as 0, 1, 2, 3, 4, 6, 7, 8, 11, 16, 18, 26, 33, 45, 50, 64, 65, 66, 67, 80, 129, 131, 135, 1456, etc. These and other specific values are for the purpose of describing examples, and the described aspects are not limited to these specific values.

[0083] Figure 2 An example of a video encoder 200 (e.g., a block - based hybrid video encoder) is illustrated. Variations of the example encoder 200 are envisioned, but for clarity, encoder 200 is described below without describing all the expected variations.

[0084] Before being encoded, the video sequence may undergo pre-encoding processing (201), for example, by performing one or more of the following: applying a color transformation to the input color picture (e.g., converting from RGB 4:4:4 to YCbCr 4:2:0) or performing remapping of the input picture components, e.g., to obtain a transmission distribution that is resilient to compression (e.g., more resilient), such as using histogram equalization of one of the color components. Metadata may be associated with the preprocessing and may be appended to the bitstream.

[0085] In the encoder 200, pictures may be encoded as described below (e.g., may be encoded by encoder elements). The picture to be encoded may be partitioned (202) and processed in units of, e.g., CUs (coding units). Each unit may be encoded using, e.g., an intra mode or an inter mode. When a unit is encoded in the intra mode, intra prediction (260) may be performed. In the inter mode, motion estimation (275) and motion compensation (270) may be performed. The encoder may determine (205) whether one of the intra mode or the inter mode will be used to encode the CU, and the intra / inter decision may be indicated, e.g., by a prediction mode indicator (e.g., a prediction mode flag) (e.g., by the encoder). The prediction residual may be calculated, e.g., by subtracting (210) the prediction block from the original image block. In an intra frame, the CU may be intra predicted (e.g., in an intra (I) frame), while in an inter frame, the CU may be intra predicted or inter predicted.

[0086] The prediction residual may be transformed at 225 and quantized at 230. One or more of the quantized transform coefficients, motion vectors, or other syntax elements (e.g., picture partition information) may be entropy encoded at 245 to output a bitstream. The encoder may apply quantization directly to (e.g., and skip transformation of) the non-transformed residual for transmission. The transformation and quantization may be bypassed (e.g., by the encoder). For example, the residual may be encoded (e.g., directly encoded without applying the transformation or quantization process).

[0087] The encoded block may be decoded (e.g., by the encoder) to provide a reference (e.g., a reference for further prediction). The quantized transform coefficients may be dequantized at 240 and inverse transformed at 250 (e.g., inverse transformed to decode the prediction residual). The decoded prediction residual and the prediction block may be combined at 255, and the image block may be reconstructed. An in-loop filter at 265 may be applied to the reconstructed picture to perform, e.g., deblocking / SAO (sample adaptive offset) / ALF (adaptive loop filter) filtering (e.g., to reduce encoding artifacts). At 280, the filtered image may be stored in the reference picture buffer.

[0088] Figure 3A block diagram illustrating an example video decoder 300 is shown. In decoder 300, a bitstream can be decoded as described herein (e.g., by decoder elements). Video decoder 300 can perform a decoding pass that is inverse to the encoding pass as Figure 2 described. As described herein, encoder 200 can perform video decoding as part of encoding video data.

[0089] Specifically, the input to the video decoder can include video data (e.g., a video bitstream), which can be generated by video encoder 200. The bitstream can be entropy decoded at 330 (e.g., to obtain one or more transform coefficients, prediction modes, motion vectors, or other encoded information). Picture partitioning information can indicate how a picture is partitioned. At 355, the decoder can partition the picture according to the decoded picture partitioning information. The transform coefficients can be dequantized at 340 and inverse transformed at 350 to decode the prediction residuals. A prediction block can be obtained at 370 from intra prediction at 360 or motion compensated prediction at 375 (e.g., inter prediction). The decoded prediction residuals and the prediction block can be combined at 355, and an image block can be reconstructed. At 365, an in-loop filter can be applied to the reconstructed image. At 380, the filtered image can be stored in a reference picture buffer. The content of reference picture buffer 380 on the decoder side can be the same as the content of reference picture buffer 280 on the encoder 200 side (e.g., for a picture).

[0090] The decoded picture can further undergo post-decoding processing at 385, such as one or more of an inverse color transformation (e.g., a conversion from YCbCr 4:2:0 to RGB 4:4:4) or an inverse remapping (e.g., performing the inverse of the remapping technique performed in the pre-encoding processing at 201). The post-decoding processing can use metadata derived in the pre-encoding processing and can be signaled in the video data (e.g., the bitstream). In an example, the decoded image (e.g., after applying the in-loop filter 365 and / or after the post-decoding processing 385, if post-decoding processing is used) can be sent to a display device for presentation to a user.

[0091] Figure 4FIG. is an example of a diagram showing a system in which various aspects and examples described herein can be implemented. System 400 can be embodied as a device including various components described below and is configured to perform one or more aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, and servers. The elements of System 400 can be embodied individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one example, the processing and encoder / decoder elements of System 400 are distributed across multiple ICs and / or discrete components. In various examples, System 400 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various examples, System 400 is configured to implement one or more aspects described in this document.

[0092] System 400 includes at least one processor 410, which is configured to execute instructions loaded therein for implementing, for example, various aspects described in this document. Processor 410 can include embedded memory, input / output interfaces, and various other circuits known in the art. System 400 includes at least one memory 420 (e.g., a volatile memory device and / or a non-volatile memory device). System 400 includes a storage device 440, which can include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 440 can include internal storage devices, attached storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices.

[0093] System 400 includes an encoder / decoder module 430, which is configured to, for example, process data to provide encoded video or decoded video, and encoder / decoder module 430 can include its own processor and memory. Encoder / decoder module 430 represents a module that can be included in a device to perform encoding and / or decoding functions. As is known, a device can include one or both of an encoding and a decoding module. Additionally, encoder / decoder module 430 can be implemented as a separate element of System 400 or can be incorporated within processor 410 as a combination of hardware and software known to those skilled in the art.

[0094] The program code to be loaded onto the processor 410 or the encoder / decoder 430 to perform the various aspects described in this document can be stored in the storage device 440 and subsequently loaded onto the memory 420 for execution by the processor 410. According to various examples, one or more of the processor 410, the memory 420, the storage device 440, and the encoder / decoder module 430 can store one or more of the various items during the execution of the processes described in this document. Such stored items can include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operation logic.

[0095] In some examples, the memory internal to the processor 410 and / or the encoder / decoder module 430 is used to store instructions and provide working memory for the processing required during encoding or decoding. However, in other examples, memory external to the processing device (e.g., the processing device can be the processor 410 or the encoder / decoder module 430) is used for one or more of these functions. The external memory can be the memory 420 and / or the storage device 440, such as dynamic volatile memory and / or non-volatile flash memory. In several examples, the external non-volatile flash memory is used to store, for example, the operating system of a television set. In at least one example, fast external dynamic volatile memory such as RAM is used as the working memory for video encoding and decoding operations.

[0096] As shown in block 445, input to the elements of the system 400 can be provided through various input devices. Such input devices include, but are not limited to, (i) a radio frequency (RF) section that receives, for example, RF signals transmitted over the air by a broadcaster, (ii) component (COMP) input terminals (or a set of COMP input terminals), (iii) universal serial bus (USB) input terminals, and / or (iv) high-definition multimedia interface (HDMI) input terminals. Figure 4 Other examples not shown include composite video.

[0097] In various examples, the input device of block 445 has corresponding input processing elements associated therewith as known in the art. For example, the RF section may be associated with elements adapted to: (i) select a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band), (ii) down-convert the selected signal, (iii) again band-limit to a narrower band to select a signal band that may be referred to as a channel in some examples, (iv) demodulate the down-converted and band-limited signal, (v) perform error correction, and / or (vi) de-multiplex to select a desired data packet stream. The RF section of various examples includes one or more elements for performing these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, filters, a down-converter, a demodulator, an error corrector, and a de-multiplexer. The RF section may include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or down-converting to baseband. In one example of a set-top box, the RF section and its associated input processing elements receive an RF signal transmitted through a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and filtering again to a desired frequency band. Various examples re-order the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as inserting an amplifier and an analog-to-digital converter. In various examples, the RF section includes an antenna.

[0098] The USB and / or HDMI terminals may include corresponding interface processors for connecting system 400 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or within processor 410 as needed. Similarly, aspects of USB or HDMI interface processing may be implemented within a separate interface IC or within processor 410 as needed. The demodulated, error-corrected, and de-multiplexed stream is provided to various processing elements, including, for example, processor 410 and encoder / decoder 430, which operate in conjunction with memory and storage elements to process the data stream as needed for presentation on an output device.

[0099] The various elements of system 400 may be disposed within an integrated housing in which the various elements may be interconnected and data may be transmitted between them using a suitable connection arrangement 425 (e.g., internal buses known in the art, including inter-integrated circuit (I2C) buses, wiring, and printed circuit boards).

[0100] System 400 includes a communication interface 450 that enables communication with other devices via a communication channel 460. The communication interface 450 can include, but is not limited to, a transceiver configured to send and receive data over the communication channel 460. The communication interface 450 can include, but is not limited to, a modem or a network card, and the communication channel 460 can be implemented, for example, within a wired and / or wireless medium.

[0101] In various examples, a wireless network (such as a Wi-Fi network, e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers)) is used to stream or otherwise provide data to system 400. Wi-Fi signals for these examples are received via a communication channel 460 and a communication interface 450 adapted for Wi-Fi communication. The communication channel 460 for these examples is typically connected to an access point or router that provides access to an external network (including the Internet) to allow streaming applications and other over-the-top communications. Other examples use a set-top box to provide streamed data to system 400, and the set-top box delivers data via an HDMI connection of input block 445. Other examples use an RF connection of input block 445 to provide streamed data to system 400. As indicated above, various examples provide data in a non-streaming manner. Additionally, various examples use wireless networks other than Wi-Fi, such as cellular networks or networks.

[0102] System 400 can provide output signals to various output devices, including a display 475, speakers 485, and other peripheral devices 495. The display 475 for various examples includes, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 475 can be used for a television, a tablet, a laptop computer, a mobile phone (cell phone), or other devices. The display 475 can also be integrated with other components (e.g., as in a smart phone) or be separate (e.g., an external monitor for a laptop computer). In various examples, the other peripheral devices 495 include one or more of a standalone digital video disc (or digital versatile disc) (DVD, both terms), a disc player, a stereo system, and / or a lighting system. Various examples use one or more of the peripheral devices 495 that perform functions based on the output of system 400. For example, a disc player performs the function of playing the output of system 400.

[0103] In various examples, signaling using a communication protocol such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable device-to-device control with or without user intervention conveys control signals between system 400 and display 475, speaker 485, or other peripheral device 495. The output devices can be communicatively coupled to system 400 via dedicated connections through respective interfaces 470, 480, and 490. Alternatively, the output devices can be connected to system 400 via communication interface 450 using communication channel 460. Display 475 and speaker 485 can be integrated in a single unit with other components of system 400 in an electronic device (e.g., a television). In various examples, display interface 470 includes a display driver, such as, for example, a timing controller (T_CON) chip.

[0104] For example, if the RF portion of input 445 is part of a separate set-top box, display 475 and speaker 485 can optionally be separated from one or more other components. In various examples where display 475 and speaker 485 are external components, output signals can be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.

[0105] Examples can be executed by computer software implemented by processor 410 or by hardware or by a combination of hardware and software. As a non-limiting example, an example can be implemented by one or more integrated circuits. Memory 420 can be of any type suitable for the technical environment and can be implemented using any appropriate data storage technology, as non-limiting examples, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. Processor 410 can be of any type suitable for the technical environment and can encompass, as non-limiting examples, one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.

[0106] Various implementations involve decoding. As used in this application, "decoding" can cover, for example, all or part of the processing performed on a received encoded sequence to produce a final output suitable for display. In various examples, such a process includes one or more of the processes typically performed by a decoder, such as, entropy decoding, inverse quantization, inverse transform, and differential decoding. In various examples, such a process also or alternatively includes processes performed by the decoders of the various embodiments described in this application, such as, receiving an indication of a first intra prediction mode associated with a first resolution for a video block; based on the first intra prediction mode, identifying a second intra prediction mode associated with a second resolution; evaluating the first intra prediction mode and the second intra prediction mode on a template of the video block; selecting a refined intra prediction mode for the video block from the first intra prediction mode and the second intra prediction mode; and decoding the video block based on the refined intra prediction mode.

[0107] As a further example, in one example, "decoding" refers only to entropy decoding, in another example, "decoding" refers only to differential decoding, and in another example, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally to a broader decoding process will be clear based on the context of the specific description and is considered well understood by those skilled in the art.

[0108] Various embodiments involve encoding. In a manner similar to the above discussion regarding "decoding", "encoding" as used in this application can cover, for example, all or part of the processing performed on an input video sequence to produce an encoded bitstream. In various examples, such a process includes one or more of the processes typically performed by an encoder, such as, partitioning, differential encoding, transform, quantization, and entropy encoding. In various examples, such a process also or alternatively includes processes performed by the encoders of the various embodiments described in this application, such as, determining a first intra prediction mode associated with a first resolution for a video block; based on the first intra prediction mode, identifying a second intra prediction mode associated with a second resolution; evaluating the first intra prediction mode and the second intra prediction mode on a template of the video block; selecting a refined intra prediction mode for the video block from the first intra prediction mode and the second intra prediction mode; including an indication of the first intra prediction mode in the video data; and encoding the video block based on the refined intra prediction mode.

[0109] As a further example, in one example, "encoding" refers only to entropy encoding, in another example, "encoding" refers only to differential encoding, and in another example, "encoding" refers to a combination of differential encoding and entropy encoding. Whether the phrase "encoding process" is intended to specifically refer to a subset of operations or generally to a broader encoding process will be clear based on the context of the specific description and is considered well understood by those skilled in the art.

[0110] Note that, as used herein, grammatical elements such as, for example, the coding grammar regarding templates, coding blocks, reference samples, refinement values, resolution factors, fractions, modes (e.g., prediction mode, intra mode), number of partitioning modes, number of intra prediction mode candidates, number of partitioning mode candidates, block size, slice type, etc. are descriptive terms. Accordingly, they do not exclude the use of other grammatical element names.

[0111] When the accompanying drawings are presented as flowcharts, it should be understood that they also provide block diagrams of the corresponding devices. Similarly, when the accompanying drawings are presented as block diagrams, it should be understood that they also provide flowcharts of the corresponding methods / processes.

[0112] The embodiments and aspects described herein can be implemented in, for example, a method or process, a device, a software program, a data stream, or a signal. Even if discussed only in the context of a single implementation form (e.g., only as a method), the implementation of the features discussed can be in other forms (e.g., a device or a program). The device can be implemented in, for example, appropriate hardware, software, and firmware. The method can be implemented in, for example, a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes communication devices such as, for example, a computer, a cellular phone, a portable / personal digital assistant (“PDA”), and other devices that facilitate information communication between end users.

[0113] References to “an example” or “example” or “an implementation” or “implementation” and other variations thereof mean that the specific features, structures, characteristics, etc. described in connection with that example are included in at least one example. Thus, the appearances of the phrases “in an example” or “in an example” or “in an implementation” or “in an implementation” and any other variations throughout this application do not necessarily refer to the same example.

[0114] In addition, this application may refer to “determining” various information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from a memory. Obtaining may include receiving, retrieving, constructing, generating, and / or determining.

[0115] In addition, this application may refer to “accessing” various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0116] In addition, the present application may refer to "receiving" each piece of information. Like "access", receiving is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from a memory). In addition, during operations such as, for example, storing information, processing information, sending information, moving information, copying information, erasing information, computing information, determining information, predicting information, or estimating information, "receiving" is generally involved in one way or another.

[0117] It should be understood that, for example, in the case of "A / B", "A and / or B", and "at least one of A and B", the use of any one of the following, namely " / ", "and / or", and "at least one of", is intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As another example, in the case of "A, B, and / or C" and "at least one of A, B, and C", such wording is intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or only the first and second-listed options (A and B), or only the first and third-listed options (A and C), or only the second and third-listed options (B and C), or the selection of all three options (A and B and C). As will be apparent to those of ordinary skill in the art and related fields, this can be extended to as many items as are listed.

[0118] In addition, as used herein, the word "signal" particularly refers to indicating something to a corresponding decoder. An encoder signal may include, for example, an intra prediction mode candidate, the number of segmentation mode candidates, a block size, a slice type, etc. In this way, in an example, the same parameters are used at both the encoder side and the decoder side. Thus, for example, the encoder may send (explicit signaling) a specific parameter to the decoder such that the decoder may use the same specific parameter. Conversely, if the decoder already has a specific parameter and other parameters, signaling may be used without sending (implicit signaling) to simply allow the decoder to know and select the specific parameter. By avoiding the transmission of any actual functionality, bit savings are achieved in various examples. It should be understood that signaling may be done in various ways. By way of example, in various examples, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder. Although the foregoing relates to the verb form of the word "signal", the word "signal" may also be used as a noun herein.

[0119] It will be apparent to those of ordinary skill in the art that the implementation can generate various signals that are formatted to carry information that can be stored or transmitted, for example. The information can include, for example, instructions for performing a method or data generated by one of the described embodiments. By way of example, a signal can be formatted to carry the bitstream of the described example. Such a signal can be formatted as, for example, an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. As is known, signals can be transmitted over a variety of different wired or wireless links. The signal can be stored on a processor-readable medium or accessed or received from a processor-readable medium.

[0120] Numerous examples are described herein. The features of the examples can be provided individually or in any combination across various claim categories and types. Additionally, an example can include one or more of the features, devices, or aspects described herein across various claim categories and types, individually or in any combination. For example, the features described herein can be implemented in a bitstream or signal that includes the information generated as described herein. The information can allow a decoder to decode the bitstream, encoder, bitstream, and / or decoder according to any of the described embodiments. For example, the features described herein can be implemented by creating and / or sending and / or receiving and / or decoding a bitstream or signal. For example, the features described herein can be implemented with a method, process, device, medium storing instructions, medium storing data, or signal. For example, the features described herein can be implemented by a TV, set-top box, cellular phone, tablet computer, or other electronic device that performs decoding. The TV, set-top box, cellular phone, tablet computer, or other electronic device can display (e.g., using a monitor, screen, or other type of display) the resulting image (e.g., an image reconstructed from the residuals of a video bitstream). The TV, set-top box, cell phone, tablet computer, or other electronic device can receive a signal that includes an encoded image and perform decoding.

[0121] Intra-sample prediction can include predicting the pixels of a target CU based on a set of reference samples. The prediction modes can include planar and DC prediction modes, which can be used to predict smooth and gradually changing regions. Angular prediction modes (e.g., angles from 45 degrees to -135 degrees in a clockwise direction) can be used to capture directional structures. For square blocks, a directional prediction mode (e.g., 33 directional modes for square blocks) can be used, which can be indexed (e.g., indexed from 2 to 34). The prediction mode can correspond to the prediction direction illustrated in the upper left side as Figure 5 The angular prediction mode can correspond to an angular direction (e.g., 65 angular prediction modes can correspond to 33 angular directions), and the angular direction (e.g., 32 additional angular directions) can correspond to the intermediate direction between adjacent pairs, asFigure 5 as described in the upper right side of

[0122] Figure 5 Example in-frame prediction directions are illustrated in the upper left of . The numbers may represent prediction mode indices associated with the corresponding directions. Modes (e.g., 2 to 17) may indicate horizontal prediction (H - 26 to H + 32), and modes 18 to 34 may indicate vertical prediction (V - 32 to V + 32). Figure 5 In-frame prediction of square blocks is illustrated in the upper right of . Modes less than 34 may indicate horizontal prediction. Modes greater than 34 may indicate vertical prediction. Figure 5 Available in-frame prediction directions are illustrated at the bottom of . The dashed lines may indicate wide-angle in-frame prediction mode (WAIP). Figure 5 Indices -1 to -14 illustrated in may be remapped to 1 to -12 (e.g., such that the angular mode indices are consecutive). In some examples, mode -15 (e.g., remapped to -13) and 81 may or may not be present in Figure 5 since the block size (e.g., no permitted block size) may or may not use mode -15 (e.g., remapped to -13) and 81. Mode -15 (e.g., remapped to -13) and 81 may be handled by the reference code.

[0123] Template-based in-frame mode derivation (TIMD) may be performed to derive the prediction mode for the decoding block. In-frame prediction mode derivation via TIMD may be applied to both the encoder and decoder sides (e.g., in the same way for both the encoder and decoder sides) for luminance, e.g., CB 103, as shown in Figure 6 (a) of . The in-frame prediction modes in the MPM list of the luminance CB (e.g., supplemented with the default mode) may be used to calculate the prediction of templates (100 and 101) of the luminance CB from the decoded reference samples of the template (102). The sum of absolute transform differences (SATD) between the prediction and the template of the luminance CB may be calculated. The (e.g., two) in-frame prediction modes with the minimum (e.g., smallest) SATD may be selected as the TIMD mode. The set of directional in-frame prediction modes (e.g., for TIMD) may be extended (e.g., from 65 to 129), e.g., by inserting directions between the solid lines and adjacent dashed arrows in Figure 5 The set of possible in-frame prediction modes derived via TIMD may aggregate modes (e.g., 131 modes). After retaining the in-frame prediction modes (e.g., two (2) in-frame prediction modes from the first-pass test involving the MPM list supplemented with the default mode), for non-planar or DC modes, TIMD may test its closest (e.g., two (2) closest) extended directional in-frame prediction modes according to the prediction SATD.

[0124] Figure 6Shows an example template of the current luminance CB and the decoded reference samples of the template used in TIMD. In Figure 6 In (a) of, the template of the luminance CB does not exceed the boundary of the current frame. The current W×H luminance CB 103 can be surrounded by its fully available template, which consists of the w t ×H part at its left side 100 and the W×h t part at its upper side 101. During the TIMD derivation step, the tested intra prediction mode can be predicted for the template of the current luminance CB from the set of 1 + 2w t + 2W + 2h t + 2H decoded reference samples at 102 of the template. If W ≤ 8, w t can be equal to two (2); otherwise w t can be equal to 4. If H ≤ 8, h t can be equal to two (2); otherwise h t can be equal to 4.

[0125] Figure 6 (b) of and Figure 6 (c) of show examples where at least one (e.g., one) part of the template of the luminance CB exceeds the boundary of the current frame. In Figure 6 (b) of, the current W×H luminance CB 103 can be surrounded by its template, and the W×h t part at its upper side 101 is available. During the TIMD derivation step, the tested intra prediction mode can be predicted for the template of the current luminance CB from the set of 1 + 2W + 2h t + 2H decoded reference samples of the template. In Figure 6 (c) of, the current W×H luminance CB 103 can be surrounded by its template, and only the w t ×H part at its left side 100 is available. During the TIMD derivation step, the tested intra prediction mode can be predicted for the template of the current luminance CB from the set of 1 + 2w t + 2W + 2H decoded reference samples of the template.

[0126] The current luminance CB can be predicted via TIMD, for example, by fusing (e.g., two) predictions of the luminance CB calculated based on (e.g., two) TIMD modes generated by (e.g., two) test passes with weights (e.g., after applying PDPC). The weights used can depend on the predicted SATD of (e.g., two) TIMD modes.

[0127] An executable decoder-side intra mode derivation (DIMD) can be performed to derive an intra prediction mode for a coding block. For example, two intra modes can be derived from the reconstructed neighboring samples. Two predictors can be combined with a planar mode predictor having weights derived from gradients. The division operation in the weight derivation can be performed using the same lookup table (LUT)-based integerization scheme used by the cross-component linear model (CCLM). For example, the division operation in the orientation calculation

[0128] Orient = G y / G x

[0129] can be calculated by the following LUT-based scheme:

[0130] x = Floor(Log2(Gx))

[0131] normDiff = ((Gx << 4) >> x) & 15

[0132] x += (3 + (normDiff!= 0)? 1 : 0)

[0133] Orient = (Gy * (DivSigTable[normDiff] | 8) + (1 << (x - 1))) >> x,

[0134] where DivSigTable

[16] = {0, 7, 6, 5, 5, 4, 4, 3, 3, 2, 2, 1, 1, 1, 1, 0}.

[0135] The derived intra mode can be included in the main list of the intra most probable mode (MPM) list, and the DIMD technique can be performed before constructing the MPM list. The main derived intra mode of the DIMD block can be stored with the block and can be used for the MPM list construction of neighboring blocks.

[0136] Figure 7 The neighboring reconstructed samples for the DIMD chroma mode are illustrated. The DIMD chroma mode can use DIMD derivation to derive the chroma intra prediction mode of the current block based on the neighboring reconstructed Y, Cb, and Cr samples in the second neighboring row and column, as Figure 7 shown. Horizontal and vertical gradients can be calculated for the collocated reconstructed luma samples and the reconstructed Cb and Cr samples of the current chroma block to construct an oriented gradient histogram (HoG). The intra prediction mode with the maximum histogram magnitude value can be used to perform the chroma intra prediction of the current chroma block.

[0137] When the intra prediction mode derived from the DIMD chroma mode is the same as the intra prediction mode derived from the direct chroma mode (DM), a refined intra prediction mode can be selected. This selection can be based on the determination that the histogram magnitude value corresponding to the refined intra prediction mode is the highest (e.g., the largest) among a plurality of histogram magnitude values (e.g., a plurality of candidate histogram magnitude values). The refined intra prediction mode (e.g., the highest histogram magnitude value) can be used as the DIMD chroma mode. A CU-level flag can be signaled to indicate whether the DIMD chroma mode is applied.

[0138] In an example, chroma intra prediction can use the same mode as the collocated luma prediction unit (PU). In DIMD, the derived chroma DIMD mode can be the same as the direct chroma mode (DM mode) (e.g., the same as the intra mode of the collocated luma PU), and the second mode from the histogram can be used for the chroma DIMD mode (e.g., to avoid redundant decoding). A signal can be sent to indicate (e.g., by the encoder / decoder) whether the DM mode is equal to the chroma DIMD mode).

[0139] As described herein, TIMD can use extended intra directions. In an example, 131 modes can be used (e.g., instead of 65 modes). 131 modes can be used to improve the prediction quality, and a higher number of directions can result in a finer prediction. When increasing the directions, additional signaling can be used.

[0140] In an example, in order to use an increased number of directions without increasing the signaling overhead, a reconstructed block template can be used to refine the prediction direction, e.g., using a reconstructed block template. Similar techniques (e.g., the same techniques) can be applied to both the encoder and decoder sides.

[0141] In an example, template-based intra direction refinement can be performed (e.g., see Figure 6 ). The encoder can find the best intra prediction mode and can signal the best intra prediction mode to the decoder. The resolution of the signaled prediction mode can be increased by inserting additional modes (e.g., inserting additional modes between the modes (e.g., between every two modes)). In an example, an angle can be inserted between the modes (e.g., two modes, as in the TIMD process). The encoder and decoder can analyze the reconstructed template (e.g., to decide which mode to select in the increased resolution).

[0142] In an example, the current resolution of the intra mode can be 67. In an example, the current resolution of the intra mode can include 65 angles with planar and DC modes (e.g., see Figure 6)。In TIMD, for example, the resolution can be 131 and can include 129 angles with planar and DC modes. The resolution can be doubled in TIMD (e.g., compared to an intra prediction mode that is not TIMD).

[0143] In an example involving a refinement algorithm, the resolution increase can be described by a factor "n". When "n" equals one (1), one (1) additional intra prediction angle can be inserted between two adjacent angles (e.g., every two (2) adjacent angles). When "n" equals two (2), two (2) additional intra prediction angles can be inserted between two (2) adjacent angles (e.g., every two (2) adjacent angles). In Table 1, "n = 1" and "n = 2" can be used.

[0144] Table 1: Additional intra angle patterns represented by "x" depending on the resolution factor.

[0145]

[0146] In Table 2, an algorithm description for refining the intra mode is provided.

[0147] Table 2: Algorithm description for refining the intra mode.

[0148]

[0149] The function AnalyzeTemplate can be used to calculate the score / distance of the prediction mode in the increased resolution to favor the intra prediction mode with the highest score / lowest distance.

[0150] In an example based on DIMD, a form can generate a gradient histogram on a reconstructed template to produce an optimal mode (e.g., the prediction quality of the optimal mode can be tested based on the directionality of the reconstructed template). The histogram can include directions given by the current intra mode to (n + the current intra mode). In an example based on TIMD, modes can be tested from the current mode to (n + the current mode), and the prediction error of the mode can be calculated (e.g., the cost of the mode can be calculated).

[0151] A video decoding device (e.g., a video encoding and / or decoding device) may calculate a first prediction of a template of a video block based on a first intra prediction mode. The device may calculate a second prediction of the template of the video block based on a second intra prediction mode. The device may obtain a first prediction error and a second prediction error corresponding to the first prediction mode and the second prediction mode respectively based on the first and second predictions. The device may select a refined intra prediction mode for the video block based on one or more of the first prediction error or the second prediction error. For example, the device may calculate a third prediction of the template of the video block based on the second intra prediction mode. The device may obtain first, second, and third errors corresponding to the first, second, and third intra prediction modes respectively based on the first, second, and third predictions. The best mode may be the mode that minimizes the prediction error. The refined intra prediction mode may be selected based on the intra prediction error corresponding to the refined intra prediction mode being determined to be the lowest among the prediction errors. The intra mode with refinement based on TIMD may be activated at the sequence parameter set (SPS) level, picture level, or slice level.

[0152] On the decoder side, refinement techniques may be used to find a second (e.g., additional) mode at a higher resolution and prediction may be performed.

[0153] In an example, a video decoding device (e.g., a video decoding device) may receive an indication of an intra prediction mode associated with a first resolution (e.g., 67 resolution) for a video block, and the indication of the intra prediction mode associated with the first resolution may be included in the video data (e.g., by the encoder). Based on the indicated intra prediction mode, the device may identify a second (e.g., additional) intra prediction mode associated with a second resolution (e.g., 131 resolution). In an example, the device may identify a third (e.g., additional) intra prediction mode associated with the second resolution (e.g., 131 resolution). The device may evaluate the first intra prediction mode and the second intra prediction mode on the template of the video block and select a refined intra prediction mode for the video block among the first intra prediction mode and the second intra prediction mode. The device may evaluate the first intra prediction mode, the second intra prediction mode, and the third intra prediction mode on the template of the video block and select a refined intra prediction mode for the video block among the first intra prediction mode, the second intra prediction mode, and the third intra prediction mode. The video block may be decoded based on the refined intra prediction mode (see Figure 6 ). Extended intra prediction may be used. In an example, 131 modes may be used by refining the signaled mode (e.g., from 67 resolution) to 131 resolution (e.g., see Figure 6 ). Example advantages may include a unified resolution of intra prediction modes (e.g., and other modes) in TIMD. The resolution may be increased by a factor greater than one (1). The prediction technical complexity may be proportional to the number of intra prediction modes.

[0154] Extended intra modes for DIMD can be used. In an example, DIMD can generate a histogram of directions of available intra modes (e.g., 65 available intra modes) by analyzing the gradients of the reconstructed template. A video decoding device (e.g., a video encoding and / or decoding device) can generate a gradient histogram on the template. The histogram can include directions associated with a first intra prediction mode and a second intra prediction mode. In the first intra prediction mode and the second intra prediction mode (e.g., additional intra prediction modes), a refined intra mode can be selected based on the histogram magnitude values associated with the first intra prediction mode and the second intra prediction mode. In an example, the histogram can include directions associated with a first, a second, and a third intra prediction mode. In the first, second, and third intra prediction modes (e.g., the second and third intra prediction modes can be additional intra prediction modes), a refined intra mode can be selected based on the histogram magnitude values associated with the first, second, and third intra prediction modes. A refinement technique (e.g., a TIMD-based technique) can be used to refine the resulting mode into 131 modes. This technique (e.g., refinement) can be applied to chrominance DIMD.

[0155] Figure 8 An example block diagram for refinement is illustrated. In an example, an intra mode can be decoded (e.g., an angular mode from 0 to 65). A template cost or a histogram can be used to determine the best (e.g., better) mode in a subsampled range. The best mode can be used to apply intra prediction.

[0156] Although the features and elements have been described above in specific combinations, one of ordinary skill in the art will understand that each feature or element can be used alone or in any combination with other features and elements. Additionally, the methods described herein can be implemented in a computer program, software, or firmware incorporated in a computer-readable medium for execution by a computer or a processor. Examples of computer-readable media include electronic signals (transmitted via wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, buffer memories, semiconductor storage devices, magnetic media (e.g., internal hard disks and removable disks), magneto-optical media, and optical media (e.g., CD-ROM disks and digital versatile disks (DVDs)). A processor associated with software can be used to implement a radio frequency transceiver used in a WTRU, UE, terminal, base station, RNC, or any host computer.

Claims

1. A video decoding device, comprising: a processor configured to: receive an indication of a first intra prediction mode associated with a first resolution for a video block; identify a second intra prediction mode associated with a second resolution based on the first intra prediction mode; evaluate the first intra prediction mode and the second intra prediction mode on a template of the video block; select a refined intra prediction mode for the video block from the first intra prediction mode and the second intra prediction mode; and decode the video block based on the refined intra prediction mode.

2. The device according to claim 1, wherein The processor is further configured to: obtain a first prediction of the template of the video block based on the first intra prediction mode; obtain a second prediction of the template of the video block based on the second intra prediction mode; calculate a first prediction error corresponding to the first intra prediction mode and a second prediction error corresponding to the second intra prediction mode, at least in part based on the first prediction and the second prediction; and select the refined intra prediction mode for the video block based on the first prediction error and the second prediction error.

3. The device according to claim 2, wherein, The refined intra prediction mode is selected based on a determination that the prediction error of the refined intra prediction mode is the lowest among the first prediction error and the second prediction error.

4. The device according to claim 1, wherein, The processor is further configured to: identify a third intra prediction mode associated with the second resolution based on the first intra prediction mode; evaluate the first intra prediction mode, the second intra prediction mode, and the third intra prediction mode on a template of the video block; obtain a first prediction of the template of the video block based on the first intra prediction mode; obtain a second prediction of the template of the video block based on the second intra prediction mode; obtain a third prediction of the template of the video block based on the third intra prediction mode; calculate a first prediction error corresponding to the first intra prediction mode, a second prediction error corresponding to the second intra prediction mode, and a third prediction error corresponding to the third intra prediction mode, at least in part based on the first prediction, the second prediction, and the third prediction; and select the refined intra prediction mode for the video block based on the first prediction error, the second prediction error, and the third prediction error.

5. The device according to claim 4, wherein, The refined intra prediction mode is selected based on a determination that the prediction error of the refined intra prediction mode is the lowest among the first prediction error, the second prediction error, and the third prediction error.

6. The device according to claim 1, wherein, The processor is further configured to: generate a gradient histogram associated with the template of the video block, wherein the histogram includes a plurality of histogram amplitude values associated with the first intra prediction mode and the second intra prediction mode; and select the refined intra mode from the first intra prediction mode and the second intra prediction mode based on the plurality of histogram amplitude values.

7. The device according to claim 6, wherein, The refined intra prediction mode is selected based on a determination that a histogram amplitude value corresponding to the refined intra prediction mode among the plurality of histogram amplitude values is the highest among the plurality of histogram amplitude values associated with the first intra prediction mode and the second intra prediction mode.

8. The device according to claim 1, wherein The processor is further configured to: Based on the first intra prediction mode, identify a third intra prediction mode associated with the second resolution; Evaluate the first intra prediction mode, the second intra prediction mode, and the third intra prediction mode on a template of the video block; Generate a gradient histogram associated with the template of the video block, wherein the histogram includes a plurality of histogram amplitude values associated with the first intra prediction mode, the second intra prediction mode, and the third intra prediction mode; And Based on the plurality of histogram amplitude values, select the refined intra mode among the first intra prediction mode, the second intra prediction mode, and the third intra prediction mode.

9. The apparatus according to claim 8, wherein, The refined intra prediction mode is selected based on a determination that a histogram amplitude value corresponding to the refined intra prediction mode is the highest among the plurality of histogram amplitude values associated with the first intra prediction mode, the second intra prediction mode, and the third intra prediction mode.

10. A video coding device, comprising: A processor configured to: Determine a first intra prediction mode associated with a first resolution for a video block; Based on the first intra prediction mode, identify a second intra prediction mode associated with a second resolution; Evaluate the first intra prediction mode and the second intra prediction mode on a template of the video block; Select a refined intra prediction mode for the video block among the first intra prediction mode and the second intra prediction mode; Include an indication of the first intra prediction mode in video data; And Encode the video block based on the refined intra prediction mode.

11. The device according to claim 10, wherein, The processor is further configured to: Obtain a first prediction of the template of the video block based on the first intra prediction mode; Obtain a second prediction of the template of the video block based on the second intra prediction mode; Calculate a first prediction error corresponding to the first intra prediction mode and a second prediction error corresponding to the second intra prediction mode at least partially based on the first prediction and the second prediction; And Select the refined intra prediction mode for the video block based on the first prediction error and the second prediction error.

12. The device according to claim 11, wherein, The refined intra prediction mode is selected based on a determination that a prediction error of the refined intra prediction mode is the lowest among the first prediction error and the second prediction error.

13. The device according to claim 10, wherein, The processor is further configured to: Based on the first intra prediction mode, identify a third intra prediction mode associated with the second resolution; Evaluate the first intra prediction mode, the second intra prediction mode, and the third intra prediction mode on a template of the video block; Obtain a first prediction of the template of the video block based on the first intra prediction mode; Obtain a second prediction of the template of the video block based on the second intra prediction mode; Obtain a third prediction of the template of the video block based on the third intra prediction mode; Calculate a first prediction error corresponding to the first intra prediction mode, a second prediction error corresponding to the second intra prediction mode, and a third prediction error corresponding to the third intra prediction mode, at least in part based on the first prediction, the second prediction, and the third prediction; And Select the refined intra prediction mode for the video block based on the first prediction error, the second prediction error, and the third prediction error.

14. The apparatus according to claim 13, wherein, The refined intra prediction mode is selected based on the determination that the prediction error based on the refined intra prediction mode is the lowest among the first prediction error, the second prediction error, and the third prediction error.

15. The apparatus according to claim 10, wherein, The processor is further configured to: Generate a gradient histogram associated with the template of the video block, wherein the histogram includes a plurality of histogram magnitude values associated with the first intra prediction mode and the second intra prediction mode; And Select the refined intra mode among the first intra prediction mode and the second intra prediction mode based on the plurality of histogram magnitude values.

16. The device according to claim 15, wherein, The refined intra prediction mode is selected based on the determination that the histogram magnitude value corresponding to the refined intra prediction mode is the highest among the plurality of histogram magnitude values associated with the first intra prediction mode and the second intra prediction mode.

17. The apparatus according to claim 10, wherein, The processor is further configured to: Identify a third intra prediction mode associated with the second resolution based on the first intra prediction mode; Evaluate the first intra prediction mode, the second intra prediction mode, and the third intra prediction mode on the template of the video block; Generate a gradient histogram associated with the template of the video block, wherein the histogram includes a plurality of histogram magnitude values associated with the first intra prediction mode, the second intra prediction mode, and the third intra prediction mode; And Select the refined intra mode among the first intra prediction mode, the second intra prediction mode, and the third intra prediction mode based on the plurality of histogram magnitude values.

18. The device according to claim 17, wherein, The refined intra prediction mode is selected based on the determination that the histogram magnitude value corresponding to the refined intra prediction mode is the highest among the plurality of histogram magnitude values associated with the first intra prediction mode, the second intra prediction mode, and the third intra prediction mode.

19. The apparatus according to any one of claims 1 to 18, further comprising a memory operably connected to the processor.

20. A method for a video decoding device, the method comprising: Receive an indication of a first intra prediction mode associated with a first resolution for a video block; Identify a second intra prediction mode associated with a second resolution based on the first intra prediction mode; Evaluate the first intra prediction mode and the second intra prediction mode on the template of the video block; Select a refined intra prediction mode for the video block among the first intra prediction mode and the second intra prediction mode; And Decode the video block based on the refined intra prediction mode.

21. The method according to claim 20, wherein, The method further includes: Obtain a first prediction of the template of the video block based on the first intra prediction mode; Obtain a second prediction of the template of the video block based on the second intra prediction mode; Calculate a first prediction error corresponding to the first intra prediction mode and a second prediction error corresponding to the second intra prediction mode, at least partially based on the first prediction and the second prediction; and Select the refined intra prediction mode for the video block based on the first prediction error and the second prediction error.

22. The method according to claim 21, wherein, The refined intra prediction mode is selected based on the determination that the prediction error based on the refined intra prediction mode is the lowest among the first prediction error and the second prediction error.

23. The method according to claim 20, wherein, The method further includes: Identify a third intra prediction mode associated with the second resolution based on the first intra prediction mode; Evaluate the first intra prediction mode, the second intra prediction mode, and the third intra prediction mode on the template of the video block; Obtain a first prediction of the template of the video block based on the first intra prediction mode; Obtain a second prediction of the template of the video block based on the second intra prediction mode; Obtain a third prediction of the template of the video block based on the third intra prediction mode; Calculate a first prediction error corresponding to the first intra prediction mode, a second prediction error corresponding to the second intra prediction mode, and a third prediction error corresponding to the third intra prediction mode, at least partially based on the first prediction, the second prediction, and the third prediction; and Select the refined intra prediction mode for the video block based on the first prediction error, the second prediction error, and the third prediction error.

24. The method according to claim 23, wherein The refined intra prediction mode is selected based on the determination that the prediction error based on the refined intra prediction mode is the lowest among the first prediction error, the second prediction error, and the third prediction error.

25. The method according to claim 20, wherein, The method further includes: Generate a gradient histogram associated with the template of the video block, where the histogram includes a plurality of histogram amplitude values associated with the first intra prediction mode and the second intra prediction mode; and Select the refined intra mode among the first intra prediction mode and the second intra prediction mode based on the plurality of histogram amplitude values.

26. The method according to claim 25, wherein, The refined intra prediction mode is selected based on the determination that the histogram amplitude value corresponding to the refined intra prediction mode among the plurality of histogram amplitude values is the highest among the plurality of histogram amplitude values associated with the first intra prediction mode and the second intra prediction mode.

27. The method according to claim 20, wherein The method further includes: Identify a third intra prediction mode associated with the second resolution based on the first intra prediction mode; Evaluate the first intra prediction mode, the second intra prediction mode, and the third intra prediction mode on the template of the video block; Generate a histogram of gradients associated with the template of the video block, wherein the histogram includes a plurality of histogram amplitude values associated with the first intra prediction mode, the second intra prediction mode, and the third intra prediction mode; and Based on the plurality of histogram amplitude values, select the refined intra mode among the first intra prediction mode, the second intra prediction mode, and the third intra prediction mode.

28. The method according to claim 27, wherein, The refined intra prediction mode is selected based on a determination that the histogram amplitude value corresponding to the refined intra prediction mode is the highest among the plurality of histogram amplitude values associated with the first intra prediction mode, the second intra prediction mode, and the third intra prediction mode.

29. A method for a video coding device, the method comprising: Determine a first intra prediction mode associated with a first resolution for a video block; Based on the first intra prediction mode, identify a second intra prediction mode associated with a second resolution; Evaluate the first intra prediction mode and the second intra prediction mode on a template of the video block; Select a refined intra prediction mode for the video block among the first intra prediction mode and the second intra prediction mode; Include an indication of the first intra prediction mode in video data; And Encode the video block based on the refined intra prediction mode.

30. The method according to claim 29, wherein, The method further comprises: Obtain a first prediction of the template of the video block based on the first intra prediction mode; Obtain a second prediction of the template of the video block based on the second intra prediction mode; Calculate a first prediction error corresponding to the first intra prediction mode and a second prediction error corresponding to the second intra prediction mode at least in part based on the first prediction and the second prediction; and Select the refined intra prediction mode for the video block based on the first prediction error and the second prediction error.

31. The method according to claim 30, wherein, The refined intra prediction mode is selected based on a determination that the prediction error of the refined intra prediction mode is the lowest among the first prediction error and the second prediction error.

32. The method according to claim 29, wherein The method further comprises: Based on the first intra prediction mode, identify a third intra prediction mode associated with the second resolution; Evaluate the first intra prediction mode, the second intra prediction mode, and the third intra prediction mode on a template of the video block; Obtain a first prediction of the template of the video block based on the first intra prediction mode; Obtain a second prediction of the template of the video block based on the second intra prediction mode; Obtain a third prediction of the template of the video block based on the third intra prediction mode; Calculate a first prediction error corresponding to the first intra prediction mode, a second prediction error corresponding to the second intra prediction mode, and a third prediction error corresponding to the third intra prediction mode at least in part based on the first prediction, the second prediction, and the third prediction; and Select the refined intra prediction mode for the video block based on the first prediction error, the second prediction error, and the third prediction error.

33. The method according to claim 32, wherein The refined intra prediction mode is selected based on a determination that the prediction error based on the refined intra prediction mode is the lowest among the first prediction error, the second prediction error, and the third prediction error.

34. The method according to claim 29, wherein The method further comprises: generating a gradient histogram associated with the template of the video block, wherein the histogram comprises a plurality of histogram magnitude values associated with the first intra prediction mode and the second intra prediction mode; and selecting the refined intra mode from the first intra prediction mode and the second intra prediction mode based on the plurality of histogram magnitude values.

35. The method according to claim 34, wherein The refined intra prediction mode is selected based on a determination that the histogram magnitude value corresponding to the refined intra prediction mode among the plurality of histogram magnitude values is the highest among the plurality of histogram magnitude values associated with the first intra prediction mode and the second intra prediction mode.

36. The method according to claim 29, wherein, The method further comprises: identifying a third intra prediction mode associated with the second resolution based on the first intra prediction mode; evaluating the first intra prediction mode, the second intra prediction mode, and the third intra prediction mode on the template of the video block; generating a gradient histogram associated with the template of the video block, wherein the histogram comprises a plurality of histogram magnitude values associated with the first intra prediction mode, the second intra prediction mode, and the third intra prediction mode; and selecting the refined intra mode from the first intra prediction mode, the second intra prediction mode, and the third intra prediction mode based on the plurality of histogram magnitude values.

37. The method according to claim 36, wherein, The refined intra prediction mode is selected based on a determination that the histogram magnitude value corresponding to the refined intra prediction mode is the highest among the plurality of histogram magnitude values associated with the first intra prediction mode, the second intra prediction mode, and the third intra prediction mode.

38. A computer program product, stored on a non-transitory computer-readable medium and comprising program code instructions for implementing the steps of the method according to any one of claims 20 to 37 when executed by a processor.

39. A computer program comprising program code instructions for implementing the steps of the method according to any one of claims 20 to 37 when executed by a processor.

40. A video data comprising information representing a video block encoded according to one of the methods according to any one of claims 29 to 37.