Equivalent intra mode for non-intra prediction coded blocks
By deriving the directional intra prediction mode in the video decoding device, using the gradient histogram of the reconstructed samples and adjacent samples, the problem of inefficient decoding efficiency of non-directional intra prediction mode in the prior art is solved, and more efficient intra prediction and decoding quality is achieved.
Patent Information
- Application Number
- CN202380072098.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-11
- Filing Date
- 2023-10-11
- Publication Date
- 2025-06-24
AI Technical Summary
When the existing video decoding system processes the non-directional intra prediction mode, it is difficult to effectively deduce the directional intra prediction mode, resulting in low decoding efficiency.
By deducing the directional intra prediction mode in the video decoding device, using the gradient histogram of the reconstructed samples and adjacent samples in the prediction block, multiple candidate directional intra prediction modes are tested, and the optimal mode is selected for decoding.
The processing efficiency of the video decoding system for non-directed intra prediction mode is improved, and the direction accuracy and decoding quality of intra prediction are enhanced.
Smart Images

Figure CN120202660A_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the benefit of European Provisional Patent Application No. EP22306526.9, filed on October 11, 2022, the content of which is incorporated herein by reference. Background Art
[0003] Video coding systems can be used to compress digital video signals, for example, to reduce the storage and / or transmission bandwidth required for such signals. Video coding systems can include, for example, block - based, wavelet - based, and / or object - based systems. Summary of the Invention
[0004] Systems, methods, and means for deriving equivalent intra - modes are disclosed.
[0005] An example device (e.g., a video decoding device) can determine that a current block is decoded in a non - directional intra - prediction mode. The device can derive a directional intra - prediction mode corresponding to the non - directional intra - prediction mode. The derived directional intra - prediction mode can indicate a derived intra - prediction direction. The device can decode the current block at least in part based on the derived directional intra - prediction mode.
[0006] Similarly, an example device (e.g., a video encoding device) can identify a non - directional intra - prediction mode for encoding a current block. The device can derive a directional intra - prediction mode corresponding to the non - directional intra - prediction mode. The derived directional intra - prediction mode can include a derived intra - prediction direction. The device can encode the current block at least in part based on the derived directional intra - prediction mode.
[0007] The device can obtain a predicted block of the current block using the non - directional intra - prediction mode. The device can obtain a plurality of reconstructed samples in the predicted block. The directional intra - prediction mode can be derived based on the plurality of reconstructed samples in the predicted block and a plurality of reconstructed neighboring samples of the current block.
[0008] The device can store the derived directional intra - prediction mode. The device can use the derived directional intra - prediction mode to generate a most - probable mode (MPM) list for neighboring predicted blocks. The device can determine a low - frequency non - separable transform (LFNST) transform set based on the derived directional intra - prediction mode. The current block can be decoded / encoded based on the LFNST transform set. The device can determine a multi - transform selection (MTS) transform set based on the derived directional intra - prediction mode. The current block can be decoded / encoded based on the MTS transform set.
[0009] Deriving a directional intra prediction mode may involve deriving a directional intra prediction mode based on a gradient histogram associated with reconstructed pixels neighboring a current block. Deriving a directional intra prediction mode may include: testing a plurality of candidate directional intra prediction modes on reconstructed pixels neighboring the current block; and selecting a directional intra prediction mode from the plurality of candidate directional intra prediction modes based on the testing.
[0010] The device may obtain a predicted block of the current block using a non-directional intra prediction mode. The device may obtain a plurality of reconstructed samples in the predicted block. The device may obtain a plurality of possible prediction modes. The device may calculate a plurality of predictions of the plurality of reconstructed samples in the predicted block based on the plurality of possible prediction modes. The device may calculate a plurality of prediction errors corresponding to the plurality of possible prediction modes based on the plurality of reconstructed samples and the corresponding plurality of predictions in the predicted block. The device may select a directional intra prediction mode from the plurality of possible prediction modes based on the plurality of prediction errors.
[0011] The non-directional intra prediction mode may be an inter prediction mode, a cross-component prediction mode, a palette mode, an intra block copy (IBC) mode, or an intra template matching prediction (IntraTMP) mode.
[0012] The device may select a low-frequency non-separable transform (LFNST) transform set based on the directional intra prediction mode. The device may perform an inverse transform on the residual of the current block based on the LFNST transform set.
[0013] The device may select a multi-transform selection (MTS) transform set based on the directional intra prediction mode. The device may perform an inverse transform on the residual of the current block based on the MTS transform set.
[0014] A video decoding device may include a processor configured to determine that a current block is decoded in a non-intra prediction mode (e.g., a non-directional intra prediction mode, a DC mode, or a planar mode). For example, the non-intra prediction mode may be one or more of an inter prediction mode, a cross-component prediction mode, a palette mode, an intra block copy (IBC) mode, or an intra template matching prediction (IntraTMP) mode. An intra prediction mode corresponding to the non-intra prediction mode may be derived. The current block may be decoded at least in part based on the derived intra prediction mode.
[0015] In an example, a non-intra prediction mode may be used to obtain a predicted block of the current block, and an intra prediction mode corresponding to the non-intra prediction mode may be derived based on the predicted block.
[0016] In an example, a non-intra prediction mode may be used to obtain a predicted block of the current block. Reconstructed samples in the predicted block may be obtained. An intra prediction mode may be derived based on the reconstructed samples and reconstructed neighboring samples in the predicted block.
[0017] An intra prediction mode can be derived by applying a decoder-side Derived Intra Mode Decision (DIMD) process to at least one of a reconstructed template of a current block (e.g., a template around the current block, template samples adjacent to the current block), a predicted block of the current block obtained using a non-intra prediction mode, or a reconstructed template within the predicted block. An intra prediction mode can be derived by applying a Template-based Intra Mode Decision (TIMD) process to at least one of a reconstructed template of a current block, a predicted block of the current block obtained using a non-intra prediction mode, or a reconstructed template within the predicted block.
[0018] In an example, a non-intra prediction mode can be used to obtain a predicted block of a current block. Reconstructed samples in the predicted block can be obtained. Candidate prediction modes can be obtained. A prediction of the reconstructed samples in the predicted block can be calculated based on the candidate prediction modes. A prediction error corresponding to a candidate prediction mode can be calculated based on the reconstructed samples in the predicted block and the corresponding prediction. An intra prediction mode can be selected from the candidate prediction modes based on the prediction error. The intra prediction mode can be selected based on determining that the prediction error corresponding to the intra prediction mode is the smallest among the prediction errors.
[0019] In an example, a non-intra prediction mode can be used to obtain a predicted block of a current block. Samples in the predicted block can be obtained. A directionality of the predicted block can be determined based on the samples in the predicted block. An intra prediction mode can be derived based on the determined directionality of the predicted block.
[0020] A video coding device can include a processor configured to identify a non-intra prediction mode for encoding a current block. An intra prediction mode corresponding to the non-intra prediction mode can be derived. The current block can be encoded at least in part based on the derived intra prediction mode. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In addition, like reference numerals in the drawings denote like elements, and in which:
[0022] Figure 1A is a system diagram illustrating an example communication system in which one or more of the disclosed embodiments may be implemented.
[0023] Figure 1B is an illustration of an example wireless transmit / receive unit (WTRU) that may be used within the Figure 1A illustrated communication system according to an embodiment.
[0024] Figure 1C is an illustration of an example radio access network (RAN) and an example core network (CN) that may be used within the Figure 1A illustrated communication system according to an embodiment.
[0025] Figure 1D Illustrates another example RAN and another example CN system diagram that can be used within the illustrated communication system according to an embodiment. Figure 1A Illustrates another example RAN and another example CN system diagram that can be used within the illustrated communication system according to an embodiment.
[0026] Figure 2 Illustrates an example video encoder.
[0027] Figure 3 Illustrates an example video decoder.
[0028] Figure 4 Illustrates an example of a system in which various aspects and examples can be implemented.
[0029] Figures 5A to 5C Shows an example prediction mode and prediction direction.
[0030] Figure 6 Shows an example of a template of the current luminance and decoded reference samples of the template.
[0031] Figure 7 Shows neighboring reconstructed samples for (e.g.) decoder-side intra-mode derivation (DIMD) chrominance modes.
[0032] Figure 8 Illustrates an example of the matrix weighted intra prediction (MIP) process.
[0033] Figure 9 Illustrates an example position of samples used in the cross-component linear model (CCLM) mode.
[0034] Figure 10A and Figure 10B Illustrates an example effect of the slope adjustment parameter.
[0035] Figure 11 Illustrates the spatial part of the convolutional filter.
[0036] Figure 12 Illustrates an example reference region for intra-block copy (IBC) when a coding tree unit (CTU) is coded.
[0037] Figure 13 Illustrates an example intra-template matching search region.
[0038] Figure 14 Illustrates an example of a block coded in palette mode.
[0039] Figures 15A to 15D Illustrates an example geometric partitioning mode (GPM) with inter and intra prediction.
[0040] Figure 16Is a table of positions of available neighboring blocks for intra prediction mode (IPM) candidate derivation based on the angles of GPM block boundaries.
[0041] Figure 17 Illustrates an example region of interest (ROI).
[0042] Figure 18 Illustrates an example ROI.
[0043] Figure 19 Illustrates an example mapping of intra prediction modes to low-frequency non-separable transform (LFNST) set indices.
[0044] Figure 20 Illustrates example neighboring blocks used to derive the general most probable mode (MPM) list.
[0045] Figure 21 Illustrates an example of deriving equivalent modes.
[0046] Figure 22A Illustrates an example DIMD process.
[0047] Figure 22B Illustrates MIP equivalent mode derivation.
[0048] Figure 23 Illustrates an example decoded block partition.
[0049] Figure 24 Illustrates an example flowchart for decoding a current block.
[0050] Figure 25 Illustrates an example flowchart for encoding a current block. Detailed Description
[0051] Figure 1A Is a diagram illustrating an example communication system 100 in which one or more of the disclosed embodiments may be implemented. The communication system 100 may be a multi-access system that provides content such as voice, data, video, messages, broadcasts, etc. to a plurality of wireless users. The communication system 100 may enable a plurality of wireless users to access such content through the sharing of system resources, including wireless bandwidth. For example, the communication system 100 may employ one or more channel access methods such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single carrier FDMA (SC-FDMA), zero-tail unique word DFT-spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block filtered OFDM, filter bank multicarrier (FBMC), etc.
[0052] AsFigure 1A As shown, the communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, a radio access network (RAN) 104 / 113, a core network (CN) 106 / 115, a public switched telephone network (PSTN) 108, the Internet 110, and other networks 112. However, it should be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, 102d can be any type of device configured to operate and / or communicate in a wireless environment. By way of example, the WTRUs 102a, 102b, 102c, 102d (any of which may be referred to as a "station" and / or "STA") can be configured to transmit and / or receive wireless signals and may include user equipment (UE), a mobile station, a fixed or mobile subscriber unit, a subscription-based unit, a pager, a cellular phone, a personal digital assistant (PDA), a smartphone, a laptop computer, a netbook, a personal computer, a wireless sensor, a hotspot or Mi-Fi device, an Internet of Things (IoT) device, a watch or other wearable device, a head-mounted display (HMD), a vehicle, a drone, a medical device and application (e.g., remote surgery), an industrial device and application (e.g., a robot and / or other wireless devices operating in an industrial and / or automated processing chain environment), a consumer electronic device, a device operating on a commercial and / or industrial wireless network, etc. Any one of the WTRUs 102a, 102b, 102c, and 102d may be interchangeably referred to as a UE.
[0053] The communication system 100 may also include base stations 114a and / or base stations 114b. Each of the base stations 114a, 114b can be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communication networks such as the CN 106 / 115, the Internet 110, and / or other networks 112. By way of example, the base stations 114a, 114b can be transceiver base stations (BTSs), Node Bs, evolved Node Bs, home Node Bs, home evolved Node Bs, gNBs, NR Node Bs, site controllers, access points (APs), wireless routers, etc. Although the base stations 114a, 114b are each depicted as a single element, it should be understood that the base stations 114a, 114b can include any number of interconnected base stations and / or network elements.
[0054] Base station 114a may be part of RAN 104 / 113, which may also include other base stations and / or network elements (not shown), such as a base station controller (BSC), a radio network controller (RNC), a relay node, etc. Base station 114a and / or base station 114b may be configured to transmit and / or receive wireless signals on one or more carrier frequencies, which may be referred to as cells (not shown). These frequencies may be in licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide coverage of wireless services to a specific geographical area, which may be relatively fixed or may change over time. A cell may be further divided into cell sectors. For example, the cell associated with base station 114a may be divided into three sectors. Thus, in one embodiment, base station 114a may include three transceivers, i.e., one transceiver for each sector of the cell. In an embodiment, base station 114a may employ multiple-input multiple-output (MIMO) technology and may utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.
[0055] Base stations 114a, 114b may communicate with one or more of the WTRUs 102a, 102b, 102c, 102d via air interface 116, which may be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, millimeter wave, infrared (IR), ultraviolet (UV), visible light, etc.). Any suitable radio access technology (RAT) may be used to establish air interface 116.
[0056] More specifically, as noted above, communication system 100 may be a multiple access system and may employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, base station 114a in RAN 104 / 113 and WTRUs 102a, 102b, 102c may implement a radio technology such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which may use Wideband CDMA (WCDMA) to establish air interfaces 115 / 116 / 117. WCDMA may include communication protocols such as High-Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High-Speed Downlink (DL) Packet Access (HSDPA) and / or High-Speed UL Packet Access (HSUPA).
[0057] In an embodiment, base station 114a and WTRUs 102a, 102b, 102c may implement a radio technology such as evolved UMTS terrestrial radio access (E-UTRA), which may use Long Term Evolution (LTE) and / or Long Term Evolution-Advanced (LTE-A) and / or LTE-A Pro to establish an air interface 116.
[0058] In an embodiment, base station 114a and WTRUs 102a, 102b, 102c may implement a radio technology such as NR radio access, which may use New Radio (NR) to establish an air interface 116.
[0059] In an embodiment, base station 114a and WTRUs 102a, 102b, 102c may implement multiple radio access technologies. For example, base station 114a and WTRUs 102a, 102b, 102c may implement LTE radio access and NR radio access together using, for example, the dual connectivity (DC) principle. Accordingly, the air interface utilized by WTRUs 102a, 102b, 102c may be characterized by multiple types of radio access technologies and / or transmissions to / from multiple types of base stations (e.g., eNBs and gNBs).
[0060] In other embodiments, base station 114a and WTRUs 102a, 102b, 102c may implement radio technologies such as IEEE 802.11 (i.e., Wi-Fi), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data rates for GSM Evolution (EDGE), GSM EDGE (GERAN), etc.
[0061] Figure 1AThe base station 114b therein can be, for example, a wireless router, a home Node B, a home evolved Node B, or an access point, and can utilize any suitable RAT to facilitate wireless connectivity in local areas such as business premises, homes, vehicles, campuses, industrial facilities, air corridors (e.g., for drones), roads, etc. In one embodiment, the base station 114b and the WTRUs 102c, 102d can implement a radio technology (such as IEEE 802.11) to establish a wireless local area network (WLAN). In an embodiment, the base station 114b and the WTRUs 102c, 102d can implement a radio technology (such as IEEE 802.15) to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 114b and the WTRUs 102c, 102d can utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a pico cell or a femto cell. As Figure 1A shown, the base station 114b can have a direct connection to the Internet 110. Thus, the base station 114b may not need to access the Internet 110 via the CN106 / 115.
[0062] The RAN 104 / 113 can communicate with the CN 106 / 115, which can be any type of network configured to provide voice, data, applications, and / or voice over Internet Protocol (VoIP) services to one or more of the WTRUs 102a, 102b, 102c, 102d. The data can have different quality of service (QoS) requirements, such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, etc. The CN 106 / 115 can provide call control, billing services, location-based services, prepaid calls, Internet connectivity, video distribution, etc., and / or perform advanced security functions such as user authentication. Although not shown in Figure 1A it, it should be understood that the RAN 104 / 113 and / or the CN106 / 115 can communicate directly or indirectly with other RANs that employ the same RAT as the RAN 104 / 113 or a different RAT. For example, in addition to being connected to the RAN 104 / 113 that can utilize NR radio technology, the CN 106 / 115 can also communicate with another RAN (not shown) that employs GSM, UMTS, CDMA 2000, WiMAX, E-UTRA, or WiFi radio technology.
[0063] CN 106 / 115 can also be used as a gateway for WTRU 102a, 102b, 102c, 102d to access the PSTN 108, the Internet 110, and / or other networks 112. The PSTN 108 may include a circuit-switched telephone network that provides plain old telephone service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices that use common communication protocols such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) in the TCP / IP Internet protocol suite. The network 112 may include a wired communication network and / or a wireless communication network owned and / or operated by other service providers. For example, the network 112 may include another CN connected to one or more RANs, and the one or more RANs may employ the same RAT or a different RAT as the RAN 104 / 113.
[0064] Some or all of the WTRUs in the communication system 100, such as WTRU 102a, 102b, 102c, 102d, may include multi-mode capabilities (e.g., WTRU 102a, 102b, 102c, 102d may include multiple transceivers for communicating with different wireless networks over different wireless links). For example, Figure 1A the illustrated WTRU 102c may be configured to communicate with a base station 114a that may employ a cellular-based radio technology and with a base station 114b that may employ IEEE 802 radio technology.
[0065] Figure 1B is a system diagram illustrating an example WTRU 102. As Figure 1B shown, the WTRU 102 may include a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, a non-removable memory 130, a removable memory 132, a power supply 134, a Global Positioning System (GPS) chipset 136, and / or other peripheral devices 138, etc. It should be understood that the WTRU 102 may include any sub-combination of the foregoing elements while remaining consistent with the embodiments.
[0066] The processor 118 can be a general-purpose processor, a dedicated processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. As suggested above, the processor 118 can include multiple processors. The processor 118 can perform signal decoding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 102 to operate in a wireless environment. The processor 118 can be coupled to the transceiver 120, which can be coupled to the transmit / receive element 122. Although Figure 1B the processor 118 and the transceiver 120 are depicted as separate components, it should be understood that the processor 118 and the transceiver 120 can be integrated together in an electronic package or chip.
[0067] The transmit / receive element 122 can be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) via the air interface 116. For example, in one embodiment, the transmit / receive element 122 can be an antenna configured to transmit and / or receive RF signals. In an embodiment, the transmit / receive element 122 can be a transmitter / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, the transmit / receive element 122 can be configured to transmit and / or receive both RF signals and optical signals. It should be understood that the transmit / receive element 122 can be configured to transmit and / or receive any combination of wireless signals.
[0068] Although the transmit / receive element 122 is depicted as a single element in Figure 1B the WTRU 102 can include any number of transmit / receive elements 122. More specifically, the WTRU 102 can employ MIMO technology. Thus, in one embodiment, the WTRU 102 can include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals via the air interface 116.
[0069] The transceiver 120 can be configured to modulate the signals to be transmitted by the transmit / receive element 122 and demodulate the signals received by the transmit / receive element 122. As noted above, the WTRU 102 can have multi-mode capabilities. For example, thus, the transceiver 120 can include multiple transceivers for enabling the WTRU 102 to communicate via multiple RATs (such as NR and IEEE 802.11).
[0070] The processor 118 of the WTRU 102 may be coupled to the speaker / microphone 124, keypad 126, and / or the display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light emitting diode (OLED) display unit) and may receive user input data therefrom. The processor 118 may also output user data to the speaker / microphone 124, keypad 126, and / or the display / touchpad 128. In addition, the processor 118 may access information from any type of suitable memory (such as non-removable memory 130 and / or removable memory 132) and store data in any type of suitable memory. The non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 may include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, etc. In other embodiments, the processor 118 may access information from a memory that is not physically located on the WTRU 102 (such as on a server or a home computer (not shown)) and store data in that memory.
[0071] The processor 118 may receive power from a power source 134 and may be configured to distribute and / or control power to other components in the WTRU 102. The power source 134 may be any suitable device for powering the WTRU 102. For example, the power source 134 may include one or more dry battery packs (e.g., nickel cadmium (NiCd), nickel zinc (NiZn), nickel metal hydride (NiMH), lithium ion (Li-ion), etc.), a solar cell, a fuel cell, etc.
[0072] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to or instead of the information from the GPS chipset 136, the WTRU 102 may receive location information via an air interface 116 from a base station (e.g., base stations 114a, 114b) and / or determine its location based on the timing of signals received from two or more nearby base stations. It should be understood that the WTRU 102 may obtain location information by any suitable location determination method while remaining consistent with the embodiments.
[0073] The processor 118 may also be coupled to other peripheral devices 138, which may include one or more software modules and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, the peripheral devices 138 may include an accelerometer, an electronic compass, a satellite transceiver, a digital camera (for photos and / or videos), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands-free headset, Modules, FM radio units, digital music players, media players, video game player modules, Internet browsers, virtual reality and / or augmented reality (VR / AR) devices, activity trackers, etc. The peripheral device 138 may include one or more sensors, which may be one or more of the following: gyroscopes, accelerometers, Hall effect sensors, magnetometers, orientation sensors, proximity sensors, temperature sensors, time sensors; geographical location sensors; altimeters, light sensors, touch sensors, magnetometers, barometers, gesture sensors, biometric sensors, and / or humidity sensors.
[0074] The WTRU 102 may include a full-duplex radio, for which the transmission and reception of some or all signals (e.g., associated with a specific subframe for both UL (e.g., for transmission) and downlink (e.g., for reception)) may be concurrent and / or simultaneous. The full-duplex radio may include an interference management unit 139 for reducing and / or substantially eliminating self-interference via signal processing performed by hardware (e.g., chokes) or via a processor (e.g., a separate processor (not shown) or via the processor 118). In an embodiment, the WRTU 102 may include a half-duplex radio, for which the transmission and reception of some or all signals (e.g., associated with a specific subframe for UL (e.g., for transmission) or downlink (e.g., for reception)).
[0075] Figure 1C Is a system diagram of the RAN 104 and the CN 106 according to an embodiment. As noted above, the RAN 104 may employ E-UTRA radio technology to communicate with the WTRUs 102a, 102b, 102c via the air interface 116. The RAN 104 may also communicate with the CN 106.
[0076] The RAN 104 may include evolved Node Bs 160a, 160b, 160c, but it should be understood that the RAN 104 may include any number of evolved Node Bs while remaining consistent with the embodiment. Each of the evolved Node Bs 160a, 160b, 160c may include one or more transceivers for communicating with the WTRUs 102a, 102b, 102c via the air interface 116. In one embodiment, the evolved Node Bs 160a, 160b, 160c may implement MIMO technology. Thus, the evolved Node B 160a, for example, may use multiple antennas to transmit wireless signals to and / or receive wireless signals from the WTRU 102a.
[0077] Each of the evolved Node Bs 160a, 160b, 160c may be associated with a specific cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, scheduling of users in the UL and / or DL, etc. As Figure 1C shown, the evolved Node Bs 160a, 160b, 160c may communicate with each other via the X2 interface.
[0078] Figure 1C The CN 106 shown may include a Mobility Management Entity (MME) 162, a Serving Gateway (SGW) 164, and a Packet Data Network (PDN) Gateway (or PGW) 166. Although each of the foregoing elements is depicted as part of the CN 106, it should be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.
[0079] The MME 162 may be connected to each of the evolved Node Bs 162a, 162b, 162c in the RAN 104 via the S1 interface and may act as a control node. For example, the MME 162 may be responsible for authenticating users of the WTRUs 102a, 102b, 102c, bearer activation / deactivation, selecting a specific serving gateway during the initial attachment of the WTRUs 102a, 102b, 102c, etc. The MME 162 may provide control plane functions for interworking between the RAN 104 and other RANs (not shown) employing other radio technologies such as GSM and / or WCDMA.
[0080] The SGW 164 may be connected to each of the evolved Node Bs 160a, 160b, 160c in the RAN 104 via the S1 interface. The SGW 164 may generally route and forward user data packets to / from the WTRUs 102a, 102b, 102c. The SGW 164 may perform other functions such as anchoring the user plane during handover between evolved Node Bs, triggering paging when DL data is available for the WTRUs 102a, 102b, 102c, managing and storing the context of the WTRUs 102a, 102b, 102c, etc.
[0081] The SGW 164 may be connected to the PGW 166, which may provide the WTRUs 102a, 102b, 102c with access to a packet switched network such as the Internet 110 to facilitate communication between the WTRUs 102a, 102b, 102c and IP-enabled devices.
[0082] CN 106 can facilitate communication with other networks. For example, CN 106 can provide the WTRUs 102a, 102b, 102c with access to a circuit-switched network (such as the PSTN 108) to facilitate communication between the WTRUs 102a, 102b, 102c and traditional landline communication devices. For example, CN 106 can include an IP gateway (such as an IP Multimedia Subsystem (IMS) server) that acts as an interface between CN 106 and the PSTN 108 or can communicate with the IP gateway. Additionally, CN 106 can provide the WTRUs 102a, 102b, 102c with access to other networks 112, which can include other wired and / or wireless networks owned and / or operated by other service providers.
[0083] Although the WTRU is described as a wireless terminal in Figures 1A to 1D it is contemplated that in some representative embodiments, such a terminal can (e.g., temporarily or permanently) use a wired communication interface with the communication network.
[0084] In a representative embodiment, the other network 112 can be a WLAN.
[0085] A WLAN in infrastructure basic service set (BSS) mode can have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP can have access to or an interface with a distribution system (DS) or another type of wired / wireless network that carries traffic to and / or from the BSS. Traffic originating outside the BSS and destined for an STA can reach the STA through the AP and can be delivered to the STA. Traffic originating from an STA and destined for a destination outside the BSS can be delivered to the AP for delivery to the corresponding destination. Traffic between STAs within the BSS can be delivered through the AP. For example, where the source STA can deliver traffic to the AP and the AP can deliver the traffic to the destination STA. Traffic between STAs within the BSS can be considered and / or referred to as peer-to-peer traffic. Peer-to-peer traffic can be delivered between the source STA and the destination STA (e.g., directly between them) using direct link setup (DLS). In some representative embodiments, DLS can use 802.11e DLS or 802.11z tunnel DLS (TDLS). A WLAN using independent BSS (IBSS) mode may not have an AP, and STAs within the IBSS or using the IBSS (e.g., all STAs in the IBSS) can communicate directly with each other. The IBSS communication mode can sometimes be referred to as an "ad hoc" communication mode in this document.
[0086] When using the 802.11ac infrastructure operation mode or a similar operation mode, the AP may send beacons on a fixed channel, such as the primary channel. The primary channel may be of a fixed width (e.g., 20 MHz bandwidth) or a width dynamically set via signaling. The primary channel may be the operation channel of the BSS and may be used by the STA to establish a connection with the AP. In some representative embodiments, for example, Carrier Sense Multiple Access / Collision Avoidance (CSMA / CA) may be implemented in the 802.11 system. For CSMA / CA, the STA (e.g., each STA) (including the AP) may listen to the primary channel. If the primary channel is listened to / detected and / or determined to be busy by a particular STA, the particular STA may back off. One STA (e.g., only one station) may transmit in a given BSS at any given time.
[0087] High Throughput (HT) STAs may communicate using 40 MHz wide channels, e.g., by combining the primary 20 MHz channel with an adjacent or non - adjacent 20 MHz channel to form a 40 MHz wide channel.
[0088] Very High Throughput (VHT) STAs may support 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz wide channels. The 40 MHz channel and / or 80 MHz channel may be formed by combining consecutive 20 MHz channels. The 160 MHz channel may be formed by combining eight consecutive 20 MHz channels, or by combining two non - consecutive 80 MHz channels (which may be referred to as the 80 + 80 configuration). For the 80 + 80 configuration, after channel coding, the data may pass through a segment parser that may divide the data into two streams. The Inverse Fast Fourier Transform (IFFT) processing and time - domain processing may be performed separately on each stream. These streams may be mapped to two 80 MHz channels, and the data may be transmitted by the transmitting STA. At the receiver of the receiving STA, the operations for the 80 + 80 configuration described above may be reversed, and the combined data may be delivered to the Medium Access Control (MAC).
[0089] 802.11af and 802.11ah support operation modes below 1 GHz. Compared to those used in 802.11n and 802.11ac, the channel operation bandwidth and carriers are reduced in 802.11af and 802.11ah. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV white space (TVWS) spectrum, and 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. According to a representative embodiment, 802.11ah may support meter type control / machine type communication, such as MTC devices in a macro coverage area. The MTC devices may have certain capabilities, such as limited capabilities, including supporting (e.g., only supporting) certain bandwidths and / or limited bandwidths. The MTC devices may include a battery with a battery life higher than a threshold (e.g., to maintain a very long battery life).
[0090] A WLAN system that can support multiple channels and channel bandwidths (such as 802.11n, 802.11ac, 802.11af, and 802.11ah) includes a channel that can be designated as a primary channel. The primary channel may have a bandwidth equal to the maximum common operation bandwidth supported by all STAs in a BSS. The bandwidth of the primary channel may be set and / or restricted by an STA (which supports the minimum bandwidth operation mode) from all STAs operating in the BSS. In an example of 802.11ah, for an STA that supports (e.g., only supports) the 1 MHz mode (e.g., an MTC type device), the primary channel may be 1 MHz wide, even if the AP and other STAs in the BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operation modes. Carrier sensing and / or network allocation vector (NAV) setting may depend on the state of the primary channel. If the primary channel is busy, for example, because an STA (only supporting the 1 MHz operation mode) is sending to the AP, the entire available frequency band may be considered busy even if most of the frequency bands remain idle and may be available.
[0091] In the United States, the available frequency band for 802.11ah is 902 MHz to 928 MHz. In Korea, the available frequency band is 917.5 MHz to 923.5 MHz. In Japan, the available frequency band is 916.5 MHz to 927.5 MHz. The total available bandwidth for 802.11ah is 6 MHz to 26 MHz, depending on the country code.
[0092] Figure 1D FIG. is a system diagram illustrating RAN 113 and CN 115 according to an embodiment. As noted above, RAN 113 may employ NR radio technology to communicate with WTRUs 102a, 102b, 102c via air interface 116. RAN 113 may also communicate with CN 115.
[0093] RAN 113 may include gNBs 180a, 180b, 180c, but it should be understood that while being consistent with the embodiments, RAN 113 may include any number of gNBs. Each of gNBs 180a, 180b, 180c may include one or more transceivers for communicating with WTRUs 102a, 102b, 102c via air interface 116. In one embodiment, gNBs 180a, 180b, 180c may implement MIMO technology. For example, gNBs 180a, 108b may utilize beamforming to transmit signals to and / or receive signals from gNBs 180a, 180b, 180c. Thus, gNB 180a, for example, may use multiple antennas to transmit wireless signals to WTRU 102a and / or receive wireless signals from this WTRU. In an embodiment, gNBs 180a, 180b, 180c may implement carrier aggregation technology. For example, gNB 180a may transmit multiple component carriers to WTRU 102a (not shown). A subset of these component carriers may be on unlicensed spectrum, while the remaining component carriers may be on licensed spectrum. In an embodiment, gNBs 180a, 180b, 180c may implement coordinated multi-point (CoMP) technology. For example, WTRU 102a may receive coordinated transmissions from gNB 180a and gNB 180b (and / or gNB 180c).
[0094] WTRUs 102a, 102b, 102c may communicate with gNBs 180a, 180b, 180c using transmissions associated with a parameter set that can be extended. For example, the OFDM symbol interval and / or the OFDM subcarrier interval may vary for different transmissions, different cells, and / or different parts of the radio transmission spectrum. WTRUs 102a, 102b, 102c may communicate with gNBs 180a, 180b, 180c using subframes or transmission time intervals (TTIs) of various or extendable lengths (e.g., containing different numbers of OFDM symbols and / or having an absolute time length that continuously varies).
[0095] gNBs 180a, 180b, 180c can be configured to communicate with WTRUs 102a, 102b, 102c in a stand-alone configuration and / or a non-stand-alone configuration. In the stand-alone configuration, WTRUs 102a, 102b, 102c can communicate with gNBs 180a, 180b, 180c without accessing other RANs (e.g., such as evolved Node Bs 160a, 160b, 160c). In the stand-alone configuration, WTRUs 102a, 102b, 102c can use one or more of gNBs 180a, 180b, 180c as a mobility anchor. In the stand-alone configuration, WTRUs 102a, 102b, 102c can communicate with gNBs 180a, 180b, 180c using signals in an unlicensed band. In the non-stand-alone configuration, WTRUs 102a, 102b, 102c can communicate / connect with gNBs 180a, 180b, 180c while also communicating / connecting with another RAN (such as evolved Node Bs 160a, 160b, 160c). For example, WTRUs 102a, 102b, 102c can implement the DC principle to communicate with one or more of gNBs 180a, 180b, 180c and one or more of evolved Node Bs 160a, 160b, 160c substantially simultaneously. In the non-stand-alone configuration, evolved Node Bs 160a, 160b, 160c can act as the mobility anchor for WTRUs 102a, 102b, 102c, and gNBs 180a, 180b, 180c can provide additional coverage and / or throughput for serving WTRUs 102a, 102b, 102c.
[0096] Each of gNBs 180a, 180b, 180c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, scheduling of users in UL and / or DL, support for network slicing, dual connectivity, interworking between NR and E-UTRA, routing of user plane data towards user plane functions (UPFs) 184a, 184b, routing of control plane information towards access and mobility management functions (AMFs) 182a, 182b, etc. As Figure 1D shown, gNBs 180a, 180b, 180c can communicate with each other via the Xn interface.
[0097] Figure 1DThe illustrated CN 115 may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one session management function (SMF) 183a, 183b, and possibly data networks (DN) 185a, 185b. Although each of the foregoing elements is depicted as part of CN 115, it should be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.
[0098] The AMF 182a, 182b may be connected to one or more of the gNBs 180a, 180b, 180c in the RAN 113 via the N2 interface and may act as a control node. For example, the AMF 182a, 182b may be responsible for authenticating users of the WTRU 102a, 102b, 102c, support for network slicing (e.g., handling of different PDU sessions with different requirements), selection of a particular SMF 183a, 183b, management of the registration area, termination of NAS signaling, mobility management, etc. The AMF 182a, 182b may use network slicing in order to customize CN support for the WTRU 102a, 102b, 102c based on the type of service utilized by the WTRU 102a, 102b, 102c. For example, different network slices may be established for different use cases (such as services relying on ultra-reliable low-latency (URLLC) access, services relying on enhanced mobile broadband (eMBB) access, services for machine type communication (MTC) access, etc.). The AMF 162 may provide control plane functions for switching between the RAN 113 and other RANs (not shown) employing other radio technologies such as LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies such as WiFi.
[0099] The SMF 183a, 183b may be connected to the AMF 182a, 182b in the CN 115 via the N11 interface. The SMF 183a, 183b may also be connected to the UPF 184a, 184b in the CN 115 via the N4 interface. The SMF 183a, 183b may select and control the UPF 184a, 184b and configure the traffic routing through the UPF 184a, 184b. The SMF 183a, 183b may perform other functions such as managing and allocating UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, providing downlink data notifications, etc. The PDU session type may be IP-based, non-IP-based, Ethernet-based, etc.
[0100] UPF 184a and 184b can be connected to one or more of gNBs 180a, 180b, and 180c in RAN 113 via the N3 interface. The one or more gNBs can provide access to a packet switched network (such as the Internet 110) to WTRU 102a, 102b, and 102c to facilitate communication between WTRU 102a, 102b, and 102c and IP-enabled devices. UPF 184a and 184b can perform other functions, such as routing and forwarding packets, enforcing user plane policies, supporting multi-homed PDU sessions, handling user plane QoS, buffering downlink packets, providing mobility anchoring, etc.
[0101] CN 115 can facilitate communication with other networks. For example, CN 115 can include an IP gateway (such as an IP Multimedia Subsystem (IMS) server) that acts as an interface between CN 115 and the PSTN 108 or can communicate with the IP gateway. In addition, CN 115 can provide access to other networks 112 to WTRU 102a, 102b, and 102c. The other networks can include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, WTRU 102a, 102b, and 102c can be connected to DN 185a and 185b via UPF 184a and 184b through the N3 interface to UPF 184a and 184b and the N6 interface between UPF 184a and 184b and the local data network (DN) 185a and 185b.
[0102] In view of Figures 1A to 1D and Figures 1A to 1D In view of the corresponding descriptions of, one or more or all of the functions described herein for one or more of the following can be performed by one or more emulation devices (not shown): WTRU 102a - 102d, base stations 114a - 114b, evolved Node Bs 160a - 160c, MME 162, SGW 164, PGW 166, gNBs 180a - 180c, AMF 182a - 182b, UPF 184a - 184b, SMF 183a - 183b, DN 185a - 185b, and / or any other device described herein. The emulation device can be one or more devices configured to mimic one or more or all of the functions described herein. For example, the emulation device can be used to test other devices and / or simulate network and / or WTRU functions.
[0103] A simulation device can be designed to implement one or more tests of other devices in a laboratory environment and / or in an operator network environment. For example, one or more simulation devices can perform one or more functions or all functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices within the communication network. One or more simulation devices can perform one or more functions or all functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. The simulation device can be directly coupled to another device for testing purposes and / or can perform tests using over-the-air wireless communication.
[0104] One or more simulation devices can perform one or more (including all) functions without being implemented / deployed as part of a wired and / or wireless communication network. For example, the simulation device can be used in a test laboratory and / or in a test scenario in a non-deployed (e.g., test) wired and / or wireless communication network to implement tests of one or more components. One or more simulation devices can be test equipment. Direct RF coupling and / or wireless communication via an RF circuit system (e.g., which can include one or more antennas) can be used by the simulation device to transmit and / or receive data.
[0105] This application describes multiple aspects, including tools, features, examples, models, methods, etc. Many of these aspects are specifically described and are generally described in a way that may sound restrictive, at least to illustrate individual features. However, this is for the purpose of clarity of description and does not limit the application or scope of those aspects. In fact, all different aspects can be combined and interchanged to provide further aspects. In addition, these aspects can also be combined and interchanged with aspects described in earlier applications.
[0106] The aspects described and contemplated in this application can be implemented in many different forms. Figures 5 to Figure 25 Some examples can be provided, but other examples can be contemplated. The discussion of Figures 5 to Figure 25 does not limit the breadth of the implementation. At least one of the aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting the generated or encoded bitstream. These and other aspects can be implemented as methods, devices, computer-readable storage media having instructions stored thereon for encoding or decoding video data according to any of the described methods, and / or computer-readable storage media having a bitstream generated according to any of the described methods stored thereon.
[0107] In this application, the terms "reconstructed" and "decoded" can be used interchangeably, the terms "pixel" and "sample" can be used interchangeably, and the terms "image", "picture", and "frame" can be used interchangeably.
[0108] This document describes various methods, and each method includes one or more steps or actions for implementing the described method. Unless the correct operation of the method requires steps or actions in a specific order, the order and / or use of specific steps and / or actions can be modified or combined. Additionally, terms such as "first", "second", etc. may be used in various examples to modify elements, components, steps, operations, etc., such as, for example, "first decoding" and "second decoding". Unless specifically required, the use of these terms does not imply an ordering of the modified operations. Thus, in this example, the first decoding does not need to be performed before the second decoding and can occur, for example, before, during, or in a time period overlapping with the second decoding.
[0109] The various methods and other aspects described in this application can be used to modify modules of the video encoder 200 and decoder 300 as shown in Figure 2 and Figure 3 e.g., the decoding module. Additionally, the subject matter disclosed herein can be applied to (for example) any type, format, or version of video coding, whether described in a standard or recommendation, whether pre - existing or future - developed, and to extensions of any such standards and recommendations. Unless otherwise stated or technically precluded, the aspects described in this application can be used alone or in combination.
[0110] In the examples described in this application, various numerical values are used, such as bits, bit depth, etc. These and other specific values are for the purpose of describing the examples, and the described aspects are not limited to these specific values.
[0111] Figure 2 is a diagram showing an example video encoder. Variations of the example encoder 200 are envisioned, but for clarity, encoder 200 is described below without describing all the expected variations.
[0112] Before being encoded, a video sequence can undergo pre - encoding processing (201), e.g., applying a color transformation to the input color picture (e.g., a conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing remapping of the input picture components in order to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata can be associated with the pre - processing and appended to the bitstream.
[0113] In encoder 200, as described below, pictures are encoded by encoder elements. The pictures to be encoded are partitioned (202) and processed in units such as coding units (CUs) for example. Each unit is encoded using, for example, an intra or inter mode. When the unit is encoded in the intra mode, it performs intra prediction (260). In the inter mode, motion estimation (275) and compensation (270) are performed. The encoder decides (205) which one of the intra mode or the inter mode to use to encode the unit, and indicates the intra / inter decision by, for example, a prediction mode flag. For example, the prediction residual is calculated by subtracting (210) the prediction block from the original image block.
[0114] Then the prediction residual is transformed (225) and quantized (230). The quantized transform coefficients, motion vectors, and other syntax elements are entropy encoded (245) to output a bitstream. The encoder may skip the transform and apply quantization directly to the non-transformed residual signal. The encoder may bypass both the transform and quantization, i.e., the residual is directly encoded without applying the transform or quantization process.
[0115] The encoder decodes the encoded blocks to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residual. The decoded prediction residual and the prediction block are combined (255) to reconstruct the image block. A loop filter (265) is applied to the reconstructed picture to perform, for example, deblocking / SAO (sample adaptive offset) filtering to reduce encoding artifacts. The filtered image is stored at the reference picture buffer (280).
[0116] Figure 3 is a diagram showing an example of a video decoder. In example decoder 300, the bitstream is decoded by decoder elements, as described below. Video decoder 300 generally performs a decoding channel that is reciprocal to the encoding channel as Figure 2 described. Encoder 200 generally also performs video decoding as part of encoding video data.
[0117] Specifically, the input to the decoder includes a video bitstream, which may be generated by video encoder 200. First, the bitstream is entropy decoded (330) to obtain transform coefficients, motion vectors, and other encoded information. The picture partitioning information indicates how the picture is partitioned. Thus, the decoder can partition (335) the picture according to the decoded picture partitioning information. The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residual. The decoded prediction residual and the prediction block are combined (355) to reconstruct the image block. The prediction block can be obtained (370) from intra prediction (360) or motion compensated prediction (i.e., inter prediction) (375). A loop filter (365) is applied to the reconstructed image. The filtered image is stored at the reference picture buffer (380).
[0118] The decoded picture may further undergo post - decoding processing (385), such as an inverse color transformation (e.g., a conversion from YCbCr 4:2:0 to RGB 4:4:4) or performing an inverse remapping that is the inverse of the remapping process performed in the pre - encoding processing (201). The post - decoding processing may use metadata derived in the pre - encoding processing and signaled in the bitstream. In an example, the decoded image (e.g., after applying the in - loop filter (365) and / or after the post - decoding processing (385) if post - decoding processing is used) may be sent to a display device for presentation to the user.
[0119] Figure 4 FIG. is an example of a system in which various aspects and examples described herein may be implemented. System 400 may be embodied as a device including various components described below and is configured to perform one or more aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smart phones, tablet computers, digital multimedia set - top boxes, digital television receivers, personal video recording systems, connected household appliances, and servers. The elements of system 400 may be embodied singly or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one example, the processing and encoder / decoder elements of system 400 are distributed over multiple ICs and / or discrete components. In various examples, system 400 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various examples, system 400 is configured to implement one or more aspects described in this document.
[0120] System 400 includes at least one processor 410 that is configured to execute instructions loaded therein for implementing, for example, the various aspects described in this document. Processor 410 may include embedded memory, input - output interfaces, and various other circuits known in the art. System 400 includes at least one memory 420 (e.g., a volatile memory device and / or a non - volatile memory device). System 400 includes a storage device 440, which may include non - volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read - only memory (EEPROM), read - only memory (ROM), programmable read - only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non - limiting example, storage device 440 may include internal storage devices, attached storage devices (including removable and non - removable storage devices), and / or network - accessible storage devices.
[0121] System 400 includes an encoder / decoder module 430 that is configured to, for example, process data to provide encoded video or decoded video, and the encoder / decoder module 430 can include its own processor and memory. The encoder / decoder module 430 represents a module that can be included in a device to perform encoding and / or decoding functions. As is known, a device can include one or both of an encoding and a decoding module. Additionally, the encoder / decoder module 430 can be implemented as a separate element of system 400 or can be incorporated within the processor 410 as a combination of hardware and software known to those skilled in the art.
[0122] Program code to be loaded onto the processor 410 or the encoder / decoder 430 to perform the various aspects described in this document can be stored in the storage device 440 and subsequently loaded onto the memory 420 for execution by the processor 410. According to various examples, one or more of the processor 410, the memory 420, the storage device 440, and the encoder / decoder module 430 can store one or more of the various items during the execution of the processes described in this document. Such stored items can include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operation logic.
[0123] In some examples, the memory internal to the processor 410 and / or the encoder / decoder module 430 is used to store instructions and provide a working memory for the processing required during encoding or decoding. However, in other examples, memory external to the processing device (e.g., the processing device can be the processor 410 or the encoder / decoder module 430) is used for one or more of these functions. The external memory can be the memory 420 and / or the storage device 440, such as dynamic volatile memory and / or non-volatile flash memory. In several examples, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one example, fast external dynamic volatile memory such as RAM is used as a working memory for video encoding and decoding operations.
[0124] As shown in block 445, input to the elements of system 400 can be provided through various input devices. Such input devices include, but are not limited to, (i) a radio frequency (RF) section that receives, for example, RF signals transmitted over the air by a broadcaster, (ii) component (COMP) input terminals (or a set of COMP input terminals), (iii) universal serial bus (USB) input terminals, and / or (iv) high-definition multimedia interface (HDMI) input terminals. Figure 4 Other examples not shown include composite video.
[0125] In various examples, the input device of block 445 has corresponding input processing elements associated therewith as known in the art. For example, the RF section may be associated with elements adapted to: (i) select a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a frequency band), (ii) down-convert the selected signal, (iii) again band-limit to a narrower frequency band to select a signal frequency band that may be referred to as a channel in some examples, (iv) demodulate the down-converted and band-limited signal, (v) perform error correction, and / or (vi) demultiplex to select a desired data packet stream. The RF section of various examples includes one or more elements for performing these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, filters, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF section may include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or a near-baseband frequency) or down-converting to baseband. In one set-top box example, the RF section and its associated input processing elements receive an RF signal transmitted over a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and again filtering to a desired frequency band. Various examples rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as inserting an amplifier and an analog-to-digital converter. In various examples, the RF section includes an antenna.
[0126] The USB and / or HDMI terminals may include corresponding interface processors for connecting system 400 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of the input processing, such as Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or within processor 410 as needed. Similarly, aspects of the USB or HDMI interface processing may be implemented within a separate interface IC or within processor 410 as needed. The demodulated, error-corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 410 and encoder / decoder 430, which operate in conjunction with memory and storage elements to process the data stream as needed for presentation on an output device.
[0127] The various elements of system 400 may be disposed within an integrated housing. Within the integrated housing, the various elements may be interconnected using a suitable connection arrangement 425 (e.g., internal buses known in the art, including an inter-integrated circuit (I2C) bus, wiring, and printed circuit boards) and data may be transmitted therebetween.
[0128] System 400 includes a communication interface 450 that enables communication with other devices via a communication channel 460. The communication interface 450 may include, but is not limited to, a transceiver configured to send and receive data over the communication channel 460. The communication interface 450 may include, but is not limited to, a modem or a network card, and the communication channel 460 may be implemented, for example, within a wired and / or wireless medium.
[0129] In various examples, a wireless network (such as a Wi-Fi network, e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers)) is used to stream or otherwise provide data to system 400. Wi-Fi signals for these examples are received via a communication channel 460 and a communication interface 450 adapted for Wi-Fi communication. The communication channel 460 for these examples is typically connected to an access point or a router that provides access to an external network (including the Internet) to allow streaming applications and other over-the-top communications. Other examples use a set-top box to provide streamed data to system 400, and the set-top box delivers data via an HDMI connection of the input block 445. Other examples use an RF connection of the input block 445 to provide streamed data to system 400. As indicated above, various examples provide data in a non-streaming manner. Additionally, various examples use a wireless network other than Wi-Fi, such as a cellular network or network.
[0130] System 400 may provide output signals to various output devices, including a display 475, speakers 485, and other peripheral devices 495. The display 475 for various examples includes, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 475 may be used in a television, a tablet, a laptop computer, a mobile phone (cellular phone), or other devices. The display 475 may also be integrated with other components (e.g., as in a smart phone) or be separate (e.g., an external monitor for a laptop computer). In various examples, other peripheral devices 495 include one or more of a standalone digital video disc (or digital versatile disc) (DVD, both terms), a disc player, a stereo system, and / or a lighting system. Various examples use one or more peripheral devices 495 that perform functions based on the output of system 400. For example, a disc player performs the function of playing the output of system 400.
[0131] In various examples, signaling using communication protocols such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable device - to - device control with or without user intervention conveys control signals between system 400 and display 475, speaker 485, or other peripheral devices 495. The output devices can be communicatively coupled to system 400 via dedicated connections through corresponding interfaces 470, 480, and 490. Alternatively, the output devices can be connected to system 400 via communication interface 450 using communication channel 460. Display 475 and speaker 485 can be integrated in a single unit with other components of system 400 in an electronic device (e.g., a television). In various examples, display interface 470 includes a display driver, such as, for example, a timing controller (T_CON) chip.
[0132] For example, if the RF portion of input 445 is part of a separate set - top box, display 475 and speaker 485 can optionally be separated from one or more other components. In various examples where display 475 and speaker 485 are external components, output signals can be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.
[0133] Examples can be executed by computer software implemented by processor 410 or by hardware or by a combination of hardware and software. As a non - limiting example, examples can be implemented by one or more integrated circuits. Memory 420 can be of any type suitable for the technical environment and can be implemented using any appropriate data storage technology, as non - limiting examples, such as optical memory devices, magnetic memory devices, semiconductor - based memory devices, fixed memory, and removable memory. Processor 410 can be of any type suitable for the technical environment and can encompass, as non - limiting examples, one or more of a microprocessor, a general - purpose computer, a special - purpose computer, and a processor based on a multi - core architecture.
[0134] Various implementations involve decoding. As used in this application, "decoding" can cover, for example, all or part of the processing performed on a received encoded sequence to produce a final output suitable for display. In various examples, such a process includes one or more of the processes typically performed by a decoder, for example, entropy decoding, inverse quantization, inverse transform, and differential decoding. In various examples, such a process also or alternatively includes processes performed by the decoders of the various embodiments described in this application, for example, determining that a current block is decoded in a non - directional intra - prediction mode; deriving a directional intra - prediction mode corresponding to the non - directional intra - prediction mode, where the derived directional intra - prediction mode indicates a derived intra - prediction direction; and decoding the current block at least in part based on the derived directional intra - prediction mode, etc.
[0135] As another example, in one example, "decoding" refers only to entropy decoding, in another example, "decoding" refers only to differential decoding, and in another example, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally refer to a broader decoding process will be clear based on the context of the specific description and is considered well understood by those skilled in the art.
[0136] Various embodiments relate to encoding. In a manner similar to the discussion of "decoding" above, "encoding" as used in this application can cover, for example, all or part of the processing performed on an input video sequence to produce an encoded bitstream. In various examples, such a process includes one or more processes typically performed by an encoder, such as, partitioning, differential encoding, transformation, quantization, and entropy encoding. In various examples, such a process also or alternatively includes processes performed by the encoders of the various embodiments described in this application, such as, identifying a non-directional intra prediction mode for encoding a current block; deriving a directional intra prediction mode corresponding to the non-directional intra prediction mode, where the derived directional intra prediction mode includes a derived intra prediction direction; and encoding the current block at least in part based on the derived directional intra prediction mode, etc.
[0137] As another example, in one example, "encoding" refers only to entropy encoding, in another example, "encoding" refers only to differential encoding, and in another example, "encoding" refers to a combination of differential encoding and entropy encoding. Whether the phrase "encoding process" is intended to specifically refer to a subset of operations or generally refer to a broader encoding process will be clear based on the context of the specific description and is considered well understood by those skilled in the art.
[0138] When the drawings are presented as flowcharts, it should be understood that they also provide block diagrams of the corresponding devices. Similarly, when the drawings are presented as block diagrams, it should be understood that they also provide flowcharts of the corresponding methods / processes.
[0139] The embodiments and aspects described herein can be implemented in, for example, a method or process, a device, a software program, a data stream, or a signal. Even if discussed only in the context of a single implementation form (e.g., only as a method), the implementation of the features discussed can also be implemented in other forms (e.g., a device or a program). The device can be implemented in, for example, appropriate hardware, software, and firmware. The method can be implemented in, for example, a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes communication devices, such as, for example, a computer, a cellular phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate information communication between end users.
[0140] References to "an example" or "examples" or "an implementation" or "implementations" and other variations thereof mean that the particular features, structures, characteristics, etc. described in connection with that example are included in at least one example. Thus, the appearances of the phrases "in an example" or "in examples" or "in an implementation" or "in implementations" and any other variations thereof throughout this application do not necessarily all refer to the same example.
[0141] In addition, this application may refer to "determining" various pieces of information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from a memory. Obtaining may include receiving, retrieving, constructing, generating, and / or determining.
[0142] In addition, this application may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0143] In addition, this application may refer to "receiving" various pieces of information. Like "accessing", receiving is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from a memory). In addition, during operations such as, for example, storing information, processing information, sending information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information, "receiving" is typically involved in one way or another.
[0144] It should be understood that, for example, in the case of "A / B", "A and / or B", and "at least one of A and B", the use of any one of the following, namely " / ", "and / or", and "at least one of", is intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As another example, in the case of "A, B, and / or C" and "at least one of A, B, and C", such phrasing is intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or only the first and second-listed options (A and B), or only the first and third-listed options (A and C), or only the second and third-listed options (B and C), or the selection of all three options (A and B and C). As will be clear to those of ordinary skill in the art and related fields, this can be extended to as many items as are listed.
[0145] In addition, as used herein, the term "signal" particularly refers to indicating something to a corresponding decoder. In this way, in an example, the same parameters are used at both the encoder side and the decoder side. Thus, for example, the encoder can send (explicit signaling) specific parameters to the decoder so that the decoder can use the same specific parameters. Conversely, if the decoder already has specific parameters and other parameters, signaling can be used without sending (implicit signaling) to simply allow the decoder to know and select the specific parameters. By avoiding the transmission of any actual functionality, bit savings are achieved in various examples. It should be understood that signaling can be done in various ways. For example, in various examples, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder. Although the previous discussion related to the verb form of the term "signal", the term "signal" can (e.g., also can) be used as a noun in this document.
[0146] It will be apparent to those of ordinary skill in the art that implementations can generate various signals that are formatted to carry information that can be stored or transmitted, for example. The information can include, for example, instructions for performing a method or data generated by one of the described embodiments. For example, a signal can be formatted to carry the bitstream of the described example. Such a signal can be formatted as, for example, an electromagnetic wave (e.g., using the radio frequency part of the spectrum) or a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. As is known, signals can be transmitted over various different wired or wireless links. The signal can be stored on a processor-readable medium or accessed or received from a processor-readable medium.
[0147] This document describes numerous examples. The features of the examples can be provided individually or in any combination across various claim categories and types. Additionally, an example can include one or more of the features, devices, or aspects described herein individually or in any combination across various claim categories and types. For example, the features described herein can be implemented in a bitstream or signal that includes information generated as described herein. This information can allow a decoder to decode the bitstream, encoder, bitstream, and / or decoder according to any of the described embodiments. For example, the features described herein can be implemented by creating and / or sending and / or receiving and / or decoding a bitstream or signal. For example, the features described herein can be implemented with a method, process, device, medium storing instructions, medium storing data, or signal. For example, the features described herein can be implemented by a TV, set-top box, cellular phone, tablet computer, or other electronic device that performs decoding. The TV, set-top box, cellular phone, tablet computer, or other electronic device can display (e.g., using a monitor, screen, or other type of display) the resulting image (e.g., an image reconstructed from residuals of a video bitstream). The TV, set-top box, cellular phone, tablet computer, or other electronic device can receive a signal that includes an encoded image and perform decoding.
[0148] These examples can be executed by a device having at least one processor. The device can be an encoder or a decoder. These examples can be executed by a computer program product stored on a non-transitory computer-readable medium and including program code instructions. These examples can be executed by a computer program including program code instructions. These examples can be executed by a bitstream including information representing a decoding block.
[0149] Intra-sample prediction can include predicting pixels of a target coding unit (CU) based on a set of reference samples. The prediction modes can include planar and DC prediction modes, which can be used to predict smooth and gradually changing regions. Angular prediction modes (e.g., angles defined from 45 degrees to -135 degrees in a clockwise direction) can be used to capture different directional structures. For square blocks, directional prediction modes (e.g., 33 directional modes for square blocks) can be used, which can be indexed (e.g., indexed from 2 to 34). The prediction modes can correspond to different prediction directions as Figure 5A illustrated in. The angular prediction modes can correspond to angular directions (e.g., 65 angular prediction modes can correspond to 33 angular directions), and the angular directions (e.g., another 32 angular directions) can correspond to intermediate directions between adjacent pairs, as Figure 5B illustrated in.
[0150] Figure 5AShows example intra prediction directions. The numbers may represent prediction mode indices associated with the corresponding directions. Modes 2 to 17 may indicate horizontal prediction (H - 26 to H + 32), and modes 18 to 34 may indicate vertical prediction (V - 32 to V + 32). Figure 5B Illustrates intra prediction for a square block (e.g., for a square block). Modes less than 34 may indicate horizontal prediction. Modes greater than 34 may indicate vertical prediction. Figure 5C Illustrates available (e.g., all available) intra prediction directions. The dashed lines may indicate wide - angle intra prediction mode (WAIP). Figure 5C The indices - 1 to - 14 illustrated in may be remapped to 1 to - 12 (e.g., such that the angular mode indices are consecutive). Modes - 15 (e.g., remapped to - 13) and 81 may not exist in Figure 5C because the block size (e.g., no allowed block size) may not use modes - 15 (e.g., remapped to - 13) and 81. Modes - 15 (e.g., remapped to - 13) and 81 may be handled by the reference code.
[0151] Template - based intra mode derivation (TIMD) may be performed to derive a prediction mode for a decoding block. Intra prediction mode derivation via TIMD may be applied (e.g., in the same way) on the encoder and decoder sides for a given luminance (such as Figure 6 the CB 603 shown in (a) of ). Each (e.g., every) intra prediction mode in the most - probable mode (MPM) list of the luminance CB (e.g., supplemented with a default mode) may be used to compute the prediction of the templates (600 and 601) of the luminance CB from the decoded reference samples of the template (602). The sum of absolute transform differences (SATD) between the prediction of the luminance CB and the template may be computed. One or (e.g., two) intra prediction modes with the minimum (e.g., smallest) SATD may be selected as the TIMD mode. The set of directional intra prediction modes (e.g., for TIMD) may be extended (e.g., from 65 to 129), for example, by inserting directions between each solid line and adjacent dashed - line arrow in Figure 5B The set of possible intra prediction modes derived via TIMD may aggregate modes (e.g., 131 modes). One or more (e.g., two) intra prediction modes available from a first - pass test involving the MPM list may be retained and supplemented with the default mode. For each retained intra prediction mode that is not planar or DC, the (e.g., two) closest extended directional intra prediction modes may be tested. The SATD between the prediction computed using the closest extended directional intra prediction modes and the template of the luminance CB may be computed. The intra prediction mode with the minimum (e.g., smallest) SATD may be selected as the TIMD mode.
[0152] Figure 6An example template of the current luminance CB and example decoded reference samples of the template used in TIMD are shown. In Figure 6 (a) of t , the template of the luminance CB does not extend beyond the boundaries of the current frame. The current W×H luminance CB 603 can be surrounded by its fully available template, which consists of the w t ×H part at its left side 600 and the W×h t part at its upper side 601. During the TIMD derivation step, the tested intra prediction mode can predict the template of the current luminance CB from the set of 1 + 2w t + 2W + 2h t + 2H decoded reference samples 602 at 602 of the template. If W ≤ 8, w t can be equal to 2; otherwise w t can be equal to 4. If H ≤ 8, h t can be equal to 2; otherwise h
[0153] Figure 6 (b) of Figure 6 and Figure 6 (c) of t show examples where at least one (e.g., one) part of the template of the luminance CB extends beyond the boundaries of the current frame. In t (b) of Figure 6 , the current W×H luminance CB 603 can be surrounded by its template, and the W×h t part at its upper side 601 is available. During the TIMD derivation step, the tested intra prediction mode can predict the template of the current luminance CB from the set of 1 + 2W + 2h t + 2H decoded reference samples of the template. In
[0154] (c) of
[0155] , the current W×H luminance CB 603 can be surrounded by its template, and only the w t ×H part at its left side 600 is available. During the TIMD derivation step, the tested intra prediction mode can predict the template of the current luminance CB from the set of 1 + 2w t + 2W + 2H decoded reference samples of the template.
[0154] The current luminance CB can be predicted via TIMD, e.g., by fusing (e.g., two) predictions of the luminance CB calculated based on (e.g., two) TIMD modes generated by (e.g., two) test passes with weights (e.g., after applying position-dependent prediction combination (PDPC)). The weights used can depend on the prediction SATD of (e.g., two) TIMD modes.
[0155] An executable decoder-side intra mode derivation (DIMD) can be performed to derive an intra prediction mode for a coding block. For example, two intra modes can be derived from the reconstructed neighboring samples. Two predictors can be combined with a planar mode predictor with weights derived from gradients. The division operation in weight derivation can be performed using the same lookup table (LUT)-based integerization scheme used by the cross-component linear model (CCLM). For example, the division operation in azimuth calculation
[0156] Orient = G y / G x
[0157] can be calculated by the following LUT-based scheme:
[0158] x = Floor(Log2(Gx))
[0159] normDiff = ((Gx << 4) >> x) & 15
[0160] x += (3 + (normDiff != 0 ? 1 : 0))
[0161] Orient = (Gy * (DivSigTable[normDiff] | 8) + (1 << (x - 1))) >> x, where DivSigTable
[16] = {0, 7, 6, 5, 5, 4, 4, 3, 3, 2, 2, 1, 1, 1, 1, 0}.
[0162] Deriving the intra mode can be included in the main list of the intra MPM list. The DIMD process can be performed before constructing the MPM list. The main derived intra mode of the DIMD block can be stored with the block and can be used for the construction of the MPM list of neighboring blocks.
[0163] Figure 7 Neighboring reconstructed samples for DIMD chrominance modes are illustrated. The DIMD chrominance mode can use DIMD derivation to derive the chrominance intra prediction mode of the current block based on neighboring reconstructed Y, Cb, and Cr samples in the second neighboring row and column, as Figure 7 shown. Horizontal and vertical gradients can be calculated for the collocated reconstructed luminance samples (e.g., each collocated reconstructed luminance sample) and the reconstructed Cb and Cr samples of the current chrominance block to construct a histogram of oriented gradients (HoG). The intra prediction mode with the maximum histogram magnitude value can be used to perform the chrominance intra prediction of the current chrominance block.
[0164] When the intra prediction mode derived from the DIMD chroma mode is the same as the intra prediction mode derived from the direct mode (DM), the intra prediction mode with the second largest histogram amplitude value can be used as the DIMD chroma mode. A CU-level indication (e.g., a flag) can be signaled to indicate whether the DIMD chroma mode is applied.
[0165] Figure 8 An example of the matrix weighted intra prediction (MIP) process is illustrated. To predict samples of a rectangular block with width W and height H, MIP can take as input a row of H reconstructed neighboring boundary samples to the left of the block and a row of W reconstructed neighboring boundary samples above the block. If the reconstructed samples are not available, then they can be generated in the same or a similar manner as in other intra prediction examples (e.g., conventional intra prediction). The generation of the prediction signal can be based on at least the following three steps: averaging; matrix-vector multiplication; and linear interpolation (e.g., as Figure 8 shown).
[0166] CCLM can be performed to predict coded blocks. The CCLM prediction mode can be used in video coding, e.g., to reduce cross-component redundancy. Chroma samples can be predicted (e.g.) based on reconstructed luma samples (e.g., for the same CU) by using a linear model. For example, a linear model can be constructed according to Equation 1:
[0167] pred C (i,j) = α·rec L ′(i,j)+β Equation 1
[0168] As shown in the example in Equation 1, pred C (i,j) can represent the predicted chroma samples in the CU. As shown in the example in Equation 1, rec L '(i,j) can represent the downsampled reconstructed luma samples of (e.g., the same) CU.
[0169] CCLM parameters (e.g., α and β) can be derived (e.g.) based on / using (e.g., at most four) neighboring chroma samples and the corresponding downsampled luma samples. For illustrative examples, assume the current chroma block size is W×H. In some examples, W” and H' can be set according to the following logic:
[0170] W' = W, H′ = H, e.g., if / when the LM mode is applied;
[0171] W' = W + H, e.g., if / when the LM-A mode is applied; and / or
[0172] H' = H + W, e.g., if / when the LM-L mode is applied.
[0173] In a further discussion of the example, the above adjacent positions can be expressed as S[0, -1]…S[W'-1, -1] and the left adjacent positions can be expressed as S[-1, 0]…S[-1, H'-1]. Four samples can be selected as follows (e.g., according to the example logic):
[0174] S[W’ / 4, -1], S[3*W’ / 4, -1], S[-1, H’ / 4], S[-1, 3*H’ / 4], for example, if / when the LM mode is applied and both the upper and left adjacent samples are available;
[0175] S[W’ / 8, -1], S[3*W’ / 8, -1], S[5*W’ / 8, -1], S[7*W’ / 8, -1], for example, if / when the LM-A mode is applied or only the upper adjacent samples are available; and / or
[0176] S[-1, H’ / 8], S[-1, 3*H’ / 8], S[-1, 5*H’ / 8], S[-1, 7*H’ / 8], for example, if / when the LM-L mode is applied or only the left adjacent samples are available.
[0177] In the example, the four adjacent lightness samples at the selected positions can be downsampled and compared (e.g., four times) to find (e.g., two) larger values (e.g., denoted as x 0 A and x 1 A ) and (e.g., two) smaller values (e.g., denoted as x 0 B and x 1 B ). The corresponding chroma sample values can be denoted as y 0 A 、y 1 A 、y 0 B and y 1 B . In the example, x A 、x B 、y A and y B can be derived, for example, according to equations 2a - 2d:
[0178] X a =(x 0 A +x 1 A +1)>>1 Equation 2a
[0179] X b =(x 0 B +x1 B +1) >> 1 Equation 2b
[0180] Y a = (y 0 A + y 1 A +1) >> 1 Equation 2c
[0181] Y b = (y 0 B + y 1 B +1) >> 1 Equation 2d
[0182] The linear model parameters α and β can be determined, for example, according to Equation 3 and Equation 4:
[0183]
[0184] β = Y b - α · X b Equation 4
[0185] Figure 9 Illustrates an example of the positions of the left and upper samples and the samples of the current block involved in the CCLM mode. Figure 9 Shows an example of the positions of the samples for deriving the linear model parameters α and β.
[0186] CCLM can be extended by adding three multi-model LM (MMLM) modes. In each MMLM mode, the reconstructed neighboring samples can be classified into two categories using a threshold. The threshold can be the average of the luminance-reconstructed neighboring samples. The least mean square (LMS) method can be used to derive the linear model for each category. For the CCLM mode, the LMS method can be used to derive the linear model. Slope adjustment can be applied to CCLM and MMLM predictions. The adjustment can involve tilting the linear function (e.g., which maps the lightness value to the chroma value) relative to the center point determined by the average lightness value of the reference samples.
[0187] CCLM slope adjustment can be implemented. CCLM can use a model with one or more (e.g., two) parameters to map the lightness value to the chroma value. The slope parameter "a" and the bias parameter "b" can define the mapping, for example, according to Equation 5:
[0188] chromaVal = a * lumaVal + b Equation 5
[0189] An adjustment "u" to the slope parameter can be signaled to update the model, for example, according to Equation 6:
[0190] chromaVal = a' * lumaVal + b' Equation 6
[0191] For example, the updated slope parameter can be determined according to Equations 7a and 7b:
[0192] a' = a + u Equation 7a
[0193] b' = b - u * y r Equation 7b
[0194] The mapping function can be tilted or rotated, for example based on a selection, around a point with a luminance value y r The average of the reference luminance samples used in model creation can be used as y r , for example, to provide a (e.g., meaningful) modification to the model.
[0195] Figure 10A and Figure 10B Examples illustrate the effect of the slope adjustment parameter "u". Figure 10A Examples illustrate a model created for CCLM without an updated slope parameter. Figure 10B Examples illustrate a model created for CCLM with an updated slope parameter.
[0196] This document provides features associated with the Convolutional Cross-Component Mode (CCCM). The reconstructed luminance samples to be used for chroma prediction can be filtered. The convolutional 7-tap filter can include a 5-tap plus sign-shaped spatial component, a non-linear term, and a bias term, as Figure 11 shown. The input of the spatial 5-tap component of the filter can include the center (C) luminance sample (e.g., which can be juxtaposed with the chroma sample to be predicted) and the upper / north (N), lower / south (S), left / west (W), and right / east (E) neighbors, as shown.
[0197] The non-linear term P can represent the power of two of the center luminance sample C and be scaled to the sample value range of the content:
[0198] P = (C * C + midVal) >> bitDepth Equation 8
[0199] For 10-bit content, P can be calculated as:
[0200] P = (C * C + 512) >> 10 Equation 9
[0201] The bias term B can represent a scalar offset between the input and the output (e.g., similar to the offset term in CCLM), and can be set to the intermediate chroma value (e.g., 512 for 10-bit content).
[0202] The output of the filter can be calculated as the convolution between the filter coefficients c i and the input values and trimmed to the range of valid chroma samples, as shown in Equation 10:
[0203] predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P + c6B Equation 10
[0204] Intra Block Copy (IBC) can improve the decoding efficiency of screen content material. The IBC mode can be a block-level decoding mode. Block matching (BM) can be performed at the encoder to find the optimal block vector (or motion vector) for a CU (e.g., each CU). The block vector can be used to indicate the displacement from the current block to a reference block (e.g., which has been reconstructed within the current picture). The luma block vector of a CU decoded by IBC can be in integer precision. The chroma block vector can be rounded to integer precision. When combined with AMVR, the IBC mode can switch between 1-pixel and 4-pixel motion vector precisions. A CU decoded by IBC can be regarded as a third prediction mode (e.g., different from the intra or inter prediction mode). The IBC mode can be applicable to CUs with both width and height less than or equal to 64 luma samples.
[0205] The reference region for IBC can be extended (e.g., extended to two CTU rows above). Figure 12 The reference region for decoding a Coding Tree Unit (CTU) (m, n) is illustrated. Figure 12 An example reference region for IBC when the CTU (m,n) is decoded is illustrated. The block labeled "m,n" represents the current CTU; other shaded blocks represent the reference region; and the white blocks represent invalid reference regions. For the CTU (m,n) to be decoded, the reference region can include CTUs with indices (M-2,N-2)…(W,N-2), (0,N-1)…(W,N-1), (0,n)…(m,n), where W represents the maximum horizontal index within the current tile, slice, or picture. When the CTU size is 256, the reference region can be limited to one CTU row above. This can ensure that for CTU sizes of 128 or 256, IBC does not use additional memory. The per-sample block vector search (sometimes referred to as local search) range can be limited to horizontal [-(C<<1),C>>2] and vertical [-C,C>>2] to accommodate the reference region extension, where C represents the CTU size.
[0206] This document provides examples of Intra-frame Template Matching Prediction (IntraTMP). IntraTMP is an intra-frame prediction mode that can copy the best prediction block from the reconstructed part of the current frame, whose L-shaped template matches the current template. For a predefined search range, the encoder can search for the template most similar to the current template in the reconstructed part of the current frame. For the predefined search range, the encoder can use the corresponding block as the prediction block. The encoder can signal the use of this mode and the same prediction operation can be performed on the decoder side.
[0207] Figure 13 Examples of the intra-frame template matching search area are illustrated. A prediction signal can be generated by matching the L-shaped causal neighbors of the current block with another block in the predefined search area in Figure 13 including:
[0208] R1: the current CTU
[0209] R2: the top-left CTU
[0210] R3: the upper CTU
[0211] R4: the left CTU
[0212] The Sum of Absolute Differences (SAD) can be used as a cost function. Within the region (e.g., within each region), the decoder can search for the template with the minimum SAD relative to the current SAD and use its corresponding block as the prediction block. The size of the region (SearchRange_w, SearchRange_h) can be set to be proportional to the block size (BlkW, BlkH) so that each pixel has a fixed number of SAD comparisons.
[0213] That is:
[0214] SearchRange_w = a * BlkW Equation 11
[0215] SearchRange_h = a * BlkH Equation 12
[0216] where 'a' is a constant that controls the gain / complexity trade-off. For example, 'a' can be equal to 5.
[0217] The intra-frame template matching tool can be enabled for CUs with sizes less than or equal to 64 in width and height. This maximum CU size for intra-frame template matching can be configurable. The intra-frame template matching prediction mode can be signaled at the CU level with a dedicated flag. If DIMD is not enabled (e.g., DIMD = 0), then the intra-frame template matching prediction mode can be signaled at the CU level with a dedicated flag. Although intra-frame template matching examples are described in this document, the examples in this document can also be applied to inter-frame template matching.
[0218] The palette mode can be used for encoding / decoding coding blocks. In some examples, the palette mode can be used for screen content decoding in chrominance formats supported in the 4:4:4 profile (i.e., 4:4:4, 4:2:0, 4:2:2, and monochrome). If the palette mode is enabled, then if the CU size is less than or equal to 64×64, a flag can be sent at the CU level, and the number of samples in the CU is greater than 16 to indicate whether the palette mode is used. Applying the palette mode on small CUs can introduce insignificant decoding gain and bring additional complexity to small blocks. The palette mode can be disabled for CUs with less than or equal to 16 samples. The palette-coded CU can be regarded as a prediction mode (e.g., separate from intra prediction, inter prediction, and IBC modes).
[0219] Figure 14 An example of palette mode encoding is illustrated (e.g., the palette size is 4). If the palette mode is utilized, the sample values in the CU can be represented by a set of representative color values. This set can be referred to as a palette. For positions with sampling values close to the palette colors, the palette index can be signaled. Samples outside the palette can be specified (e.g., by signaling an escape symbol). For samples decoded using the escape symbol within the CU, their component values can be signaled using the quantized component values (e.g., directly). The quantized escape symbol can be binaryized (e.g., by a fifth-order exponential Golomb binaryization process (EG5)).
[0220] The combined intra-inter prediction (CIIP) mode can be used for decoding blocks. In the CIIP mode, the predicted samples can be generated by weighting the inter prediction signal using CIIP template matching (CIIP-TM) to merge candidates and the intra prediction signal predicted using the intra prediction mode derived by TIMD. The CIIP mode can be applied to (e.g., only applied to) decoding blocks with an area less than or equal to 1024.
[0221] The TIMD derivation method can be used to derive the intra prediction mode in CIIP. Specifically, the intra prediction mode with the minimum SATD value in the TIMD mode list can be selected and mapped to one of the 67 directional intra prediction modes (e.g., conventional intra prediction modes).
[0222] If the derived intra prediction mode is an angular mode, the weights (w Intra , w Inter ) of two tests can be modified. For near-horizontal modes (e.g., 2 <= angular mode index < 34), the current block can be vertically divided. For near-vertical modes (e.g., 34 <= angular mode index <= 66), the current block can be horizontally divided.
[0223] In some examples, a Geometric Partitioning Mode (GPM) can be used in conjunction with inter - frame and intra - frame prediction. In a GPM with inter - frame and intra - frame prediction, the final prediction samples can be generated by weighting the inter - frame prediction samples and the intra - frame prediction samples for each GPM - separated region. The inter - frame prediction samples can be derived through the inter - frame GPM, while the intra - frame prediction samples can be derived through an Intra - Prediction Mode (IPM) candidate list and / or an index signaled from the encoder. The IPM candidate list size can be predefined as three. The available IPM candidates can be a parallel angle mode (parallel mode) for the GPM block boundary, a vertical angle mode (vertical mode) for the GPM block boundary, and a planar mode as shown respectively Figures 15A to 15C as shown. Figure 15D A GPM with intra - frame and inter - frame prediction is illustrated. The GPM with intra - frame and inter - frame prediction can be restricted (e.g., to reduce the signaling overhead of the IPM and / or to avoid an increase in the size of the intra - frame prediction circuit on a hardware decoder). A direct motion vector and IPM storage on the GPM hybrid region can be introduced (e.g., to further improve the coding performance).
[0224] In the IPM derivation based on DIMD and neighboring modes, the parallel mode can be registered (e.g., first). If the same IPM candidate is not in the list, up to two IPM candidates derived from the DIMD method and / or neighboring blocks can be registered. As for the neighboring mode derivation, there can be five positions (e.g., at most) for the available neighboring blocks. The positions can be restricted by the angle of the GPM block boundary (e.g., as shown Figure 16 as shown), which can be used for a GPM with template matching (GPM - TM). In Figure 16 this, A and L can represent the upper side and the left side of the prediction block, respectively.
[0225] In some examples, GPM - intra can be combined with a GPM with Motion Vector Difference Merging (GPM - MMVD). TIMD can be used for the IPM candidates of GPM - intra (e.g., to further improve the decoding performance). The parallel mode can be registered first. Subsequently, the IPM candidates of TIMD, DIMD, and neighboring blocks can be registered.
[0226] A Low - Frequency Non - Separable Transform (LFNST) can be performed. The forward LFNST can be applied to the upper - left low - frequency region, which can be referred to as the Region of Interest (ROI). If the LFNST is applied, the main transform coefficients in the region outside the ROI can be zeroed.
[0227] Figure 17Illustrates the ROI of LFNST16. The ROI of LFNST16 includes six 4×4 sub-blocks (e.g., which can be consecutive in scan order). The number of input samples can be 96. In this case, the transform matrix of the forward LFNST 16 can be R×96. 32 coefficients (two 4×4 sub-blocks) can be generated from the forward LFNST 16 (e.g., if the value of R is chosen as 32). The coefficients can be placed following the coefficient scan order.
[0228] Figure 18 Illustrates the ROI of LFNST8. The forward LFNST8 matrix can be R×64. The value of R can be 32. The generated coefficients can be positioned in the same way as LFNST16. Figure 19 Illustrates an example mapping from the intra prediction mode to the LFNST set index.
[0229] Multiple transform selection (MTS) can be used. For MTS, the DST7 and DST8 (e.g., only DST7 and DCT8) transform kernels can be utilized. The DST7 and DST8 transform kernels can be used for intra and inter coding.
[0230] Other primary transforms (e.g., including DCT5, DST4, DST1) and / or the identity transform (IDT) can be adopted. The MTS set can depend on the TU size and / or the intra mode information. 16 different TU sizes can be considered. For each TU size, depending on the intra mode information, five different categories can be considered. For each category, one, four, or six different transform pairs can be considered. The number of intra MTS candidates can be adaptively selected (e.g., between one, four, and six MTS candidates). The number of intra MTS candidates can depend on the sum of the absolute values of the transform coefficients. The sum can be compared with one or more thresholds (e.g., two fixed thresholds) to determine the total number of allowed MTS candidates. For example:
[0231] 1 candidate: sum <= th0 Equation 13
[0232] 4 candidates: th0 < sum <= th1 Equation 14
[0233] 6 candidates: sum > th1 Equation 15
[0234] Intra mode propagation can be performed. For a CU that is not decoded in intra prediction, the intra mode of the reference CU can be considered the same intra mode as the current CU. This mode can be used when constructing the most probable mode (MPM) list for other blocks. The MPM can be generated using a method. In this method, the first entry in the MPM list can be the planar mode. The remaining entries can include the intra modes of the left (L), above (A), bottom-left (BL), top-right (AR), and top-left (AL) neighboring blocks (e.g., as shown in Figure 20 ), the directional mode with an added offset from the first two available directional modes of the neighboring blocks, and / or the default mode.
[0235] If any of the neighboring blocks are inter-coded, the intra mode of the block can be obtained from the reference block (or its reference if it is also inter-coded). A buffer for the intra mode of the position can be generated (e.g., with a resolution of the minimum CU size (4×4)). After decoding the CU, the buffer can be filled with the intra mode or the reference intra mode (e.g., if inter-coded). This process can be referred to as intra mode propagation. The MIP, IntraTMP, and / or palette modes can be propagated as the planar mode. The GPM mode with an intra-inter mode can generate an MPM with three entries (e.g., similar to MPM list generation).
[0236] The intra mode can provide useful information about the statistics of the current block. The intra mode can provide information about the directionality of the block. This information can be used to design the transform in MTS and / or LFSNT (e.g., the optimal transform). The LFNST can be a transform learned by clustering the residual signal according to the intra mode of the residual signal. The intra mode can be used to construct the MPM list. The intra mode can be used for GPM MPM.
[0237] In some examples, when a block is decoded in a non-directional intra prediction mode (e.g., inter prediction mode, IBC mode, intra TMP mode, MIP, palette mode, cross-component prediction mode, etc.), the intra mode-dependent tools can be disabled. For example, the LFNST can be disabled based on a block decoded in a non-directional intra prediction mode (e.g., because the LFNST is directional mode-related). If the planar mode is considered, the MIP can be used with the LFNST. In some examples, the intra-dependency tools can use an equivalent mode. The equivalent mode (e.g., using the DIMD process) can be used for LFNST kernel selection, where a decoding gain is provided.
[0238] An equivalent mode (e.g., a directional intra prediction mode) can be derived for a block that uses a non - directional intra prediction mode (e.g., a CU that does not use regular intra coding). For example, the TIMD and / or DIMD processes can be used to derive the directional intra prediction mode. The equivalent mode can be used to select the MTS / LFSNT kernel and / or the intra mode propagation process.
[0239] In some examples, a video decoding device may determine that a current block is decoded using a non - directional intra prediction mode. A directional intra prediction mode corresponding to the non - directional intra prediction mode (e.g., which indicates the derived intra prediction direction) can be derived. The video decoding device can decode the current block at least in part based on the derived directional intra prediction mode.
[0240] The non - directional intra prediction mode can be used to obtain a predicted block of the current block. In some examples, a directional intra prediction mode corresponding to the non - directional intra prediction mode can be derived based on the predicted block. In some examples, the reconstructed samples (e.g., multiple reconstructed samples) in the predicted block can be obtained. The directional intra prediction mode can be derived based on the reconstructed samples in the predicted block and the reconstructed neighboring samples (e.g., multiple reconstructed neighboring samples) of the current block.
[0241] The directional intra prediction mode can be derived based on a gradient histogram associated with the reconstructed pixels neighboring the current block (e.g., the directional intra prediction mode can be derived by applying the DIMD process to the reconstructed template of the current block, e.g., the template, the predicted block of the current block obtained using the non - directional intra prediction mode, or the reconstructed template inside the predicted block). For example, multiple samples in the predicted block can be obtained. The directionality of the predicted block can be determined.
[0242] Figure 21 A process for deriving an equivalent mode is illustrated. For example, the DIMD process can be used to derive an equivalent mode (e.g., a directional intra prediction mode corresponding to a non - directional intra prediction mode). An equivalent directional intra prediction mode for MIP can be generated during the MIP prediction process (e.g., as Figure 21 shown). In some examples, DIMD can be applied to the reconstructed template around the current block. In some examples, the DIMD process can be applied to the predicted block (e.g., the prediction signal). For example, the DIMD process can be used to find the directionality of the predicted block generated by the MIP process. For example, the directionality of the predicted block can be determined based on multiple samples in the predicted block. The intra prediction mode can be derived based on the determined directionality of the predicted block.
[0243] Figure 22A An example DIMD process is illustrated, where a template around the current block to be decoded is used. Figure 22BIllustrates a process for deriving a MIP equivalent mode, where the template is part of a prediction block (e.g., before upsampling).
[0244] In some examples, the DIMD process can be used to analyze prediction signals generated from inter - frame prediction, IBC, CCLM / MMLM / CCCM, and / or IntraTMP. In some examples, a prediction unit (e.g., the entire prediction unit) can be analyzed to derive an equivalent directional intra - prediction mode (e.g., rather than using a template inside the prediction unit). In some examples, the default DIMD process can be used as the equivalent directional intra - prediction mode. This can be used for palette mode (e.g., because prediction signals may not be generated in palette mode).
[0245] In some examples, a directional intra - prediction mode can be derived by: testing multiple candidate directional intra - prediction modes on the reconstructed pixels adjacent to the current block; and selecting a directional intra - prediction mode from the multiple candidate directional intra - prediction modes based on the test. For example, a template - based intra - mode derivation (TIMD) process can be applied to at least one of the reconstructed template of the current block, the prediction block of the current block obtained using a non - directional intra - prediction mode, or the reconstructed template inside the prediction block.
[0246] For example, the TIMD process can be used to derive an equivalent directional intra - prediction mode (e.g., a directional intra - prediction mode corresponding to a non - directional intra - prediction mode). TIMD can be applied with a template around the current block (e.g., in the same or a similar manner as DIMD). For example, TIMD can be applied with a template around the current block using a template inside the prediction block. For example, TIMD can be applied with a template around the current block using the entire prediction block.
[0247] For example, a non - directional intra - prediction mode can be used to obtain the prediction block of the current block. In some examples, the reconstructed samples in the prediction block (e.g., multiple reconstructed samples) can be obtained. In some examples, possible prediction modes (e.g., multiple possible prediction modes) can be obtained. Predictions of the reconstructed samples in the prediction block can be calculated (e.g., multiple predictions). For example, the predictions of the reconstructed samples in the prediction block can be calculated based on the possible prediction modes. Prediction errors can be calculated (e.g., multiple prediction errors). For example, the prediction errors can be calculated based on the reconstructed samples in the prediction block and the corresponding predictions. The prediction errors can correspond to the possible prediction modes. A directional intra - prediction mode can be selected based on the prediction errors (e.g., from the possible prediction modes). In some examples, a directional intra - prediction mode can be selected based on determining that the prediction error corresponding to the directional intra - prediction mode is the smallest among the prediction errors.
[0248] In some examples, a planar or DC mode can be used as an equivalent directional intra prediction mode. For example, MIP and IntraTMP can be considered as planar modes in LFNST kernel selection. In some examples, LFNST can be activated for directional inter prediction modes and IBC modes.
[0249] In some examples, a history-based intra prediction mode (HIPM) can be used as an equivalent directional intra prediction mode. The history-based intra prediction mode can be used as an equivalent directional intra prediction mode for a CU that employs a non-directional intra prediction mode (e.g., a CU that does not employ conventional intra coding). In an example, the derivation process can be similar to that of a history-based MVP (HMVP) merge candidate. The derived directional intra prediction mode (e.g., the directional intra prediction mode of a previously conventional intra-coded block) can be stored in a table. The derived directional intra prediction mode can be used to generate an MPM list for neighboring prediction blocks. A table with multiple HIPM candidates can be used as an equivalent directional intra prediction mode for the current CU. The table with multiple HIPM candidates can be maintained during the encoding and / or decoding process. When a new CTU row is encountered, the table can be reset (e.g., cleared). If there is a CU coded with a directional intra prediction mode (e.g., a conventional intra-coded CU), the associated directional intra prediction mode can be added to the last entry of the table (e.g., as a new HIPM candidate).
[0250] The HIPM table size S can be set to a value M (e.g., indicating that up to M - 1 HIPM candidates can be added to the table). When inserting a new directional intra prediction mode candidate into the table, two options can be considered. For example, in the first option, the new HIPM can be moved to the last entry of the table. In this example, subsequent HIPM candidates (e.g., all subsequent HIPM candidates) can be moved forward. In this example, the HIPM candidate in the last entry of the table can be considered "the most recent / closest" and can be used as an equivalent directional intra prediction mode.
[0251] For example, in the second option, the occurrences of existing HIPMs in the table can be counted. It can be determined whether the same HIPM exists in the table. If found, the count of the same HIPM can be incremented. In this case, the current HIPM table can be reordered. For example, if the HIPM candidate occurrence count is higher than the last entry of the current HIPM table, the HIPM candidate can be moved to the last entry of the table. In this case, the HIPM candidate in the last entry of the table can be used as an equivalent directional intra prediction mode (e.g., because it is frequently used).
[0252] LFNST can be performed based on equivalent directional intra prediction modes. For example, the LFNST transform set can be determined / selected based on the derived directional intra prediction modes. The equivalent directional intra prediction modes can be derived as described herein. The current block can be encoded and / or decoded based on the LFNST transform set. For example, a transform or inverse transform can be performed on the residual of the current block based on the LFNST transform set.
[0253] LFNST can be activated for inter-coded CUs and IBC modes. Equivalent directional intra prediction mode derivation can be used for LFNST kernel selection. LFNST can be activated for cross-component prediction / IntraTMP (e.g., using equivalent mode derivation for LFNST kernel selection instead of assuming a planar mode).
[0254] In an example, MTS can be performed using equivalent directional intra prediction modes. For example, the MTS transform set can be determined based on the derived directional intra prediction modes. The current block can be encoded and / or decoded based on the MTS transform set. For example, a transform or inverse transform can be performed on the residual of the current block based on the MTS transform set. MTS IBC and IntraTMP modes can be activated (e.g., with equivalent mode derivation for LFNST kernel selection). In some examples, MTS for the chroma portion can be activated (e.g., with cross-component prediction derived from equivalent mode for LFNST core selection instead of assuming a planar mode). MTS kernel selection can be used for inter-frame CUs (e.g., with equivalent directional intra prediction mode derivation for kernel selection).
[0255] In some examples (e.g., for inter-frame CUs), the MTS index can be decoded independently of the directional intra prediction mode. In some examples, a table mapping the MTS index to kernels can be (pre-)defined.
[0256] Specific considerations for TU splitting can be provided. In some examples, TU splitting can be allowed. A CU can be split into multiple TUs. For example, a residual quad tree (RQT) can be used to split a CU into multiple TUs. In some examples, the RQT can be removed.
[0257] Sub-block transform (SBT) can be similar to RQT. SBT can be used on inter-coded CUs. Using SBT, a CU can be split into two parts (e.g., as illustrated in Figure 23 ). One of these parts can be zeroed. A (pre-)defined transform set can be used to transform the other part.
[0258] In the case where TU is less than CU (e.g., as in SBT), equivalent directional intra prediction mode derivation can be performed for each TU. This can result in N equivalent modes for N sub - partitions. This can allow for proper selection of the MTS / LFNST kernels and / or can provide better propagation of the equivalent directional intra prediction modes. Specifically, in SBT, since (e.g., one) partition is zeroed, the equivalent directional intra prediction mode for that partition may not be determined.
[0259] In some examples, equivalent directional intra prediction modes can be propagated. When intra mode propagation is used, equivalent directional intra prediction modes can be derived (e.g., as described herein). For example, if a non - directional intra prediction mode (e.g., IBC, inter, cross - component prediction, intra TMP, palette mode) is used, then the equivalent directional intra prediction mode can be used to fill the intra mode buffer.
[0260] Specific considerations can be provided for CIIP and / or GPM modes. For example, when using CIIP, CIIP can be used to derive the directional intra prediction mode for the intra part. For example, when using GPM intra - inter, the directional intra prediction mode can be derived. Equivalent directional intra prediction modes may not be derived for CIIP and / or GPM modes. The intra modes of the intra part of CIIP and / or GPM modes can be used for LFNST / MTS kernel selection and intra mode propagation.
[0261] A video encoding device (e.g., an encoder) can perform actions that are the same as or similar to those described above. For example, the encoder can identify a non - directional intra prediction mode for encoding a current block. The encoder can derive a directional intra prediction mode corresponding to the non - directional intra prediction mode (e.g., which includes the derived intra prediction direction). The encoder can encode the current block at least in part based on the derived directional intra prediction mode.
[0262] Figure 24 An example flowchart 2400 for decoding a current block is illustrated. At 2410, it can be determined that the current block is encoded in a non - directional intra prediction mode. At 2420, a directional intra prediction mode corresponding to the non - directional intra prediction mode can be derived. At 2430, the current block can be decoded at least in part based on the derived directional intra prediction mode.
[0263] Figure 25 An example flowchart 2500 for encoding a current block is illustrated. At 2510, a non - directional intra prediction mode for encoding the current block can be identified. At 2520, a directional intra prediction mode corresponding to the non - directional intra prediction mode can be derived. At 2530, the current block can be encoded at least in part based on the derived directional intra prediction mode.
[0264] Although the features and elements have been described above in a particular combination, one of ordinary skill in the art will appreciate that each feature or element can be used separately or in any combination with other features and elements. In addition, the methods described herein can be implemented in a computer program, software, or firmware incorporated in a computer-readable medium for execution by a computer or a processor. Examples of computer-readable media include electronic signals (transmitted via wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, buffer memories, semiconductor storage devices, magnetic media of internal hard disks and removable disks, magneto-optical media, and optical media of CD-ROM discs and digital versatile discs. A processor associated with the software can be used to implement a radio frequency transceiver used in a WTRU, UE, terminal, base station, RNC, or any host computer.
Claims
1. A device for video decoding, comprising: a processor configured to: determine that a current block is decoded in a non - directional intra - prediction mode; derive a directional intra - prediction mode corresponding to the non - directional intra - prediction mode, wherein the derived directional intra - prediction mode indicates a derived intra - prediction direction; and decode the current block at least in part based on the derived directional intra - prediction mode.
2. The device according to claim 1, wherein, The processor is further configured to: obtain a predicted block of the current block using the non - directional intra - prediction mode; and obtain a plurality of reconstructed samples in the predicted block, wherein the directional intra - prediction mode is derived based on the plurality of reconstructed samples in the predicted block and a plurality of reconstructed neighboring samples of the current block.
3. The device according to claim 1 or 2, wherein, The processor is further configured to: store the derived directional intra - prediction mode; and use the derived directional intra - prediction mode to generate a most - probable mode (MPM) list for neighboring predicted blocks.
4. The apparatus according to any one of claims 1 to 3, wherein, The processor is further configured to determine a low - frequency non - separable transform (LFNST) transform set based on the derived directional intra - prediction mode, and wherein the current block is decoded based on the LFNST transform set.
5. The apparatus according to any one of claims 1 to 3, wherein, The processor is further configured to determine a multi - transform selection (MTS) transform set based on the derived directional intra - prediction mode, and wherein the current block is decoded based on the MTS transform set.
6. The device according to any one of claims 1 to 5, wherein The processor being configured to derive the directional intra - prediction mode includes the processor being configured to derive the directional intra - prediction mode based on a gradient histogram associated with reconstructed pixels neighboring the current block.
7. The device according to any one of claims 1 to 5, wherein The processor being configured to derive the directional intra - prediction mode includes the processor being configured to: test a plurality of candidate directional intra - prediction modes on reconstructed pixels neighboring the current block; and select the directional intra - prediction mode from the plurality of candidate directional intra - prediction modes based on the test.
8. The device according to any one of claims 1 to 5, wherein, The processor is further configured to: obtain a predicted block of the current block using the non - directional intra - prediction mode; obtain a plurality of reconstructed samples in the predicted block; obtain a plurality of possible prediction modes; calculate a plurality of predictions of the plurality of reconstructed samples in the predicted block based on the plurality of possible prediction modes; calculate a plurality of prediction errors corresponding to the plurality of possible prediction modes based on the plurality of reconstructed samples in the predicted block and the corresponding plurality of predictions; and select the directional intra - prediction mode from the plurality of possible prediction modes based on the plurality of prediction errors.
9. The device according to any one of claims 1 to 8, wherein, The non - directional intra - prediction mode includes an inter - prediction mode, a cross - component prediction mode, a palette mode, an intra - block copy (IBC) mode, or an intra - template matching prediction (IntraTMP) mode.
10. The device according to any one of claims 1 to 3 and 6 to 9, wherein, The processor is further configured to: select a low - frequency non - separable transform (LFNST) transform set based on the directional intra - prediction mode; and perform an inverse transform on the residual of the current block based on the LFNST transform set.
11. The device according to any one of claims 1 to 3 and 6 to 9, wherein, The processor is further configured to: select a multi - transform selection (MTS) transform set based on the directional intra - prediction mode; and Perform an inverse transform on the residual of the current block based on the MTS transform set.
12. A method for video decoding, comprising: Determine that a current block is decoded in a non - directional intra - prediction mode; Derive a directional intra - prediction mode corresponding to the non - directional intra - prediction mode, wherein the derived directional intra - prediction mode indicates a derived intra - prediction direction; And Decode the current block at least in part based on the derived directional intra - prediction mode.
13. The method according to claim 12, wherein, The method further comprises: Obtain a predicted block of the current block using the non - directional intra - prediction mode; and Obtain a plurality of reconstructed samples in the predicted block, wherein the directional intra - prediction mode is derived based on the plurality of reconstructed samples in the predicted block and a plurality of reconstructed neighboring samples of the current block.
14. The method according to claim 12 or 13, wherein, The method further comprises: Store the derived directional intra - prediction mode; and Use the derived directional intra - prediction mode to generate a most - probable mode (MPM) list for neighboring predicted blocks.
15. The method according to any one of claims 12 to 14, wherein The method further comprises determining a low - frequency non - separable transform (LFNST) transform set based on the derived directional intra - prediction mode, and wherein the current block is decoded based on the LFNST transform set.
16. The method according to any one of claims 12 to 14, wherein The method further comprises determining a multi - transform selection (MTS) transform set based on the derived directional intra - prediction mode, and wherein the current block is decoded based on the MTS transform set.
17. The method according to any one of claims 12 to 16, wherein, Deriving the directional intra - prediction mode includes deriving the directional intra - prediction mode based on a gradient histogram associated with reconstructed pixels neighboring the current block.
18. The method according to any one of claims 12 to 16, wherein Deriving the directional intra - prediction mode includes: Testing a plurality of candidate directional intra - prediction modes on reconstructed pixels neighboring the current block; and Based on the testing, selecting the directional intra - prediction mode from the plurality of candidate directional intra - prediction modes.
19. The method according to any one of claims 12 to 16, wherein The method further comprises: Obtain a predicted block of the current block using the non - directional intra - prediction mode; Obtain a plurality of reconstructed samples in the predicted block; Obtain a plurality of possible prediction modes; Calculate a plurality of predictions for the plurality of reconstructed samples in the predicted block based on the plurality of possible prediction modes; Calculate a plurality of prediction errors corresponding to the plurality of possible prediction modes based on the plurality of reconstructed samples in the predicted block and the corresponding plurality of predictions; and Based on the plurality of prediction errors, select the directional intra - prediction mode from the plurality of possible prediction modes.
20. The method according to any one of claims 12 to 19, wherein The non - directional intra - prediction mode includes an inter - prediction mode, a cross - component prediction mode, a palette mode, an intra - block copy (IBC) mode, or an intra - template matching prediction (IntraTMP) mode.
21. The method according to any one of claims 12 to 14 and 17 to 20, wherein, The method further comprises: Select a low - frequency non - separable transform (LFNST) transform set based on the directional intra - prediction mode; and Perform an inverse transform on the residual of the current block based on the LFNST transform set.
22. The method according to any one of claims 12 to 14 and 17 to 20, wherein, The method further comprises: Select a multi - transform selection (MTS) transform set based on the directional intra - prediction mode; and Perform an inverse transform on the residual of the current block based on the MTS transform set.
23. A device for video encoding, comprising: A processor, the processor being configured to: Identify a non - directional intra - prediction mode for encoding a current block; Derive a directional intra - prediction mode corresponding to the non - directional intra - prediction mode, wherein the derived directional intra - prediction mode includes a derived intra - prediction direction; And At least partially encode the current block based on the derived directional intra - prediction mode.
24. The device according to claim 23, wherein, The processor is further configured to: Obtain a predicted block of the current block using the non - directional intra - prediction mode; and Obtain a plurality of reconstructed samples in the predicted block, wherein the directional intra - prediction mode is derived based on the plurality of reconstructed samples and a plurality of reconstructed neighboring samples in the predicted block.
25. The device according to claim 23 or 24, wherein, The processor is further configured to: Store the derived directional intra - prediction mode; and Use the derived directional intra - prediction mode to generate a most - probable mode (MPM) list for neighboring predicted blocks.
26. The device according to any one of claims 23 to 25, wherein, The processor is further configured to determine a low - frequency non - separable transform (LFNST) transform set based on the derived directional intra - prediction mode, and wherein the current block is encoded based on the LFNST transform set.
27. The apparatus according to any one of claims 23 to 25, wherein The processor is further configured to determine a multi - transform selection (MTS) transform set based on the derived directional intra - prediction mode, and wherein the current block is encoded based on the MTS transform set.
28. The device according to any one of claims 23 to 27, wherein The processor being configured to derive the directional intra - prediction mode includes the processor being configured to derive the directional intra - prediction mode based on a gradient histogram associated with reconstructed pixels neighboring the current block.
29. The apparatus according to any one of claims 23 to 27, wherein, The processor being configured to derive the directional intra - prediction mode includes the processor being configured to: Test a plurality of candidate directional intra - prediction modes on reconstructed pixels neighboring the current block; And Based on the test, select the directional intra - prediction mode from the plurality of candidate directional intra - prediction modes.
30. The apparatus according to any one of claims 23 to 27, wherein, The processor is further configured to: Obtain a predicted block of the current block using the non - directional intra - prediction mode; Obtain a plurality of reconstructed samples in the predicted block; Obtain a plurality of possible prediction modes; Calculate a plurality of predictions of the plurality of reconstructed samples in the predicted block based on the plurality of possible prediction modes; Calculate a plurality of prediction errors corresponding to the plurality of possible prediction modes based on the plurality of reconstructed samples and the corresponding plurality of predictions in the predicted block; And Based on the plurality of prediction errors, select the directional intra - prediction mode from the plurality of possible prediction modes.
31. The apparatus according to any one of claims 23 to 30, wherein, The non - directional intra - prediction mode includes an inter - prediction mode, a cross - component prediction mode, a palette mode, an intra - block copy (IBC) mode, or an intra - template matching prediction (IntraTMP) mode.
32. The apparatus according to any one of claims 23 to 25 and 28 to 31, wherein The processor is further configured to: Select a low - frequency non - separable transform (LFNST) transform set based on the directional intra - prediction mode; and Perform a transform on the residual of the current block based on the LFNST transform set.
33. The apparatus according to any one of claims 23 to 25 and 28 to 31, wherein The processor is further configured to: Select a multi - transform selection (MTS) transform set based on the directional intra - prediction mode; and Perform a transform on the residual of the current block based on the MTS transform set.
34. A method for video coding, comprising: identifying a non - directional intra - prediction mode for encoding a current block; deriving a directional intra - prediction mode corresponding to the non - directional intra - prediction mode, wherein the derived directional intra - prediction mode includes a derived intra - prediction direction; and encoding the current block at least in part based on the derived directional intra - prediction mode.
35. The method according to claim 34, wherein, The method further comprises: obtaining a predicted block of the current block using the non - directional intra - prediction mode; and obtaining a plurality of reconstructed samples in the predicted block, wherein the directional intra - prediction mode is derived based on the plurality of reconstructed samples and a plurality of reconstructed neighboring samples in the predicted block.
36. The method according to claim 34 or 35, wherein, The method further comprises: storing the derived directional intra - prediction mode; and using the derived directional intra - prediction mode to generate a most - probable mode (MPM) list for neighboring predicted blocks.
37. The method according to any one of claims 34 to 36, wherein, The method further comprises determining a low - frequency non - separable transform (LFNST) transform set based on the derived directional intra - prediction mode, and encoding the current block based on the LFNST transform set.
38. The method according to any one of claims 34 to 36, wherein, The method further comprises determining a multi - transform selection (MTS) transform set based on the derived directional intra - prediction mode, and encoding the current block based on the MTS transform set.
39. The method according to any one of claims 34 to 38, wherein, Deriving the directional intra - prediction mode is based on a gradient histogram associated with reconstructed pixels neighboring the current block.
40. The method according to any one of claims 34 to 38, wherein, Deriving the directional intra - prediction mode comprises: testing a plurality of candidate directional intra - prediction modes on reconstructed pixels neighboring the current block; and selecting the directional intra - prediction mode from the plurality of candidate directional intra - prediction modes based on the testing.
41. The method according to any one of claims 34 to 38, wherein The method further comprises: obtaining a predicted block of the current block using the non - directional intra - prediction mode; obtaining a plurality of reconstructed samples in the predicted block; obtaining a plurality of possible prediction modes; calculating a plurality of predictions of the plurality of reconstructed samples in the predicted block based on the plurality of possible prediction modes; calculating a plurality of prediction errors corresponding to the plurality of possible prediction modes based on the plurality of reconstructed samples and the corresponding plurality of predictions in the predicted block; and selecting the directional intra - prediction mode from the plurality of possible prediction modes based on the plurality of prediction errors.
42. The method according to any one of claims 34 to 41, wherein, The non - directional intra - prediction mode includes an inter - prediction mode, a cross - component prediction mode, a palette mode, an intra - block copy (IBC) mode, or an intra - template matching prediction (IntraTMP) mode.
43. The method according to any one of claims 34 to 36 and 39 to 42, wherein, The method further comprises: selecting a low - frequency non - separable transform (LFNST) transform set based on the directional intra - prediction mode; and performing a transform on the residual of the current block based on the LFNST transform set.
44. The method according to any one of claims 34 to 36 and 39 to 42, wherein, The method further comprises: selecting a multi - transform selection (MTS) transform set based on the directional intra - prediction mode; and performing a transform on the residual of the current block based on the MTS transform set.
45. A computer - readable medium comprising instructions for causing one or more processors to perform the method according to any one of claims 12 to 22 and 34 to 44.
46. A video data includes information representing an encoded current block generated by the method according to one of claims 34 to 44.