IntraCIIP mode
Incorporating IntraTMP into CIIP for video encoding systems addresses inefficiencies in prediction modes, leading to improved compression and decoding efficiency through weighted signal combinations and transformations.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- INTERDIGITALCE PATENT HLDG SAS
- Filing Date
- 2024-03-27
- Publication Date
- 2026-06-02
AI Technical Summary
Existing video encoding systems face challenges in optimizing prediction modes for video blocks, leading to inefficiencies in compression and decoding processes.
Incorporating Intra Template Matching (IntraTMP) into Combined Inter and Intra Prediction (CIIP) to enhance video encoding and decoding by merging prediction modes, using weighted combinations of intra-prediction signals and block vector-based predictions, and applying transformations like MTS, LFNST, or NSPT.
Improves video encoding efficiency by optimizing prediction modes, reducing computational complexity, and enhancing compression performance.
Smart Images

Figure 2026517604000001_ABST
Abstract
Description
Background Art
[0001] Cross-references to related applications This application claims the benefit of European Provisional Patent Application No. 23305455.0, filed on March 30, 2023, and European Provisional Patent Application No. 23306561.4, filed on September 20, 2023, and thus the content is incorporated herein by reference.
[0002] background Video encoding systems are widely used to compress digital video signals, for example, to reduce the storage and / or transmission bandwidth required for the above signals. For example, video encoding systems may include block-based, wavelet-based, and / or object-based systems.
Summary of the Invention
[0003] Systems, methods, and instrumentalities are disclosed herein with respect to video encoding and / or decoding using an intra combined intra-intra prediction mode. The prediction process of video encoding and / or decoding may incorporate IntraTMP (Intra Template Matching) into CIIP (Combined Inter and Intra Prediction). The incorporation described above may, for example, add a new mode that merges the prediction from intra prediction and the prediction from IntraTMP (e.g., using the merge mechanism of CIIP).
[0004] A device, for example, a video decoding device, may perform (or be configured to perform) one or more of the following operations: The device may acquire a first intra-prediction signal associated with a first prediction mode and a second intra-prediction signal associated with a block vector-based prediction mode (for example, a second prediction mode) for a video block. The device may generate prediction samples associated with the video block based on the acquired intra-prediction signals (for example, the first and second intra-prediction signals). For example, the device may merge the acquired prediction signals to generate prediction samples. The device may decode the video block based on the predicted samples.
[0005] In the example, the first intra-prediction signal may be associated with the most probable mode (MPM) list of the video block. In the example, the block vector-based prediction mode may be either intra-template matching prediction (intra-TMP) or intra-block copy (IBC).
[0006] The device may determine a first weight associated with the intra-prediction signal associated with a first prediction mode and a second weight associated with the intra-prediction signal associated with a block vector-based prediction mode. Prediction samples associated with a video block may be generated based on a first weight applied to the first intra-prediction signal and a second weight applied to the second intra-prediction signal. The first and second weights may be determined based on whether the intra-TMP is associated with one or more of the video block and its neighbors (first and second blocks).
[0007] In the example, the decoder and / or encoder may determine a first weight and a second weight. For example, based on the fact that one of the first neighbor block prediction modes or the second neighbor block prediction mode is an intra-TMP mode, and that one of the first neighbor block prediction modes or the second neighbor block prediction mode is a normal intra-prediction mode, the decoder and / or encoder may determine that the first weight and the second weight are weighted equally (e.g., each has a weight equal to 2). For example, based on the fact that both the first and second neighbor block prediction modes are intra-TMP modes, the decoder or encoder may determine that the first weight is 3 and the second weight is 1. For example, based on the fact that both the first and second neighbor block prediction modes are normal intra-prediction modes, the decoder or encoder may determine that the first weight is 1 and the second weight is 3.
[0008] The device may sort multiple MPM modes in the MPM list (for example, when a first intra-prediction signal is associated with a video block's most probable mode (MPM) list) based on the respective costs associated with each MPM mode. The device may determine a first prediction mode based on the sorted multiple MPM modes. For example, the first intra-prediction signal may be obtained based on a first prediction mode (for example, from the sorted multiple MPM modes). For example, the first prediction mode may be determined to be the MPM mode with the lowest respective cost. The device may obtain an index associated with the MPM list. The device may determine a first prediction mode based on the index (for example, using the sorted multiple MPM modes). For example, the first prediction mode may be selected from the sorted MPM list based on the index.
[0009] The device may acquire multiple block vectors. The device may select a block vector from among multiple block vectors. A second intra-prediction signal may be acquired based on the selected block vector. The device may acquire an index. The selected block vector may be based on the index. The multiple block vectors may include one or more of the multiple block vectors acquired by performing intra-template matching prediction (intra-TMP) on the video block, or multiple block vectors acquired from the video block and at least one neighboring block.
[0010] A video decoding method may include obtaining a first intra-prediction signal associated with a first prediction mode and a second intra-prediction signal associated with a block vector-based prediction mode (e.g., a second prediction mode) for a video block. For example, the block vector-based prediction mode may be intra-TMP or IBC. The method may include generating prediction samples associated with the video block based on the first and second intra-prediction signals. The method may include decoding the video block based on the predicted samples.
[0011] The method may include determining a first weight associated with a first intra-prediction signal and a second weight associated with a second intra-prediction signal. Prediction samples associated with a video block may be generated based on the first weight applied to the first intra-prediction signal and the second weight applied to the second intra-prediction signal. The first and / or second weights may be determined based on whether the intra-TMP is associated with one or more of the video block and its neighbors (first and second blocks).
[0012] A device, for example, a video encoding device, may perform (or be configured to perform) one or more of the following operations: The device may obtain a first intra-prediction signal associated with a first prediction mode (e.g., a normal prediction mode) and a second intra-prediction signal associated with a block vector-based prediction mode (e.g., a second prediction mode) for a video block. The device may generate prediction samples associated with the video block based on the first and second intra-prediction signals. The device may encode the video block based on the predicted samples.
[0013] The first intra-prediction signal may be associated with the most probable mode (MPM) list of the video block. The block vector-based prediction mode may be either intra-template matching prediction (intra-TMP) or intra-block copy (IBC).
[0014] The device may determine a first weight associated with the intra-prediction signal associated with a first prediction mode and a second weight associated with the intra-prediction signal associated with a block vector-based prediction mode. Prediction samples associated with a video block may be generated based on a first weight applied to the first intra-prediction signal and a second weight applied to the second intra-prediction signal. The first and second weights may be determined based on whether the intra-TMP is associated with one or more of the video block and its neighbors (first and second blocks).
[0015] The device may obtain an index associated with the most probable mode (MPM) list (for example, when the first intra-prediction signal is associated with the most probable mode (MPM) list of the video block). The device may sort the multiple MPM modes in the MPM list based on the respective costs associated with each MPM mode. The device may determine the first prediction mode based on the sorted multiple MPM modes and their indices. The first intra-prediction signal may be obtained based on the first prediction mode.
[0016] The device may acquire multiple block vectors and / or indices. The device may select a block vector from among the multiple block vectors (for example, based on its index). A second prediction signal may be acquired based on the selected block vector.
[0017] The device may select a transformation for a video block based on at least one of a first prediction mode and a second intra-prediction mode. For example, the transformation for a video block may be one or more of MTS (multiple transform selection), LFNST (low frequency non-separable transform), or NSPT (non-separable primary transform).
[0018] A video encoding method may include obtaining a first intra-prediction signal associated with a first prediction mode and a second intra-prediction signal associated with a block vector-based prediction mode (e.g., intra-TMP or IBC) for a video block. The method may include generating prediction samples associated with the video block based on the first and second intra-prediction signals. The method may include encoding the video block based on the predicted samples.
[0019] The method may include determining a first weight associated with a first intra-prediction signal and a second weight associated with a second intra-prediction signal. Prediction samples associated with a video block may be generated based on the first weight applied to the first intra-prediction signal and the second weight applied to the second intra-prediction signal. The first and / or second weights may be determined based on whether the intra-TMP is associated with one or more of the video block and its neighbors (first and second blocks).
[0020] The method may include selecting a transformation for a video block based on at least one of a first intra-prediction mode and a second intra-prediction mode, where the transformation for the video block is one or more of MTS (multiple transform selection), LFNST (low frequency non-separable transform), or NSPT (non-separable primary transform).
[0021] The systems, methods, and apparatus described herein may include decoders. In some examples, the systems, methods, and apparatus described herein may include encoders. In some examples, the systems, methods, and apparatus described herein may include signals (e.g., from and / or received by encoders). Video data (e.g., video bitstreams) may include video block / current block indications encoded according to intra-intra CIIP. Computer-readable media may include instructions for causing one or more processors to perform the methods described herein. Computer program products may include instructions that, when the program is executed by one or more processors, cause one or more processors to perform the methods described herein. [Brief explanation of the drawing]
[0022] [Figure 1A] This is a system diagram of an exemplary communication system in which one or more of the disclosed embodiments may be implemented. [Figure 1B] This is a system diagram illustrating an example of a wireless transmit / receive unit (WTRU) that may be used in a communication system as illustrated in Figure 1A, which relates to one embodiment. [Figure 1C] This is a system diagram illustrating exemplary radio access networks (RANs) and core networks (CNs) that may be used within a communication system as illustrated in Figure 1A, which relates to one embodiment. [Figure 1D] This is a system diagram illustrating further exemplary RANs and further exemplary CNs that may be used within the communication system illustrated in Figure 1A relating to one embodiment. [Figure 2] This figure shows an example video encoder. [Figure 3] A diagram showing an exemplary video decoder. [Figure 4] A diagram showing an example of a system in which various aspects and examples may be implemented. [Figure 5A] A diagram showing an example of a division related to an angle mode. [Figure 5B] A diagram showing an example of a division related to an angle mode. [Figure 6A] A diagram showing an example of a geometric partitioning mode (GPM) by inter and intra prediction. [Figure 6B] A diagram showing an example of a geometric partitioning mode (GPM) by inter and intra prediction. [Figure 6C] A diagram showing an example of a geometric partitioning mode (GPM) by inter and intra prediction. [Figure 6D] A diagram showing an example of a geometric partitioning mode (GPM) by inter and intra prediction. [Figure 7] A diagram showing an example of an intra template matching prediction (intra TMP) search area. [Figure 8] A diagram showing an example of a combined inter-intra prediction (CIIP) having intra-intra prediction. [Figure 9] A diagram showing an example of CIIP in which an intra prediction derivation process is replaced by an intra TMP mode. [Figure 10] A diagram showing an example of CIIP having intra-intra prediction.
Modes for Carrying Out the Invention
[0023] A more detailed understanding can be obtained from the following explanation, which is provided by examples associated with the attached drawings.
[0024] Figure 1A illustrates an exemplary communication system 100 in which one or more disclosed embodiments may be implemented. The communication system 100 may be a multiple-access system that provides content such as voice, data, video, messaging, and broadcast to multiple wireless users. The communication system 100 may enable multiple wireless users to access the above content through the sharing of system resources, including wireless bandwidth. For example, the communication system 100 may employ one or more channel access methods such as CDMA (Code Division Multiple Access), TDMA (Time Division Multiple Access), FDMA (Frequency Division Multiple Access), OFDMA (Orthogonal Frequency Division Multiple Access), SC-FDMA (Single Carrier FDMA), ZT UW DTS-s OFDM (zero-tail unique-word DFT-Spread OFDM), UW-OFDM (unique word OFDM), resource block-filtered OFDM, and FBMC (filter bank multicarrier).
[0025] As shown in Figure 1A, the communication system 100 may include WTRUs (Wireless Transmitter / Receiver Units) 102a, 102b, 102c, 102d, RAN 104 / 113, CN 106 / 115, Public Switched Telephone Network (PSTN) 108, the Internet 110, and other networks 112, but the disclosed embodiments will be understood to anticipate several WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, and 102d may be any type of device configured to operate and / or communicate in a wireless environment. For example, WTRU102a, 102b, 102c, and 102d may all be referred to as “station” and / or “STA,” and may be configured to transmit and / or receive wireless signals, including UEs (User Equipment), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, PDAs (Personal Digital Assistants), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, IoT (Internet of Things) devices, watches or other wearables, HMDs (Head-Mounted Displays), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in the context of industrial and / or automated processing chains), consumer electronics devices, and devices operating on commercial and / or industrial wireless networks. Any of WTRU102a, 102b, 102c, and 102d may be interchangeable with UE.
[0026] Furthermore, the communication system 100 may also include base stations 114a and / or base stations 114b. Each of the base stations 114a and 114b may be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, and 102d to facilitate access to one or more communication networks, such as CN106 / 115, the Internet 110, and / or other networks 112. As an example, base stations 114a and 114b may be a BTS (Radio Base Station Equipment), Node-B, eNode B, Home Node B, Home eNode B, gNB, NR Node B, Site Controller, AP (Access Point), Wireless Router, etc. While base stations 114a and 114b are each depicted as single elements, it will be understood that base stations 114a and 114b may include any number of interconnected base stations and / or network elements.
[0027] Base station 114a may also be part of RAN 104 / 113, which may include other base stations and / or network elements (not shown), such as BSC (Base Station Control Unit), RNC (Radio Network Control Unit), and relay nodes. Base station 114a and / or base station 114b may be configured to transmit and / or receive wireless signals on one or more carrier frequencies, which may be called a cell (not shown). The frequencies mentioned may be permitted spectrum, unpermitted spectrum, or a combination of permitted and unpermitted spectrum. A cell may provide wireless service coverage to a particular geographical area that may be relatively fixed or may change in the future. Furthermore, a cell may be divided into sectors. For example, a cell associated with base station 114a may be divided into three sectors. Thus, in one embodiment, base station 114a may include three transceivers, i.e., one for each sector of the cell. In the embodiment, the base station 114a may employ MIMO (multiple-input multiple output) technology and utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.
[0028] Base stations 114a and 114b may communicate with one or more WTRUs 102a, 102b, 102c, and 102d via an air interface 116, which may be any suitable wireless communication link (e.g., RF (radio frequency), microwave, centimeter wave, micrometer wave, IR (infrared), UV (ultraviolet), visible light, etc.). The air interface 116 may be established using any suitable RAT (radio access technology).
[0029] More specifically, as described above, the communication system 100 may be a multiple access system and may employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, and the like. For example, base stations 114a and WTRU 102a, 102b, and 102c in RAN 104 / 113 may implement radio technologies such as UTRA (UMTS (Universal Mobile Telecommunications System) Terrestrial Radio Access), which may establish air interfaces 115 / 116 / 117 using WCDMA (wideband CDMA). WCDMA may include communication protocols such as HSPA (High-Speed Packet Access) and / or HSPA+ (Evolved HSPA). HSPA may include HSDPA (High-Speed DL (Downlink) Packet Access) and / or HSUPA (High-Speed UL Packet Access).
[0030] In the embodiment, base stations 114a and WTRUs 102a, 102b, and 102c may implement radio technologies such as E-UTRA (Evolved UMTS Terrestrial Radio Access), which may establish an air interface 116 using, for example, LTE (Long Term Evolution) and / or LTE-A (LTE-Advanced) and / or LTE-A Pro (LTE-Advanced Pro).
[0031] In the embodiment, base stations 114a and WTRUs 102a, 102b, and 102c may implement radio technologies such as radio access for NR (New Radio), for example, to establish an air interface 116 using NR.
[0032] In the embodiment, base stations 114a and WTRUs 102a, 102b, and 102c may implement multiple radio access technologies. For example, base stations 114a and WTRUs 102a, 102b, and 102c may implement both LTE radio access and NR radio access using, for example, the DC (dual connectivity) principle. Thus, the air interface used by WTRUs 102a, 102b, and 102c may be characterized by multiple types of radio access technologies and / or transmissions sent to and from multiple types of base stations (e.g., eNBs and gNBs).
[0033] In other embodiments, base stations 114a and WTRUs 102a, 102b, 102c may implement wireless technologies such as IEEE 802.11 (i.e., WiFi (Wireless Fidelity)), IEEE 802.16 (i.e., WiMAX (Worldwide Interoperability for Microwave Access)), CDMA2000, CDMA2000 1X, CDMA2000EV-DO, IS-2000 (Interim Standard 2000), IS-95 (Interim Standard 95), IS-856 (Interim Standard 856), GSM (Global System for Mobile communications), EDGE (Enhanced Data rates for GSM Evolution), GERAN (GSM EDGE), and similar technologies.
[0034] In Figure 1A, base station 114b may be, for example, a wireless router, home Node B, home eNode B, or access point, and may utilize any RAT suitable for facilitating wireless connectivity in localized areas such as businesses, homes, vehicles, campuses, industrial facilities, aerial walkways (e.g., for drone use), and roadways. In one embodiment, base station 114b and WTRU 102c, 102d may implement wireless technology such as IEEE 802.11 to establish a WLAN (wireless local area network). In another embodiment, base station 114b and WTRU 102c, 102d may implement wireless technology such as IEEE 802.15 to establish a WPAN (wireless personal area network). In yet another embodiment, base stations 114b and WTRUs 102c, 102d may utilize cellular-based RATs (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish picocells or femtocells. As shown in Figure 1A, base station 114b may have a direct connection to the internet 110. Therefore, base station 114b may not be required to access the internet 110 via CNs 106 / 115.
[0035] RAN104 / 113 may be in communication with CN106 / 115 and may be any type of network configured to provide voice, data, applications, and / or VoIP (Voice over Internet Protocol) services to one or more of WTRU102a, 102b, 102c, and 102d. The data may have various QoS (Quality of Service) requirements, such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, and mobility requirements. CN106 / 115 may provide call control, billing services, mobile location-based services, prepaid calling, internet connectivity, video distribution, and / or high-level security functions, such as user authentication. Although not shown in Figure 1A, it will be understood that RAN104 / 113 and / or CN106 / 115 may be in direct or indirect communication with other RANs employing the same or different RAT as RAN104 / 113. For example, in addition to being connected to RAN104 / 113, which may be using NR radio technology, CN106 / 115 may also be in communication with another RAN (not shown) employing GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.
[0036] Furthermore, CN106 / 115 may also serve as a gateway for WTRU102a, 102b, 102c, and 102d to access PSTN108, the Internet 110, and / or other networks 112. PSTN108 may include a circuit-switched telephone network providing POTS (plain old telephone service). The Internet 110 may include a global system of interconnected computer networks and devices using common communication protocols such as TCP (transmission control protocol), UDP (user datagram protocol), and / or IP in the TCP / IP (Internet Protocol) suite. Network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, network 112 may include another CN connected to one or more RANs that may employ the same RAT as RAN104 / 113, or a different RAT.
[0037] Some or all of the WTRUs 102a, 102b, 102c, and 102d in the communication system 100 may include multimode capabilities (for example, WTRUs 102a, 102b, 102c, and 102d may include multiple transceivers to communicate with separate wireless networks via separate wireless links). For example, WTRU 102c, shown in Figure 1A, may be configured to communicate with base station 114a, which may employ cellular-based radio technology, and base station 114b, which may employ IEEE 802 radio technology.
[0038] Figure 1B is a system diagram illustrating an exemplary WTRU 102. As shown in Figure 1B, the WTRU 102 may include, among others, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, a non-removable memory 130, a removable memory 132, a power supply 134, a GPS (Global Positioning System) chipset 136, and / or other peripherals 138. It will be understood that the WTRU 102 may include any sub-combination of the elements described above, without inconsistency with the embodiment.
[0039] The processor 118 could be a general-purpose processor, a dedicated processor, a conventional processor, a DSP (digital signal processor), multiple microprocessors, one or more microprocessors with a DSP core, a controller, a microcontroller, an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) circuit, any other type of IC (integrated circuit), a state machine, etc. As suggested above, the processor 118 may include multiple processors. The processor 118 may perform signal coding, data processing, power control, input / output processing, and / or any other functions that enable the WTRU 102 to operate in a wireless environment. The processor 118 may be coupled to a transceiver 120 which may be coupled to a transmit / receive element 122. While Figure 1B depicts the processor 118 and transceiver 120 as separate components, it will be understood that the processor 118 and transceiver 120 may be integrated together in an electronic package or chip.
[0040] The transmit / receive element 122 may be configured to transmit or receive signals to or from a base station (e.g., base station 114a) via the air interface 116. For example, in one embodiment, the transmit / receive element 122 may be an antenna configured to transmit and / or receive RF signals. In an embodiment, the transmit / receive element 122 may be an emitter / detector configured to transmit and / or receive, for example, IR signals, UV signals, or visible light signals. In yet another embodiment, the transmit / receive element 122 may be configured to transmit and / or receive both RF and optical signals. It will be understood that the transmit / receive element 122 may be configured to transmit and / or receive any combination of wireless signals.
[0041] Although the transmit / receive element 122 is depicted as a single element in Figure 1B, the WTRU 102 may contain any number of transmit / receive elements 122. More specifically, the WTRU 102 may employ MIMO technology. Therefore, in one embodiment, the WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) to transmit and receive wireless signals via the air interface 116.
[0042] The transceiver 120 may be configured to modulate a signal that is to be transmitted by the transmit / receive element 122, and to demodulate a signal that is received by the transmit / receive element 122. As mentioned above, the WTRU 102 may have multimode capabilities. Therefore, for example, the transceiver 120 may include multiple transceivers to enable the WTRU 102 to communicate by multiple RATs, such as NR and IEEE 802.11.
[0043] The processor 118 of the WTRU102 may be connected to and receive user input data via a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., an LCD (liquid crystal display) display unit or an OLED (organic light-emitting diode) display unit). Furthermore, the processor 118 may output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. In addition, the processor 118 may access information and store data in any suitable type of memory, such as a non-removable memory 130 and / or removable memory 132. The non-removable memory 130 may include RAM (random-access memory), ROM (read-only memory), a hard disk, or any other type of memory storage device. The removable memory 132 may include a SIM (subscriber identity module) card, a Memory Stick, an SD (secure digital) memory card, and similar. In other embodiments, the processor 118 may access information and store data in memory, for example, a server or home computer (not shown), which is not physically located in the WTRU 102.
[0044] The processor 118 may receive power from a power supply 134 and may be configured to distribute and / or control power to other components in the WTRU 102. The power supply 134 can be any device suitable for supplying power to the WTRU 102. For example, the power supply 134 may include one or more dry cell batteries (e.g., NiCd (nickel-cadmium), NiZn (nickel-zinc), NiMH (nickel-metal hydride), Li-ion (lithium-ion), etc.), a solar cell, a fuel cell, etc.
[0045] Furthermore, the processor 118 may be coupled to a GPS chipset 136 which may be configured to provide location information (e.g., longitude and latitude) about the current location of the WTRU 102. In addition to or instead of the information from the GPS chipset 136, the WTRU 102 may determine its location based on receiving location information from base stations (e.g., base stations 114a, 114b) via the air interface 116 and / or based on the timing of signals received from two or more neighboring base stations. It will be understood that the WTRU 102 may acquire location information through any appropriate location-determination method, without being inconsistent with the embodiments.
[0046] Furthermore, the processor 118 may be coupled to other peripherals 138 which may include one or more software modules and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, peripherals 138 may include an accelerometer, e-compass, satellite transceiver, digital camera (for photos and / or video), USB (Universal Serial Bus) port, vibration device, television transceiver, hands-free headset, Bluetooth module, FM (frequency modulated) radio unit, digital music player, media player, video game player module, internet browser, VR / AR (virtual reality and / or augmented reality) device, activity tracker, etc. Peripherals 138 may include one or more sensors, which may be one or more of the following: gyroscope, accelerometer, Hall effect sensor, magnetometer, compass sensor, proximity sensor, temperature sensor, time sensor, geolocation sensor, altimeter, light sensor, touch sensor, magnetometer, barometer, gesture sensor, biometric sensor, and / or humidity sensor.
[0047] WTRU102 may include a full-duplex radio where some or all of the transmission and reception of signals (e.g., associated with a particular subframe with respect to both UL (e.g., for transmission) and downlink (e.g., for reception) may be in parallel and / or simultaneous. The full-duplex radio may include an interference management unit to reduce and / or substantially eliminate self-interference by either hardware (e.g., chokes) or signal processing by a processor (e.g., separate processors (not shown), or by processor 118). In an embodiment, WRTU102 may include a half-duplex radio where some or all of the transmission and reception of signals (e.g., associated with a particular subframe with respect to either UL (e.g., for transmission) or downlink (e.g., for reception) may be in parallel and / or simultaneously.
[0048] Figure 1C is a system diagram illustrating RAN104 and CN106 according to an embodiment. As described above, RAN104 may employ E-UTRA's wireless technology to communicate with WTRU102a, 102b, and 102c via the air interface 116. Furthermore, RAN104 may also be in communication with CN106.
[0049] RAN104 may include eNode-B160a, 160b, and 160c, but it will be understood that RAN104 may include any number of eNode-B without being inconsistent with the embodiment. Each of eNode-B160a, 160b, and 160c may include one or more transceivers to communicate with WTRU102a, 102b, and 102c via the air interface 116. In one embodiment, eNode-B160a, 160b, and 160c may implement MIMO technology. For example, eNode-B160a may use multiple antennas to transmit and / or receive wireless signals to WTRU102a.
[0050] Each of the eNode-B160a, 160b, and 160c may be associated with a specific cell (not shown) and may be configured to handle decisions regarding radio resource management, handover decisions, user scheduling in UL and / or DL, etc. As shown in Figure 1C, the eNode-B160a, 160b, and 160c may communicate with each other via the X2 interface.
[0051] The CN106 shown in Figure 1C may include an MME (mobility management entity) 162, an SGW (serving gateway) 164, and a PDN (packet data network) gateway (or PGW) 166. While each of the above elements is depicted as part of CN106, it should be understood that any of the elements described may be owned and / or operated by an entity other than the CN operator.
[0052] MME162 may be connected to each of the eNode-B162a, 162b, and 162c in RAN104 via the S1 interface and may act as a control node. For example, MME162 may be responsible for authenticating users of WTRU102a, 102b, and 102c, activating / deactivating bearers, selecting a specific serving gateway during the initial attachment of WTRU102a, 102b, and 102c, and similar matters. MME162 may provide control plane functionality for switching between RAN104 and other RANs (not shown) employing other radio technologies such as GSM and / or WCDMA.
[0053] SGW164 may be connected to each of the eNode B160a, 160b, and 160c in RAN104 via the S1 interface. Generally, SGW164 may route and forward user data packets to WTRU102a, 102b, and 102c. SGW164 may also perform other functions, such as fixing the user plane during eNode B handovers, triggering paging when DL data is available to WTRU102a, 102b, and 102c, and managing and storing the context of WTRU102a, 102b, and 102c.
[0054] SGW164 may be connected to PGW166, which may provide WTRU102a, 102b, and 102c with access to a packet-switched network, such as the Internet 110, in order to facilitate communication between WTRU102a, 102b, and 102c and IP-enabled devices.
[0055] CN106 may facilitate communication with other networks. For example, CN106 may provide WTRU102a, 102b, and 102c with access to a circuit-switched network, such as PSTN108, to facilitate communication between WTRU102a, 102b, and 102c and conventional terrestrial communication line communication devices. For example, CN106 may include, or communicate with, an IP gateway (e.g., an IMS (IP Multimedia Subsystem) server) that acts as an interface between CN106 and PSTN108. Furthermore, CN106 may provide WTRU102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.
[0056] Although the WTRU is described as a wireless terminal in Figures 1A to 1D, in a typical embodiment, the terminal may use a wired communication interface with a communication network (for example, temporarily or permanently).
[0057] In a typical embodiment, the other network 112 may be a WLAN.
[0058] In the BSS (Basic Service Set) mode of infrastructure, a WLAN may have an AP (Access Point) for the BSS and one or more STAs (Stations) associated with the AP. The AP may have access to or interfaces to a DS (Distribution System), or another type of wired / wireless network that carries traffic entering and leaving the BSS. Traffic originating outside the BSS and destined for the STA may arrive through the AP and be delivered to the STA. Traffic originating from the STA and destined for destinations outside the BSS may be sent to the AP and delivered to their respective destinations. Traffic between STAs within the BSS may be sent through the AP, for example, when the originating STA sends traffic to the AP, and the AP delivers the traffic to the destination STA. Traffic between STAs within the BSS may be considered and / or referred to as peer-to-peer traffic. Peer-to-peer traffic may be sent between the originating and destination STAs (for example, directly between them) via a DLS (direct link setup). In a typical embodiment, the DLS may be 802.11e DLS or 802.11z TDLS (tunneled DLS). A WLAN using IBSS (Independent BSS) mode may not have APs, and STAs within or using IBSS (e.g., all STAs) may communicate directly with each other. The IBSS mode of communication is sometimes referred to herein as the "ad-hoc" mode of communication.
[0059] When using the 802.11ac infrastructure mode of operation or a similar mode of operation, an AP may transmit beacons on a fixed channel, such as the primary channel. The primary channel may have a fixed width (e.g., a 20 MHz bandwidth) or a width dynamically set by signaling. The primary channel may be the operating channel of the BSS and may be used by the STA to establish a connection with the AP. In one typical embodiment, CSMA / CA (Carrier Sensing Multiple Access / Collision Avoidance) may be implemented in the 802.11 system, for example. With respect to CSMA / CA, an STA, including the AP (e.g., any STA), may sense the primary channel. If the primary channel is sensed / detected and / or determined to be busy by a particular STA, that particular STA may back off. A single STA (e.g., just one station) may transmit at any given time in a given BSS.
[0060] HT (High Throughput) STAs may use a 40MHz wide channel for communication by combining a 20MHz primary channel with adjacent or non-adjacent 20MHz channels, for example, to form a 40MHz wide channel.
[0061] A VHT (Very High Throughput) STA may support channels with widths of 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz. 40 MHz and / or 80 MHz channels may be constructed by combining contiguous 20 MHz channels. A 160 MHz channel may be constructed by combining eight contiguous 20 MHz channels, or by combining two non-contiguous 80 MHz channels, which may be in an 80+80 configuration. In the 80+80 configuration, data may proceed to a segment parser that splits the data into two streams after channel encoding. Inverse Fast Fourier Transform (IFFT) processing and time-domain processing may be performed separately on each stream. The streams may be mapped onto two 80 MHz channels, and data may be transmitted by the transmitting STA. In the receiving STA receiver, the operation for the 80+80 configuration described above may be inverted, and the combined data may be sent to MAC (Media Access Control).
[0062] A sub-1GHz mode of operation is supported by 802.11af and 802.11ah. The operating bandwidth and carrier of the channel are reduced in 802.11af and 802.11ah compared to those used in 802.11n and 802.11ac. 802.11af supports 5MHz, 10MHz, and 20MHz bandwidths in the TVWS (TV White Space) spectrum, while 802.11ah supports 1MHz, 2MHz, 4MHz, 8MHz, and 16MHz bandwidths using the non-TVWS spectrum. In a typical embodiment, 802.11ah may support Meter Type Control / Machine-Type Communication, for example, MTC devices in macro coverage areas. MTC devices may have limited performance, such as support for certain and / or limited bandwidths (e.g., support only). MTC devices may contain batteries with battery life exceeding a threshold (for example, to maintain a very long battery life).
[0063] A WLAN system that may support multiple channels and channel bandwidths, such as 802.11n, 802.11ac, 802.11af, and 802.11ah, may include a channel that can be designated as the primary channel. The primary channel may have a bandwidth equal to the largest common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel may be set and / or limited by the STA that supports the smallest bandwidth operating mode among all STAs operating in the BSS. In the case of 802.11ah, even if the AP and other STAs in the BSS support operating modes of 2MHz, 4MHz, 8MHz, 16MHz, and / or other channel bandwidths, the primary channel may be 1MHz wide for an STA (e.g., an MTC type device) that supports (e.g., only supports) the 1MHz mode. Carrier sensing and / or NAV (Network Allocation Vector) settings may depend on the state of the primary channel. If the primary channel is busy, for example, due to an STA (which only supports a 1MHz operating mode) transmitting to the AP, then the entire available frequency band may be considered busy, even if a large portion of the frequency band remains idle and available.
[0064] In the United States, the available frequency band that can be used by 802.11ah is from 902 MHz to 928 MHz. In South Korea, the available frequency band is from 917.5 MHz to 923.5 MHz. In Japan, the available frequency band is from 916.5 MHz to 927.5 MHz. The total bandwidth available for 802.11ah is from 6 MHz to 26 MHz, depending on the country code.
[0065] Figure 1D is a system diagram illustrating RAN113 and CN115 according to an embodiment. As described above, RAN113 may employ NR radio technology to communicate with WTRU102a, 102b, and 102c via the air interface 116. Furthermore, RAN113 may also be in communication with CN115.
[0066] RAN113 may include gNB180a, 180b, and 180c, but it will be understood that RAN113 may include any number of gNBs without being inconsistent with the embodiment. Each of the gNB180a, 180b, and 180c may include one or more transceivers to communicate with WTRU102a, 102b, and 102c via the air interface 116. In one embodiment, the gNB180a, 180b, and 180c may implement MIMO technology. For example, the gNB180a and 108b may utilize beamforming to transmit and / or receive signals to the gNB180a, 180b, and 180c. Thus, for example, the gNB180a may use multiple antennas to transmit and / or receive wireless signals to the WTRU102a. In embodiments, gNB180a, 180b, and 180c may implement carrier aggregation techniques. For example, gNB180a may transmit multiple component carriers to WTRU102a (not shown). The subset of component carriers described above may lie on the unallowed spectrum while the remaining component carriers may lie on the allowed spectrum. In embodiments, gNB180a, 180b, and 180c may implement Coordinated Multi-Point (CoMP) techniques. For example, WTRU102a may receive coordinated transmissions from gNB180a and gNB180b (and / or gNB180c).
[0067] WTRU102a, 102b, and 102c may communicate with gNB180a, 180b, and 180c using transmissions associated with scalable numerology. For example, OFDM symbol spacing and / or OFDM subcarrier spacing may vary for separate transmissions, separate cells, and / or separate portions of the spectrum of wireless transmissions. WTRU102a, 102b, and 102c may communicate with gNB180a, 180b, and 180c using subframes or TTI (Transmission Time Interval) of varying or scalable lengths (e.g., including varying numbers of OFDM symbols and / or persistent absolute times of varying lengths).
[0068] gNB180a, 180b, and 180c may be configured to communicate with WTRU102a, 102b, and 102c in standalone and / or non-standalone configurations. In a standalone configuration, WTRU102a, 102b, and 102c may communicate with gNB180a, 180b, and 180c without accessing other RANs (e.g., eNode-B160a, 160b, and 160c). In a standalone configuration, WTRU102a, 102b, and 102c may use one or more of gNB180a, 180b, and 180c as a mobility anchor point. In a standalone configuration, WTRU102a, 102b, and 102c may communicate with gNB180a, 180b, and 180c using signals in unauthorized bandwidths. In non-standalone configurations, WTRU102a, 102b, and 102c may communicate with / connect to gNB180a, 180b, and 180c while simultaneously communicating with / connecting to another RAN, such as eNode-B160a, 160b, and 160c. For example, WTRU102a, 102b, and 102c may implement DC principles to communicate substantially simultaneously with one or more gNB180a, 180b, and 180c and one or more eNode-B160a, 160b, and 160c. In non-standalone configurations, eNode-B160a, 160b, and 160c may act as mobility anchors for WTRU102a, 102b, and 102c, while gNB180a, 180b, and 180c may provide additional coverage and / or throughput to service WTRU102a, 102b, and 102c.
[0069] Each of the gNB180a, 180b, and 180c may be associated with a specific cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, support for network slicing, dual connectivity, interworking between NR and E-UTRA, routing of user plane data to UPF (User Plane Function) 184a and 184b, and routing of control plane information to AMF (Access and Mobility Management Function) 182a and 182b. As shown in Figure 1D, the gNB180a, 180b, and 180c may communicate with each other via the Xn interface.
[0070] The CN115 shown in Figure 1D may include at least one AMF182a, 182b, at least one UPF184a, 184b, at least one SMF (Session Management Function)183a, 183b, and possibly a DN (Data Network)185a, 185b. While each of the above elements is depicted as part of the CN115, it will be understood that any of the elements described may be owned and / or operated by an entity other than the CN operator.
[0071] AMF182a and 182b may be connected to one or more of gNB180a, 180b, and 180c in RAN113 via the N2 interface and may act as control nodes. For example, AMF182a and 182b may be responsible for authenticating users of WTRU102a, 102b, and 102c, supporting network slicing (e.g., handling sessions of separate PDUs with distinct requirements), selecting specific SMF183a and 183b, managing registration areas, terminating NAS signaling, and mobility management. Network slicing may be used by AMF182a and 182b to customize CN support for WTRU102a, 102b, and 102c based on the type of services used by WTRU102a, 102b, and 102c. For example, separate network slices may be established for separate use cases, such as services that rely on URLLC (ultra-high reliability, low latency) access, services that rely on eMBB (enhanced massive mobile broadband) access, or services related to MTC (machine-type communication) access. The AMF162 may provide control plane functionality for switching between RAN113 and other RANs (not shown) employing other radio technologies such as LTE, LTE-A, LTE-A Pro and / or non-3GPP access technologies such as WiFi.
[0072] SMF183a and 183b may be connected to AMF182a and 182b in CN115 via the N11 interface. Furthermore, SMF183a and 183b may be connected to UPF184a and 184b in CN115 via the N4 interface. SMF183a and 183b may select and control UPF184a and 184b and configure the routing of traffic passing through them. SMF183a and 183b may also perform other functions, such as managing and assigning IP addresses to UEs, managing PDU sessions, controlling policy enforcement and QoS, and providing downlink data notifications. PDU session types can be IP-based, non-IP-based, Ethernet-based, etc.
[0073] UPF184a and 184b may be connected to one or more of gNB180a, 180b, and 180c in RAN113 via the N3 interface, and may provide WTRU102a, 102b, and 102c with access to a packet-switched network, such as the Internet 110, to facilitate communication between WTRU102a, 102b, and 102c and IP-enabled devices. UPF184 and 184b may perform other functions, such as routing and forwarding packets, enforcing user plane policies, supporting sessions for multi-homed PDUs, handling user plane QoS, buffering downlink packets, and providing mobility anchoring.
[0074] CN115 may facilitate communication with other networks. For example, CN115 may include, or communicate with, an IP gateway (e.g., an IMS (IP Multimedia Subsystem) server) that acts as an interface between CN115 and PSTN108. In addition, CN115 may provide WTRU102a,102b,102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, WTRU102a,102b,102c may be connected to local DN185a,185b via UPF184a,184b through an N3 interface to UPF184a,184b and an N6 interface between UPF184a,184b and DN (Data Network) 185a,185b.
[0075] In terms of Figures 1A-1D and the descriptions corresponding to Figures 1A-1D, any WTRU102a-d, base stations 114a-b, eNode-B160a-c, MME162, SGW164, PGW166, gNB180a-c, AMF182a-b, UPF184a-b, SMF183a-b, DN185a-b, and / or any other device(s) described herein may, in relation to one or more of them, perform one or more or all of the functions described herein by one or more emulation devices (not shown). An emulation device may be one or more devices configured to emulate one or more or all of the functions described herein. For example, an emulation device may be used to test other devices and / or to simulate network and / or WTRU functions.
[0076] Emulation devices may be designed to implement one or more tests of other devices in a lab environment and / or an operator's network environment. For example, one or more emulation devices may perform one, more, or all functions while being implemented and / or deployed as part of a wired and / or wireless communication network, either entirely or partially, to test other devices in a communication network. One or more emulation devices may perform one, more, or all functions while being temporarily implemented and / or deployed as part of a wired and / or wireless communication network. Emulation devices may be directly coupled to another device for testing purposes and / or testing may be performed using over-the-air (OTA) wireless communication.
[0077] One or more emulation devices may perform one or more functions, including all of the above, while not implemented / deployed as part of a wired and / or wireless communication network. For example, an emulation device may be used in a testing scenario in a testing laboratory and / or a wired and / or wireless communication network that is not deployed (e.g., for testing) to implement testing of one or more components. One or more emulation devices may be test equipment. Direct RF coupling and / or wireless communication via RF circuitry (which may include, for example, one or more antennas) may be used by emulation devices to transmit and / or receive data.
[0078] This application describes various aspects, including tools, features, examples, models, and approaches. Many of these aspects are described in a specific manner and, at least to illustrate their individual characteristics, often in a way that may give the impression of limitation. However, this is for the purpose of clarity in the description and does not limit the application or scope of these aspects. Indeed, all the different aspects may be combined and interchangeable to provide further aspects. Furthermore, these aspects may also be combined and interchangeable with aspects described in earlier applications.
[0079] The embodiments described and envisioned in this application may be implemented in many different ways. Figures 5–9 described herein may provide some examples, but other examples are envisioned. The discussion of Figures 5–9 is not intended to limit the scope of implementation. Generally, at least one of the embodiments relates to video encoding and decoding, and generally, at least one other embodiment relates to transmitting the generated or encoded bitstream. The embodiments described herein and other embodiments may be implemented as computer-readable recording media storing methods, apparatus, instructions for encoding or decoding video data according to any of the described methods, and / or computer-readable recording media storing the bitstream generated according to any of the described methods.
[0080] In this application, the terms “reconstructed” and “decoded” may be used interchangeably, the terms “pixel” and “sample” may be used interchangeably, and the terms “image,” “picture,” and “frame” may be used interchangeably.
[0081] Various methods are described herein, each of which includes one or more steps or actions to achieve the described method. Unless a particular order of steps or actions is required for the inherent operation of the method, the order and / or use of certain steps and / or actions may be modified or combined. In addition, terms such as “first,” “second,” etc., may be used in various examples of modifying elements, components, steps, actions, etc., such as “first decode” and “second decode.” The use of such terms does not imply ordering the modified operations unless specifically required. Thus, in the examples given, the first decode does not need to be performed before the second decode, and may occur, for example, before the second decode, during the second decode, or in a time period overlapping with the second decode.
[0082] Various methods and other embodiments described herein may be used to modify modules of the video encoder 200 and decoder 300, such as the decoding module, as shown in Figures 2 and 3. Furthermore, the subject matter disclosed herein may apply to any type, format, or version of video encoding, whether existing or future development, whether described in standards or recommendations, and to extensions of the above standards and recommendations. Unless otherwise indicated or unless particularly technically impeded, the embodiments described herein may be used individually or in combination.
[0083] For example, various numerical values such as bit count and bit depth are used in the examples described in this application. The values mentioned above and other specific values are for illustrative purposes only, and the embodiments described are not limited to the specific values mentioned above.
[0084] Figure 2 shows an exemplary video encoder. Variations of the exemplary encoder 200 are assumed, but encoder 200 is described below for clarity without describing all expected variations.
[0085] Before encoding, a video sequence may undergo pre-encoding (201), which may involve, for example, applying color conversions to the input color picture (e.g., RGB4:4:4 to YCbCr4:2:0 conversion) or remapping input picture components to make the signal distribution more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata may be associated with pre-processing and linked to the bitstream.
[0086] In encoder 200, the picture is encoded by the encoder elements described below. The picture to be encoded is divided (202) and processed by units of coding units (CUs), for example. Each unit is encoded using either intra-mode or inter-mode, for example. When a unit is encoded in intra-mode, intra-prediction (260) is performed. In inter-mode, motion estimation (275) and motion compensation (270) are performed. The encoder decides (205) whether to use intra-mode or inter-mode to encode the unit, and indicates the intra / inter decision by, for example, a prediction mode flag. The prediction residual is calculated by, for example, subtracting the predicted block from the original image block (210).
[0087] Next, the predicted residual is transformed (225) and quantized (230). The quantized transformed coefficients, as well as the motion vector and other syntax elements, are entropy-encoded to output a bitstream (245). The encoder can skip the transformation and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transformation and quantization, i.e., the residual is encoded directly without any transformation or quantization processes applied.
[0088] The encoder decodes the encoded blocks to provide a reference for further prediction. The quantized transformation coefficients are inversely quantized (240) and inversely transformed (250) to decode the prediction residuals. The decoded prediction residuals and the predicted blocks are combined (255) to reconstruct the image blocks. An in-loop filter (265) is applied to the reconstructed picture to perform deblocking / SAO (Sample Adaptive Offset) filtering, for example, to reduce encoding artifacts. The filtered image is stored in a reference picture buffer (280).
[0089] Figure 3 shows an example of a video decoder. In the exemplary decoder 300, the bitstream is decoded by decoder elements as described below. Generally, the video decoder 300 performs a decoding pass which is the reverse of the encoding pass described in Figure 2. Furthermore, generally, the encoder 200 also performs video decoding as part of encoding the video data.
[0090] In particular, the decoder input includes a video bitstream that may be generated by the video encoder 200. First, the bitstream is entropy-decoded (330) to obtain transformation coefficients, motion vectors, and other encoded information. Picture partitioning information indicates how the picture will be partitioned. Therefore, the decoder may partition (335) the picture according to the decoded picture partitioning information. The transformation coefficients are inversely quantized (340) and inversely transformed (350) to decode the prediction residuals. The image blocks are reconstructed by combining (355) the decoded prediction residuals with the predicted blocks. The predicted blocks may be obtained (370) from intra-predictions (360) or motion-compensated predictions (i.e., inter-predictions) (375). An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (380).
[0091] Furthermore, the decoded picture can undergo post-decoding (385), such as reverse color conversion (e.g., conversion from YCbCr4:2:0 to RGB4:4:4) or reverse remapping, which is the reverse of the remapping performed in pre-encoding (201). Post-decoding can use metadata signaled in the bitstream derived in pre-encoding. In the example, the decoded image (e.g., after applying an in-loop filter (365) and / or after post-decoding (385) if post-decoding is used) may be sent to a display device for rendering to the user.
[0092] Figure 4 shows an example of a system in which various embodiments and examples described herein may be implemented. System 400 may be embodied as a device including various components described below and configured to perform one or more embodiments described herein. Examples of the above device include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home electrical appliances, and servers. The elements of System 400 may be embodied individually or in combination as a single integrated circuit (IC), multiple ICs, and / or separate components. For example, in at least one example, the processing and encoder / decoder elements of System 400 are distributed across multiple ICs and / or separate components. In various examples, System 400 is communicated to one or more other systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various examples, System 400 is configured to implement one or more embodiments described herein.
[0093] System 400 includes, for example, at least one processor 410 configured to execute instructions loaded to implement various embodiments described in this document. The processor 410 may include embedded memory, input / output interfaces, and various other circuits as known in the Art. System 400 includes at least one memory 420 (e.g., a volatile memory device and / or a non-volatile memory device). System 400 includes a storage device 440 which may include non-volatile memory and / or volatile memory, including, but not limited to, EEPROM (Electrically Erasable Programmable Read-Only Memory), ROM (Read-Only Memory), PROM (Programmable Read-Only Memory), RAM (Random Access Memory), DRAM (Dynamic Random Access Memory), SRAM (Static Random Access Memory), flash, magnetic disk drives, and / or optical disk drives. Storage device 440 may include, in non-limiting examples, internal storage devices, attached storage devices (including detachable and non-detachable storage devices), and / or network-accessible storage devices.
[0094] System 400 includes, for example, an encoder / decoder module 430 configured to process data that provides encoded or decoded video, the encoder / decoder module 430 which may include its own processor and memory. Encoder / decoder module 430 represents a module(s) that may be included in a device that performs encoding and / or decoding functions. As is known, the device may include one or both of the encoding and decoding modules. In addition, the encoder / decoder module 430 may be implemented as a separate element of System 400, or it may be incorporated into the processor 410 as a combination of hardware and software as known to those skilled in the art.
[0095] Program code loaded onto a processor 410 or encoder / decoder 430 performing various actions described herein may be stored in a storage device 440 and subsequently loaded into memory 420 for execution by the processor 410. According to various examples, one or more processors 410, memory 420, storage device 440, and encoder / decoder modules 430 may store one or more items during the execution of the processes described herein. These stored items may include, but are not limited to, input video, decoded video, or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.
[0096] In some examples, the internal memory of the processor 410 and / or the encoder / decoder module 1030 is used to store instructions and provide working memory for the processing required during encoding or decoding. However, in other embodiments, external memory of the processing device (for example, the processing device may be either the processor 410 or the encoder / decoder module 430) is used for one or more of the functions described above. The external memory may be memory 420 and / or storage device 440, for example, dynamic volatile memory and / or non-volatile flash memory. In some examples, external non-volatile flash memory is used to store, for example, the television's operating system. In at least one example, high-speed external dynamic volatile memory, such as RAM, is used as working memory for video encoding and decoding operations.
[0097] Inputs to the elements of System 400 may be provided via various input devices, as shown in Block 445. These input devices include, but are not limited to, (i) an RF section for receiving, for example, RF (radio frequency) signals transmitted by a broadcasting station via radio waves, (ii) a COMP (Component) input terminal (or a set of COMP input terminals), (iii) a USB (Universal Serial Bus) input terminal, and / or (iv) an HDMI (High Definition Multimedia Interface) input terminal. Other examples not shown in Figure 4 include composite video.
[0098] In various examples, the input device of block 445 associates various input processing elements known in the art. For example, the RF portion may be associated with an element suitable for (i) selecting a desired frequency (also called selecting a signal or band-limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) again band-limiting a single frequency band, which in some examples may be called a channel, to a narrower frequency band, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and / or (vi) demultiplexing to select a desired stream for data packets. The RF portion of various examples includes one or more elements that perform the functions described above, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error collectors, and demultiplexers. The RF portion may also include a tuner that performs various functions, such as down-converting a received signal to a lower frequency (e.g., an intermediate frequency or a frequency close to the baseband) or to the baseband. In one example set-top box, the RF section and associated input processing elements receive, filter, down-convert, and filter again to the desired frequency band, thereby performing frequency selection. Various examples involve rearranging the order of the elements described above (and others), removing some of the elements mentioned, and / or adding other elements that perform similar or different functions. Adding elements can include inserting elements between existing elements, such as inserting amplifiers and analog-to-digital converters. In various examples, the RF section includes an antenna.
[0099] USB and / or HDMI terminals may include their respective interface processors to connect system 400 to other electronic devices over USB and / or HDMI connections. It is understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented, for example, in a separate input processing IC or, if necessary, in processor 11010. Similarly, aspects of USB or HDMI interface processing may be implemented in a separate interface IC or, if necessary, in processor 11010. Demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, for example, processor 410 and encoder / decoder 430, which works in cooperation with memory and storage elements to process the data stream for submission to output devices as necessary.
[0100] Various elements of system 400 may be provided within an integrated housing, where the various elements may be interconnected and transmit data between them using an appropriate interconnection array 425, such as an I2C (Inter-IC) bus, wiring, and an internal bus known in the art, including a printed circuit board.
[0101] System 400 includes a communication interface 450 that enables communication with other devices via a communication channel 460. The communication interface 450 may, but is not limited to, include a transceiver configured to transmit and receive data via the communication channel 460. The communication interface 450 may, but is not limited to, include a modem or a network card, and the communication channel 460 may be implemented, for example, in a wired and / or wireless medium.
[0102] In various examples, data is streamed to system 400 or otherwise provided using a wireless network, such as a Wi-Fi network, e.g., IEEE 802.11 (IEEE stands for Institute of Electrical and Electronics Engineers). In the example described above, the Wi-Fi signal is received via a communication channel 460 and a communication interface 450, which are modified to accommodate Wi-Fi communication. Typically, the communication channel 460 in the example described above is connected to an access point or router that provides access to an external network, including the Internet, to enable streaming applications and other over-the-top communications. In other examples, the streamed data is provided to system 400 using a set-top box that distributes the data via the HDMI connection of input block 445. Still, in other examples, the streamed data is provided to system 440 using the RF connection of input block 445. As shown above, various examples provide data in ways other than streaming. In addition, various examples use wireless networks other than Wi-Fi, e.g., cellular networks or Bluetooth networks.
[0103] System 400 can provide output signals to various output devices, including a display 475, a speaker 485, and other peripheral devices 495. Various examples of the display 475 include, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 475 may be for a television, tablet, laptop, cell phone (mobile phone), or other device. Furthermore, the display 475 may be integrated with other components (e.g., in a smartphone) or separate (e.g., with an external monitor for a laptop). Various examples of the other peripheral devices 495 include one or more of a standalone digital video disc (or digital versatile disc) (DVD for both terms), a disc player, a stereo system, and / or a lighting system. Various examples utilize one or more peripheral devices 495 that provide functionality based on the output of System 400. For example, a disc player plays the role of playing the output of System 400.
[0104] In various examples, control signals are communicated between the system 400 and the display 475, speaker 485, or other peripheral devices 495 using signaling such as AV.Link, CEC (Consumer Electronics Control), or other communication protocols that enable device-to-device control with or without user intervention. Output devices may be connected to the system 400 via dedicated connections through their respective interfaces 470, 480, and 490. Alternatively, output devices may be connected to the system 400 via communication interface 450 using communication channel 460. The display 475 and speaker 485 may be integrated into a single unit with other components of the system 400, for example, in an electronic device such as a television. In various examples, the display interface 470 includes a display driver, such as a timing controller (TCon) chip.
[0105] Alternatively, the display 475 and speaker 485 can be isolated from one or more other components, for example, if the RF portion of input 445 is part of a separate set-top box. In various examples where the display 475 and speaker 485 are external components, the output signals may be submitted via a dedicated output connection, such as an HDMI port, a USB port, or a COMP output.
[0106] The example may be implemented by computer software implemented by the processor 410, or by hardware, or by a combination of hardware and software. In a non-limiting example, the example may be implemented by one or more integrated circuits. The memory 420 may be of any type appropriate for the technical environment and may be implemented using any suitable data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, for example, in a non-limiting example. The processor 410 may be of any type appropriate for the technical environment and may include, in a non-limiting example, one or more of a microprocessor, a general-purpose computer, a dedicated computer, and a processor based on a multi-core architecture.
[0107] Various implementations include decoding. “Decoding,” as used in this application, can encompass all or part of the processing performed on the received encoded sequence to produce a final output suitable for display. In various examples, the processing includes one or more of the processing typically performed by a decoder, e.g., entropy decoding, inverse quantization, inverse transform, and differential decoding. In various examples, the processing may further include, or may not include, processing performed by a decoder in any of the various implementations described in this application, e.g., obtaining a first intra-prediction signal associated with a first intra-prediction mode for a block, obtaining a second intra-prediction signal associated with a second intra-prediction mode for a block, generating a prediction sample associated with a block by determining a first weight associated with the first intra-prediction signal and a second weight associated with the second intra-prediction signal, wherein the first and second weights are determined by a first neighbor block prediction mode and a second neighbor block prediction mode, and decoding the block based on the prediction sample.
[0108] For further examples, in one example, "decoding" refers only to "entropy decoding," in another example, "decoding" refers only to differential decoding, and in yet another example, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to refer specifically to a subset of operations or to refer more broadly in general, it is believed that decoding processes will become clear from the context of a particular description and will be well understood by those skilled in the art.
[0109] Various implementations include encoding. In a manner similar to the above explanation of "decoding," "encoding" as used in this application can encompass, for example, all or part of the processing performed on the input video sequence to generate the encoded bitstream. In various examples, the above processing includes one or more of the processing typically performed by an encoder, such as splitting, differential encoding, transformation, quantization, and entropy encoding. In various examples, the above processing may further include, or may be performed by encoders of various implementations described herein, for example, obtaining a first intra-prediction signal associated with a first intra-prediction mode for a block; obtaining a second intra-prediction signal associated with a second intra-prediction mode for a block; generating a predicted sample associated with a block by determining a first weight associated with the first intra-prediction signal and a second weight associated with the second intra-prediction signal, wherein the first and second weights are determined by a first neighbor block prediction mode and a second neighbor block prediction mode; deriving a residual sample associated with a block based on the predicted sample; and encoding a block based on the derived residual sample.
[0110] For further examples, in one example, "encoding" refers only to "entropy decoding," in another example, "encoding" refers only to differential decoding, and in yet another example, "encoding" refers to a combination of differential encoding and entropy encoding. Whether the phrase "encoding process" is intended to refer specifically to a subset of operations or to refer more broadly in general, it is believed that encoding processes will become clear from the context of a particular description and will be well understood by those skilled in the art.
[0111] When a diagram is given as a flowchart, it should be understood that it also provides a block diagram of the corresponding device. Similarly, when a diagram is given as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / process.
[0112] The implementations and embodiments described herein may be implemented, for example, in the form of methods or processes, apparatus, software programs, data streams, or signals. Even when described only in the context of a single form of implementation (e.g., only as a method), the implementation of the described features may also be implemented in other forms (e.g., apparatus or programs). Apparatus may be implemented, for example, in appropriate hardware, software, and firmware. Methods may be implemented in processors, for example, generally referring to processing devices, including computers, microprocessors, integrated circuits, or programmable logic devices. Furthermore, processors also include communication devices, such as computers, cell phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate the communication of information between end users.
[0113] References to “one example,” “example,” “one implementation,” or “implementation” mean, as with other variations, that the specific features, structures, properties, etc. described in relation to the example are included in at least one example. Therefore, the appearances of the phrase “in one example,” “in one example,” “in one implementation,” or “in implementation,” which appear in various places throughout this application, do not necessarily refer to all of the same examples, as with any other variations.
[0114] In addition, this application may use the term “determining” various parts of information. Determining information may include, for example, one or more of the following: estimating information, calculating information, predicting information, or retrieving information from memory. Obtaining may include receiving, searching, constructing, generating, and / or retrieving.
[0115] Furthermore, this application may use the term “accessing” various parts of information. Accessing information can include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0116] In addition, this application may use the term “receiving” various parts of information. Receiving is intended to be a broad term, as with respect to “accessing.” Receiving information can include, for example, accessing information or retrieving information (for example, from memory). Furthermore, “receiving” usually includes, in some way or another, actions such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0117] It should be understood that the use of any of the following " / ", "and / or", and "at least one of" is intended to include, for example, "A / B", "A and / or B", and "at least one of A and B", which include selecting only the first enumerated option (A), or only the second enumerated option (B), or both options (A and B). As further examples, in the case of "A, B, and / or C" and "at least one of A, B, and C", the above phrases are intended to include selecting only the first enumerated option (A), or only the second enumerated option (B), or only the third enumerated option (C), or only the first and second enumerated options (A and B), or only the first and third enumerated options (A and C), or the second and third enumerated options (B and C), or all three options (A, B, and C). What has been stated above may be extended to matters merely listed, as will be obvious to those skilled in the art and related businesses.
[0118] Furthermore, as used herein, the term "signaling" refers, in particular, to indicating something to a corresponding decoder. Encoder signals may include, for example, an encoding function for the input of a block using precision coefficients. In this way, the same parameters are used on both the encoder and decoder sides. Therefore, for example, an encoder can transmit specific parameters to a decoder (explicit signaling), and the decoder can use the same specific parameters. Conversely, if the decoder already has the same specific parameters, signaling may be used without transmission (implicit signaling) simply to allow the decoder to know and select the specific parameters. By avoiding transmission in any real-world function, bit savings are achieved in various examples. It is understood that signaling can be performed in various ways. For example, one or more syntax elements, flags, etc., are used in various examples to signal information to the corresponding decoder. The above discussion concerns the verb form of the word "signaling," but the word "signaling" may also be used as a noun in this specification.
[0119] As will be apparent to those skilled in the art, implementations may generate various signals formatted to carry information that may be stored or transmitted. For example, the information may include instructions for performing a method or data generated by one of the implementations described. For example, a signal may be formatted to carry the bitstream of the example described. For example, the above signal may be formatted as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. For example, formatting may include encoding a data stream and modulating a carrier wave with the encoded data stream. For example, the information carried by a signal may be analog or digital information. Signals may be transmitted over various separate wired or wireless links, as is known. Signals may be stored, accessed, or retrieved in a processor-readable medium.
[0120] Many examples are described herein. The features of the examples may be provided individually or in any combination across various claim categories and types. Furthermore, the examples may include one or more of the features, devices, or embodiments described herein, individually or in any combination, across various claim categories and types. For example, the functions described herein may be implemented in a bitstream or signal containing information generated as described herein. The information may enable a decoder to decode the bitstream, which is an encoder, bitstream, and / or decoder, according to any of the embodiments described herein. For example, the functions described herein may be implemented by creating and / or transmitting and / or receiving and / or decoding a bitstream or signal. For example, the functions described herein may be implemented in a method, process, apparatus, medium for storing instructions, medium for storing data, or signal. For example, the functions described herein may be implemented by a television, set-top box, cell phone, tablet, or other electronic device that performs decoding. Televisions, set-top boxes, cell phones, tablets, or other electronic devices may display the resulting image (e.g., an image from the residual reconstruction of a video bitstream) (e.g., using a monitor, screen, or other type of display). Televisions, set-top boxes, cell phones, tablets, or other electronic devices receive and decode signals containing encoded images.
[0121] The examples described above may be performed by a device having at least one processor. The device could be an encoder or a decoder. The examples described above may be performed by a computer program product containing program code instructions stored on a non-temporary computer-readable medium. The examples described above may be performed by a computer program containing program code instructions.
[0122] Examples of intra-CIIP are provided herein. Intra-CIIP uses a mechanism similar to or the same as CIIP, but may replace interprediction with an intra-template prediction (intra-TMP) mode. What has been described above may enable the use of CIIP for intra-coded blocks and intra-slices, and may provide meaningful coding gain. Intra-CIIP may be performed using blending weights determined based on the intra-prediction mode(s) of neighboring blocks(s). Intra-CIIP may be performed by transform coding. Intra-CIIP may be performed by IBC. Intra-CIIP may be performed by sign prediction. Intra-CIIP may be performed by multiple intra-modes. Intra-CIIP may be performed by multiple TMP modes. Intra-CIIP may be performed by MPM lists.
[0123] Template-based intra-mode derivation (TIMD) may be performed. For intra-prediction modes in the MPM (most probable mode) list (e.g., each intra-prediction mode), the SATD (sum of absolute transformed difference) between the template's predicted sample and the reconstructed sample may be calculated. The intra-prediction mode with the smallest SATD (e.g., the first two intra-prediction modes) may be selected as the TIMD mode. The TIMD mode (e.g., the two TIMD modes mentioned above) may be fused with weights, and the weighted intra-prediction described above may be used to encode the current CU. Position-dependent intra-prediction combination (PDPC) may be included in the derivation of the TIMD mode.
[0124] The costs of the two selected modes may be compared to a threshold, and in the test, a cost coefficient of 2 may be applied as costMode2 < 2*costMode1. If this condition is true, fusion is applied; otherwise, only mode1 is used. The weights of the modes may be calculated from the SATD costs as follows: weight1 = costMode2 / (costMode1+ costMode2) weight2 = 1 - weight1 An example of DIMD (decoder side intra-mode derivation) is provided herein. When DIMD is applied, intra-modes (e.g., two intra-modes) may be derived from reconstructed neighboring samples, and their predictors (e.g., two predictors) may be coupled with planar mode predictors having weights derived from gradients (Gx, Gy) calculated on the reconstructed neighboring samples. Division operations in weight derivation may be performed using a lookup table (LUT) (e.g., the same LUT-based integerization scheme used in CCLM mode). In the example, division in orientation calculation Orient = Gy / Gx This is calculated using the following LUT-based scheme. x = Floor( Log2( Gx ) ) normDiff = ( ( Gx<< 4 ) >> x ) & 15 x +=( 3 + ( normDiff != 0 ) ? 1 : 0 ) Orient = (Gy* ( DivSigTable[ normDiff ] | 8 ) + ( 1<<( x-1 ) )) >> x however, DivSigTable
[16] = { 0, 7, 6, 5 ,5, 4, 4, 3, 3, 2, 2, 1, 1, 1, 1, 0}. Since the derived intra-mode may be included in the primary list of the intra-MPM, DIMD processing may occur before the MPM list is constructed. The primary derived intra-mode of a DIMD block is saved with the block and may be used to construct the MPM list of neighboring blocks.
[0125] Examples of combining CIIP with TIMD and template matching merge are provided herein. In CIIP, prediction samples may be generated by weighting inter-prediction signals predicted using CIIP-TM merge candidates with intra-prediction signals predicted using TIMD-derived intra-prediction modes. The combination may (for example, may only) be applied to coded blocks with an area of 1024 or less.
[0126] The derivation of TIMD is sometimes used to derive the intra-prediction mode of CIIP. The intra-prediction mode with the smallest SATD value in the TIMD mode list may be selected and mapped to one of the 67 normal intra-prediction modes.
[0127] Figures 5A and 5B illustrate examples of partitioning based on angular modes. The weights for the two tests (wIntra, wInter) may be changed (for example, if the derived intra predictive mode is an angular mode). Figure 5A shows a vertically partitioned block (e.g., the current block), which may be applied to modes closer to horizontal (2 <= angular mode index < 34). Figure 5B shows a horizontal partition of the block (e.g., the current block), which can be applied to modes closer to vertical modes (34 <= angular mode index <= 66).
[0128] Table 1 below lists the modified weights used for the angular modes for the different subblocks (wIntra, wInter).
[0129] [Table 1]
[0130] In the CIIP-template matching example, a list of merge candidates for CIIP-template matching may be constructed for CIIP-template matching mode. Merge candidates may be refined by template matching. Merge candidates for CIIP templates may be reordered as regular merge candidates by ARMC (adaptive re-ordering as merge candidate). The maximum number of CIIP-template matching merge candidates may be equal to 2.
[0131] Figures 6A–6C show examples of GPM with inter and intra prediction. In examples of GPM with inter and intra prediction, the final prediction samples may be generated by weighting inter and intra prediction samples for regions separated by the GPM (e.g., each region generated by the GPM). Inter prediction samples are derived by the inter GPM, and intra prediction samples are derived by an intra-prediction mode (IPM) candidate list and an index signaled from the encoder. The size of the IPM candidate list may be defined as 3 beforehand. Available IPM candidates are at least one of the following: angular modes parallel to the GPM block boundary (parallel modes), angular modes perpendicular to the GPM block boundary (perpendicular modes), or planar modes as shown in Figures 6A–6C. Intra prediction and GPM with intra prediction, as shown in Figure 6D, may be limited to reduce the signaling overhead of the IPM and to avoid increasing the size of the intra prediction circuit on the hardware decoder. Encoding performance may be improved by directly introducing motion vectors and IPM storage into the GPM blend.
[0132] In IPM derivation based on DIMD and adjacent mode, parallel mode may be registered first. Two IMP candidates (e.g., up to two IMP candidates) may be derived from DIMD, and if there are no identical IPM candidates in the list, a neighboring block may be registered. In the case of neighbor mode derivation, there may be (e.g., at most) five available neighboring block positions. The positions may be limited by the angle of the GPM block boundary, as shown in Table 2 below, which may already be used in GPM by template matching (GPM-TM). As shown in Table 2, the positions of available neighboring blocks for an IPM candidate may be derived based on the angle of the GPM block boundary. A and L represent the upper and left sides of the predicted block.
[0133] [Table 2]
[0134] GPM-Intra can be combined with GPM-MMVD (GPM with merge with different motion vectors). TIMD may be used as an IPM candidate for GPM-Intra, which can improve coding performance. First, the parallel mode is registered, followed by TIMD, DIMD, and neighboring block IPM candidates.
[0135] Figure 7 shows an example of the template matching search area used in intra-TMP. Intra-TMP is an intra-prediction mode in which the L-shaped template may copy the best prediction block from the reconfigured portion of the current picture (e.g., the current frame) where the L-shaped template matches the current template. Within a predefined search range, the encoder may search the reconfigured portion of the current picture for the template most similar to the current template and use the corresponding block as the prediction block. The encoder may notify the use of this mode (e.g., next). The decoder may also perform similar prediction calculations.
[0136] The prediction signal may be generated by matching the L-shaped causal neighborhood of the current block with another block in a predefined search region, as shown in Figure 7, which includes the following: R1: Current CTU R2: CTU in the upper left R3: Upper CTU R4: CTU on the left The sum of absolute difference (SAD) can also be used as the cost function. Within a region (for example, within each region), the decoder may search for the template with the smallest SAD relative to the current template and use its corresponding block as the prediction block. The dimensions of the region ((SearchRange_w, SearchRange_h)) may be set proportionally to the block dimensions (BlkW, BlkH) which have a fixed number of SAD comparisons per pixel, as follows: SearchRange_w = a * BlkW SearchRange_h = a * BlkH Here, "a" is a constant that controls the trade-off between gain and complexity. For example, "a" may be equal to 5.
[0137] The intra-TMP tool may be enabled for CUs with a width and height of 64 or less. The maximum CU size for this intra-TMP is configurable. Intra-TMP mode is signaled at the CU level through a dedicated flag.
[0138] Figures 8-9 show examples of CIIPs with intra-intraprediction. As shown in Figure 8, in intra-intraprediction, the inter portion of the CIIP may be replaced with an intra-TMP mode. The first intra-prediction signal may be associated with the intra-TMP mode. As shown in Figure 8, in intra-intraprediction, the intra portion of the CIIP may use a TIMD mode derivation. In the example, the normal intra-prediction mode is derived by the TIMD mode derivation. The second intra-prediction signal may be associated with (e.g., acquired based on) the normal intra-prediction mode.
[0139] As shown in Figure 9, in intra-intra prediction, the intra portion of CIIP may be replaced with an intra-TMP mode. The first intra-prediction signal may be associated with the intra-TMP mode. As shown in Figure 9, in intra-intra, the inter portion of CIIP may use template matching merge prediction. In the example, a CIIP-template matching merge candidate list may be constructed for the CIIP-template matching mode. The merge candidates may be refined by template matching. The second intra-prediction signal may be associated with (e.g., obtained based on) the template matching prediction.
[0140] Figure 10 shows an example of a CIIP with intra-intraprediction. Examples of blend weights associated with prediction signals are provided herein. In CIIP, the blend mode (e.g., wIntra and wInter) may depend on the intra prediction mode. For non-square modes (e.g., DC or planar) or small blocks (e.g., width less than 4 or height less than 4), the following weights from Table 3 below may be used.
[0141] [Table 3]
[0142] As shown in Table 3 above, if both neighboring blocks (left CU and upper CU) are intra-encoded (isIntra=true), the intra-prediction (wIntra) may be weighted three times more than the inter-prediction (wInter). If both adjacent blocks are inter-encoded (isIntra=false), the inter-prediction may be three times heavier than the intra-prediction. If the predictions of neighboring blocks are different (for example, one is intra-encoded and the other is inter-encoded), both predictions may be given a weight of 2. The weights mentioned above may be ultimately normalized by dividing the final prediction by 4.
[0143] In the case of intraCIIP, the weights are calculated based on whether or not intraTMP is used in neighboring blocks (or multiple blocks) (for example, based on whether or not interpretation is used). The following weights in Table 4 may be used.
[0144] [Table 4]
[0145] As shown in Table 4 above, if both neighboring blocks (left CU and upper CU) are intra-TMP encoded (isIntraTMP=true), the intra-prediction (wIntraTMP) may be weighted more than three times more than the normal intra-prediction (wIntra). If neither adjacent block is intraIMP encoded (isIntraTMP=false) (for example, using the normal intra-prediction), the normal intra-prediction may be weighted three times more than the intra-TMP prediction. If the neighboring block predictions are different (for example, one is intra and the other is intra-TMP), a weight of 2 may be assigned to both predictions. The weights mentioned above may be finalized by dividing the final prediction by 4.
[0146] In the example, a video decoder or video encoder may acquire a first intra-prediction signal and a second intra-prediction signal for a block. The first intra-prediction signal may be associated with a first intra-prediction mode, which may be an intra-template prediction mode (intra-TMP) mode. The second prediction signal may be associated with a second intra-prediction mode, which may be a normal intra-prediction mode. The decoder or encoder may generate predicted samples associated with the block by determining a first weight associated with the first intra-prediction signal and a second weight associated with the second intra-prediction signal. The first and second weights may be determined based on a first neighbor block prediction mode and a second neighbor block prediction mode. The encoder may derive residual samples associated with the block based on the predicted samples. The block is decoded or encoded based on the residual samples.
[0147] As described above (for example, in the example shown in Table 4), if one of the first neighbor block prediction mode or the second neighbor block prediction mode is an intra-template prediction mode (intra-TMP) mode, and one of the first neighbor block prediction mode or the second neighbor block prediction mode is a normal intra-prediction mode, the decoder or encoder may determine that the first weight and the second weight are weighted equally (for example, each having a weight equal to 2). If both the first neighbor block prediction mode and the second neighbor block prediction mode are intra-template prediction (intra-TMP) modes, the decoder or encoder may determine that the first weight is 3 and the second weight is 1. If both the first neighbor block prediction mode and the second neighbor block prediction mode are normal intra-prediction modes, the decoder or encoder may determine that the first weight is 1 and the second weight is 3.
[0148] For example, a blending process equivalent to IBC-CIIP may be used. The IBC-CIIP blending process is defined as follows: Pred(x,y) = (a*PredReg + b*PredIbc + offset) >> shift Here, >> represents a downshift, PredReg corresponds to predictions obtained by a normal intra process, and PredIbc corresponds to predictions obtained by an IBC process. The variables a, b, shift, and offset may be calculated as follows: b=13, a=3, shift=4 if merge is used; otherwise, a=b=1, shift=1. The offset is calculated as 1 << (shift - 1) (e.g., both with and without merge).
[0149] PredIbc can be replaced with intra-TMP prediction. The following weights may be used: b=13, a=3, shift=4 if any of the neighbors are encoded with a block vector (intra-TMP, IBC, or newer CIIP mode); otherwise, a=b=1, shift=1.
[0150] Examples of transform coding associated with intraCIIP are provided herein. Transform coding may have at least one or three components: MTS (multiple transform selection), LFNST (low frequency non-separable transform), or NSPT (non-separable primary transform). MTS (e.g., primary transform) may include triangular transforms (e.g., DCT and DST as an alternative to the default DCT2 transform). LFNST (e.g., quadratic transform) may be applied non-separably to the low-frequency portion after the primary transform (on the encoder side). NSPT is a transform that can be applied directly to the residual in an inseparable manner. Due to the large number of multiplications, NSPT may be limited to small blocks.
[0151] The above transformations may depend on the intra-mode used. This is likely because training is performed offline based on datasets grouped by intra-prediction mode. To adapt intra-CIIP with the modes and transformations described above, at least one of the following may be used: treat intra-CIIP as a planar mode, use the corresponding intra-mode of intra-CIIP, or derive an equivalent mode. To treat intra-CIIP as a planar mode, MTS, LFNST, and NSPT may use transformation kernels specific to planar modes. Regarding using the corresponding intra-mode of intra-CIIP, the intra-prediction mode may be used in the selection of the transformation kernel, since CIIP may already use intra-prediction (e.g., TIMD mode). For the derivation of an equivalent mode, an equivalent mode may be derived for an intra-CIIP mode. That is, a prediction signal may be taken from intra-CIIP, and a prediction mode similar to that prediction signal may be derived. The DIMD process is used to obtain a prediction mode from a prediction signal.
[0152] Examples of intra-CIIPs interacting with IBCs are described herein. For example, when intra-TMPs are used, their block vectors may be used as merge candidates for IBCs. With respect to IBC merges, there are five spatial candidates: upper PUs, left PUs, upper right PUs, lower left PUs, and upper left PUs. To generate the merge candidate list, if a PU is IBC encoded or intra-TMP encoded, the PU may be treated equally (for example, since it may have a block vector that can be used in both modes). In the example, if any of the five PUs are encoded in an intra-CIIP, its corresponding block vector may be used to construct the merge candidate list.
[0153] In chroma direct block vector (chromaDBV) mode, intra-TMP may be used. In this mode, IBC and intra-TMP block vectors (for example, only usable for rumors) may be used for chroma collocation blocks. Syntax elements may be signaled to indicate the use of this mode. In the case of intra-CIIP, if the collated rumor is encoded with intra-CIIP, chromaDBVD mode is used, and the corresponding block vector is used for the chroma.
[0154] An example of an intra-CIIP interacting with an MPM list is provided herein. The MPM construction process may consider the following five adjacent PUs: the PU above, the PU to the left, the PU above and to the right, the PU below and to the left. The MPM may construct an intra-mode from these PUs if they are coded in a normal intra-mode (e.g., planar, DC, or angular). If an intra-CIIP is used, at least one of the following options may be considered: using the intra-part of the intra-CIIP, or using an equivalent mode. If an equivalent mode is used, the equivalent mode may be derived from the predicted signal using DIMD. This may lead to an improvement in MPM list construction. In the example, both the intra-mode and equivalent mode of the intra-CIIP may be used in the intra-CIIP. This is possible by placing one of the modes at the end of the MPM list and the other at the beginning of the secondary MPM list.
[0155] An example of the interaction between intra-CIIP and code prediction is provided herein. The coefficient code prediction mode can be disabled in intra-CIIP (e.g., disabled when CIIP is being used). This is likely because this mode requires multiple template analyses / processes. Templates may be used to compute (e.g., required for computation) template analyses for intra-TMP prediction, TIMP mode derivation, potentially equivalent modes for transform kernel selection, and code prediction. Since this combination may not exhibit a practical gain-complexity trade-off, code prediction may be disabled when CIIP is being used.
[0156] Examples of intra-CIIPs that interact with multiple intra-modes are provided herein. For example, intra-modes may be derived from a TIMD process. Multiple intra-modes may be allowed. For example, N possible intra-mode candidates may be built. Intra-modes may be signaled to a decoder, and predictions may be made according to the signaled modes. A certain number of intra-modes (e.g., only a certain number) may be used because allowing all modes may require high overhead. For example, MPM modes may be used in intra-CIIPs. For example, modes may be sorted according to template cost.
[0157] When using MPM mode (e.g., all MPM modes), the same signal mechanism as the default mode may be used. For example, MPM-based signaling may be derived from neighboring PUs. The encoder may select the best mode.
[0158] When sorting modes according to template cost, some or all modes may be tested on a reconstructed template and sorted according to template cost (e.g., similar to the TIMD process). Blending processes or CIIP may be considered. When testing intra-mode candidates, the blending capability with intra-TMP modes may be tested on a template where weighting between intra-modes and intra-TMPs is used. The tested predictions result in matching costs on the template (e.g., SSD, SAD, SATD), and modes are sorted accordingly. For example, all modes may be used for selection. For example, sorting may be limited to MPM modes. When N is equal to 1, this corresponds to the TIMD process.
[0159] An example of an intra-CIIP interacting with multiple intra-TMP candidates is provided herein. For example, the intra-TMP search process may yield multiple candidates along with their associated template distances. The encoder may select the best candidate from among the N candidates and signal that best candidate to the decoder. In the example, block vectors from neighboring blocks, which may be encoded in IBC or intra-TMP, may be used as candidates for the current intra-CIIP mode (e.g., additional candidates). That is, in addition to the current block vector obtained by the intra-TMP process, other block vectors obtained from neighboring blocks may be used. N intra-TMP candidates may be tested along with M intra-modes, resulting in L sorted modes (e.g., each mode may contain both an intra-TMP candidate and an intra-mode candidate). The encoder may select the best mode and signal it to the decoder.
[0160] Examples of how intraCIIP interacts with horizontal and vertical template selections are provided herein. For example, an intraTMP may use only the top or only the left-side template. That is, the encoder may choose to find the best intraTMP candidate from matching with left, top, or top-left candidates. IntraCIIP allows one or more of the following: When the top template is used, horizontal TIMD may be used for the intra portion. When the left template is used, vertical TIMD may be used for the intra portion. In a special mode called S GPM (spatial GPM), horizontal and vertical TIMD may be performed. In this mode (in addition to the default TIMD process, for example), two other TIMD candidates may be derived. In the example, the horizontal TIMD may be derived from the top template, and the vertical TIMD may be derived from the left-side template.
[0161] A device (for example, a video encoding device) may perform (or be configured to perform) one or more actions. For example, the device may obtain a first intra-prediction signal associated with a first prediction mode for a video block, obtain a second intra-prediction signal associated with a block vector-based prediction mode for a video block, generate prediction samples associated with the video block based on the first and second intra-prediction signals, and encode the video block based on the prediction samples.
[0162] For example, block vector-based prediction modes may be intra-template matching prediction (intra-TMP) or intra-block copy (IBC).
[0163] The device may determine a first weight associated with a first intra-prediction signal and a second weight associated with a second intra-prediction signal. Prediction samples associated with a video block may be generated based on the first weight applied to the first intra-prediction signal and the second weight applied to the second intra-prediction signal. For example, prediction samples are generated based on the merging of the first and second intra-prediction signals. The first and / or second weights may be determined based on whether the intra-TMP is associated with one or more of the video block and its neighbors, namely a first block and a second block. For example, if both neighboring blocks are intra-encoded (e.g., the intra-TMP is not associated with any neighboring blocks), the first weight might be 3 and the second weight might be 1.
[0164] The first intra-prediction signal may be associated with the most probable mode (MPM) list of the video block. The device retrieves the index associated with the MPM list, sorts the multiple MPM modes in the MPM list based on the respective costs associated with each MPM mode, determines the first prediction mode based on the sorted multiple MPM modes and index, and the first intra-prediction signal is obtained based on the first prediction mode.
[0165] The device may acquire multiple block vectors and indices, select a block vector from the multiple block vectors based on the index, and acquire a second predictive signal based on the selected block vector.
[0166] The device may be configured to select a transformation for a video block based on a first prediction mode and / or a second intra-prediction mode. The transformation selected for a video block may be one of multiple transform selections (MTS), low frequency non-separable transforms (LFNST), or non-separable primary transforms (NSPT).
[0167] In the example, a video decoder or video encoder may obtain a first intra-prediction signal and a second intra-prediction signal for a video block. The first prediction signal may be associated with a first intra-prediction mode (e.g., a non-template-based prediction mode). The second prediction signal may be associated with a second, block-vector-based intra-prediction mode (e.g., intra-TMP or IBC). The decoder or encoder may generate a combined prediction for the video block by applying a first weight to the first intra-prediction signal and a second weight to the second intra-prediction signal. The first and second weights may be determined based on a first neighbor block prediction mode and a second neighbor block prediction mode. Based on the combined prediction, the encoder may derive a residual block associated with the video block. The block is encoded based on the derived residual block. The decoder may obtain a residual block from the video data and may reconstruct the block based on the residual block and the generated combined prediction.
[0168] While features and elements are described above in specific combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with other features and elements. In addition, the methods described herein may be implemented in computer programs, software, or firmware embedded on computer-readable media for execution by a computer or processor. Examples of computer-readable media include electrical signals (transmitted via wired or wireless connections) and computer-readable recording media. Examples of computer-readable recording media include, but are not limited to, ROM (described), RAM (random access memory), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media such as CD-ROM discs and DVDs (digital versatile disks). Processors associated with software may be used to implement radio frequency transceivers for use in UEs, WTRUs, terminals, base stations, RNCs, or any host computer.
Claims
1. For a video block, acquire a first intra-prediction signal associated with a first prediction mode. For the aforementioned video block, a second intra-prediction signal associated with a block vector-based prediction mode is acquired. Based on the first intra-prediction signal and the second intra-prediction signal, predict samples associated with the video block are generated. The video block is decoded based on the predicted sample. Processor configured in such a way A video decoding device characterized by having the following features.
2. The video decoding device according to claim 1, characterized in that the block vector-based prediction mode is intra-template matching prediction (intra-TMP).
3. The video decoding device according to claim 1, characterized in that the block vector-based prediction mode is an intra-block copy (IBC).
4. The aforementioned processor, Determine a first weight associated with the first intra-prediction signal and a second weight associated with the second intra-prediction signal, and the prediction samples associated with the video block are generated based on the first weight applied to the first intra-prediction signal and the second weight applied to the second intra-prediction signal. The video decoding device according to any one of claims 1 to 3, further characterized in that it is configured as follows.
5. The video decoding device according to claim 4, characterized in that the first weight and the second weight are determined based on whether the intra TMP is associated with one or more of the video block and neighboring first blocks and neighboring second blocks.
6. The first intra-prediction signal is associated with the most probable mode (MPM) list of the video block, and the processor, Based on the respective costs associated with each MPM mode, the multiple MPM modes in the MPM list are sorted. The first prediction mode is determined based on the sorted plurality of MPM modes, and the first intra-prediction signal is acquired based on the first prediction mode. The video decoding device according to any one of claims 1 to 3, further characterized in that it is configured as follows.
7. The aforementioned processor, The index associated with the MPM list is obtained, and the first prediction mode is determined based on the index and the sorted MPM list. The video decoding device according to claim 6, further configured as follows.
8. The aforementioned processor, Obtain multiple block vectors and indices, Based on the aforementioned index, a block vector is selected from the plurality of block vectors, and the second intra-prediction signal is obtained based on the selected block vector. The video decoding device according to any one of claims 1 to 7, further characterized in that it is configured as follows.
9. The aforementioned plurality of block vectors are, Multiple block vectors obtained by performing intra-template matching prediction (intra-TMP) on the aforementioned video block, or Multiple block vectors obtained from the aforementioned video block and at least one neighboring block. The video decoding device according to claim 8, characterized by including at least one of the following.
10. For a video block, a first intra-prediction signal associated with a first prediction mode is acquired, For the aforementioned video block, a second intra-prediction signal associated with a block vector-based prediction mode is acquired, Based on the first intra-prediction signal and the second intra-prediction signal, predict samples associated with the video block are generated, Decoding the video block based on the predicted sample and A video decoding method characterized by comprising:
11. The video decoding method according to claim 10, characterized in that the block vector-based prediction mode is an intra-TMP.
12. The video decoding method according to claim 10, characterized in that the block vector-based prediction mode is IBC.
13. The method involves determining a first weight associated with the first intra-prediction signal and a second weight associated with the second intra-prediction signal, wherein the prediction sample associated with the video block is generated based on the first weight applied to the first intra-prediction signal and the second weight applied to the second intra-prediction signal. The video decoding method according to any one of claims 10 to 12, further comprising:
14. The video decoding method according to claim 13, characterized in that the first weight and the second weight are determined based on whether the intra TMP is associated with one or more of the video block and neighboring first blocks and neighboring second blocks.
15. A computer program product comprising program code instructions for implementing any one of the methods described in claims 10 to 14, which are stored in a non-temporary computer-readable medium and executed by a processor.