Combining intra-template prediction and intra-block copying with other coding tools
By integrating intra-template prediction and intra-block copy modes with other coding tools, the video coding systems achieve improved compression efficiency and reduced bandwidth requirements through enhanced prediction signal generation and decoding processes.
Patent Information
- Application Number
- JP2025530683
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-09-20
- Filing Date
- 2023-12-18
- Publication Date
- 2026-01-14
AI Technical Summary
Existing video coding systems face challenges in efficiently compressing digital video signals, particularly in achieving optimal prediction accuracy and reducing bandwidth requirements.
Incorporating intra-template prediction (intra-TMP) and intra-block copy (IBC) modes with other coding tools, utilizing block-based and inter prediction modes, and employing decoder-side intra-mode derivation (DIMD) and template-based intra-mode derivation (TIMD) to enhance prediction signal generation and decoding/encoding processes.
Improves video compression efficiency by optimizing prediction signals, reducing storage and transmission bandwidth, and enhancing decoding accuracy.
Smart Images

Figure 2026501085000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of European Patent Application Publication No. 22307040.0, filed December 23, 2022, European Patent Application Publication No. 23306112.6, filed June 30, 2023, and European Patent Application Publication No. 23306559.8, filed September 20, 2023, the disclosures of which are incorporated herein by reference in their entireties. [Background technology]
[0002] background Video coding systems may be used to compress digital video signals, for example, to reduce the storage and / or transmission bandwidth required for such signals. Video coding systems may include, for example, block-based, wavelet-based, and / or object-based systems. Summary of the Invention
[0003] overview SUMMARY
[0003] Disclosed herein are systems, methods, and apparatus in the field of video compression.
[0004] In an example, a video decoder or video encoder may obtain a first prediction signal for a current block using a block vector-based intra-prediction mode (e.g., an intra-template prediction (intra-TMP) mode or an intra-block copy (IBC) mode). A second prediction mode may be used to obtain a second prediction signal for the current block. A prediction block may be generated by weighting the first prediction signal using the block-based intra-prediction mode and the second prediction signal using the second prediction mode. The current block may be decoded or encoded based on the prediction block.
[0005]
[0005] In an example, the second prediction mode may be a normal intra prediction mode.
[0005] In an example, the second prediction mode may be an inter prediction mode.
[0006] In an example, a video decoder or encoder may determine that a current block is coded using a decoder-side intra-mode derivation (DIMD) mode. A second set of prediction modes may be derived using DIMD based on a histogram of gradients of template samples associated with the current block. In an example, a video decoder or encoder may determine that a current block is coded using a template-based intra-mode derivation (TIMD) mode. A second set of prediction modes may be derived using TIMD based on testing a most likely mode on a template associated with the current block. A prediction block may be generated by weighting a first prediction signal using a block-based intra-prediction mode and a second set of prediction signals.
[0007]
[0007] These examples may be performed by a device having a processor. The device may be an encoder or a decoder. These examples may be performed by a computer program product stored on a non-transitory computer-readable medium and including program code instructions. These examples may be performed by a computer program including program code instructions.
[0008] The systems, methods, and devices described herein may include a decoder. In some examples, the systems, methods, and devices described herein may involve an encoder. In some examples, the systems, methods, and devices described herein may involve a signal (e.g., from an encoder and / or received by a decoder). A computer-readable medium may include instructions for causing one or more processors to perform the methods described herein. A computer program product may include instructions that, when executed by one or more processors, cause the one or more processors to perform the methods described herein. [Brief explanation of the drawings]
[0009] BRIEF DESCRIPTION OF THE DRAWINGS [Figure 1A] FIG. 1 is a system diagram illustrating an example communication system in which one or more disclosed embodiments may be implemented. [Figure 1B]
[0010] 1B is a system diagram illustrating an example wireless transmit / receive unit (WTRU) that may be used within the communication system illustrated in FIG. 1A, according to one embodiment. [Figure 1C]
[0011] 1B is a system diagram illustrating an example radio access network (RAN) and an example core network (CN) that may be used within the communication system illustrated in FIG. 1A, according to one embodiment. [Figure 1D]
[0012] 1B is a system diagram illustrating a further exemplary RAN and a further exemplary CN that may be used within the communication system illustrated in FIG. 1A, according to one embodiment. [Figure 2]
[0013] 1 illustrates an exemplary video encoder. [Figure 3]
[0014] 1 illustrates an exemplary video decoder. [Figure 4]
[0015] 1 illustrates an example of a system in which various aspects and examples may be implemented. [Figure 5A]
[0016] 1 illustrates an example of division for angular modes. [Figure 5B]
[0016] An example of division for angular modes is illustrated. [Figure 6A]
[0017] 1 illustrates an example of a geometric partitioning mode (GPM) with inter-prediction and intra-prediction. [Figure 6B]
[0017] An example of a geometric partitioning mode (GPM) with inter-prediction and intra-prediction is illustrated. [Figure 6C]
[0017] An example of a geometric partitioning mode (GPM) with inter-prediction and intra-prediction is illustrated. [Figure 6D]
[0017] An example of a geometric partitioning mode (GPM) with inter-prediction and intra-prediction is illustrated. [Figure 7]
[0018] 1 illustrates an example of an intra-template matching prediction (intra-TMP) search area. [Figure 8]
[0019] 1 illustrates an example of CIIP with intra-intra prediction. [Figure 9]
[0020] 1 illustrates an example of CIIP with inter-intra prediction. [Figure 10]
[0021] Illustrates an example of using a DIMD mode in combination with a block vector-based intra prediction mode (eg, intra-TMP mode and / or IBC mode). [Figure 11]
[0022] 1 illustrates an example of TIMD combined with a block vector-based intra prediction mode (eg, intra-TMP mode and / or IBC mode). DETAILED DESCRIPTION OF THE INVENTION
[0010] Detailed Description
[0023] A more detailed understanding may be had from the following description, given by way of example in conjunction with the accompanying drawings, in which:
[0011]
[0024] 1A is a diagram illustrating an example communication system 100 in which one or more disclosed embodiments may be implemented. The communication system 100 may be a multiple-access system that provides content, such as voice, data, video, messaging, broadcasts, and the like, to multiple wireless users. The communication system 100 may enable multiple wireless users to access such content through the sharing of system resources, including wireless bandwidth. For example, the communication system 100 may employ one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single-carrier FDMA (SC-FDMA), zero-tailed unique word DFT spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block filtered OFDM, filter bank multicarrier (FBMC), etc.
[0012]
[0025] 1A, communications system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RANs 104 / 113, CNs 106 / 115, public switched telephone network (PSTN) 108, the Internet 110, and other networks 112, although it should be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of WTRUs 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a wireless environment. By way of example, the WTRUs 102a, 102b, 102c, 102d (any of which may be referred to as a “station” and / or “STA”) may be configured to transmit and / or receive wireless signals and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription base units, pagers, mobile phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearables, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated processing chain contexts), consumer electronics devices, devices operating in commercial and / or industrial wireless networks, etc. Any of the WTRUs 102a, 102b, 102c, and 102d may be referred to interchangeably as UEs.
[0013]
[0026] The communications system 100 may also include a base station 114a and / or a base station 114b. Each of the base stations 114a, 114b may be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communications networks, such as the CN 106 / 115, the Internet 110, and / or other networks 112. By way of example, the base stations 114a, 114b may be a base transceiver station (BTS), a Node B, an eNode B, a Home Node B, a Home eNode B, a gNB, an NR Node B, a site controller, an access point (AP), a wireless router, etc. Although the base stations 114a, 114b are each depicted as a single element, it should be understood that the base stations 114a, 114b may include any number of interconnected base stations and / or network elements.
[0014]
[0027] The base station 114a may be part of the RAN 104 / 113, which may also include other base stations and / or network elements (not shown) (e.g., a base station controller (BSC), a radio network controller (RNC), relay nodes, etc.). The base station 114a and / or base station 114b may be configured to transmit and / or receive wireless signals on one or more carrier frequencies, which may be referred to as a cell (not shown). These frequencies may reside in the licensed spectrum, the unlicensed spectrum, or a combination of the licensed and unlicensed spectrum. A cell may provide wireless service coverage for a particular geographic area, which may be relatively fixed or which may change over time. A cell may be further divided into cell sectors. For example, the cell associated with the base station 114a may be divided into three sectors. Thus, in one embodiment, the base station 114a may include three transceivers, i.e., one transceiver for each sector of the cell. In one embodiment, the base station 114a may employ multiple-input multiple-output (MIMO) technology and utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in desired spatial directions.
[0015]
[0028] The base stations 114a, 114b may communicate with one or more of the WTRUs 102a, 102b, 102c, 102d over an air interface 116, which may be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 116 may be established using any suitable radio access technology (RAT).
[0016]
[0029] More specifically, as noted above, the communication system 100 may be a multiple-access system and may employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, the base stations 114a and WTRUs 102a, 102b, 102c in the RAN 104 / 113 may implement a radio technology such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which may establish the air interface 115 / 116 / 117 using Wideband CDMA (WCDMA). WCDMA may include communication protocols such as High Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High Speed Downlink (DL) Packet Access (HSDPA) and / or High Speed UL Packet Access (HSUPA).
[0017]
[0030] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which may establish the air interface 116 using Long Term Evolution (LTE), and / or LTE-Advanced (LTE-A), and / or LTE-Advanced Pro (LTE-A Pro).
[0018]
[0031] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as New Radio (NR) radio access, which may establish the air interface 116 using NR.
[0019]
[0032] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement multiple radio access technologies. For example, the base station 114a and the WTRUs 102a, 102b, 102c may jointly implement LTE radio access and NR radio access, e.g., using a dual connectivity (DC) principle. Thus, the radio interface utilized by the WTRUs 102a, 102b, 102c may be characterized by multiple types of radio access technologies and / or transmissions sent to and from multiple types of base stations (e.g., eNBs and gNBs).
[0020]
[0033] In other embodiments, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as IEEE 802.11 (i.e., Wireless Fidelity, WiFi), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access, WiMAX), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile communications (GSM), Enhanced Data rates for GSM Evolution (EDGE), GSM EDGE (GERAN), or the like.
[0021]
[0034] 1A may be, for example, a wireless router, a Home Node B, a Home eNode B, or an access point, and may utilize any suitable RAT to facilitate wireless connectivity in a local area, such as a business, a home, a vehicle, a campus, an industrial facility, an air corridor (e.g., for use by drones), a road, etc. In one embodiment, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In an embodiment, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 114b and the WTRUs 102c, 102d may utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a picocell or femtocell. 1A, the base station 114b may have a direct connection to the Internet 110. Thus, the base station 114b may not need to access the Internet 110 via the CN 106 / 115.
[0022]
[0035] The RAN 104 / 113 may communicate with the CN 106 / 115, which may be any type of network configured to provide voice, data, application, and / or voice over internet protocol (VoIP) services to one or more of the WTRUs 102a, 102b, 102c, 102d. The data may have various quality of service (QoS) requirements, such as different throughput, latency, error tolerance, reliability, data throughput, and mobility requirements. The CN 106 / 115 may provide call control, billing services, mobile location-based services, prepaid calling, Internet connectivity, video distribution, and / or perform high-level security functions such as user authentication. Although not shown in FIG. 1A , it should be understood that the RAN 104 / 113 and / or the CN 106 / 115 may communicate directly or indirectly with other RANs employing the same RAT as the RAN 104 / 113 or a different RAT. For example, in addition to being connected to the RAN 104 / 113, which may be utilizing NR radio technology, the CN 106 / 115 may also communicate with another RAN (not shown) that employs GSM, UMTS, CDMA 2000, WiMAX, E-UTRA, or WiFi radio technology.
[0023]
[0036] The CN 106 / 115 may also serve as a gateway for the WTRUs 102a, 102b, 102c, 102d to access the PSTN 108, the Internet 110, and / or other networks 112. The PSTN 108 may include a circuit-switched telephone network providing plain old telephone service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices that use common communication protocols such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) in the TCP / IP Internet protocol suite. The network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, the network 112 may include another CN connected to one or more RANs, which may employ the same RAT as the RAN 104 / 113 or a different RAT.
[0024]
[0037] Some or all of the WTRUs 102a, 102b, 102c, 102d in the communications system 100 may include multi-mode capabilities (e.g., the WTRUs 102a, 102b, 102c, 102d may include multiple transceivers for communicating with different wireless networks over different wireless links). For example, the WTRU 102c shown in FIG. 1A may be configured to communicate with a base station 114a, which may employ a cellular-based wireless technology, and a base station 114b, which may employ an IEEE 802.2 wireless technology.
[0025]
[0038] 1B is a system diagram illustrating an example WTRU 102. As shown in FIG. 1B, the WTRU 102 may include, among other things, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power source 134, a global positioning system (GPS) chipset 136, and / or other peripherals 138. It should be understood that the WTRU 102 may include any sub-combination of the above elements while remaining consistent with an embodiment.
[0026]
[0039] The processor 118 may be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. As alluded to above, the processor 118 may include multiple processors. The processor 118 may perform signal coding, data processing, power control, input / output processing, and / or any other function that enables the WTRU 102 to operate in a wireless environment. The processor 118 may be coupled to the transceiver 120, which may be coupled to the transmit / receive element 122. While FIG. 1B depicts the processor 118 and the transceiver 120 as separate components, it should be understood that the processor 118 and the transceiver 120 may be integrated together in an electronic package or chip.
[0027]
[0040] The transmit / receive element 122 may be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) via the air interface 116. For example, in one embodiment, the transmit / receive element 122 may be an antenna configured to transmit and / or receive RF signals. In one embodiment, the transmit / receive element 122 may be an emitter / detector configured to transmit and / or receive IR, UV, or visible light signals, for example. In yet another embodiment, the transmit / receive element 122 may be configured to transmit and / or receive both RF and light signals. It should be understood that the transmit / receive element 122 may be configured to transmit and / or receive any combination of wireless signals.
[0028]
[0041] 1B, the transmit / receive element 122 is depicted as a single element, the WTRU 102 may include any number of transmit / receive elements 122. More specifically, the WTRU 102 may employ MIMO technology. Thus, in one embodiment, the WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 116.
[0029]
[0042] The transceiver 120 may be configured to modulate signals transmitted by the transmit / receive element 122 and demodulate signals received by the transmit / receive element 122. As mentioned above, the WTRU 102 may have multi-mode capabilities. Thus, the transceiver 120 may include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs, such as NR and IEEE 802.11.
[0030]
[0043] The processor 118 of the WTRU 102 may be coupled to and may receive user input data from a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light emitting diode (OLED) display unit). The processor 118 may also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. Additionally, the processor 118 may access information from and store data in any type of suitable memory, such as non-removable memory 130 and / or removable memory 132. The non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 may include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, etc. In other embodiments, the processor 118 may access information from and store data in memory that is not physically located on the WTRU 102, such as on a server or home computer (not shown).
[0031]
[0044] The processor 118 may receive power from the power source 134 and may be configured to distribute and / or control power to other components within the WTRU 102. The power source 134 may be any suitable device for powering the WTRU 102. For example, the power source 134 may include one or more dry batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel-metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.
[0032]
[0045] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to, or instead of, information from the GPS chipset 136, the WTRU 102 may receive location information from a base station (e.g., base stations 114a, 114b) via the air interface 116 and / or may determine its location based on the timing of signals received from two or more nearby base stations. It will be appreciated that the WTRU 102 may acquire location information by way of any suitable location-determination method while remaining consistent with an embodiment.
[0033]
[0046] The processor 118 may further be coupled to other peripherals 138, which may include one or more software and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, the peripherals 138 may include an accelerometer, an e-compass, a satellite transceiver, a digital camera (for photos and / or videos), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth module, a frequency modulation (FM) radio unit, a digital music player, a media player, a video game player module, an internet browser, a virtual reality and / or augmented reality (VR / AR) device, an activity tracker, etc. The peripherals 138 may include one or more sensors, which may be one or more of a gyroscope, an accelerometer, a Hall effect sensor, a magnetometer, a direction sensor, a proximity sensor, a temperature sensor, a time sensor, a geolocation sensor, an altimeter, a light sensor, a touch sensor, a magnetometer, a barometer, a gesture sensor, a biometric sensor, and / or a humidity sensor.
[0034]
[0047] The WTRU 102 may include a full-duplex radio that may transmit and receive some or all of the signals associated with a particular subframe (e.g., for both UL (e.g., for transmission) and downlink (e.g., for reception) in parallel and / or simultaneously). The full-duplex radio may include an interference management unit to reduce and or substantially eliminate self-interference through either hardware (e.g., a choke) or processor-mediated signal processing (e.g., via a separate processor (not shown) or processor 118). In one embodiment, the WTRU 102 may include a half-duplex radio that transmits and receives some or all of the signals (e.g., associated with a particular subframe for either the UL (e.g., for transmission) or downlink (e.g., for reception)).
[0035]
[0048] 1C is a system diagram illustrating the RAN 104 and the CN 106, according to one embodiment. As mentioned above, the RAN 104 may employ E-UTRA radio technology to communicate with the WTRUs 102a, 102b, 102c over the air interface 116. The RAN 104 may also communicate with the CN 106.
[0036]
[0049] While the RAN 104 may include eNode-Bs 160a, 160b, and 160c, it should be understood that the RAN 104 may include any number of eNode-Bs while remaining consistent with an embodiment. The eNode-Bs 160a, 160b, and 160c may each include one or more transceivers for communicating with the WTRUs 102a, 102b, and 102c over the air interface 116. In one embodiment, the eNode-Bs 160a, 160b, and 160c may implement MIMO technology. Thus, for example, the eNode-B 160a may use multiple antennas to transmit wireless signals to and / or receive wireless signals from the WTRU 102a.
[0037]
[0050] Each of the eNode-Bs 160a, 160b, 160c may be associated with a particular cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, scheduling of users in the UL and / or DL, etc. As shown in FIG. 1C, the eNode-Bs 160a, 160b, 160c may communicate with one another over an X2 interface.
[0038]
[0051] 1C may include a mobility management entity (MME) 162, a serving gateway (SGW) 164, and a packet data network (PDN) gateway (or PGW) 166. Although each of the above elements is depicted as part of the CN 106, it should be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.
[0039]
[0052] The MME 162 may be connected to each of the eNode-Bs 162a, 162b, 162c in the RAN 104 via an S1 interface and may function as a control node. For example, the MME 162 may be responsible for authenticating users of the WTRUs 102a, 102b, 102c, bearer activation / deactivation, selecting a particular serving gateway during initial attach of the WTRUs 102a, 102b, 102c, etc. The MME 162 may provide a control plane function for switching between the RAN 104 and other RANs (not shown) that employ other radio technologies such as GSM and / or WCDMA.
[0040]
[0053] The SGW 164 may be connected to each of the eNode Bs 160a, 160b, 160c in the RAN 104 via an S1 interface. The SGW 164 may generally route and forward user data packets to and from the WTRUs 102a, 102b, 102c. The SGW 164 may perform other functions such as anchoring the user plane during inter-eNode B handovers, triggering paging when DL data is available for the WTRUs 102a, 102b, 102c, and managing and storing the context of the WTRUs 102a, 102b, 102c.
[0041]
[0054] The SGW 164 may be connected to a PGW 166, which may provide the WTRUs 102a, 102b, 102c with access to packet-switched networks, such as the Internet 110, to facilitate communication between the WTRUs 102a, 102b, 102c and IP-enabled devices.
[0042]
[0055] The CN 106 may facilitate communications with other networks. For example, the CN 106 may provide the WTRUs 102a, 102b, 102c with access to circuit-switched networks, such as the PSTN 108, to facilitate communications between the WTRUs 102a, 102b, 102c and traditional land-based communications devices. For example, the CN 106 may include or communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between the CN 106 and the PSTN 108. In addition, the CN 106 may provide the WTRUs 102a, 102b, 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.
[0043]
[0056] Although the WTRUs are described in Figures 1A-1D as wireless terminals, it is contemplated that in certain representative embodiments such terminals may use a wired communication interface (e.g., temporary or permanent) with a communication network.
[0044]
[0057] In a representative embodiment, the other network 112 may be a WLAN.
[0045]
[0058] A WLAN in Infrastructure Basic Service Set (BSS) mode may have an access point (AP) of the BSS and one or more stations (STAs) associated with the AP. The AP may have access to or interface with a distribution system (DS) or another type of wired / wireless network that carries traffic to and / or from the BSS. Traffic originating from outside the BSS to a STA may arrive via the AP and be delivered to the STA. Traffic originating from a STA to a destination outside the BSS may be sent to the AP and delivered to the respective destination. Traffic between STAs within the BSS may be transmitted via the AP; for example, a source STA may transmit traffic to the AP, and the AP may deliver the traffic to the destination STA. Traffic between STAs within the BSS may be considered and / or referred to as peer-to-peer traffic. Peer-to-peer traffic may be transmitted (e.g., directly) between a source STA and a destination STA using direct link setup (DLS). In certain representative embodiments, the DLS may use 802.11e DLS or 802.11z Tunneled DLS (TDLS). A WLAN using an Independent BSS (IBSS) mode may not have an AP, and STAs within or using the IBSS (e.g., all STAs) may communicate directly with each other. The IBSS communication mode is sometimes referred to herein as an "ad hoc" communication mode.
[0046]
[0059] When using the 802.11ac infrastructure mode of operation or a similar mode of operation, an AP may transmit beacons on a fixed channel, such as a primary channel. The primary channel may be a fixed width (e.g., a wide bandwidth of 20 MHz) or may be dynamically set via signaling. The primary channel may be the operating channel of the BSS and may be used by STAs to establish a connection with the AP. In certain representative embodiments, for example, in an 802.11 system, carrier sense multiple access with collision avoidance (CSMA / CA) may be implemented. With CSMA / CA, STAs (e.g., all STAs), including the AP, may sense the primary channel. If the primary channel is sensed / detected by a particular STA and / or determined to be busy, the particular STA may back off. In a given BSS, one STA (e.g., only one station) may transmit at any given time.
[0047]
[0060] High-throughput (HT) STAs may use 40 MHz wide channels for communication, for example, by combining a primary 20 MHz channel with adjacent or non-adjacent 20 MHz channels to form the 40 MHz wide channel.
[0048]
[0061] A very high throughput (VHT) STA may support 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz wide channels. A 40 MHz and / or 80 MHz channel may be formed by combining contiguous 20 MHz channels. A 160 MHz channel may be formed by combining eight contiguous 20 MHz channels or two non-contiguous 80 MHz channels, which may be referred to as an 80+80 configuration. In the 80+80 configuration, the channel-encoded data may be passed through a segment parser that may split the data into two streams. Inverse fast Fourier transform (IFFT) processing and time-domain processing may be performed separately for each stream. The streams may be mapped to two 80 MHz channels, and the data may be transmitted by the transmitting STA. At the receiver of the receiving STA, the operations described above for the 80+80 configuration may be reversed, and the combined data may be transmitted to the medium access control (MAC).
[0049]
[0062] Sub-1 GHz operating modes are supported by 802.11af and 802.11ah. Channel operating bandwidths and carriers are reduced in 802.11af and 802.11ah compared to those used in 802.11n and 802.11ac. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV White Space (TVWS) spectrum, while 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. According to representative embodiments, 802.11ah may support meter-type control / machine-type communications, such as MTC devices, within a macro coverage area. MTC devices may have limited functionality, including specific features, such as support for (e.g., support only) specific and / or limited bandwidths. MTC devices may include batteries with above-threshold battery life (e.g., to maintain very long battery life).
[0050]
[0063] WLAN systems that can support multiple channels and channel bandwidths, such as 802.11n, 802.11ac, 802.11af, and 802.11ah, include a channel that can be designated as a primary channel. The primary channel can have a bandwidth equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel can be configured and / or limited by a STA that supports the smallest bandwidth operating mode among all STAs operating in the BSS. In an 802.11ah example, the primary channel can be 1 MHz wide for STAs (e.g., MTC-type devices) that support (e.g., only support) 1 MHz mode, even if the AP and other STAs in the BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operating modes. Carrier sensing and / or network allocation vector (NAV) configuration can depend on the status of the primary channel. For example, if the primary channel is busy because a STA (that only supports 1 MHz operating mode) is transmitting to the AP, the entire available frequency band may be considered busy, even though most of the frequency band may remain idle and available for use.
[0051]
[0064] In the United States, the available frequency bands for use with 802.11ah are 902MHz to 928MHz. In South Korea, the available frequency bands are 917.5MHz to 923.5MHz. In Japan, the available frequency bands are 916.5MHz to 927.5MHz. The total available bandwidth for 802.11ah is 6MHz to 26MHz depending on the country code.
[0052]
[0065] 1D is a system diagram illustrating the RAN 113 and the CN 115 according to one embodiment. As mentioned above, the RAN 113 may employ NR radio technology and communicate with the WTRUs 102a, 102b, 102c over the air interface 116. The RAN 113 may also communicate with the CN 115.
[0053]
[0066] While the RAN 113 may include gNBs 180a, 180b, and 180c, it should be understood that the RAN 113 may include any number of gNBs while remaining consistent with an embodiment. The gNBs 180a, 180b, and 180c may each include one or more transceivers for communicating with the WTRUs 102a, 102b, and 102c over the air interface 116. In one embodiment, the gNBs 180a, 180b, and 180c may implement MIMO technology. For example, the gNBs 180a and 180b may utilize beamforming to transmit signals to and / or receive signals from the gNBs 180a, 180b, and 180c. Thus, for example, the gNB 180a may use multiple antennas to transmit wireless signals to and / or receive wireless signals from the WTRU 102a. In one embodiment, the gNBs 180a, 180b, 180c may implement carrier aggregation technology. For example, the gNB 180a may transmit multiple component carriers to the WTRU 102a (not shown). A subset of these component carriers may be on an unlicensed spectrum, and the remaining component carriers may be on a licensed spectrum. In one embodiment, the gNBs 180a, 180b, 180c may implement Coordinated Multi-Point (CoMP) technology. For example, the WTRU 102a may receive coordinated transmissions from the gNBs 180a and 180b (and / or 180c).
[0054]
[0067] The WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using transmissions associated with scalable numerology. For example, the OFDM symbol spacing and / or OFDM subcarrier spacing may vary for different transmissions, different cells, and / or different portions of the wireless transmission spectrum. The WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using subframes or transmission time intervals (TTIs) of various or scalable lengths (e.g., including various numbers of OFDM symbols and / or lasting various lengths of absolute time).
[0055]
[0068] The gNBs 180a, 180b, 180c may be configured to communicate with the WTRUs 102a, 102b, 102c in a standalone configuration and / or a non-standalone configuration. In a standalone configuration, the WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c without accessing another RAN (e.g., eNode-Bs 160a, 160b, 160c, etc.). In a standalone configuration, the WTRUs 102a, 102b, 102c may utilize one or more of the gNBs 180a, 180b, 180c as mobility anchor points. In a standalone configuration, the WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using signals in unlicensed bands. In a non-standalone configuration, the WTRUs 102a, 102b, 102c may communicate / connect with a gNB 180a, 180b, 180c while also communicating / connecting with another RAN, such as an eNode-B 160a, 160b, 160c. For example, the WTRUs 102a, 102b, 102c may implement a DC principle to communicate with one or more gNBs 180a, 180b, 180c and one or more eNode-Bs 160a, 160b, 160c substantially simultaneously. In a non-standalone configuration, the eNode-Bs 160a, 160b, 160c may act as a mobility anchor for the WTRUs 102a, 102b, 102c, and the gNBs 180a, 180b, 180c may provide additional coverage and / or throughput for serving the WTRUs 102a, 102b, 102c.
[0056]
[0069] Each of the gNBs 180a, 180b, 180c may be associated with a particular cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, scheduling of users in the UL and / or DL, support for network slicing, dual connectivity, interworking between NR and E-UTRA, routing of user plane data towards user plane functions (UPFs) 184a, 184b, routing of control plane information towards access and mobility management functions (AMFs) 182a, 182b, etc. As shown in FIG. 1D , the gNBs 180a, 180b, 180c may communicate with each other via an Xn interface.
[0057]
[0070] 1D may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and possibly a Data Network (DN) 185a, 185b. While each of the above elements is depicted as part of the CN 115, it should be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.
[0058]
[0071] The AMF 182a, 182b may be connected to one or more of the gNBs 180a, 180b, 180c in the RAN 113 via an N2 interface and may function as a control node. For example, the AMF 182a, 182b may be responsible for authenticating users of the WTRUs 102a, 102b, 102c, supporting network slicing (e.g., handling different PDU sessions with different requirements), selecting a particular SMF 183a, 183b, managing registration areas, terminating NAS signaling, mobility management, etc. Network slicing may be used by the AMF 182a, 182b to customize the CN support of the WTRUs 102a, 102b, 102c based on the type of service being utilized by the WTRUs 102a, 102b, 102c. For example, different network slices may be established for different use cases, such as services relying on Ultra-Reliable Low Latency (URLLC) access, services relying on enhanced Massive Mobile Broadband (eMBB) access, services for Machine Type Communications (MTC) access, etc. The AMF 162 may provide a control plane function for switching between the RAN 113 and other RANs (not shown) that employ other radio technologies, such as LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies, such as WiFi.
[0059]
[0072] The SMFs 183a and 183b may be connected to the AMFs 182a and 182b in the CN 115 via an N11 interface. The SMFs 183a and 183b may also be connected to the UPFs 184a and 184b in the CN 115 via an N4 interface. The SMFs 183a and 183b may select and control the UPFs 184a and 184b and configure the routing of traffic through the UPFs 184a and 184b. The SMFs 183a and 183b may perform other functions such as managing and assigning UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, providing downlink data notification, etc. The PDU session type may be IP-based, non-IP-based, Ethernet-based, etc.
[0060]
[0073] The UPFs 184a, 184b may be connected to one or more of the gNBs 180a, 180b, 180c in the RAN 113 via an N3 interface, which may provide the WTRUs 102a, 102b, 102c with access to packet-switched networks such as the Internet 110 to facilitate communications between the WTRUs 102a, 102b, 102c and IP-enabled devices. The UPFs 184, 184b may perform other functions such as routing and forwarding packets, enforcing user plane policy, supporting multi-homed PDU sessions, handling user plane QoS, buffering downlink packets, providing mobility anchoring, etc.
[0061]
[0074] The CN 115 may facilitate communication with other networks. For example, the CN 115 may include or communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between the CN 115 and the PSTN 108. In addition, the CN 115 may provide the WTRUs 102a, 102b, 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, the WTRUs 102a, 102b, 102c may be connected to the local data networks (DNs) 185a, 185b via the UPFs 184a, 184b via an N3 interface to the UPFs 184a, 184b and an N6 interface between the UPFs 184a, 184b and the DNs 185a, 185b.
[0062]
[0075] 1A-1D and the corresponding descriptions thereof, one or more or all of the functions described herein with respect to one or more of the WTRUs 102a-d, base stations 114a-b, eNode-Bs 160a-c, MME 162, SGW 164, PGW 166, gNBs 180a-c, AMFs 182a-b, UPFs 184a-b, SMFs 183a-b, DNs 185a-b, and / or any other devices described herein may be performed by one or more emulation devices (not shown). The emulation devices may be one or more devices configured to emulate one or more or all of the functions described herein. For example, the emulation devices may be used to test other devices and / or to simulate network and / or WTRU functionality.
[0063]
[0076] The emulation device may be designed to perform one or more tests of other devices in a laboratory environment and / or an operator network environment. For example, one or more emulation devices may perform one or more or all functions while fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices in the communication network. One or more emulation devices may perform one or more or all functions while temporarily implemented / deployed as part of a wired and / or wireless communication network. The emulation device may be directly coupled to another device for testing purposes and / or may perform testing using over-the-air wireless communications.
[0064]
[0077] One or more emulation devices may perform one or more functions, including all functions, without being implemented / deployed as part of a wired and / or wireless communication network. For example, the emulation devices may be utilized in a test lab and / or in a test scenario in an undeployed (e.g., test) wired and / or wireless communication network to perform testing of one or more components. One or more emulation devices may be test equipment. Direct RF coupling and / or wireless communication via RF circuitry (which may include, e.g., one or more antennas) may be used by the emulation devices to transmit and / or receive data.
[0065]
[0078] This application describes various aspects, including tools, features, examples, models, techniques, and the like. Many of these aspects are described with specificity, often in a manner that may sound limiting, at least to illustrate their individual characteristics. However, this is for the purpose of clarity of description and does not limit the applicability or scope of the aspects. In fact, all of the different aspects may be combined and interchanged to provide further aspects. Moreover, these aspects may also be combined and interchanged with aspects described in prior applications.
[0066]
[0079] Aspects described and contemplated in this application may be implemented in many different forms. Figures 5-11 described herein may provide some examples, but other examples are contemplated. Discussion of Figures 5-11 does not limit the breadth of implementations. At least one of the aspects relates generally to video encoding and decoding, and at least one other aspect relates generally to transmitting a generated or encoded bitstream. These and other aspects may be implemented as a method, an apparatus, a computer-readable storage medium having stored thereon instructions for encoding or decoding video data according to any of the described methods, and / or a computer-readable storage medium having stored thereon a bitstream generated according to any of the described methods.
[0067]
[0080] In this application, the terms "reconstructed" and "decoded" may be used interchangeably, the terms "pixel" and "sample" may be used interchangeably, and the terms "image," "picture," and "frame" may be used interchangeably.
[0068]
[0081] Various methods are described herein, each of which includes one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. Additionally, terms such as “first,” “second,” etc. may be used in various examples to modify elements, components, steps, operations, etc. (e.g., “first decode” and “second decode,” etc.). The use of such terms does not imply an ordering to the modified operations unless specifically required. Thus, in this example, the first decode need not be performed before the second decode, but may occur before, during, or during an overlapping period with the second decode.
[0069]
[0082] 2 and 3, various methods and other aspects described herein may be used to modify modules, such as decoding modules, of video encoder 200 and decoder 300. Moreover, the subject matter disclosed herein may apply to any type, format, or version of video coding, whether described in a standard or recommendation, whether existing or developed in the future, and to any extensions of such standards and recommendations. Unless otherwise indicated or technically precluded, aspects described herein may be used individually or in combination.
[0070]
[0083] Various numerical values of bits, bit depth, etc. are used in the examples described in this application. These and other specific values are for illustrative purposes only, and the aspects described are not limited to these specific values.
[0071]
[0084] 2 illustrates an exemplary video encoder. While variations of exemplary encoder 200 are contemplated, encoder 200 is described below for clarity and without describing all anticipated variations.
[0072]
[0085] Before being encoded, a video sequence may undergo encoding pre-processing (201), such as applying a color transformation to an input color picture (e.g., converting from RGB 4:4:4 to YCbCr 4:2:0) or performing a remapping of the input picture components to obtain a signal distribution that is more resistant to compression (e.g., using histogram equalization of one of the color components). Metadata may be associated with the pre-processing and added to the bitstream.
[0073]
[0086] In encoder 200, a picture is encoded by encoder elements as described below. The picture to be encoded is divided (202) into units, e.g., coding units (CUs), and processed. Each unit is coded, e.g., using either intra mode or inter mode. When a unit is coded in intra mode, intra prediction is performed (260). In inter mode, motion estimation (275) and compensation (270) are performed. The encoder determines (205) whether intra mode or inter mode should be used to code the unit, and indicates the intra / inter decision, e.g., by a prediction mode flag. A prediction residual is calculated, e.g., by subtracting (210) the prediction block from the original image block.
[0074]
[0087] The prediction residual is then transformed (225) and quantized (230). The quantized transform coefficients, as well as motion vectors and other syntax elements, are entropy coded (245) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can also bypass both the transform and quantization, i.e., the residual is coded directly without applying the transform or quantization processes.
[0075]
[0088] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residual. The decoded prediction residual is combined with the prediction block (255) to reconstruct an image block. An in-loop filter (265) is applied to the reconstructed picture to perform, for example, deblocking / sample adaptive offset (SAO) filtering to reduce coding artifacts. The filtered image is stored in a reference picture buffer (280).
[0076]
[0089] Figure 3 illustrates an example video decoder. In decoder 300, a bitstream is decoded by decoder elements as described below. Video decoder 300 generally performs a decoding pass that is the reverse of the encoding pass as described in Figure 2. Encoder 200 also generally performs video decoding as part of encoding the video data.
[0077]
[0090] In particular, the decoder's input includes a video bitstream, which may be generated by the video encoder 200. The bitstream is first entropy decoded (330) to obtain transform coefficients, motion vectors, and other coding information. Picture partition information indicates how the picture is partitioned. Thus, the decoder may partition the picture according to the decoded picture partition information (335). The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residual. The decoded prediction residual is combined with a prediction block (355) to reconstruct an image block. The prediction block may be obtained from intra prediction (360) or motion-compensated prediction (i.e., inter prediction) (375) (370). An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (380).
[0078]
[0091] The decoded picture may further undergo post-decoding processing (385), such as an inverse color transform (e.g., converting from YCbCr 4:2:0 to RGB 4:4:4) or an inverse remapping that performs the reverse of the remapping process performed in the pre-encoding processing (201). The post-decoding processing may use metadata derived in the pre-encoding processing and signaled in the bitstream. In one example, the decoded image (e.g., after application of an in-loop filter (365) and / or post-decoding processing (385), if post-decoding processing is used) may be sent to a display device for rendering to a user.
[0079]
[0092] FIG. 4 illustrates an example of a system in which various aspects and examples described herein may be implemented. System 400 may be embodied as a device including various components described below and configured to perform one or more of the aspects described herein. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. The elements of system 400, singly or in combination, may be embodied in a single integrated circuit (IC), multiple ICs, and / or separate components. For example, in at least one example, the processing elements and encoder / decoder elements of system 400 are distributed across multiple ICs and / or separate components. In various examples, system 400 is communicatively coupled to one or more other systems or other electronic devices, for example, via a communication bus or via dedicated input and / or output ports. In various examples, system 400 is configured to implement one or more of the aspects described herein.
[0080]
[0093] The system 400 includes at least one processor 410 configured to execute loaded instructions, for example, to implement various aspects described herein. The processor 410 may include embedded memory, input / output interfaces, and various other circuitry known in the art. The system 400 includes at least one memory 420 (e.g., a volatile memory device and / or a non-volatile memory device). The system 400 includes a storage device 440, which may include non-volatile memory and / or volatile memory, including, but not limited to, electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash, magnetic disk drives, and / or optical disk drives. The storage device 440 may include, by way of non-limiting example, an internal storage device, an attached storage device (including removable and non-removable storage devices), and / or a network-accessible storage device.
[0081]
[0094] System 400 includes, for example, an encoder / decoder module 430 configured to process data to provide encoded or decoded video, which may include its own processor and memory. Encoder / decoder module 430 represents a module that may be included in a device that performs encoding and / or decoding functions. As is known, a device may include one or both of an encoding module and a decoding module. Additionally, encoder / decoder module 430 may be implemented as a separate element of system 400 or may be incorporated within processor 410 as a combination of hardware and software, as known to those skilled in the art.
[0082]
[0095] Program code loaded onto the processor 410 or the encoder / decoder 430 to perform various aspects described herein may be stored in the storage device 440 and then loaded onto the memory 420 for execution by the processor 410. According to various examples, one or more of the processor 410, the memory 420, the storage device 440, and the encoder / decoder module 430 may store one or more of various items during performance of the processes described herein. Such stored items include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and arithmetic logic.
[0083]
[0096] In some examples, memory internal to the processor 410 and / or the encoder / decoder module 430 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other examples, memory external to the processing device (e.g., the processing device may be either the processor 410 or the encoder / decoder module 430) is used for one or more of these functions. The external memory may be the memory 420 and / or the storage device 440, such as dynamic volatile memory and / or non-volatile flash memory. In some examples, external non-volatile flash memory is used to store, for example, the television's operating system. In at least one example, high-speed external dynamic volatile memory, such as RAM, is used as working memory for video encoding and decoding operations.
[0084]
[0097] Input to the elements of system 400 may be provided via various input devices, as shown in block 445. Such input devices may include, but are not limited to, (i) a radio frequency (RF) section that receives, for example, RF signals broadcast by a broadcast station, (ii) a component (COMP) input terminal (or set of COMP input terminals), (iii) a universal serial bus (USB) input terminal, and / or (iv) a high definition multimedia interface (HDMI) input terminal. Another example, not shown in FIG. 4, is composite video.
[0085]
[0098] In various examples, the input devices of block 445 have associated respective input processing elements, as is known in the art. For example, the RF section may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal or bandlimiting a signal to a band of frequencies), (ii) downconverting the selected signal, (iii) bandlimiting again to a narrower frequency band to select a signal frequency band, which in particular examples may be referred to as a channel (for example), (iv) demodulating the downconverted and bandlimited signal, (v) performing error correction, and / or (vi) demultiplexing to select a desired data packet stream. The RF section in various examples includes one or more elements for performing these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a downconverter, a demodulator, an error corrector, and a demultiplexer. The RF section may include, for example, a tuner for performing various of these functions, including downconverting a received signal to a lower frequency (e.g., an intermediate frequency or a frequency near baseband) or to baseband. In one example set-top box, the RF section and its associated input processing elements receive RF signals transmitted over a wired (e.g., cable) medium and perform frequency selection by filtering, downconverting, and filtering again to a desired frequency band. Various examples rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements can include inserting elements between existing elements, such as inserting an amplifier and an analog-to-digital converter. In various examples, the RF section includes an antenna.
[0086]
[0099] The USB and / or HDMI terminals may include respective interface processors for connecting system 400 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of the input processing, e.g., Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or within processor 410, as desired. Similarly, aspects of the USB or HDMI interface processing may be implemented, as desired, within a separate interface IC or within processor 410. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 410 and encoder / decoder 430, which operate in combination with memory and storage elements to process the data stream as desired for presentation on an output device.
[0087]
[0100] The various elements of system 400 may be provided within a unified housing in which the various elements may be interconnected and data may be transmitted between them using suitable connection arrangements 425, such as internal buses known in the art, including the Inter-IC (I2C) bus, wiring, and printed circuit boards.
[0088]
[0101] System 400 includes a communication interface 450 that enables communication with other devices over a communication channel 460. Communication interface 450 may include, but is not limited to, a transceiver configured to transmit and receive data over communication channel 460. Communication interface 450 may include, but is not limited to, a modem or a network card, and communication channel 460 may be implemented in a wired and / or wireless medium, for example.
[0089]
[0102] In various examples, data is streamed or otherwise provided to system 400 using a wireless network such as a Wi-Fi network, e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signal in these examples is received via communication channel 460 and communication interface 450 adapted for Wi-Fi communication. Communication channel 460 in these examples is typically connected to an access point or router that provides access to external networks, including the Internet, to enable streaming applications and other over-the-top communications. Another example provides streamed data to system 400 using a set-top box that delivers data via an HDMI connection in input block 445. Yet another example provides streamed data to system 400 using an RF connection in input block 445. As noted above, various examples provide data in a non-streaming manner. Additionally, various examples use wireless networks other than Wi-Fi, such as a cellular network or a Bluetooth network.
[0090]
[0103] The system 400 can provide output signals to various output devices, including a display 475, speakers 485, and other peripheral devices 495. The display 475 in various examples includes, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 475 may be for a television, a tablet, a laptop, a mobile phone, or other device. The display 475 can also be integrated with other components (e.g., as in a smartphone) or separate (e.g., an external monitor for a laptop). The other peripheral devices 495, in various examples, include one or more of a standalone digital video disc (or digital versatile disc) (both terms DVD), a disc player, a stereo system, and / or a lighting system. Various examples use one or more peripheral devices 495 that provide functionality based on the output of the system 400. For example, a disc player performs the function of playing the output of the system 400.
[0091]
[0104] In various examples, control signals are communicated between system 400 and display 475, speakers 485, or other peripheral devices 495 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that allow control between devices with or without user intervention. Output devices may be communicatively coupled to system 400 via dedicated connections through respective interfaces 470, 480, and 490. Alternatively, output devices may be connected to system 400 via communication interface 450 using communication channel 460. Display 475 and speakers 485 may be integrated into a single unit with other components of system 400, for example, in an electronic device such as a television. In various examples, display interface 470 includes a display driver, such as, for example, a timing controller (T Con) chip.
[0092]
[0105] For example, if the RF portion of input 445 is part of a separate set-top box, display 475 and speakers 485 may alternatively be separate from one or more of the other components. In various examples where display 475 and speakers 485 are external components, the output signal may be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.
[0093]
[0106] These examples may be performed by computer software implemented by processor 410, by hardware, or by a combination of hardware and software. As a non-limiting example, these examples may be implemented by one or more integrated circuits. Memory 420 may be of any type suitable for the technology environment and may be implemented using any suitable data storage technology, such as, by way of non-limiting examples, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. Processor 410 may be of any type suitable for the technology environment and may include, by way of non-limiting examples, one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.
[0094]
[0107] Various implementations involve decoding. As used herein, "decoding" may encompass all or part of the processes performed on a received encoded sequence, for example, to produce a final output suitable for display. In various examples, such processes include one or more of the processes typically performed by a decoder (e.g., entropy decoding, inverse quantization, inverse transform, and differential decoding). In various examples, such processes may also, or alternatively, include processes performed by decoders of various implementations described herein, such as obtaining a first prediction signal for a current block using a block vector-based intra prediction mode, obtaining a second prediction signal for the current block using a second prediction mode, generating a prediction block based on at least the first prediction signal obtained using the block vector-based intra prediction and the second intra prediction signal obtained using the prediction mode, and decoding the current block based on the prediction block.
[0095]
[0108] As a further example, in one example, "decoding" may refer only to entropy decoding, in another example, "decoding" may refer only to differential decoding, and in another example, "decoding" may refer to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to refer specifically to a subset of operations or to the broader decoding process generally should be clear based on the context of the specific description and is believed to be well understood by those skilled in the art.
[0096]
[0109] Various implementations involve encoding. Similar to what was discussed above with respect to “decoding,” as used herein, “encoding” can encompass, for example, all or part of the processes performed on an input video sequence to produce an encoded bitstream. In various examples, such processes include one or more of the processes typically performed by an encoder (e.g., partitioning, differential encoding, transforming, quantizing, and entropy coding). In various examples, such processes may also, or alternatively, include processes performed by an encoder of various implementations described herein, such as obtaining a first prediction signal for a current block using a block vector-based intra prediction mode, obtaining a second prediction signal for the current block using a second prediction mode, generating a prediction block based on at least the first prediction signal obtained using the block vector-based intra prediction and the second intra prediction signal obtained using the prediction mode, and encoding the current block based on the prediction block.
[0097]
[0110] As a further example, in one example, "encoding" may refer only to entropy encoding, in another example, "encoding" may refer only to differential encoding, and in another example, "encoding" may refer to a combination of differential encoding and entropy encoding. Whether the phrase "encoding process" is intended to refer specifically to a subset of operations or to the broader encoding process in general should be clear based on the context of the specific description and is believed to be well understood by those skilled in the art.
[0098]
[0111] When a figure is presented as a flow diagram, it should be understood that the figure also provides a block diagram of the corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that the figure also provides a flow diagram of the corresponding method / process.
[0099]
[0112] Implementations and aspects described herein may be implemented in, for example, a method or process, an apparatus, a software program, a data stream, or a signal. Even if discussed only in the context of a single implementation (e.g., discussed only as a method), the implementation of the discussed features may also be implemented in other forms (e.g., an apparatus or a program). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. A method may be implemented in, for example, a processor, which refers to a general processing device including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include, for example, communication devices such as computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end users.
[0100]
[0113] References to "one example" or "one example" or "one implementation" or "an implementation," as well as other variations thereof, mean that a particular feature, structure, characteristic, etc. described in connection with that example is included in at least one example. Thus, appearances of the phrases "in one example" or "in one example" or "in one implementation" or "in one implementation," as well as any other variations thereof, appearing in various places throughout this application are not necessarily all referring to the same example.
[0101]
[0114] Additionally, the application may refer to "determining" various information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from memory. Obtaining may include receiving, retrieving, constructing, generating, and / or determining.
[0102]
[0115] Additionally, the application may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, moving information, replicating information, calculating information, determining information, predicting information, or estimating information.
[0103]
[0116] Additionally, the present application may refer to "receiving" various information. Receiving, like "accessing," is intended to be a broad term. Receiving information can include, for example, one or more of accessing information or retrieving information (e.g., from memory). Furthermore, "receiving" typically involves in some way an operation such as, for example, storing information, processing information, transmitting information, moving information, replicating information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0104]
[0117] It should be understood that the use of any of the following terms " / ," "and / or," and "at least one of" is intended to encompass the selection of only the first listed alternative (A), or the selection of only the second listed alternative (B), or the selection of both alternatives (A and B), for example, "A / B," "A and / or B," and "at least one of A and B." As a further example, for "A, B, and / or C" and "at least one of A, B, and C," such phrases are intended to encompass the selection of only the first listed alternative (A), or the selection of only the second listed alternative (B), or the selection of only the third listed alternative (C), or the selection of only the first and second listed alternatives (A and B), or the selection of only the first and third listed alternatives (A and C), or the selection of only the second and third listed alternatives (B and C), or the selection of all three alternatives (A, B, and C). This can be expanded as many times as there are listed items, as would be apparent to one of ordinary skill in the art.
[0105]
[0118] Also, as used herein, the term "signal" refers, among other things, to indicating something to a corresponding decoder. An encoder signal may include, for example, a coding function for a block's input using a precision factor, etc. In this way, in one example, the same parameters are used on both the encoder and decoder sides. Thus, for example, the encoder may send specific parameters to the decoder so that the decoder can use the same specific parameters (explicit signaling). Conversely, if the decoder already has specific parameters and other parameters, signaling may be used without sending them, simply to allow the decoder to know and select the specific parameters (implicit signaling). By avoiding sending any actual function, bit savings are realized in various examples. It should be understood that signaling can be achieved in various ways. For example, in various examples, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder. Although the above relates to the verb form of the term "signal," the term "signal" may also be used herein as a noun (e.g., as a noun).
[0106]
[0119] As will be apparent to those skilled in the art, multiple implementations can produce a variety of signals formatted to carry information that can be, for example, stored or transmitted. The information can include, for example, instructions for performing a method or data produced by one of the described implementations. For example, a signal can be formatted to carry a bit stream of the described example. Such a signal can be formatted, for example, as an electromagnetic wave (e.g., using a radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. As is known, the signal can be transmitted over a variety of different wired or wireless links. The signal can be stored on, accessed from, or received from a processor-readable medium.
[0107]
[0120] Many examples are described herein. Example features may be provided, alone or in any combination, across various claim categories and types. Furthermore, examples may include one or more of the features, devices, or aspects described herein, alone or in any combination, across various claim categories and types. For example, features described herein may be implemented in a bitstream or signal including information generated as described herein. This information may enable a decoder to decode the bitstream, where the encoder, bitstream, and / or decoder are according to any of the described embodiments. For example, features described herein may be implemented by creating and / or transmitting and / or receiving and / or decoding a bitstream or signal. For example, features described herein may be implemented as a method, process, apparatus, medium storing instructions, medium storing data, or signal. For example, features described herein may be implemented by a television, a set-top box, a mobile phone, a tablet, or other electronic device that performs decoding. A television, set-top box, mobile phone, tablet, or other electronic device may display (e.g., using a monitor, screen, or other type of display) the resulting image (e.g., an image from the residual reconstruction of the video bitstream). A television, set-top box, mobile phone, tablet, or other electronic device may receive a signal containing the encoded image and perform decoding.
[0108]
[0121] These examples may be performed by a device having at least one processor. The device may be an encoder or a decoder. These examples may be performed by a computer program product stored on a non-transitory computer-readable medium and including program code instructions. These examples may be performed by a computer program including program code instructions.
[0109]
[0122] Here, examples of intra template prediction (intra-TMP) and intra block copy (IBC) are provided. Intra-TMP can be enabled for at least one of camera capture or screen content. Intra-TMP may provide a trade-off between gain and coding time. In examples, intra-TMP can be combined with other coding modes (e.g., like other intra modes, e.g., enhanced combined inter-intra prediction (CIIP) mode can combine inter-prediction and intra-prediction). Intra-TMP and / or IBC may be combined with other intra-prediction coding modes. In examples, at least one of the following may apply: CIIP may use (e.g., may be combined with) intra-TMP and / or IBC, geometric partitioning mode (GPM) may use (e.g., may be combined with) intra-TMP and / or IBC, decoder-side intra-mode derivation (DIMD) may use (e.g., may be combined with) intra-TMP and / or IBC, or template-based intra-mode derivation (TIMD) may use (e.g., may be combined with) intra-TMP and / or IBC.
[0110]
[0123] Here, we provide an example of template-based intra mode derivation (TIMD). For each intra prediction mode (e.g., each intra prediction mode) in a most probable mode (MPM) list, a difference such as the sum of absolute transform differences (SATD) between the template's predicted sample and the reconstructed sample can be calculated. The intra prediction mode with the smallest SATD (e.g., the first two intra prediction modes) can be selected as the TIMD mode. The TIMD modes (e.g., these two TIMD modes) can be weighted and the weighted intra prediction can be used to code the current CU. Position-dependent intra prediction combining (PDPC) can be included in the derivation of the TIMD mode.
[0111]
[0124] The costs of the two selected modes can be compared to a threshold, and the test applies a cost factor of 2 as follows: cost mode 2 < 2 * cost mode 1. If this condition is true, fusion can be applied, otherwise only mode 1 can be used. The weights of the modes can be calculated from their SATD costs as follows: Weight1 = CostMode2 / (CostMode1 + CostMode2) Weight 2 = 1 - Weight 1
[0112]
[0125] Here, we provide an example of decoder-side intra-mode derivation (DIMD). When DIMD is applied, intra-modes (e.g., two intra-modes) may be derived from reconstructed adjacent samples, and their predictors (e.g., two predictors) may be combined with a planar mode predictor with weights derived from gradients (Gx, Gy) calculated on the reconstructed adjacent samples. The division operation in the weight derivation may be accomplished using a look-up table (LUT) (e.g., the same LUT-based integerization scheme as used by CCLM mode). In the example, the division operation in the orientation calculation is as follows: Direction=Gy / Gx can be calculated by the following LUT-based scheme: x=Floor(Log2(Gx)) normDiff=((Gx<<4)>>x)&15 x+=(3+(normDiff!=0)?1:0) Azimuth=(Gy*(DivSigTable[normDiff]|8)+(1<<(x-1)))>>x During the ceremony, DivSigTable
[16] ={0,7,6,5,5,4,4,3,3,2,2,1,1,1,1,0} is.
[0113]
[0126] The derived intra modes can be included in the primary list of the intra MPM so that the DIMD process can be performed before building the MPM list. The primary derived intra modes of a DIMD block can be stored with the block and used to build the MPM lists of neighboring blocks.
[0114]
[0127] Here, we provide an example of combining CIIP with TIMD and template matching merging. In CIIP, a prediction sample can be generated by weighting an inter-prediction signal predicted using a CIIP-TM merge candidate and an intra-prediction signal predicted using a TIMD-derived intra-prediction mode. This combination can be applied to coded blocks with an area of 1024 or less (e.g., can only be applied to coded blocks with an area of 1024 or less).
[0115]
[0128] TIMD derivation may be used to derive intra-prediction modes in the CIIP. The intra-prediction mode with the smallest SATD value in the TIMD mode list may be selected and mapped to one of the 67 normal intra-prediction modes.
[0116]
[0129] 5A-5B illustrate examples of splitting for angular modes. The weights of the two tests (wintra, winter) can be modified (e.g., when the derived intra-prediction mode is an angular mode). FIG. 5A shows a vertically split block (e.g., the current block) that can be applied to a near-horizontal mode (2≦angular mode index<34). FIG. 5B shows a horizontally split block (e.g., the current block) that can be applied to a near-vertical mode (34≦angular mode index≦66).
[0117]
[0130] The (wintra, winter) for different sub-blocks are shown in Table 1 below, which lists the modified weights used in angle mode.
[0118] [Table 1]
[0119]
[0131] In the CIIP template matching example, a CIIP template matching merge candidate list can be constructed for the CIIP template matching mode. The merge candidates can be narrowed down by template matching. The CIIP template merge candidates can be sorted (e.g., reordered) by adaptive sorting of merge candidates (ARMC) as regular merge candidates. The maximum number of CIIP template matching merge candidates can be equal to 2.
[0120]
[0132] 6A-6C illustrate an example of a GPM with inter-prediction and intra-prediction. In the example of a GPM with inter-prediction and intra-prediction, the final predicted sample may be generated by weighting the inter-predicted sample and the intra-predicted sample for a GPM separation region (e.g., each GPM generation region). The inter-predicted sample may be derived by the inter GPM, while the intra-predicted sample may be derived by an intra-prediction mode (IPM) candidate list and an index signaled from the encoder. The size of the IPM candidate list may be predefined as three. The available IPM candidates may be at least one of a parallel angle mode (parallel mode) relative to the GPM block boundary, a perpendicular angle mode (vertical mode) relative to the GPM block boundary, or a planar mode, as shown in FIGS. 6A-6C. A GPM with intra-prediction and intra-prediction, as shown in FIG. 6D, may be limited to reduce the signaling overhead of the IPM, which can avoid increasing the size of the intra-prediction circuit on the hardware decoder. Direct motion vector and IPM storage in GPM blending can be introduced to improve (eg, further improve) coding performance.
[0121]
[0133] In DIMD and neighboring mode-based IPM derivation, a parallel mode may be registered first. Two IMP candidates (e.g., a maximum of two IMP candidates) may be derived from DIMD, and / or neighboring blocks may be registered if the same IPM candidate does not exist in the list. For neighboring mode derivation, there may be (e.g., a maximum of) five available neighboring block positions. As shown in Table 2 below, the positions may be limited by the angle of the GPM block boundary, which may already be used for GPM with template matching (GPM-TM). As shown in Table 2, the available neighboring block positions for IPM candidates may be derived based on the angle of the GPM block boundary. A and L indicate the upper and left sides of the prediction block.
[0122] [Table 2]
[0123]
[0134] GPM-Intra can be combined with GPM with motion vector merging (GPM-MMVD). TIMD may be used on the IPM candidates of GPM-Intra, which may improve (e.g., further improve) coding performance. Parallel modes may be registered first, and then TIMD, DIMD, and IPM candidates of neighboring blocks may be registered.
[0124]
[0135] 7 illustrates an example of an intra-template matching search area used in intra-TMP. Intra-TMP is an intra-prediction mode that can copy the best predicted block from a reconstructed portion of a current picture (e.g., a current frame), whose L-shaped template matches the current template. For a predefined search range, the encoder can search for a template that is most similar to the current template in the reconstructed portion of the current picture and use the corresponding block as the predicted block. The encoder can (e.g., then) signal the use of this mode. The same prediction operation can be performed at the decoder side.
[0125]
[0136] The prediction signal may be generated by matching the L-shaped causal neighborhood of the current block with another block within a predefined search area in FIG. 7, which includes: R1: Current CTU R2: Upper left CTU R3: Upper CTU R4: Left CTU
[0126]
[0137] The sum of absolute differences (SAD) may be used as a cost function. Within a region (e.g., within each region), the decoder may search for the template with the smallest SAD relative to the current template and use the corresponding block as the predicted block. The region dimensions (SearchRange_w, SearchRange_h) may be set proportional to the block dimensions (BlkW, BlkH) so that the number of SAD comparisons per pixel is constant. That is, SearchRange_w=a*BlkW SearchRange_h=a*BlkH where "a" may be a constant that controls the gain / complexity tradeoff. For example, "a" may be equal to 5.
[0127]
[0138] The intra-TMP tool can be enabled for CUs with width and height sizes less than or equal to 64. The maximum CU size for this intra-TMP may be configurable. The intra-TMP mode can be signaled at the CU level through a dedicated flag.
[0128]
[0139] The block vector-based intra prediction mode may be combined with other prediction modes (e.g., CIIP, DIMD, TIMD, GPM) to provide more efficient and improved prediction. In an example, a video decoding or encoding device may obtain a first prediction signal for a current block using the block vector-based intra prediction mode. A second prediction signal for the current block may be obtained using the second prediction mode. A prediction block may be generated by weighting the first prediction signal using the block vector-based intra prediction mode and the second prediction signal using the second prediction mode. In an example, weighting the first and second prediction signals may include multiplying the first and second prediction signals by a weighting factor. For an N×M block, samples (e.g., each sample) within the block may be multiplied by a weighting factor. The current block may be decoded or encoded based on the prediction block. In an example, the block vector-based intra prediction mode may be an intra template prediction (intra-TMP) mode or an intra block copy (IBC) mode.
[0129]
[0140] Here, an example is provided in which a block vector-based intra prediction mode (e.g., intra-TMP mode and / or IBC mode) is used as one of the modes of CIIP. CIIP may be a combination of inter prediction and intra prediction. Predictions (e.g., each prediction) may be weighted by a predefined weighting factor. CIIP cannot be applied to an I slice because inter prediction information is unavailable. Therefore, the inter prediction mode may be replaced with a block vector-based intra prediction mode, allowing both predictions to use the intra prediction mode.
[0130]
[0141] FIG. 8 illustrates an example of a CIIP with intra-intra prediction. In intra-intra prediction, the inter portion of the CIIP may be replaced with a block vector-based intra prediction mode (e.g., intra-TMP mode and / or IBC mode). A first prediction signal may be obtained by using the block vector-based intra prediction mode. As shown in FIG. 8, in intra-intra prediction, the intra portion of the CIIP may use TIMD mode derivation. In an example, a normal intra prediction mode may be derived from the TIMD mode derivation. In an example, the normal intra prediction mode may also be block vector-based intra prediction. A second prediction signal may be obtained using the normal intra prediction mode. A prediction block may be generated by weighting the first prediction signal derived using the block vector-based intra prediction and the second prediction signal derived using the normal intra prediction mode.
[0131]
[0142] FIG. 9 illustrates an example of a CIIP with inter-intra prediction. In inter-intra prediction, the intra portion of the CIIP can be replaced with block vector-based intra prediction (e.g., intra-TMP mode and / or IBC mode). A first prediction signal can be obtained by using intra prediction based on block vectors. In inter-intra, the inter portion of the CIIP can use a template matching merge prediction mode. In an example, a CIIP template matching merge candidate list can be constructed for the CIIP template matching mode. Merge candidates can be narrowed down by template matching. A second intra prediction signal can be obtained using an inter prediction mode (e.g., template matching prediction mode). A predicted sample block can be generated by weighting the first prediction signal using the block vector-based intra prediction and the second prediction signal using the inter prediction mode.
[0132]
[0143] FIG. 10 illustrates an example of using a DIMD mode in combination with a block vector-based intra prediction mode (e.g., intra-TMP mode and / or IBC mode). A video decoder or encoder may determine that a current block is coded using a DIMD mode. A first prediction signal may be obtained by using the block vector-based intra prediction mode. A second set of prediction modes may be derived using DIMD based on a histogram of gradients as described herein. The second set of prediction modes may be up to five intra prediction modes. The second set of intra prediction modes may include a normal angular prediction mode. A second set of prediction signals may be obtained using a second set of intra prediction modes. A prediction block may be generated by weighting the first prediction signal and the second set of prediction signals.
[0133]
[0144] In an example, the block vector-based intra prediction mode may be blended with DIMD (e.g., a second set of prediction modes derived using DIMD) rather than with planar mode (e.g., the block vector-based intra prediction mode may replace the planar mode). This may provide more efficient and improved prediction. In an example, an indication (e.g., a flag) may be signaled indicating that the block vector-based intra prediction mode (e.g., rather than the planar mode) is used for blending with DIMD. In an example, both the planar mode and the block vector-based intra prediction mode may be tested on a reconstructed template (e.g., similar to a TIMD process). The video decoder or encoder may determine whether to blend the planar mode or the block vector-based intra prediction mode based on, for example, minimizing a difference (e.g., SATD) between the reconstructed template and the prediction template of the current block. The block vector-based intra prediction mode is used to obtain a first prediction signal based on determining that the block vector-based intra prediction mode minimizes the distance (e.g., SATD).
[0134]
[0145] Here, we provide an example of combining intra-TMP and / or IBC with TIMD. In TIMD, modes (e.g., all modes) in an MPM list are tested on templates surrounding the block. A first intra-prediction signal can be obtained for the block using intra-TMP and / or IBC. A second intra-prediction signal can be obtained for the block using TIMD.
[0135]
[0146] FIG. 11 illustrates an example of TIMD combined with a block vector-based intra prediction mode (e.g., intra-TMP mode and / or IBC mode). A video decoder or encoder may determine that a current block is coded using a TIMD mode. A first prediction signal may be obtained by using the block vector-based intra prediction mode. A second set of prediction modes may be derived using TIMD based on testing the most likely modes in the MPM list tested on a template associated with the current block. The second set of prediction modes may be two normal intra prediction modes. The block vector-based intra prediction mode may be blended with the two normal intra prediction modes in TIMD. A second set of prediction signals may be obtained using the second set of intra prediction modes. A prediction block may be generated by weighting the first prediction signal and the second set of prediction signals.
[0136]
[0147] In an example, using a block vector-based intra prediction mode with TIMD can be achieved by including intra TMP and / or IBC as one of the candidates to be tested on the template. Thus, the intra TMP mode and / or IBC can be tested (for example, in addition to the MPM mode). The SATD can be calculated in the TIMD, and based on which the intra TMP mode and / or IBC is selected, the intra TMP mode and / or IBC mode can be used and blended with the TIMD. In an example, the two normal intra prediction modes in TIMD can always be blended with the block vector-based intra prediction mode without testing the intra TMP and / or IBC as one of the candidates to be tested on the template.
[0137]
[0148] The video encoder and / or decoder may identify neighboring blocks and the current block. The neighboring blocks may be associated with block vectors. In addition to deriving a block vector-based intra-prediction mode (e.g., intra-TMP and / or IBC) for the current block, the neighboring information (e.g., via the block vectors of the neighboring blocks) may be used. The neighboring information (e.g., the block vectors of the neighboring blocks) may generate a prediction signal (e.g., a first prediction signal and / or a second prediction signal) using an intra-prediction mode (e.g., intra-TMP, IBC, DIMD, TIMD, etc.). That is, if any of the neighboring blocks uses intra-TMP, its block vector may be used as a prediction candidate for TIMD. Specifically, the block vector may be used to obtain a reference sample for predicting the current template. The template cost may be used to obtain a TIMD best mode. If any of the best modes is intra-TMP, the same block vector may be used to generate a prediction for the current block.
[0138]
[0149] If any of the neighboring blocks use IBC, then its block vectors can be used in the same way as intra-TMP: the block vectors can be used to predict the template, obtain the template cost, and generate the prediction signal if selected by the TIMD process.
[0139]
[0150] In examples, both intra-TMP and IBC may generate multiple block vectors. That is, when using bidirectional IBC, two block vectors can be associated with the block. Similarly, when intra-TMP fusion is used, multiple block vectors may be used. For TIMD, three options may be considered: use a single block vector (e.g., always use), test a vector (e.g., all vectors) in the TIMD process, or use a vector with an appropriate fusion example.
[0140]
[0151] To test block vectors (e.g., all vectors) in the TIMD process, each block vector can be considered independent and used to generate a prediction signal. Predictions can be tested using the TIMD process and selected accordingly. The same prediction mechanism can be used to use vectors with their appropriate fusion instances. Bidirectional IBC or fusion-based intra-TMP can be tested on templates and selected according to the TIMD process.
[0141]
[0152] The block vectors tested in TIMD may be refined using template costs (e.g., to further improve coding performance). That is, a block vector (e.g., a block vector obtained from a neighboring block) may be tested on the reconstructed template of the current block to obtain a current template cost (e.g., measured by SATD). A refinement range (e.g., around a rectangle formed by (-1, 1) horizontally and vertically) may be used to obtain a better block vector with a lower template cost compared to the current block vector. This may improve prediction quality because a better block vector may be used, which may have a shorter template distance than a block vector obtained directly from a neighboring block. The refinement step may be performed at the end (e.g., only at the end) of the TIMD process when a block vector is selected.
[0142]
[0153] The intra-TMP search process may be simplified for blocks employing TIMD with intra-TMP: for example, if a block vector is selected, refinement in the intra-TMP may be bypassed and refinement may be performed in the TIMD process.
[0143]
[0154] If a block vector is selected, a block vector derived from TIMD may be used for block vector propagation. For example, if the TIMD process results in the use of a block vector, the block vector may be stored for use (e.g., further use) in an IBC merge mode, a chroma block vector, or a TIMD mode (e.g., a further TIMD mode) that uses a further adjacent block vector.
[0144]
[0155] In examples, a directional intra mode can be derived and used for transform selection. Because prediction can be a mixture of normal vector-based and block vector-based, the selection of the transform kernel can be unclear (e.g., it can be intra-mode dependent). In examples, prediction can be considered as planar prediction, and a transform kernel corresponding to the planar mode can be used. In some examples, the intra mode can be derived using DIMD on the prediction signal (e.g., already used for MIP mode and intra-TMP mode). For example, if TIMD with intra-TMP is used, a DIMD process can be applied to the prediction signal to obtain a directional mode. Then, a transform kernel corresponding to the directional intra mode derived from DIMD can be used.
[0145]
[0156] While features and elements are described above in particular combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with the other features and elements. In addition, the methods described herein may be implemented in a computer program, software, or firmware embodied in a computer-readable medium for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted via wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random-access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs). A processor, together with software, may be used to implement a radio frequency transceiver for use in a WTRU, UE, terminal, base station, RNC, or any host computer.
Claims
1. 1. A device for video decoding, the device comprising:
1. A processor, comprising: Obtaining a first prediction signal for the current block using a block vector-based intra prediction mode; obtaining a second prediction signal for the current block using a second prediction mode; generating a prediction block based on at least the first prediction signal obtained using the block vector-based intra prediction and the second intra prediction signal obtained using the prediction mode; Decoding the current block based on the predicted block. A processor configured to A device comprising:
2. the second prediction mode is a normal intra prediction mode; the prediction block is generated by weighting the first prediction signal using the block vector-based intra prediction mode and the second prediction signal using the normal intra prediction mode. The device of claim 1 .
3. the second prediction mode is an inter prediction mode; the prediction block is generated by weighting the first prediction signal using the block vector-based intra prediction mode and the second prediction signal using the inter prediction mode. The device of claim 1 .
4. The second prediction mode is one of a set of second prediction modes, and the second prediction signal is one of a set of second prediction signals, and the processor: determining that the current block is coded in decoder-side intra-mode derivation (DIMD) mode; deriving the second set of prediction modes based on a histogram of gradients; obtaining the second set of prediction signals using the second set of prediction modes; weighting the first prediction signal and the second set of prediction signals to generate the prediction block; The device of claim 1 further configured to:
5. The processor: determining whether a planar mode or the block vector-based intra prediction mode minimizes a difference between a reconstructed template and a prediction template for the block; obtaining the first prediction signal by using the block vector-based intra prediction mode based on determining that the block vector-based intra prediction mode minimizes the distance; The device of claim 4 further configured to:
6. The second prediction mode is one of a set of second prediction modes, and the second prediction signal is one of a set of second prediction signals, and the processor: determining that the current block is coded using a template-based intra-mode derivation (TIMD) mode; deriving the second set of prediction modes based on testing a most probable mode on a template associated with the current block; obtaining the second set of prediction signals using the second set of prediction modes; weighting the first prediction signal and the second set of prediction signals to generate the prediction block; The device of claim 1 further configured to:
7. The device according to any one of claims 1 to 6, wherein the block vector-based intra prediction mode is an intra template prediction (intra-TMP) mode or an intra block copy (IBC) mode.
8. 1. A method for video decoding, the method comprising: Obtaining a first prediction signal for a current block using a block vector-based intra prediction mode; obtaining a second prediction signal for the current block using a second prediction mode; and generating a prediction block based on at least the first prediction signal obtained using the block vector-based intra prediction and the second intra prediction signal obtained using the prediction mode; and decoding the current block based on the predicted block.
9. the second prediction mode is a normal intra prediction mode; the prediction block is generated by weighting the first prediction signal using the block vector-based intra prediction mode and the second prediction signal using the normal intra prediction mode. The method of claim 8.
10. the second prediction mode is an inter prediction mode; the prediction block is generated by weighting the first prediction signal using the block vector-based intra prediction mode and the second prediction signal using the inter prediction mode. The method of claim 8.
11. the second prediction mode is one of a set of second prediction modes, and the second prediction signal is one of a set of second prediction signals; determining that the current block is coded in decoder-side intra-mode derivation (DIMD) mode; deriving the second set of prediction modes based on a histogram of gradients; obtaining the second set of prediction signals using the second set of prediction modes; weighting the first prediction signal and the second set of prediction signals to generate the prediction block; The method of claim 8 further comprising:
12. determining whether a planar mode or the block vector-based intra prediction mode minimizes a difference between a reconstructed template and a prediction template for the block; obtaining the first prediction signal by using the block vector-based intra prediction mode based on determining that the block vector-based intra prediction mode minimizes the distance; and The method of claim 11 further comprising:
13. the second prediction mode is one of a set of second prediction modes, and the second prediction signal is one of a set of second prediction signals; determining that the current block is coded using a template-based intra-mode derivation (TIMD) mode; deriving the second set of prediction modes based on testing most probable modes on a template associated with the current block; and obtaining the second set of prediction signals using the second set of prediction modes; weighting the first prediction signal and the second set of prediction signals to generate the prediction block; The method of claim 8 further comprising:
14. The method according to any one of claims 8 to 13, wherein the block vector based intra prediction mode is an intra template prediction (intra-TMP) mode or an intra block copy (IBC) mode.
15. 1. A device for video encoding, said device comprising:
1. A processor, comprising: Obtaining a first prediction signal for the current block using a block vector-based intra prediction mode; obtaining a second prediction signal for the current block using a second prediction mode; generating a prediction block based on at least the first prediction signal obtained using the block vector-based intra prediction and the second intra prediction signal obtained using the prediction mode; encoding the current block based on the predicted block; A processor configured to A device comprising:
16. the second prediction mode is a normal intra prediction mode; the prediction block is generated by weighting the first prediction signal using the block vector-based intra prediction mode and the second prediction signal using the normal intra prediction mode.
16. The device of claim 15.
17. the second prediction mode is an inter prediction mode; the prediction block is generated by weighting the first prediction signal using the block vector-based intra prediction mode and the second prediction signal using the inter prediction mode.
16. The device of claim 15.
18. The second prediction mode is one of a set of second prediction modes, and the second prediction signal is one of a set of second prediction signals, and the processor: determining that the current block is coded in decoder-side intra-mode derivation (DIMD) mode; deriving the second set of prediction modes based on a histogram of gradients; obtaining the second set of prediction signals using the second set of prediction modes; weighting the first prediction signal and the second set of prediction signals to generate the prediction block; The device of claim 15 further configured to:
19. The processor: determining whether a planar mode or the block vector-based intra prediction mode minimizes a difference between a reconstructed template and a prediction template for the block; obtaining the first prediction signal by using the block vector-based intra prediction mode based on determining that the block vector-based intra prediction mode minimizes the distance; 20. The device of claim 18, further configured to:
20. The second prediction mode is one of a set of second prediction modes, and the second prediction signal is one of a set of second prediction signals, and the processor: determining that the current block is coded using a template-based intra-mode derivation (TIMD) mode; deriving the second set of prediction modes based on testing a most probable mode on a template associated with the current block; obtaining the second set of prediction signals using the second set of prediction modes; weighting the first prediction signal and the second set of prediction signals to generate the prediction block; The device of claim 15 further configured to:
21. The device according to any one of claims 15 to 20, wherein the block vector-based intra prediction mode is an intra template prediction (intra-TMP) mode or an intra block copy (IBC) mode.
22. 1. A method for video encoding, the method comprising: Obtaining a first prediction signal for a current block using a block vector-based intra prediction mode; obtaining a second prediction signal for the current block using a second prediction mode; and generating a prediction block based on at least the first prediction signal obtained using the block vector-based intra prediction and the second intra prediction signal obtained using the prediction mode; encoding the current block based on the predicted block.
23. the second prediction mode is a normal intra prediction mode; the prediction block is generated by weighting the first prediction signal using the block vector-based intra prediction mode and the second prediction signal using the normal intra prediction mode.
23. The method of claim 22.
24. the second prediction mode is an inter prediction mode; the prediction block is generated by weighting the first prediction signal using the block vector-based intra prediction mode and the second prediction signal using the inter prediction mode.
23. The method of claim 22.
25. the second prediction mode is one of a set of second prediction modes, and the second prediction signal is one of a set of second prediction signals; determining that the current block is coded in decoder-side intra-mode derivation (DIMD) mode; deriving the second set of prediction modes based on a histogram of gradients; obtaining the second set of prediction signals using the second set of prediction modes; weighting the first prediction signal and the second set of prediction signals to generate the prediction block; 23. The method of claim 22, further comprising:
26. determining whether a planar mode or the block vector-based intra prediction mode minimizes a difference between a reconstructed template and a prediction template for the block; obtaining the first prediction signal by using the block vector-based intra prediction mode based on determining that the block vector-based intra prediction mode minimizes the distance; and 26. The method of claim 25, further comprising:
27. the second prediction mode is one of a set of second prediction modes, and the second prediction signal is one of a set of second prediction signals; determining that the current block is coded using a template-based intra-mode derivation (TIMD) mode; deriving the second set of prediction modes based on testing most probable modes on a template associated with the current block; and obtaining the second set of prediction signals using the second set of prediction modes; weighting the first prediction signal and the second set of prediction signals to generate the prediction block; 23. The method of claim 22, further comprising:
28. The method according to any one of claims 22 to 27, wherein the block vector based intra prediction mode is an intra template prediction (intra-TMP) mode or an intra block copy (IBC) mode.
29. A computer program product stored on a non-transitory computer readable medium and comprising program code instructions for performing the steps of the methods according to at least one of claims 8 to 14 and claims 22 to 28 when executed by at least one processor.
30. A computer readable medium comprising program code instructions for performing the steps of the method according to at least one of claims 8 to 14 and claims 22 to 28 when executed by a processor.
31. Video data comprising information representative of an encoded output produced according to the method of any one of claims 22 to 28.