Intrablock copy and intratemplate matching predictive filtering
By applying regression-based mean squared error minimization techniques for filter parameter determination in intra-block copy mode, the video encoding and decoding processes are enhanced, addressing inefficiencies in existing video coding systems.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- INTERDIGITALCE PATENT HLDG SAS
- Filing Date
- 2024-04-05
- Publication Date
- 2026-06-02
AI Technical Summary
Existing video coding systems face challenges in efficiently encoding and decoding video data using intra-block copy and intra-template matching prediction, particularly in determining optimal filter parameters for improved prediction accuracy.
The implementation of filter parameters based on regression-based mean squared error minimization techniques for intra-block copy mode, using smoothing filters and multiple of 2, to enhance prediction block filtering and encoding/decoding processes.
Improves the accuracy and efficiency of video encoding and decoding by optimizing filter parameters for intra-block copy and intra-template matching prediction, leading to better compression and transmission of video data.
Smart Images

Figure 2026517647000001_ABST
Abstract
Description
[Background technology]
[0001] Cross-reference of related applications This application claims the benefit of European Patent Application No. 23305530.0, filed on April 7, 2023, the entirety of which is incorporated herein by reference.
[0002] A video coding system can be used to compress digital video signals, for example, to reduce the storage bandwidth and / or transmission bandwidth required for such signals. Video coding systems can include, for example, block-based, wavelet-based, and / or object-based systems. [Overview of the Initiative] [Means for solving the problem]
[0003] Systems, methods, and means for video encoding and / or decoding using intra-block copy (IBC) and / or intra-template matching prediction (intraTMP) filtering are disclosed herein. In an example, a video decoder or video encoder may determine that the current block is encoded in IBC mode. Based on the current block being encoded in intra-block copy mode, filter parameters may be determined based on the template samples of the prediction block and the template samples of the current block. In an example, at least one of the filter parameters may be a smoothing filter and a multiple of 2. In an example, the filter parameters may be further determined using a regression-based mean squared error (MSE) minimization technique. In an example, the prediction block may be the best candidate prediction block from a list of candidate prediction blocks, and the best candidate prediction block may be determined based on template cost.
[0004] The samples of the prediction block can be filtered based on the determined filter parameters. In the example, a subset of the filtered samples can be filtered based on the determined filter parameters. In the example, the filtered subset can be the upper and left prediction samples of the prediction block, and the template samples of the current block can be the upper and left templates of the current block. The upper prediction samples can be filtered using the upper template of the current block, and the left prediction samples can be filtered using the left template of the current block. The current block can be encoded or decoded based on the filtered samples.
[0005] These examples can be implemented by a device having a processor. The device may be an encoder or a decoder. These examples can be implemented by a computer program product, which includes program code instructions and is stored on a non-temporary computer-readable medium. These examples can be implemented by a computer program comprising program code instructions.
[0006] The systems, methods, and means described herein may involve decoders. In some examples, the systems, methods, and means described herein may involve encoders. In some examples, the systems, methods, and means described herein may involve signals (for example, from and / or received by encoders). Computer-readable media may include instructions for causing one or more processors to perform the methods described herein. Computer program products may include instructions that cause one or more processors to perform the methods described herein when the program is executed by one or more processors. [Brief explanation of the drawing]
[0007] [Figure 1A] This is a system diagram showing an exemplary communication system in which one or more disclosed embodiments may be implemented. [Figure 1B] This figure shows an exemplary wireless transceiver unit (WTRU) used in the communication system shown in Figure 1A, according to one embodiment. [Figure 1C] This is a system diagram showing an exemplary radio access network (RAN) and core network (CN) used in the communication system shown in Figure 1A, according to one embodiment. [Figure 1D] This is a system diagram showing further exemplary RAN and CN used in the communication system of Figure 1A according to one embodiment. [Figure 2]This is a diagram illustrating an example video encoder. [Figure 3] This is a diagram illustrating an exemplary video decoder. [Figure 4] This figure shows an example of a system in which various forms and examples can be implemented. [Figure 5] This figure shows an example of discontinuity in the predicted block relative to the current block template when copying a predicted block in intrablock copy (IBC) or intratemplate matching prediction (intraTMP). [Figure 6] This figure shows an example of an intraTMP search area. [Figure 7A] This figure shows an example of an IBC reference region corresponding to block prediction. [Figure 7B] This figure shows an example of an IBC reference region corresponding to block prediction. [Figure 7C] This figure shows an example of an IBC reference region corresponding to block prediction. [Figure 7D] [Figure 8] This figure shows an example of a reference area for IBC when an encoded tree unit (CTU) (m,n) is encoded. [Figure 9A] This figure shows an exemplary sample used by position-dependent intra-prediction combinations (PDPCs) applied to diagonal and adjacent angular intra-modes. [Figure 9B] This figure shows an exemplary sample used by position-dependent intra-predictive combinations (PDPC) applied to diagonal and adjacent angle intra-modes. [Figure 9C] This figure shows an exemplary sample used by position-dependent intra-predictive combinations (PDPC) applied to diagonal and adjacent angle intra-modes. [Figure 9D]A diagram showing exemplary samples used by a position-dependent intra prediction combination (PDPC) applied to diagonal and adjacent angle intra modes. [Figure 10] A diagram showing an example of a spatial part of a convolutional filter. [Figure 11] A diagram showing an example of a reference area used to derive filter coefficients. [Figure 12] A diagram showing an example of filtering a single line (P00 to P03 and P10 to P30 samples). [Figure 13] A diagram showing an example of the filter shape of a reference block and a training area. [Figure 14] A diagram showing an example of calculating filter parameters between a reference template and a current template. [Figure 15] A diagram showing an example of calculating filter parameters between a reference template and a reference block. [Figure 16] A diagram showing an example of calculating filter parameters between a current template and a reference block.
Best Mode for Carrying Out the Invention
[0008] A more detailed understanding can be obtained from the following description, given by way of example together with the accompanying drawings.
[0009] Figure 1A is a system diagram showing an exemplary communication system 100 in which one or more disclosed embodiments may be implemented. The communication system 100 may be a multiple access system that provides content such as voice, data, video, messaging, and broadcast to multiple radio users. The communication system 100 can enable multiple radio users to access such content through the sharing of system resources, including radio bandwidth. For example, the communication system 100 may employ one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single-carrier FDMA (SC-FDMA), zero-tail (ZT) unique-word (UW) discrete Fourier transform (DFT) spread OFDM (ZT UW DTS-s OFDM), unique-word OFDM (UW-OFDM), resource block filtering OFDM, and filter bank multicarrier (FBMC).
[0010] As shown in Figure 1A, the communication system 100 may include radio transceiver units (WTRUs) 102a, 102b, 102c, 102d, radio access networks (RANs) 104 / 113, core networks (CNs) 106 / 115, public switched telephone networks (PSTNs) 108, the Internet 110, and other networks 112, but it will be understood that the disclosed embodiments intend any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, and 102d may be any type of device configured to operate and / or communicate in a radio environment. For example, WTRU102a, 102b, 102c, and 102d may all be referred to as “stations” and / or “STAs” and may be configured to transmit and / or receive radio signals, and may include (or be) user equipment (UEs), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearables, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in an industrial and / or automated processing chain context), consumer electronics devices, and devices operating on commercial and / or industrial wireless networks. Any of WTRU102a, 102b, 102c, and 102d may interchangeably be referred to as UEs.
[0011] The communication system 100 may also include base stations 114a and / or base stations 114b. Each of the base stations 114a and 114b may be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, and 102d to facilitate access to one or more communication networks, such as CN 106 / 115, the Internet 110, and / or network 112. As an example, base stations 114a and 114b may be any of the following: base station transceiver station (BTS), node B (NB), e-node B (eNB), home node B (HNB), home e-node B (HeNB), g-node B (gNB), NR node B (NR NB), site controller, access point (AP), wireless router, etc. Although base stations 114a and 114b are shown as single elements, it will be understood that base stations 114a and 114b may include any number of interconnected base stations and / or network elements.
[0012] Base station 114a may be part of RAN 104 / 113, which may also include other base stations and / or network elements (not shown), such as a base station controller (BSC), a radio network controller (RNC), and relay nodes. Base station 114a and / or base station 114b may be configured to transmit and / or receive radio signals on one or more carrier frequencies, which may be called cells (not shown). These frequencies may be licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell can provide coverage for radio services to a particular geographic area that may be relatively fixed or change over time. A cell may be further divided into cell sectors. For example, a cell associated with base station 114a may be divided into three sectors. Thus, in one embodiment, base station 114a may include three transceivers, i.e., one for each sector of the cell. In one embodiment, base station 114a may employ multiple-input multiple-output (MIMO) technology, which may utilize multiple transceivers for each sector of the cell or any sector. For example, beamforming can be used to transmit and / or receive signals in a desired spatial direction.
[0013] Base stations 114a and 114b can communicate with one or more WTRUs 102a, 102b, 102c, and 102d via an air interface 116, the air interface 116 may be any suitable radio communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 116 may be established using any suitable radio access technology (RAT).
[0014] More specifically, as described above, the communication system 100 may be a multiple access system and may employ one or more channel access schemes such as CDMA, TDMA, FDMA, OFDMA, and SC-FDMA. For example, base stations 114a and WTRU 102a, 102b, and 102c in RAN 104 / 113 may implement radio technologies such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which can establish an air interface 116 using broadband CDMA (WCDMA). WCDMA may include communication protocols such as High Speed Packet Access (HSPA) and / or Advanced HSPA (HSPA+). HSPA may include High Speed Downlink Packet Access (HSDPA) and / or High Speed Uplink Packet Access (HSUPA).
[0015] In one embodiment, base stations 114a and WTRUs 102a, 102b, and 102c can implement radio technologies such as Advanced UMTS Terrestrial Radio Access (E-UTRA), which can establish an air interface 116 using Long-Term Evolution (LTE) and / or LTE Advanced (LTE-A) and / or LTE Advanced Pro (LTE-A Pro).
[0016] In one embodiment, base stations 114a and WTRUs 102a, 102b, and 102c can implement radio technologies such as NR radio access, which can establish an air interface 116 using New Radio (NR).
[0017] In one embodiment, base stations 114a and WTRUs 102a, 102b, and 102c can implement multiple radio access technologies. For example, base stations 114a and WTRUs 102a, 102b, and 102c can implement LTE radio access and NR radio access together, for example, using the dual connectivity (DC) principle. Thus, the air interface utilized by WTRUs 102a, 102b, and 102c may be characterized by multiple types of radio access technologies and / or transmissions from / to multiple types of base stations (e.g., eNBs and gNBs).
[0018] In one embodiment, base stations 114a and WTRUs 102a, 102b, and 102c can implement wireless technologies such as IEEE 802.11 (i.e., Wireless Fidelity (Wi-Fi)), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile Communications (GSM), GSM Advanced Data Rate (EDGE), and GSM EDGE (GERAN).
[0019] In Figure 1A, base station 114b may be, for example, a wireless router, home node B, home enode B, or access point, and can utilize any suitable RAT to facilitate wireless connectivity in localized areas such as offices, homes, vehicles, premises, industrial facilities, aerial corridors (for use by drones, for example), and roads. In one embodiment, base station 114b and WTRU 102c, 102d can implement wireless technologies such as IEEE 802.11 to establish a wireless local area network (WLAN). In one embodiment, base station 114b and WTRU 102c, 102d can implement wireless technologies such as IEEE 802.15 to establish a wireless personal area network (WPAN). In one embodiment, base station 114b and WTRU 102c, 102d can utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish any small cell, picocell, or femtocell. As shown in Figure 1A, base station 114b may have a direct connection to the internet 110. Therefore, base station 114b may not be required to access the internet 110 via CN 106 / 115.
[0020] RAN104 / 113 may communicate with CN106 / 115, which may be any type of network configured to provide voice, data, applications, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRU102a, 102b, 102c, and 102d. The data may have various Quality of Service (QoS) requirements, including different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, and mobility requirements. CN106 / 115 may provide call control, billing services, mobile location-based services, prepaid calling, internet connectivity, video distribution, and / or implement high-level security functions, such as user authentication. Although not shown in Figure 1A, it will be understood that RAN104 / 113 and / or CN106 / 115 may communicate directly or indirectly with other RANs employing the same or different RATs as RAN104 / 113. For example, in addition to being connected to RAN104 / 113, which may utilize NR radio technology, CN106 / 115 may also communicate with another RAN (not shown) employing one of the following technologies: GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or Wi-Fi radio technology.
[0021] CN106 / 115 can also act as a gateway for WTRU102a, 102b, 102c, and 102d to access PSTN108, the Internet 110, and / or other networks 112. PSTN108 may include a circuit-switched telephone network providing plain old telephone service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices using common communication protocols such as TCP, User Datagram Protocol (UDP), and / or IP in the Transmission Control Protocol / Internet Protocol (TCP / IP) Internet Protocol Suite. Network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, network 112 may include another CN connected to one or more RANs, which may employ the same RAT as RAN104 / 114 or a different RAT.
[0022] Some or all of the WTRUs 102a, 102b, 102c, and 102d in the communication system 100 can include multimode capability (for example, WTRUs 102a, 102b, 102c, and 102d can include multiple transceivers for communicating with different radio networks via different radio links). For example, WTRU 102c shown in Figure 1A may be configured to communicate with base station 114a which can employ cellular-based radio technology and may be configured to communicate with base station 114b which can employ IEEE 802 radio technology.
[0023] Figure 1B is a system diagram showing an exemplary WTRU 102. As shown in Figure 1B, the WTRU 102 may include, in particular, a processor 118, a transceiver 120, a transceiver element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power supply 134, a Global Positioning System (GPS) chipset 136, and / or other elements / peripherals 138. It will be understood that the WTRU 102 may include any partial combination of the above elements while remaining consistent with one embodiment.
[0024] The processor 118 may be a general-purpose processor, a dedicated processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. The processor 118 can perform signal coding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 102 to operate in a wireless environment. The processor 118 may be coupled to the transceiver 120, which may be coupled to the transceiver element 122. Although Figure 1B shows the processor 118 and the transceiver 120 as separate components, it will be understood that the processor 118 and the transceiver 120 may be integrated together, for example, in an electronic package or chip.
[0025] The transmitting / receiving element 122 may be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) via the air interface 116. For example, in one embodiment, the transmitting / receiving element 122 may be an antenna configured to transmit and / or receive RF signals. In one embodiment, the transmitting / receiving element 122 may be an emitter / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In one embodiment, the transmitting / receiving element 122 may be configured to transmit and / or receive both RF signals and optical signals. It will be understood that the transmitting / receiving element 122 may be configured to transmit and / or receive any combination of radio signals.
[0026] Although the transmit / receive element 122 is shown as a single element in Figure 1B, the WTRU 102 can include any number of transmit / receive elements 122. For example, the WTRU 102 can employ MIMO technology. Thus, in one embodiment, the WTRU 102 can include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving radio signals via the air interface 116.
[0027] The transceiver 120 may be configured to modulate the signal to be transmitted by the transmitting / receiving element 122 and to demodulate the signal to be received by the transmitting / receiving element 122. As described above, the WTRU 102 may have multimode capability. Therefore, the transceiver 120 may include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs, such as NR and IEEE 802.11.
[0028] The processor 118 of the WTRU102 may be coupled to a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (for example, a liquid crystal display (LCD) display unit or an organic light-emitting diode (OLED) display unit) and may receive user input data from them. The processor 118 may also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. Furthermore, the processor 118 may access information from any type of suitable memory, such as non-removable memory 130 and / or removable memory 132, and store data therein. Non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. Removable memory 132 may include a subscriber identification module (SIM) card, a memory stick, a secure digital (SD) memory card, etc. In other embodiments, the processor 118 can access information from memory not physically located on the WTRU 102, such as on a server or home computer (not shown), and store data therein.
[0029] The processor 118 may be configured to receive power from the power supply 134 and distribute and / or control power to other components in the WTRU 102. The power supply 134 can be any suitable device for supplying power to the WTRU 102. For example, the power supply 134 may include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), a solar cell, a fuel cell, etc.
[0030] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to, or instead of, the information from the GPS chipset 136, the WTRU 102 may receive location information from base stations (e.g., base stations 114a, 114b) via the air interface 116 and / or determine its location based on the timing of when signals are received from two or more nearby base stations. It will be understood that the WTRU 102 may acquire location information via any preferred location determination method while remaining consistent with one embodiment.
[0031] The processor 118 may further be coupled to other elements / peripherals 138, which may include one or more software and / or hardware modules / units that provide additional features, functionality and / or wired or wireless connectivity. For example, the elements / peripherals 138 may include an accelerometer, an electronic compass, a satellite transceiver, a digital camera (for photos and / or videos), a Universal Serial Bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth® module, a frequency-modulated (FM) radio unit, a digital music player, a media player, a video game player module, an internet browser, a virtual reality and / or augmented reality (VR / AR) device, an activity tracker, and the like. The element / peripheral device 138 may include one or more sensors, the sensors being one or more of the following: gyroscope, accelerometer, Hall effect sensor, magnetometer, compass sensor, proximity sensor, temperature sensor, time sensor, geolocation sensor, altimeter, light sensor, touch sensor, magnetometer, barometer, gesture sensor, biometric sensor, and / or humidity sensor.
[0032] WTRU102 may include a full-duplex radio where the transmission and reception of some or all of a signal may be parallel and / or simultaneous, associated with a specific subframe for both an uplink (for transmission, for example) and a downlink (for reception, for example). The full-duplex radio may include an interference management unit for reducing and / or substantially eliminating self-interference via signal processing either through hardware (e.g., chokes) or through a processor (e.g., a separate processor (not shown) or via processor 118). In one embodiment, WTRU102 may include a half-duplex radio, which is for the transmission and reception of some or all of a signal (e.g., associated with a specific subframe for either an uplink (for transmission, for example) or a downlink (for reception, for example).
[0033] Figure 1C is a system diagram showing RAN104 and CN106 according to one embodiment. As described above, RAN104 can employ E-UTRA radio technology to communicate with WTRU102a, 102b, and 102c via the air interface 116. RAN104 may also communicate with CN106.
[0034] RAN104 may include enodes B160a, 160b, and 160c, but it will be understood that RAN104 may include any number of enodes B while remaining consistent with one embodiment. Each of enodes B160a, 160b, and 160c may include one or more transceivers for communicating with WTRU102a, 102b, and 102c via the air interface 116. In one embodiment, enodes B160a, 160b, and 160c can implement MIMO technology. Thus, enode B160a may, for example, use multiple antennas to transmit radio signals to and receive radio signals from WTRU102a.
[0035] Each of the e-nodes B160a, 160b, and 160c may be associated with a specific cell (not shown) and may be configured to handle wireless resource management decisions, handover decisions, user scheduling on uplink (UL) and / or downlink (DL), etc. As shown in Figure 1C, the e-nodes B160a, 160b, and 160c can communicate with each other via the X2 interface.
[0036] The CN106 shown in Figure 1C may include a Mobility Management Entity (MME) 162, a Serving Gateway (SGW) 164, and a Packet Data Network (PDN) Gateway (PGW) 166. Although each of the above elements is shown as part of CN106, it will be understood that any one of these elements may be owned and / or operated by an entity other than the CN operator.
[0037] The MME162 can be connected to each of the e-nodes B160a, 160b, and 160c in RAN104 via the S1 interface and can act as a control node. For example, the MME162 can be responsible for authenticating users of WTRU102a, 102b, and 102c, activating / deactivating bearers, and selecting a specific serving gateway during the initial attachment of WTRU102a, 102b, and 102c. The MME162 can provide control plane functionality for switching between RAN104 and other RANs (not shown) employing other radio technologies such as GSM and / or WCDMA.
[0038] The SGW164 can be connected to each of the e-nodes B160a, 160b, and 160c in RAN104 via the S1 interface. The SGW164 can generally route and forward user data packets to and from WTRU102a, 102b, and 102c. The SGW164 can perform other functions, such as anchoring the user plane during e-node B handovers, triggering paging when DL data is available for WTRU102a, 102b, and 102c, and managing and remembering the context of WTRU102a, 102b, and 102c.
[0039] SGW164 may be connected to PGW166, which can provide WTRU102a, 102b, and 102c with access to a packet-switched network such as the Internet 110 to facilitate communication between WTRU102a, 102b, and 102c and IP-enabled devices.
[0040] CN106 can facilitate communication with other networks. For example, CN106 can provide WTRU102a, 102b, and 102c with access to circuit-switched networks such as PSTN108, thereby facilitating communication between WTRU102a, 102b, and 102c and legacy landline communication devices. For example, CN106 may include or communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between CN106 and PSTN108. Furthermore, CN106 can provide WTRU102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.
[0041] Although the WTRU is described as a wireless terminal in Figures 1A to 1D, in certain representative embodiments, such a terminal is intended to be able to use a wired communication interface with a communication network (for example, temporarily or permanently).
[0042] In a typical embodiment, the other network 112 may be a WLAN.
[0043] In Infrastructure Basic Service Set (BSS) mode, a WLAN may have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP may have access to or interfaces with a distributed system (DS) or another type of wired / wireless network that carries traffic during and / or from within the BSS. Traffic originating outside the BSS to the STA may arrive through the AP and be delivered to the STA. Traffic originating from the STA to destinations outside the BSS may be sent to the AP to be delivered to their respective destinations. Traffic between STAs within the BSS may be sent through the AP, for example, here, the source STA can send traffic to the AP, and the AP can deliver the traffic to the destination STA. Traffic between STAs within the BSS is considered and / or sometimes referred to as peer-to-peer traffic. Peer-to-peer traffic may be sent between the source STA and the destination STA (for example, directly between them) by a direct link setup (DLS). In some typical embodiments, the DLS may be an 802.11e DLS or an 802.11z tunnel DLS (TDLS). A WLAN using Independent BSS (IBSS) mode may not have access points (APs), and STAs within or using IBSS (for example, all STAs) can communicate directly with each other. The IBSS communication mode is sometimes referred to as the “ad-hoc” communication mode in this specification.
[0044] When using the 802.11ac infrastructure operating mode or a similar operating mode, an AP can transmit beacons on a fixed channel, such as a primary channel. The primary channel can be a fixed width (e.g., a 20 MHz bandwidth) or a dynamically set width via signaling. The primary channel can be the operating channel of the BSS, which can be used by STAs to establish a connection with the AP. In some typical embodiments, Carrier sense multiple access with collision avoidance (CSMA / CA) can be implemented, for example, in an 802.11 system. In CSMA / CA, an STA, including the AP (e.g., any STA), can sense the primary channel. If the primary channel is sensed / detected and / or determined to be busy by a particular STA, that STA can backoff. One STA (e.g., only one station) can transmit at any given time within a given BSS.
[0045] A high-throughput (HT) STA can use a 40MHz wide channel for communication, for example, via a combination of a primary 20MHz channel and adjacent or non-adjacent 20MHz channels to form a 40MHz wide channel.
[0046] Ultra-high throughput (VHT) STAs can support 20MHz, 40MHz, 80MHz, and / or 160MHz wide channels. 40MHz channels and / or 80MHz channels can be formed by combining consecutive 20MHz channels. 160MHz channels can be formed by combining eight consecutive 20MHz channels, or by combining two discontinuous 80MHz channels, sometimes referred to as an 80+80 configuration. In the 80+80 configuration, data can be passed through a segment parser that, after channel encoding, can split the data into two streams. Inverse fast Fourier transform (IFFT) processing and time-domain processing can be performed separately for each stream. The streams can be mapped onto two 80MHz channels, and the data can be transmitted by a transmitting STA. At the receiver of a receiving STA, the operation described above for the 80+80 configuration can be reversed, and the combined data can be sent to a media access control (MAC) layer, entities, etc.
[0047] Sub-1GHz operating modes are supported by 802.11af and 802.11ah. Channel operating bandwidth and carrier are reduced in 802.11af and 802.11ah compared to those used in 802.11n and 802.11ac. 802.11af supports 5MHz, 10MHz, and 20MHz bandwidths in the TV white space (TVWS) spectrum, while 802.11ah supports 1MHz, 2MHz, 4MHz, 8MHz, and 16MHz bandwidths using the non-TVWS spectrum. According to a typical embodiment, 802.11ah can support meter-type control / machine-type communications (MTC), such as MTC devices in a macro coverage area. MTC devices may have limited capabilities, including support for some and / or limited bandwidths (e.g., support only for that). MTC devices may include batteries with above-threshold battery life (e.g., to maintain very long battery life).
[0048] A WLAN system that can support multiple channels and channel bandwidths, such as 802.11n, 802.11ac, 802.11af, and 802.11ah, includes a channel that can be designated as the primary channel. The primary channel may have a bandwidth equal to the largest common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel may be set and / or limited by the STA that supports the minimum bandwidth operating mode from among all STAs operating in the BSS. In the 802.11ah example, the primary channel may be 1 MHz wide for an STA (e.g., an MTC type device) that supports (e.g., only) 1 MHz mode, even if other STAs in the AP and BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operating modes. Carrier detection and / or network allocation vector (NAV) settings may depend on the status of the primary channel. For example, if the primary channel is busy because an STA (which only supports 1MHz operating mode) is transmitting to the AP, the entire available frequency band may be considered busy, even though a large portion of the frequency band remains idle and could be available.
[0049] In the United States, the available frequency band that can be used by 802.11ah is from 902 MHz to 928 MHz. In South Korea, the available frequency band is from 917.5 MHz to 923.5 MHz. In Japan, the available frequency band is from 916.5 MHz to 927.5 MHz. The total bandwidth available for 802.11ah is from 6 MHz to 26 MHz, depending on the country code.
[0050] Figure 1D is a system diagram showing RAN113 and CN115 according to one embodiment. As described above, RAN113 can employ NR radio technology to communicate with WTRU102a, 102b, and 102c via the air interface 116. RAN113 may also communicate with CN115.
[0051] RAN113 may include gNB180a, 180b, and 180c, but it will be understood that RAN113 may include any number of gNBs while remaining consistent with one embodiment. Each of the gNB180a, 180b, and 180c may include one or more transceivers for communicating with WTRU102a, 102b, and 102c via the air interface 116. In one embodiment, the gNB180a, 180b, and 180c can implement MIMO technology. For example, the gNB180a and 180b can utilize beamforming to transmit signals to and / or receive signals from the WTRU102a, 102b, and 102c. Thus, the gNB180a can, for example, use multiple antennas to transmit radio signals to and / or receive radio signals from the WTRU102a. In one embodiment, gNB180a, 180b, and 180c can implement carrier aggregation technology. For example, gNB180a can transmit multiple component carriers to WTRU102a (not shown). A subset of these component carriers may be on the unlicensed spectrum, while the remaining component carriers may be on the licensed spectrum. In one embodiment, gNB180a, 180b, and 180c can implement coordinated multi-point (CoMP) technology. For example, WTRU102a can receive coordinated transmissions from gNB180a and gNB180b (and / or gNB180c).
[0052] WTRU102a, 102b, and 102c can communicate with gNB180a, 180b, and 180c using transmissions associated with scalable numerology. For example, OFDM symbol intervals and / or OFDM subcarrier intervals may differ for different transmissions, different cells, and / or different parts of the radio transmission spectrum. WTRU102a, 102b, and 102c can communicate with gNB180a, 180b, and 180c using subframes or transmit time intervals (TTIs) of varying or scalable lengths (including, for example, a varying number of OFDM symbols and / or a varying length of absolute time that persists).
[0053] gNB180a, 180b, and 180c can be configured to communicate with WTRU102a, 102b, and 102c in standalone and / or non-standalone configurations. In a standalone configuration, WTRU102a, 102b, and 102c can communicate with gNB180a, 180b, and 180c without accessing other RANs (such as e-nodes B160a, 160b, and 160c). In a standalone configuration, WTRU102a, 102b, and 102c can utilize one or more of gNB180a, 180b, and 180c as mobility anchor points. In a standalone configuration, WTRU102a, 102b, and 102c can communicate with gNB180a, 180b, and 180c using signals in unlicensed bands. In a non-standalone configuration, WTRU102a, 102b, and 102c can communicate with gNB180a, 180b, and 180c while also communicating with other RANs such as enodes B160a, 160b, and 160c. For example, WTRU102a, 102b, and 102c can implement DC principles to communicate substantially simultaneously with one or more gNB180a, 180b, and 180c, and one or more enodes B160a, 160b, and 160c. In a non-standalone configuration, enodes B160a, 160b, and 160c can act as mobility anchors for WTRU102a, 102b, and 102c, and gNB180a, 180b, and 180c can provide additional coverage and / or throughput to service WTRU102a, 102b, and 102c.
[0054] Each of the gNB180a, 180b, and 180c may be associated with a specific cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, support for network slicing, dual connectivity, interworking between NR and E-UTRA, routing of user plane data to user plane functions (UPF) 184a and 184b, routing of control plane information to access and mobility management functions (AMF) 182a and 182b, etc. As shown in Figure 1D, the gNB180a, 180b, and 180c can communicate with each other via the Xn interface.
[0055] The CN115 shown in Figure 1D may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and at least one Data Network (DN) 185a, 185b. While each of the above elements is shown as part of the CN115, it will be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.
[0056] AMF182a and 182b can be connected to one or more gNB180a, 180b, and 180c in RAN113 via the N2 interface and can act as control nodes. For example, AMF182a and 182b can be responsible for user authentication of WTRU102a, 102b, and 102c, support for network slicing (e.g., handling different protocol data unit (PDU) sessions with different requirements), selection of specific SMF183a and 183b, management of registration areas, termination of NAS signaling, mobility management, etc. Network slicing can be used by AMF182a and 182b to customize CN support for WTRU102a, 102b, and 102c based on the type of service being used by WTRU102a, 102b, and 102c. For example, different network slices may be established for different use cases, such as services relying on ultra-high reliability low latency (URLLC) access, services relying on enhanced massive mobile broadband (eMBB) access, and services for MTC access. AMF182a, 182b can provide control plane functionality for switching between RAN113 and other RANs (not shown) employing other radio technologies such as LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies such as Wi-Fi.
[0057] SMF183a and 183b can be connected to AMF182a and 182b in CN115 via the N11 interface. SMF183a and 183b can also be connected to UPF184a and 184b in CN115 via the N4 interface. SMF183a and 183b can select and control UPF184a and 184b and configure the routing of traffic through UPF184a and 184b. SMF183a and 183b can perform other functions such as managing and allocating UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, and providing downlink data notifications. PDU session types can be IP-based, non-IP-based, Ethernet-based, etc.
[0058] UPF184a and 184b may be connected via the N3 interface to one or more gNB180a, 180b, and 180c in RAN113, which can provide WTRU102a, 102b, and 102c with access to a packet-switched network, such as the Internet 110, to facilitate communication between WTRU102a, 102b, and 102c and IP-enabled devices. UPF184a and 184b can perform other functions, such as routing and forwarding packets, enforcing user plane policies, supporting multi-homed PDU sessions, handling user plane QoS, buffering downlink packets, and providing mobility anchoring.
[0059] CN115 can facilitate communication with other networks. For example, CN115 may include or be able to communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between CN115 and PSTN108. Furthermore, CN115 can provide WTRU102a,102b,102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, WTRU102a,102b,102c may be connected to DN185a,185b through UPF184a,184b via an N3 interface to UPF184a,184b, and an N6 interface between UPF184a,184b and local data networks (DN) 185a,185b.
[0060] In view of Figures 1A to 1D and their corresponding descriptions, one or more of the functions described herein with respect to WTRU 102a to d, base stations 114a to 1b, e-nodes B160a to 1c, MME 162, SGW 164, PGW 166, gNB 180a to 1c, AMF 182a to 1b, UPF 184a to 1b, SMF 183a to 1b, DN 185a to 1b, and / or any other (one or more) elements / devices described herein may be implemented by one or more emulation elements / devices (not shown). An emulation device may be one or more devices configured to emulate one or more of the functions described herein, or all of them. For example, an emulation device may be used to test other devices and / or to simulate network and / or WTRU functions.
[0061] Emulation devices may be designed to implement one or more tests of other devices in a laboratory environment and / or a carrier network environment. For example, one or more emulation devices may perform one or more, or all, of the functions while fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices in a communication network. One or more emulation devices may perform one or more, or all, of the functions while temporarily implemented / deployed as part of a wired and / or wireless communication network. Emulation devices may be directly coupled to another device for testing purposes and / or tests may be performed using over-the-air wireless communication.
[0062] One or more emulation devices can perform one or more functions, including all of the above, while not implemented / deployed as part of a wired and / or wireless communication network. For example, an emulation device may be used in a test laboratory and / or in a test scenario in a non-deployed (e.g., test) wired and / or wireless communication network to implement testing of one or more components. One or more emulation devices may be test equipment. Direct RF coupling and / or wireless communication via RF circuitry (e.g., which may include one or more antennas) may be used by an emulation device to transmit and / or receive data.
[0063] This application describes various embodiments, including tools, features, examples, models, and methods. Many of these embodiments are described in detail and, at least to illustrate their individual characteristics, are often described in a manner that may seem restrictive. However, this is for the sake of clarity of description and does not limit the application or scope of those embodiments. In fact, all of the different embodiments can be combined and interchangeable to provide further embodiments. Furthermore, embodiments can be combined and interchangeable with embodiments described in previous applications.
[0064] The embodiments described and intended in this application can be implemented in many different forms. Figures 5 to 16 described herein can provide some examples, but other examples are intended. The description of Figures 5 to 16 is not intended to limit the scope of implementation. At least one of the embodiments relates generally to video encoding and decoding, and at least one other embodiment relates generally to transmitting a generated or encoded bitstream. These and other embodiments can be implemented as a computer-readable storage medium storing instructions for encoding or decoding video data according to any of the methods, apparatus, or described methods, and / or a computer-readable storage medium storing a bitstream generated according to any of the described methods.
[0065] In this application, the terms “reconstructed” and “decoded” may be used interchangeably, the terms “pixel” and “sample” may be used interchangeably, and the terms “image,” “picture,” and “frame” may be used interchangeably.
[0066] Various methods are described herein, each of which includes one or more steps or actions to achieve the method described. Unless a particular order of steps or actions is required for the proper operation of the method, the order and / or use of any particular steps and / or actions may be modified or combined. Furthermore, terms such as “first,” “second,” etc., may be used in various examples to modify elements, components, steps, operations, etc., such as “first decoding” and “second decoding.” The use of such terms does not imply ordering of the modified operations unless specifically required. Thus, in this example, the first decoding does not need to be performed before the second decoding, and may occur, for example, before the second decoding, during the second decoding, or during a time period overlapping with the second decoding.
[0067] Various methods and other embodiments described herein can be used to modify modules, such as the decoding module, of the video encoder 200 and decoder 300 as shown in Figures 2 and 3. Furthermore, the subject matter disclosed herein can be applied to any type, format, or version of video encoding, whether described in a standard or recommendation, existing or to be developed in the future, and to any extensions of such standards and recommendations. Unless otherwise specified or technically excluded, the embodiments described herein can be used individually or in combination.
[0068] In the examples described in this application, various numerical values are used, such as bits and bit depth. These and other specific values are for illustrative purposes only, and the embodiments described are not limited to these specific values.
[0069] Figure 2 shows an exemplary video encoder. While variations of the exemplary encoder 200 are intended, the encoder 200 is described below for clarity without describing all possible variations.
[0070] Before encoding, the video sequence may undergo encoding preprocessing (201), such as applying a color conversion to the input color picture (e.g., a conversion from RGB4:4:4 to YCbCr4:2:0), or performing a remapping of the input picture components to obtain a more resilient signal distribution to compression (e.g., using histogram equalization of one of the color components). Metadata may be associated with the preprocessing and attached to the bitstream.
[0071] In encoder 200, the picture is encoded by encoder elements as described below. The picture to be encoded is partitioned (202) and processed in units, for example, encoding units (CUs). Each unit is encoded using either intra-mode or inter-mode, for example. When a unit is encoded in intra-mode, it performs intra-prediction (260). In inter-mode, motion estimation (275) and motion compensation (270) are performed. The encoder determines whether to use intra-mode or inter-mode to encode a unit (205), and indicates the intra / inter determination, for example, by a prediction mode flag. The prediction residual is calculated, for example, by subtracting the predicted blocks from the original image blocks (210).
[0072] The predicted residual is then transformed (225) and quantized (230). The quantized transformation coefficients, as well as the motion vector and other syntax elements, are entropy encoded (245) to output a bitstream. The encoder can skip the transformation and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transformation and quantization, i.e., the residual is encoded directly without the application of any transformation or quantization process.
[0073] The encoder decodes the encoded blocks to provide a reference for further prediction. The quantized transformation coefficients are dequantized (240) and inversely transformed (250) to decode the prediction residuals. Combining the decoded prediction residuals with the predicted blocks (255) reconstructs the image blocks. An in-loop filter (265) is applied to the reconstructed picture to reduce encoding artifacts, for example, by performing deblocking / SAO (sample adaptive offset) filtering. The filtered image is stored in a reference picture buffer (280).
[0074] Figure 3 shows an example of a video decoder. In the exemplary decoder 300, the bitstream is decoded by the decoder elements as described below. The video decoder 300 generally performs a decoding path that is the reverse of the encoding path shown in Figure 2. The encoder 200 also generally performs video decoding as part of encoding the video data.
[0075] In particular, the decoder input includes a video bitstream that can be generated by the video encoder 200. The bitstream is first entropy-decoded to obtain transformation coefficients, motion vectors, and other encoded information (330). Picture partition information indicates how the picture is partitioned. Thus, the decoder can partition the picture according to the decoded picture partition information (335). The transformation coefficients are dequantized (340) and inversely transformed (350) to decode the prediction residuals. Combining the decoded prediction residuals with the predicted blocks (355), the image blocks are reconstructed. The predicted blocks can be obtained from intra-predictions (360) or motion-compensated predictions (i.e., inter-predictions) (375) (370). An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (380).
[0076] The decoded picture may undergo further decoded post-processing (385), such as inverse color conversion (e.g., conversion from YCbCr4:2:0 to RGB4:4:4), or inverse remapping, which performs the reverse of the remapping process performed in the encoding pre-processing (201). The decoded post-processing may use metadata derived in the encoding pre-processing and signaled in the bitstream. For example, the decoded image (e.g., after applying the in-loop filter (365) and / or after decoded post-processing (385) if decoded post-processing is used) can be sent to a display device for rendering to the user.
[0077] Figure 4 shows an example of a system in which various embodiments and examples can be implemented. System 400 can be embodied as a device comprising various components described below and configured to implement one or more of the embodiments described herein. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. The elements of System 400 can be embodied individually or in combination in a single integrated circuit (IC), multiple ICs, and / or individual components. For example, in at least one example, the processing elements and / or encoder / decoder elements of System 400 are distributed across multiple ICs and / or individual components. In various examples, System 400 is communicably coupled to one or more other systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various examples, System 400 is configured to implement one or more of the embodiments described herein.
[0078] System 400 includes, for example, at least one processor 410 configured to execute instructions loaded therein in order to implement various embodiments described herein. The processor 410 may include embedded memory, input / output interfaces, and various other circuits known in the Art. System 400 includes at least one memory 420 (for example, a volatile memory device and / or a non-volatile memory device). System 400 includes a storage device 440 which may include, but is not limited to, electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, magnetic disk drives, and / or optical disk drives, as well as non-volatile and / or volatile memory. The storage device 440 may, in non-limiting examples, include internal storage devices, accessory storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices.
[0079] System 400 includes, for example, an encoder / decoder module 430 configured to process data and provide encoded or decoded video, the encoder / decoder module 430 which may include its own processor and memory. The encoder / decoder module 430 represents one or more modules that may be included in the device for performing encoding and / or decoding functions. As is known, the device may include one or both of the encoding module and the decoding module. Furthermore, the encoder / decoder module 430 may be implemented as a separate element of System 400, as is known to those skilled in the art, or it may be incorporated into the processor 410 as a combination of hardware and software.
[0080] Program code to be loaded onto the processor 410 or encoder / decoder 430 to implement the various embodiments described herein can be stored in the storage device 440 and then loaded onto the memory 420 for execution by the processor 410. According to various examples, one or more of the processor 410, memory 420, storage device 440, and encoder / decoder module 430 can store one or more of various items during the implementation of the processes described herein. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, expressions, actions, and operational logic.
[0081] In some examples, internal memory of the processor 410 and / or encoder / decoder module 430 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other examples, memory outside the processing device (for example, the processing device can be either the processor 410 or the encoder / decoder module 430) is used for one or more of these functions. The external memory can be memory 420 and / or storage device 440, for example, dynamic volatile memory and / or non-volatile flash memory. In some examples, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one example, high-speed external dynamic volatile memory, such as RAM, is used as working memory for video encoding and decoding operations.
[0082] Inputs to the elements of System 400 can be provided through various input devices, as shown in Block 445. Such input devices include, but are not limited to, (i) a radio frequency (RF) portion that receives RF signals, for example, transmitted over the air by a broadcaster; (ii) a component (COMP) input terminal (or set of COMP input terminals); (iii) a universal serial bus (USB) input terminal; and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Other examples not shown in Figure 4 include composite video.
[0083] In various examples, the input device of block 445 has associated input processing elements known in the art. For example, the RF portion may be associated with elements suitable for (i) selecting a desired frequency (also called selecting a signal or band-limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) further band-limiting to a narrower frequency band in order to select a signal frequency band that may (for example) be called a channel in a particular example, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF portion of various examples may include one or more elements for performing these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion may include tuners for performing various functions, such as down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or baseband. In one set-top box example, the RF section and its associated input processing elements receive an RF signal transmitted over a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and further filtering to a desired frequency band. Various examples rearrange the order of the elements described above (and others), remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements can include inserting elements between existing elements, such as inserting an amplifier and an analog-to-digital converter. In various examples, the RF section includes an antenna.
[0084] Furthermore, the USB and / or HDMI terminals may include their respective interface processors for connecting the system 400 to other electronic devices over USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, can be implemented as needed, for example, within a separate input processing IC or within the processor 410. Similarly, aspects of USB or HDMI interface processing can be implemented as needed, within a separate interface IC or within the processor 410. Demodulated, error-corrected, and multiplexed streams are provided to various processing elements, including the processor 410 and an encoder / decoder 430 operating in conjunction with memory and storage elements, for processing the data stream as needed for presentation on an output device, for example.
[0085] Various elements of system 400 can be provided within an integrated housing. Within the integrated housing, the various elements are interconnected and can transmit data between them using an internal bus, such as those known in the art, including a suitable connection arrangement 425, for example, an inter-IC (I2C) bus, wiring, and a printed circuit board.
[0086] System 400 includes a communication interface 450 that enables communication with other devices via a communication channel 460. The communication interface 450 may include, but is not limited to, transceivers configured to transmit and receive data via the communication channel 460. The communication interface 450 may include, but is not limited to, a modem or network card, and the communication channel 460 may be implemented, for example, within a wired and / or wireless medium.
[0087] In various examples, data is streamed to system 400 or otherwise provided using a wireless network, such as a Wi-Fi network, e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signal in these examples is received via a communication channel 460 and communication interface 450 adapted for Wi-Fi communication. In these examples, communication channel 460 is typically connected to an access point or router that provides access to an external network, including the Internet, to enable streaming applications and other over-the-top communications. Other examples provide the streamed data to system 400 using a set-top box that distributes data via an HDMI connection on input block 445. Yet another example provides the streamed data to system 400 using an RF connection on input block 445. As described above, various examples provide data in a non-streaming manner. Furthermore, various examples use wireless networks other than Wi-Fi, e.g., cellular networks or Bluetooth® networks.
[0088] System 400 can provide output signals to various output devices, including a display 475, a speaker 485, and other peripheral devices 495. Various examples of the display 475 include, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 475 may be for a television, tablet, laptop, cell phone (mobile phone), or other device. The display 475 may also be integrated with other components (for example, in the case of a smartphone) or be separate (for example, an external monitor for a laptop). Various examples of the other peripheral devices 495 include, in various examples, one or more of a standalone digital video disc (disc) (or digital multipurpose disc (disc)) (DVD for both terms), a disc player, a stereo system, and / or a lighting system. Various examples use one or more peripheral devices 495 that provide functionality based on the output of System 400. For example, a disc player performs the function of playing the output of System 400.
[0089] In various examples, control signals are communicated between the system 400 and the display 475, speaker 485, or other peripheral devices 495 using signaling, such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable inter-device control with or without user intervention. Output devices can be coupled to the system 400 communicably via dedicated connections through their respective interfaces 470, 480, and 490. Alternatively, output devices can be connected to the system 400 using communication channel 460 via communication interface 450. The display 475 and speaker 485 can be integrated into a single unit along with other components of the system 400 in an electronic device, such as a television. In various examples, the display interface 470 includes a display driver, such as a timing controller (TCon) chip.
[0090] The display 475 and speaker 485 can, alternatively, be separate from one or more of the other components, for example, if the RF portion of input 445 is part of a separate set-top box. In various examples where the display 475 and speaker 485 are external components, the output signals can be provided via dedicated output connections, including, for example, an HDMI port, a USB port, or a COMP output.
[0091] These examples can be implemented by computer software implemented by the processor 410, by hardware, or by a combination of hardware and software. In a non-limiting example, these examples can be implemented by one or more integrated circuits. Memory 420 can be of any type appropriate to the technical environment and, in a non-limiting example, can be implemented using any suitable data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. Processor 410 can be of any type appropriate to the technical environment and, in a non-limiting example, can include one or more of microprocessors, general-purpose computers, dedicated computers, and processors based on multi-core architectures.
[0092] Various implementations are involved in decoding. As used in this application, “decoding” can encompass all or part of a process performed on a received encoded sequence to produce a final output suitable for display. In various examples, such a process may include one or more processes commonly performed by decoders, such as entropy decoding, inverse quantization, inverse transform, and difference decoding. In various examples, such a process may also, or alternatively, include processes performed by decoders in various implementations described in this application, which may include, for example, determining that the current block is encoded in intra-block copy mode, determining filter parameters based on template samples of the predicted block and the current block based on the current block being encoded in intra-block copy mode, filtering the samples of the predicted block based on the determined filter parameters, and decoding the current block based on the filtered samples.
[0093] As further examples, in one example, “decoding” refers only to entropy decoding; in another example, “decoding” refers only to differential decoding; and in yet another example, “decoding” refers to a combination of entropy decoding and differential decoding. Whether the phrase “decoding process” is intended to refer specifically to a subset of operations or to the broader decoding process as a whole will become clear from the context of the detailed description and should be well understood by those skilled in the art.
[0094] Various implementations are involved in encoding. As with the above description of “decoding,” “encoding” as used in this application may encompass all or part of the process performed on an input video sequence to produce an encoded bitstream. In various examples, such a process may include one or more of the processes commonly performed by an encoder, such as segmentation, differential encoding, transformation, quantization, and entropy encoding. In various examples, such a process may also, or alternatively, include processes performed by an encoder in the various implementations described in this application, which may include, for example, determining that the current block is encoded in intra-block copy mode, determining filter parameters based on template samples of a predicted block and template samples of the current block based on the current block being encoded in intra-block copy mode, filtering the samples of the predicted block based on the determined filter parameters, and encoding the current block based on the filtered samples.
[0095] As further examples, in one example, “encoding” refers only to entropy encoding; in another example, “encoding” refers only to differential encoding; and in yet another example, “encoding” refers to a combination of differential encoding and entropy encoding. Whether the phrase “encoding process” is intended to refer specifically to a subset of operations or to the broader encoding process as a whole will become clear from the context of the detailed description and should be well understood by those skilled in the art.
[0096] When a diagram is presented as a flow chart, it should be understood that it also provides a block diagram of the corresponding device. Similarly, when a diagram is presented as a block diagram, it should be understood that it also provides a flow chart of the corresponding method / process.
[0097] The implementations and embodiments described herein may be implemented, for example, in methods or processes, apparatus, software programs, data streams, or signals. Even when a feature is discussed only in the context of a single form of implementation (for example, only as a method), the implementation of the feature discussed may also be implemented in other forms (for example, apparatus or programs). Apparatus may be implemented, for example, with appropriate hardware, software, and firmware. Methods may be implemented in, for example, processors, which generally refer to processing devices, including computers, microprocessors, integrated circuits, or programmable logic devices. Processors also include communication devices, such as computers, cell phones, portable / mobile personal digital assistants ("PDAs"), and other devices that facilitate the communication of information between end users.
[0098] The phrases "one example" or "one implementation" or "one implementation," as well as any other variations thereof, mean that the specific features, structures, properties, etc. described in relation to the example are included in at least one example. Therefore, the appearance of the phrases "one example" or "one implementation" or "one implementation," as well as any other variations thereof, in various places throughout this application, does not necessarily all refer to the same example.
[0099] Furthermore, this application may refer to "determining" various types of information. Determining information may include, for example, one or more of the following: estimating information, calculating information, predicting information, or retrieving information from memory. Acquiring may include receiving, retrieving, constructing, generating, and / or determining.
[0100] Furthermore, this application may refer to “accessing” various types of information. Accessing information may include, for example, receiving information, retrieving information (for example, from memory), storing information, moving information, copying information, computing information, determining information, predicting information, or estimating information.
[0101] Furthermore, this application may refer to "receiving" various types of information. Receiving is intended to be a broad term, similar to "accessing." Receiving information may include, for example, accessing information or retrieving information (for example, from memory). Moreover, "receiving" generally involves, in some way, being involved in an operation such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0102] For example, in the cases of "A / B", "A and / or B", and "at least one of A and B", it should be understood that the use of " / ", "and / or", and "at least one of" is intended to encompass the selection of only the first enumerated option (A), only the second enumerated option (B), or both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such phrasing is intended to encompass the selection of only the first enumerated option (A), only the second enumerated option (B), only the third enumerated option (C), only the first and second enumerated options (A and B), only the first and third enumerated options (A and C), only the second and third enumerated options (B and C), or all three options (A, B, and C). This can be extended to the same number of items as listed, as will be obvious to those skilled in the art.
[0103] Furthermore, the term "signal" as used herein specifically refers to indicating something to the corresponding decoder. Encoder signals can include encoding functions for an input, such as a precision factor, for example. In this way, in one example, the same parameters are used on both the encoder and decoder sides. Thus, for example, the encoder can send certain parameters to the decoder so that the decoder can use the same certain parameters (explicit signaling). Conversely, if the decoder already has certain parameters and others, signaling can be used without transmission, simply to allow the decoder to know and select the certain parameters (implicit signaling). Bit saving is achieved in various examples by avoiding the transmission of arbitrary actual functions. It should be understood that signaling can be performed in various ways. For example, in various examples, one or more syntax elements, flags, etc., are used to signal information to the corresponding decoder. The above concerns the verb form of the word "signal," but the word "signal" may also be used as a noun in this specification (for example, it may be used as a noun).
[0104] As will be apparent to those skilled in the art, implementations can produce a variety of signals formatted to carry information, which can be stored or transmitted. The information may include, for example, instructions for performing a method, or data produced by one of the implementations described. For example, a signal can be formatted to carry the bitstream of the example described. Such a signal may be formatted, for example, as an electromagnetic wave (for example, using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored in a processor-readable medium, or accessed or received from a processor-readable medium.
[0105] Many examples are described herein. The features of the examples may be provided individually or in any combination across various claim categories and types. Furthermore, the examples may include, individually or in any combination across various claim categories and types, one or more of the features, devices, or embodiments described herein. For example, the features described herein may be implemented in a bitstream or signal containing information generated as described herein. This information may enable a decoder to decode the bitstream, or an encoder, bitstream, and / or decoder according to any of the embodiments described. For example, the features described herein may be implemented by creating and / or transmitting and / or receiving and / or decoding a bitstream or signal. For example, the features described herein may be implemented as a method, process, apparatus, medium for storing instructions, medium for storing data, or signal. For example, the features described herein may be implemented by a TV, set-top box, cellphone, tablet, or other electronic device performing decoding. A TV, set-top box, cell phone, tablet, or other electronic device can display the resulting image (for example, an image from residual reconstruction of a video bitstream) (for example, using a monitor, screen, or other type of display). A TV, set-top box, cell phone, tablet, or other electronic device can receive a signal containing an encoded image and perform decoding.
[0106] These examples can be implemented by a device having at least one processor. The device may be an encoder or a decoder. These examples can be implemented by a computer program product, which includes program code instructions and is stored on a non-temporary computer-readable medium. These examples can be implemented by a computer program comprising program code instructions.
[0107] Figure 5 shows an example of discontinuity in the predicted block relative to the current block template when copying the predicted block in intrablock copy (IBC) or intratemplate matching prediction (intraTMP). The copied predicted block can be obtained by decoding the block vector via IBC or by inferring it via intraTMP. The copied predicted block may have distinctive properties that present some discontinuity when compared to the current block template (for example, as shown in Figure 5).
[0108] Discontinuities between predicted blocks and current blocks can lead to inappropriate predictions. This may be true for natural content compared to screen or gaming content, which can be characterized by sharp edges. A predicted signal that presents a natural flow of the scene, without abrupt changes between previous samples (e.g., samples associated with a template) and current samples (e.g., samples associated with a predicted block), may be desired. In examples, an intra-prediction process (e.g., planar, DC, angle prediction) may be followed by a filtering process (e.g., smoothing), such as a position-dependent intra-prediction combination (PDPC). Examples of filtering processes following IBC or IntraTMP predictions (e.g., as provided herein) can smooth the predictions and improve their quality, which can provide coding gain.
[0109] Figure 6 shows an example of an intra-template matching search area used in intraTMP. intraTMP is an intra-prediction mode in which the best predicted block can be copied from the reconstructed part of the current picture (e.g., the current frame) whose L-shaped template matches the current template. Within a predefined search range, the encoder searches the reconstructed part of the current picture for the template most similar to the current template and can use the corresponding block as the predicted block. The encoder can then signal the use of this mode. The same prediction operation can be performed on the decoder side.
[0110] The prediction signal can be generated in Figure 6 by matching the L-shaped causal neighbor of the current block with another block in a predefined search area, and includes the following: R1:Current CTU R2: Upper left CTU R3: Upper CTU R4: Left CTU
[0111] The absolute difference sum (SAD) can be used as the cost function. Within a region (for example, within each region), the decoder can search for the template with the minimum SAD for the current template and use its corresponding block as the prediction block. The dimensions of the region (SearchRange_w, SearchRange_h) can be set proportionally to the block dimensions (BlkW, BlkH) so that there is a fixed number of SAD comparisons per pixel. That is, SearchRange_w=a*BlkW SearchRange_h=a*BlkH Here, "a" can be a constant that controls the gain / complexity trade-off. For example, "a" may be equal to 5.
[0112] The intraTMP tool can be enabled for CUs with a width and height of 64 or less. This maximum CU size for intraTMP can be configurable. The intraTMP mode can be signaled at the CU level through a dedicated flag.
[0113] Figures 7A-7C show examples of IBC reference regions corresponding to current block prediction. In the examples, the encoder or decoder can use template matching in the IBC for both IBC merge mode and IBC advanced motion vector prediction (AMVP) mode. The intra-block copy-template matching (IBC-TM) merge list can be modified compared to the one used by the normal IBC merge mode (for example, so that candidates can be selected according to a pruning method with motion distances between candidates, as in the normal TM merge mode). Ending zero motion fulfillment can be replaced by motion vectors to the left (-W,0), up (0,-H), and top-left (-W,-H), where W can be the width of the current CU and H can be the height.
[0114] In IBC-TM merge mode, selected candidates can be refined using template matching before rate distortion optimization (RDO) or the decoding process. IBC-TM merge mode may conflict with normal IBC merge mode and TM merge flags, which can be signaled.
[0115] In IBC-TM AMVP mode, up to three candidates can be selected from the IBC-TM merge list. Each of these three selected candidates can be refined using template matching and sorted according to their resulting template matching costs. The first two (for example, only the first two) can be considered in the motion estimation process (for example, then the others).
[0116] For template matching refinement in both IBC-TM merge mode and AMVP mode, the IBC motion vector may be constrained to (i) be an integer and (ii) be within the reference region. In IBC-TM merge mode, refinement (e.g., all refinements) can be performed with integer precision. In IBC-TM AMVP mode, refinement (e.g., all refinements) can be performed with either integer or 4Pell precision, depending on the AMVR value. Such refinement accesses samples (e.g., samples only) without interpolation. The refined motion vector and the template used in parts of the refinement process (e.g., each part) may respect the constraints of the reference region (e.g., in both IBC-TM merge mode and IBC-TM AMVP mode).
[0117] Figure 8 shows an example of a reference area for IBC when a coded tree unit (CTU) (m,n) is coded. As shown in Figure 8, the (m,n) block can currently represent a CTU, the (m+1,n-1), (m+1,n-2), (m,n-1), (m,n-2), (m-1,n), (m-1,n-1), (m-1,n-2), (m-2,n), (m-2,n-1), (m-2,n-2) blocks can represent reference areas, and blank white blocks can represent invalid reference areas. The reference area for IBC can be extended up to two CTU rows above. Figure 8 shows the reference area for coding CTU(m,n). More specifically, for a CTU(m,n) to be encoded, the reference area may contain CTUs with indices (m-2,n-2)...(W,n-2), (0,n-1)...(W,n-1), (0,n)...(m,n), where W can indicate the highest horizontal index in the current tile, slice, or picture. If the CTU size is 256, the reference area may be limited to the CTU row one level up. This setting ensures that the IBC may not require extra memory if the CTU size is 128 or 256. The sample-by-sample block vector search (e.g., local search) range may be limited to [-(C<<1),C>>2] horizontally and [-C,C>>2] vertically with respect to the reference area expansion, where C can indicate the CTU size.
[0118] Figures 9A–9D show exemplary samples used by position-dependent intra-prediction combinations (PDPCs) applied to diagonal and adjacent angle intra-modes. The results of intra-predictions in planar modes can be modified (e.g., further modified) by PDPC. PDPC is an intra-prediction tool that can invoke combinations of intra-predictions with unfiltered boundary reference samples and filtered boundary reference samples. PDPC can be applied without signaling to at least one of the following intra-modes: planar, DC, horizontal, vertical, lower-left angle mode and its eight adjacent angle modes, or upper-right angle mode and its eight adjacent angle modes.
[0119] The predicted sample pred(x,y) can be predicted using a linear combination of the intra-prediction mode (e.g., DC, plane, angle, etc.) and the reference sample, according to the following formula: pred(x,y)=(wL×R -1,y +wT×R x,-1 -wTL×R -1,-1 +(64-wL-wT+wTL)×pred(x,y)+32)>>6 And here, R x,-1 , R -1,y These can represent reference samples located above and to the left of the current sample (x,y), respectively, in R -1,-1 This can represent a reference sample currently located in the upper left corner of the block.
[0120] When PDPC is applied to DC, planar, horizontal, and / or vertical intra modes, a boundary filter (e.g., an additional boundary filter) may not be required (e.g., this may be required in some DC mode boundary filters or horizontal / vertical mode edge filters). The PDPC processes for DC mode and planar mode can be the same, and the clipping operation can be avoided. In the case of the angular mode, the PDPC scale factor may not require range checking and can be adjusted such that the angle for enabling PDPC is removed (scale >= 0 is used). The PDPC weights can be (e.g., additionally) based on 32 in the case of the angular mode (e.g., all angular modes).
[0121] Reference samples (R x,-1 、R -1,y 、and R -1,-1 ) for PDPC applied across various prediction modes are defined in FIG. 8. The prediction sample pred(x’, y’) can be located at (x’, y’) within the prediction block. In the example, for the diagonal mode, the coordinate x of the reference sample R x,-1 can be given by x = x’ + y’ + 1, and the coordinate y of the reference sample R -1,y can similarly be given by y = x’ + y’ + 1. In the case of other annular modes, the reference samples R x,-1 and R -1,y can be located at fractional sample positions. (For example, in this case) the sample value at the nearest integer sample location can be used.
[0122] The PDPC weights may depend on the prediction mode. An example of PDPC weights by prediction mode is shown in Table 1 below.
[0123]
Table 1
[0124] The convolutional cross-component model (CCCM) can be applied to predict chroma samples from reconstructed lumen samples, in a similar manner to that performed by the current cross-component linear model (CCLM) mode. The reconstructed lumen samples can be downsampled to match lower-resolution chroma grids if chroma subsampling is used (e.g., as with CCLM). There may be options to use a single-model or multi-model variant of the CCCM (e.g., as with CCLM). The multi-model variant can use two models, one of which can be derived for samples above the mean lumen reference value, and the other model can be derived for the rest of the samples. The multi-model CCCM mode can be selected for PUs with at least 128 reference samples available.
[0125] Figure 10 shows an example of a spatial portion of a convolutional filter. A convolutional 7-tap filter can include a 5-tap plus sign-shape spatial component, a nonlinear term, and a bias term. The input to the filter's 5-tap spatial component can include a central (C) lumen sample that can be collated with the chroma sample to be predicted, as shown in Figure 10, and its upper / north (N), lower / south (S), left / west (W), and right / east (E) neighborhoods. The nonlinear term P can be expressed as a power of 2 of the central lumen sample C and can be scaled to the sample value range of the content, i.e., P = (C * C + midVal) >> bitDepth.
[0126] For 10-bit content, the content can be calculated as P=(C*C+512)>>10. The bias term B can represent a scalar offset between the input and output (similar to the offset term in CCLM, for example) and can be set to an intermediate chroma value (e.g., 512 for 10-bit content).
[0127] The filter output can be calculated as a convolution between the filter coefficients ci and the input values, and can be clipped to the range of valid chroma samples. The calculation can be as follows: predChromaVal=c0C+c1N+c2S+c3E+c4W+c5P+c6B
[0128] Figure 11 shows an example of a reference area used to derive the filter coefficients. The filter coefficients ci can be calculated by minimizing the mean squared error (MSE) between the predicted chroma samples and the reconstructed chroma samples in the reference area. As shown in Figure 11, the reference area can include chroma samples in the six lines above and to the left of the PU. The reference area can be extended to the right by one PU width and down by one PU height from the PU boundary. This area can be adjusted to include only available samples. The extension to the area shown in blue may be necessary to support side samples of a plus-shaped spatial filter and can be padded if they are in an unavailable area.
[0129] Minimizing the MSE can be performed by computing the autocorrelation matrix of the lumern input and the cross-correlation vector between the lumern input and the chroma output. The autocorrelation matrix can be decomposed into LDL, and the final filter coefficients can be computed using back substitution. This computation can follow the exemplary computation of adaptive loop filtering (ALF) filter coefficients, but LDL decomposition can be chosen (for example, instead of Cholesky decomposition) to avoid the use of square root operations.
[0130] Figure 12 shows an example of filtering a single line (samples P00 to P03 and P10 to P30). In the example, the encoder or decoder can obtain the predicted block of the current block in the current picture. The encoder or decoder can obtain multiple samples in the predicted block. The encoder or decoder can determine that the current block is encoded in IBC mode. Based on the current block being encoded in IBC mode, filter parameters can be determined based on multiple template samples of the predicted block and template samples of the current block. Multiple template samples of the predicted block can be filtered based on the determined filter parameters. In the example, a subset of multiple template samples of the predicted block can be filtered based on the determined filter parameters. The current block can be encoded or decoded based on the filtered samples. In the example, some filtered samples may include boundary samples (e.g., only boundary samples). Boundary samples can be the top and left samples in the predicted block (e.g., as shown by P00 to P03 and P10 to P30 in Figure 12). This is because boundary samples may represent the greatest discontinuity, and filtering these samples may have less computational complexity compared to complete filtering. Filtering can be one-dimensional (e.g., horizontal and vertical) or two-dimensional. In one dimension, the filtering parameter can be Fil = {2, 4, 2} / 8.
[0131] In the example, the filtered samples (a subset of filtered samples) can be the upper and left predicted samples of the prediction block. The upper predicted samples can be filtered using the upper template of the current block, and the left predicted samples can be filtered using the left template of the current block. The filtering parameters to be applied to filter subsets of multiple samples can be based on being a smoothing filter and being a multiple of 2 to reduce computational complexity. In the example, filter parameters such as {1,6,1} / 8 can be used. In 2D filtering, both the upper and left samples can be used to filter some or all of the samples in the prediction block. The filter parameters can be as follows: Fil={1,1,1 1,8,1 1,1,1} / 16 As explained above, the filtering parameters to be applied to filter several samples can be based on the fact that they are smoothing filters and are multiples of 2.
[0132] In the example, all of the multiple template samples in the prediction block can be filtered based on the determined filter parameters. This can allow for further smoothing of the prediction signal and ensure better insertion of the prediction signal.
[0133] PDPC can be used to smooth the prediction signal of IBC or IntraTMP. In the example, PDPC filtering can be applied if the prediction is considered a DC / planar prediction. In the example, PDPC filtering can derive an equivalent mode for IBC / IntraTMP and the corresponding PDPC process can be used. Decoder-side intra-mode derivation (DIMD) can be employed to derive the equivalent intra-mode. For MIP or intra-coded blocks, DIMD can be applied to the prediction signal to derive the intra-mode that is most similar to the current prediction. This mode can be used for conversion kernel selection. The same procedure can be used to derive the intra-mode corresponding to the prediction signal obtained from IBC or intraTMP. The PDPC process with that mode can be used to smooth the current prediction signal.
[0134] In the example, a spatial intra-prediction mode that yields minimum strain (SAD, Sum of Absolute Transformed Differences (SATD), HAD, etc.) can be used with predicted blocks acquired by IBC or intraTMP. A PDPC process corresponding to this mode can be used to filter the IBC or intraTMP predicted blocks.
[0135] In the example, filter parameters for filtering several samples can be learned from reconstructed samples (e.g., template samples of the prediction block). The filtering process can be defined only for intraTMP. That is, to determine the filter parameters, reconstructed lumen samples of the template area of the prediction block (e.g., the reference block) can be used as input to the filter during training, and the corresponding lumen samples in the current block's template can be the target. The filter parameters can be derived using regression-based MSE minimization techniques (e.g., LDL decomposition). In the example, a 6-tap filter can be used, taking spatial samples from the reconstructed reference block and using an intermediate value of the operating bit depth for bias (e.g., 512 for 10-bit content). The formula for each new prediction sample can be as follows: predVal=c0C+c1N+c2S+c3W+c4E+c5b
[0136] Figure 13 shows an example of the filter shape and training area for a reference block. Inputs to the filter's five spatial tap components can include (for example, as shown in Figure 13) the central (C) luma sample and its up / north (N), down / south (S), left / west (W), and right / east (E) neighbors. An example of an intraTMP SAD-based search is provided for creating a list of candidate prediction blocks (e.g., the reference block) to be filtered. Entries to the candidate prediction block (e.g., the reference block) list can be based on the SAD cost of the unfiltered template. Filter parameters can be calculated for each candidate prediction block (e.g., the reference block) in the list. The candidate prediction block (e.g., the reference block) with the best performance in terms of the filtered template cost is selected and can be used as the final candidate prediction block (e.g., the reference block). The final prediction for a block (e.g., the current block) can be generated (e.g., then generated) by applying the derived filter (e.g., filter parameters) for the best candidate prediction block (e.g., the reference block).
[0137] Figure 14 shows an example of calculating filter parameters between a reference template and a current template. This can be used for a current block encoded in IBC mode (e.g., IBC prediction). Based on the fact that the current block is encoded in IBC mode, the filter parameters can be determined based on the template samples of the prediction block and the template samples of the current block. The filter parameters can be calculated to minimize the difference between the current template and the reference template. The samples of the prediction block can be filtered based on the determined filter parameters. The current block can be encoded or decoded based on the filtered samples.
[0138] Figure 15 shows an example of calculating filter parameters between a reference template and a reference block. This allows for modification of the filtering learning process so that the optimal filter can be found between the reference template and the reference block (for example, not between the reference template and the current template). This example can minimize the distance between a given template and its corresponding block. Filtering can be performed between the current template and the current prediction mode. The example provided in Figure 15 can be applied to both IBC prediction and intraTMP prediction.
[0139] Figure 16 shows an example of calculating the filter parameters between the current template and the reference block. In the example provided in Figure 16, the learning process can find the optimal filter between the current template and the reference block used for IBC prediction.
[0140] In the example, the encoder can adaptively select whether or not to use the filtering process. That is, the encoder may have the option to switch the filtering process on or off. This switching can be at a higher or lower level. A higher level of switching (e.g., SPS, PH, SH) may allow the encoder to completely switch off the process for sequences that do not require smooth prediction (e.g., screen content or gaming content). A lower level of switching (e.g., CU, PU, TU) may give the encoder flexibility in deciding whether or not to perform filtering depending on block characteristics and / or RD criteria (e.g., signaling at a lower level may be required).
[0141] While features and elements have been described above in specific combinations, those skilled in the art will understand that each feature or element may be used alone or in any combination with other features and elements. Furthermore, the methods described herein may be implemented in computer programs, software, or firmware embedded in computer-readable media for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted via wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital multipurpose disks (DVDs). A software-related processor may be used to implement a radio frequency transceiver for use in a WTRU, UE, terminal, base station, RNC, or any host computer.
Claims
1. A device for video decoding, It is determined that the block is currently encoded in intra-block copy mode. Based on the fact that the current block is encoded in the intra-block copy mode, the filter parameters are determined based on the template sample of the prediction block and the template sample of the current block. Based on the determined filter parameters, the sample of the prediction block is filtered, Based on the filtered sample, decode the current block. Processor configured in such a way A device equipped with the following features.
2. The device according to claim 1, wherein the filtered subset of the prediction block is filtered based on the determined filter parameters.
3. The subset of the filtered samples is the upper and left predicted samples of the prediction block, The template sample of the current block is the top template and left template of the current block, The device according to claim 2, wherein the upper prediction sample is filtered using the upper template of the current block, and the left prediction sample is filtered using the left template of the current block.
4. The device according to claim 1, wherein at least one of the filter parameters is a smoothing filter and a multiple of 2.
5. The device according to claim 1, wherein the filter parameters are further determined using regression based on a mean squared error (MSE) minimization technique.
6. The device according to claim 1, wherein the prediction block is the best candidate prediction block from a list of candidate prediction blocks, and the best candidate prediction block is determined based on the template cost.
7. A method for video decoding, The steps include determining that the block is currently encoded in intra-block copy mode, Based on the fact that the current block is encoded in the intra-block copy mode, the filter parameters are determined based on the template sample of the prediction block and the template sample of the current block. The steps include filtering the sample of the prediction block based on the determined filter parameters, The steps include: decoding the current block based on the filtered sample; A method for providing this.
8. The method according to claim 7, wherein the filtered subset of the prediction block is filtered based on the determined filter parameters.
9. The subset of the filtered samples is the upper and left predicted samples of the prediction block, The template sample of the current block is the top template and left template of the current block, The method according to claim 8, wherein the upper prediction sample is filtered using the upper template of the current block, and the left prediction sample is filtered using the left template of the current block.
10. The method according to claim 7, wherein at least one of the filter parameters is a smoothing filter and a multiple of 2.
11. The method according to claim 7, wherein the filter parameters are further determined using regression based on a mean squared error (MSE) minimization technique.
12. The method according to claim 7, wherein the prediction block is the best candidate prediction block from a list of candidate prediction blocks, and the best candidate prediction block is determined based on the template cost.
13. A device for video encoding, It is determined that the block is currently encoded in intra-block copy mode. Based on the fact that the current block is encoded in the intra-block copy mode, the filter parameters are determined based on the template sample of the prediction block and the template sample of the current block. Based on the determined filter parameters, the sample of the prediction block is filtered, Based on the filtered sample, encode the current block. A processor configured in such a way A device equipped with the following features.
14. The device according to claim 13, wherein the filtered subset of the prediction block is filtered based on the determined filter parameters.
15. The subset of the filtered samples is the upper and left predicted samples of the prediction block, The template sample of the current block is the top template and left template of the current block, The device according to claim 14, wherein the upper prediction sample is filtered using the upper template of the current block, and the left prediction sample is filtered using the left template of the current block.
16. The device according to claim 13, wherein at least one of the filter parameters is a smoothing filter and a multiple of 2.
17. The device according to claim 13, wherein the filter parameters are further determined using regression based on a mean squared error (MSE) minimization technique.
18. The device according to claim 13, wherein the prediction block is the best candidate prediction block from a list of candidate prediction blocks, and the best candidate prediction block is determined based on the template cost.
19. A method for video encoding, The steps include determining that the block is currently encoded in intra-block copy mode, Based on the fact that the current block is encoded in the intra-block copy mode, the filter parameters are determined based on the template sample of the prediction block and the template sample of the current block. The steps include filtering the sample of the prediction block based on the determined filter parameters, The steps include encoding the current block based on the filtered sample and A method for providing this.
20. The method according to claim 19, wherein the filtered subset of the prediction block is filtered based on the determined filter parameters.
21. The subset of the filtered samples is the upper and left predicted samples of the prediction block, The template sample of the current block is the top template and left template of the current block, The method according to claim 20, wherein the upper prediction sample is filtered using the upper template of the current block, and the left prediction sample is filtered using the left template of the current block.
22. The method according to claim 19, wherein at least one of the filter parameters is a smoothing filter and a multiple of 2.
23. The method according to claim 19, wherein the filter parameters are further determined using regression based on a mean squared error (MSE) minimization technique.
24. The method according to claim 19, wherein the prediction block is the best candidate prediction block from a list of candidate prediction blocks, and the best candidate prediction block is determined based on the template cost.
25. A computer program product stored in a non-temporary computer-readable medium, comprising, when executed by a processor, program code instructions for implementing the steps of any one of claims 7 to 12 and 19 to 24.
26. A computer-readable medium comprising program code instructions for implementing a step of the method according to any one of claims 7 to 12 and 19 to 24, when executed by a processor.
27. Video data including information representing an encoded output generated according to one of the methods described in any one of claims 19 to 24.