Combination of Intra-frame Prediction Based on Extrapolation Filter with CIIP and GPM Modes
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2026-08-11
Smart Images

Figure CN122556069A_ABST
Abstract
Description
[0001] Cross-references to related applications This application claims the benefit of European Patent Application No. 24305056.4, filed on January 9, 2024, the disclosure of which is incorporated herein by reference in its entirety. Background Technology
[0002] Video codec systems can be used to compress digital video signals, for example, by reducing the storage and / or transmission bandwidth required for such signals. Video codec systems can include, for example, block-based, wavelet-based, and / or object-based systems. Summary of the Invention
[0003] Systems, methods, and means for performing extrapolation filter-based intra-frame prediction (EIP) by combining other prediction modes are disclosed. For example, EIP can be combined with one or more of combined inter-frame intra-frame prediction (CIIP), geometric segmentation mode (GPM), spatial geometric segmentation mode (SGPM), or combined intra-frame block copy geometric segmentation mode (IBC-GPM). Predictions obtained from EIP can be examined by template analysis and / or mixed according to the selected mode.
[0004] A video decoding / encoding device may include a processor. The video decoding / encoding device may obtain a first prediction associated with the current block. The first prediction may be obtained based on an extrapolation filter. The device may obtain a second prediction associated with the current block. The device may generate a mixed prediction based on the first and / or second predictions. The device may decode / encode the current block based on the mixed predictions.
[0005] The device may include one or more features. A second prediction may be obtained based on an inter-frame prediction mode. A second prediction may be obtained based on an intra-frame prediction mode. The device may obtain an extrapolation filter based on a reconstruction template associated with the current block. The second prediction may be at least one of the following: an inter-frame prediction signal obtained based on a merge candidate list, an intra-frame prediction signal obtained based on an intra-frame prediction mode (IPM) list and / or an IPM index of signaling, or an intra-frame prediction signal obtained based on a template-based intra-frame mode derivation (TIMD) result. The device may determine adjacent reconstruction regions of the current block and / or the filter shape of the current block. The device may determine extrapolation filter coefficients associated with the current block based on adjacent reconstruction regions and / or filter shapes. A second prediction associated with the current block may be obtained based on the extrapolation filter coefficients.
[0006] Video decoding / encoding methods may include obtaining a first prediction associated with the current block. The first prediction may be obtained based on an extrapolation filter. The method may include obtaining a second prediction associated with the current block. The method may include generating a hybrid prediction based on the first and / or second predictions. The method may include decoding / encoding the current block based on the hybrid predictions.
[0007] This method may include one or more features. A second prediction may be obtained based on an inter-frame prediction mode. A second prediction may be obtained based on an intra-frame prediction mode. This method may include obtaining an extrapolation filter based on a reconstruction template associated with the current block. The second prediction may be at least one of the following: an inter-frame prediction signal obtained based on a merge candidate list, an intra-frame prediction signal obtained based on an intra-frame prediction mode (IPM) list and / or an IPM index of signaling, or an intra-frame prediction signal obtained based on a template-based intra-frame mode derivation (TIMD) result. This method may include determining adjacent reconstruction regions of the current block and / or the filter shape of the current block. This method may include determining extrapolation filter coefficients associated with the current block based on adjacent reconstruction regions and / or filter shapes. A second prediction associated with the current block may be obtained based on the extrapolation filter coefficients.
[0008] Computer program products are stored on non-transitory computer-readable media and may include program code instructions for implementing the steps of the methods described herein when executed by a processor.
[0009] The video decoding device can obtain a first prediction signal based on an extrapolation filter and a second prediction signal. The video decoding device can generate a hybrid prediction based on the obtained prediction signal and decode the current block based on the hybrid prediction.
[0010] The second prediction signal may depend on the prediction mode that the EIP can be combined with. For example, the second prediction signal may be an inter-frame prediction signal based on a merge candidate list (e.g., when combined with CIIP). For example, the second prediction signal may be an intra-frame prediction signal based on an IPM list and an IPM index of signaling (e.g., when used as an EIP-CIIP mode). The second prediction signal may be an intra-frame prediction signal based on the result of template-based intra-frame mode derivation (TIMD).
[0011] The video decoding device can determine the type and filter shape of the reconstructed region of the current block based on the video bitstream. The video decoding device can then determine the extrapolated filter coefficients associated with the current block based on the reconstructed region and filter shape. Finally, the video decoding device can determine the predicted sample of the current block based on the extrapolated filter coefficients.
[0012] The video coding device can obtain a first prediction signal based on an extrapolation filter and a second prediction signal. The video coding device can generate a hybrid prediction based on the obtained prediction signal and encode the current block based on the hybrid prediction.
[0013] The systems, methods, and means described herein may relate to a decoder. In some examples, the systems, methods, and means described herein may relate to an encoder. In some examples, the systems, methods, and means described herein may relate to a signal (e.g., from an encoder and / or received by a decoder). A computer-readable medium may include instructions for causing one or more processors to perform the methods described herein. A computer program product may include instructions that, when executed by one or more processors, cause one or more processors to perform the methods described herein. Attached Figure Description
[0014] Figure 1A This is a system diagram illustrating an example communication system in which one or more of the disclosed embodiments may be implemented.
[0015] Figure 1B This illustrates a method according to one embodiment. Figure 1A The system diagram shown is of an example wireless transmit / receive unit (WTRU) used in the communication system.
[0016] Figure 1C This illustrates a method according to one embodiment. Figure 1A The system diagram shows an example radio access network (RAN) and an example core network (CN) used in the communication system shown.
[0017] Figure 1D This illustrates a method according to one embodiment. Figure 1A The system diagram shows another example RAN and another example CN used in the communication system shown.
[0018] Figure 2 An example video encoder is shown.
[0019] Figure 3 An example video decoder is shown.
[0020] Figure 4 Examples of systems in which various aspects and examples can be implemented are shown.
[0021] Figure 5 The example intra-frame template matching search region used is described.
[0022] Figures 6A-6B An example partitioning method for angular patterns is described.
[0023] Figures 7A-7C Example geometric segmentation modes (GPM) with inter-frame and intra-frame predictions are described.
[0024] Figure 7D An example GPM with intra-frame and intra-frame prediction is described.
[0025] Figure 8 The example space GPM candidates are described.
[0026] Figure 9 A sample GPM template is described.
[0027] Figure 10 An example of an extrapolation-based intra-prediction (EIP) filter shape is described.
[0028] Figure 11 This describes an example definition type for the reconstructed region.
[0029] Figure 12 An example is described that generates predictions for different predictions in the current block in diagonal order.
[0030] Figure 13 An example method for combining EIP with other prediction models is described. Detailed Implementation
[0031] A more detailed understanding can be obtained by combining the accompanying drawings with the following description, which is given by way of example.
[0032] Figure 1A This diagram illustrates an example communication system 100 in which one or more of the disclosed embodiments may be implemented. The communication system 100 may be a multiple access system that provides content such as voice, data, video, messaging, and broadcasting to multiple wireless users. The communication system 100 enables multiple wireless users to access this content by sharing system resources, including wireless bandwidth. For example, the communication system 100 may employ one or more channel access methods, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal FDMA (OFDMA), Single Carrier FDMA (SC-FDMA), Zero Tail Unique Word DFT Spread Spectrum OFDM (ZT UWDTS-s OFDM), Unique Word OFDM (UW-OFDM), Resource Block Filtered OFDM, Filter Bank Multicarrier (FBMC), etc.
[0033] like Figure 1AAs shown, the communication system 100 may include wireless transceiver units (WTRUs) 102a, 102b, 102c, 102d, RAN 104 / 113, CN 106 / 115, Public Switched Telephone Network (PSTN) 108, Internet 110, and other networks 112. However, it should be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, and 102d may be any type of device configured to operate and / or communicate in a wireless environment. For example, WTRUs 102a, 102b, 102c, and 102d (any of which may be referred to as a “station” and / or “STA”) may be configured to transmit and / or receive wireless signals and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearable devices, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in the context of industrial and / or automated processing chains), consumer electronics devices, devices operating on commercial and / or industrial wireless networks, etc. Any of WTRUs 102a, 102b, 102c, and 102d may be interchangeably referred to as a UE.
[0034] The communication system 100 may also include base station 114a and / or base station 114b. Each of base stations 114a and 114b may be any type of device configured to wirelessly interface with at least one of WTRUs 102a, 102b, 102c, and 102d to facilitate access to one or more communication networks (e.g., CN 106 / 115, Internet 110, and / or other networks 112). For example, base stations 114a and 114b may be base transceiver stations (BTS), node Bs, eNodeBs, home node Bs, home eNodeBs, gNBs, NRNodeBs, site controllers, access points (APs), wireless routers, etc. Although base stations 114a and 114b are each described as a single element, it should be understood that base stations 114a and 114b may include any number of interconnected base station and / or network elements.
[0035] Base station 114a may be part of RAN 104 / 113, and may also include other base stations and / or network elements (not shown), such as base station controllers (BSCs), radio network controllers (RNCs), relay nodes, etc. Base station 114a and / or base station 114b may be configured to transmit and / or receive radio signals on one or more carrier frequencies, which may be referred to as cells (not shown). These frequencies may be in licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide coverage of a specific geographic area, which may be relatively fixed or may change over time. A cell may also be divided into cell sectors. For example, the cell associated with base station 114a may be divided into three sectors. Thus, in one embodiment, base station 114a may include three transceivers, i.e., one transceiver per sector of the cell. In one embodiment, base station 114a may employ multiple-input multiple-output (MIMO) technology and may utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.
[0036] Base stations 114a and 114b can communicate with one or more of WTRUs 102a, 102b, 102c, and 102d via air interface 116. Air interface 116 can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). Any suitable radio access technology (RAT) can be used to establish air interface 116.
[0037] More specifically, as described above, the communication system 100 can be a multiple access system and can employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, base stations 114a and WTRUs 102a, 102b, and 102c in RAN 104 / 113 can implement radio technologies such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which can establish air interfaces 115 / 116 / 117 using Wideband CDMA (WCDMA). WCDMA can include communication protocols such as High-Speed Packet Access (HSPA) and / or evolved HSPA (HSPA+). HSPA can include High-Speed Downlink (DL) Packet Access (HSDPA) and / or High-Speed UL Packet Access (HSUPA).
[0038] In one embodiment, base station 114a and WTRUs 102a, 102b, 102c can implement radio technologies such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which can use Long Term Evolution (LTE) and / or Advanced LTE (LTE-A) and / or Advanced LTE Pro (LTE-A Pro) to establish air interface 116.
[0039] In one embodiment, base station 114a and WTRUs 102a, 102b, 102c can implement radio technologies such as NR radio access, which can establish an air interface 116 using a new radio (NR).
[0040] In one embodiment, base station 114a and WTRUs 102a, 102b, and 102c can implement multiple radio access technologies. For example, base station 114a and WTRUs 102a, 102b, and 102c can jointly implement LTE radio access and NR radio access, for example, using the dual connectivity (DC) principle. Therefore, the air interface used by WTRUs 102a, 102b, and 102c can be characterized by multiple types of radio access technologies and / or transmissions sent to / from multiple types of base stations (e.g., eNBs and gNBs).
[0041] In other embodiments, base station 114a and WTRUs 102a, 102b, and 102c can implement radio technologies such as IEEE 802.11 (i.e., Wi-Fi), IEEE 802.16 (i.e., WiMAX), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Provisional Standard 2000 (IS-2000), Provisional Standard 95 (IS-95), Provisional Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rate GSM Evolution (EDGE), and GSM EDGE (GERAN).
[0042] For example, Figure 1ABase station 114b can be a wireless router, home node B, home eNodeB, or access point, and can utilize any suitable RAT to facilitate wireless connectivity in a local area, such as commercial locations, homes, vehicles, campuses, industrial facilities, air corridors (e.g., for drone use), roads, etc. In one embodiment, base station 114b and WTRUs 102c, 102d can implement radio technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In one embodiment, base station 114b and WTRUs 102c, 102d can implement radio technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, base station 114b and WTRUs 102c, 102d can utilize cellular-based RATs (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish picocells or femtocells. Figure 1A As shown, base station 114b can be directly connected to Internet 110. Therefore, it is not required that base station 114b access Internet 110 via CN 106 / 115.
[0043] RAN 104 / 113 can communicate with CN 106 / 115, which can be any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRUs 102a, 102b, 102c, and 102d. Data can have different Quality of Service (QoS) requirements, such as different throughput requirements, latency requirements, fault tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, etc. CN 106 / 115 can provide call control, billing services, location-based services, prepaid calling, internet connectivity, video distribution, and / or perform advanced security functions such as user authentication. Although in Figure 1A Although not shown, it should be understood that RAN 104 / 113 and / or CN 106 / 115 can communicate directly or indirectly with other RANs using the same RAT as RAN 104 / 113 or a different RAT. For example, in addition to connecting to RAN 104 / 113, which may utilize NR radio technology, CN 106 / 115 can also communicate with another RAN (not shown) using GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.
[0044] CN 106 / 115 can also serve as a gateway for WTRU 102a, 102b, 102c, 102d to access PSTN 108, the Internet 110, and / or other networks 112. PSTN 108 may include a circuit-switched telephone network providing Common Old-Style Telephone Service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices using common communication protocols such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) from the TCP / IP Internet Protocol suite. Network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, network 112 may include another CN connected to one or more RANs, which may use the same RAT as RAN 104 / 113 or a different RAT.
[0045] Some or all of the WTRUs 102a, 102b, 102c, and 102d in the communication system 100 may include multi-mode capabilities (e.g., WTRUs 102a, 102b, 102c, and 102d may include multiple transceivers for communicating with different wireless networks via different wireless links). For example... Figure 1A The WTRU 102c shown can be configured to communicate with a base station 114a that can employ cellular-based radio technology and with a base station 114b that can employ IEEE 802 radio technology.
[0046] Figure 1B This is a system diagram illustrating example WTRU 102. (See diagram below.) Figure 1B As shown, WTRU 102 may include a processor 118, a transceiver 120, a transmitting / receiving element 122, a speaker / microphone 124, a keyboard 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power supply 134, a Global Positioning System (GPS) chipset 136, and / or other peripheral devices 138, etc. It should be understood that WTRU 102 may include any sub-combination of the foregoing elements while remaining consistent with the embodiments.
[0047] Processor 118 may be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. Processor 118 may perform signal encoding / decoding, data processing, power control, input / output processing, and / or any other functions that enable WTRU 102 to operate in a wireless environment. Processor 118 may be coupled to transceiver 120, and transceiver 120 may be coupled to transmitting / receiving element 122. Although Figure 1B While processor 118 and transceiver 120 are described as separate components, it should be understood that processor 118 and transceiver 120 may be integrated together in an electronic package or chip.
[0048] Transmitting / receiving element 122 can be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) via air interface 116. For example, in one embodiment, transmitting / receiving element 122 can be an antenna configured to transmit and / or receive RF signals. In one embodiment, transmitting / receiving element 122 can be, for example, a transmitter / detector configured to transmit and / or receive IR, UV, or visible light signals. In yet another embodiment, transmitting / receiving element 122 can be configured to transmit and / or receive both RF and optical signals. It should be understood that transmitting / receiving element 122 can be configured to transmit and / or receive any combination of wireless signals.
[0049] Although the transmitting / receiving element 122 is in Figure 1B While described as a single element, WTRU 102 may include any number of transmitting / receiving elements 122. More specifically, WTRU 102 may employ MIMO technology. Thus, in one embodiment, WTRU 102 may include two or more transmitting / receiving elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals via air interface 116.
[0050] Transceiver 120 can be configured to modulate signals to be transmitted by transmitting / receiving element 122 and demodulate signals received by transmitting / receiving element 122. As described above, WTRU 102 can have multi-mode capability. Therefore, for example, transceiver 120 may include multiple transceivers to enable WTRU 102 to communicate via multiple RATs, such as NR and IEEE 802.11.
[0051] The processor 118 of WTRU 102 can be coupled to a speaker / microphone 124, a keyboard 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) unit or an organic light-emitting diode (OLED) display unit) and can receive user input data therefrom. The processor 118 can also output user data to the speaker / microphone 124, keyboard 126, and / or display / touchpad 128. Furthermore, the processor 118 can access and store information from any type of suitable memory (e.g., non-removable memory 130 and / or removable memory 132). Non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. Removable memory 132 may include a subscriber identity module (SIM) card, a memory stick, a secure digital storage (SD) card, etc. In other embodiments, the processor 118 can access and store information from memory that is not physically located on WTRU 102 (e.g., on a server or home computer (not shown)).
[0052] The processor 118 can receive power from the power supply 134 and can be configured to distribute and / or control power to other components in the WTRU 102. The power supply 134 can be any suitable device for powering the WTRU 102. For example, the power supply 134 may include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.
[0053] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) about the current location of the WTRU 102. In addition to, or instead of, information from the GPS chipset 136, the WTRU 102 may receive location information from base stations (e.g., base stations 114a, 114b) via air interface 116, and / or determine its location based on the timing of signals received from two or more nearby base stations. It should be understood that the WTRU 102 may acquire location information using any suitable location determination method while remaining consistent with the embodiments.
[0054] The processor 118 may be further coupled to other peripheral devices 138, which may include one or more software and / or hardware modules providing additional features, functions, and / or wired or wireless connectivity. For example, peripheral devices 138 may include accelerometers, electronic compasses, satellite transceivers, digital cameras (for photos and / or video), Universal Serial Bus (USB) ports, vibration devices, television transceivers, hands-free headsets, Bluetooth® modules, FM radio units, digital music players, media players, video game player modules, internet browsers, virtual reality and / or augmented reality (VR / AR) devices, activity trackers, etc. Peripheral devices 138 may include one or more sensors, which may be gyroscopes, accelerometers, Hall effect sensors, magnetometers, orientation sensors, proximity sensors, temperature sensors, time sensors; geolocation sensors; altimeters, light sensors, touch sensors, magnetometers, barometers, attitude sensors, biosensors, and / or humidity sensors.
[0055] WTRU 102 may include a full-duplex radio for which the transmission and reception of some or all signals (e.g., associated with a specific subframe of both UL (e.g., for transmission) and downlink (e.g., for reception)) may be concurrent and / or simultaneous. The full-duplex radio may include an interference management unit to reduce and / or substantially eliminate self-interference via hardware (e.g., a choke) or via signal processing (e.g., a separate processor (not shown) or via processor 118). In one embodiment, WTRU 102 may include a half-duplex radio for which the transmission and reception of some or all signals (e.g., associated with a specific subframe of either UL (e.g., for transmission) or downlink (e.g., for reception) may be concurrent and / or simultaneous.
[0056] Figure 1C This is a system diagram illustrating RAN 104 and CN 106 according to one embodiment. As described above, RAN 104 can communicate with WTRUs 102a, 102b, and 102c via air interface 116 using E-UTRA radio technology. RAN 104 can also communicate with CN 106.
[0057] RAN 104 may include eNode-B 160a, 160b, 160c; however, it should be understood that RAN 104 may include any number of eNode-Bs while remaining consistent with the embodiments. eNode-B 160a, 160b, 160c may each include one or more transceivers for communicating with WTRU 102a, 102b, 102c via air interface 116. In one embodiment, eNode-B 160a, 160b, 160c may implement MIMO technology. Therefore, for example, eNode-B 160a may use multiple antennas to transmit and / or receive radio signals from WTRU 102a.
[0058] Each of the eNode-B 160a, 160b, and 160c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, and user scheduling in the UL and / or DL, etc. Figure 1C As shown, eNode-B 160a, 160b, and 160c can communicate with each other via the X2 interface.
[0059] Figure 1C The CN 106 shown may include a Mobility Management Entity (MME) 162, a Serving Gateway (SGW) 164, and a Packet Data Network (PDN) Gateway (or PGW) 166. While each of the foregoing elements is described as part of CN 106, it should be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.
[0060] The MME 162 can connect to each eNode-B 162a, 162b, 162c in RAN 104 via the S1 interface and can be used as a control node. For example, the MME 162 can be responsible for authenticating users of WTRUs 102a, 102b, 102c, bearer activation / deactivation, selecting a specific serving gateway during the initial attachment of WTRUs 102a, 102b, 102c, etc. The MME 162 can provide control plane functions for handover between RAN 104 and other RANs (not shown) employing other radio technologies (such as GSM and / or WCDMA).
[0061] The SGW 164 can connect to each eNode B 160a, 160b, or 160c in RAN 104 via the S1 interface. The SGW 164 can typically route and forward user data packets to / from WTRUs 102a, 102b, or 102c. The SGW 164 can perform other functions, such as anchoring the user plane during inter-eNode B handover; triggering paging when DL data is available for WTRUs 102a, 102b, or 102c; and managing and storing the context of WTRUs 102a, 102b, or 102c.
[0062] The SGW 164 can connect to the PGW 166, which can provide WTRU 102a, 102b, and 102c with access to packet-switched networks such as Internet 110, to facilitate communication between WTRU 102a, 102b, 102c and IP-enabled devices.
[0063] CN 106 can facilitate communication with other networks. For example, CN 106 can provide WTRU 102a, 102b, and 102c with access to a circuit-switched network such as PSTN 108 to facilitate communication between WTRU 102a, 102b, and 102c and traditional landline communication equipment. For example, CN 106 may include, or be able to communicate with, an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between CN 106 and PSTN 108. Furthermore, CN 106 can provide WTRU 102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.
[0064] Despite WTRU in Figure 1A-1D While described as a wireless terminal, it is conceivable that, in some representative embodiments, such a terminal may use (e.g., temporarily or permanently) a wired communication interface with a communication network.
[0065] In a representative embodiment, the other network 112 may be a WLAN.
[0066] A WLAN in Infrastructure Basic Services Set (BSS) mode can have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP can access or peer into a distributed system (DS) or another type of wired / wireless network that carries traffic into and / or out of the BSS. Traffic originating outside the BSS destined for a STA can be delivered to the AP via it. Traffic originating from a STA destined for a destination outside the BSS can be sent to the AP for delivery to the appropriate destination. For example, traffic between STAs within the BSS can be sent via the AP, where the source STA can send traffic to the AP, and the AP can deliver traffic to the destination STA. Traffic between STAs within the BSS can be considered and / or referred to as peering traffic. Peering traffic can be sent between source and destination STAs (e.g., directly between them) using Direct Link Establishment (DLS). In some representative embodiments, the DLS can use 802.11e DLS or 802.11z Tunneled DLS (TDLS). A WLAN using the Standalone BSS (IBSS) mode may not have an access point (AP), and STAs within the IBSS or using the IBSS (e.g., all STAs) can communicate directly with each other. The IBSS communication mode is sometimes referred to as the "self-organizing" communication mode in this document.
[0067] When operating in 802.11ac infrastructure mode or a similar mode, the AP can transmit beacons on a fixed channel (e.g., the primary channel). The primary channel can be of a fixed width (e.g., a 20 MHz bandwidth) or dynamically set via signaling. The primary channel can be the operating channel of the BSS and can be used by the STA to establish a connection with the AP. In some representative embodiments, Carrier Sense Multiple Access with Collision Avoidance (CSMA / CA) can be implemented, for example in an 802.11 system. For CSMA / CA, each STA, including the AP, can sense the primary channel. If a particular STA senses / detects and / or determines that the primary channel is busy, that particular STA can back off. A single STA (e.g., only one station) can transmit at any given time within a given BSS.
[0068] High-throughput (HT) STAs can communicate using a 40MHz wide channel, for example, by combining a primary 20MHz channel with adjacent or non-adjacent 20MHz channels.
[0069] Very High Throughput (VHT) STAs can support channels with widths of 20MHz, 40MHz, 80MHz, and / or 160MHz. 40MHz and / or 80MHz channels can be formed by combining consecutive 20MHz channels. A 160MHz channel can be formed by combining eight consecutive 20MHz channels, or by combining two non-consecutive 80MHz channels, which can be referred to as an 80+80 configuration. For the 80+80 configuration, after channel coding, the data passes through a segment resolver, which splits the data into two streams. Each stream can be processed separately using Inverse Fast Fourier Transform (IFFT) and time-domain processing. These streams can be mapped onto the two 80MHz channels, and the data can be transmitted by the transmitting STA. At the receiver of the receiving STA, the operation of the 80+80 configuration described above can be reversed, and the combined data can be sent to the Media Access Control (MAC).
[0070] 802.11af and 802.11ah support operating modes below 1 GHz. The channel operating bandwidth and carrier in 802.11af and 802.11ah are reduced compared to those used in 802.11n and 802.11ac. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV Blank (TVWS) spectrum, while 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz using non-TVWS. According to a representative embodiment, 802.11ah can support metering-type control / machine-type communications, such as MTC devices in macro coverage areas. MTC devices may have certain capabilities, such as limited capabilities, including support (e.g., only support) certain and / or limited bandwidths. MTC devices may include batteries with a battery life exceeding a threshold (e.g., to maintain very long battery life).
[0071] WLAN systems that can support multiple channels and channel bandwidths (e.g., 802.11n, 802.11ac, 802.11af, and 802.11ah) include a channel that can be designated as the primary channel. The bandwidth of the primary channel can be equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel can be set and / or limited by one of the STAs operating in the BSS that supports the minimum bandwidth operating mode. In the 802.11ah example, for STAs that support (e.g., only support) the 1MHz mode (e.g., MTC type devices), the primary channel can be 1MHz wide, even if the AP and other STAs in the BSS support 2MHz, 4MHz, 8MHz, 16MHz, and / or other channel bandwidth operating modes. Carrier Sense and / or Network Allocation Vector (NAV) settings can depend on the status of the primary channel. If the primary channel is busy, for example due to STAs (only supporting the 1MHz operating mode) sending to the AP, the entire available band can be considered busy, even if most of the band remains idle and may be available.
[0072] In the United States, the available frequency band for 802.11ah is from 902MHz to 928MHz. In South Korea, the available frequency band is from 917.5MHz to 923.5MHz. In Japan, the available frequency band is from 916.5MHz to 927.5MHz. The total available bandwidth for 802.11ah is 6MHz to 26MHz, depending on the country code.
[0073] Figure 1D This is a system diagram illustrating RAN 113 and CN 115 according to one embodiment. As described above, RAN 113 can communicate with WTRUs 102a, 102b, and 102c via air interface 116 using NR radio technology. RAN 113 can also communicate with CN 115.
[0074] RAN 113 may include gNBs 180a, 180b, and 180c; however, it should be understood that RAN 113 may include any number of gNBs while remaining consistent with the embodiments. gNBs 180a, 180b, and 180c may each include one or more transceivers for communicating with WTRUs 102a, 102b, and 102c via air interface 116. In one embodiment, gNBs 180a, 180b, and 180c may implement MIMO technology. For example, gNBs 180a and 180b may utilize beamforming to transmit signals to and / or receive signals from gNBs 180a, 180b, and 180c. Therefore, for example, gNB 180a may use multiple antennas to transmit and / or receive radio signals from WTRU 102a. In one embodiment, gNBs 180a, 180b, and 180c can implement carrier aggregation technology. For example, gNB 180a can transmit multiple component carriers (not shown) to WTRU 102a. A subset of these component carriers may be on unlicensed spectrum, while the remaining component carriers may be on licensed spectrum. In one embodiment, gNBs 180a, 180b, and 180c can implement Coordinated Multipoint (CoMP) technology. For example, WTRU 102a can receive coordinated transmissions from gNBs 180a and 180b (and / or gNB 180c).
[0075] WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using transmissions associated with scalable digitization. For example, OFDM symbol spacing and / or OFDM subcarrier spacing can vary for different transmissions, different cells, and / or different portions of the radio transmission spectrum. WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using subframes or transmission time intervals (TTIs) of various or scalable lengths (e.g., containing different numbers of OFDM symbols and / or continuously varying lengths of absolute time).
[0076] gNBs 180a, 180b, and 180c can be configured to communicate with WTRUs 102a, 102b, and 102c in standalone and / or non-standalone configurations. In standalone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c without simultaneously accessing other RANs (e.g., eNode-Bs 160a, 160b, and 160c). In standalone configuration, WTRUs 102a, 102b, and 102c can utilize one or more gNBs 180a, 180b, and 180c as mobility anchors. In standalone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using signals in unlicensed frequency bands. In a non-standalone configuration, WTRUs 102a, 102b, and 102c can communicate / connect with gNBs 180a, 180b, and 180c, while also communicating / connecting with another RAN such as eNode-Bs 160a, 160b, and 160c. For example, WTRUs 102a, 102b, and 102c can implement DC principles to communicate substantially simultaneously with one or more gNBs 180a, 180b, and 180c, as well as one or more eNode-Bs 160a, 160b, and 160c. In a non-standalone configuration, eNode-Bs 160a, 160b, and 160c can be used as mobility anchors for WTRUs 102a, 102b, and 102c, and gNBs 180a, 180b, and 180c can provide additional coverage and / or throughput for serving WTRUs 102a, 102b, and 102c.
[0077] Each of gNBs 180a, 180b, and 180c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, network slicing support, dual connectivity, interoperability between NR and E-UTRA, routing user plane data to User Plane Functions (UPF) 184a and 184b, and routing control plane information to Access and Mobility Management Functions (AMF) 182a and 182b, etc. Figure 1D As shown, gNB 180a, 180b, and 180c can communicate with each other via the Xn interface.
[0078] Figure 1DThe CN 115 shown may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and possibly a Data Network (DN) 185a, 185b. Although each of the foregoing elements is described as part of the CN 115, it should be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.
[0079] AMF 182a and 182b can connect to one or more gNBs 180a, 180b, and 180c in RAN 113 via the N2 interface and can be used as control nodes. For example, AMF 182a and 182b can be responsible for authenticating users of WTRU 102a, 102b, and 102c, supporting network slicing (e.g., handling different PDU sessions with different requirements), selecting specific SMF 183a and 183b, managing registration areas, terminating NAS signaling, mobility management, etc. AMF 182a and 182b can use network slicing to customize CN support for WTRU 102a, 102b, and 102c based on the service type used by WTRU 102a, 102b, and 102c. For example, different network slices can be established for different use cases (e.g., services relying on Ultra Reliable Low Latency (URLLC) access, services relying on Enhanced Massive Mobile Broadband (eMBB) access, services for Machine Type Communication (MTC) access, etc.). AMF 162 can provide control plane functions for handover between RAN 113 and other RANs (not shown) that employ other radio technologies, such as LTE, LTE-A, LTE-A Pro and / or non-3GPP access technologies, such as WiFi.
[0080] SMFs 183a and 183b can connect to AMFs 182a and 182b in CN 115 via the N11 interface. SMFs 183a and 183b can also connect to UPFs 184a and 184b in CN 115 via the N4 interface. SMFs 183a and 183b can select and control UPFs 184a and 184b, and configure the routing of services through UPFs 184a and 184b. SMFs 183a and 183b can perform other functions, such as managing and allocating UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, and providing downlink data notifications. PDU session types can be IP-based, non-IP-based, Ethernet-based, etc.
[0081] UPF 184a and 184b can be connected to one or more gNBs 180a, 180b, and 180c in RAN 113 via the N3 interface. This interface can provide WTRU 102a, 102b, and 102c with access to a packet-switched network (e.g., Internet 110) to facilitate communication between WTRU 102a, 102b, 102c and IP-enabled devices. UPF 184 and 184b can perform other functions such as routing and forwarding packets, enforcing user plane policies, supporting multi-destination PDU sessions, handling user plane QoS, buffering downlink packets, and providing mobility anchoring.
[0082] CN 115 can facilitate communication with other networks. For example, CN 115 may include, or be able to communicate with, an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) serving as an interface between CN 115 and PSTN 108. Furthermore, CN 115 can provide WTRUs 102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, WTRUs 102a, 102b, and 102c can be connected to local data networks (DNs) 185a and 185b via UPFs 184a and 184b through their N3 interfaces and the N6 interface between UPFs 184a and 184b and DNs 185a and 185b.
[0083] Given Figure 1A-1D as well as Figure 1A-1D As described in the corresponding descriptions herein, one or all of the functions described for one or more of the WTRU 102a-d, base station 114a-b, eNode-B 160a-c, MME 162, SGW 164, PGW 166, gNB 180a-c, AMF 182a-b, UPF 184a-b, SMF183a-b, DN 185a-b, and / or any other device described herein (one or more) may be performed by one or more emulation devices (not shown). An emulation device may be one or more devices configured to emulate one or more of the functions described herein. For example, an emulation device may be used to test other devices and / or simulate network and / or WTRU functions.
[0084] Simulation devices can be designed to perform one or more tests on other devices in laboratory and / or carrier network environments. For example, one or more simulation devices can perform one or more or all functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices within the communication network. One or more simulation devices can perform one or more or all functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. Simulation devices can be directly coupled to another device for testing and / or testing purposes that can be performed using over-the-air wireless communication.
[0085] One or more emulation devices may perform one or more functions, including all functions, rather than being implemented / deployed as part of a wired and / or wireless communication network. For example, emulation devices may be used in test scenarios outside of deployment (e.g., testing) wired and / or wireless communication networks and / or test laboratories to implement testing of one or more components. One or more emulation devices may be test equipment. Emulation devices may transmit and / or receive data using direct RF coupling and / or wireless communication via RF circuitry (e.g., which may include one or more antennas).
[0086] This application describes various aspects, including tools, features, examples, models, methods, etc. Many of these aspects are described in detail, and often in a manner that may sound restrictive, at least to illustrate individual characteristics. However, this is for the purpose of clarity and does not limit the application or scope of these aspects. In fact, all the different aspects can be combined and interchanged to provide other aspects. Furthermore, aspects can also be combined and interchanged with those described in earlier applications.
[0087] The aspects described and envisioned in this application can be implemented in many different forms. Figure 5-13 Some examples can be provided, but other examples can also be considered. Figure 5-13 The discussion does not limit the breadth of implementation. At least one aspect generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting the generated or encoded bitstream. These and other aspects can be implemented as methods, apparatus, computer-readable storage media having instructions stored thereon for encoding or decoding video data according to any of the described methods, and / or computer-readable storage media having bitstreams generated according to any of the described methods stored thereon.
[0088] In this application, the terms “reconstruction” and “decoding” are used interchangeably, the terms “pixel” and “sample” are used interchangeably, and the terms “image”, “picture” and “frame” are used interchangeably.
[0089] This document describes various methods, and each method includes one or more steps or actions for implementing the described method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions can be modified or combined. Furthermore, terms such as "first," "second," etc., can be used in various examples to modify elements, components, steps, operations, etc., such as "first decoding" and "second decoding," for example. Unless specifically required, the use of these terms does not imply a sequence of modified operations. Therefore, in this example, the first decoding does not need to be performed before the second decoding and can occur, for example, before, during, or in a time period overlapping with the second decoding.
[0090] The various methods and other aspects described in this application can be used to modify, for example... Figure 2 and Figure 3 The illustrated video encoder 200 and decoder 300 modules are, for example, decoding modules. Furthermore, the subject matter disclosed herein can be applied to, for example, any type, format, or version of video encoding (whether described in standards or recommendations, whether pre-existing or future-developed, and any extensions to such standards and recommendations). Unless otherwise stated or technically excluded, the aspects described in this application may be used individually or in combination.
[0091] Various numerical values, such as block size limits, are used in the examples described in this application. These and other specific values are for illustrative purposes only, and the aspects described are not limited to these specific values.
[0092] Figure 2 This is a diagram illustrating an example video encoder. Variations of the example encoder 200 are envisioned, but for clarity, encoder 200 is described below without describing all anticipated variations.
[0093] Before being encoded, the video sequence may undergo pre-coding (201), such as applying color transformations to the input color image (e.g., a conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing remapping of the input image components to obtain a more resilient signal distribution to compression (e.g., using histogram equalization with one of the color components). Metadata may be associated with pre-processing and appended to the bitstream.
[0094] In encoder 200, the image is encoded by encoder elements as described below. The image to be encoded is segmented (202) and processed in units, for example, coding units (CUs). Each unit is encoded using, for example, intra-frame or inter-frame modes. When a unit is encoded in intra-frame mode, it performs IntraTMP (260). In inter-frame mode, motion estimation (275) and compensation (270) are performed. The encoder determines (205) which of the intra-frame or inter-frame modes to use to encode the unit and indicates the intra-frame / inter-frame decision by, for example, a prediction mode flag. For example, the prediction residual is calculated by subtracting (210) the prediction block from the original image block.
[0095] The predicted residual is then transformed (225) and quantized (230). The quantized transform coefficients, along with the motion vector and other syntax elements, are entropy encoded (245) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass the transform and quantization, i.e., the residual is directly encoded without applying the transform or quantization process.
[0096] The encoder decodes the coded blocks to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inversely transformed (250) to decode the prediction residuals. The decoded prediction residuals and prediction blocks are combined (255) to reconstruct the image blocks. An in-loop filter (265) is applied to the reconstructed image to perform, for example, deblocking / SAO (Sample Adaptive Shift) filtering, thereby reducing coding artifacts. The filtered image is stored in a reference image buffer (280).
[0097] Figure 3 This is a diagram illustrating an example video decoder. In the example decoder 300, the bitstream is decoded by decoder elements, as described below. The video decoder 300 typically performs the same operations as... Figure 2 The encoding process described herein is the opposite of the decoding process. Encoder 200 typically also performs video decoding as part of the encoded video data.
[0098] Specifically, the decoder's input includes a video bitstream, which can be generated by the video encoder 200. The bitstream is first entropy decoded (330) to obtain transform coefficients, motion vectors, and other encoded information. Image segmentation information indicates how the image is segmented. Therefore, the decoder can segment (335) the image based on the decoded image segmentation information. The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residuals. The decoded prediction residuals and prediction blocks are combined (355) to reconstruct image blocks. The prediction blocks can be obtained from IntraTMP (360) or motion-compensated prediction (i.e., inter-frame prediction) (375) (370). An in-loop filter (365) is applied to reconstruct the image. The filtered image is stored in a reference image buffer (380).
[0099] The decoded image may undergo further post-decoding processing (385), such as inverse color transformation (e.g., a conversion from YCbCr4:2:0 to RGB4:4:4) or inverse remapping, which is the inverse of the remapping process performed in the pre-encoding process (201). Post-decoding processing may use metadata derived in the pre-encoding process and signaled in the bitstream. In the example, the decoded image (e.g., after applying an in-loop filter (365), and / or after post-decoding processing (385), if post-decoding processing is used) may be sent to a display device for presentation to the user.
[0100] Figure 4 This is a diagram illustrating an example of a system in which the various aspects and examples described herein can be implemented. System 400 can be implemented as a device including the various components described below and configured to perform one or more aspects described herein. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. The elements of system 400 may be embodied individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one example, the processing and encoder / decoder elements of system 400 are distributed across multiple ICs and / or discrete components. In various examples, system 400 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various examples, system 400 is configured to implement one or more aspects described herein.
[0101] System 400 includes at least one processor 410 configured to execute instructions loaded thereon for implementing various aspects described herein, such as those described herein. Processor 410 may include embedded memory, input / output interfaces, and various other circuitry known in the art. System 400 includes at least one memory 420 (e.g., a volatile memory device and / or a non-volatile memory device). System 400 includes a storage device 440 which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 440 may include internal storage devices, additional storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices.
[0102] System 400 includes an encoder / decoder module 430 configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 430 may include its own processor and memory. The encoder / decoder module 430 represents a module that can be included in a device to perform encoding and / or decoding functions. It is well known that a device may include one or both encoding and decoding modules. Furthermore, the encoder / decoder module 430 may be implemented as a separate element of system 400, or it may be included within processor 410 as a combination of hardware and software known to those skilled in the art.
[0103] Program code to be loaded onto processor 410 or encoder / decoder 430 to execute the various aspects described herein may be stored in storage device 440 and subsequently loaded onto memory 420 for execution by processor 410. According to various examples, one or more of processor 410, memory 420, storage device 440, and encoder / decoder module 430 may store one or more various items during the execution of the processes described herein. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.
[0104] In some examples, the memory within processor 410 and / or encoder / decoder module 430 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other examples, external memory (e.g., the processing device could be processor 410 or encoder / decoder module 430) is used for one or more of these functions. External memory could be memory 420 and / or storage device 440, such as volatile memory and / or non-volatile flash memory. In several examples, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one example, fast external volatile memory, such as RAM, is used as working memory for video encoding and decoding operations.
[0105] As shown in block 445, inputs can be provided to the components of system 400 through various input devices. Such input devices include, but are not limited to, (i) a radio frequency (RF) section that receives, for example, RF signals transmitted over the air by a broadcasting company, (ii) component (COMP) input terminals (or a set of COMP input terminals), (iii) universal serial bus (USB) input terminals, and / or (iv) high-definition multimedia interface (HDMI) input terminals. Figure 4 Other examples not shown include composite video.
[0106] In various examples, the input devices of block 445 have associated respective input processing elements, as known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or limiting a signal band to a band), (ii) down-converting the selected signal, (iii) further band-limiting to a narrower band to select, for example, a signal band, which in some examples may be referred to as a channel, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and / or (vi) demultiplexing to select a desired data packet stream. The RF section of various examples includes one or more elements performing these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include a tuner performing various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or baseband. In one set-top box example, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to the desired frequency band. Various examples rearrange the order of the aforementioned (and other) components, remove some of these components, and / or add other components that perform similar or different functions. Adding components may include inserting components between existing components, such as, for example, inserting amplifiers and analog-to-digital converters. In various examples, the RF section includes an antenna.
[0107] USB and / or HDMI terminals may include their respective interface processors for connecting system 400 to other electronic devices across USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) may be implemented as needed, for example, within a separate input processing IC or within processor 410. Similarly, as needed, various aspects of USB or HDMI interface processing may be implemented within a separate interface IC or within processor 410. The demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, processor 410 and encoder / decoder 430, which operate in combination with memory and storage elements to process the data streams as needed for presentation on the output device.
[0108] Various components of system 400 can be housed within an integrated housing. Within the integrated housing, various components can be interconnected and transmit data therebetween using a suitable connection arrangement 425 (e.g., internal buses known in the art, including inter-IC (I2C) buses, wiring, and printed circuit boards).
[0109] System 400 includes a communication interface 450 capable of communicating with other devices via a communication channel 460. The communication interface 450 may include, but is not limited to, a transceiver configured to send and receive data via the communication channel 460. The communication interface 450 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 460 may be implemented, for example, in a wired and / or wireless medium.
[0110] In various examples, data is streamed or otherwise provided to system 400 using a wireless network such as Wi-Fi (e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers)). In these examples, the Wi-Fi signal is received via a communication channel 460 and a communication interface 450 suitable for Wi-Fi communication. The communication channel 460 in these examples is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other examples use a set-top box to provide streaming data to system 400, with the set-top box transmitting data via an HDMI connection to input block 445. Still other examples use an RF connection to input block 445 to provide streaming data to system 400. As mentioned above, various examples provide data in a non-streaming manner. Furthermore, various examples use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth® networks.
[0111] System 400 can provide output signals to various output devices, including a display 475, a speaker 485, and other peripheral devices 495. Various examples of the display 475 include one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a flexible display, and / or a foldable display. The display 475 can be used in televisions, tablets, laptops, cellular phones (mobile phones), or other devices. The display 475 can also be integrated with other components (e.g., as in a smartphone) or standalone (e.g., an external monitor for a laptop). In various examples, other peripheral devices 495 include one or more of a standalone digital video disc (or digital multifunction disc) (DVD, for both terms), a disc player, a stereo system, and / or a lighting system. Various examples use one or more peripheral devices 495 that provide functionality based on the output of system 400. For example, a disc player performs the function of playing the output of system 400.
[0112] In various examples, signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols capable of enabling device-to-device control with or without user intervention is used to transmit control signals between system 400 and display 475, speaker 485, or other peripheral devices 495. Output devices can be communicatively coupled to system 400 via dedicated connections through their respective interfaces 470, 480, and 490. Alternatively, output devices can be connected to system 400 via communication interface 450 using communication channel 460. Display 475 and speaker 485 can be integrated into a single unit with other components of system 400 in electronic devices such as televisions. In various examples, display interface 470 includes display drivers, such as, for example, a timing controller (TCon) chip.
[0113] For example, if the RF section of input 445 is part of a standalone set-top box, then display 475 and speaker 485 can alternatively be separated from one or more other components. In various examples where display 475 and speaker 485 are external components, the output signal can be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.
[0114] These examples can be implemented by computer software implemented by processor 410, or by hardware, or by a combination of hardware and software. As a non-limiting example, these examples can be implemented by one or more integrated circuits. Memory 420 can be of any type suitable for the technical environment and can be implemented using any suitable data storage technology, such as, as a non-limiting example, optical storage devices, magnetic storage devices, semiconductor-based memory devices, fixed memory, and removable memory. As a non-limiting example, processor 410 can be of any type suitable for the technical environment and can include one or more of microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.
[0115] Various implementations involve decoding. As used in this application, "decoding" can include, for example, all or part of a process performed on a received coded sequence to generate a final output suitable for display. In various examples, these processes include one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various examples, these processes also include, or alternatively include, processes performed by a decoder of the various implementations described in this application, such as obtaining a first prediction signal based on an extrapolation filter, obtaining a second prediction signal, generating a mixed prediction based on weighted prediction samples and the prediction signal, decoding the current block based on the mixed prediction, and so on.
[0116] As further examples, in one example, "decoding" refers only to entropy decoding; in another example, "decoding" refers only to differential decoding; and in yet another example, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or to refer to a broader decoding process will be clear based on the specific context of the description and is considered well understood by those skilled in the art.
[0117] Various implementations involve encoding. Similar to the discussion of “decoding” above, “encoding” as used in this application can include, for example, all or part of a process performed on an input video sequence to generate an encoded bitstream. In various examples, these processes include one or more processes typically performed by an encoder, such as segmentation, differential coding, transform, quantization, and entropy coding. In various examples, these processes also include, or alternatively include, processes performed by an encoder of the various implementations described in this application, such as obtaining a first prediction signal based on an extrapolation filter, obtaining a second prediction signal, generating a mixed prediction based on a weighted prediction sample and the prediction signal, encoding the current block based on the mixed prediction, and so on.
[0118] As further examples, in one example, "encoding" refers only to entropy encoding; in another example, "encoding" refers only to differential encoding; and in yet another example, "encoding" refers to a combination of differential and entropy encoding. Whether the phrase "encoding process" is intended to specifically refer to a subset of operations or to refer to a broader encoding process will be clear based on the specific context of the description and is considered well understood by those skilled in the art.
[0119] Note that the syntax elements used in this article (such as the encoding syntax for indicators (e.g., cu_eip_flag, eip_merge_flag), indexes (e.g., eip_merge_idx), etc.) are descriptive terms. Therefore, they do not preclude the use of other syntax element names.
[0120] When a diagram is presented as a flowchart, it should be understood that it also provides a block diagram of the corresponding device. Similarly, when a diagram is presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / process.
[0121] The implementations and aspects described herein can be implemented, for example, in methods or processes, apparatuses, software programs, data streams, or signals. Even if discussed only in the context of a single implementation (e.g., discussed only as a method), the implementation of the features in question can be implemented in other forms (e.g., apparatuses or programs). Apparatuses can be implemented, for example, in suitable hardware, software, and firmware. Methods can be implemented, for example, in a processor, which generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as, for example, computers, cellular phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate communication of information between end users.
[0122] References to “an example” or “an example” or “an implementation” or “an implementation” and their other variations mean that the specific features, structures, characteristics, etc., described in the example are included in at least one example. Therefore, the appearance of the phrase “in an example” or “in the example” or “in an implementation” or “in the implementation”, and any other variations appearing in various places throughout this application, do not necessarily refer to the same example.
[0123] Furthermore, this application may relate to "determining" various information fragments. Determining information may include, for example, one or more of estimated information, calculated information, predicted information, or information retrieved from memory. Obtaining may include receiving, retrieving, constructing, generating, and / or determining.
[0124] Furthermore, this application may relate to "accessing" various information fragments. Accessing information may include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information, or one or more of these.
[0125] Furthermore, this application may relate to "receiving" various pieces of information. Like "access," receiving is intended to be a broad term. Receiving information may include, for example, accessing information or retrieving information (e.g., from memory) one or more. Moreover, "receiving" is generally referred to in one way or another during operations such as, for example, storing information, processing information, sending information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0126] It should be understood that any use of " / ", "and / or", and "at least one" (e.g., in the cases of "A / B", "A and / or B", and "at least one of A and B") is intended to include selecting only the first listed option (A), or only the second listed option (B), or both options (A and B). As another example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such wording is intended to include selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A, B, and C). This can be extended to as many items as listed, as will be apparent to those skilled in the art and related fields.
[0127] Furthermore, as used herein, the word "signal" specifically refers to instructing the corresponding decoder to do something. Encoder signals may include, for example, intra-prediction (EIP) related syntax based on extrapolation filters. Thus, in one example, the same parameters are used on both the encoder and decoder sides. Therefore, for example, the encoder can send (explicit signaling) specific parameters to the decoder so that the decoder can use the same specific parameters. Conversely, if the decoder already has specific parameters and others, signaling can be used without sending them (implicit signaling) to simply allow the decoder to know and select specific parameters. Bit savings are achieved in various examples by avoiding the transmission of any actual functionality. It should be understood that signaling can be done in many ways. For example, in various examples, one or more syntax elements, flags, etc., are used to signal information to the corresponding decoder. Although the verb form of the word "signaling" has been used above, the word "signaling" can also be used as a noun in this article.
[0128] It will be apparent to those skilled in the art that implementations can generate various signals that are formatted to carry information, for example, that can be stored or transmitted. The information may include, for example, instructions for performing a method, or data generated by one of the described implementations. For example, a signal may be formatted to carry a bitstream of the described example. Such a signal may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or baseband signals. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. It is well known that signals can be transmitted via a variety of different wired or wireless links. Signals may be stored on, or accessed or received from, a processor-readable medium.
[0129] This document describes numerous examples. Features of the examples may be provided individually or in any combination across various claim classes and types. Furthermore, examples may include one or more of the features, devices, or aspects described herein (individually or in any combination across various claim classes and types). For example, features described herein may be implemented in a bitstream or signal that includes information generated as described herein. This information may allow a decoder to decode the bitstream, and an encoder, bitstream, and / or decoder may be implemented according to any of the embodiments described. For example, features described herein may be implemented by creating and / or transmitting and / or receiving and / or decoding a bitstream or signal. For example, features described herein may implement a method, process, apparatus, medium storing instructions, medium storing data, or signal. For example, features described herein may be implemented by a TV, set-top box, cellular phone, tablet computer, or other electronic device performing decoding. The TV, set-top box, cellular phone, tablet computer, or other electronic device may display (e.g., using a monitor, screen, or other type of display) a resulting image (e.g., an image reconstructed from the residual of a video bitstream). The TV, set-top box, cellular phone, tablet computer, or other electronic device may receive a signal including an encoded image and perform decoding.
[0130] Extrapolation-based intra-frame prediction (EIP) can be a coding / decoding tool. EIP learns extrapolation filters from a reconstruction template and applies them to the current block (e.g., to generate a prediction signal). EIP can provide significant coding / decoding gains. EIP can be used less frequently than regular prediction modes (e.g., because EIP can be used as a single predictor without being fused / mixed with other prediction modes). Mixing modes can combine EIP with other intra-frame / intra-prediction modes, for example, to improve prediction quality and / or coding / decoding gains.
[0131] It can perform decoder-side intra-frame mode derivation (DIMD).
[0132] When DIMD is applied, intra-frame modes (e.g., up to five intra-frame modes) can be derived from reconstructed neighboring samples, and predictors (e.g., these five predictors) can be combined with planar mode predictors, such as with weights derived from gradient histograms. The lookup table (LUT)-based integration scheme used by CCLM (e.g., the same LUT-based integration scheme) can be used to perform division operations in weight derivation. For example, division operations in orientation calculation: The following LUT-based scheme can be used to calculate it: in: .
[0133] For a block of size W×H, the weights of each of the five derivation patterns can be modified, for example, if one of the upper or left histogram magnitudes is twice as large as another. In this case, the weights can be position-dependent and calculated as follows.
[0134] If the histogram above is twice the size of the one on the left, then: .
[0135] If the histogram on the left is twice the size of the histogram on the top, then: in It is the unmodified uniform weight of the selected DIMD, and It can be predefined (e.g., set to 10).
[0136] The derived intra-frame modes can be included in the main list (MPM) of the most probable intra-frame modes, so the DIMD process can be performed before the MPM list is constructed. The main derived intra-frame modes of a DIMD block can be stored with the block and / or used for the construction of the MPM list of adjacent blocks.
[0137] The regions used to compute the gradient histogram of adjacent reconstructed samples can be modified, for example, depending on the availability of reconstructed samples. The region of the decoded reference sample of the current WxH brightness CB can be extended upward to the right (e.g., up to W additional columns) and / or downward to the left (e.g., up to H additional rows) if available.
[0138] DIMD merging mode can be performed. When using DIMD merging, DIMD information extracted from neighboring blocks can be used to compute intra-prediction for the current block. A new merged gradient histogram (MHoG) can be computed for the current block based on the HoG of neighboring blocks. Only neighboring blocks encoded with DIMD or DIMD merging can be considered.
[0139] When a single DIMD or a DIMD-merged neighboring block is available, its gradient histogram can be used to form the MHoG for the current block. If more than one DIMD or DIMD-merged neighboring block is available, the corresponding histograms can be combined, for example, by deriving the MHoG through magnitude averaging. CUs surrounding the current block (e.g., up to 13 CUs) can be considered to extract DIMD information.
[0140] MHoG can be used to compute intra-frame prediction modes and / or weights (e.g., as in DIMD). The five directional modes corresponding to the highest amplitudes in the MHoG and their weights can be selected, and the corresponding predictors can be mixed (e.g., as in DIMD).
[0141] Template-based intra-mode derivation (TIMD) fusion can be performed. For intra-prediction modes in the MPM (e.g., each intra-prediction mode) and wide-angle modes (if top-right and / or bottom-left reference samples are available), the SATD between the template's predictions and reconstructed samples can be computed. Two intra-prediction modes with the minimum SATD can be selected as TIMD modes. After applying the PDPC procedure, these two TIMD modes can be fused with weights. Weighted intra-prediction can be used to encode the current CU. Position-dependent intra-prediction combination (PDPC) can be included in the derivation of the TIMD modes.
[0142] The costs of the two selection modes are compared with a threshold. In the test, the cost factor of 2 can be applied as follows: .
[0143] If the condition is true, fusion can be applied; otherwise, mode 1 can be used (e.g., mode 1 only).
[0144] The weights of the patterns can be calculated based on their SATD costs as follows: .
[0145] Division operations can be performed (e.g., using the same lookup table (LUT)-based integration scheme used by CCLM).
[0146] Intra-frame prediction fusion can be performed. This intra-frame prediction method derives prediction samples as a weighted combination of multiple predictors generated from different reference lines. In this process, multiple intra-frame predictors can be generated, and these predictors can be fused by weighted averaging. The process of deriving the predictors to be used in the fusion process can include one or more of the following.
[0147] For the intra-frame prediction mode of a single mode including TIMD and DIMD, the proposed method can be applied by representing it as... Intra-prediction is derived by weighting the intra-prediction obtained from multiple reference lines, where... It can be an intra-frame prediction from the default reference line, and This can be a prediction from a line above the default reference line. The weights can be set to... and .
[0148] For TIMD modes with hybridity, Available in first mode , Available in second mode .
[0149] For DIMD patterns with a mix, the number of predictors selected for the weighted average can be increased from 3 to 6.
[0150] Intra-prediction fusion methods can be applied to luma blocks, for example, when the intra-angle mode has a non-integer slope (e.g., required reference sample interpolation) and the block size is sufficient (e.g., greater than 16). Intra-prediction fusion methods can be used with MRL (e.g., and not applied to blocks encoded / decoded by ISP). In the method studied in subtest a, PDPC can be applied to the intra-prediction mode using the reference line closest to the current block.
[0151] Intra-template matching can be performed. Intra-template matching prediction (IntraTMP) is an intra-prediction mode that copies the best prediction block from the reconstructed portion of the current frame, whose template (e.g., an L-shaped template) matches the current template. For a predefined search range, the encoder can search for the template most similar to the current template in the reconstructed portion of the current frame and use the corresponding block as the prediction block. The encoder can signal the use of this mode, and the same prediction operation can be performed on the decoder side.
[0152] Figure 5 The example intra-frame template matching search region used is described. This can be achieved by matching the L-shaped, top-only, and / or left-only causally adjacent blocks of the current block with another block in a predefined search region (e.g., such as...). Figure 5 (As shown) matching is used to generate the predicted signal. There may be a predefined search region (e.g., as shown) Figure 5 The six predefined search regions (R1 to R6) may include reconstruction samples from the top-left CTU as well as partial reconstruction samples located above, to the left, to the bottom left, and / or to the top right of the current CTU.
[0153] The sum of absolute differences (SAD) can be used as the cost function. A given search order of regions can be utilized (e.g., 6 regions, such as the order of R4, R5, R6, R1, R2, and R3). Within a region (e.g., each region), the decoder can construct a candidate list of template-matched block vectors (e.g., up to 19 template-matched block vectors), which can be ordered in ascending order, for example, according to the template cost (e.g., SAD). One or more of the following modes can be supported: single predictor, fusion of multiple predictors, subpixel precision, or linear filter model.
[0154] For a single predictor, you can select a single predictor from the candidate list.
[0155] For the fusion of multiple predictors, multiple predictors can be combined to derive the final prediction block. The combined weights can be calculated based on the template matching cost of each predictor and / or using a Wiener filter-based weight derivation method.
[0156] For subpixel precision (e.g., when using a single predictor), subpixel precision can be used in conjunction with 1 / 2 pixel precision, 1 / 4 pixel precision, and / or 3 / 4 pixel precision, with 8 possible directions (e.g., each with 8 possible directions).
[0157] For a linear filter model, the linear filter can be learned, for example, between a reference template and the current template, and can be applied to the reference block by the linear model. This pattern can be used for a single predictor (e.g., when subpixel precision is not used).
[0158] The size of the region (SearchRange_w, SearchRange_h) can be set to be proportional to the block size (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. That is: , Here, 'a' is a constant that controls the gain / complexity tradeoff. For example, 'a' can be equal to 5.
[0159] The search range of the search region can be subsampled by a factor of 3 (e.g., to speed up the template matching process). For example, after finding the best match, a refinement process can be performed. Refinement can be done by a second template matching search around the best match (e.g., with a narrowed range).
[0160] Intra-frame template matching can be enabled for a CU. For example, intra-frame template matching can be enabled for CUs with a width and height less than or equal to 64. The maximum CU size used for intra-frame template matching is configurable.
[0161] Intra-frame template matching prediction mode can be implemented in CU-level signaling, for example, via dedicated flags (e.g., when DIMD is not used in the current CU).
[0162] Combining TIMD and TM merging can be used to perform combined inter-frame intra-frame prediction (CIIP). In CIIP mode, prediction samples can be generated by weighting the inter-frame prediction signal predicted using CIIP-TM merging candidates and the intra-frame prediction signal predicted using the intra-frame prediction mode derived using TIMD. This method can be applied to (e.g., only to) codec blocks with an area less than or equal to 1024.
[0163] The TIMD derivation method can be used to derive intra-prediction modes in CIIP. Intra-prediction modes with the smallest SATD value can be selected from the TIMD mode list and mapped to intra-prediction modes (e.g., one of the 67 regular intra-prediction modes).
[0164] If the derived intra-prediction mode is angular mode, the weights (wIntra, wInter) can be modified for the two tests. Figures 6A-6B An example partitioning method for angle patterns is described. For patterns that are close to horizontal (e.g., 2 ≤ angle pattern index < 34), the current block can be partitioned vertically, such as... Figure 6A As shown. For near-vertical patterns (34 ≤ angle pattern exponent ≤ 66), the current block can be divided horizontally, as... Figure 6B As shown.
[0165] The different sub-blocks (wIntra, wInter) are shown in Table 1.
[0166] Table 1: Example weight modifications for angle mode: .
[0167] CIIP-TM can be used to create a CIIP-TM merge candidate list for CIIP-TM patterns. Merge candidates can be refined using template matching. CIIP-TM merge candidates can be reordered (e.g., also reordered) into regular merge candidates using the ARMC method. The maximum number of CIIP-TM merge candidates can be equal to 2.
[0168] It can perform GPM with inter-frame and intra-frame prediction. Figures 7A-7C An example GPM with inter-frame and intra-frame prediction is described. Figure 7D An example GPM with intra-frame and intra-frame prediction is described.
[0169] In a GPM with inter-frame and intra-frame prediction, the final prediction samples can be generated by weighting the inter-frame prediction samples and intra-frame prediction samples of the GPM separation regions (e.g., each GPM separation region). Inter-frame prediction samples can be derived from inter-frame GPMs. Intra-frame prediction samples can be derived from an intra-frame prediction mode (IPM) candidate list and an index from encoder signaling. The size of the IPM candidate list can be predefined (e.g., set to 3). Available IPM candidates can be parallel angle modes (e.g., parallel mode) for GPM block boundaries, vertical angle modes (e.g., vertical mode) for GPM block boundaries, and planar modes, such as... Figure 7A , 7B As shown in 7C. Figure 7D As shown, GPM with intra-frame and intra-frame prediction can be limited to reduce the signaling overhead of IPM and / or avoid increasing the size of intra-frame prediction circuitry on the hardware decoder. Direct motion vectors and IPM storage on the GPM mixing region can be introduced (e.g., additionally) to further improve encoding and decoding performance.
[0170] In DIMD-based and neighbor-mode-based IPM derivation, parallel modes can be registered first. Therefore, up to two IPM candidates derived from the decoder-side intra-frame mode derivation (DIMD) method and / or neighboring blocks can be registered, for example, if no identical IPM candidates are in the list. For neighbor-mode derivation, there are available neighboring block locations (e.g., up to five locations), but they may be limited by the angle of the GPM block boundaries as shown in Table 2, which can be used for GPM with template matching (GPM-TM).
[0171] Table 2: Positions of available neighboring blocks derived from IPM candidate derivation based on the angle of the GPM block boundary. A and L can represent the top and left sides of the predicted block, respectively: .
[0172] GPM-intraframe can be combined with GPM with motion vector difference combining (GPM-MMVD). TIMD can be used as an IPM candidate (e.g., on) within GPM-intraframe to improve encoding / decoding performance. Parallel modes can be registered first, followed by TIMD, DIMD, and IPM candidates for adjacent blocks.
[0173] It can execute the Spatial Geometry Partitioning (SGPM) mode. Figure 8 Example space GPM candidates are described. SGPM is an intra-frame mode that can be similar to inter-frame coding / decoding tools for GPM, for example, where two prediction parts can be generated from the intra-frame prediction process. In this mode, a candidate list can be constructed where each entry includes a partition split and two intra-frame prediction modes, such as... Figure 8 As shown. For example, a combination can be formed using partitioned mode and three intra-frame prediction modes. The length of the candidate list can be (pre-)configured (e.g., set to equal to 16). The selected candidate index can be signaled.
[0174] The GPM candidate in the example space can be represented as follows: .
[0175] The parameter spgm_cand_idx can include or represent one or more of partition_mode_idx, intra_pred_mode0_idx, or intra_pred_mode1_idx.
[0176] Figure 9 A sample GPM template is described. For example... Figure 9 As shown, a template can be used to reorder a list, where the SAD between the template's prediction and reconstruction is used for sorting. The template size can be fixed (e.g., fixed at 1).
[0177] For partitioned modes (e.g., per partition mode), an IPM list can be derived for each partition using, for example, an intra-inter-frame GPM list derivation (e.g., as described herein). The IPM list size can be (pre-)configured (e.g., set to 3). In the list, the TIMD derivation pattern can be replaced by two derivation patterns with horizontal and vertical orientations.
[0178] The SGPM pattern can be applied with limited block sizes, for example: 4≤width≤64, 4≤height≤64, width<height*8, height<width*8, width*height≥32.
[0179] The PPS flag can be encoded / decoded to indicate whether blending of two intra-frame predictions is allowed. When the PPS flag is set to false, the following adaptive blending can be used (e.g., also for spatial GPM), where the blending depth τ can be derived as follows: .
[0180] Otherwise (e.g., the PPS flag can be set to true), 1 / 4τ can be used (e.g., always used) for spatial GPM codec blocks to ensure that blending is not used when SGPM blocks have completely horizontal or vertical partition angles, and a narrower blending width is used when SGPM blocks have other partition angles. In some Common Test Conditions (CTC) for screen content video, this flag can be set (e.g., set to true).
[0181] It can perform combined intra-block copying and intra-prediction (IBC-CIIP). IBC-CIIP is an encoding / decoding tool for CUs that uses IBC and intra-prediction to obtain two prediction signals, which can then be weighted and summed to generate the final prediction, as follows: .
[0182] in and These can be represented as IBC prediction signal and intra-frame prediction signal, respectively. For IBC combining mode and IBC AMVP mode, It can be set to (13, 4) and (1, 1).
[0183] An intra-prediction mode (IPM) candidate list can be used to generate intra-prediction signals, and the size of the IPM candidate list can be (pre)defined (e.g., defined as 2). The IPM index can be signaled to indicate which IPM to use.
[0184] IBC with geometric partitioning mode (IBC-GPM) can be executed. IBC-GPM is a codec tool that geometrically divides the CU into two sub-partitions. Predicted signals for the two sub-partitions can be generated using IBC and intra-frame prediction. IBC-GPM can be applied to regular IBC merging mode or IBC TM merging mode. A candidate list of intra-frame prediction modes (IPM) can be constructed, for example, using the same method as GPM with inter-frame and intra-frame prediction for intra-frame prediction, and the size of the IPM candidate list can be (pre)defined (e.g., defined as 3). There are a total of 48 geometric partitioning modes, which can be divided into two sets of geometric partitioning modes, as follows.
[0185] Table 3: Geometric segmentation patterns in the first geometric segmentation pattern set: .
[0186] Table 4: Geometric segmentation patterns in the second geometric segmentation pattern set: .
[0187] When using IBC-GPM, the IBC-GPM geometric segmentation mode set flag can be signaled to indicate whether to select the first or second geometric segmentation mode set, followed by the geometric segmentation mode index. The IBC-GPM intra-frame flag can be signaled to indicate whether intra-frame prediction is used for the first sub-partition. When intra-frame prediction is used for a sub-partition, the intra-frame prediction mode index can be signaled. When IBC is used for a sub-partition, the merge index can be signaled.
[0188] In bidirectional prediction IBC GPM, two flags can be signaled to indicate the prediction mode for the two partitions. The first flag indicates whether the first partition is intra-predictive; if not, the second flag can be signaled to indicate whether intra-prediction is used for the second partition. This method can be applied to screen content coding (SCC) (e.g., only to SCC).
[0189] Intra-frame prediction (EIP) based on extrapolation filters can be performed. EIP can be processed in three parts. The extrapolation filter coefficients can be derived from adjacent reconstructed regions of the current block or inherited from previous EIP blocks. The extrapolation process (e.g., then) can generate a prediction signal from top left to bottom right within the current block. (e.g., then) the intra-frame prediction angle can be derived by analyzing the gradient of the prediction block, and the corresponding intra-frame mode can be used to select the MTS, NSPT, and LFNST kernels for transformation.
[0190] Based on size and / or components, the application of EIP can be limited to blocks. For example, EIP can be applied to blocks no larger than 32×32 and only the luma component.
[0191] An EIP filter can be obtained. Figure 10 An example EIP filter shape with fifteen inputs and one output is described.
[0192] There are two possible ways to obtain the filter coefficients of the current CU. The coefficients can be derived from adjacent reconstructed pixels, and / or the coefficients can be inherited from previously decoded blocks.
[0193] EIP coefficients can be derived. The decoder decodes relevant syntax elements to determine the reconstruction region and filter shape for the current block's selection type. The selected filter can be moved horizontally or vertically within the selected reconstruction region (e.g., with a one-pixel stride) to construct the autocorrelation matrix and cross-correlation vector. The coefficients calculated from the autocorrelation matrix and cross-correlation vector can be similar to (e.g., identical) those in the Convolutional Cross Component Model (CCCM).
[0194] Figure 11 This describes an example definition type for the refactored region. The size of the refactored region can depend on... And / or the selected filter shape. For example, when the current block is an 8×16 block and the selected filter shape is 4×4, the aboveSize of the reconstructed region can be equal to min(8,16)+4–1=11, and the leftSize of the reconstructed region can be equal to min(8,16)+4–1=11.
[0195] EIP filters can be inherited. EIP merging patterns can be executed. Filter shapes and filter coefficients can be inherited from previously decoded blocks using EIPs or EIP merging patterns. The decoder can decode EIP merging flags to determine whether the proposed merging pattern should be used when the current block uses an EIP pattern. For example, when the EIP merging flag is true, the merging index can be decoded (e.g., further decoded). The EIP merging list can include spatially neighboring and non-neighboring candidates, temporal candidates, and / or historical candidates. The constructed EIP merging list can include up to 12 candidates, and the list can be reduced to up to 6 candidates through a reordering process, for example, based on the SAD cost measured on an L-shaped template with a column width and row height of 1. In SAD computation, predictions for template regions by EIP filters can be generated (e.g., generated only) from reconstructed samples (e.g., neighboring samples and template samples), which allows EIP filters to be applied in parallel (e.g., rather than sequentially).
[0196] Spatial proximity, temporal proximity, non-proximity, temporal shift, and historical candidate location and inclusion order can be similar to (e.g., identical to) those defined for CCP merging prediction candidates.
[0197] The current block can be predicted. Figure 12 An example is described where predictions are generated for different predictions in the current block in diagonal order. For example... Figure 12 As shown, the EIP mode can generate the prediction value of the current block by predicting the diagonal position from the top left to the bottom right.
[0198] The predicted value in this contribution is calculated as follows: .
[0199] in, It could be the predicted value at (x, y) in the current block. This can be the i-th coefficient of the selected EIP filter, and the coefficient index can be from 0 to 14. It can be a reconstruction or prediction of the current location. and It can be a position offset relative to the current position along the x and y directions, respectively.
[0200] Low-frequency non-separable transform (LFNST) / non-separable master transform (NSPT) / multiple transform selection (MTS) sets can be mapped.
[0201] The DIMD process can be used to derive the intra-prediction mode of the current block based on EIP prediction samples. Horizontal and vertical gradients can be computed for each prediction sample to construct a gradient histogram (HoG). The intra-prediction mode corresponding to the maximum histogram count can be used (e.g., and then can be used to) determine the LFNST, NSPT, and / or MTS transform sets.
[0202] Signaling CU-level syntax is possible. For example, EIP-related syntax can be used in CU-level signaling. Table 5 shows examples of EIP-related syntax.
[0203] Table 5: Example EIP related syntax: .
[0204] Figure 13 An example method for combining EIP with other forecasting models (e.g., CIIP and / or GPM) is described. Forecasts obtained from EIP can be examined through template analysis and blended according to the selected model. EIP can be combined with one or more of the following models: CIIP, GPM, SGPM, or IBC-GPM.
[0205] EIP can be used in combination with CIIP. CIIP mode combines inter-frame prediction with intra-frame prediction obtained from TIMD mode. The mixing process can be (pre-)defined according to TIMD mode. IBC-CIIP can be used (e.g., in conjunction with CIIP). Some differences may be that IBC-CIIP can use IBC prediction (e.g., instead of inter-frame prediction) and a list of 3 candidate intra-frame predictions, and the mixing weights can differ from the CIIP mixing weights (e.g., the default CIIP mixing weights).
[0206] To combine EIPs with CIIPs, an EIP can be combined with a CIIP (e.g., the default CIIP) and / or an IBC-CIIP (e.g., the default IBC-CIIP). When an EIP is combined with the default CIIP, the EIP mode can be used instead of the TIMD mode for intra-frame prediction. When an EIP is combined with the default IBC-CIIP, the EIP mode can replace the default mode in the intra-frame prediction list.
[0207] When EIPs are combined (e.g., with CIIPs or IBC-CIIPs), three modes of the EIPs (full, horizontal, and vertical) can be considered. Furthermore, candidate EIP modes can be used (e.g., may also be merged). Template analysis can be used to determine the optimal mode. For example, predictions can be performed on the reconstructed template, and cost metrics (e.g., SATD) can be used. Template cost can be used to compare different EIP modes with regular intra-frame modes. The mode with the lowest cost can be used and can replace the intra-frame portion of CIIPs or IBC-CIIPs. IBC-CIIPs can be generalized to IntraTMP-CIIPs, where IBC predictions can be replaced by IntraTMP predictions.
[0208] The EIP-CIIP mode can be used. The video encoding device can determine whether to use EIP-CIIP (e.g., for blocks) and can include an EIP-CIIP mode indication in the video data based on that determination. For example, the EIP-CIIP mode can be signaled (e.g., named EIP-CIIP). The video decoder can determine whether to execute EIP-CIIP (e.g., for blocks) based on this indication. For example, if the mode is signaled, one or more of the following can be executed.
[0209] EIP prediction can be performed based on the selected EIP mode. In some examples, intra-frame prediction can be performed based on TIMD results, and the default CIIP blending can be used. In some examples, intra-frame prediction can be performed based on the IPM list and the IPM index of the signaling, and IBC-CIIP blending can be used. In these examples, the inter-frame portion or the IBC portion can be replaced by EIP. Predicted blocks obtained via EIP can be blended with predicted blocks obtained via intra-frame prediction (e.g., via the IPM list and IPM index).
[0210] In some examples, when EIP is mixed with intra-frame prediction, if the intra-frame prediction is in angular intra-frame mode, the EIP filtering process can be performed using the scan path direction derived from the angular direction (or a function of the angular direction). For example: if the angular direction is close to vertical, EIP filtering can be performed by scanning from top to bottom, column by column. For example, if the angular direction is close to horizontal, EIP filtering can be performed by scanning from left to right, row by row. For example, if the angular direction is close to diagonal, EIP filtering can be performed from the top left to the bottom right position according to the diagonal prediction order.
[0211] In some examples, the scan path can be perpendicular to the dividing line direction.
[0212] EIP can be used in conjunction with GPM. In GPM with intra-frame and inter-frame prediction, the intra-frame portion can be selected from an IPM list including modes perpendicular to the segmentation, modes parallel to the segmentation, and planar modes. These three modes can be replaced by EIP modes obtained from the current block and / or merge candidate blocks.
[0213] Template cost (e.g., SATD) can be used to compare EIP modes and regular modes. That is, prediction can be performed, and the template cost of reconstructing the template can be measured. Template cost can be used to determine the three EIP modes. Furthermore, the default regular modes (vertical, parallel, and planar) can be compared to determine which mode to mix. TIMD can be used for intra-frame IPM candidates to improve (e.g., further improve) encoding / decoding performance. Parallel modes can be registered before TIMD, DIMD, and IPM candidates for adjacent blocks.
[0214] In some examples, the EIP filtering process can be performed using a scan order direction derived from the segmentation line direction (or a function of the segmentation line direction). For example, if the segmentation line direction is nearly vertical, EIP filtering can be performed using a top-to-bottom, column-by-column scan order. If the segmentation line direction is nearly horizontal, EIP filtering can be performed using a left-to-right, row-by-row scan order. If the segmentation line direction is nearly diagonal, EIP filtering can be performed from the top-left to the bottom-right position according to the diagonal prediction order.
[0215] In some examples, the scanning order can be perpendicular to the dividing line direction.
[0216] EIP can be used in conjunction with SGPM. In SGPM, template analysis can be used to derive two intra-frame modes with minimum template cost and 16 candidate segmentation lines. For each segmentation direction, three intra-frame modes can be defined.
[0217] To combine EIP and SGPM, EIP modes can be added to the intra-candidate list and / or can replace intra-candidate modes in the intra-candidate list.
[0218] For example, the intra-candidate list in SGPM can include three candidates, and one or more EIP modes can be added. An EIP can have three modes, and can (e.g., may also) consider other modes from merged candidates. To limit encoding / decoding complexity, only good candidates (e.g., those with the lowest template cost) can be considered. Template cost can be used to pre-select several candidates (e.g., those with the lowest template cost) to add to the internal candidate list.
[0219] In the second option, the same number of intra-frame modes can be kept in the intra-frame candidate list (e.g., 3), and one or more intra-frame modes can be replaced by those in the EIP. This reduces encoding / decoding complexity because the number of combinations remains unchanged. To improve encoding / decoding performance, template analysis can be employed. If the EIP template cost is less than that of a regular mode, template analysis can be performed to selectively replace some intra-frame candidates with EIP modes.
[0220] In some examples, the EIP filtering process can be performed using a scan order direction derived from the segmentation line direction (or a function of the segmentation line direction).
[0221] In some examples, the new EIP intra-GPM mode can be used. For instance, one sub-partition may have only EIP mode, while another sub-partition may have only regular intra-frame mode.
[0222] EIP can be used in combination with IBC-GPM. The combination of EIP and IBC-GPM can complement the examples and variations described above. IBC-GPM can be considered (e.g., in place of normal GPM). The intra-frame portion of IBC-GPM can be replaced by the EIP portion. Adaptations can be similar to (e.g., identical) those considered in the previous examples and variations described herein.
[0223] Although the features and elements have been described above in specific combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with other features and elements. Furthermore, the methods described herein can be implemented in computer programs, software, or firmware, including in computer-readable media for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted via wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor storage devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROMs and digital multifunction discs (DVDs). A processor associated with the software can be used to implement a radio frequency transceiver used in a WTRU, UE, terminal, base station, RNC, or any host computer.
Claims
1. A video decoding device, comprising: The processor is configured as follows: Obtain a first prediction associated with the current block, wherein the first prediction is obtained based on an extrapolation filter; Obtain the second prediction associated with the current block; A hybrid prediction is generated based on the first prediction and the second prediction; and Decode the current block based on hybrid prediction.
2. A video encoding device, comprising: The processor is configured as follows: Obtain a first prediction associated with the current block, wherein the first prediction is obtained based on an extrapolation filter; Obtain the second prediction associated with the current block; A hybrid prediction is generated based on the first prediction and the second prediction; and The current block is encoded based on hybrid prediction.
3. The video device according to any one of claims 1 or 2, wherein, The second prediction is obtained based on the inter-frame prediction mode.
4. The video device according to any one of claims 1 or 2, wherein, The second prediction is obtained based on the intra-frame prediction mode.
5. The video device according to any one of claims 1-4, wherein, The processor is also configured to: The extrapolation filter is obtained based on the reconstruction template associated with the current block.
6. The video device according to any one of claims 1 or 2, wherein, The second prediction is at least one of the following: an inter-frame prediction signal obtained based on a merged candidate list, an intra-frame prediction signal obtained based on an intra-frame prediction mode (IPM) list and an IPM index of signaling, or an intra-frame prediction signal obtained based on a template-based intra-frame mode derivation (TIMD) result.
7. The video device according to any one of claims 1-6, wherein, The processor is also configured to: Determine the adjacent reconstruction regions of the current block and the filter shape of the current block; and The extrapolation filter coefficients associated with the current block are determined based on the adjacent reconstruction regions and the filter shape, wherein the second prediction associated with the current block is obtained based on the extrapolation filter coefficients.
8. A video decoding method, comprising: Obtain a first prediction associated with the current block, wherein the first prediction is obtained based on an extrapolation filter; Obtain the second prediction associated with the current block; A hybrid prediction is generated based on the first and second predictions; and Decode the current block based on hybrid prediction.
9. A video coding method, comprising: Obtain a first prediction associated with the current block, wherein the first prediction is obtained based on an extrapolation filter; Obtain the second prediction associated with the current block; A hybrid prediction is generated based on the first and second predictions; and The current block is encoded based on hybrid prediction.
10. The method according to any one of claims 8 or 9, wherein, The second prediction is obtained based on the inter-frame prediction mode.
11. The method according to any one of claims 8 or 9, wherein, The second prediction is obtained based on the intra-frame prediction mode.
12. The method according to any one of claims 8-11, wherein the method further comprises: The extrapolation filter is obtained based on the reconstruction template associated with the current block.
13. The method according to any one of claims 8-9, wherein the second prediction is at least one of the following: an inter-frame prediction signal obtained based on a merged candidate list, an intra-frame prediction signal obtained based on an intra-frame prediction mode (IPM) list and an IPM index of signaling, or an intra-frame prediction signal obtained based on a template-based intra-frame mode derivation (TIMD) result.
14. The method according to any one of claims 8-13, wherein the method further comprises: Determine the adjacent reconstruction regions of the current block and the filter shape of the current block; and The extrapolation filter coefficients associated with the current block are determined based on the adjacent reconstruction regions and the filter shape, wherein the second prediction associated with the current block is obtained based on the extrapolation filter coefficients.
15. A computer program product stored on a non-transitory computer-readable medium and comprising program code instructions that, when executed by a processor, are used to implement the steps of the method according to any one of claims 8 to 14.