Prediction and transform dependencies
By selecting appropriate candidate mode template metrics and filters in the video coding system, the prediction and coding process of video blocks is optimized, solving the problem of low coding efficiency in existing technologies and achieving more efficient video coding.
Patent Information
- Application Number
- CN202480024539.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-07
- Filing Date
- 2024-04-04
- Publication Date
- 2025-12-02
AI Technical Summary
Existing video coding systems struggle to effectively utilize transform types to select appropriate prediction and coding methods when processing video blocks, resulting in low coding efficiency.
Different candidate mode template metrics, such as sum of absolute differences (SAD) or sum of absolute transform differences (SATD), are selected based on the residual type of the video block, and filters such as spread filters, bilateral filters, or adaptive loop filters are applied to optimize the prediction and coding process.
It improves the efficiency and quality of video coding by selecting appropriate candidate mode template metrics and filters, optimizing the prediction and coding process of video blocks, and improving coding efficiency and image quality.
Smart Images

Figure CN121058239A_ABST
Abstract
Description
[0001] Cross-reference to related applications This application claims the benefit of European Provisional Patent Application No. 23305527.6, filed on 7 April 2023, the contents of which are hereby incorporated herein by reference. Background Technology
[0002] Video coding systems can be used to compress digital video signals, for example, to reduce the storage and / or transmission bandwidth required for such signals. Video coding systems can include, for example, block-based, wavelet-based, and / or object-based systems. Summary of the Invention
[0003] This paper describes systems, devices, and methods for addressing prediction and transform dependencies. Prediction of video blocks can be improved based on the type of transform used to encode them. For example, candidate metrics for deriving template-based tools can be selected based on the transform used for the video block. Smoothing steps can be added to the prediction (e.g., if a regular transform is used for the video block).
[0004] Devices such as video decoding devices can perform (e.g., are configured to perform) one or more of the following actions: The device can determine the residual type associated with a video block. The device can select a candidate mode template metric for the video block based on the associated residual type. The device can derive a candidate mode based on the selected candidate mode template metric. The device can decode the video block based on the derived candidate mode.
[0005] The candidate pattern template metric can be either the Sum of Absolute Transform Differences (SATD) or the Sum of Absolute Differences (SAD). For example, if the device determines that the residual type associated with a video block is transform skipped, the device can select the Sum of Absolute Differences (SAD) as the candidate pattern template metric for the video block. Similarly, if the device determines that the residual type associated with a video block is no residual, the device can select the Sum of Absolute Differences (SAD) as the candidate pattern template metric for the video block. Finally, if the device determines that the residual type associated with a video block is residual, the device can select the Sum of Absolute Transform Differences (SATD) as the candidate pattern template metric for the video block.
[0006] In the example, the candidate pattern template metric can be a template-based intra-frame pattern derivation (TIMD) metric, and the candidate pattern can be a TIMD candidate. In the example, the candidate pattern template metric can also be an intra-frame template matching prediction (IntraTMP) search metric, and the candidate pattern can be an IntraTMP candidate.
[0007] The device can obtain a predictor based on the derived candidate patterns. The device can also apply filters to the predictor associated with the derived candidate patterns based on an indication that the residual type is associated with the video block. The filters can be one of a spread filter, a bilateral filter, or an adaptive loop filter.
[0008] A video decoding method may include: determining a residual type associated with a video block. The method may include: selecting a candidate mode template metric for the video block based on the residual type. The method may include: deriving a candidate mode based on the selected candidate mode template metric. The method may include: decoding the video block based on the derived candidate mode.
[0009] In the example, if / when the transform is skipped or no residual is associated with the video block, the sum of absolute differences (SAD) can be selected as the candidate pattern template metric for the video block. In the example, if / when the residual is associated with the video block, the sum of absolute transform differences (SATD) can be selected as the candidate pattern template metric for the video block.
[0010] The method may include obtaining a predictor based on derived candidate patterns. The method may also include applying a filter to the predictor associated with the derived candidate patterns (e.g., indicating the association of residuals with video blocks based on a determined residual type). The filter may be a spread filter, a bilateral filter, or an adaptive loop filter.
[0011] Devices such as video encoding devices can perform (e.g., be configured to perform) one or more of the following actions: The device can derive a first candidate mode associated with a first candidate mode template metric based on a first residual type. The device can derive a second candidate mode associated with a second candidate mode template metric based on a second residual type. The device can select either the first or second candidate mode as a candidate mode and encode video blocks based on the selected candidate mode. For example, the device can select a candidate mode based on (e.g., rate distortion cost analysis of the first and second candidate modes).
[0012] The device can indicate the candidate pattern template metric in the residual type and / or video data (e.g., with encoded video blocks) associated with the selected candidate pattern.
[0013] In the example, the first candidate mode template metric can be the sum of absolute transform differences (SATD), and the second candidate mode template metric can be the sum of absolute differences (SAD). The device can obtain a first predictor based on the first candidate mode and a second predictor based on the second candidate mode. The device can apply filters to the first predictor associated with the first candidate mode. The filters can be spread filters, bilateral filters, or adaptive loop filters. The device can select candidate modes based on the filtered first and second predictors.
[0014] The device can select the sum of absolute transformation differences (SATD) as the first candidate pattern template metric based on determining the first residual type as the residual.
[0015] The device can select the sum of absolute differences (SAD) as the first candidate pattern template metric based on determining the second residual type as either no residual or transform skip.
[0016] A video coding method may include: deriving a first candidate mode associated with a first candidate mode template metric based on a first residual type; and deriving a second candidate mode associated with a second candidate mode template metric based on a second residual type. The method may include: selecting either the first or second candidate mode as a candidate mode (e.g., based on rate distortion cost analysis). The method may include: encoding video blocks based on the selected candidate mode.
[0017] In the example, the first candidate mode template metric may be the sum of absolute transform differences (SATD), and the second candidate mode template metric may be the sum of absolute differences (SAD). The method may include: obtaining a first predictor based on the first candidate mode; and obtaining a second predictor based on the second candidate mode. The method may include: (e.g., when the first candidate mode template metric is SATD) applying a filter to the first predictor associated with the first candidate mode. The candidate mode may be selected based on the filtered first and second predictors.
[0018] If / when the first residual type is determined to be a residual, the first candidate pattern template metric can be SATD.
[0019] If / when the second residual type is determined to be no residual or transformation skipped, the second candidate pattern template metric can be SAD.
[0020] The systems, methods, and instruments described herein may relate to a decoder. In some examples, the systems, methods, and instruments described herein may relate to an encoder. In some examples, the systems, methods, and instruments described herein may relate to signals (e.g., from an encoder and / or received by a decoder). Video data (e.g., a video bitstream) may include indications of video blocks / current blocks encoded according to intra-intra-CIIP. A computer-readable medium may include instructions for causing one or more processors to perform the methods described herein. A computer program product may include instructions that, when executed by one or more processors, cause the one or more processors to implement the methods described herein. Attached Figure Description
[0021] Figure 1A This is a system diagram illustrating an example communication system in which one or more of the disclosed embodiments may be implemented.
[0022] Figure 1B The illustration shows a device that can be used according to one embodiment. Figure 1A The diagram shows a system diagram of an example wireless transmit / receive unit (WTRU) used in a communication system.
[0023] Figure 1C The illustration shows a device that can be used according to one embodiment. Figure 1A The diagram shows a system diagram of an example radio access network (RAN) and an example core network (CN) used in a communication system.
[0024] Figure 1D The illustration shows a device that can be used according to one embodiment. Figure 1A The diagram shows a system diagram of a further example RAN and a further example CN used within the communication system.
[0025] Figure 2 This is a diagram illustrating an example block-based video encoder.
[0026] Figure 3 This is a diagram showing an example video decoder.
[0027] Figure 4 This is a diagram illustrating examples of systems in which various aspects and examples can be implemented.
[0028] Figure 5 The illustration shows an example of an intra-template matching search region used in IntraTMP (IntraTMP) prediction.
[0029] Figure 6 The illustration shows an example of choosing between two optimal predictors for IntraTMP.
[0030] Figure 7 Examples of the left and top template shapes are shown in the illustration.
[0031] Figure 8A The illustration shows an example template region and a reference sample used to derive template-based intra-frame mode derivation (TIMD) modes and associated weights.
[0032] Figure 8B The illustration shows an example reference sample used to derive intra-prediction for the current block.
[0033] Figure 9 The illustration shows an example evaluation of the intraTMP mode at the encoder, where the metric used to derive the candidate depends on the residual type.
[0034] Figure 10 The illustration shows an example derivation of the optimal intraTMP predictor at the decoder, which depends on the residual type.
[0035] Figure 11 The illustration shows an example derivation of the predictor, in which a smoothing step is added when using regular residuals.
[0036] Figure 12 The illustration shows an example derivation of the predictor at the decoder, with a smoothing step added when using regular residual coding.
[0037] Figure 13 An example derivation of the optimal intraTMP predictor is illustrated, in which a smoothing step is added when using regular residuals.
[0038] Figure 14 The illustration shows an example derivation of the intraTMP predictor at the decoder, with a smoothing step added when using regular residual coding. Detailed Implementation
[0039] A detailed description of illustrative embodiments will now be described with reference to various figures. While this description provides detailed examples of possible implementations, it should be noted that the details are intended to be exemplary and in no way limit the scope of the application.
[0040] Figure 1AThis diagram illustrates an example communication system 100 in which one or more of the disclosed embodiments may be implemented. The communication system 100 may be a multiple access system that provides content (such as voice, data, video, messaging, broadcasting, etc.) to multiple wireless users. The communication system 100 enables multiple wireless users to access such content by sharing system resources (including wireless bandwidth). For example, the communication system 100 may employ one or more channel access methods, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal FDMA (OFDMA), Single Carrier FDMA (SC-FDMA), Zero-Tail Unique Word DFT Extended OFDM (ZT UW DTS-s OFDM), Unique Word OFDM (UW-OFDM), Resource Block Filtered OFDM, Filter Bank Multicarrier (FBMC), etc.
[0041] like Figure 1A As shown, the communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RAN 104 / 113, CN 106 / 115, Public Switched Telephone Network (PSTN) 108, Internet 110, and other networks 112. Although it will be appreciated, the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, and 102d can be any type of device configured to operate and / or communicate in a wireless environment. As an example, WTRUs 102a, 102b, 102c, and 102d (any of which may be referred to as a “station” and / or “STA”) may be configured to transmit and / or receive wireless signals and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearable devices, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated processing chain scenarios), consumer electronics devices, devices operating on commercial and / or industrial wireless networks, etc. Any of WTRUs 102a, 102b, 102c, and 102d may be interchangeably referred to as a UE.
[0042] The communication system 100 may also include base station 114a and / or base station 114b. Each of base stations 114a and 114b may be any type of device configured to wirelessly interface with at least one of WTRUs 102a, 102b, 102c, and 102d to facilitate access to one or more communication networks, such as CN 106 / 115, the Internet 110, and / or other networks 112. As an example, base stations 114a and 114b may be any of a base transceiver station (BTS), Node-B, eNode B, home node B, home eNode B, gNB, NR NodeB, site controller, access point (AP), wireless router, etc. Although base stations 114a and 114b are depicted as single elements, it will be understood that base stations 114a and 114b may include any number of interconnected base stations and / or network elements.
[0043] Base station 114a may be part of RAN 104 / 113, which may also include other base stations and / or network elements (not shown), such as base station controllers (BSCs), radio network controllers (RNCs), relay nodes, etc. Base station 114a and / or base station 114b may be configured to transmit and / or receive radio signals on one or more carrier frequencies, which may be referred to as cells (not shown). These frequencies may be licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide coverage for a specific geographic area for a radio service, which may be relatively fixed or may change over time. A cell may be further divided into cell sectors. For example, the cell associated with base station 114a may be divided into three sectors. Therefore, in one embodiment, base station 114a may include three transceivers, i.e., one transceiver for each sector of the cell. In one embodiment, base station 114a may employ multiple-input multiple-output (MIMO) technology, and multiple transceivers may be used for each sector of the cell. For example, beamforming can be used to transmit and / or receive signals in a desired spatial direction.
[0044] Base stations 114a and 114b can communicate with one or more of WTRUs 102a, 102b, 102c, and 102d via air interface 116, which can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, millimeter wave, infrared (IR), ultraviolet (UV), visible light, etc.). Air interface 116 can be established using any suitable radio access technology (RAT).
[0045] More specifically, as noted above, communication system 100 can be a multiple access system and can employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, base station 114a in RAN 104 / 113, and WTRUs 102a, 102b, and 102c can implement radio technologies such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which can use Wideband CDMA (WCDMA) to establish air interfaces 115 / 116 / 117. WCDMA can include communication protocols such as High-Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA can include High-Speed Downlink (DL) Packet Access (HSDPA) and / or High-Speed UL Packet Access (HSUPA).
[0046] In one embodiment, base station 114a and WTRUs 102a, 102b, 102c can implement radio technologies such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which can use Long Term Evolution (LTE) and / or Advanced LTE (LTE-A) and / or Advanced LTE Pro (LTE-A Pro) to establish air interface 116.
[0047] In one embodiment, base station 114a and WTRUs 102a, 102b, 102c can implement radio technologies such as NR radio access, which can use a new radio (NR) to establish an air interface 116.
[0048] In one embodiment, base station 114a and WTRUs 102a, 102b, and 102c can implement multiple radio access technologies. For example, base station 114a and WTRUs 102a, 102b, and 102c can jointly implement LTE radio access and NR radio access, for example, using the dual connectivity (DC) principle. Therefore, the air interface utilized by WTRUs 102a, 102b, and 102c can be characterized by multiple types of radio access technologies and / or transmissions sent to / from multiple types of base stations (e.g., eNBs and gNBs).
[0049] In other embodiments, base station 114a and WTRUs 102a, 102b, 102c can implement the following radio technologies, such as IEEE 802.11 (i.e., WiFi), IEEE 802.16 (i.e., WiMAX), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Provisional Standard 2000 (IS-2000), Provisional Standard 95 (IS-95), Provisional Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rate GSM Evolution (EDGE), GSMEDGE (GERAN), etc.
[0050] Figure 1A Base station 114b can be, for example, a wireless router, a home node B, a home eNode B, or an access point, and can utilize any suitable RAT to facilitate wireless connectivity in a local area, such as a commercial area, home, vehicle, campus, industrial facility, air corridor (e.g., for drone use), road, etc. In one embodiment, base station 114b and WTRUs 102c, 102d can implement radio technologies such as IEEE 802.11 to establish a wireless local area network (WLAN). In one embodiment, base station 114b and WTRUs 102c, 102d can implement radio technologies such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, base station 114b and WTRUs 102c, 102d can utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a picocell or femtocell. Figure 1A As shown, base station 114b may have a direct connection to Internet 110. Therefore, base station 114b may not be required to access Internet 110 via CN 106 / 115.
[0051] RAN 104 / 113 can communicate with CN 106 / 115, which can be any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRUs 102a, 102b, 102c, and 102d. Data may have different Quality of Service (QoS) requirements, such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, etc. CN 106 / 115 can provide call control, billing services, location-based services, prepaid calling, internet connectivity, video distribution, and / or perform advanced security functions, such as user authentication. Although... Figure 1AAlthough not shown, it will be understood that RAN104 / 113 and / or CN106 / 115 can communicate directly or indirectly with other RANs that use the same RAT as or a different RAT than RAN 104 / 113. For example, in addition to being connected to RAN 104 / 113, which can utilize NR radio technology, CN106 / 115 can also communicate with another RAN (not shown) that uses GSM, UMTS, CDMA 2000, WiMAX, E-UTRA, or WiFi radio technology.
[0052] CN 106 / 115 may also act as a gateway for WTRU 102a, 102b, 102c, 102d to access PSTN 108, the Internet 110, and / or other networks 112. PSTN 108 may include a circuit-switched telephone network providing Common Old-Style Telephone Service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices using common communication protocols such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) from the TCP / IP Internet Protocol suite. Network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, network 112 may include another CN connected to one or more RANs, which may use the same RAT as RAN 104 / 113 or a different RAT.
[0053] Some or all of the WTRUs 102a, 102b, 102c, and 102d in communication system 100 may include multi-mode capabilities (e.g., WTRUs 102a, 102b, 102c, and 102d may include multiple transceivers for communicating with different wireless networks via different wireless links). For example, Figure 1A The WTRU 102c shown can be configured to communicate with a base station 114a that can use cellular-based radio technology and a base station 114b that can use IEEE 802 radio technology.
[0054] Figure 1B This is a system diagram illustrating example WTRU 102. (Example:) Figure 1B As shown, WTRU 102 may include a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power supply 134, a global positioning system (GPS) chipset 136, and / or other peripherals 138, etc. It will be appreciated that WTRU 102 may include any sub-combination of the above-described elements while remaining consistent with the embodiments.
[0055] Processor 118 may be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. Processor 118 may perform signal encoding, data processing, power control, input / output processing, and / or any other functions that enable WTRU 102 to operate in a wireless environment. Processor 118 may be coupled to transceiver 120, which may be coupled to transmitting / receiving element 122. Although... Figure 1B The processor 118 and transceiver 120 are depicted as separate components, but it will be understood that the processor 118 and transceiver 120 can be integrated together in an electronic package or chip.
[0056] Transmitting / receiving element 122 can be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) via air interface 116. For example, in one embodiment, transmitting / receiving element 122 can be an antenna configured to transmit and / or receive RF signals. In one embodiment, transmitting / receiving element 122 can be a transmitter / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, transmitting / receiving element 122 can be configured to transmit and / or receive both RF and optical signals. It will be appreciated that transmitting / receiving element 122 can be configured to transmit and / or receive any combination of wireless signals.
[0057] Although the transmitting / receiving element 122 is in Figure 1B While depicted as a single element, WTRU 102 may include any number of transmitting / receiving elements 122. More specifically, WTRU 102 may employ MIMO technology. Thus, in one embodiment, WTRU 102 may include two or more transmitting / receiving elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals via air interface 116.
[0058] Transceiver 120 can be configured to modulate signals to be transmitted by transmitting / receiving element 122 and demodulate signals received by transmitting / receiving element 122. As noted above, WTRU 102 can have multi-mode capability. Thus, for example, transceiver 120 may include multiple transceivers for enabling WTRU 102 to communicate via multiple RATs (such as NR and IEEE 802.11).
[0059] The processor 118 of WTRU 102 can be coupled to the speaker / microphone 124, keypad 126, and / or display / touchpad 128 (e.g., a liquid crystal display (LCD) unit or an organic light-emitting diode (OLED) display unit), and can receive user input data from them. The processor 118 can also output user data to the speaker / microphone 124, keypad 126, and / or display / touchpad 128. Additionally, the processor 118 can access information from any type of suitable memory (such as non-removable memory 130 and / or removable memory 132), and store data in that memory. Non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), hard disk, or any other type of memory storage device. Removable memory 132 may include a subscriber identity module (SIM) card, memory stick, secure digital storage (SD) card, etc. In other embodiments, processor 118 may access information from memory that is not physically located on WTRU 102 (such as on a server or home computer (not shown)) and store data in that memory.
[0060] The processor 118 can receive power from the power supply 134 and can be configured to distribute and / or control the power going to other components in the WTRU 102. The power supply 134 can be any suitable device for powering the WTRU 102. For example, the power supply 134 may include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.
[0061] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) about the current location of the WTRU 102. In addition to, or instead of, information from the GPS chipset 136, the WTRU 102 may receive location information from base stations (e.g., base stations 114a, 114b) via air interface 116 and / or determine its location based on the timing of signals received from two or more nearby base stations. It will be understood that the WTRU 102 may acquire location information using any suitable location determination method, while remaining consistent with the embodiments.
[0062] The processor 118 may be further coupled to other peripherals 138, which may include one or more software and / or hardware modules providing additional features, functions, and / or wired or wireless connectivity. For example, peripherals 138 may include accelerometers, electronic compasses, satellite transceivers, digital cameras (for photos and / or videos), Universal Serial Bus (USB) ports, vibration devices, television transceivers, hands-free headsets, Bluetooth® modules, FM radio units, digital music players, media players, video game player modules, internet browsers, virtual reality and / or augmented reality (VR / AR) devices, activity trackers, etc. Peripherals 138 may include one or more sensors, which may be one or more of the following: gyroscopes, accelerometers, Hall effect sensors, magnetometers, orientation sensors, proximity sensors, temperature sensors, time sensors, geolocation sensors, altimeters, light sensors, touch sensors, magnetometers, barometers, gesture sensors, biometric sensors, and / or humidity sensors.
[0063] WTRU 102 may include a full-duplex radio, for which the transmission and reception of some or all signals (e.g., associated with specific subframes for both UL (e.g., for transmission) and downlink (e.g., for reception)) may be concurrent and / or simultaneous. The full-duplex radio may include an interference management unit to reduce and / or substantially eliminate self-interference via hardware (e.g., a choke) or via signal processing (e.g., a separate processor (not shown) or via processor 118). In one embodiment, WTRU 102 may include a half-duplex radio, for which the transmission and reception of some or all signals (e.g., associated with specific subframes for UL (e.g., for transmission) or downlink (e.g., for reception)) may be concurrent and / or simultaneous.
[0064] Figure 1C The diagram illustrates a system diagram of RAN 104 and CN 106 according to an embodiment. As noted above, RAN 104 can employ E-UTRA radio technology to communicate with WTRUs 102a, 102b, and 102c via air interface 116. RAN 104 can also communicate with CN 106.
[0065] RAN 104 may include eNode-Bs 160a, 160b, and 160c, although it will be understood that RAN 104 may include any number of eNode-Bs while remaining consistent with the embodiments. eNode-Bs 160a, 160b, and 160c may each include one or more transceivers for communicating with WTRUs 102a, 102b, and 102c via air interface 116. In one embodiment, eNode-Bs 160a, 160b, and 160c may implement MIMO technology. Therefore, eNode-B 160a may, for example, use multiple antennas to transmit radio signals to and / or receive radio signals from WTRU 102a.
[0066] Each of the eNode-B 160a, 160b, and 160c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, and user scheduling in the UL and / or DL, etc. Figure 1C As shown, eNode-B 160a, 160b, and 160c can communicate with each other via the X2 interface.
[0067] Figure 1C The CN 106 shown may include a Mobility Management Entity (MME) 162, a Serving Gateway (SGW) 164, and a Packet Data Network (PDN) Gateway (or PGW) 166. Although each of the foregoing elements is depicted as part of CN 106, it will be understood that any of these elements may be owned and / or operated by an entity other than a CN operator.
[0068] The MME 162 can connect to each of the eNode-Bs 162a, 162b, and 162c in RAN104 via the S1 interface and can act as a control node. For example, the MME 162 can be responsible for authenticating users of WTRUs 102a, 102b, and 102c, activating / deactivating bearers, and selecting a specific serving gateway during the initial attachment of WTRUs 102a, 102b, and 102c. The MME 162 can provide control plane functions for handover between RAN104 and other RANs (not shown) employing other radio technologies, such as GSM and / or WCDMA.
[0069] The SGW 164 can connect to each of the eNode Bs 160a, 160b, and 160c in RAN104 via the S1 interface. The SGW 164 can typically route and forward user data packets to / from WTRUs 102a, 102b, and 102c. The SGW 164 can perform other functions, such as anchoring the user plane during inter-eNode B handover, triggering paging when DL data is available for WTRUs 102a, 102b, and 102c, and managing and storing the context of WTRUs 102a, 102b, and 102c.
[0070] SGW 164 can be connected to PGW 166, which can provide WTRU 102a, 102b, 102c with access to packet-switched networks (such as Internet 110) to facilitate communication between WTRU 102a, 102b, 102c and IP-enabled devices.
[0071] CN 106 facilitates communication with other networks. For example, CN 106 can provide WTRUs 102a, 102b, and 102c with access to a circuit-switched network (such as PSTN 108) to facilitate communication between WTRUs 102a, 102b, and 102c and conventional terrestrial line communication equipment. For example, CN 106 may include or communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between CN 106 and PSTN 108. Additionally, CN 106 can provide WTRUs 102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.
[0072] Despite WTRU in Figures 1A to 1D While described as a wireless terminal, it is envisioned that, in some representative embodiments, such a terminal may use (e.g., temporarily or permanently) a wired communication interface with a communication network.
[0073] In a representative embodiment, the other network 112 may be a WLAN.
[0074] In an Infrastructure Basic Services Set (BSS) mode, a WLAN may have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP may have access or an interface to a distribution system (DS) or another type of wired / wireless network that carries traffic into and / or out of the BSS. Traffic originating outside the BSS destined for a STA can be delivered to the STA via the AP. Traffic from a STA to a destination outside the BSS can be sent to the AP for delivery to the appropriate destination. Traffic between STAs within the BSS can be sent via the AP, for example, where a source STA can send traffic to the AP, and the AP can deliver the traffic to the destination STA. Traffic between STAs within the BSS can be considered and / or referred to as peer-to-peer traffic. Peer-to-peer traffic can be sent between source and destination STAs (e.g., directly between them) using a direct link setup (DLS). In some representative embodiments, the DLS may use 802.11e DLS or 802.11z Tunneled DLS (TDLS). A WLAN using the Standalone BSS (IBSS) mode may not have an access point (AP), and STAs within or using the IBSS (e.g., all STAs) can communicate directly with each other. The IBSS communication mode may sometimes be referred to as a "self-organizing" communication mode in this document.
[0075] When using 802.11ac infrastructure operation mode or a similar operation mode, the AP can transmit beacons on a fixed channel, such as a primary channel. The primary channel can be of fixed width (e.g., a bandwidth of 20 MHz) or dynamically set via signaling. The primary channel can be the operating channel of the BSS and can be used by STAs to establish connections with the AP. In some representative embodiments, Carrier Sense Multiple Access with Collision Avoidance (CSMA / CA) can be implemented, for example, in an 802.11 system. For CSMA / CA, STAs including the AP (e.g., each STA) can sense the primary channel. If the primary channel is sensed / detected and / or determined to be busy by a particular STA, that STA can back off. A STA (e.g., only one station) can transmit at any given time within a given BSS.
[0076] High-throughput (HT) STAs can communicate using a 40MHz wide channel, for example, by combining a primary 20MHz channel with adjacent or non-adjacent 20MHz channels to form a 40MHz wide channel.
[0077] Very High Throughput (VHT) STAs can support channels with widths of 20MHz, 40MHz, 80MHz, and / or 160MHz. 40MHz and / or 80MHz channels can be formed by combining consecutive 20MHz channels. A 160MHz channel can be formed by combining eight consecutive 20MHz channels or by combining two non-consecutive 80MHz channels (which can be referred to as an 80+80 configuration). For the 80+80 configuration, after channel coding, data is transmitted via a segment resolver that divides the data into two streams. Inverse Fast Fourier Transform (IFFT) processing and time-domain processing can be performed separately on each stream. The streams can be mapped onto two 80MHz channels, and the data can be transmitted by the transmitting STA. At the receiver of the receiving STA, the above operations for the 80+80 configuration can be reversed, and the combined data can be sent to the Media Access Control (MAC).
[0078] Operating modes below 1 GHz are supported by 802.11af and 802.11ah. The channel operating bandwidth and carrier are reduced in 802.11af and 802.11ah compared to those used in 802.11n and 802.11ac. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV white space (TVWS) spectrum, while 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. According to representative embodiments, 802.11ah can support instrument-type control / machine-type communications, such as MTC devices in macro coverage areas. MTC devices may have certain capabilities, such as limited capabilities, including support for (e.g., only support) certain and / or limited bandwidths. MTC devices may include batteries with a lifespan exceeding a threshold (e.g., to maintain a very long battery life).
[0079] WLAN systems that can support multiple channels and channel bandwidths (such as 802.11n, 802.11ac, 802.11af, and 802.11ah) include a channel that can be designated as the primary channel. The primary channel can have a bandwidth equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel can be set and / or limited by the STA that supports the minimum bandwidth operating mode among all STAs operating in the BSS. In the 802.11ah example, for STAs that support (e.g., only support) the 1MHz mode (e.g., MTC type devices), the primary channel can be 1MHz wide, even if the AP and other STAs in the BSS support 2MHz, 4MHz, 8MHz, 16MHz, and / or other channel bandwidth operating modes. Carrier sensing and / or Network Allocation Vector (NAV) settings can depend on the status of the primary channel. If the primary channel is busy, for example because an STA (which only supports the 1MHz operating mode) is transmitting to the AP, the entire available band may be considered busy, even if most of the band is still idle and could be available.
[0080] In the United States, the available frequency band for 802.11ah is from 902MHz to 928MHz. In South Korea, the available frequency band is from 917.5MHz to 923.5MHz. In Japan, the available frequency band is from 916.5MHz to 927.5MHz. Depending on the country code, the total available bandwidth for 802.11ah is 6MHz to 26MHz.
[0081] Figure 1D The diagram illustrates a system diagram of RAN 113 and CN 115 according to an embodiment. As noted above, RAN 113 may employ NR radio technology to communicate with WTRUs 102a, 102b, and 102c via air interface 116. RAN 113 may also communicate with CN 115.
[0082] RAN 113 may include gNBs 180a, 180b, and 180c, although it will be understood that RAN 113 may include any number of gNBs while remaining consistent with the embodiments. gNBs 180a, 180b, and 180c may each include one or more transceivers for communicating with WTRUs 102a, 102b, and 102c via air interface 116. In one embodiment, gNBs 180a, 180b, and 180c may implement MIMO technology. For example, gNBs 180a and 180b may utilize beamforming to transmit signals to and / or receive signals from gNBs 180a, 180b, and 180c. Therefore, gNB 180a may, for example, use multiple antennas to transmit radio signals to and / or receive radio signals from WTRU 102a. In one embodiment, gNBs 180a, 180b, and 180c can implement carrier aggregation technology. For example, gNB 180a can transmit multiple component carriers to WTRU 102a (not shown). A subset of these component carriers can be on unlicensed spectrum, while the remaining component carriers can be on licensed spectrum. In one embodiment, gNBs 180a, 180b, and 180c can implement Cooperative Multipoint (CoMP) technology. For example, WTRU 102a can receive cooperative transmissions from gNBs 180a and 180b (and / or gNB 180c).
[0083] WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using transmissions associated with scalable digitization. For example, OFDM symbol spacing and / or OFDM subcarrier spacing can vary depending on different transmissions, different cells, and / or different portions of the radio transmission spectrum. WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using subframes of various or scalable lengths or transmission time intervals (TTIs) (e.g., containing different numbers of OFDM symbols and / or absolute times of varying durations).
[0084] gNBs 180a, 180b, and 180c can be configured to communicate with WTRUs 102a, 102b, and 102c in standalone and / or non-standalone configurations. In standalone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c without also accessing other RANs (e.g., eNode-Bs 160a, 160b, and 160c). In standalone configuration, WTRUs 102a, 102b, and 102c can utilize one or more of gNBs 180a, 180b, and 180c as mobility anchors. In standalone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using signals in unlicensed frequency bands. In a non-standalone configuration, WTRUs 102a, 102b, and 102c can communicate with / connect to gNBs 180a, 180b, and 180c, and simultaneously communicate with / connect to another RAN (such as eNode-Bs 160a, 160b, and 160c). For example, WTRUs 102a, 102b, and 102c can implement DC principles to communicate substantially simultaneously with one or more gNBs 180a, 180b, and 180c and one or more eNode-Bs 160a, 160b, and 160c. In a non-standalone configuration, eNode-B 160a, 160b, and 160c can act as mobility anchors for WTRU 102a, 102b, and 102c, and gNB 180a, 180b, and 180c can provide additional coverage and / or throughput to serve WTRU 102a, 102b, and 102c.
[0085] Each of gNBs 180a, 180b, and 180c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, network slicing support, dual connectivity, interoperability between NR and E-UTRA, routing of user plane data to User Plane Functions (UPF) 184a and 184b, routing of control plane information to Access and Mobility Management Functions (AMF) 182a and 182b, etc. Figure 1D As shown, gNB 180a, 180b, and 180c can communicate with each other via the Xn interface.
[0086] Figure 1DThe CN 115 shown may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and possibly a Data Network (DN) 185a, 185b. Although each of the foregoing elements is depicted as part of the CN 115, it will be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.
[0087] AMF 182a and 182b can connect to one or more of the gNBs 180a, 180b, and 180c in RAN 113 via the N2 interface and can act as control nodes. For example, AMF 182a and 182b can be responsible for authenticating users of WTRU 102a, 102b, and 102c, supporting network slicing (e.g., handling different PDU sessions with different requirements), selecting specific SMF183a and 183b, managing registration areas, terminating NAS signaling, mobility management, etc. Network slices can be used by AMF182a and 182b to customize CN support for WTRU 102a, 102b, and 102c based on the service types utilized by WTRU 102a, 102b, and 102c. For example, different network slices can be built for different use cases, such as services that rely on Ultra Reliable Low Latency (URLLC) access, services that rely on Enhanced Massive Mobile Broadband (eMBB) access, and services for Machine Type Communication (MTC) access. AMF 162 can provide control plane functions for handover between RAN 113 and other RANs (not shown) employing other radio technologies, such as LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies, such as WiFi.
[0088] SMFs 183a and 183b can connect to AMFs 182a and 182b in CN 115 via the N11 interface. SMFs 183a and 183b can also connect to UPFs 184a and 184b in CN 115 via the N4 interface. SMFs 183a and 183b can select and control UPFs 184a and 184b, and configure traffic routing through UPFs 184a and 184b. SMFs 183a and 183b can perform other functions, such as managing and allocating UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, and providing downlink data notifications. PDU session types can be IP-based, non-IP-based, Ethernet-based, etc.
[0089] UPF 184a and 184b can connect to one or more of gNBs 180a, 180b, and 180c in RAN 113 via the N3 interface. These gNBs can provide WTRU 102a, 102b, and 102c with access to a packet-switched network (such as the Internet 110) to facilitate communication between WTRU 102a, 102b, 102c and IP-enabled devices. UPF 184a and 184b can perform other functions such as routing and forwarding packets, enforcing user plane policies, supporting multi-homed PDU sessions, handling user plane QoS, buffering downlink packets, and providing mobility anchoring.
[0090] CN 115 can facilitate communication with other networks. For example, CN 115 may include or communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between CN 115 and PSTN 108. Additionally, CN 115 can provide WTRUs 102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, WTRUs 102a, 102b, and 102c can connect to local data networks (DNs) 185a and 185b via the N3 interface to UPFs 184a and 184b and the N6 interface between UPFs 184a and 184b and DNs 185a and 185b.
[0091] Given Figures 1A to 1D as well as Figures 1A to 1D The corresponding descriptions may be performed by one or more emulation devices (not shown) to perform one or more of the functions described herein with respect to one or more of the following: WTRU102a-d, base station 114a-b, eNode-B160a-c, MME 162, SGW 164, PGW 166, gNB180a-c, AMF 182a-b, UPF 184a-b, SMF 183a-b, DN185a-b, and / or one or more other devices described herein. An emulation device may be one or more devices configured to emulate one or more of the functions described herein. For example, an emulation device may be used to test other devices and / or simulate network and / or WTRU functions.
[0092] Simulation devices can be designed to perform tests on one or more other devices in a laboratory environment and / or a carrier network environment. For example, the one or more simulation devices can perform one or more or all of their functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices within the communication network. The one or more simulation devices can perform one or more or all of their functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. Simulation devices can be directly coupled to another device for testing purposes and / or can use over-the-air wireless communication to perform tests.
[0093] The one or more simulation devices can perform one or more (including all) functions without being implemented / deployed as part of a wired and / or wireless communication network. For example, the simulation devices can be used in test scenarios in a test laboratory and / or in non-deployed (e.g., testing) wired and / or wireless communication networks to perform testing on one or more components. The one or more simulation devices can be test rigs. Direct RF coupling and / or wireless communication via RF circuitry (e.g., which may include one or more antennas) can be used by the simulation devices to transmit and / or receive data.
[0094] This application describes various aspects, including tools, features, examples or embodiments, models, schemes, etc. Many of these aspects are described in detail and are often described in a manner that may sound limiting, at least to illustrate individual characteristics. However, this is for clarity of purpose and does not limit the application or scope of those aspects. Indeed, all the different aspects can be combined and interchanged to provide further aspects. Furthermore, this aspect can also be combined and interchanged with aspects described in earlier filings.
[0095] The aspects described and conceived in this application can be implemented in many different forms. Figure 5-14 Some embodiments may be provided, but other embodiments are conceived. Figure 5-14 The discussion does not limit the breadth of implementation methods. At least one aspect generally relates to video encoding and decoding, and at least one other aspect generally relates to the transmission of a generated or encoded bitstream. These and other aspects can be implemented as methods, apparatus, computer-readable storage media having instructions stored thereon for encoding or decoding video data according to any of the described methods, and / or computer-readable storage media having bitstreams generated according to any of the described methods stored thereon.
[0096] In this application, the terms “reconstructed” and “decoded” may be used interchangeably, as may the terms “pixel” and “sample”, and may the terms “image”, “picture” and “frame” be used interchangeably.
[0097] The terms HDR (High Dynamic Range) and SDR (Standard Dynamic Range) may be used in this disclosure. These terms often convey specific values of dynamic range to those skilled in the art. However, additional embodiments are contemplated where references to HDR are understood to mean “higher dynamic range” and references to SDR are understood to mean “lower dynamic range.” Such additional embodiments are not constrained by any specific values of dynamic range that may often be associated with the terms “high dynamic range” and “standard dynamic range.”
[0098] This document describes various methods, and each method includes one or more steps or actions for implementing the described method. Unless a specific order of steps or actions is required for the proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. Additionally, terms such as "first," "second," etc., may be used in various embodiments to modify elements, components, steps, operations, etc., such as, for example, "first decoding" and "second decoding." The use of such terms does not imply a sequence of modified operations unless specifically required. Thus, in this example, the first decoding need not be performed before the second decoding, but may occur, for example, before, during, or in the time period overlapping with the second decoding.
[0099] The various methods and other aspects described in this application can be used to modify, respectively, such as Figure 2 and Figure 3 The modules of the video encoder 200 and decoder 300 shown herein include, for example, intra-frame prediction and entropy coding and / or decoding modules (260, 360, 245, 330). Furthermore, the subject matter disclosed herein presents aspects not limited to VVC or HEVC and can be applied to, for example, any type, format, or version of video coding, whether or not described in a standard or recommendation, whether pre-existing or future-developed, and any extensions to such standards and recommendations (e.g., including VVC and HEVC). Unless otherwise indicated or technically excluded, the aspects described herein may be used individually or in combination.
[0100] Various numerical values, such as coefficients, block sizes, etc., are used in the examples described in this application. These and other specific values are used for the purpose of describing the examples, and the aspects described are not limited to these specific values.
[0101] Figure 2This is a diagram illustrating an example video encoder (e.g., an example block-based hybrid video encoder) 200. Variations of the example encoder 200 are conceivable, but the encoder 200 is described below for clarity without depicting all expected variations.
[0102] Before being encoded, the video sequence may undergo pre-encoding processing (201), such as applying color transformations to the input color image (e.g., a conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing remapping of the input image components to obtain a more resilient signal distribution for compression (e.g., using histogram equalization with one of the color components). Metadata may be associated with preprocessing and attached to the bitstream.
[0103] In encoder 200, the image is encoded by encoder elements as described below. The image to be encoded is partitioned (202) and processed in units such as coding units (CUs). Each unit is encoded using, for example, an intra-frame or inter-frame mode. When a unit is encoded in an intra-frame mode, it performs intra-frame prediction (260). In an inter-frame mode, motion estimation (275) and compensation (270) are performed. The encoder determines (205) which of the intra-frame or inter-frame modes to use for encoding the unit and indicates the intra-frame / inter-frame decision by, for example, a prediction mode flag. The prediction residual is calculated, for example, by subtracting (210) the prediction block from the original image block.
[0104] The predicted residual is then transformed (225) and quantized (230). The quantized transform coefficients, motion vectors, and other syntax elements are entropy encoded (245) to output a bit stream. The encoder can skip the transform and apply the quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., encode the residual directly without applying the transform or quantization process.
[0105] The encoder decodes the encoded blocks to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residuals. The decoded prediction residuals and the predicted blocks are combined (255) to reconstruct the image blocks. A loop filter (265) is applied to the reconstructed image to perform, for example, deblocking / SAO (sample adaptive offset) filtering to reduce coding artifacts. The filtered image is stored in a reference image buffer (280).
[0106] Figure 3 This is a diagram illustrating an example of a video decoder 300. In the example decoder 300, the bitstream is decoded by decoder elements, as described below. The video decoder 300 generally performs operations similar to... Figure 2The encoding passes described herein are mutually decoded passes. Encoder 200 generally also performs video decoding as part of the encoding of video data.
[0107] Specifically, the input to the decoder includes a video bitstream, which can be generated by the video encoder 200. The bitstream is first entropy-decoded (330) to obtain transform coefficients, motion vectors, and other encoded information. Image partitioning information indicates how the image should be partitioned. Therefore, the decoder can partition (335) the image based on the decoded image partitioning information. The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residuals. The decoded prediction residuals and the predicted blocks are combined (355) to reconstruct the image blocks. The predicted blocks (370) can be obtained from intra-frame prediction (360) or motion-compensated prediction (i.e., inter-frame prediction) (375). A loop filter (365) is applied to the reconstructed image. The filtered image is stored at a reference image buffer (380). For a given image, the contents of the reference image buffer 380 on the decoder 300 can be the same as the contents of the reference image buffer 280 on the encoder 200 side for the same image.
[0108] The decoded image can undergo further post-decoding processing (385), such as inverse color transformation (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4), or inverse remapping, which is the inverse operation of the remapping process performed in pre-encoding processing (201). Post-decoding processing can utilize metadata derived in pre-encoding processing and signaled in the bitstream.
[0109] Figure 4 This is a diagram illustrating examples of systems in which the various aspects and embodiments described herein can be implemented. System 400 may be embodied as a device including the various components described below and configured to perform one or more of the aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. The elements of system 400 may be embodied individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one example, the processing and encoder / decoder elements of system 400 are distributed across multiple ICs and / or discrete components. In various embodiments, system 400 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 400 is configured to implement one or more of the aspects described in this document.
[0110] System 400 includes: at least one processor 410 configured to execute instructions loaded therein for implementing various aspects, such as those described in this document. Processor 410 may include embedded memory, input / output interfaces, and various other circuitry as known in the art. System 400 includes at least one memory 420 (e.g., a volatile memory device and / or a non-volatile memory device). System 400 includes: a storage device 440 which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. Storage device 440 may include internal storage devices, attached storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices, as non-limiting examples.
[0111] System 400 includes an encoder / decoder module 430 configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 430 may include its own processor and memory. The encoder / decoder module 430 represents one or more modules that may be included in a device to perform encoding and / or decoding functions. It is well known that a device may include one or both encoding and decoding modules. Furthermore, the encoder / decoder module 430 may be implemented as a separate element of system 400, or may be incorporated into processor 410 as a combination of hardware and software as known to those skilled in the art.
[0112] Program code to be loaded onto processor 410 or encoder / decoder 430 to execute the various aspects described herein may be stored in storage device 440 and subsequently loaded onto memory 420 for execution by processor 410. According to various embodiments, one or more of processor 410, memory 420, storage device 440, and encoder / decoder module 430 may store one or more items of various kinds during the execution of the processes described herein. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from processing of equations, formulas, operations, and operational logic.
[0113] In some embodiments, the memory within processor 410 and / or encoder / decoder module 430 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, external memory (e.g., processor 410 or encoder / decoder module 430) is used for one or more of these functions. External memory may be memory 420 and / or storage device 440, such as volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one embodiment, a fast external dynamic volatile memory, such as RAM, is used as working memory for video encoding and decoding operations, such as, for example, MPEG-2 (MPEG stands for Moving Picture Experts Group; MPEG-2 is also known as ISO / IEC 13818, and 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC stands for High Efficiency Video Coding; also known as H.265 and MPEG-H Part 2), or VVC (Universal Video Coding: a new standard developed by the Joint Video Experts Team JVET)).
[0114] Input to the components of system 400 can be provided through various input devices as indicated in box 445. Such input devices include, but are not limited to: (i) a radio frequency (RF) section that receives RF signals transmitted over the air, for example, by a broadcaster; (ii) component (COMP) input terminals (or a collection of COMP input terminals); (iii) a universal serial bus (USB) input terminal; and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Figure 4 Other examples not shown include composite video.
[0115] In various embodiments, the input device of block 445 has associated corresponding input processing elements as known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also referred to as selecting a signal or limiting a signal to a frequency band); (ii) down-converting the selected signal; (iii) further limiting the frequency band to a narrower band to select, for example, a signal band that may be referred to as a channel in some embodiments; (iv) demodulating the down-converted and band-limited signal; (v) performing error correction; and (vi) demultiplexing to select the desired stream of data packets. The RF section in various embodiments includes one or more elements to perform these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or to baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted over a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and filtering again to the desired frequency band. Various embodiments rearrange the order of the components described above (and others), remove some of these components, and / or add other components that perform similar or different functions. Adding components may include inserting components between existing components, such as, for example, inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.
[0116] Additionally, the USB and / or HDMI endpoints may include corresponding interface processors for connecting system 400 to other electronic devices across USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or, if necessary, within processor 410. Similarly, aspects of USB or HDMI interface processing may be implemented, either within a separate interface IC or, if necessary, within processor 410. The demodulated, error-corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 410 and encoder / decoder 430, which operate in conjunction with memory and storage elements to process the data stream as needed for presentation on an output device.
[0117] Various components of the system 400 can be provided within an integrated housing, in which the various components can be interconnected and data can be transmitted therebetween using a suitable connection arrangement 425, such as an internal bus as known in the art, including inter-IC (I2C) bus, wiring and printed circuit board.
[0118] System 400 includes a communication interface 450 that enables communication with other devices via a communication channel 460. The communication interface 450 may include, but is not limited to, a transceiver configured to transmit and receive data on the communication channel 460. The communication interface 450 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 460 may be implemented over, for example, wired and / or wireless media.
[0119] In various embodiments, data is streamed or otherwise provided to system 400 using a wireless network such as WiFi (e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers)). In these examples, the Wi-Fi signal is received on a communication channel 460 and a communication interface 450 adapted for Wi-Fi communication. The communication channel 460 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box that delivers data over an HDMI connection in input box 445 to provide streaming data to system 400. Still other embodiments use an RF connection in input box 445 to provide streaming data to system 400. As indicated above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.
[0120] System 400 can provide output signals to various output devices, including a display 475, a speaker 485, and other peripheral devices 495. The display 475 in various embodiments includes one or more of the following: for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 475 can be used in televisions, tablets, laptops, cellular phones (mobile phones), or other devices. The display 475 can also be integrated with other components (e.g., as in a smartphone) or separate (e.g., an external monitor for a laptop). In various examples of embodiments, other peripheral devices 495 include one or more of a standalone digital video disc (or digital multifunction disc) (DVR, for both terms), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 495 that provide functionality based on the output of system 400. For example, a disc player performs the function of playing the output of system 400.
[0121] In various embodiments, signaling such as AV is used to transmit control signals between system 400 and display 475, speaker 485, or other peripheral devices 495. Device-to-device control links, consumer electronics control (CEC), or other communication protocols are implemented with or without user intervention. Output devices can be communicatively coupled to system 400 via dedicated connections through corresponding interfaces 470, 480, and 490. Alternatively, output devices can be connected to system 400 via communication interface 450 using communication channel 460. Display 475 and speaker 485 can be integrated into a single unit with other components of system 400 in electronic devices such as, for example, televisions. In various embodiments, display interface 470 includes display drivers, such as, for example, a timing controller (TCon) chip.
[0122] Display 475 and speaker 485 can alternatively be separated from one or more other components, for example, if the RF section of input 445 is part of a separate set-top box. In various embodiments where display 475 and speaker 485 are external components, output signals can be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.
[0123] The embodiments may be implemented by computer software, hardware, or a combination of hardware and software, implemented by processor 410. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. As a non-limiting example, memory 420 may be of any type suitable for the technical environment and may be implemented using any suitable data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. As a non-limiting example, processor 410 may be of any type suitable for the technical environment and may encompass one or more of microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.
[0124] Various implementations involve decoding. As used herein, “decoding” can encompass all or part of a process performed, for example, on a received encoded sequence to produce a final output suitable for display. In various embodiments, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, dequantization, inverse transform, and differential decoding. In various embodiments, such a process may also or alternatively include processes performed by a decoder of the various implementations described herein, such as: determining a transform mode associated with the current block; selecting a candidate mode template metric for the current block from a plurality of candidate mode template metrics based on the transform mode; deriving a candidate mode based on the candidate mode template metric; and / or decoding the block based on the candidate mode.
[0125] As a further example, in one example, "decoding" refers only to entropy decoding; in another embodiment, "decoding" refers only to differential decoding; and in yet another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally to a broader decoding process will be clear based on the specific context of the description and is considered to be well understood by those skilled in the art.
[0126] Various implementations involve encoding. In a manner similar to the above discussion of “decoding,” the term “encoding,” as used herein, can encompass all or part of a process performed, for example, on an input video sequence to produce an encoded bitstream. In various embodiments, such a process includes one or more processes typically performed by an encoder, such as partitioning, differential coding, transform, quantization, and entropy coding. In various embodiments, such a process also, or alternatively, includes processes performed by an encoder of the various implementations described herein, such as: deriving a first candidate mode based on a first candidate mode template metric associated with a first candidate mode; deriving a second candidate mode based on a second candidate mode template metric associated with a second candidate mode; selecting a candidate mode from the first and second candidate modes; and including in the video data an indication of a transform mode associated with the selected candidate mode.
[0127] As a further example, in one embodiment, “encoding” refers only to entropy encoding; in another embodiment, “encoding” refers only to differential encoding; and in yet another embodiment, “encoding” refers to a combination of differential and entropy encoding. Whether the phrase “encoding process” is intended to specifically refer to a subset of operations or generally to a broader encoding process will be clear based on the specific context of the description and is considered well understood by those skilled in the art.
[0128] When a diagram is presented as a flowchart, it should be understood that it also provides a block diagram of the corresponding device. Similarly, when a diagram is presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / process.
[0129] Various embodiments mention rate distortion optimization. Specifically, during the encoding process, a balance or trade-off between rate and distortion is typically considered, often due to computational complexity constraints. Rate distortion optimization is generally formatted as minimizing a rate distortion function, which is a weighted sum of rate and distortion. Different schemes exist for solving the rate distortion optimization problem. For example, a scheme can be based on extensive testing of all encoding options, including all considered modes or encoding parameter values, with a complete evaluation of their encoding costs and associated distortions for the reconstructed signal after encoding and decoding. Faster schemes can also be used to save encoding complexity, particularly regarding the computation of approximate distortions based on predictions or predictions of residual signals rather than the reconstructed signal. A hybrid of these two schemes can also be used, such as by using approximate distortions for only some of the possible encoding options and full distortions for others. Other schemes evaluate only a subset of the possible encoding options. More generally, many schemes employ any of a variety of techniques to perform optimization, but optimization is not necessarily a complete evaluation of both encoding costs and associated distortions.
[0130] The implementations and aspects described herein can be implemented, for example, in methods or processes, apparatus, software programs, data streams, or signals. Even if discussed only in the context of a single form of implementation (e.g., discussed only as a method), the features in question can be implemented in other forms (e.g., apparatus or program). Apparatus can be implemented, for example, in appropriate hardware, software, and firmware. Methods can be implemented, for example, in a processor, where processor generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as, for example, computers, cellular phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate communication of information between end users.
[0131] References to “an embodiment,” “an example,” “an implementation,” or “an implementation,” and other variations thereof, mean that a particular feature, structure, characteristic, etc., described in connection with an embodiment is included in at least one embodiment. Therefore, the phrases “in an embodiment,” “in an example,” “in an implementation,” or “in an implementation,” and any variations appearing throughout this application, do not necessarily all refer to the same embodiment or example.
[0132] Additionally, this application may refer to "determining" various pieces of information. Determining information may include one or more of the following: estimated information, calculated information, predicted information, or information retrieved from memory. Obtaining may include receiving, retrieving, constructing, generating, and / or determining.
[0133] Furthermore, this application may refer to "accessing" various pieces of information. Accessing information may include one or more of the following: receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0134] Additionally, this application may refer to "receiving" various pieces of information. As with "access," "receiving" is intended to be a broad term. Receiving information may include one or more of, for example, accessing information or retrieving information (e.g., from memory). Further, "receiving" typically refers to actions performed during operation, such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0135] It should be understood that the use of any of the following “ / ”, “and / or”, and “…at least one of” (e.g., in the cases of “A / B”, “A and / or B”, and “at least one of A and B”) is intended to cover the selection of only the first listed option (A), or only the second listed option (B), or the selection of both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C”, this phrase is intended to cover the selection of only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or the selection of all three options (A, B, and C). This can be extended to as many items as are listed, as will be clear to those skilled in the art and related fields.
[0136] Moreover, as used herein, the term “signaling” refers, among other things, to instructing the corresponding decoder to do something. In this way, in embodiments, the same parameters are used on both the encoder and decoder sides. Thus, for example, the encoder can transmit specific parameters (explicit signaling) to the decoder so that the decoder can use the same specific parameters. Conversely, if the decoder already has the specific parameters as well as other parameters, signaling can be used without transmission (implicit signaling) to simply allow the decoder to know and select the specific parameters. Bit saving is achieved in various embodiments by avoiding the transmission of any actual functionality. It should be understood that signaling can be accomplished in a variety of ways. For example, in various embodiments, one or more syntax elements, tags, etc., are used to signal information to the corresponding decoder. Although the signature refers to the verb form of the term “signaling,” the term “signaling” may also be used as a noun herein.
[0137] As will be apparent to those skilled in the art, implementations can generate various signals formatted to carry, for example, information that can be stored or transmitted. The information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, the signal may be formatted to carry a bit stream of the described embodiments. Such a signal may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of a spectrum) or baseband signals. Formatting may include, for example, encoding the data stream and modulating a carrier wave using the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is well known. The signal may be stored on a processor-readable medium.
[0138] This document describes numerous embodiments. Features of the embodiments may be provided individually or in any combination across various claim classes and types. Further, embodiments may include one or more of the features, devices, or aspects described individually or in any combination across various claim classes and types. For example, the features described herein may be implemented in a bitstream or signal that includes information generated as described herein. This information may allow a decoder to decode a bitstream, encoder, bitstream, and / or decoder according to any of the described embodiments. For example, the features described herein may be implemented by creating and / or transmitting and / or receiving and / or decoding a bitstream or signal. For example, the features described herein may be implemented using methods, processes, apparatus, media storing instructions, media storing data, or signals. For example, the features described herein may be implemented by a TV, set-top box, cellular phone, tablet, or other electronic device performing decoding. The TV, set-top box, cellular phone, tablet, or other electronic device may display (e.g., using a monitor, screen, or other type of display) an image obtained (e.g., a signal reconstructed from a residual of a video bitstream). The TV, set-top box, cellular phone, tablet, or other electronic device may receive a signal including an encoded image and perform decoding.
[0139] Figure 5 The illustration shows an example of an intra-template matching search region used in intraTMP. IntraTMP is an intra-prediction mode that copies the best prediction block from a reconstructed portion of the current image (e.g., the current frame), with its L-shaped template matching the current template. For a predefined search range, the encoder can search for the template most similar to the current template in the reconstructed portion of the current image and use the corresponding block as the prediction block. The encoder can (e.g., then can) signal the use of this mode. The same prediction operation can be performed on the decoder side.
[0140] By connecting the L-shaped causal neighbors of the current block with Figure 5The prediction signal is generated by matching another block within a predefined search region, which includes: R1: Current CTU R2: Top Left CTU R3: Up CTU R4: Left CTU.
[0141] The sum of absolute differences (SAD) can be used as a cost function. Within a region (e.g., within each region), the decoder can search for the template with the minimum SAD relative to the current template and use its corresponding block as the prediction block. The region size (SearchRange_w, SearchRange_h) can be set to be proportional to the block size (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. That is: SearchRange_w = a * BlkW SearchRange_h = a * BlkH Here, "a" can be a constant that controls the gain / complexity tradeoff. For example, "a" can be equal to 5.
[0142] The intraTMP tool can be enabled for CUs with a width and height of 64 or less. This maximum CU size for intraTMP can be configurable. IntraTMP mode can be signaled at the CU level using a dedicated flag.
[0143] Multiple candidates can be selected for IntraTMP. The index used to indicate the best candidate can be encoded into a bitstream.
[0144] The fusion method of IntraTMP can be used (e.g., where the fusion weights are derived from template cost or a fixed value). The number of fused blocks can be determined by a threshold. Up to three candidate blocks can be fused. If one block is selected (e.g., only one) for fusion, the final predictor can be a fusion of the selected matching block and an intra-predictor derived from the planar pattern.
[0145] Subpixel precision can be supported in IntraTMP. Subpixel positions around integer pixel positions obtained by template matching can be enabled for IntraTMP. DCT-IF filters can be used for subpixel interpolation.
[0146] Figure 6 The illustration shows an example of choosing between two optimal predictors for IntraTMP.
[0147] Adaptive IntraTMP fusion can be used (e.g., with up to two prediction blocks corresponding to the best and second-best TM costs), such as... Figure 6 As shown, P1 is the prediction data with the lowest TM cost and P2 is the prediction data with the second lowest TM cost. Two conditions can be applied to determine whether to apply fusion. The weight derivation can be the same as in TIMD. The fused prediction (e.g., the final fused prediction) can be a weighted sum of P1 and P2.
[0148] Other (e.g., three additional) IntraTMP modes (e.g., left template mode, top template mode, and L-shaped fusion mode) can be used. The left and top template modes can be used (e.g., only) on the left and top sides respectively to derive template matching candidates. The L-shaped fusion mode can fuse the two best template matching candidates according to a cost-based linear combination formula. Figure 7 Examples of the left and top template shapes are shown in the illustration.
[0149] This paper provides one or more features associated with Template-Based Intra-Frame Pattern Derivation (TIMD).
[0150] Figure 8A The example template region (820) and reference sample (810) used to derive TIMD patterns and association weights are illustrated. Figure 8B The illustration shows an example reference sample (830) used to derive intra-prediction for the current block.
[0151] For each intra-prediction mode in the MPM (e.g., each intra-prediction mode) (and possibly supplemented by default modes PLANA and DC), the SATD between the lumen prediction of the reconstructed reference sample (810) with its built-in template and the lumen prediction of the reconstructed sample (820) with its built-in template is calculated. The two intra-prediction modes (IPMs) with the minimum SATD can be selected as TIMD modes. These two TIMD modes can be fused with weights (weight1, weight2) after the PDPC procedure is applied. The weighted intra-prediction can be used to encode the current CU. Position-dependent intra-prediction combination (PDPC) can be included in the derivation of the TIMD modes.
[0152] The costs of the two selected modes can be compared with a threshold. Cost factor 2 can be applied as follows: costMode2 < 2 * costMode1. If this condition is true, fusion can be applied. If the condition is not true, mode1 can be used (e.g., only mode1).
[0153] The weights of the SATD cost calculation model can be determined based on the following: Where weight2 = 1 – weight1.
[0154] This article provides one or more features associated with diffusion filters.
[0155] Filters can be applied to blocks predicted intra-frame or inter-frame.
[0156] For example, pred This can be a predicted signal over a given block (e.g., obtained from intra-frame or motion-compensated prediction). To handle boundary points for the filter, the predicted signal can be extended to the predicted signal. pred ext The expanded prediction signal can be formed by adding rows (e.g., a row) of reconstructed samples from the left and top of the block to the prediction signal. The signal (e.g., the resulting signal) can be mirrored in one or more (e.g., all) directions. This can be achieved by combining the prediction signal with a fixed mask (e.g., as defined below). h I The given method uses convolution to implement a uniform diffusion filter. This can be achieved using the boundary extension and other techniques described in this paper. h I * pred Alternative Predictive Signal pred Filter mask h I It can be given as: .
[0157] Prediction can be performed based on the type of transform used to encode the current block. A device (e.g., a video encoder, video decoder) can select candidate metrics for deriving a template-based tool based on the transform used for the current block. The device can smooth the prediction (e.g., if a regular transform is used for the current block). Prediction with a type of transform that depends on the type used can be combined with smoothed prediction (e.g., if a regular transform is used for the current block).
[0158] The device can select a derivation metric (e.g., candidate pattern derivation metric / candidate pattern template metric) depending on the type of residual coding used for the current block.
[0159] The IntraTMP tool is used as an example to illustrate the use of different metrics for template-based analysis based on the transformation pattern used for the current block. It can be understood that one or more other metrics can be used (e.g., independently) to derive TIMD or other tools for predicting candidates using a template from the current block.
[0160] Figure 9The illustration shows an example evaluation of the intraTMP mode at the encoder (e.g., where the metric used to derive the candidate depends on the residual type). As illustrated, template-based mode-based evaluation can be performed. The best candidate can be derived using SATD. The best candidate derived using SATD can be evaluated using regular residual coding. The best candidate can be derived using SAD. The best candidate derived using SAD can be evaluated using transform skipping or no residuals. The best candidate can be selected based on the associated rate distortion cost (RDCost).
[0161] The metric used to derive intraTMP candidates can be determined based on the type of residual encoded for the current block (e.g., inference). If no residual or transform skip is selected, SAD can be used as the intraTMP search metric. If regular residuals are used, SATD can be used as the intraTMP search metric (e.g., as...). Figure 9 (Illustrated in the figure). At the encoder, the best candidate derived using the SATD metric can be evaluated using a regular transformation. If (e.g., during evaluation) the encoder estimate does not require a residual for the current block, the candidate can be discarded.
[0162] In some examples, candidates can be derived at the decoder side (e.g., only one candidate). The decoder can know (e.g., at the parsing stage) whether a regular transformation, a transformation skipped, or no transformation will be used for the current block. The decoder can derive candidates (e.g., only candidates) using metrics related to the type of residual used for the current block.
[0163] Figure 10 The illustration shows an example derivation of the optimal intraTMP predictor at the decoder (e.g., depending on the residual type). As illustrated, template-based modes can be decoded. If transform skipping and no-residual modes are not used, SATD can be used to derive the optimal candidate. If transform skipping or no-residual modes are used, SAD can be used to derive the optimal candidate.
[0164] This paper provides one or more features associated with smoothing predictions for conventional residual encoding. If the prediction does not depend on the residual type, a smoothing step can be applied to the predictor (e.g., in the case of conventional residuals).
[0165] Figure 11The diagram illustrates an example derivation of the predictor, where smoothing is performed if regular residuals are used. Evaluation of regular patterns (e.g., merge patterns, affine patterns, etc.) can be performed. Prediction can be performed for the current candidate. Candidates can be evaluated using transform skipping or no residuals. Smoothing can be performed on the predictor. Candidates can be evaluated using regular residual encoding. An optimal candidate can be selected (e.g., from candidates evaluated using transform skipping or no residuals and candidates evaluated using regular residual encoding). The optimal candidate can be selected based on RDCost.
[0166] Smoothing can involve linear filters. For example, smoothing can involve diffusion filters. Smoothing can involve the detection of one or more particles (e.g., where a particle sample is detected and replaced by the average of its neighboring samples). Other filters can be applied (e.g., bilateral filters (BIF) or adaptive loop filters (ALF)).
[0167] Figure 12 The diagram illustrates an example derivation of the predictor at the decoder, with a smoothing step added when using regular residual coding. As illustrated, regular modes (e.g., merge mode, affine mode, etc.) can be used to predict the coded block. Prediction can be performed for the current candidate. Without using transform skipping and no residual mode, SATD can be used to derive the best candidate. In this case, predictor smoothing can be performed. If transform skipping or no residual mode is used, SAD can be used to derive the best candidate.
[0168] This paper provides one or more features associated with smoothing template-based predictions for conventional residual encoding. Smoothing can be performed between predictions and residuals (e.g., if using conventional residuals and / or if the predictions differ depending on the residual type, such as...). Figure 11 and 12 (See illustration in the image).
[0169] Figure 13 The illustration shows an example derivation of the optimal intraTMP predictor, with a smoothing step added when using regular residuals. As illustrated, template-based patterns can be evaluated. The optimal candidate can be derived using SATD. Smoothing of the predictor can be performed. Candidates derived using SATD can be evaluated using regular residual encoding. The optimal candidate can be derived using SAD. For example, candidates derived using SAD can be evaluated using transform skipping or no residuals. The optimal candidate can be selected (e.g., from candidates evaluated using transform skipping or no residuals and candidates evaluated using regular residual encoding). The optimal candidate can be selected based on RDCost.
[0170] Figure 14The illustration shows an example derivation of the intraTMP predictor at the decoder, where smoothing is performed if regular residual coding is used. As shown, template-based patterns can be decoded. Without transform skipping and no-residual patterns, SATD can be used to derive the best candidate. Smoothing of the predictor can be performed. If transform skipping or no-residual patterns are used, SAD can be used to derive the best candidate.
[0171] An example device for video decoding can determine the transform mode associated with the current block. The device can select a candidate mode template metric for the current block from multiple candidate mode template metrics based on the transform mode. The device can derive candidate modes based on the candidate mode template metrics. The device can decode the block based on the candidate modes.
[0172] When the transform mode is a regular residual coding mode, the device can smooth the predictor associated with the current block. Based on whether the transform mode associated with the current block is transform skipped or has no residual, the device can select a first candidate mode template metric as the candidate mode template metric for the current block. Based on whether the transform mode associated with the current block is a residual coding mode, the device can select a second candidate mode template metric as the candidate mode template metric for the current block.
[0173] The example device for video coding can derive a first candidate mode based on a first candidate mode template metric associated with a first candidate mode. The device can derive a second candidate mode based on a second candidate mode template metric associated with a second candidate mode. The device can select a candidate mode from the first and second candidate modes. The device can include an indication of the transform mode associated with the selected candidate mode in the video data.
[0174] A device for video encoding may (for example, be configured to) perform one or more of the following: derive a first candidate mode associated with a first candidate mode template metric based on a first residual type; derive a second candidate mode associated with a second candidate mode template metric based on a second residual type; select a candidate mode from the first candidate mode and the second candidate mode; and encode a video block based on the selected candidate mode.
[0175] The device can indicate the type of residual associated with the selected candidate pattern in the video data. The device can also indicate the candidate pattern template metric in the video data.
[0176] The device can select candidate modes based on rate distortion cost analysis.
[0177] The first candidate mode template metric can be the sum of absolute transform differences (SATD), and the second candidate mode template metric can be the sum of absolute differences (SAD). The device can obtain a first predictor based on the first candidate mode; apply a filter to the first predictor associated with the first candidate mode; and obtain a second predictor based on the second candidate mode. The candidate modes can be selected based on the filtered first and second predictors. The filter can be one of a spread filter, a bilateral filter, or an adaptive loop filter.
[0178] Based on the determination of the first residual type as the residual, the device can select the sum of absolute transformation differences (SATD) as the first candidate mode template metric.
[0179] Based on determining the second residual type as either no residual or transform skip, the device can select the sum of absolute differences (SAD) as the second candidate pattern template metric.
[0180] Although features and elements have been described above in specific combinations, those skilled in the art will appreciate that each feature or element may be used individually or in any combination with other features and elements. Furthermore, the methods described herein can be implemented in a computer program, software, or firmware incorporated in a computer-readable medium for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted via wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media (such as internal hard disks and removable disks), magnetic-optical media, and optical media (such as CD-ROMs and digital multifunction discs (DVDs)). The processor associated with the software can be used to implement a radio frequency transceiver for a WTRU, UE, terminal, base station, RNC, or any host computer.
Claims
1. A device for video decoding, comprising: The processor is configured as follows: Determine the type of residual associated with the current block; Based on the residual type, a candidate pattern template metric for the current block is selected from multiple candidate pattern template metrics; Candidate patterns are derived based on the candidate pattern template metric. as well as The current block is decoded based on the candidate pattern.
2. The device of claim 1, wherein the plurality of candidate pattern template metrics includes at least one of Sum of Absolute Transform Differences (SATD) or Sum of Absolute Differences (SAD).
3. The device of claim 1, wherein selecting a candidate pattern template metric for the current block based on the residual type further comprises: Based on the residual type associated with the current block being transform skipped, the sum of absolute differences (SAD) for the current block is selected.
4. The device of claim 1, wherein selecting a candidate pattern template metric for the current block based on the residual type further comprises: Based on the fact that the residual type associated with the current block is no residual, the sum of absolute differences (SAD) for the current block is selected.
5. The device of claim 1, wherein determining the residual type comprises: Based on the fact that the residual type associated with the current block is a residual, the sum of absolute transformation differences (SATD) for the current block is selected.
6. The device as claimed in any one of claims 1-5, wherein the processor is further configured to: The predictor is obtained based on the derived candidate patterns; and Based on the determination of the residual type indicating the association of the residual with the current block, the filter is applied to the predictor associated with the derived candidate pattern.
7. The device of claim 6, wherein the filter comprises at least one of a diffuse filter, a bilateral filter, or an adaptive loop filter.
8. The device of any one of claims 1-7, wherein the candidate mode template metric is a template-based intra-frame mode derivation (TIMD) metric, and the candidate mode is a TIMD candidate.
9. The device of any one of claims 1-7, wherein the candidate pattern template metric is an IntraTMP search metric, and the candidate pattern is an IntraTMP candidate.
10. A method for decoding video, comprising: Determine the type of residual associated with the current block; Based on the residual type, a candidate pattern template metric for the current block is selected from multiple candidate pattern template metrics; Candidate patterns are derived based on the candidate pattern template metric. as well as The current block is decoded based on the candidate pattern.
11. The method of claim 10, wherein selecting a candidate mode template metric for the current block based on the residual type further comprises: Based on the determination of the transformation skip associated with the current block, the sum of absolute differences (SAD) for the current block is selected.
12. The method of claim 10, wherein selecting a candidate pattern template metric for the current block based on the residual type further comprises: Based on the determination that the residual is associated with the current block, the sum of absolute transformation differences (SATD) for the current block is selected.
13. The method of any one of claims 10-12, further comprising: The predictor is obtained based on the derived candidate patterns; as well as Based on the determination of the residual type indicating the association of the residual with the current block, the filter is applied to the predictor associated with the derived candidate pattern.
14. The method of claim 13, wherein the filter is at least one of a diffuse filter, a bilateral filter, or an adaptive loop filter.
15. A computer program product stored on a non-transient computer-readable medium and comprising program code instructions for implementing any one of the methods of claims 10-14 when executed by a processor.