Regression-Based Intra-Prediction Blending
Regression-based intra-prediction blending techniques optimize intra-prediction modes in video coding systems, enhancing decoding and encoding efficiency and compression performance.
Patent Information
- Application Number
- JP2025530400
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-23
- Filing Date
- 2023-12-20
- Publication Date
- 2026-01-14
AI Technical Summary
Existing video coding systems face challenges in efficiently determining intra-prediction modes for video decoding and encoding, leading to suboptimal compression and transmission of digital video signals.
The implementation of regression-based intra-prediction blending techniques, which utilize weights derived from reconstructed and predicted samples to optimize intra-prediction models, allowing for improved prediction accuracy and compression efficiency.
Enhances the accuracy and efficiency of video decoding and encoding by optimizing intra-prediction modes, resulting in improved compression and reduced bandwidth requirements.
Smart Images

Figure 2026501079000001_ABST
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of European Provisional Patent Application No. 22307027.7, filed December 23, 2022, the contents of which are incorporated herein by reference. [Background technology]
[0002] Video coding systems can be used to compress digital video signals, for example, to reduce the storage and / or transmission bandwidth of such signals. Video coding systems may include, for example, block-based, wavelet-based, and / or object-based systems. Summary of the Invention
[0003] The video decoding device may be configured to obtain a template associated with the block, including a reconstructed sample of the template and a corresponding predicted sample. The video decoding device may determine weights based on minimizing a difference between the predicted sample and the corresponding reconstructed sample. The video decoding device may apply the weights to respective parameters, for example, to determine an intra-prediction model. The parameters may include reconstructed samples neighboring the sample position and / or iteratively derived predicted samples (e.g., obtained based on an intra-prediction mode). The video decoding device may determine intra-prediction samples for the block based on the intra-prediction model. The predicted sample, the corresponding reconstructed sample, and the intra-prediction sample may share a component type. For example, the predicted sample, the corresponding reconstructed sample, and the intra-prediction sample may be luma samples. For example, the predicted sample, the corresponding reconstructed sample, and the intra-prediction sample may be chroma samples. The video decoding device may decode the block based on the intra-prediction samples.
[0004] The video decoding device may obtain reference samples (e.g., predicted samples) of the template and may optimize weights corresponding to the reference samples of the template based on minimizing the difference between the sum of the weighted reference samples and the reconstructed samples of the template.
[0005] The intra-prediction sample may be associated with a sample position within the block, and the parameters may include predicted samples neighboring the sample position. The video decoding device may determine the intra-prediction sample for the block based on the intra-prediction model, for example, by applying the determined weights to the predicted samples neighboring the sample position.
[0006] The intra-prediction model may include applying a weight to each reconstructed sample neighboring the sample position to determine an intra-prediction sample. The intra-prediction sample may be associated with a sample position within the block. The parameters may include the reconstructed samples neighboring the sample position. For a second sample position within the block, the video decoding device may identify predicted samples neighboring the second sample position. The video decoding device may apply the determined weight to predicted samples neighboring the second sample position to determine a second intra-prediction sample for the block. The block may be decoded further based on the second intra-prediction sample of the block. The predicted samples neighboring the second sample position may include the first intra-prediction sample. For example, the derivation of the intra-prediction sample may be regression-based and / or iterative.
[0007] The intra-prediction model may include applying weights to respective prediction samples obtained based on different derived prediction modes. The video decoding device may derive the intra-prediction mode based on the reconstructed samples of the template and the reference samples. The parameters may correspond to the prediction samples obtained based on the derived intra-prediction mode. Determining the intra-prediction samples for the block based on the intra-prediction model may include applying weights to the prediction samples obtained based on the derived intra-prediction mode.
[0008] The video decoding device may obtain a histogram of gradients associated with samples of the template. The video decoding device may derive an intra-prediction mode based on the histogram of gradients associated with the template. Prediction samples of the template may be obtained based on each derived intra-prediction mode, and the parameters may correspond to the prediction samples obtained based on the derived intra-prediction mode.
[0009] The intra-prediction model may include applying weights to each prediction block obtained based on different derived prediction modes. The video decoding device may derive the intra-prediction mode based on reconstructed samples of a template, for example. The video decoding device may obtain predictive samples for the block based on the derived intra-prediction mode. The video decoding device may blend the predictive samples based on the weights.
[0010] The video decoding device may, for example, derive an intra-prediction mode based on the reconstructed samples of the template. The video decoding device may obtain a weighted and blended prediction of the template based on the derived intra-prediction mode. In an example, each of the weighted and blended predictions of the template samples may correspond to a respective reconstructed sample of the reconstructed samples in the template. The video decoding device may optimize the weights corresponding to the derived intra-prediction mode based on minimizing the difference between the weighted and blended prediction of the template and the reconstructed samples in the template.
[0011] The video encoding device may be configured to obtain a template associated with the block. The video encoding device may obtain predicted samples and corresponding reconstructed samples of the template. The video encoding device may determine weights based on minimizing differences between the predicted samples of the template and the corresponding reconstructed samples. The video encoding device may determine an intra-prediction model including applying the weights to respective parameters. The video encoding device may determine intra-prediction samples for the block based on the intra-prediction model. The predicted samples, the corresponding reconstructed samples, and the intra-prediction samples share a component type (e.g., chroma, luma). The video encoding device may encode the block based on the intra-prediction samples.
[0012] The video encoding device may obtain reference samples of the template. The weights may correspond to the reference samples of the template. The video encoding device may optimize the weights based on minimizing a difference between a sum of the weighted reference samples and a reconstructed sample of the template.
[0013] The intra-prediction sample may be associated with a sample position within the block, and the parameters may include predicted samples neighboring the sample position. The video encoding device may determine the intra-prediction sample for the block based on the intra-prediction model, for example, by applying the determined weights to the predicted samples neighboring the sample position.
[0014] The intra-predicted sample may be a first intra-predicted sample and may be associated with a first sample position within the block. The parameters may include reconstructed samples adjacent to the first sample position. For a second sample position within the block, the video encoding device may identify predicted samples adjacent to the second sample position. The video encoding device may apply the determined weights to the predicted samples to determine a second intra-predicted sample for the block. The block may be decoded further based on the second intra-predicted sample of the block. The predicted samples adjacent to the second sample position may include the first intra-predicted sample.
[0015] The video encoding device may derive an intra-prediction mode based on the reconstructed samples of the template and the reference samples of the template. The parameters may correspond to prediction samples obtained based on the derived intra-prediction mode. The video encoding device may determine intra-prediction samples for the block based on an intra-prediction model, for example, by applying weights to the prediction samples obtained based on the derived intra-prediction mode.
[0016] The video encoding device may obtain a histogram of gradients associated with samples of the template. The video encoding device may derive an intra-prediction mode based on the histogram of gradients associated with the template. Prediction samples of the template may be obtained based on each derived intra-prediction mode, and / or the parameters may correspond to prediction samples obtained based on the derived intra-prediction mode.
[0017] The video encoding device may derive an intra-prediction mode based at least on the reconstructed samples of the template. The video encoding device may obtain predictive samples for the block based on the derived intra-prediction mode. The video encoding device may blend the predictive samples based on weights.
[0018] The video encoding device may derive an intra-prediction mode based at least on the reconstructed samples of the template. The video encoding device may obtain a weighted and blended prediction of the template based on the derived intra-prediction mode. In an example, each of the weighted and blended predictions of the template may correspond to a respective reconstructed sample of the reconstructed samples in the template. The video encoding device may optimize the weights corresponding to the derived intra-prediction mode based on minimizing a difference between the weighted and blended prediction of the template and the reconstructed samples in the template.
[0019] Systems, methods, and means for performing regression-based intra prediction mode (IPM) blending are disclosed. Intra prediction blending weights may be derived for template-based intra mode derivation (TIMD) and decoder-side intra mode derivation (DIMD), for example, using a regression-based method. Intra prediction samples may be generated from regression-based locally calculated parameters for neighboring samples. Regression-based weights may be derived for TIMD. IPM blending weights may be derived for a selected TIMD mode (e.g., having the smallest sum of absolute transformed differences (SATD) costs). Regression-based weights may be derived for DIMD. IPM blending weights for a selected intra DIMD mode (e.g., having the highest histogram bar) and weights for PLANAR mode may be determined using a regression-based method. Regression-based intra prediction may be calculated as a linear combination of template-derived parameters. The parameters derived from the template may include a reconstructed reference sample (e.g., chroma, luma) in the same column or line as the sample in the template, or the position (e.g., value) of the reconstructed reference sample. Regression-based intra prediction may be iterative. One or more parameters (e.g., all parameters) may be iteratively derived from previously predicted sample values of the current block. For example, the parameters may include predicted samples neighboring the current intra prediction sample. For example, the intra prediction parameters may depend on (e.g., only on) neighboring reference samples. The intra prediction sample may be a block, a sub-block, or a sub-sub-block (e.g., one pixel or multiple pixels). Regression-based model derivation may be performed using local adaptation.Intra prediction may be spatially adapted, for example, according to a histogram of oriented gradient (HoG) cost (e.g., DIMD) and / or a SATD cost (e.g., TIMD). Spatial adaptation may be performed using spatial weighting. Weighting adaptation may be combined with various implementations, for example, using spatially adapted weighting in IPM blending.
[0020] The systems, methods, and means described herein may involve a decoder. In some examples, the systems, methods, and means described herein may involve an encoder. In some examples, the systems, methods, and means described herein may involve a signal (e.g., from an encoder and / or received by a decoder). A computer-readable medium may include instructions for causing one or more processors to perform the methods described herein. A computer program product may include instructions that, when executed by one or more processors, cause the one or more processors to perform the methods described herein. [Brief explanation of the drawings]
[0021] [Figure 1A] FIG. 1 is a system diagram illustrating an example communication system in which one or more disclosed embodiments may be implemented. [Figure 1B] 1B is a system diagram illustrating an exemplary wireless transmit / receive unit (WTRU) that may be used within the communication system shown in FIG. 1A, according to an embodiment. [Figure 1C] 1B is a system diagram illustrating an example radio access network (RAN) and an example core network (CN) that may be used within the communication system shown in FIG. 1A, according to an embodiment. [Figure 1D]1B is a system diagram illustrating a further exemplary RAN and a further exemplary CN that may be used within the communication system shown in FIG. 1A, according to an embodiment. [Figure 2] 1 illustrates an exemplary video encoder. [Figure 3] 1 illustrates an exemplary video decoder. [Figure 4] 1 illustrates an example system in which various aspects and embodiments may be implemented. [Figure 5] 10 shows an exemplary definition of samples used by the PDPC applied to the diagonal-top-right mode. [Figure 6A] 1 shows examples of template areas and reference samples that can be used to derive template-based intra-mode derivation (TIMD) modes and associated weights. [Figure 6B] 10 illustrates examples of reference samples that may be used to derive intra predictions for the current block. [Figure 7] 10 shows an example of derivation of blending weights for decoder-side intra-mode derivation (DIMD). [Figure 8] 10 shows an example of using adjacent reconstructed samples for DIMD chroma mode. [Figure 9] 1 shows an example of the spatial part of a convolution filter. [Figure 10] 10 shows an example of a reference area that can be used to derive filter coefficients. [Figure 11] 1 shows an example of spatial samples used for GL-CCCM. [Figure 12] 1 shows an example of a contiguous region used to derive a CC model. [Figure 13A] 10 shows an example of using template samples and reference samples to derive an intra-prediction model. [Figure 13B] 10 illustrates an example of using a reference sample to apply an intra-prediction model. [Figure 14] An example of iteratively applying a predictive model is shown. [Figure 15]An example of using spatially varying weights to derive a blending model is shown. [Figure 16] 10 shows an example of deriving blending parameters using spatially varying blending weights. DETAILED DESCRIPTION OF THE INVENTION
[0022] A more detailed understanding may be had from the following description, given by way of example in conjunction with the accompanying drawings, in which:
[0023] 1A illustrates an exemplary communication system 100 in which one or more disclosed embodiments may be implemented. Communication system 100 may be a multiple-access system that provides content, such as voice, data, video, messaging, broadcasts, etc., to multiple wireless users. Communication system 100 may enable multiple wireless users to access such content through the sharing of system resources, including wireless bandwidth. For example, the communication system 100 may use one or more channel access methods such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single-carrier FDMA (SC-FDMA), zero-tail unique-word DFT-Spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block filtered OFDM, filter bank multicarrier (FBMC), etc.
[0024] 1A, communications system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RANs 104 / 113, CNs 106 / 115, a public switched telephone network (PSTN) 108, the Internet 110, and other networks 112, although it will be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of WTRUs 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a wireless environment. By way of example, the WTRUs 102a, 102b, 102c, 102d (any of which may be referred to as a “station” and / or “STA”) may be configured to transmit and / or receive wireless signals and may include user equipment (UE), a mobile station, a fixed or mobile subscriber unit, a subscription-based unit, a pager, a mobile phone, a personal digital assistant (PDA), a smartphone, a laptop, a netbook, a personal computer, a wireless sensor, a hotspot or Mi-Fi device, an Internet of Things (IoT) device, a watch or other wearable, a head-mounted display (HMD), a vehicle, a drone, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in the context of an industrial and / or automated processing chain), a consumer electronics device, a device operating in a commercial and / or industrial wireless network, etc. Any of the WTRUs 102a, 102b, 102c, and 102d may be referred to interchangeably as a UE.
[0025] The communications system 100 may also include a base station 114a and / or a base station 114b. Each of the base stations 114a, 114b may be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communications networks, such as the CN 106 / 115, the Internet 110, and / or other networks 112. By way of example, the base stations 114a, 114b may be a base transceiver station (BTS), a Node B, an eNodeB, a Home Node B, a Home eNodeB, a gNB, an NR Node B, a site controller, an access point (AP), a wireless router, etc. Although the base stations 114a, 114b are each shown as a single element, it will be understood that the base stations 114a, 114b may include any number of interconnected base stations and / or network elements.
[0026] The base station 114a may be part of the RAN 104 / 113, which may also include other base stations and / or network elements (not shown), such as a base station controller (BSC), a radio network controller (RNC), a relay node, etc. The base station 114a and / or base station 114b may be configured to transmit and / or receive radio signals on one or more carrier frequencies, which may be referred to as a cell (not shown). These frequencies may be licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide wireless service coverage for a particular geographic area, which may be relatively fixed or may change over time. A cell may be further divided into cell sectors. For example, the cell associated with the base station 114a may be divided into three sectors. Thus, in one embodiment, the base station 114a may include three transceivers, i.e., one transceiver for each sector of the cell. In some embodiments, the base station 114a may employ multiple-input multiple output (MIMO) technology and may utilize multiple transceivers per sector of the cell, for example, using beamforming to transmit and / or receive signals in desired spatial directions.
[0027] The base stations 114a, 114b may communicate with one or more of the WTRUs 102a, 102b, 102c, 102d over an air interface 116, which may be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 116 may be established using any suitable radio access technology (RAT).
[0028] More specifically, as noted above, the communications system 100 may be a multiple-access system, but may use one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, the base stations 114a and WTRUs 102a, 102b, 102c of the RANs 104 / 113 may implement a radio technology such as Universal Mobile Telecommunications System (UMTS) Universal Terrestrial Radio Access (UTRA), which may establish the air interfaces 115 / 116 / 117 using wideband CDMA (WCDMA). WCDMA may include communications protocols such as High-Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High-Speed Downlink (DL) Packet Access (HSDPA) and / or High-Speed Uplink Packet Access (HSUPA).
[0029] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which may establish the air interface 116 using Long Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-Advanced Pro (LTE-A Pro).
[0030] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as NR radio access, which may establish the air interface 116 using New Radio (NR).
[0031] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement multiple radio access technologies. For example, the base station 114a and the WTRUs 102a, 102b, 102c may jointly perform LTE radio access and NR radio access, e.g., using dual connectivity (DC) principles. Thus, the air interface utilized by the WTRUs 102a, 102b, 102c may be characterized by transmissions sent to and from multiple types of radio access technologies and / or multiple types of base stations (e.g., eNBs and gNBs).
[0032] In other embodiments, the base station 114a and the WTRUs 102a, 102b, 102c may implement a wireless technology such as IEEE 802.11 (i.e., Wireless Fidelity, WiFi), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access, WiMAX), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile communications (GSM), Enhanced Data rates for GSM Evolution (EDGE), GSM EDGE (GSM EDGE, GERAN), or the like.
[0033] 1A may be, for example, a wireless router, a Home NodeB, a Home eNodeB, or an access point, but may utilize any suitable RAT to facilitate wireless connectivity in a local area such as a business, home, vehicle, campus, industrial facility, air corridor (e.g., for use by drones), road, or other location. In one embodiment, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In an embodiment, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 114b and the WTRUs 102c, 102d may establish a picocell or a femtocell using a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.). As shown in FIG. 1A, the base station 114b may be directly connected to the Internet 110. Thus, the base station 114b may not need to access the Internet 110 through the CN 106 / 115.
[0034] The RAN 104 / 113 may communicate with the CN 106 / 115, which may be any type of network configured to provide voice, data, application, and / or voice over internet protocol (VoIP) services to one or more of the WTRUs 102a, 102b, 102c, 102d. The data may have various quality of service (QoS) requirements, such as different throughput, latency, error tolerance, reliability, data throughput, and mobility requirements. The CN 106 / 115 may provide call control, billing services, mobile location-based services, prepaid calling, Internet connectivity, video distribution, and / or perform high-level security functions such as user authentication. Although not shown in FIG. 1A , it will be understood that the RAN 104 / 113 and / or the CN 106 / 115 may communicate directly or indirectly with other RANs that use the same RAT as the RAN 104 / 113 or a different RAT. For example, in addition to being connected to the RAN 104 / 113, which may utilize NR radio technology, the CN 106 / 115 may also communicate with another RAN (not shown) that uses GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.
[0035] The CN 106 / 115 may also serve as a gateway for the WTRUs 102a, 102b, 102c, 102d to access the PSTN 108, the Internet 110, and / or other networks 112. The PSTN 108 may include a circuit-switched telephone network providing plain old telephone service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices that use common communication protocols, such as the transmission control protocol (TCP), the user datagram protocol (UDP), and / or the internet protocol (IP) of the TCP / IP Internet protocol suite. The network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, the network 112 may include another CN connected to one or more RANs, which may use the same RAT as the RANs 104 / 113 or a different RAT.
[0036] Some or all of the WTRUs 102a, 102b, 102c, 102d in the communications system 100 may include multi-mode capabilities (e.g., the WTRUs 102a, 102b, 102c, 102d may include multiple transceivers for communicating with different wireless networks over different wireless links.) For example, the WTRU 102c shown in FIG. 1A may be configured to communicate with a base station 114a that can use cellular-based wireless technology and a base station 114b that can use IEEE 802 wireless technology.
[0037] 1B is a system diagram illustrating an example WTRU 102. As shown in FIG. 1B, the WTRU 102 may include, among other things, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power source 134, a global positioning system (GPS) chipset 136, and / or other peripherals 138. It will be understood that the WTRU 102 may include any sub-combination of the foregoing elements while remaining consistent with an embodiment.
[0038] The processor 118 may be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. The processor 118 may perform signal coding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 102 to operate in a wireless environment. The processor 118 may be coupled to the transceiver 120, which may be coupled to the transmit / receive element 122. While FIG. 1B depicts the processor 118 and the transceiver 120 as separate components, it will be understood that the processor 118 and the transceiver 120 may be integrated together in an electronic package or chip.
[0039] The transmit / receive element 122 may be configured to transmit or receive signals to or from a base station (e.g., base station 114a) over the air interface 116. For example, in one embodiment, the transmit / receive element 122 may be an antenna configured to transmit and / or receive RF signals. In an embodiment, the transmit / receive element 122 may be an emitter / detector configured to transmit and / or receive, for example, IR signals, UV signals, or visible light signals. In yet another embodiment, the transmit / receive element 122 may be configured to transmit and / or receive both RF signals and light signals. It will be understood that the transmit / receive element 122 may be configured to transmit and / or receive any combination of wireless signals.
[0040] 1B as a single element, the WTRU 102 may include any number of transmit / receive elements 122. More specifically, the WTRU 102 may use MIMO technology. Thus, in one embodiment, the WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 116.
[0041] The transceiver 120 may be configured to modulate signals transmitted by the transmit / receive element 122 and demodulate signals received by the transmit / receive element 122. As noted above, the WTRU 102 may have multi-mode capabilities. Thus, the transceiver 120 may include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs, such as, for example, NR and IEEE 802.11.
[0042] The processor 118 of the WTRU 102 may be coupled to and may receive user-entered data from a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light-emitting diode (OLED) display unit). The processor 118 may also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. Additionally, the processor 118 may access information from and store data in any type of suitable memory, such as non-removable memory 130 and / or removable memory 132. The non-removable memory 130 may include random-access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 may include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, etc. In other embodiments, the processor 118 may access information from and store data in memory that is not physically located on the WTRU 102, such as on a server or home computer (not shown).
[0043] The processor 118 may receive power from the power source 134 and may be configured to distribute and / or control the power to other components in the WTRU 102. The power source 134 may be any suitable device for providing power to the WTRU 102. For example, the power source 134 may include one or more dry batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.
[0044] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to, or instead of, information from the GPS chipset 136, the WTRU 102 may receive location information from base stations (e.g., base stations 114a, 114b) over the air interface 116 and / or determine its location based on the timing of signals received from two or more nearby base stations. It will be appreciated that the WTRU 102 may obtain location information by way of any suitable location determination method while remaining consistent with an embodiment.
[0045] The processor 118 may further be coupled to other peripherals 138, which may include one or more software and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, the peripherals 138 may include an accelerometer, an electronic compass, a satellite transceiver, a digital camera (for photos and / or videos), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth module, a frequency modulated (FM) radio unit, a digital music player, a media player, a video game player module, an internet browser, a virtual reality and / or augmented reality (VR / AR) device, an activity tracker, etc. The peripheral device 138 may include one or more sensors, which may be one or more of a gyroscope, an accelerometer, a Hall effect sensor, a magnetometer, a direction sensor, a proximity sensor, a temperature sensor, a time sensor, a geolocation sensor, an altimeter, a light sensor, a touch sensor, a magnetometer, a barometer, a gesture sensor, a biometric sensor, and / or a humidity sensor.
[0046] The WTRU 102 may include a full-duplex radio for transmitting and receiving some or all of the signals (e.g., associated with a particular subframe) on both the UL (e.g., for transmission) and downlink (e.g., for reception) in parallel and / or simultaneously. The full-duplex radio may include an interference management unit for reducing and / or substantially eliminating self-interference through either hardware (e.g., a choke) or signal processing via a processor (e.g., a separate processor (not shown) or via processor 118). In an embodiment, the WTRU 102 may include a half-duplex radio for transmitting and receiving some or all of the signals (e.g., associated with a particular subframe on either the UL (e.g., for transmission) or downlink (e.g., for reception) in parallel and / or simultaneously.
[0047] 1C is a system diagram illustrating the RAN 104 and the CN 106 in accordance with an embodiment. As noted above, the RAN 104 may communicate with the WTRUs 102a, 102b, 102c over the air interface 116 using E-UTRA radio technology. The RAN 104 may also communicate with the CN 106.
[0048] The RAN 104 may include eNodeBs 160a, 160b, and 160c, although it will be understood that the RAN 104 may include any number of eNodeBs while remaining consistent with an embodiment. The eNodeBs 160a, 160b, and 160c may each include one or more transceivers for communicating with the WTRUs 102a, 102b, and 102c over the air interface 116. In one embodiment, the eNodeBs 160a, 160b, and 160c may implement MIMO technology. Thus, the eNodeB 160a may, for example, use multiple antennas to transmit wireless signals to and / or receive wireless signals from the WTRU 102a.
[0049] Each of the eNodeBs 160a, 160b, 160c may be associated with a particular cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, scheduling of users in the UL and / or DL, etc. As shown in FIG. 1C, the eNodeBs 160a, 160b, 160c may communicate with one another over an X2 interface.
[0050] 1C may include a mobility management entity (MME) 162, a serving gateway (SGW) 164, and a packet data network (PDN) gateway (or packet gateway, PGW) 166. Although each of the foregoing elements is shown as part of the CN 106, it will be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.
[0051] The MME 162 may be connected to each of the eNodeBs 162a, 162b, 162c in the RAN 104 via an S1 interface and may function as a control node. For example, the MME 162 may be responsible for authenticating users of the WTRUs 102a, 102b, 102c, activating / deactivating bearers, selecting a particular serving gateway during initial connection of the WTRUs 102a, 102b, 102c, etc. The MME 162 may provide a control plane function for switching between the RAN 104 and other RANs (not shown) that employ other radio technologies such as GSM and / or WCDMA.
[0052] The SGW 164 may be connected to each of the eNodeBs 160a, 160b, 160c in the RAN 104 via an S1 interface. The SGW 164 may generally route and forward user data packets to and from the WTRUs 102a, 102b, 102c. The SGW 164 may perform other functions such as anchoring the user plane during inter-eNodeB handovers, triggering paging when DL data is available to the WTRUs 102a, 102b, 102c, managing and storing context for the WTRUs 102a, 102b, 102c, etc.
[0053] The SGW 164 may be connected to a PGW 166, which may provide the WTRUs 102a, 102b, 102c with access to packet-switched networks such as the Internet 110 to facilitate communication between the WTRUs 102a, 102b, 102c and IP-enabled devices.
[0054] The CN 106 may facilitate communications with other networks. For example, the CN 106 may provide the WTRUs 102a, 102b, 102c with access to circuit-switched networks, such as the PSTN 108, to facilitate communications between the WTRUs 102a, 102b, 102c and traditional landline communications devices. For example, the CN 106 may include or communicate with an IP gateway (e.g., an IP multimedia subsystem (IMS) server) that serves as an interface between the CN 106 and the PSTN 108. In addition, the CN 106 may provide the WTRUs 102a, 102b, 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.
[0055] Although the WTRU is depicted in FIGS. 1A-1D as a wireless terminal, in certain representative embodiments it is contemplated that such a terminal may use a wired communication interface (e.g., temporarily or permanently) with the communication network.
[0056] In an exemplary embodiment, the other network 112 may be a WLAN.
[0057] A WLAN in infrastructure Basic Service Set (BSS) mode may have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP may have access to or interface with a distribution system (DS) or another type of wired / wireless network that transmits traffic into and / or out of the BSS. Traffic to a STA originating from outside the BSS may arrive through the AP and be delivered to the STA. Traffic originating from a STA to a destination outside the BSS may be sent to the AP and transmitted to the respective destination. Traffic between STAs within a BSS may be transmitted, for example, through the AP, where the source STA may send traffic to the AP, and the AP may deliver the traffic to the destination STA. Traffic between STAs within a BSS may be considered and / or referred to as peer-to-peer traffic. Peer-to-peer traffic may be transmitted between (e.g., directly between) a source STA and a destination STA using a direct link setup (DLS). In certain representative embodiments, the DLS may use 802.11e DLS or 802.11z tunneled DLS (TDLS). A WLAN using an Independent BSS (IBSS) mode may not have an AP, and STAs within or using the IBSS (e.g., all of the STAs) may communicate directly with each other. The IBSS communication mode is sometimes referred to herein as an "ad hoc" communication mode.
[0058] When using 802.11ac infrastructure mode operation or a similar mode of operation, an AP may transmit beacons on a fixed channel, such as a primary channel. The primary channel may be a fixed width (e.g., a wide 20 MHz bandwidth) or a width dynamically configured via signaling. The primary channel may be the operating channel of the BSS and may be used by STAs to establish a connection with the AP. In certain representative embodiments, Carrier Sense Multiple Access with Collision Avoidance (CSMA / CA) may be implemented, for example, in an 802.11 system. With CSMA / CA, STAs (e.g., all STAs), including the AP, may sense the primary channel. If the primary channel is detected / sensed and / or determined to be busy by a particular STA, the particular STA may back off. One STA (e.g., only one station) may transmit at any given time in a given BSS.
[0059] High throughput (HT) STAs may use 40 MHz wide channels for communication, for example, through combination of a primary 20 MHz channel with adjacent or non-adjacent 20 MHz channels to form a 40 MHz wide channel.
[0060] A Very High Throughput (VHT) STA may support channels with widths of 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz. A 40 MHz and / or 80 MHz channel may be formed by combining contiguous 20 MHz channels. A 160 MHz channel may be formed by combining eight contiguous 20 MHz channels or two non-contiguous 80 MHz channels, which may be referred to as an 80+80 configuration. For the 80+80 configuration, after channel encoding, the data may be passed through a segment parser that may separate the data into two streams. Inverse Fast Fourier Transform (IFFT) processing and time-domain processing may be performed separately on each stream. The streams may be mapped to two 80 MHz channels, and the data may be transmitted by the transmitting STA. At the receiver of the receiving STA, the above operations for the 80+80 configuration may be reversed, and the combined data may be transmitted to the Medium Access Control (MAC).
[0061] Sub-1 GHz operating modes are supported by 802.11af and 802.11ah. The channel operating bandwidths and carriers are reduced in 802.11af and 802.11ah compared to those used in 802.11n and 802.11ac. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV White Space (TVWS) spectrum, while 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. According to representative embodiments, 802.11ah may support metered control / machine-type communications, such as MTC devices, within a macro coverage area. MTC devices may have limited functionality, including specific features, such as support for (e.g., support only) specific bandwidths and / or limited bandwidths. MTC devices may include batteries with above-threshold battery life (e.g., to maintain very long battery life).
[0062] WLAN systems that can support multiple channels and channel bandwidths, such as 802.11n, 802.11ac, 802.11af, and 802.11ah, include a channel that can be designated as a primary channel. The primary channel can have a bandwidth equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel may be configured and / or limited by the STA from among all STAs operating in the BSS that support the smallest bandwidth operating mode. In the 802.11ah example, the primary channel of a STA (e.g., an MTC-type device) that supports (e.g., only supports) 1 MHz mode may be 1 MHz wide, even if the AP and other STAs in the BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operating modes. Carrier sensing and / or Network Allocation Vector (NAV) configuration may depend on the status of the primary channel. For example, if the primary channel is busy because a STA (that only supports 1 MHz mode of operation) is transmitting to the AP, the entire available frequency band may be considered busy, even though most of the frequency band may remain idle and available.
[0063] In the United States, the available frequency band that can be used with 802.11ah is 902MHz to 928MHz. In South Korea, the available frequency band is 917.5MHz to 923.5MHz. In Japan, the available frequency band is 916.5MHz to 927.5MHz. The total available bandwidth for 802.11ah is 6MHz to 26MHz depending on the country code.
[0064] 1D is a system diagram illustrating the RAN 113 and the CN 115 in accordance with an embodiment. As described above, the RAN 113 may communicate with the WTRUs 102a, 102b, and 102c over the air interface 116 using NR radio technology. The RAN 113 may also communicate with the CN 115.
[0065] The RAN 113 may include gNBs 180a, 180b, and 180c, although it will be understood that the RAN 113 may include any number of gNBs while remaining consistent with an embodiment. The gNBs 180a, 180b, and 180c may each include one or more transceivers for communicating with the WTRUs 102a, 102b, and 102c over the air interface 116. In one embodiment, the gNBs 180a, 180b, and 180c may implement MIMO technology. For example, the gNBs 180a and 180b may utilize beamforming to transmit signals to and / or receive signals from the gNBs 180a, 180b, and 180c. Thus, the gNB 180a may transmit wireless signals to and / or receive wireless signals from the WTRU 102a using, for example, multiple antennas. In an embodiment, the gNBs 180a, 180b, and 180c may implement carrier aggregation technology. For example, the gNB 180a may transmit multiple component carriers to the WTRU 102a (not shown). A subset of these component carriers may be on unlicensed spectrum, and the remaining component carriers may be on licensed spectrum. In an embodiment, the gNBs 180a, 180b, and 180c may implement Coordinated Multi-Point (CoMP) technology. For example, the WTRU 102a may receive coordinated transmissions from the gNBs 180a and 180b (and / or 180c).
[0066] The WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using transmissions associated with scalable numerology. For example, the OFDM symbol spacing and / or OFDM subcarrier spacing may vary for different transmissions, different cells, and / or different portions of the wireless transmission spectrum. The WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using subframes or transmission time intervals (TTIs) of different or scalable lengths (e.g., different lengths of absolute time including and / or lasting different numbers of OFDM symbols).
[0067] The gNBs 180a, 180b, 180c can be configured to communicate with the WTRUs 102a, 102b, 102c in a standalone configuration and / or a non-standalone configuration. In a standalone configuration, the WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c without accessing another RAN (e.g., eNodeBs 160a, 160b, 160c, etc.). In a standalone configuration, the WTRUs 102a, 102b, 102c may utilize one or more of the gNBs 180a, 180b, 180c as mobility anchor points. In a standalone configuration, the WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using signals in unlicensed bands. In a non-standalone configuration, the WTRUs 102a, 102b, 102c may communicate with and connect to a gNB 180a, 180b, 180c while also communicating with and connecting to another RAN, such as an eNodeB 160a, 160b, 160c. For example, the WTRUs 102a, 102b, 102c may implement the DC principle to communicate with one or more gNBs 180a, 180b, 180c and one or more eNodeBs 160a, 160b, 160c substantially simultaneously. In a non-standalone configuration, the eNodeBs 160a, 160b, 160c may act as mobility anchors for the WTRUs 102a, 102b, 102c, and the gNBs 180a, 180b, 180c may provide additional coverage and / or throughput for serving the WTRUs 102a, 102b, 102c.
[0068] Each of the gNBs 180a, 180b, 180c may be associated with a particular cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, scheduling of users in the UL and / or DL, support for network slicing, dual connectivity, interworking between NR and E-UTRA, routing of user plane data to User Plane Functions (UPFs) 184a, 184b, access and routing of control plane information to Mobility Management Functions (AMFs) 182a, 182b, etc. As shown in FIG. 1D , the gNBs 180a, 180b, 180c may communicate with each other via an Xn interface.
[0069] 1D may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and possibly a Data Network (DN) 185a, 185b. While each of the foregoing elements is shown as part of the CN 115, it will be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.
[0070] The AMF 182a, 182b may be connected to one or more gNBs 180a, 180b, 180c in the RAN 113 via an N2 interface and may function as a control node. For example, the AMF 182a, 182b may be responsible for authenticating users of the WTRUs 102a, 102b, 102c, supporting network slicing (e.g., handling different PDU sessions with different requirements), selecting a particular SMF 183a, 183b, registration area, termination of NAS signaling, mobility management, etc. Network slicing may be used by the AMF 182a, 182b to customize the CN support of the WTRUs 102a, 102b, 102c based on the type of service being utilized by the WTRUs 102a, 102b, 102c. Different network slices may be established for different use cases, such as services relying on ultra-reliable low latency (URLLC) access, services relying on enhanced massive mobile broadband (eMBB) access, services with machine type communication (MTC) access, etc. The AMF 162 may provide a control plane function for switching between the RAN 113 and other RANs (not shown) that employ other radio technologies, such as LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies like WiFi.
[0071] The SMFs 183a and 183b may be connected to the AMFs 182a and 182b in the CN 115 via an N11 interface. The SMFs 183a and 183b may also be connected to the UPFs 184a and 184b in the CN 115 via an N4 interface. The SMFs 183a and 183b may select and control the UPFs 184a and 184b and configure the routing of traffic through the UPFs 184a and 184b. The SMFs 183a and 183b may perform other functions, such as managing and allocating UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, providing downlink data notification, etc. The PDU session type may be IP-based, non-IP-based, Ethernet-based, etc.
[0072] The UPFs 184a, 184b may be connected to one or more gNBs 180a, 180b, 180c in the RAN 1113 via an N3 interface, which may provide the WTRUs 102a, 102b, 102c with access to packet-switched networks such as the Internet 110 to facilitate communications between the WTRUs 102a, 102b, 102c and IP-enabled devices. The UPFs 184, 184b may perform other functions such as routing and forwarding packets, enforcing user plane policy, supporting multi-homed PDU sessions, handling user plane QoS, buffering downlink packets, providing mobility anchors, etc.
[0073] The CN 115 may facilitate communication with other networks. For example, the CN 115 may include or communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between the CN 115 and the PSTN 108. Additionally, the CN 115 may provide the WTRUs 102a, 102b, 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, the WTRUs 102a, 102b, 102c may be connected to local data networks (DNs) 185a, 185b via UPFs 184a, 184b via an N3 interface to the UPFs 184a, 184b and an N6 interface between the UPFs 184a, 184b and the DNs 185a, 185b.
[0074] 1A-1D and the corresponding description thereof, one or more, or all, of the functions described herein with respect to one or more of the WTRUs 102a-d, base stations 114a-b, eNodeBs 160a-c, MME 162, SGW 164, PGW 166, gNBs 180a-c, AMFs 182a-b, UPFs 184a-b, SMFs 183a-b, DNs 185a-b, and / or any other devices described herein may be performed by one or more emulation devices (not shown). The emulation devices may be one or more devices configured to emulate one or more, or all, of the functions described herein. For example, the emulation devices may be used to test other devices and / or simulate network and / or WTRU functions.
[0075] The emulation device may be designed to perform one or more tests of other devices in a lab environment and / or an operator network environment. For example, one or more emulation devices may perform one or more or all functions while fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices in the communication network. One or more emulation devices may perform one or more or all functions while temporarily implemented / deployed as part of a wired and / or wireless communication network. The emulation device may be directly coupled to another device for testing purposes and / or may perform testing using over-the-air wireless communication.
[0076] One or more emulation devices may perform one or more functions, inclusive, while not being implemented / deployed as part of a wired and / or wireless communication network. For example, the emulation devices may be utilized in test scenarios in a test lab and / or in an undeployed (e.g., test) wired and / or wireless communication network to perform testing of one or more components. One or more emulation devices may be test equipment. Direct RF coupling and / or wireless communication via RF circuitry (e.g., which may include one or more antennas) may be used by the emulation devices to transmit and / or receive data.
[0077] This application describes various aspects, including tools, features, examples, models, approaches, and the like. Many of these aspects are described with specificity, often in a definitive manner, to at least illustrate their individual characteristics. However, this is for clarity of description and does not limit the application or scope of these aspects. In fact, all of the different aspects may be combined and interchanged to provide further aspects. Furthermore, these aspects may be combined and interchanged with aspects described in previous applications.
[0078] Aspects described and contemplated in this application may be implemented in many different forms. Figures 5-16 described herein may provide some examples, but other examples are contemplated. The discussion of Figures 5-16 is not intended to limit the breadth of implementations. At least one of these aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a generated or encoded bitstream. These and other aspects may be implemented as a method, an apparatus, a computer-readable storage medium having stored thereon instructions for encoding or decoding video data according to any of the described methods, and / or a computer-readable storage medium having stored thereon a bitstream generated according to any of the described methods.
[0079] In this application, the terms "reconstructed" and "decoded" may be used interchangeably, the terms "pixel" and "sample" may be used interchangeably, and the terms "image," "picture," and "frame" may be used interchangeably.
[0080] Various methods are described herein, each of which includes one or more steps or acts for achieving the described method. Unless a specific order of steps or acts is required for the successful operation of the method, the order and / or use of specific steps and / or acts may be modified or combined. Additionally, terms such as “first,” “second,” etc. may be used in various embodiments to modify elements, components, steps, operations, etc., e.g., “first decode” and “second decode.” The use of such terms does not imply a modified order of operations unless specifically required. Thus, in this example, the first decode need not be performed before the second decode, but could occur, for example, before, during, or during an overlapping period with the second decode.
[0081] Various methods and other aspects described herein can be used to modify modules, such as the decoding modules of video encoder 200 and decoder 300, as shown in Figures 2 and 3. Furthermore, the subject matter disclosed herein may apply to any type, format, or version of video coding, whether described in a standard or recommendation, existing or developed in the future, and extensions of any such standard or recommendation. Unless otherwise indicated or technically excluded, aspects described herein may be used individually or in combination.
[0082] Various numerical values are used in the examples described in this application, such as the number of weights, weight values, number of taps in the filter, number of bits in the content, scalar offsets, matrix dimensions, etc. These and other specific values are for purposes of describing the examples, and the described aspects are not limited to these specific values.
[0083] 2 illustrates an exemplary video encoder. While variations of exemplary encoder 200 are contemplated, encoder 200 is described below for clarity without describing all possible variations.
[0084] Before being encoded, the video sequence may undergo a pre-encoding process (201), such as applying a color transformation to the input color picture (e.g., from RGB 4:4:4 to YCbCr 4:2:0) or performing a remapping of the input picture components to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata may be associated with the pre-processing and may be attached to the bitstream.
[0085] In the encoder 200, a picture is coded by the encoder elements as described below. The picture to be coded is divided (202) and processed in units, e.g., coding units (CUs). Each unit is coded, e.g., using either intra mode or inter mode. If the unit is coded in intra mode, intra prediction is performed (260). In inter mode, motion estimation (275) and motion compensation (270) are performed. The encoder decides (205) whether to use intra mode or inter mode to code the unit, and indicates the intra / inter decision, e.g., by a prediction mode flag. A prediction residual is calculated, e.g., by subtracting (210) the predicted block from the original image block.
[0086] The prediction residual is then transformed (225) and quantized (230). The quantized transform coefficients, plus motion vectors and other syntax elements, are entropy coded (245) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., the residual is coded directly without applying a transform or quantization process.
[0087] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inverse transformed (250), and the prediction residual is decoded. The decoded prediction residual is combined (255) with the prediction block to reconstruct an image block. An in-loop filter (265) is applied to the reconstructed picture, for example, to perform deblocking / sample adaptive offset (SAO) filtering to reduce coding artifacts. The filtered image is stored in a reference picture buffer (280).
[0088] Figure 3 illustrates an example of a video decoder. In the exemplary decoder 300, the bitstream is decoded by elements of the decoder as described below. The video decoder 300 generally performs a decoding pass that is the inverse of the encoding pass as described in Figure 2. The encoder 200 also generally performs video decoding as part of encoding the video data.
[0089] In particular, the decoder's input includes a video bitstream, such as may be generated by the video encoder 200. First, the bitstream is entropy decoded (330) to obtain transform coefficients, motion vectors, and other coded information. Picture partition information indicates how the picture is partitioned. The decoder may then partition the picture according to the decoded picture partition information (335). The transform coefficients are inverse quantized (340) and inverse transformed (350), and the prediction residual is decoded. The decoded prediction residual is combined with a predicted block (355) to reconstruct an image block. The predicted block may be obtained from intra prediction (360) or motion-compensated prediction (i.e., inter prediction) (375) (370). An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (380).
[0090] The decoded picture may further undergo a post-decoding process (385), such as an inverse color transform (e.g., converting from YCbCr 4:2:0 to RGB 4:4:4) or an inverse remapping that performs the inverse of the remapping process performed in the pre-encoding process (201). The post-decoding process may use metadata derived in the pre-encoding process and signaled in the bitstream. In one example, the decoded image (e.g., after application of the in-loop filter (365) and / or the post-decoding process (385), if a post-decoding process is used) may be sent to a display device for rendering to a user.
[0091] FIG. 4 illustrates an example system in which various aspects and embodiments described herein may be implemented. System 400 may be embodied as a device including various components described below and configured to perform one or more of the aspects described herein. Examples of such devices include, but are not limited to, various electronic devices, such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television sets, personal video recording systems, connected home appliances, and servers. The elements of system 400, singly or in combination, may be embodied in a single integrated circuit (IC), multiple ICs, and / or separate components. For example, in at least one example, the processing and encoder / decoder elements of system 400 are distributed across multiple ICs and / or separate components. In various examples, system 400 is communicatively coupled to one or more other systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various examples, system 400 is configured to perform one or more of the aspects described herein.
[0092] The system 400 includes at least one processor 410 configured to execute loaded instructions, for example, to implement various aspects described herein. The processor 410 may include embedded memory, input / output interfaces, and various other circuitry known in the art. The system 400 includes at least one memory 420 (e.g., a volatile memory device and / or a non-volatile memory device). The system 400 includes a storage device 440, which may include non-volatile memory and / or volatile memory, including, but not limited to, Electrically Erasable Programmable Read-Only Memory (EEPROM), Read-Only Memory (ROM), Programmable Read-Only Memory (PROM), Random Access Memory (RAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Flash, magnetic disk drives, and / or optical disk drives. Storage devices 440 may include, by way of non-limiting example, internal storage devices, attached storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices.
[0093] System 400 includes an encoder / decoder module 430 configured to process data to provide, for example, encoded video or decoded video, which may include its own processor and memory. Encoder / decoder module 430 represents a module that may be included in a device that performs encoding and / or decoding functions. As is known, a device may include one or both of an encoding and decoding module. Additionally, encoder / decoder module 430 may be implemented as a separate element of system 400 or may be incorporated within processor 410 as a combination of hardware and software, as known to those skilled in the art.
[0094] Program code loaded into the processor 410 or the encoder / decoder 430 to perform various aspects described herein may be stored in the storage device 440 and subsequently loaded into the memory 420 for execution by the processor 410. According to various examples, one or more of the processor 410, the memory 420, the storage device 440, and the encoder / decoder module 430 may store one or more of various items during the execution of processes described herein. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, expressions, operations, and operational logic.
[0095] In some examples, memory internal to the processor 410 and / or the encoder / decoder module 430 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other examples, memory external to the processing device (e.g., the processing device can be either the processor 410 or the encoder / decoder module 430) is used for one or more of these functions. The external memory can be the memory 420 and / or the storage device 440, and can be, for example, dynamic volatile memory and / or non-volatile flash memory. In some examples, external non-volatile flash memory is used to store, for example, the television's operating system. In at least one example, high-speed external dynamic volatile memory, such as RAM, is used as working memory for video encoding and decoding operations.
[0096] Input to the elements of system 400 may be provided through various input devices, as shown in block 445. Such input devices include, but are not limited to, (i) a radio frequency (RF) section that receives, for example, RF signals transmitted over the air by a broadcast station, (ii) a component (COMP) input terminal (or set of COMP input terminals), (iii) a universal serial bus (USB) input terminal, and / or (iv) a High Definition Multimedia Interface (HDMI) input terminal. Other examples not shown in FIG. 4 include composite video.
[0097] In various examples, the input devices of block 445 have associated respective input processing elements as known in the art. For example, the RF section may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal or bandlimiting a signal to a frequency band), (ii) downconverting the selected signal, (iii) bandlimiting again to a narrower frequency band to select a signal frequency band, which in certain examples may be referred to as a channel (for example), (iv) demodulating the downconverted, bandlimited signal, (v) performing error correction, and / or (vi) demultiplexing to select a desired data packet stream. The RF section in various examples includes one or more elements that perform these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a downconverter, a demodulator, an error corrector, and a demultiplexer. The RF section may include, for example, a tuner that performs various of these functions, including downconverting a received signal to a lower frequency (e.g., an intermediate frequency or a frequency near baseband) or to baseband. In one set-top box example, the RF section and its associated input processing elements perform frequency selection by receiving, filtering, downconverting, and re-filtering RF signals transmitted over a wired (e.g., cable) medium to a desired frequency band. In various examples, the order of the above (and other) elements is rearranged, some of these elements are removed, and / or other elements that perform similar or different functions are added. Adding elements can include inserting elements between existing elements, such as inserting amplifiers and analog-to-digital converters. In various examples, the RF section includes an antenna.
[0098] The USB and / or HDMI terminals may include respective interface processors for connecting system 400 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing, e.g., Reed-Solomon error correction, may be performed, for example, in a separate input processing IC or within processor 410, as desired. Similarly, aspects of USB or HDMI interface processing may be performed, as desired, in a separate interface IC or within processor 410. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 410 and an encoder / decoder 430, which operates in combination with memory and storage elements to process the data stream as desired for display on an output device.
[0099] The various elements of the system 400 may be provided within a unitary housing in which the various elements are interconnected and capable of transmitting data between them using a suitable connection arrangement 425, for example, an internal bus as known in the art, including an Inter-IC (I2C) bus, wiring, and printed circuit boards.
[0100] System 400 includes a communication interface 450 that enables communication with other devices over a communication channel 460. Communication interface 450 may include, but is not limited to, a transceiver configured to transmit and receive data over communication channel 460. Communication interface 450 may include, but is not limited to, a modem or a network card, and communication channel 460 may be implemented in a wired and / or wireless medium, for example.
[0101] Data, in various examples, is streamed or otherwise provided to system 400 using a wireless network such as a Wi-Fi network, e.g., IEEE 802.11 (IEEE (Institute of Electrical and Electronics Engineers) refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signal in these examples is received via communication channel 460 and communication interface 450 adapted for Wi-Fi communication. Communication channel 460 in these examples is typically connected to an access point or router that provides access to external networks, including the Internet, enabling streaming applications and other over-the-top communications. Another example provides streamed data to system 400 using a set-top box that delivers data via an HDMI connection in input block 445. Yet another example provides streamed data to system 400 using an RF connection in input block 445. As noted above, various examples provide data in a non-streaming manner. Additionally, various examples use wireless networks other than Wi-Fi, such as a cellular network or a Bluetooth® network.
[0102] The system 400 can provide output signals to various output devices, including a display 475, speakers 485, and other peripheral devices 495. Various example displays 475 include, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 475 may be for a television, a tablet, a laptop, a mobile phone, or other device. The display 475 may also be integrated with other components (e.g., as in a smartphone) or may be separate (e.g., an external monitor for a laptop). The other peripheral devices 495, in various examples, include one or more of a standalone digital video disc (or digital versatile disc) (both terms DVD, digital versatile disc), a disc player, a stereo system, and / or a lighting system. Various examples use one or more peripheral devices 495 to provide functionality based on the output of the system 400. For example, a disc player performs the function of playing the output of the system 400.
[0103] In various examples, control signals are communicated between system 400 and display 475, speakers 485, or other peripheral devices 495 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable inter-device control with or without user intervention. Output devices may be communicatively coupled to system 400 via dedicated connections through respective interfaces 470, 480, and 490. Alternatively, output devices may be connected to system 400 using communication channel 460 via communication interface 450. Display 475 and speakers 485 may be integrated into a single unit with other components of system 400 in an electronic device such as a television. In various examples, display interface 470 includes a display driver, such as a timing controller (TCon) chip.
[0104] Display 475 and speakers 485 may alternatively be separate from one or more of the other components, for example, if the RF portion of input 445 is part of a separate set-top box. In various examples where display 475 and speakers 485 are external components, output signals may be provided via dedicated output connections, including, for example, an HDMI port, a USB port, or a COMP output.
[0105] The examples may be implemented by computer software executed by the processor 410, or by hardware, or by a combination of hardware and software. As a non-limiting example, the examples may be implemented by one or more integrated circuits. The memory 420 may be of any type appropriate to the technological environment and may be implemented using any suitable data storage technology, such as, by way of non-limiting examples, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. The processor 410 may be of any type appropriate to the technological environment and may include, by way of non-limiting examples, one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.
[0106] Various implementations involve decoding. As used herein, "decoding" can encompass all or some of the processes performed on a received encoded sequence to generate a final output suitable for display, for example. In various examples, such processes include one or more of the processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding.
[0107] As a further example, in one example, "decoding" refers only to entropy decoding, in another example, "decoding" refers only to differential decoding, and in another example, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to refer specifically to a subset of operations or to the broader decoding process generally will be clear based on the context of the particular description and will be well understood by one of ordinary skill in the art.
[0108] Various implementations involve encoding. Similar to the above discussion of "decoding," "encoding," as used herein, can encompass all or some of the processes performed on an input video sequence to generate, for example, an encoded bitstream. In various examples, such processes include one or more processes typically performed by an encoder, such as segmentation, differential encoding, transform, quantization, and entropy coding.
[0109] As a further example, in one example, "encoding" refers only to entropy encoding, in another example, "encoding" refers only to differential encoding, and in another example, "encoding" refers to a combination of differential encoding and entropy encoding. Whether the phrase "encoding process" is intended to refer specifically to a subset of operations or to the broader encoding process in general will be clear based on the context of the particular description and will be well understood by one of ordinary skill in the art.
[0110] It should be noted that the syntax elements used herein are descriptive terms, and therefore they do not preclude the use of other syntax element names.
[0111] When a figure is presented as a flow diagram, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flow diagram of the corresponding method / process.
[0112] Implementations and aspects described herein may be realized in, for example, a method or process, an apparatus, a software program, a data stream, or a signal. Even if discussed in the context of only one type of implementation (e.g., discussed only as a method), the implementation of the discussed features may also be realized in other forms (e.g., an apparatus or a program). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. The method may be performed by, for example, a processor, which generally refers to a processing device including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, for example, computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end users.
[0113] References to "one example" or "example," or "one implementation" or "an implementation," as well as other variations thereof, mean that a particular feature, structure, characteristic, etc. described in connection with the example is included in at least one example. Thus, appearances of the phrases "in one example" or "in an example," or "in one implementation" or "in an implementation," as well as any other variations, in various places throughout this application are not necessarily all referring to the same example.
[0114] Additionally, the application may refer to "determining" various information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from memory. Obtaining may include receiving, retrieving, constructing, generating, and / or determining.
[0115] Additionally, the application may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0116] Additionally, the present application may refer to "receiving" various information. Receiving, like "accessing," is intended to be a broad term. Receiving information can include, for example, one or more of accessing information or retrieving information (e.g., from memory). Furthermore, "receiving" typically involves in some manner, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0117] For example, in the case of "A / B," "A and / or B," and "at least one of A and B," it should be understood that the use of any of the following " / ," "and / or," and "at least one of" is intended to encompass the selection of only the first listed option (A), or the selection of only the second listed option (B), or the selection of both options (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C," such language is intended to encompass the selection of only the first listed option (A), or the selection of only the second listed option (B), or the selection of only the third listed option (C), or the selection of only the first and second listed options (A and B), or the selection of only the first and third listed options (A and C), or the selection of only the second and third listed options (B and C), or the selection of all three options (A, B, and C). This may be expanded as many times as the number of listed items, as would be apparent to one of ordinary skill in the art.
[0118] Also, as used herein, the word "signal" refers to, among other things, indicating something to a corresponding decoder. The encoder signal may include, for example, a residual signal, metadata, an intra-prediction signal, a selected region for reference samples, etc. In this way, in one example, the same parameters are used on both the encoder and decoder sides. Thus, for example, the encoder can transmit specific parameters to the decoder (explicit signaling) so that the decoder can use the same specific parameters. Conversely, if the decoder already has specific parameters as well as other parameters, signaling can be used without transmission (implicit signaling) to enable the decoder to easily recognize and select specific parameters. By avoiding the transmission of any actual functions, bit savings are realized in various examples. It should be understood that signaling can be achieved in various manners. For example, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder in various examples. Although the above relates to the verb form of the word "signal," the word "signal" can also be used as a noun in this specification.
[0119] As will be apparent to those skilled in the art, implementations may generate a wide variety of signals formatted to carry information that may be stored or transmitted, for example. The information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal may be formatted to carry a bit stream of the described examples. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal may be transmitted over a wide variety of different wired or wireless links, as is known. The signal may be stored on, accessed from, or received from a processor-readable medium.
[0120] Many examples are described herein. Features of the examples may be provided alone or in any combination across various claim categories and types. Furthermore, examples may include one or more of the features, devices, or aspects described herein, alone or in any combination across various claim categories and types. For example, features described herein may be embodied in a bitstream or signal including information generated as described herein. The information enables a decoder to decode the bitstream, and the encoder, bitstream, and / or decoder are according to any of the described embodiments. For example, features described herein may be implemented by creating and / or transmitting and / or receiving and / or decoding a bitstream or signal. For example, features described herein may be implemented by a method, process, apparatus, medium storing instructions, medium storing data, or signal. For example, features described herein may be implemented by a TV, a set-top box, a mobile phone, a tablet, or other electronic device that performs decoding. A TV, set-top box, mobile phone, tablet, or other electronic device may display (e.g., using a monitor, screen, or other type of display) the resulting image (e.g., an image from the residual reconstruction of the video bitstream). A TV, set-top box, mobile phone, tablet, or other electronic device may receive a signal including the encoded image and perform decoding.
[0121] The results of intra prediction of DC, planar, and / or other angle modes may be modified, for example, by a position dependent intra prediction combination (PDPC) method. PDPC is an intra prediction method. PDPC may invoke a combination of boundary reference samples and intra prediction using filtered boundary reference samples. PDPC may be applied, for example, without signaling, to one or more of the following intra modes: planar, DC, intra angles less than or equal to horizontal, or intra angles greater than or equal to vertical and less than or equal to 80. PDPC may not be applied, for example, if the current block is in block-based delta pulse code modulation (BDPCM) mode and / or if the multiple reference line (MRL) index is greater than zero.
[0122] The prediction sample pred(x',y') may be predicted using a linear combination of the intra prediction mode (eg, DC, planar, angular) and the reference, for example, according to equation (1).
number
[0123] Referring to equation (1), R x,-1 , R -1,y may represent the reference samples located at the upper and left boundaries of the current sample (x, y), respectively.
[0124] Boundary filters (e.g., additional boundary filters) may not be needed, for example, when PDPC is applied to DC, planar, horizontal, and / or vertical intra modes (as they may be in the case of DC mode boundary filters or horizontal / vertical mode edge filters). The PDPC process for DC mode and planar mode may be similar (e.g., identical). If the current angular mode is HOR_IDX or VER_IDX, respectively, the left or top reference sample may not be used. The PDPC weight and scale factor may depend on the prediction mode and / or block size. PDPC may be applied to blocks with width and height greater than or equal to 4.
[0125] 5 shows an exemplary definition of samples used by PDPC applied to the diagonal-top-right mode. In particular, FIG. 5 illustrates an example of the definition of reference samples (e.g., Rx,-1 and R-1,y) for PDPC applied across various prediction modes. A prediction sample pred(x', y') may be located at (x', y') within a prediction block. As an example, the coordinate x of the reference sample Rx,-1 is obtained by x=x'+y'+1, and the coordinate y of the reference sample R-1,y is obtained by y=x'+y'+1 (e.g., for the diagonal mode). The reference sample R x,-1 and R -1,y may be located at fractional sample positions for other angle modes. The sample value at the nearest integer sample position may be used (e.g., for fractional sample positions).
[0126] Template-based intra mode derivation (TIMD) may be performed. Figure 6A shows an example 600 of template areas 620 (e.g., in light gray) and reference samples 610 (e.g., in dark gray) that may be used to derive TIMD modes and associated weights. Figure 6B shows an example of reference samples (e.g., in dark gray) that may be used to derive intra prediction for a current block.
[0127] A sum of absolute transform differences (SATD) between a luma prediction constructed using reconstructed reference samples of template 610 and reconstructed samples of template 620 may be calculated for each intra-prediction mode (e.g., each intra-prediction mode) in a most probable mode (MPM) (e.g., supplemented with default modes PLANAR and DC, if necessary). The first (e.g., two) intra-prediction mode (IPM) with the smallest SATD may be selected as the TIMD mode. The TIMD modes (e.g., two TIMD modes) may be fused with weights (e.g., weight1, weight2), for example, after applying a PDPC process. The weighted intra-prediction may be used to code the current CU. PDPC may be included in the derivation of the TIMD mode.
[0128] The costs of the (e.g., two) selected modes (e.g., costMode1, costMode2) may be compared to a threshold. For example, a cost factor for TIMD Mode 2 may be applied as follows: A condition may be, for example, costMode2<2*costMode1. Fusion may be applied, for example, if the condition is true. Mode 1 may be used, for example, if the condition is not true. The weights of the modes may be calculated from their SATD costs, for example, according to equations (2) and (3) as follows: Weight 1 = costMode2 / (costMode1+costMode2) (2) Weight 2 = 1 - Weight 1 (3)
[0129] In some examples, the number of weights may be increased to N>2. A can be determined as the IPM with the smallest SATD using the template above, and / or L may be determined as the IPM with the smallest SATD using the left template. The fusion weights may vary spatially, for example, as a linear function of the distance to the top / left edge. The fusion weights may be determined, for example, according to Equations (4), (5), and (6) as follows: SATDA >K.SATD L Then,
number
number
number
[0130] Referring to equations (4) and (5), K may be a predetermined value (eg, equal to 2), and (W×H) may be the current block size.
[0131] Decoder-side intra mode derivation (DIMD) may be performed. The intra modes (e.g., two intra modes) may be derived, for example, from a histogram of gradients (HoG) of reconstructed neighboring samples when DIMD is applied. The HoG calculation may be performed, for example, by applying horizontal and vertical Sobel filters to pixels within a template (e.g., of width 3) around the block. The intra prediction mode corresponding to the highest (e.g., two) histogram bars may be selected for the block.
[0132] 7 shows an example 700 of derivation of blending weights for DIMD. As shown, (e.g., two) predictors may be combined with a planar mode predictor. Weights may be derived from the gradient, for example, as follows: (i) The planar weight may be fixed at 2 1 / 64 (approximately 1 / 3). (ii) The remaining weight of 4 3 / 64 (approximately 2 / 3) may be shared between (e.g., two) HoG IPMs in proportion to the amplitude of their HoG bars (e.g., as shown in FIG. 7).
[0133] The derived intra mode may be included in a primary list of intra most probable modes (MPMs). The DIMD process may, for example, be performed before the MPM list is built. The primary derived intra mode of a DIMD block may be stored with the block. The primary derived intra mode of a DIMD block is used for building the MPM lists of neighboring blocks.
[0134] Multiple HoGs may be calculated. In some examples, three HoGs may be calculated, such as a HoG with a top template, a HoG with a left template, and a HoG with a top+left template, which may allow for determining, for an IPM (e.g., each of two selected IPMs), whether the IPM depends on a particular template region. i The location dependency of may be defined, for example, according to all or part of the following logic: IPMi may be dependent on region ABOVE if, for example, (Habove[IPMi]>2.Hleft[IPMi]). i For example, (H left [IPM i ]>2.H above [IPM i ]), it may depend on the region LEFT, and / or IPMi may not be location-dependent otherwise, for example.
[0135] For example, if IPMi is position (top / left) dependent, the value of weighti can be adjusted (e.g., as a linear function of the distance to the top / left edge). i If σ is position-dependent, it may vary spatially (eg, according to equation (7)).
number
[0136] The weights are, for example, IPM i If is position-dependent, it may vary spatially (e.g., according to equation (8)).
number
[0137] Referring to equations (7) and (8), Δ i may be a predefined weight range variation, and (W×H) may be the current block size.
[0138] Figure 8 shows an example (800) of using neighboring reconstructed samples for a DIMD chroma mode. As shown in Figure 8, the DIMD chroma mode can use the DIMD derivation method to derive a chroma intra prediction mode for a current block based on neighboring reconstructed Y, Cb, and Cr samples in a second neighboring row and column. Horizontal and vertical gradients can be calculated for (e.g., each) co-located reconstructed luma sample and reconstructed Cb and Cr samples of the current chroma block. The gradients can be used to construct an HoG. The intra prediction mode with the largest histogram amplitude value can then be used to perform chroma intra prediction for the current chroma block.
[0139] A convolutional cross-component model (CCCM) may be implemented for intra prediction. CCCM may predict chroma samples from reconstructed luma samples, similar to CCLM, for example. The reconstructed luma samples may be downsampled (e.g., as with CCLM) to match a lower resolution chroma grid, for example, if / when chroma subsampling is used.
[0140] Single-model or multi-model variants of CCCM may be used (e.g., similar to CCLM). The multi-model variant may use multiple (e.g., two) models, which may include a model derived for samples above the average luma reference value and another model for the rest of the samples (e.g., similar to CCLM). The multi-model CCCM mode may be selected for PUs that have at least 128 reference samples available.
[0141] 9 shows an example of the spatial portion of a convolution filter (900). CCCM may use a convolution 7-tap filter that may include a 5-tap plus sign shape spatial component, a nonlinear term, and a bias term. The input to the spatial 5-tap component of the filter may include, for example, a chroma sample to be predicted and a center (C) luma sample that may be co-located with its above / north (N), below / south (S), left / west (W), and right / east (E) neighbors, as shown by the example of FIG. 9.
[0142] The nonlinear term P may be expressed as a power of two of the central luma sample C, eg, scaled to the sample value range of the content, eg, according to equation (9). P=(C*C+midVal)>>bitDepth (9)
[0143] In the example of 10-bit content, the nonlinear term P may be calculated as P=(C*C+512)>>10. The bias term B may represent a scalar offset between the input and the output (e.g., similar to the offset term in CCLM). The bias term B may be set to the midpoint chroma value, midVal (e.g., 512 for 10-bit content), or some other value (e.g., 256 for 10-bit content).
[0144] The output of the filter is the filter coefficient c i and the input value clipped to the range of valid chroma samples, for example according to equation (10).
number
[0145] filter coefficient c i may be calculated by minimizing the mean square error (MSE) between predicted samples in the reference area and reconstructed samples (e.g., chroma samples). FIG. 10 shows an example reference area (e.g., with its padding) and a PU 1000. The reference area (e.g., with its padding) may be used to derive filter coefficients. An example reference area 1010 (e.g., as shown in light gray) may include six lines / columns of chroma samples above and to the left of the PU. The reference area may extend, for example, one PU width to the right and one PU height below the PU boundary. The area may be adjusted to include (e.g., only) available samples. Extension to the reference area 1020 (e.g., as shown in dark gray) may be needed to support the "side samples" of a plus-shaped spatial filter. The extension may be padded, for example, if / when it falls into unavailable areas.
[0146] The MSE minimization may be performed, for example, by calculating the autocorrelation matrix of the luma input and the cross-correlation vector between the luma input and the chroma output, which is a i where a i =C,N,S,E,W,P,B, and can be expressed, for example, as follows.
number
[0147] The autocorrelation matrix may be subjected to LDL decomposition. The final filter coefficients may be calculated, for example, using backsubstitution. This process may be similar to the calculation of the ALF filter coefficients. LDL decomposition may be used (e.g., instead of Cholesky decomposition) to avoid using square root operations, for example. The calculation may use integer arithmetic (e.g., only integer arithmetic).
[0148] In some examples (e.g., gradient and location based convolutional cross-component model (GL-CCCM)), the GL-CCCM filter for prediction may follow equation (11):
number
[0149] 11 shows an example of spatial samples used in GL-CCCM. Referring to Equation (11), Gy and Gx may be vertical and horizontal gradients, respectively. Gy and Gx may be calculated according to Equation (12) and Equation (13), respectively. Gy = (2N + NW + NE) - (2S + SW + SE) (12) Gx = (2W + NW + SW) - (2 E +NE+SE) (13)
[0150] The Y and X parameters may be the vertical and horizontal positions of the central luma sample. The Y and X parameters may be calculated relative to the top-left coordinate of the block.
[0151] FIG. 12 shows an example of a contiguous region used to derive a CC model.
[0152] In some examples, the neighboring region used to select the reference sample for deriving the CC model may be selected from among a set of predefined regions. For example, the bitstream may signal the region selected from {top+left, top, left} corresponding to modes {CHROMA_IDX, T_IDX, L_IDX} using information single model mode or multi-model mode {MDLM, MMLM}, respectively, as shown by the example of FIG. 12.
[0153] In some examples, the derivation of DIMD and / or TIMD fusion weights may be performed independently for each intra mode. In some examples, intra prediction blending weights for TIMD and / or DIMD processes may be derived using, for example, a regression-based method. Intra prediction samples may be generated from regression-based locally calculated parameters for neighboring samples.
[0154] The reconstructed samples of the template and the corresponding predicted samples may be used to determine blending weights. For example, the weights may be based on minimizing the difference between the predicted samples of the template and the corresponding reconstructed samples.
[0155] Regression-based weights may be derived for the TIMD. For example, N intra-prediction modes may be derived based on multiple reconstructed samples of the template. Prediction blocks for the block corresponding to the derived intra-prediction modes may be obtained.
[0156] The selected N intra-modes M k The N IPM blending weights used in the weighted linear combination of a k} k=0,...N-1 The intra prediction model may support / enable deriving the value of intraPred(x) at position "x" in the template according to equation (14).
number
[0157] Weighted and blended predictions of the template corresponding to the derived intra-prediction modes may be obtained, and the video decoding device may optimize the weights corresponding to each derived intra-prediction mode based on minimizing a difference between the weighted blended prediction of the template and the reconstructed template.
[0158] For example, IPM blending weight a k The value of x may be derived (e.g., calculated) by, for example, minimizing the difference (e.g., MSE) between predicted samples and reconstructed samples (x) (e.g., luma samples) in a reference area. The reference area may be, for example, a template area (e.g., as shown at 620 in FIG. 6A) or an area used for CCCM (e.g., as shown at 1010 in FIG. 10).
[0159] In some examples, an additional bias term B may be added (e.g., as a scalar offset). The bias term B may be set to the intermediate luma value (e.g., 512 for 10-bit content).
[0160] The MSE minimization is performed by, for example, calculating the autocorrelation matrix for the reconstructed luma samples
number
number
[0161] The intra prediction samples for the current block may be obtained, for example, by applying an intra prediction model (e.g., as shown in equation (14)) using N intra predictions constructed using normal reference samples.
[0162] Regression-based weights may be derived for DIMD. For example, the (e.g., two) IPM blending weights (e.g., {a0, a1}) for the selected intra-DIMD mode with the highest histogram bar and the weight {a2} for the PLANAR mode may be derived using, for example, a regression-based method (e.g., as described herein). The reference area may be the area used to calculate the HoG.
[0163] In some examples, the weights for PLANAR may be fixed predetermined values and / or the matrix may be 2 x 2. The value of intraPred(x) may be determined, for example, according to equation (15).
number
[0164] In some instances, a bias may be added.
[0165] The regression-based intra prediction may be calculated as a linear combination of template-derived parameters P(x), for example, according to equation (16).
number
[0166] P k The value of (x) can be, for example, the reconstructed reference sample component (e.g., sample luma) in the same column R(col) (e.g., or the same line R(lig)) as the sample in the template (e.g., col=column x, lig=line x). P k The value of (x) can be, for example, the position (eg, column (col) or line (lig)) of the reconstructed reference sample.
[0167] For example, the parameters may include reconstructed and / or predicted samples neighboring the sample location. Determining intra-predicted samples for a block or a portion of a block (e.g., a unit, a sub-block, a sub-sub-block) may include applying the determined weights to reconstructed and / or predicted samples neighboring the sample location.
[0168] For example, a histogram of gradients associated with samples of the template may be obtained, an intra-prediction mode may be derived based on the histogram of gradients associated with the template, prediction samples of the template may be obtained based on each derived intra-prediction mode, and parameters may correspond to the prediction samples.
[0169] In the example, the weights (e.g., a0.a n ) may correspond to multiple reference samples of the template. The weights may be optimized based on minimizing the difference between the sum of the weighted reference samples and the reconstructed samples of the template.
[0170] Figure 13A shows an example of using a template sample and a reference sample to derive an intra-prediction model. Figure 13B shows an example of using a reference sample to apply an intra-prediction model. Figure 14 shows an example 1400 of iteratively applying a prediction model.
[0171] In some examples, the reference sample that can be used to apply the model can be the same as the reference sample used to derive the model (see, for example, FIG. 13A). In some examples, the reference sample that can be used to apply the model can be a normal intra-reference sample (see, for example, FIG. 13B). In some examples, the reference sample that can be used to derive the model can be the same as the reference sample used to apply the model (see, for example, FIG. 13B).
[0172] An advantage of calculating the intra prediction as a linear combination of template-derived parameters P(x) is that the intra prediction parameters can depend (e.g., only) on neighboring reference samples. This method can be used, for example, to replace the PLANAR mode and / or as an additional intra prediction mode.
[0173] Regression-based iterative intra prediction may be performed. One or more (e.g., several) parameters P(x) may be iteratively derived from previously predicted sample values of the current block, as shown by the example of FIG. 14. For example, the parameters P(x) may include (e.g., comprise entirely) previously predicted samples. In some examples, P k The value of (x) can be a neighboring reconstructed sample (e.g., if "x" is close to the reference sample) or a predicted sample. k The value of (x) can (for example, alternatively) be the local H / V gradient.
[0174] The regression-based model derivation may be performed using local adaptation. As described herein, intra prediction may be spatially adapted, for example, depending on HoG cost (e.g., DIMD) and / or SATD cost (e.g., TIMD). Spatial adaptation may be performed using spatial weighting. For example, three best intra modes may be derived for TIMD, for example, using only the top template, only the left template, and both the top and left templates. The three templates may be fused with weights that may vary locally depending, for example, on the distance of the current sample from a given template.
[0175] DIMD may use the position (e.g., left or top) of the reference sample that contributed to the HoG peak to derive weights that vary locally depending on, for example, the distance of the current sample to the left or top. In some examples (e.g., as a variant of TIMD), left and top templates may be used to select (e.g., two) intra-modes that minimize SATD on the left and top templates, respectively. For example, a final blending of (e.g., two) intra-modes may be performed using weights that vary locally depending on the distance of the current sample to the left or top.
[0176] Weight adaptation can be combined with other examples (e.g., as described herein). Figure 15 shows an example 1500 of using spatially varying weights to derive a blending model. Figure 16 shows an example 1600 of deriving blending {ai} parameters using spatially varying blending weights {wi(x)}.
[0177] 16, IPMi may be derived at 1610. At 1620, IPMi may be evaluated to determine whether it is location dependent.
[0178] If IPMi is location dependent, locally dependent weights wi(x) may be derived at 1630. Blending {ai} parameters may be derived at 1640. Blending may be applied at 1650 (e.g., according to equation (17) below).
[0179] If IPMi is not location-dependent, blending {ai} parameters may be derived (e.g., without locally dependent weights) at 1660. Blending may be applied (e.g., according to equation (16) above) at 1670.
[0180] Previous embodiments have been described, for example, by using spatially adapted weightings w in the IPM blending equations. k(x) can be used to extend (e.g., any) spatial adaptation (e.g., as referenced by 1650 in FIG. 16). The IPM blending equation can determine intraPred(x), for example, according to equation (17).
number
[0181] IPM blending parameters {a i Spatial weighting can be included while deriving the autocorrelation matrix and cross-correlation vector
number
number
number
[0182] w i The value of (x) may correspond to, for example, in the case of a sample of the left template, the value of the spatial weight adaptation of the TIMD or DIMD blending for the first column of the current block (e.g., as referenced by 1510 in FIG. 15). i The value of (x) may correspond, for example, in the case of a sample of the top template, to the value of spatial weight adaptation of TIMD or DIMD blending for the first line of the current block (e.g., as referenced by 1520 in Figure 15).
[0183] Although features and elements are described above in particular combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with the other features and elements. In addition, the methods described herein may be implemented in a computer program, software, or firmware embodied in a computer-readable medium for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted via wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random-access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, optical media such as CD-ROM disks, and digital versatile disks (DVDs). A processor in association with software may be used to implement a radio frequency transceiver for use in a WTRU, UE, terminal, base station, RNC, or any host computer.
Claims
1. 1. A video decoding device, comprising: a processor, the processor comprising: Get the template associated with the block, obtaining a plurality of predicted samples and a plurality of corresponding reconstructed samples of the template; determining a plurality of weights based on minimizing a difference between the predicted samples and the corresponding reconstructed samples of the template; determining an intra prediction model, the intra prediction model including applying the plurality of weights to a respective plurality of parameters; determining intra-predicted samples for the block based on the intra-prediction model, where the predicted samples, the corresponding reconstructed samples, and the intra-predicted samples share a component type; and decoding the block based on the intra-predicted samples.
2. The processor: obtaining a plurality of reference samples of the template; 2. The video decoding device of claim 1, further configured to optimize the weights corresponding to the reference samples of the template based on minimizing a difference between a sum of weighted reference samples and the reconstructed samples of the template.
3. The intra-prediction sample is associated with a sample position within the block, and the plurality of parameters includes a plurality of predicted samples neighboring the sample position, and determining the intra-prediction sample for the block based on the intra-prediction model comprises: The video decoding device of claim 2 , comprising: applying the determined weights to the predicted samples adjacent to the sample position.
4. the intra-predicted sample is a first intra-predicted sample associated with a first sample position within the block, the plurality of parameters includes a plurality of reconstructed samples adjacent to the first sample position, and the processor: for a second sample position within the block, identifying a plurality of predicted samples adjacent to the second sample position; 3. The video decoding device of claim 2, further configured to: apply the determined weights to the predicted samples to determine a second intra-predicted sample for the block, wherein the block is decoded further based on the second intra-predicted sample of the block.
5. The video decoding device of claim 4 , wherein the plurality of predicted samples adjacent to the second sample position includes the first intra-predicted sample.
6. The processor: The method is further configured to derive a plurality of intra prediction modes based on a plurality of reconstructed samples of the template and a plurality of reference samples of the template, wherein the plurality of parameters correspond to a plurality of prediction samples obtained based on the derived intra prediction modes, and determining the intra prediction samples for the block based on the intra prediction model comprises: The video decoding device of claim 1 , further comprising: applying the weights to the prediction samples obtained based on the derived intra-prediction modes.
7. The processor: obtaining a histogram of gradients associated with a plurality of samples of the template; 2. The video decoding device of claim 1, further configured to: derive a plurality of intra-prediction modes based on a histogram of the gradients associated with the template, wherein the plurality of prediction samples of the template are obtained based on the respective derived intra-prediction modes, and the plurality of parameters correspond to a plurality of prediction samples obtained based on the derived intra-prediction modes.
8. The processor: deriving a plurality of intra-prediction modes based at least on a plurality of reconstructed samples of the template; Obtaining a plurality of prediction blocks for the block based on the derived plurality of intra-prediction modes; The video decoding device of claim 1 , further configured to blend the plurality of prediction blocks based on the plurality of weights.
9. The processor: deriving a plurality of intra-prediction modes based at least on the template associated with the block; obtaining a plurality of weighted and blended predictions of the template based on the derived plurality of intra-prediction modes; 2. The video decoding device of claim 1, further configured to: optimize the weights corresponding to the derived intra-prediction modes based on minimizing a difference between the template and the weighted blended predictions of the template.
10. 1. A decoding method comprising: Retrieving the template associated with the block; obtaining predicted samples and corresponding reconstructed samples of the template; determining a plurality of weights based on minimizing a difference between the predicted samples and the corresponding reconstructed samples of the template; determining an intra prediction model, including applying the weights to a respective one of a plurality of parameters; determining intra-predicted samples for the block based on the intra-prediction model, where the predicted samples, the corresponding reconstructed samples, and the intra-predicted samples share a component type; decoding the block based on the intra-predicted samples; A decoding method comprising:
11. obtaining a plurality of reference samples of the template; optimizing the weights corresponding to the reference samples of the template based on minimizing a difference between a sum of weighted reference samples and the reconstructed sample of the template; The decoding method of claim 10 further comprising:
12. The intra-prediction sample is associated with a sample position within the block, and the plurality of parameters includes a plurality of predicted samples neighboring the sample position, and determining the intra-prediction sample for the block based on the intra-prediction model comprises:
12. The method of decoding of claim 11, comprising applying a plurality of the determined weights to the plurality of predicted samples adjacent to the sample position.
13. the intra-predicted sample is a first intra-predicted sample associated with a first sample position within the block, and the plurality of parameters includes a plurality of reconstructed samples adjacent to the first sample position; for a second sample position within the block, identifying a plurality of predicted samples adjacent to the second sample position; applying the determined weights to the predicted samples to determine a second intra-predicted sample for the block; and The decoding method of claim 11 further comprising:
14. The decoding method of claim 13 , wherein the plurality of predicted samples adjacent to the second sample position includes the first intra-predicted sample.
15. and determining the intra prediction samples for the block based on the intra prediction model, the intra prediction samples being derived from the plurality of parameters corresponding to prediction samples obtained based on the plurality of derived intra prediction modes. The decoding method of claim 10 , comprising applying the weights to the prediction samples obtained based on the derived intra-prediction modes.
16. 11. The decoding method of claim 10, further comprising: deriving a plurality of intra-prediction modes based on a plurality of reconstructed samples of the template, the plurality of prediction samples of the template being obtained based on the respective derived intra-prediction modes, and the plurality of parameters corresponding to the plurality of prediction samples obtained based on the plurality of derived intra-prediction modes.
17. deriving a plurality of intra-prediction modes based at least on a plurality of reconstructed samples of the template; obtaining a plurality of prediction blocks for the block based on the derived plurality of intra-prediction modes; The decoding method of claim 10 , further comprising blending the plurality of prediction blocks based on the plurality of weights.
18. deriving a plurality of intra-prediction modes based at least on the template associated with the block; and obtaining a plurality of weighted and blended predictions of the template based on the derived plurality of intra-prediction modes; 11. The decoding method of claim 10, further comprising: optimizing the weights corresponding to the derived intra-prediction modes based on minimizing differences between the weighted blended predictions of the template and the template.
19. 1. A video encoding device, comprising: a processor, the processor comprising: Get the template associated with the block, obtaining a plurality of predicted samples and a plurality of corresponding reconstructed samples of the template; determining a plurality of weights based on minimizing a difference between the predicted samples and the corresponding reconstructed samples of the template; determining an intra prediction model, the intra prediction model including applying the plurality of weights to a respective plurality of parameters; determining intra-predicted samples for the block based on the intra-prediction model, where the predicted samples, the corresponding reconstructed samples, and the intra-predicted samples share a component type; and encoding the block based on the intra-predicted samples.
20. The processor: obtaining a plurality of reference samples of the template; 20. The video encoding device of claim 19, further configured to optimize the weights corresponding to the reference samples of the template based on minimizing a difference between a sum of weighted reference samples and the reconstructed samples of the template.
21. The intra-prediction sample is associated with a sample position within the block, and the plurality of parameters includes a plurality of predicted samples neighboring the sample position, and determining the intra-prediction sample for the block based on the intra-prediction model comprises: The video encoding device of claim 20 , further comprising applying the determined weights to the predicted samples adjacent to the sample position.
22. the intra-predicted sample is a first intra-predicted sample associated with a first sample position within the block, the plurality of parameters includes a plurality of reconstructed samples adjacent to the first sample position, and the processor: for a second sample position within the block, identifying a plurality of predicted samples adjacent to the second sample position; 21. The video encoding device of claim 20, further configured to: apply the determined weights to the predicted samples to determine a second intra-predicted sample for the block, wherein the block is decoded further based on the second intra-predicted sample of the block.
23. 23. The video encoding device of claim 22, wherein the plurality of predicted samples adjacent to the second sample position includes the first intra-predicted sample.
24. The processor: and deriving a plurality of intra prediction modes based on a plurality of reconstructed samples of the template and a plurality of reference samples of the template, wherein the plurality of parameters correspond to a plurality of prediction samples obtained based on the derived intra prediction modes, and determining the intra prediction samples for the block based on the intra prediction model comprises: The video encoding device of claim 19 , further comprising applying the weights to the prediction samples obtained based on the derived intra-prediction modes.
25. The processor: obtaining a histogram of gradients associated with a plurality of samples of the template; 20. The video encoding device of claim 19, further configured to: derive a plurality of intra-prediction modes based on a histogram of the gradients associated with the template, wherein the plurality of prediction samples of the template are obtained based on the respective derived intra-prediction modes, and the plurality of parameters correspond to a plurality of prediction samples obtained based on the derived intra-prediction modes.
26. The processor: deriving a plurality of intra-prediction modes based at least on a plurality of reconstructed samples of the template; Obtaining a plurality of prediction blocks for the block based on the derived plurality of intra-prediction modes; The video encoding device of claim 19 , further configured to blend the plurality of prediction blocks based on the plurality of weights.
27. The processor: deriving a plurality of intra-prediction modes based at least on the template associated with the block; obtaining a plurality of weighted and blended predictions of the template based on the derived plurality of intra-prediction modes; 20. The video encoding device of claim 19, further configured to optimize the weights corresponding to the derived intra-prediction modes based on minimizing a difference between the template and the weighted blended predictions of the template.
28. 1. An encoding method comprising: Retrieving the template associated with the block; obtaining predicted samples and corresponding reconstructed samples of the template; determining a plurality of weights based on minimizing a difference between the predicted samples and the corresponding reconstructed samples of the template; determining an intra prediction model, including applying the weights to a respective one of a plurality of parameters; determining intra-predicted samples for the block based on the intra-prediction model, where the predicted samples, the corresponding reconstructed samples, and the intra-predicted samples share a component type; encoding the block based on the intra-predicted samples; 10. An encoding method comprising:
29. obtaining a plurality of reference samples of the template; optimizing the weights corresponding to the reference samples of the template based on minimizing a difference between a sum of weighted reference samples and the reconstructed sample of the template; 29. The encoding method of claim 28, further comprising:
30. The intra-prediction sample is associated with a sample position within the block, and the plurality of parameters includes a plurality of predicted samples neighboring the sample position, and determining the intra-prediction sample for the block based on the intra-prediction model comprises:
30. The method of claim 29, comprising applying a plurality of the determined weights to the plurality of predicted samples adjacent to the sample position.
31. the intra-predicted sample is a first intra-predicted sample associated with a first sample position within the block, and the plurality of parameters includes a plurality of reconstructed samples adjacent to the first sample position; for a second sample position within the block, identifying a plurality of predicted samples adjacent to the second sample position; applying the determined weights to the predicted samples to determine a second intra-predicted sample for the block; and 30. The encoding method of claim 29, further comprising:
32. 32. The encoding method of claim 31 , wherein the plurality of predicted samples adjacent to the second sample position includes the first intra-predicted sample.
33. and determining the intra prediction samples for the block based on the intra prediction model, the intra prediction samples being derived from the plurality of parameters corresponding to prediction samples obtained based on the plurality of derived intra prediction modes.
29. The encoding method of claim 28, comprising applying the weights to the prediction samples obtained based on the derived intra-prediction modes.
34. 29. The encoding method of claim 28, further comprising: deriving a plurality of intra-prediction modes based on a plurality of reconstructed samples of the template, the plurality of prediction samples of the template being obtained based on the respective derived intra-prediction modes, and the plurality of parameters corresponding to the plurality of prediction samples obtained based on the plurality of derived intra-prediction modes.
35. deriving a plurality of intra-prediction modes based at least on a plurality of reconstructed samples of the template; obtaining a plurality of prediction blocks for the block based on the derived plurality of intra-prediction modes; 29. The encoding method of claim 28, further comprising blending the plurality of prediction blocks based on the plurality of weights.
36. deriving a plurality of intra-prediction modes based at least on the template associated with the block; and obtaining a plurality of weighted and blended predictions of the template based on the derived plurality of intra-prediction modes; 29. The encoding method of claim 28, further comprising: optimizing the weights corresponding to the derived intra-prediction modes based on minimizing differences between the weighted blended predictions of the template.
37. A computer program product stored on a non-transitory computer readable medium and comprising program code instructions for performing the steps of the method of any one of claims 10 to 18 or claims 28 to 36 when executed by a processor.
38. A computer program comprising program code instructions for carrying out the steps of the method according to any one of claims 10 to 18 or claims 28 to 36 when the computer program is executed by a processor.
39. Video data comprising information representing video blocks encoded according to one of the methods of any one of claims 28 to 36.