Internal chroma format increase

The internal chroma format increase (ICFI) in video encoding upsamples chroma signals to match luma signals, enhancing chroma prediction accuracy and reducing processing demands without increasing data volume.

WO2025163017A1PCT designated stage Publication Date: 2025-08-07INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/052296
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-02
Filing Date
2025-01-29
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Existing video encoding methods impair the efficiency and accuracy of intra-frame and inter-frame chroma prediction by compressing chroma components relative to luma components, leading to loss of high-frequency information and increased processing requirements.

Method used

Implementing an internal chroma format increase (ICFI) by upsampling the chroma format of reconstructed frames to match the luma format, allowing residuals to be coded in the original format or with adjusted quantization parameters to maintain data rates, and using rate-distortion control to optimize encoding.

Benefits of technology

Improves the representation and quality of chroma signals by maintaining high-frequency information and reducing processing resources, while avoiding increased data volume.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025052296_07082025_PF_FP_ABST
    Figure EP2025052296_07082025_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods for upscaling a chroma format of chroma signals internally (i.e. internal chroma format increase (ICFI)), to improve an inner representation of a reconstructed chroma signal are provided. The chroma format of one or more reconstructed frames may be increased in comparison to one or more original frames. The one or more original frames may be upscaled before coding. In some implementations, one or more residuals may be coded in an original (compressed and / or lower resolution) chroma format (e.g. not upscaled) to avoid increasing an amount of encoded data. In some implementations, the chroma format of the one or more residuals may be increased, but a quantization parameter (QP) for quantizing one or more chroma residuals may be increased to maintain data rates. In some implementations, rate-distortion may be modified at an encoder to account for upscaling the original chroma format.
Need to check novelty before this filing date? Find Prior Art

Description

INTERNAL CHROMA FORMAT INCREASETECHNICAL FIELD

[0001] The present embodiments generally relate to systems and methods for video encoding and decoding. In particular, the present disclosure is directed to systems and methods for video coding with an internal chroma format increase.BACKGROUND

[0002] Video signals may be encoded with separate chroma and luma information, sometimes referred to as a Y’CbCr encoding scheme, with Y’ representing a luma component, and Cb and Cr representing one or more color difference components, i.e. , chroma components. A human eye is more sensitive to the luma information. So, a video signal may be compressed by reducing a resolution of the chroma components while maintaining a resolution of the luma component. Significant bandwidth savings may be realized, with compression ratios up to 0.5 frequently applied.

[0003] Additionally, the video signal may be compressed by encoding one or more temporal differences with motion compensation relative to a reference frame (e.g., inter-frame coding) and / or relative to other portions of a same frame (e.g., intra-frame coding), rather than providing full details of every frame. Such compression may be applied to both the luma and chroma components.

[0004] However, when the chroma components are compressed relative to the luma components, efficiency and / or accuracy of the intra-frame and / or inter-frame chroma prediction may be impaired. For example, as a part of performing the intra-frame and / or inter-frame motion compensation prediction, the chroma components may need to be upscaled, requiring additional processing resources. In another example, when the chroma components are coded with a cross-component mode, one or more co-located reconstructed luma components may be downscaled for building a chroma prediction. In such instances, high frequency information of the co-located luma components may be removed and lost and may not contribute to chroma prediction. In another example, when the chroma components are predicted with the motioncompensation, a codec interpolates one or more chroma samples of one or more reference pictures with one or more down-sampled chroma samples and consequently reduces an amount of high frequencies.SUMMARY

[0005] In some embodiments, the present disclosure is directed to systems and methods for upscaling a format (e.g. a number of samples) of a chroma signal internally, i.e. internal chroma format increase (ICFI), to improve an inner representation of a reconstructed chroma signal. In some embodiments, a chroma format of one or more reconstructed frames may be increased compared to one or more original frames. In some such embodiments, the one or more original frames may be upscaled before coding. In some embodiments, one or more residuals may still be coded in an original (e.g. compressed and / or lower resolution) chroma format (e.g. not upscaled) to avoid increasing an amount of encoded data. In other embodiments, a chroma format of the residuals may be increased, but a quantization parameter (QP) for quantizing the one or more chroma residuals may be increased to maintain data rates. In some implementations, rate distortion control may be modified at an encoder to account for increase in the chroma format.

[0006] In one or more embodiments, a decoder is provided. The decoder comprises a memory, a transceiver, and a processor. The transceiver and the processor are configured to receive a bitstream comprising a coded representation of one or more input images at an original chroma format. An original chroma component size is lower than an original luma component size in the original chroma format. The transceiver and the processor are configured to generate, based on the received bitstream, one or more prediction blocks and one or more residuals using one or more reference images stored in a buffer in the memory at an upscaled chroma format. An upscaled chroma component size in the upscaled chroma format is larger than the original chroma component size. The transceiver and the processor are configured to generate the one or more reconstructed images based on the one or more prediction blocks and the one ormore residuals. The transceiver and the processor are configured to store the one or more reconstructed images in the buffer in the upscaled chroma format.

[0007] In one or more embodiments, a method performed by a decoder is provided. The method includes receiving a bitstream comprising a coded representation of one or more input images at an original chroma format. In an example, the coded representation (e.g., a number of blocks, etc....) may be at the original chroma format or at the upscaled chroma format. An original chroma component size is lower than a luma component size in the original chroma format. The method includes generating, based on the received bitstream, one or more prediction blocks and one or more residuals using one or more reference images stored in a buffer at an upscaled chroma format. An upscaled chroma component size in the upscaled chroma format is larger than the original chroma component size. The method includes generating the one or more reconstructed images based on the one or more prediction blocks and the one or more residuals. The method includes storing the one or more reconstructed images in the buffer in the upscaled chroma format.

[0008] In an embodiment, the decoder generate one or more output images in the upscaled chroma format.

[0009] In an embodiment, decoder downscales the one or more reference images to the original chroma format.

[0010] In an embodiment, the decoder generates one or more output images in the original chroma format.

[0011] In an embodiment, the upscaled chroma format is 4:4:4 and the original chroma format is at least one of: 4:2:0 or 4:2:2.

[0012] In an embodiment, the original chroma format includes upscaling the original chroma component size to match the original luma component size.

[0013] In an embodiment, the one or more prediction blocks are generated at the original chroma format, and wherein the one or more residuals are generated at the original chroma format. In an embodiment, the decoder upscales the one or more reconstructed images to the upscaled chroma format.

[0014] In an embodiment, the one or more prediction blocks are generated at the upscaled chroma format, and wherein the one or more residuals are generated at the upscaled chroma format.

[0015] In an embodiment, the one or more prediction blocks are generated at the upscaled chroma format, and wherein the one or more residuals are generated at the original chroma format.

[0016] In an embodiment, the decoder upscales the one or more residuals to the upscaled chroma format. The decoder reconstructs the images at the upscaled chroma format.

[0017] In one or more embodiments, a method performed by an encoder is provided. The method includes receiving one or more input images in an original chroma format. In an example, the encoder may receive a coded representation (e.g. , a number of blocks, etc.... ) of the one or more images at the original chroma format or at the upscaled chroma format. An original chroma component size is smaller than an original luma component size in the original chroma format. The method includes generating, based on the one or more input images, one or more prediction blocks and one or more residuals. The method includes generating one or more reconstructed images based on the one or more prediction blocks and the one or more residuals. The method includes storing the one or more reconstructed images in a buffer as one or more reference images in an upscaled chroma format. An upscaled chroma component size in the upscaled chroma format is larger than the original chroma component size. The method includes encoding the one or more residuals.

[0018] In an embodiment, the one or more residuals are in the original chroma format.

[0019] In an embodiment, the one or more residuals are in the upscaled chroma format.

[0020] In an embodiment, encoding the one or more residuals comprises one of: downscaling the one or more residuals in the original chroma format, or transforming the one or more residuals at the original chroma format.

[0021] In an embodiment, the encoding the one or more residuals comprises: transforming the one or more residuals for generating one or moretransformation coefficients, and downscaling the one or more transformation coefficients at the original chroma format.

[0022] In an embodiment, encoding the one or more residuals further comprises: dynamically determining a modified quantization parameter based on the upscaled chroma format, and quantizing the one or transformation coefficients based on the modified quantization parameter.

[0023] In an embodiment, the encoder determines a distortion value based on a reconstructed block in the upscaled chroma format and an input block in the upscaled chroma format. The encoder derives, based on the distortion value, a cost and a number of bits to encode. The encoder selecting a coding mode for encoding the input block based on the cost.

[0024] In an embodiment, the encoder determines a distortion value based on a down-scaled reconstructed block in the original chroma format and an input block in the original chroma format. The encoder derives a cost based on distortion and number of bits to encode. The encoder selects a coding mode for encoding the input block based on the cost.BRIEF DESCRIPTION OF THE DRAWINGS

[0025] A more detailed understanding may be had from the following description, given by way of example in conjunction with the accompanying drawings, wherein like reference numerals in the figures indicate like elements, and wherein:

[0026] FIG. 1 illustrates a block diagram of an embodiment of a video encoder in which various aspects of the embodiments may be implemented;

[0027] FIG. 2 illustrates a block diagram of an embodiment of a video decoder in which various aspects of the embodiments may be implemented;

[0028] FIG. 3 illustrates a block diagram of a system within which aspects of the present embodiments may be implemented;

[0029] FIG. 4 illustrates a block diagram of an example encoder reference picture rescaler, according to some implementations;

[0030] FIG. 5 illustrates a block diagram of an example decoder reference picture rescaler, according to some implementations;

[0031] FIG. 6 illustrates examples of chroma formats used in various implementations of video compression;

[0032] FIG. 7 illustrates examples of chroma sample phase vs. luma samples for some implementations of chroma formats used in video compression;

[0033] FIG. 8 illustrates an example of chroma component upscaling for an internal chroma format increase, according to one or more embodiments;

[0034] FIG. 9 illustrates an example of chroma residuals coefficients, and reducing chroma residuals coefficients according to one or more embodiments;

[0035] FIG. 10 is a flowchart illustrating an example video decoding process according to one or more embodiments;

[0036] FIG. 11 is a flowchart illustrating an example video encoding process according to one or more embodiments;

[0037] FIG. 12 is a flowchart illustrating an example video decoding process including upscaling one or more reconstructed images according to one or more embodiments;

[0038] FIG. 13 is a flowchart illustrating an example video encoding process including upscaling one or more reconstructed images according to one or more embodiments;

[0039] FIG. 14 is a flowchart illustrating an example video decoding process including decoding one or more residuals at an original chroma format and reconstructing one or more images at an upscaled chroma format according to one or more embodiments; and

[0040] FIG. 15 is a flowchart illustrating an example encoding process including upscaling one or more images before encoding and coding one or more residuals at an original chroma format according to one or more embodiments.DETAILED DESCRIPTION

[0041] Before discussing specifics of the implementations of systems and methods discussed herein, it may be helpful to briefly describe video encodersand decoders and computing devices suitable for practicing these implementations. FIG. 1 illustrates a block diagram of an example video encoder 100 in which various aspects of the embodiments may be implemented. Variations of this encoder 100 are contemplated, but the encoder 100 is described below for purposes of clarity without describing all expected variations.

[0042] Before being encoded, a video signal may go through a preencoding processing (101 ), for example, applying a color transform to an input color picture (e.g., conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing a remapping of one or more input picture components in order to get a signal distribution more resilient to compression (for instance using a histogram equalization of one color components). Metadata can be associated with the preprocessing, and attached to a video bitstream.

[0043] In the encoder 100, the picture is encoded by the encoder elements as described below. The picture to be encoded is partitioned (102) and processed in one or more units, for example, coding units (CUs). Each unit is encoded using, for example, either an intra mode and / or inter mode. When the unit is encoded in the intra mode, it performs intra prediction (160). In the inter mode, motion estimation (175) and / or motion compensation (170) are performed. The encoder decides (105) which one of the intra mode or inter mode to use for encoding the unit, and indicates the intra and / or inter decision by, for example, a prediction mode flag. One or more prediction residuals are calculated, for example, by subtracting (110) a predicted block from an original image block.

[0044] The one or more prediction residuals are then transformed (125) and / or quantized (130). The one or more quantized transform coefficients, as well as on or more motion vectors and other syntax elements, are entropy coded (145) to output a bitstream. In an example, the encoder can skip the transform and apply quantization directly to a non-transform ed residual signal. In an example, the encoder can bypass both transform and quantization, i.e., the residual may be coded directly without the application of the transform and / or quantization processes.

[0045] The encoder decodes an encoded block to provide a reference for one or more further predictions. The quantized transform coefficients are dequantized (140) and inverse transformed (150) to decode prediction residuals. Combining (155) the decoded prediction residuals and the predicted block, an image block is reconstructed. In-loop filters (165) are applied to the reconstructed picture to perform, for example, deblocking and / or Sample Adaptive Offset (SAO) filtering to reduce encoding artifacts. The filtered image is stored at a reference picture buffer (180).

[0046] FIG. 2 illustrates a block diagram of a video decoder 200. In the decoder 200, the bitstream is decoded by one or more decoder elements as described below. The video decoder 200 generally performs a decoding pass reciprocal to an encoding pass as described in FIG. 1 . The encoder 100 also generally performs video decoding as part of encoding video data.

[0047] In particular, an input of the decoder includes the video bitstream, which can be generated by the video encoder 100. The bitstream is first entropy decoded (230) to obtain transform one or more coefficients, motion vectors, and other coded information. The picture partition information indicates how the picture is partitioned. The decoder may therefore divide (235) the picture according to a decoded picture partitioning information. The transform coefficients are de-quantized (240) and inverse transformed (250) to decode the one or more prediction residuals. Combining (255) the one or more decoded prediction residuals and the predicted block, an image block is reconstructed. The predicted block can be obtained (270) from an intra prediction (260) and / or a motion-compensated prediction (i.e., inter prediction) (275). One or more in-loop filters (265) are applied to the reconstructed image. The filtered image is stored at a reference picture buffer (280).

[0048] The decoded picture can further go through post-decoding processing (285), for example, an inverse color transform (e.g. conversion from YCbCr 4:2:0 to RGB 4:4:4) or an inverse remapping performing the inverse of the remapping process performed in the pre-encoding processing (101). The postdecoding processing can use metadata derived in the pre-encoding processing and signaled in the bitstream.

[0049] FIG. 3 illustrates a block diagram of an example system in which various aspects and embodiments can be implemented. System 300 may be embodied as a device including the various components described below and is configured to perform one or more of the aspects described in this application. Examples of such devices, include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 300, singly or in combination, may be embodied in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 300 are distributed across multiple ICs and / or discrete components. In various embodiments, the system 300 is communicatively coupled to other systems, or to other electronic devices, via, for example, a communications bus or through dedicated input and / or output ports. In various embodiments, the system 300 is configured to implement one or more of the aspects described in this application.

[0050] The system 300 includes at least one processor 310 configured to execute instructions loaded therein for implementing, for example, the various aspects described in this application. The processor 310 may include embedded memory, input output interface, and various other circuitries as known in the art. The system 300 includes at least one memory 320 (e.g., a volatile memory device, and / or a non-volatile memory device). The system 300 includes a storage device 340, which may include non-volatile memory and / or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drive, and / or optical disk drive. The storage device 340 may include an internal storage device, an attached storage device, and / or a network accessible storage device, as non-limiting examples.

[0051] The system 300 includes an encoder / decoder module 330 configured, for example, to process data to provide an encoded video / 3D object or decoded video / 3D object, and the encoder / decoder module 330 may include its own processor and memory. The encoder / decoder module 330 represents one or more modules that may be included in a device to perform the encoding and / ordecoding functions. As is known, a device may include one or both of the encoding and decoding modules. Additionally, encoder / decoder module 330 may be implemented as a separate element of system 300 or may be incorporated within the processor 310 as a combination of hardware and software as known to those skilled in the art.

[0052] A program code to be loaded onto processor 310 or encoder / decoder 330 to perform the various aspects described in this application may be stored in storage device 340 and subsequently loaded onto memory 320 for execution by processor 310. In accordance with various embodiments, one or more of processor 310, memory 320, storage device 340, and encoder / decoder module 330 may store one or more of various items during the performance of the processes described in this application. Such stored items may include, but are not limited to, the input video / 3D object, the decoded video / 3D object or portions of the decoded video / 3D object, the bitstream, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.

[0053] In several embodiments, a memory inside of the processor 310 and / or the encoder / decoder module 330 is used to store instructions and to provide working memory for processing that is needed during encoding and / or decoding. In other embodiments, however, a memory external to the processing device (for example, the processing device may be either the processor 310 or the encoder / decoder module 330) is used for one or more of these functions. The external memory may be the memory 320 and / or the storage device 340, for example, a dynamic volatile memory and / or a non-volatile flash memory. In several embodiments, an external non-volatile flash memory is used to store the operating system of a television. In at least one embodiment, a fast external dynamic volatile memory such as a RAM is used as working memory for coding and decoding operations, such as for instance MPEG-2, HEVC, or 1C.

[0054] The input to the elements of system 300 may be provided through various input devices as indicated in block 305. Such input devices include, but are not limited to, (i) an RF portion that receives an RF signal transmitted, forexample, over the air by a broadcaster, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.

[0055] In various embodiments, the input devices of block 305 have associated respective input processing elements as known in the art. For example, the RF portion may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band of frequencies), (ii) down converting the selected signal, (iii) band-limiting again to a narrower band of frequencies to select (for example) a signal frequency band which may be referred to as a channel in certain embodiments, (iv) demodulating the down converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired stream of data packets. The RF portion of various embodiments includes one or more elements to perform these functions, for example, frequency selectors, signal selectors, band-limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion may include a tuner that performs various of these functions, including, for example, down converting the received signal to a lower frequency (for example, an intermediate frequency or a near-baseband frequency) or to baseband. In one set-top box embodiment, the RF portion and its associated input processing element receives an RF signal transmitted over a wired (for example, cable) medium, and performs frequency selection by filtering, down converting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the abovedescribed (and other) elements, remove some of these elements, and / or add other elements performing similar or different functions. Adding elements may include inserting elements in between existing elements, for example, inserting amplifiers and an analog-to-digital converter. In various embodiments, the RF portion includes an antenna.

[0056] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 300 to other electronic devices across USB and / or HDMI connections. It is to be understood that various aspects of input processing, for example, Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or withinprocessor 310 as necessary. Similarly, aspects of USB or HDMI interface processing may be implemented within separate interface ICs or within processor 310 as necessary. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 310, and encoder / decoder 330 operating in combination with the memory and storage elements to process the data stream as necessary for presentation on an output device.

[0057] Various elements of system 300 may be provided within an integrated housing, Within the integrated housing, the various elements may be interconnected and transmit data therebetween using suitable connection arrangement 315, for example, an internal bus as known in the art, including the I2C bus, wiring, and printed circuit boards.

[0058] The system 300 includes communication interface 350 that enables communication with other devices via communication channel 390. The communication interface 350 may include, but is not limited to, a transceiver configured to transmit and to receive data over communication channel 390. The communication interface 350 may include, but is not limited to, a modem or network card and the communication channel 390 may be implemented, for example, within a wired and / or a wireless medium.

[0059] Data is streamed to the system 300, in various embodiments, using a Wi-Fi network such as IEEE 802.11 . A Wi-Fi signal of these embodiments is received over the communications channel 390 and the communications interface 350 which are adapted for Wi-Fi communications. The communications channel 390 of these embodiments is typically connected to an access point or router that provides access to outside networks including the Internet for allowing streaming applications and other over-the-top communications. Other embodiments provide streamed data to the system 300 using a set-top box that delivers the data over the HDMI connection of the input block 305. Still other embodiments provide streamed data to the system 300 using the RF connection of the input block 305.

[0060] The system 300 may provide an output signal to various output devices, including a display 365, speakers 375, and other peripheral devices 385.The other peripheral devices 385 include, in various examples of embodiments, one or more of a stand-alone DVR, a disk player, a stereo system, a lighting system, and other devices that provide a function based on the output of the system 300. In various embodiments, one or more control signals are communicated between the system 300 and the display 365, speakers 375, or other peripheral devices 385 using signaling such as AV.Link, CEC, or other communications protocols that enable device- to-device control with or without user intervention. The output devices may be communicatively coupled to system 300 via dedicated connections through respective interfaces 360, 370, and 380. Alternatively, the output devices may be connected to system 300 using the communications channel 390 via the communications interface 350. The display 365 and speakers 375 may be integrated in a single unit with the other components of system 300 in an electronic device, for example, a television. In various embodiments, the display interface 360 includes a display driver, for example, a timing controller (T Con) chip.

[0061] The display 365 and speaker 375 may alternatively be separate from one or more of the other components, for example, if the RF portion of input 305 is part of a separate set-top box. In various embodiments in which the display 365 and speakers 375 are external components, the output signal may be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.

[0062] In some implementations, for example as part of pre-encoding processing 101 in an encoder and as part of post-processing decoding 285 in a decoder, picture-based re-scaling of original pictures to encode or of reconstructed pictures, may be utilized to improve compression trade-off. Then the motion compensation (encoder 170 or decoder 275) may include implicit rescaling to improve motion compensation accuracy when the reference picture and the current picture have different size. For example, in the Versatile Video Coding (WC) standard promulgated as ISO / IEC 23090-3 and ITU-T H.266, each of which are incorporated by reference herein, the re-scaling feature implicitly done at the motion compensation stage is named Reference Picture Resampling (RPR). Given an original video sequence composed of pictures of size {width xheight}, the encoder may choose for each frame which resolution (picture size) to use for coding the frame. Different picture parameter sets (PPS) are coded in the bit-stream with the possible sizes of the pictures and the slice / picture header indicates which PPS to use to decode the current video coding layer (VCL) encapsulated into network abstraction layer (NAL) unit. The SPS contains the maximum picture size.

[0063] FIGs. 4 and 5 are illustrations of a block diagram of an example encoder 400 and an example decoder 500 with reference picture resampler (430 and 530), respectively, according to some implementations. As discussed above, a down-sampler 440 and an up-sampler 540 may be provided by pre-encoding processing 101 before a core encoder 410 (which may comprise the encoder 100 and / or portions of the encoder 101 subsequent to the pre-encoding processing 101 , such as the image partitioning 102 through the entropy coding 145) and the post-decoding processing 285 after the core decoder 510 (which may comprise the decoder 200 and / or portions of the decoder 200 prior to post-decoding processing such as the entropy decoding 230 through the in-loop filters 265 (including the buffer 280 and the motion compensation 275).

[0064] For each frame, the encoder chooses whether to encode at an original resolution or at a down-sized resolution (ex: picture width and / or height divided by 2). The choice can be made with two passes encoding and / or considering spatial and / or temporal activity in one or more original pictures. Consequently, the buffer (e.g., a decoded picture buffer (DPB)) 420, 520, which may be provided by one or more reference picture buffers 180, 280, can include pictures with a different size from a current picture size.

[0065] In an example, the terms pictures, frames, or images may be used interchangeably.

[0066] In case a reference picture in the buffer 420, 520 has a size different from the current picture size, in reference picture rescaling (RPR), the resamplers 430, 530 (sometimes referred to as re-samplers, scalers, or by similar terms) may upscale or downscale the reference block as needed to build a prediction block during the motion compensation process (170, 275) using an appropriate phase of the sample to interpolate. For example, if the ratio betweenthe reference picture size and the current picture size is 2 and the motion vector is full pel, then the MC phase may be set to 0 or 0.5 alternatively for consecutive prediction samples. By contrast, if the reference and the current picture sizes are identical and the motion vector is full pel, then the MC phase may be set to 0 for all prediction samples.

[0067] As discussed above, in many implementations, the video signal may be encoded with separate chroma and luma information, sometimes referred to as a Y’CbCr encoding scheme, with Y’ representing the luma component, and Cb and Cr representing color difference components. Because the human eye is more sensitive to the luma information, the video signal may be compressed by reducing the resolution of the chroma components while maintaining the resolution of the luma information. FIG. 6 is an illustration 600 of some examples of chroma formats used in various implementations of video compression. The chroma formats are typically described as a three part ratio J:a:b (or four parts, if an alpha channel is present), that describe the ratios of luminance and chrominance samples in a region that is J pixels wide by 2 pixels high. For example, at left in FIG. 6 is shown an uncompressed chroma format, referred to as 4:4:4, in which, for every 4 luminance samples, there are 4 chrominance samples in each of the two rows of pixels. As a result, the Y, Cr, and Cb effective frames are all the same size.

[0068] By contrast, at center is an example of 4:2:2 subsampling, in which for every 4 luma samples, there are 2 chroma samples in the first row and two chroma samples in the second row (providing 1 / 2 horizontal chroma resolution and full vertical chroma resolution). At right is an example of 4:2:0 subsampling, with 4 luma samples, 2 chroma samples in the first row, and 0 chroma samples in the second row (the Cr and Cb channels are typically sampled on alternate lines), providing 1 / 2 horizontal and vertical chroma resolution. As shown, 4:2:2 subsampling results in a 2:3 compression ratio compared to the uncompressed video, and 4:2:0 subsampling results in a 1 :2 compression ratio.

[0069] Four variants of 4:2:0 chroma sub-sampling schemes 700 are depicted in FIG. 7, with different horizontal and vertical sampling siting relative to the 2x2 “square” of the original (luminance) input size (shown as open squaresin pixel positions 710). In JPEG / JFIF, H.261 , and MPEG-1 , illustrated in FIG. 7(a), Cb and Cr chrominance samples (shown as round dots 720) are taken at the center of 2x2 the square. In other words, they are offset one-half pixel to the right and one-half pixel down compared to the top-left pixel. This is sometimes referred to as “center” sub-sampling.

[0070] In MPEG-2, MPEG-4, and AVC, illustrated in FIG. 7(b), Cb and Cr samples are taken on midpoint of the left-edge of the 2x2 square. In other words, they have the same horizontal location as the top-left pixel, but is shifted one-half pixel down vertically. This may be referred to as “left” sub-sampling.

[0071] In HEVC for BT.2020 and BT.2100 content (in particular on Blu-Ray discs), illustrated in FIG. 7(c), Cb and Cr are sampled at the same location as the group’s top-left Y pixel (referred to as “co-sited”, “co-located”, or similar terms, and sometimes referred to as “top-left” sub-sampling). An analogous co-sited sampling is used in MPEG-2 4:2:2.

[0072] Similarly, in 4:2:0 PAL-DV (IEC 61834-2), Cb is sampled at the same location as the group's top-left Y pixel (as in FIG. 7(c)), but Cr is sampled one pixel down (not illustrated in FIG. 7). It may also be referred to as “top-left” sub-sampling in mpeg.

[0073] In some other implementations, Cb and Cr samples may be made at the same vertical location as the top-left pixel, but shifted one-half pixel right horizontally, as shown in FIG. 7(d). This may be referred to as “top-center” subsampling.

[0074] In some implementations, RPR may be used to support different scaling ratios for both the luma and chroma components relatively to an original picture size (ex: signaled in the SPS). Some coded pictures may have different size (ex: signaled in the PPS) than the original frame size and may build inter prediction with reference pictures of different size. However, in some implementations, such as when video is sub-sampled to a 4:2:0 format, RPR may result in a higher peak signal to noise ratio (PSNR) drop for the chroma components of a frame than for the luma component, at least in part because RPR performs an additional downscaling filtering on samples that were already filtered from the original canonical 4:4:4 content to create the 4:2:0 format. Inparticular, whenever the chroma format is different from 4:4:4, the chroma components may be reconstructed at lower precision than the luma component. This may jeopardize the efficiency of the subsequent intra- or inter-frame chroma predictions.

[0075] Internal bit-depth increase (I BDI) is a technique that increases a bit depth of the reconstructed pictures stored in the buffer (e.g., the DPB). By increasing the internal bit depth, an accuracy of various internal processes, including motion compensation, interpolation filtering, and / or deblocking filtering etc. for example, is increased. The bit depth may generally refer to a number of bits of information for a given sample (e.g., luma and / or chroma sample values) of video data. For example, in some implementations, the encoder may expand the bit depth of a sample being coded from a first number of bits (e.g., “M” bits) to a second, increased number of bits (e.g., “N” bits, with N > M). The greater bit depth may reduce rounding errors in internal calculations by increasing arithmetic precision.

[0076] For example, in some implementations, the encoder may increase the bit depth of the received picture and store decoded video data at the increased bit depth into the buffer during coding (e.g., for use as reference data for predictive coding), and may perform one or more rounding operations prior to decreasing the bit depth and / or outputting the picture at the original bit depth.

[0077] To provide a similar increased accuracy and efficiency in coding chroma samples, in one or more embodiments, the present disclosure provides implementations of systems and methods for an internal chroma format increase (ICFI), which may be used in addition to or separately from I BDI . In one or more embodiments, the chroma format of reconstructed frames may be increased (e.g., upscaled) compared to one or more original frames. In one or more embodiments, the original pictures may be upscaled before coding. In some implementations, the residuals may still be coded in an original (e.g., compressed and / or lower resolution etc.) chroma format (e.g. not upscaled) to avoid increasing an amount of encoded data. In other implementations, the chroma format of the residuals may be increased, but a quantization parameter (QP) for quantizing chroma residuals may be increased to maintain data rates trade-off.In some implementations, rate-distortion control may be modified at the encoder to account for the chroma format increase.

[0078] As discussed above, in some implementations of video coding standards, a coding unit CU (or prediction unit (PU)) may be composed of 3 coding blocks corresponding to the luma and chroma components (e.g., Y’,Cb,Cr). The following implementations are primarily discussed in terms of an original input sequence with 4:2:0 sub-sampling and an internal chroma format of 4:4:4. However, the systems and methods discussed herein may be utilized with any other input or output chroma formats (e.g., 4:2:2, 4:1 :1 , etc.).

[0079] In some implementations of ICFI, the reconstructed chroma format of the block or image may be increased, such that the width and / or height of the reconstructed chroma block and / or image is greater than the original chroma size. For example, in some implementations, the reconstructed chroma format may be 4:4:4, regardless of whatever the original chroma format was (e.g., input chroma format may be 4:2:0, 4:1 :1 , etc.). The reference pictures in the buffer (e.g., the DPB) may be stored at the increased chroma format. For example, the chroma component arrays may have the same size as the luma component arrays, and equal to the original input luma component size. Just as the chroma format is increased in the buffer, the one or more chroma components can be down sampled to match the original picture format when the picture is output from the buffer. However, in some implementations, the output picture format does not need to match the input picture format (e.g. the input may be 4:2:0 subsampled, and upscaled to 4:4:4 internally, but the output may remain at 4:4:4). Accordingly, in such implementations, the upscaled resolution in the chroma components is maintained in the output pictures, thereby increasing quality and reducing processing steps (and potential artifacts or noise).

[0080] In another implementation of the ICFI, the chroma format may be increased to 4:4:4 subsampling before coding, and chroma residuals may similarly be processed at an uncompressed 4:4:4 subsampling level. For example, FIG. 8 illustrates an example of chroma component upscaling 800 for the ICFI, according to some implementations. The original chroma samples 810 are not changed (in grey in the figure) but the missing chroma samples 820 (inwhite in the figure) are generated via interpolation. This is similar to 2:1 chroma up-sampling for collocated chroma samples, as in processing the co-located subsampling implementation of FIG. 7(c). In case the chroma is not co-located (e.g. for AVC or MPEG2 4:2:0 chroma sampling, as shown in FIG. 7(b)), using interpolation as shown in FIG. 8 can result in distortions in the chroma components. It can also result in misalignment between the luma and chroma components, and reduce the performance of cross-component prediction tools. In such implementations, such as where the chroma sample is not collocated, one or more upscaling filters that are based on a chroma location and / or one or more original chroma samples may be used and all the chroma samples may be generated (or interpolated). In such cases, the reconstructed chroma samples may be re-phased before display, as shown in block 540 of FIG. 5.

[0081] In another implementation of ICFI, one or more chroma residual coefficients may be scaled relative to the upscaled chroma format. For example, in some such implementations, the sub-sampling format used for the prediction may be 4:4:4, with a chroma block prediction using the same size as a luma block prediction. The reconstructed chroma blocks are accordingly 4:4:4.

[0082] For inter-frame and / or inter-block prediction, since the reference pictures in the buffer (e.g., the DPB) are 4:4:4, the motion compensation of the chroma block is also 4:4:4. In case of cross-component prediction, the crosscomponent model is applied on the reconstructed luma block (without downscaling the reconstructed luma block) so that the chroma prediction is 4:4:4. In a variant, the reconstructed luma block may be filtered to re-phase with the chroma phase if necessary. In case of intra-frame and / or intra-block prediction, since the current picture is 4:4:4, the neighboring reconstructed chroma samples in the same frame are also at 4:4:4. However, the transformed chroma residuals size may be the same as the original input chroma format (for example, the residuals may still be subsampled at 4:2:0 if the input pictures are similarly subsampled at 4:2:0). At the encoder, this may be obtained in different ways as follows.

[0083] The chroma residuals are first computed at 4:4:4, and then are downscaled (e.g. by using a linear filter or sub-sampling) to the original inputchroma format (4:2:0) and then transformed (4:2:0). The downscaling is performed via at least one downscaling filter. In some implementations, padding and / or mirroring may be used for applying downscaling filters at the block border.

[0084] In another implementation, the chroma residuals are first computed at 4:4:4, then are transformed (4:4:4) and downscaled (e.g., via a linear filter).

[0085] In another implementation, the chroma residuals may first be computed at 4:4:4, then are transformed (4:4:4) and zeroed-out to 4:2:0, for example, using a Hamming window. For example, FIG. 9 illustrates an example (900) of coefficients of transformed 8x8 chroma residuals (920) and of coefficients of transformed down-sampled 4x4 chroma residuals (910), which can be obtained by setting to zero the coefficients out of a hamming window (930) in the 8x8 transform 930 (the coefficients of the high frequencies 920 are un-forced to zero).

[0086] In another implementation, the chroma residuals may first be computed at 4:4:4, then transformed into 4:2:0 coefficients directly, for example by dropping terms corresponding to high frequencies during the calculation of the transform coefficients.

[0087] In some implementations, at the decoder side, the decoded (inverse quantized) transform coefficients of the residuals may be 4:2:0. The 4:4:4 residuals may be obtained in several different ways as follows.

[0088] In some implementations, the inverse transform allows building 4:2:0 residuals, which may then be upscaled to obtain 4:4:4 residuals using linear up-sampling for example.

[0089] In some implementations, the transform coefficients block may be upscaled (using linear up-sampling, for example) to obtain a 4:4:4 matrix. The inverse transform may be used to build 4:4:4 residuals.

[0090] In some implementations, the transform coefficients may be zero- padded to obtain a 4:4:4 matrix as depicted in FIG. 9, discussed below. The inverse transform then allows building 4:4:4 residuals.

[0091] In some implementations, the inverse transform of the coefficients may be used to derive 4:4:4 residuals directly.

[0092] In some implementations, the chroma residuals may not be upscaled, and instead a different quantization parameter could be used for suchvideo signals. This may help avoid an increase in encoded data due to an artificial increase in the chroma residuals resolution. In such implementations, the quantization parameter of the chroma residuals may be increased when the internal chroma format is increased. For example, an additional offset of +6 may be applied to the quantization parameter of the chroma residuals when the internal chroma format is increased.

[0093] In still another implementation, a rate-distortion optimization may be applied. The rate-distortion optimization balances an amount of distortion and / or loss of video quality against the data rate or the amount of data required to encode the video signal. For example, at the encoder, for a given block, one or more coding modes and associated parameters (e.g., motion vectors, quantized coefficient values, etc.) may be selected using rate-distortion optimization:Cost = dist + lambda * nbits equation (1 ) with ‘dist’ denoting the distortion (for example, a sum of absolute differences (SAD), mean-removed sum of absolute difference (MRSAD), etc.), ‘nbits’ denoting the entropy coding cost or number of coded bits in the bitstream, and ‘lambda’ identifying a Lagrange multiplier (constraint minimization).

[0094] Regarding the chroma components, a distortion value (as a difference between reconstructed and original block) may be computed in different ways. For example, in some implementations, the distortion value may be calculated based on a reconstructed increased chroma format (ex: 4:4:4) and an original upscaled block. For example, if the distortion function used is a sum of absolute differences (SAD), then the distortion value is the sum of absolute differences between the upscaled original block samples and the reconstructed block samples. In other implementations, the distortion value may be calculated from a downscaled reconstructed block (ex: 4:2:0) and the original block. For example, if the distortion function is SAD, then distortion value is the sum of absolute differences between the original block samples and the reconstructed downscaled block samples. In still other implementations, the distortion value may be calculated from a sub-sampled reconstructed increased chroma format and the original block. For example, if the distortion function is SAD, thereconstructed block is 4:4:4 and the original is 4:2:0, then the distortion value is the sum of absolute differences between original samples and one out of four reconstructed block samples. The value of lambda may be adjusted to account for the chroma format used in computation of the distortion. The distortion may be scaled with a pre-determined scaling value so that the overall cost of one luma and chroma block obtained with IFCI may be compared with the overall cost of one luma and chroma block obtained without using ICFI.

[0095] Accordingly, in various implementations, the systems and methods discussed herein provide for increasing format of chroma components (e.g., chroma signals) internally to improve the inner representation of the reconstructed chroma components.

[0096] In a first aspect, the present disclosure is directed to a method for video coding. The method includes receiving, by a video encoder of a device, an input video data having a chroma sampling format at a first resolution (e.g., the original chroma format including the original chroma components and the luma components). The method also includes generating, by the video encoder, an intermediate video data by increasing the chroma subsampling format of the input video data to a second resolution (e.g., the upscaled chroma format including the upscaled chroma components and the luma components), the second resolution higher than the first resolution. The method also includes performing inter-frame, intra-frame, inter-block, and / or intra-block coding etc., by the video encoder, using the intermediate video data.

[0097] In some implementations, the first resolution comprises 4:2:2 or 4:2:0 chroma subsampling, and the second resolution comprises 4:4:4 chroma subsampling. In some implementations, increasing (e.g., upscaling) the chroma subsampling format comprises interpolating, by the video encoder, the one or more chroma samples at one or more locations between one or more existing chroma samples in the input video data. In some implementations, the method includes storing, by the video encoder, the reconstructed video data in a reference picture buffer at the second resolution.

[0098] In some implementations, performing inter-frame, intra-frame, interblock, and / or intra-block coding etc. further comprises calculating, by the videoencoder, a chroma residual having the second resolution. In a further implementation, the method includes downscaling, by the video encoder, the chroma residual to the first resolution. In another further implementation, the method includes increasing, by the video encoder, a quantization parameter of the chroma residual.

[0099] In some implementations, the method includes generating, by the video encoder, the output video data by decreasing (e.g., downscaling) the chroma subsampling format of the reconstructed video data to the first resolution. In some implementations, the method includes outputting, by the video encoder, an output video data with a chroma subsampling format at the second resolution.

[0100] Referring now to FIG. 10, a flowchart illustrating an example video decoding process 1000 is shown according to one or more embodiments. The video decoding process 1000 may be performed by a video decoder.

[0101] At 1010, the video decoder receives the bitstream (e.g. a coded representation) of one or more input blocks and / or images.

[0102] At 1020, the video decoder builds the prediction of the one or more input blocks and / or images.

[0103] At 1030, the video decoder decodes the one or more residuals.

[0104] At 1040, the video decoder adds the one or more residuals to the block prediction and reconstructs the one or more blocks and / or images.

[0105] At 1050, the video decoder stores the one or more decoded and / or reconstructed images at the upscaled chroma format in the buffer (e.g., the DPB).

[0106] At 1060, the video decoder displays the one or more output images in the upscaled chroma format.

[0107] Referring now to FIG. 11 , a flowchart illustrating an example video encoding process 1100 is shown according to one or more embodiments. The video encoding process 1100 is performed by a video encoder.

[0108] At 1110, the video encoder receives one or more input images in the original chroma format.

[0109] At 1120, the video encoder builds the prediction and the one or more residuals of input blocks and / or images.

[0110] At 1130, the video encoder adds the one or more residuals to the block prediction and reconstructs the one or more blocks and / or images.

[0111] At 1140, the video encoder stores the one or more decoded and / or reconstructed images at the upscaled chroma format in the buffer.

[0112] In the FIGs. 12-13 the one or more images are decoded / encoded at the original chroma format and upscaled to the upscaled chroma format before storage in the buffer (e.g., the DPB).

[0113] Referring now to FIG. 12, a flowchart illustrating an example decoding process 1200 including upscaling one or more reconstructed images is shown according to one or more embodiments. The video decoding process 1200 is performed by a decoder.

[0114] At 1210, the decoder receives a bitstream (e.g. the coded representation) of one or more input images in the original chroma format. The original chroma format is indicative of one or more original chroma components in the original chroma component size. The original chroma format is also indicative of one or more luma component sizes. In an example, the original chroma format can be any one of: 4:4:4, 4:2:0, or 4:2:2 etc.

[0115] At 1220, the decoder reconstructs the one or more input images at original chroma format

[0116] At 1230, the decoder generates, based on the reconstruction, one or more images in the upscaled chroma format. In an embodiment, the decoder upscales the original chroma component size to match the luma component size.

[0117] At 1240, the decoder stores the one or more reconstructed images as one or more reference images in the buffer in the memory of the decoder.

[0118] At 1250, the decoder generates and displays one or more output images in the upscaled chroma format.

[0119] In an embodiment, the decoder may further downscale the one or more reconstructed images to the original chroma format. The decoder may generate the one or more output images in the original chroma format.

[0120] Referring now to FIG. 13, a flowchart illustrating an example encoding process 1300 including upscaling one or more reconstructed images isshown according to one or more embodiments. The video encoding process 1300 is performed by an encoder.

[0121] At 1310, the encoder receives one or more input images in the original chroma format.

[0122] At 1320, the encoder encodes the one or more input images at the original chroma format and generates the bitstream.

[0123] At 1330, the encoder upscales the one or more reconstructed images.

[0124] At 1340, the encoder stores the one or more reconstructed images in the buffer.

[0125] In FIGs. 14-15 the one or more images are encoded and / or decoded at the upscaled chroma format, the one or more residuals are encoded and / or decoded at the original chroma format. The one or more reconstructed images are stored at the upscaled chroma format in the buffer (e.g., the DPB).

[0126] Referring now to FIG. 14, a flowchart illustrating an example decoding process 1400 including decoding the one or more residuals at the original chroma format and reconstructing the one or more images at the upscaled chroma format is shown according to one or more embodiments. The video decoding process 1400 is performed by a video decoder.

[0127] At 1410, the video decoder receives the bitstream (e.g. the coded representation) of one or more input blocks and / or images at the upscaled chroma format.

[0128] At 1420, the video decoder builds the prediction of the one or more input blocks and / or images at the upscaled chroma format and decodes the one or more residuals at the original chroma formats.

[0129] At 1430, the video decoder upscales the one or more residuals.

[0130] At 1440, the video decoder adds the one or more residuals to the block prediction and reconstructs the one or more blocks and / or images.

[0131] At 1450, the video decoder stores the one or more decoded and / or reconstructed images at the upscaled chroma format in the buffer.

[0132] At 1460, the video decoder displays the one or more output images in the upscaled chroma format.

[0133] Referring now to FIG. 15, a flowchart illustrating an example encoding process 1500 including upscaling one or more images before encoding and coding the one or more residuals at the original chroma format is shown according to one or more embodiments. The video encoding process 1500 is performed by a video encoder.

[0134] At 1510, the video encoder receives the one or more input images in the original chroma format.

[0135] At 1520, the video encoder upscales the one or more input images to the upscaled chroma format.

[0136] At 1530, the video encoder generates the one or more residuals based on the one or more reference images at the upscaled chroma format.

[0137] At 1540, the video encoder downscales and encodes the one or more residuals.

[0138] At 1550, the video encoder transmits the encoded residuals at the original chroma format.

[0139] In an embodiment, the one or more residuals can be in the upscaled chroma format. The encoder downscales the one or more residuals in the original chroma format. The encoder transforms the one or more residuals at the original chroma format.

[0140] In another embodiment, the encoder transforms the one or more residuals for generating one or more transformation coefficients. The encoder downscales the one or more transformation coefficients at the original chroma format.

[0141] In an embodiment, the encoder can dynamically determine a modified quantization parameter based on the upscaled chroma format. The encoder can quantize the one or more transformation coefficients based on the modified quantization parameter.

[0142] In an embodiment, the encoder can determine a distortion value based on a reconstructed block in the upscaled chroma format and an input block in the upscaled chroma format.

[0143] In an embodiment, the encoder can determine a distortion value based on the reconstructed block in the original chroma format and the input block in the original chroma format.

[0144] The encoder can use the distortion value and the number of bits to code the block (residuals and coding mode) to derive the cost and select the coding mode for encoding the input block with a lowest cost value.

[0145] In an embodiment, any of the video decoding processes and / or the video encoding processes of FIGs. 10-15 may be performed by a video codec.

[0146] In another aspect, the present disclosure is directed to a media consumption device, a network device, a computing device, or an integrated circuit configured to perform any of the methods discussed above.

[0147] In another aspect, the present disclosure is directed to at least one processor operatively connected to at least one transceiver, the at least one processor and at least one transceiver configured to perform any of the methods discussed above.

[0148] In another aspect, the present disclosure is directed to a non- transitory computer readable medium comprising instructions which when executed by a processing device cause the processing device to perform any of the methods discussed above.

[0149] In the present application, the terms “reconstructed” and “decoded” may be used interchangeably, the terms “encoded” or “coded” may be used interchangeably, the terms “pixel” or “sample” may be used interchangeably, and the terms “image,” “picture” and “frame” may be used interchangeably. Usually, but not necessarily, the term “reconstructed” is used at the encoder side while “decoded” is used at the decoder side.

[0150] Various methods are described herein, and each of the methods comprises one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. Additionally, terms such as “first”, “second”, etc. may be used in various embodiments to modify an element, component, step, operation, etc., for example, a “first decoding” and a “second decoding”. Use of such terms does notimply an ordering to the modified operations unless specifically required. So, in this example, the first decoding need not be performed before the second decoding, and may occur, for example, before, during, or in an overlapping time period with the second decoding.

[0151] Various methods and other aspects described in this application can be used to modify modules, for example, the motion compensation (170, 275), motion estimation (175), entropy coding, intra (160, 260) and / or decoding modules (145, 230), of a video encoder 100 and decoder 200 as shown in FIG. 1 and FIG. 2.

[0152] Moreover, the present aspects are not limited to H.261 , AVC, HEVC, or PAL-DV, and can be applied, for example, to other standards and recommendations, and extensions of any such standards and recommendations. Unless indicated otherwise, or technically precluded, the aspects described in this application can be used individually or in combination. Various numeric values are used in the present application. The specific values are for example purposes and the aspects described are not limited to these specific values.

[0153] Various implementations involve decoding. “Decoding,” as used in this application, may encompass all or part of the processes performed, for example, on a received encoded sequence in order to produce a final output suitable for display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, for example, entropy decoding, inverse quantization, inverse transformation, and differential decoding. Whether the phrase “decoding process” is intended to refer specifically to a subset of operations or generally to the broader decoding process will be clear based on the context of the specific descriptions and is believed to be well understood by those skilled in the art.

[0154] Various implementations involve encoding. In an analogous way to the above discussion about “decoding”, “encoding” as used in this application may encompass all or part of the processes performed, for example, on an input video sequence in order to produce an encoded bitstream.

[0155] The implementations and aspects described herein may be implemented in, for example, a method or a process, an apparatus, a softwareprogram, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method), the implementation of features discussed may also be implemented in other forms (for example, an apparatus or program). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. The methods may be implemented in, for example, an apparatus, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, for example, computers, cell phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate communication of information between end-users.

[0156] Reference to “one embodiment” or “an embodiment” or “one implementation” or “an implementation”, as well as other variations thereof, means that a particular feature, structure, characteristic, and so forth described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrase “in one embodiment” or “in an embodiment” or “in one implementation” or “in an implementation”, as well any other variations, appearing in various places throughout this application are not necessarily all referring to the same embodiment. Additionally, this application may refer to “determining” various pieces of information. Determining the information may include one or more of, for example, estimating the information, calculating the information, predicting the information, or retrieving the information from memory.

[0157] Further, this application may refer to “accessing” various pieces of information. Accessing the information may include one or more of, for example, receiving the information, retrieving the information (for example, from memory), storing the information, moving the information, copying the information, calculating the information, determining the information, predicting the information, or estimating the information.

[0158] Additionally, this application may refer to “receiving” various pieces of information. Receiving is, as with “accessing”, intended to be a broad term. Receiving the information may include one or more of, for example, accessing the information, or retrieving the information (for example, from memory). Further,“receiving” is typically involved, in one way or another, during operations, for example, storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information.

[0159] It is to be appreciated that the use of any of the following 7”, “and / or”, and “at least one of’, for example, in the cases of “A / B”, “A and / or B” and “at least one of A and B”, is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C”, such phrasing is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of the third listed option (C) only, or the selection of the first and the second listed options (A and B) only, or the selection of the first and third listed options (A and C) only, or the selection of the second and third listed options (B and C) only, or the selection of all three options (A and B and C). This may be extended, as is clear to one of ordinary skill in this and related arts, for as many items as are listed.

[0160] Also, as used herein, the word “signal” refers to, among other things, indicating something to a corresponding decoder. For example, in certain embodiments the encoder signals a quantization matrix for de-quantization. In this way, in an embodiment the same parameter is used at both the encoder side and the decoder side. Thus, for example, an encoder can transmit (explicit signaling) a particular parameter to the decoder so that the decoder can use the same particular parameter. Conversely, if the decoder already has the particular parameter as well as others, then signaling can be used without transmitting (implicit signaling) to simply allow the decoder to know and select the particular parameter. By avoiding transmission of any actual functions, a bit savings is realized in various embodiments. It is to be appreciated that signaling can be accomplished in a variety of ways.

[0161] For example, one or more syntax elements, flags, and so forth are used to signal information to a corresponding decoder in various embodiments.While the preceding relates to the verb form of the word “signal”, the word “signal” can also be used herein as a noun.

[0162] As will be evident to one of ordinary skill in the art, implementations may produce a variety of signals formatted to carry information that may be, for example, stored or transmitted. The information may include, for example, instructions for performing a method, or data produced by one of the described implementations. For example, a signal may be formatted to carry the bitstream of a described embodiment. Such a signal may be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal. The formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information that the signal carries may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is known. The signal may be stored on a processor-readable medium.

Claims

CLAIMSWhat is Claimed:

1. A decoder, comprising: a memory; a transceiver; and a processor, wherein the transceiver and the processor are configured to: receive a bitstream comprising a coded representation of one or more input images at an original chroma format, wherein an original chroma component size is lower than an original luma component size in the original chroma format, generate, based on the received bitstream, one or more prediction blocks and one or more residuals using one or more reference images stored in a buffer in the memory at an upscaled chroma format, wherein an upscaled chroma component size in the upscaled chroma format is larger than the original chroma component size, generate one or more reconstructed images based on the one or more prediction blocks and the one or more residuals, and store the one or more reconstructed images in the buffer in the upscaled chroma format.

2. The decoder of claim 1 , wherein the processor is further configured to: generate one or more output images in the upscaled chroma format.

3. The decoder of claim 1 , wherein the processor is further configured to: downscale the one or more reference images to the original chroma format.

4. The decoder of any of claim 1 or 3, wherein the processor is further configured to: generate one or more output images in the original chroma format.

5. The decoder of any of claims 1 to 4, wherein the upscaled chroma format is 4:4:4 and the original chroma format is at least one of: 4:2:0 or 4:2:2.

6. The decoder of any of claims 1 to 5, wherein upscaling the original chroma format includes upscaling the original chroma component size to match the original luma component size.

7. The decoder of claim 1 , wherein the one or more prediction blocks are generated at the original chroma format, and wherein the one or more residuals are generated at the original chroma format.

8. The decoder of claim 7, wherein the processor is further configured to: upscale the one or more reconstructed images to the upscaled chroma format.

9. The decoder of claim 1 , wherein the one or more prediction blocks are generated at the upscaled chroma format, and wherein the one or more residuals are generated at the upscaled chroma format.

10. The decoder of claim 1 , wherein the one or more prediction blocks are generated at the upscaled chroma format, and wherein the one or more residuals are generated at the original chroma format.

11. The decoder of claim 10, wherein the processor is further configured to: upscale the one or more residuals to the upscaled chroma format, and generate the one or more reconstructed images at the upscaled chroma format.

12. A method performed by a decoder, the method comprising: receiving a bitstream comprising a coded representation of one or more input images at an original chroma format, wherein an original chroma component size is lower than a luma component size in the original chroma format; generating, based on the received bitstream, one or more prediction blocks and one or more residuals using one or more reference images stored in a buffer at an upscaled chroma format, wherein an upscaled chroma component size in the upscaled chroma format is larger than the original chroma component size; generating one or more reconstructed images based on the one or more prediction blocks and the one or more residuals; and storing the one or more reconstructed images in the buffer in the upscaled chroma format.

13. A method performed by an encoder, the method comprising: receiving one or more input images in an original chroma format, wherein an original chroma component size is smaller than an original luma component size in the original chroma format; generating, based on the one or more input images, one or more prediction blocks and one or more residuals; generating one or more reconstructed images based on the one or more prediction blocks and the one or more residuals; storing the one or more reconstructed images in a buffer as one or more reference images in an upscaled chroma format, wherein an upscaled chroma component size in the upscaled chroma format is larger than the original chroma component size; and encoding the one or more residuals.

14. The method of claim 13, wherein the one or more residuals are in the original chroma format.

15. The method of claim 13, wherein the one or more residuals are in the upscaled chroma format.

16. The method of any of claim 13 or 15, wherein encoding the one or more residuals comprises one of: downscaling the one or more residuals in the original chroma format; or transforming the one or more residuals at the original chroma format.

17. The method of any of claim 13 or 15, wherein encoding the one or more residuals comprises: transforming the one or more residuals for generating one or more transformation coefficients; and downscaling the one or more transformation coefficients at the original chroma format.

18. The method of any of claims 13 to 17, wherein encoding the one or more residuals further comprises: dynamically determining a modified quantization parameter based on the upscaled chroma format; and quantizing the one or transformation coefficients based on the modified quantization parameter.

19. The method of any of claims 13 to 18, the method further comprising: determining a distortion value based on a reconstructed block in the upscaled chroma format and an input block in the upscaled chroma format; deriving, based on the distortion value, a cost and a number of bits to encode; and selecting a coding mode for encoding the input block based on the cost.

20. The method of any of claims 13 to 18, the method further comprising: determining a distortion value based on a down-scaled reconstructed block in the original chroma format and an input block in the original chroma format; deriving a cost based on distortion and a number of bits to encode; and selecting a coding mode for encoding the input block based on the cost.

Citation Information

Patent Citations

  • Video or image coding based on mapping of LUMA samples and scaling of chroma samples

    US20230024786A1

  • Video encoding and decoding based on resampling chroma signals

    US20230058283A1

  • Method and apparatus for video encoding and decoding with chroma residuals sampling

    WO2023041317A1