Method, device, and non-transitory computer-readable medium for a chroma from lumintra prediction mode

The method optimizes chroma from luma intra prediction by selecting appropriate downsampling filters, addressing inefficiencies in existing video coding technologies and improving accuracy and complexity in video coding formats.

JP2025524755APending Publication Date: 2025-08-01TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024546156
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-09
Filing Date
2022-11-15
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Existing video coding technologies, such as AV1 and HEVC, face challenges in efficiently performing chroma from luma (CfL) intra prediction due to suboptimal downsampling filters, which can lead to inaccuracies and increased computational complexity.

Method used

A method and device for performing CfL intra prediction by selecting an appropriate downsampling filter based on a syntax element from the coded bitstream, determining luma sample positions, and reconstructing the chroma block using downsampled luma samples, with support for multiple downsampling filters to accommodate different chroma downsampling formats.

Benefits of technology

Improves prediction accuracy and reduces computational complexity by optimizing the downsampling process, enhancing the efficiency of video coding, particularly in formats like 4:2:0 and 4:2:2.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025524755000001_ABST
    Figure 2025524755000001_ABST
Patent Text Reader

Abstract

A method and apparatus for performing Chroma from Luma (CfL) intra prediction, comprising: receiving a current chroma block from a coded bitstream; determining whether a CfL intra prediction mode is enabled for the coded bitstream; determining a color format corresponding to the coded bitstream; based on determining that the CfL intra prediction mode is enabled and based on the determined color format, obtaining a syntax element indicating a downsampling filter used in the CfL intra prediction mode from the coded bitstream; selecting a downsampling filter to be used for the chroma block in the CfL intra prediction mode from among a plurality of downsampling filters; determining a luma sample position associated with the current chroma block based on the selected downsampling filter; downsampling a plurality of luma samples at the luma sample position, wherein pixels within each downsampled luma sample are co-located with corresponding pixels within the current chroma block; and reconstructing the current chroma block based at least on the plurality of downsampled luma samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Cross - Reference to Related Applications] This application claims the benefit of U.S. Provisional Application No. 63 / 390,714, filed with the United States Patent and Trademark Office on July 20, 2022, and U.S. Application No. 17 / 983,936, filed with the United States Patent and Trademark Office on November 9, 2022, the disclosures of which are hereby incorporated by reference in their entirety.

[0002] [Technical Field] Embodiments of the present disclosure relate to a set of advanced video coding techniques, and more particularly, to a signaling downsampling filter for cross - component intra prediction mode.

Background Art

[0003] AOMedia Video1 (AV1) is an open video coding format designed for video transmission over the Internet. It was developed by the Alliance for Open Media (AOMedia) as a successor to VP9. AOMedia is a group established in 2015 that includes semiconductor companies, video - on - demand providers, video content producers, software development companies, and web browser vendors. Many of the components of the AV1 project were derived from previous research results by Alliance members. Individual contributors had started experimental technical platforms several years earlier as follows: Daala by Xiph / Mozilla (registered trademark) had its code published in 2010, Google (registered trademark)'s experimental VP9 evolution project VP10 was announced on September 12, 2014, and Cisco (registered trademark)'s Thor was published on August 11, 2015. AV1, built on the VP9 codebase, incorporates additional technologies, some of which were developed in these experimental formats. The first version of the AV1 reference codec, version 0.1.0, was published on April 7, 2016. The Alliance released the AV1 bitstream specification on March 28, 2018, along with a software - based reference encoder and decoder. On June 25, 2018, a verified version 1.0.0 of the specification was released. On January 8, 2019, the "AV1 Bitstream & Decoding Process Specification", a verified version 1.0.0 that includes errata sheet 1 of the specification, was released. The AV1 bitstream specification includes a reference video codec. The "AV1 Bitstream & Decoding Process Specification" (version 1.0.0 and errata sheet 1), Alliance for Open Media (January 8, 2019) is hereby incorporated by reference in its entirety into this document.

[0004] The High Efficiency Video Coding (HEVC) standard was jointly developed by the standardization organizations of ITU-T Video Coding Experts Group (VCEG) and ISO / IEC Moving Picture Experts Group (MPEG). To develop the HEVC standard, these two standardization organizations cooperate in a partnership known as the Joint Collaborative Team on Video Coding (JCT-VC). The first version of the HEVC standard was completed in January 2013, resulting in a harmonized text issued by both ITU-T and ISO / IEC. Subsequently, additional work was organized to extend the standard to support several additional application scenarios, including extended range of use with enhanced accuracy and support for color formats, scalable video coding, 3D / stereo / multi-view video coding. In ISO / IEC, the HEVC standard became MPEG-H Part 2 (ISO / IEC 23008-2), and in ITU-T it became ITU-T Recommendation H.265. The specification of the HEVC standard "SERIES H: AUDIOVISUAL AND MULTIMEDIA SYSTEMS, Infrastructure of audiovisual services - Coding of moving video", ITU-T H.265, International Telecommunication Union (April 2015) is hereby incorporated by reference in its entirety into this specification.

[0005] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC1 / SC29 / WG11) published the H.265 / HEVC (High Efficiency Video Coding) standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Since then, they have been considering the potential need for standardization of future video coding technologies that may significantly outperform HEVC in terms of compression capabilities. In October 2017, they announced a joint call for proposals on video compression (CfP) with capabilities beyond HEVC. By February 15, 2018, 22 CfP responses regarding standard dynamic range (SDR), 12 CfP responses regarding high dynamic range (HDR), and 12 CfP responses regarding 360 video categories were submitted respectively. In April 2018, all received CfP responses were evaluated at the 122nd MPEG / Joint Video Exploration Team-Joint Video Expert Team (JVET) meeting. After careful evaluation, JVET officially started the standardization of the next-generation video coding beyond HEVC, namely, the so-called Versatile Video Coding (VVC). The specification of the VVC standard, "Versatile Video Coding (Draft 7)", JVET-P2001-vE, Joint Video Experts Team (October 2019), is hereby incorporated by reference in its entirety. Another specification of the VVC standard, "Versatile Video Coding (Draft10)", JVET-S2001-vE, Joint Video Experts Team (July 2020), is hereby incorporated by reference in its entirety.

Summary of the Invention

[0006] According to an aspect of the present disclosure, a method for performing chroma from luma (CfL) intra prediction includes: receiving a current chroma block from a coded bitstream; determining whether a CfL intra prediction mode is enabled for the coded bitstream; determining a color format corresponding to the coded bitstream; obtaining, based on determining that the CfL intra prediction mode is enabled and based on the determined color format, a syntax element indicating a downsampling filter used for the CfL intra prediction mode from the coded bitstream; selecting, from a plurality of downsampling filters, a downsampling filter to be used for the chroma block in the CfL intra prediction mode; determining a luma sample position associated with the current chroma block based on the selected downsampling filter; downsampling a plurality of luma samples at the luma sample position, wherein pixels within each downsampled luma sample are co-located with corresponding pixels within the current chroma block; and reconstructing the current chroma block based at least on the plurality of downsampled luma samples.

[0007] According to an aspect of the present disclosure, a device for performing Chroma from Luma (CfL) intra prediction, the device includes at least one memory configured to store program code, and at least one processor configured to access the program code and operate as instructed by the program code, and the program code includes: reception code configured to cause the at least one processor to receive a current chroma block from a coded bitstream; first determination code configured to cause the at least one processor to determine whether a CfL intra prediction mode is enabled for the coded bitstream; second determination code configured to cause the at least one processor to determine a color format corresponding to the coded bitstream; acquisition code configured to cause the at least one processor to obtain a syntax element indicating a downsampling filter used for the CfL intra prediction mode from the coded bitstream based on a determination that the CfL intra prediction mode is enabled and based on the determined color format; selection code configured to cause the at least one processor to select a downsampling filter to be used for the chroma block in the CfL intra prediction mode from among a plurality of downsampling filters; third determination code configured to cause the at least one processor to determine a luma sample position associated with the current chroma block based on the selected downsampling filter; downsampling code configured to cause the at least one processor to downsample a plurality of luma samples at the luma sample position, wherein pixels within each downsampled luma sample are co-located with corresponding pixels within the current chroma block; and reconstruction code configured to cause the at least one processor to reconstruct the current chroma block based at least on the plurality of downsampled luma samples.

[0008] According to an aspect of the present disclosure, a non-transitory computer-readable medium stores instructions that, when executed by one or more processors of a device for performing Chroma from Luma (CfL) intra prediction, cause the one or more processors to: receive a current chroma block from a coded bitstream; determine whether a CfL intra prediction mode is enabled for the coded bitstream; determine a color format corresponding to the coded bitstream; based on determining that the CfL intra prediction mode is enabled and based on the determined color format, obtain a syntax element indicating a downsampling filter used for the CfL intra prediction mode from the coded bitstream; select a downsampling filter to be used for the chroma block in the CfL intra prediction mode from among a plurality of downsampling filters; determine a luma sample position associated with the current chroma block based on the selected downsampling filter; downsample a plurality of luma samples at the luma sample position, wherein pixels within each downsampled luma sample are co-located with corresponding pixels within the current chroma block; and reconstruct the current chroma block based at least on the plurality of downsampled luma samples, including one or more instructions.

Brief Description of the Drawings

[0009] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.

[0010]

Figure 1

[0011]

Figure 2

[0012]

Figure 3

[0013]

Figure 4

[0014]

Figure 5

[0015]

Figure 6

[0016]

Figure 7

[0017]

Figure 8

[0018]

Figure 9

[0019]

Figure 10

[0020]

Figure 11A

Figure 11B

Figure 11C

Figure 11D

Figure 11E

[0021]

Figure 12

[0022]

Figure 13

[0023]

Figure 14

DETAILED DESCRIPTION OF THE INVENTION

[0024] In the present disclosure, the term "block" may be interpreted as a prediction block, a coding block, or a coding unit (CU). In this specification, the term "block" may also be used to refer to a transform block.

[0025] In the present disclosure, the term "transform set" refers to a group of transform kernel (or candidate) options. A transform set may include one or more transform kernel (or candidate) options. According to embodiments of the present disclosure, when two or more transform options are available, an index may be signaled to indicate which of the transform options within the transform set is applied to the current block.

[0026] In the present disclosure, the term "prediction mode set" refers to a group of prediction mode options. A prediction mode set may include one or more prediction mode options. According to embodiments of the present disclosure, when two or more prediction mode options are available, an index may be further signaled to indicate which of the prediction mode options within the prediction mode set is applied to the current block for performing prediction.

[0027] In the present disclosure, the term "adjacent reconstructed sample set" refers to a group of reconstructed samples from previously decoded adjacent blocks or reconstructed samples within a previously decoded picture.

[0028] In the present disclosure, the term "neural network" refers to the general concept of a data processing structure having one or more layers, as described herein with respect to "deep learning for video coding". According to embodiments of the present disclosure, any neural network may be configured to implement the embodiments.

[0029] FIG. 1 shows a simplified block diagram of a communication system (100) according to an embodiment of the present disclosure. The system (100) may include at least two terminals (110, 120) interconnected via a network (150). In a one-way data transmission, the first terminal (110) may code video data at a local location for transmission to another terminal (120) via the network (150). The second terminal (120) may receive the coded video data of another terminal from the network (150) and decode the coded data to display the recovered video data. One-way data transmission may be common in media delivery applications and the like.

[0030] FIG. 1 shows a second pair of terminals (130, 140) provided to support two-way transmission of coded video, which may occur, for example, during a video conference. In a two-way data transmission, each terminal (130, 140) may code video data captured at a local location for transmission to another terminal via the network (150). Each terminal (130, 140) may also receive the coded video data transmitted by another terminal, may decode the coded data, and may display the recovered video data on a local display device.

[0031] In FIG. 1, the terminals (110-140) may be illustrated as servers, personal computers, and smart phones, and / or any other type of terminal. For example, the terminals (110-140) may be laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. The network (150) represents any number of networks that transmit coded video data between the terminals (110-140), including, for example, wired and / or wireless communication networks. The communication network (150) may exchange data over circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, the architecture and topology of the network (150) may not be important for the operation of the present disclosure, unless otherwise described herein below.

[0032] FIG. 2 illustrates the placement of video encoders and decoders in a streaming environment as an example of the application of the disclosed subject matter. The disclosed subject matter may be equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storage of compressed video on digital media such as CDs and DVDs, memory sticks, etc.

[0033] As shown in FIG. 2, the streaming system (200) may include a capture subsystem (213) that can include a video source (201) and an encoder (203). The video source (201) may be, for example, a digital camera and may be configured to create an uncompressed video sample stream (202). The uncompressed video sample stream (202) may provide a high data volume compared to an encoded video bitstream and may be processed by an encoder (203) coupled to the video source (201), and the video source (201) may be, for example, a camera. The encoder (203) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as will be described in more detail below. The encoded video bitstream (204) may include a lower data volume compared to the sample stream and may be stored in a streaming server (205) for future use. One or more streaming clients (206) may access the streaming server (205) to retrieve a video bitstream (209) that may be a copy of the encoded video bitstream (204).

[0034] In embodiments, the streaming server (205) may also function as a media-aware network element (MANE). For example, the streaming server (205) may be configured to prune the encoded video bitstream (204) to adapt potentially different bitstreams to one or more of the streaming clients (206). In embodiments, the MANE may be provided separately from the streaming server (205) within the streaming system (200).

[0035] The streaming client (206) can include a video decoder (210) and a display (212). The video decoder (210) can decode, for example, a video bitstream (209) that is an input copy of an encoded video bitstream (204) and create an output video sample stream (211) that can be rendered on the display (212) or another rendering device (not shown). In some streaming systems, the video bitstreams (204, 209) can be encoded according to a particular video coding / compression standard. Examples of such standards include, but are not limited to, ITU-T Recommendation H.265. In one example, a video coding standard under development is informally known as VVC (Versatile Video Coding). The disclosed subject matter can be used in the context of VVC.

[0036] FIG. 3 illustrates an exemplary functional block diagram of a video decoder (210) attached to a display (212) according to an embodiment of the present disclosure.

[0037] The video decoder (210) can include a channel (312), a receiver (310), a buffer memory (315), an entropy decoder / parser (320), a scaler / inverse transform unit (351), an intra prediction unit (352), a motion compensation prediction unit (353), an aggregator (355), a loop filter unit (356), a reference picture memory (357), and a current picture memory (358). In at least one embodiment, the video decoder (210) can include an integrated circuit, a series of integrated circuits, and / or other electronic circuits. The video decoder (210) can also be embodied, in part or in whole, as software executed on one or more CPUs having associated memory.

[0038] Receiver (310) may receive one or more coded video sequences to be decoded by decoder (210), receiving one coded video sequence at a time, in which case the decoding of each coded video sequence is independent of other coded video sequences. The coded video sequence may be received from channel (312), which may be a hardware / software link to a storage device storing the encoded video data. Receiver (310) may receive the encoded video data together with other data, such as coded audio data and / or auxiliary data streams, and these data may be transferred to their respective usage entities (not shown). Receiver (310) may separate the coded video sequence from other data. To eliminate network jitter, buffer memory (315) may be coupled between receiver (310) and entropy decoder / parser (320) (hereinafter, "parser (320)"). Buffer memory (515) may not be used or may be made small when receiver (310) is receiving data from a store / forward device having sufficient bandwidth and controllability or from an isosynchronous network. For use in a best effort packet network such as the Internet, buffer memory (315) may be required and its size may be relatively large and may be an adaptive size.

[0039] Video decoder (210) may include a parser (320) to reconstruct symbols (321) from an entropy-coded video sequence. The categories of these symbols include, for example, information used to manage the operation of the decoder (210) and, potentially, information to control a rendering device such as a display (212) that may be coupled to the decoder illustrated in FIG. 2. The control information for the rendering device may be in the form of, for example, a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (320) may syntax analyze / entropy decode the received, coded video sequence. The coding of the coded video sequence may follow video coding techniques or standards and may follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (320) may extract a set of subgroup parameters for at least one subgroup of a subgroup of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. Subgroups can include Groups of Pictures (GOPs), pictures, tiles, slices, macroblocks, Coding Units (CUs), blocks, Transform Units (TUs), Prediction Units (PUs), etc. The parser (320) may also extract from coded video sequence information such as transform coefficients, quantization parameter values, motion vectors, etc.

[0040] The parser (320) may perform an entropy decoding / syntax analysis operation on the video sequence received from the buffer memory (315) to create symbols (321).

[0041] The reconstruction of symbol (321) can involve multiple different units depending on the type of the coded video picture or a part thereof (e.g., inter and intra pictures, inter and intra blocks) and other factors. Which units are involved and how they are involved can be controlled by subgroup control information parsed by parser (320) from the coded video sequence. The flow of such subgroup control information between parser (320) and the multiple units below is not shown for clarity.

[0042] In addition to the functional blocks already described, decoder (210) can be conceptually subdivided into multiple functional units as described below. In a practical implementation operating under commercial constraints, many of these units can interact closely with each other and can be at least partially integrated with each other. However, for the purpose of explaining the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.

[0043] One unit can be a scaler / inverse transform unit (351). The scaler / inverse transform unit (351) can receive, as symbol (321) from parser (320), not only control information including which transform to use, block size, quantization coefficients, quantization scaling matrix, etc., but also quantized transform coefficients. The scaler / inverse transform unit (351) can output a block including sample values that can be input to aggregator (355).

[0044] In some cases, the output samples of the scaler / inverse transform unit (351) may be related to intra-coded blocks; that is, blocks that do not use prediction information from previously reconstructed pictures but can use prediction information from previously reconstructed parts of the current picture. Such prediction information can be provided by the intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) uses surrounding already reconstructed information fetched from the current (partially reconstructed) picture in the current picture memory (358) to generate blocks of the same size and shape as the block being reconstructed. The aggregator (355) may, in some cases, add the prediction information generated by the intra prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351) on a sample-by-sample basis.

[0045] In other cases, the output samples of the scaler / inverse transform unit (351) may be related to inter-coded blocks that are potentially motion compensated. In such cases, the motion compensation prediction unit (353) can access the reference picture memory (357) to fetch the samples used for prediction. According to the symbol (321) related to the block, after motion compensating the fetched samples, these samples can be added by the aggregator (355) to the output of the scaler / inverse transform unit (351) (in this case, called the residual samples or residual signal) to generate the output sample information. The address in the reference picture memory (357) where the motion compensation prediction unit (353) fetches the prediction samples can be controlled by the motion vector. The motion vector can be available to the motion compensation prediction unit (353) in the form of a symbol (321) that can have, for example, X, Y and reference picture components. Motion compensation can also include interpolation of the sample values fetched from the reference picture memory (357) when an exact motion vector of sub-samples is used, a motion vector prediction mechanism, etc.

[0046] The output samples of the aggregator (355) may be subject to various loop filtering techniques within the loop filter unit (356). Video compression techniques can be controlled by parameters included in the coded video stream and can include in-loop filter techniques made available to the loop filter unit (356) as symbols (321) from the parser (320), but can also respond to meta information obtained during the decoding of previous parts (in decoding order) of the coded picture or coded video sequence, and can respond to previously reconstructed and loop filtered sample values.

[0047] The output of the loop filter unit (356) can be an output sample stream that can be output to a rendering device such as the display (212) and can be stored in the reference picture memory (357) for use in future inter-picture prediction.

[0048] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. When the current coded picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (320)), the current reference picture can become part of the reference picture memory (357), and the fresh current picture memory can be reallocated before starting the reconstruction of subsequent coded pictures.

[0049] The video decoder (210) may perform a decoding operation according to a predetermined video compression technique that can be documented by a standard such as ITU-T Rec. H.265. The coded video sequence may conform to the syntax specified by the video compression technique or standard being used, in the sense of conforming to the syntax of the video compression technique or standard as specified in the video compression technique document and standard, particularly in the profile document therein. Also, for compliance with some video compression techniques and standards, the complexity of the coded video sequence may be within the range defined by the level of the video compression technique or standard. In some cases, the level limits, for example, the maximum picture size, the maximum frame rate, the maximum reconstruction sample rate (measured, for example, in megasamples per second), the maximum reference pixel size, etc. The limits set by the level may, in some cases, be further restricted through the HRD specifications and metadata for buffer management of the Hypothetical Reference Decoder (HRD) signaled in the coded video sequence.

[0050] In one embodiment, the receiver (310) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The additional data can be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0051] FIG. 4 illustrates an exemplary functional block diagram of a video encoder (203) associated with a video source (201) according to an embodiment of the present disclosure.

[0052] The video encoder (203) may include an encoder such as, for example, a source coder (430), a coding engine (432), a (local) decoder (433), a reference picture memory (434), a predictor (435), a transmitter (440), an entropy coder (445), a controller (450), and a channel (460).

[0053] The encoder (203) may receive video samples from a video source (201) (which is not part of the encoder) that may capture a video image to be coded by the encoder (203).

[0054] The video source (201) may provide a source video sequence to be coded by the encoder (203) in the form of a digital video sample stream that can have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits,...), any color space (e.g., BT.601 YCrCb, RGB,...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media supply system, the video source (201) may be a storage device that stores pre-prepared video. In a video conferencing system, the video source (201) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that convey motion when viewed in sequence. Each picture itself may be organized as a spatial array of pixels, in which case each pixel can include one or more samples depending on the sampling structure, color space, etc. in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.

[0055] According to one embodiment, the encoder (203) can code and compress the pictures of the source video sequence in real time or under any other time constraints required by the application to obtain a coded video sequence (443). Implementing an appropriate coding speed is one function of the controller (450). The controller (450) may also control other functional units and can be functionally coupled to these units, as will be described below. This coupling is not shown for clarity. The parameters set by the controller (450) can include rate control related parameters (picture skip, quantizer, lambda value of rate distortion optimization techniques,...), picture size, layout of the group of pictures (GOP), maximum motion vector search range, etc. Those skilled in the art can easily identify other functions of the controller (450) since they may be related to the video encoder (203) optimized for a specific system design.

[0056] Some video encoders operate in what those skilled in the art would readily recognize as a "coding loop." As an overly simplified explanation, the coding loop may consist of an encoding portion of a source coder (430) (responsible for creating symbols based on, for example, input pictures and reference pictures to be coded), and a (local) decoder (433) embedded in the encoder (203), which reconstructs symbols to create sample data that a (remote) decoder would also create when the compression between the symbols and the coded video bitstream is reversible with a particular video compression technique. The reconstructed sample stream may be input into a reference picture memory (434). Since the decoding of the symbol stream yields bit-exact results independent of the decoder location (local or remote), the contents of the reference picture memory are also bit-exact between the local encoder and the remote encoder. In other words, the prediction portion of the encoder "sees" the same sample values as reference picture samples that the decoder "sees" when the decoder uses prediction during decoding. This basic principle of the synchronicity of reference pictures (and the resulting drift in the case where synchronicity cannot be maintained, for example, due to channel errors) is known to those skilled in the art.

[0057] The operation of the "local" decoder (433) can be made the same as that of the "remote" decoder (210), which has already been described above in connection with FIG. 3. However, since the symbols are available and the encoding / decoding of the symbols into the coded video sequence by the entropy coder (445) and the parser (320) can be reversible, the entropy decoding portion of the decoder (210) including the channel (312), the receiver (310), the buffer memory (315), and the parser (320) may not be fully implemented in the local decoder (433).

[0058] The observations that can be made at this point are that any decoder technology other than the parsing / entropy decoding present in the decoder may need to exist in substantially the same functional form in the corresponding encoder. For this reason, the disclosed subject matter focuses on the operation of the decoder. The description of encoder technology may be omitted since it may be the opposite of the decoder technology described comprehensively. More detailed descriptions will be required and provided below only in certain areas.

[0059] As part of its operation, the source coder (430) may perform motion compensation predictive coding, which predictively codes an input frame in relation to one or more previously coded frames from a video sequence designated as a "reference frame". In this way, the coding engine (432) codes the difference between a pixel block of the input frame and a pixel block of a reference frame that can be selected as a prediction reference for the input frame.

[0060] The local video decoder (433) may decode the coded video data of a frame that can be designated as a reference frame based on the symbols created by the source coder (430). The operation of the coding engine (432) may advantageously be an irreversible process. When the coded video data can be decoded by a video decoder (not shown in FIG. 4), the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder (433) may replicate the decoding process that can be performed by the video decoder with respect to the reference frame and store the reconstructed reference frame in the reference picture memory (434). In this way, the encoder (203) may locally store a copy of the reconstructed reference frame having common content as the reconstructed reference frame obtained by the remote video decoder (without transmission errors).

[0061] Predictor (435) may perform predictive search for the coding engine (432). That is, for a new frame to be coded, predictor (435) may search the reference picture memory (434) for sample data (as candidate reference pixel blocks) or specific metadata such as reference picture motion vectors, block shapes, etc., which may function as appropriate prediction references for the new picture. Predictor (435) may operate on a sample block-by-pixel block basis to find an appropriate prediction reference. In some cases, the input picture may have prediction references drawn from a plurality of reference pictures stored in the reference picture memory (434), as determined by the search results obtained by predictor (435).

[0062] Controller (450) may manage the coding operations of video coder (430), including setting parameters and subgroup parameters used, for example, to encode video data.

[0063] Outputs of all of the aforementioned functional units may be subject to entropy coding in entropy coder (445). The entropy coder converts symbols generated by various functional units into a coded video sequence by reversibly compressing the symbols according to techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc.

[0064] The transmitter (440) may buffer the coded video sequence created by the entropy coder (445) to prepare for transmission via the communication channel (460), which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (440) may merge the coded video data from the video coder (430) with other data to be transmitted, such as coded audio data and / or an auxiliary data stream (source not shown).

[0065] The controller (450) may manage the operation of the encoder (203). During coding, the controller (450) may assign a specific coded picture type to each coded picture, which may affect the coding applicable to each picture. For example, a picture may often be assigned as an intra picture (I picture), a predicted picture (P picture), or a bi-directionally predicted picture (B picture):

[0066] An intra picture (I picture) may be coded and decoded without using any other frames in the sequence as a source of prediction. Some video codecs allow different types of intra pictures, such as an Independent Decoder Refresh (IDR) picture. Those skilled in the art are aware of these variations of I pictures and their respective uses and characteristics.

[0067] A predicted picture (P picture) may be coded and decoded using intra prediction or inter prediction, using at most one motion vector and a reference index, to predict the sample values of each block.

[0068] A bidirectional predicted picture (B picture) can be coded and decoded using intra prediction or inter prediction, using up to two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple-predictive pictures can use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0069] The source picture is typically subdivided spatially into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8 or 16×16 samples each) and can be coded block by block. The blocks can be coded predictively in relation to other (already coded) blocks, as determined by the coding assignment applied to each picture of the block. For example, blocks of an I picture may be coded non-predictively, or they may be coded predictively in relation to already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be coded non-predictively via spatial prediction or via temporal prediction in relation to one previously coded reference picture. Blocks of a B picture can be coded non-predictively via spatial prediction or via temporal prediction in relation to one or two previously coded reference pictures.

[0070] The video coder (203) can perform coding operations according to a predetermined video coding technology or standard such as ITU-T Rec. H.265. In that operation, the video coder (203) can perform various compression operations, including predictive coding operations that utilize the temporal and spatial redundancies in the input video sequence. The coded video data can thus conform to the syntax specified by the video coding technology or standard being used.

[0071] In one embodiment, the transmitter (440) may transmit additional data along with the encoded video. The video coder (430) may include such data as part of the coded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, supplementary enhancement information (SEI) messages, visual user ability information (VUI) parameter set fragments, and the like.

[0072] [AV1's Directional Intra Prediction] VP9 supports eight direction modes corresponding to angles from 45 degrees to 207 degrees. To more diversely utilize the spatial redundancy of the directional texture, in AV1, the directional intra mode is extended to a finer-grained set of angles. The original eight angles are slightly modified and created as nominal angles, which are named V_PRED(542), H_PRED(543), D45_PRED(544), D135_PRED(545), D113_PRED(5446), D157_PRED(547), D203_PRED(548), and D67_PRED(549), and are illustrated in FIG. 5 in relation to the current block (541). For each nominal angle, there are seven finer angles, so AV1 has a total of 56 direction angles. The predicted angle is represented by adding an angle delta to the nominal intra angle, which is -3 to 3 times the step size of 3 degrees. In AV1, eight nominal modes and five non-directional smooth modes are signaled first. Then, if the current mode is an angle mode, an index is further signaled, indicating the angle delta for the corresponding nominal angle. To implement the directional prediction mode of AV1 in a general way, all 56 directional intra prediction modes of AV1 are implemented by an integrated direction predictor that projects each pixel to a reference sub-pixel position and interpolates the reference pixels by a 2-tap bilinear filter.

[0073] [AV1's Non-Directional Smooth Intra Predictor] In AV1, there are five non-directional smooth intra prediction modes: DC, PAETH, SMOOTH, SMOOTH_V, and SMOOTH_H. In DC prediction, the average of the left and upper adjacent samples is used as the predictor for the block to be predicted. In PAETH prediction, the upper, left, and upper-left reference samples are first fetched, and then the value closest to (upper + left - upper-left) is set as the predictor for the pixel to be predicted. FIG. 6 illustrates the positions of the upper sample (554), left sample (556), and upper-left sample (558) of the pixel (552) within the current block (550). In the SMOOTH, SMOOTH_V, and SMOOTH_H modes, the current block (550) is predicted using vertical or horizontal quadratic interpolation or the average in both directions.

[0074] [Chroma predicted from luma] In addition to the above modes, Chroma from Luma (CfL) is an intra predictor for chroma only that models the luminance pixels as a linear function of the simultaneously reconstructed luma pixels. CfL prediction may be expressed as shown in Equation (1) below: [Equation] In Equation (1), L AC represents the AC contribution of the luma component, α represents the parameter of the linear model, and DC represents the DC contribution of the chroma component.

[0075] FIG. 7 provides a graph of the linear function described by formula (1). As can be seen from FIG. 7 and formula (1), the reconstructed luma pixels are subsampled to chroma resolution and then the average value is subtracted to form the AC contribution. Instead of requiring the decoder to calculate a scaling parameter as in some prior arts to approximate the chroma AC component from the AC contribution, AV1 CfL can determine the parameter α based on the original chroma pixels and signal these in the bitstream. This reduces the complexity of the decoder and provides a more accurate prediction. For the DC contribution of the chroma component, an intra DC mode may be used for calculation, which is sufficient for most chroma content and has a mature and fast implementation.

[0076] In CfL mode, when some samples within a co-located luma block are outside the picture boundary, these samples can be padded and used to calculate the average of the luma samples. As shown in FIG. 8, the samples within area 802 of the current block 800 are outside the picture as indicated by the picture boundary 804, and these samples can be padded by copying the values of the nearest available samples within the current block 800, for example, the samples included within area 806 of the current block 800.

[0077] In CfL mode, the luma subsampling step is combined with the average subtraction step as shown in FIG. 7. In this way, not only is the formula simplified, but also the subsampling division and the corresponding rounding error are removed. An example of the formula corresponding to the combination of both steps is given in formula (2), which can be simplified to form formula (3):

Equation

[0078] Note that both equations use integer division. MxN represents a matrix of pixels in the luma plane.

[0079] Based on the supported chroma subsampling, it can be shown that for Sx×Sy ε {1,2,4}, and since both M and N are powers of 2, M×N is also a power of 2.

[0080] For example, in the context of 4:2:0 chroma subsampling, instead of applying a box filter, all that is required with the proposed approach is to sum four reconstructed luma pixels that match the chroma pixel. That is, to align the chroma resolution, a 4-tap {1 / 4,1 / 4,1 / 4,1 / 4} filter is used to downsample the co-located luma samples. Then, when CfL scales that luma pixel to improve prediction accuracy, the approach of some embodiments described herein can scale by only 2.

[0081] [Chroma Downsampling Format] In embodiments, there may be different YUV formats depending on different chroma downsampling phases, i.e., chroma formats, an example of which is shown in FIG. 9. Different chroma formats define different downsampling grids or phases for different color components. In the 4:2:0 format, as shown in FIG. 9, there may be two different downsampling formats that can be referred to as 4:2:0 MPEG1 or 4:2:0 MPEG2.

[0082] [Downsampling Filter] For the luma downsampling filter in AV1, the following equation (4) is applied to derive the reconstructed luma samples: [Equation]

[0083] In the above equation (4) and the following equations, rec’ Lrepresents the pixel value of the downsampled reconstructed luma pixel at position (i, j), where position (i, j) can be the position of the corresponding chroma pixel. Additionally, rec L (x, y) represents the reconstructed luma sample value located at position (x, y).

[0084] The downsampling filter in AV1 assumes the chroma downsampling format 1000 shown in FIG. 10, which may correspond to the 4:2:0 MPEG1 downsampling format of FIG. 9. Specifically, FIG. 10 shows chroma sample 1002 and four corresponding luma samples 1004 indexed from 0 to 3. In an embodiment, the chroma sample may be a chroma pixel, and the luma sample may be a luma pixel. In an embodiment, the luma sample may be used to determine the value of the downsampled luma sample corresponding to the chroma sample, or the value of the downsampled luma pixel corresponding to the chroma pixel, according to the above formula (4).

[0085] In some implementations of the CfL mode, only one downsampling filter is supported, but for some content or different chroma downsampling formats, the current unique downsampling filter may not be the optimal filter.

[0086] Therefore, an embodiment may provide support for multiple downsampling filters for luma reconstruction samples when a cross-component prediction mode such as the CfL prediction mode is selected.

[0087] The embodiment may also be applied to modes other than the CfL prediction mode, for example, another prediction mode that uses one color component to predict another color component, and where downsampling is required for one or more color components. Therefore, the embodiment may also be applied, for example, by replacing luma with one specific color component (e.g., R) and chroma with another specific color component (e.g., G or B).

[0088] An example of a downsampling filter according to an embodiment is provided below.

[0089] Example 1 According to Example 1, a 6-tap filter can be supported for a cross-component downsampling process, such as a CfL luma downsampling process. FIG. 11A shows a chroma downsampling format 1110 corresponding to a 6-tap filter according to an embodiment. As can be seen from FIG. 11A, the chroma downsampling format 1110 includes a chroma sample 1112 and six corresponding luma samples 1114 indexed from 0 to 5. In an embodiment, the luma samples 1114 may be reconstructed luma samples, which can be used to derive downsampled luma pixels corresponding to the chroma sample 1112 that may be located at position (i,j) based on the following formula (5).

Equation

[0090] In the above formula (5) and other formulas discussed herein, "rounding" may represent a rounded value. In formula (5), the rounded value may be, for example, 0 or 4. In an embodiment, the 6-tap downsampling filter may assume that the chroma downsampling format corresponds to the 4:2:0 MPEG2 downsampling format.

[0091] Example 2 According to Example 2, a 5-tap filter can be supported for a cross-component downsampling process, such as a CfL luma downsampling process. FIG. 11B shows a chroma downsampling format 1120 corresponding to a 5-tap filter according to an embodiment. As can be seen from FIG. 11B, the chroma downsampling format 1110 includes a chroma sample 1122 and five corresponding luma samples 1124 indexed from 0 to 4. In an embodiment, the luma sample 1124 corresponding to index 4 can be juxtaposed with the chroma sample 1122. In an embodiment, the luma sample 1124 may be a reconstructed luma sample, which can be used to derive a downsampled luma pixel corresponding to the chroma sample 1122 that can be placed at position (i,j) based on one of the following equations (6) and (7).

Number

[0092] In an embodiment, the rounding value may be 0 or non-zero. For example, in Equation (6), the rounding value may be 4, and in Equation (7), the rounding value may be 8.

[0093] In an embodiment, when samples are not available from the upper row, the current row pixel (e.g., the luma sample 1124 located at index 4) may be used to pad the upper sample (e.g., the luma sample 1124 located at index 0). In an embodiment, when the current sample is located at a superblock / CTU boundary, the current row pixel may be used to pad the sample located at the sample position of index 0.

[0094] Example 3 According to Example 3, a 4-tap filter can be supported for the cross-component downsampling process, such as the CfL luma downsampling process. FIG. 11C shows a chroma downsampling format 1130 corresponding to a 4-tap filter according to an embodiment. As can be seen from FIG. 11C, the chroma downsampling format 1130 includes a chroma sample 1132 and four corresponding luma samples 1134 indexed from 0 to 3. In an embodiment, the luma sample 1124 corresponding to index 3 can be juxtaposed with the chroma sample 1132. In an embodiment, the luma samples 1134 may be reconstructed luma samples, which can be used to derive the downsampled luma pixels corresponding to the chroma sample 1132 that can be arranged at position (i,j) based on one of the following equations (8) and (9).

Number

[0095] In an embodiment, the rounding value may be 0 or non-zero. For example, in Equation (8), the rounding value may be 4, and in Equation (9), the rounding value may be 8.

[0096] Example 4 According to Example 4, another 4-tap filter can be supported for the cross-component downsampling process, such as the CfL luma downsampling process. FIG. 11D shows a chroma downsampling format 1140 corresponding to a 4-tap filter according to an embodiment. As can be seen from FIG. 11D, the chroma downsampling format 1140 includes a chroma sample 1142 and four corresponding luma samples 1144 indexed from 0 to 3. In an embodiment, the luma sample 1144 corresponding to index 3 can be juxtaposed with the chroma sample 1142. In an embodiment, the luma samples 1144 can be reconstructed luma samples, which can be used to derive the downsampled luma pixels corresponding to the chroma sample 1142 that can be placed at position (i,j) based on one of the following equations (10) and (11).

Number

[0097] In an embodiment, the rounding value may be 0 or non-zero. For example, in Equation (10), the rounding value may be 4, and in Equation (11), the rounding value may be 8.

[0098] Example 5 According to Example 5, a 3-tap filter can be supported for the cross-component downsampling process, such as the CfL luma downsampling process. FIG. 11E shows a chroma downsampling format 1150 corresponding to a 3-tap filter according to an embodiment. As can be seen from FIG. 11E, the chroma downsampling format 1150 includes a chroma sample 1152 and three corresponding luma samples 1154 indexed from 0 to 2. In an embodiment, the luma sample 1154 corresponding to index 1 can be juxtaposed with the chroma sample 1152. In an embodiment, the luma sample 1154 can be a reconstructed luma sample, which can be used to derive a downsampled luma pixel corresponding to the chroma sample 1152 that can be placed at position (i,j) based on one of the following equations (12) and (13).

Number

[0099] In an embodiment, the rounding value may be 0 or non-zero. For example, in Equation (12), the rounding value may be 4, and in Equation (13), the rounding value may be 8.

[0100] In an embodiment, when a downsampling process is required for the luma reconstruction sample, in addition to the filter corresponding to Equation (4) and the downsampling format 1000 of AV1, the filters discussed in Examples 1-5 can be supported.

[0101] In one embodiment, N filters may be used, where N can be 1, 2, 3, 4, 5, or 6. In an embodiment, when N is 4, the filters corresponding to Examples 1 to 3 and the AV1 filter corresponding to Equation (4) may be used. In an embodiment, when N is 3, the filters corresponding to Example 1 and Example 2 and the AV1 filter corresponding to Equation (4) may be used. In an embodiment, when N is 3, the filters corresponding to Example 1 and Example 3 and the AV1 filter corresponding to Equation (4) may also be used.

[0102] In an embodiment, a high-level syntax flag / index can be signaled to indicate a downsampling filter used in a CfL prediction mode or a cross-component intra prediction mode such as another downsampling process, which is involved in an encoding / decoding process that requires downsampling of luminance to a lower resolution and matching with chroma, such as other cross-component prediction methods. In an embodiment, the high-level syntax flag / index can be signaled in at least one of a sequence header or a sequence parameter set (SPS), a picture parameter set (PPS), an adaptive parameter set (APS), a video parameter set (VPS), a slice header (SH), a picture header (PH), a frame header, a tile header, a coding tree unit (CTU) header, a superblock header, and a block having a specific predefined block size (e.g., 32×32, 64×64).

[0103] In an embodiment, the high-level syntax flag / index for downsampling filter selection may not be signaled when one or all of the cross-component coding tools such as the CfL mode are invalid or when the color format is neither 4:2:0 nor 4:2:2.

[0104] In an embodiment, when the color format of the video sequence is 4:4:4 or 4:0:0, the high-level syntax flag / index for downsampling filter selection may not be signaled.

[0105] In an embodiment, the downsampling filter can be signaled on a per-pixel basis, and the tap coefficients at nine positions around the neighboring luma pixels may be signaled. In one embodiment, when signaling the filter coefficients, the filter coefficients corresponding to the neighboring positions may be associated with a larger size. FIG. 12 shows an example of nine positions indexed from 0 to 8. In an embodiment, the central position (index 4) may be juxtaposed with the chroma sample.

[0106] In an embodiment, the filter coefficients may be signaled in at least one of a sequence header, SPS, PPS, APS, VPS, SH, PH, frame header, tile header, CTU / superblock header, and a block having a specific predefined block size (e.g., 32×32, 64×64).

[0107] FIG. 13 is a flowchart of a process 1300 for performing chroma from luma (CfL) intra prediction according to an embodiment.

[0108] As shown in FIG. 13, in operation 1302, process 1300 includes receiving a current chroma block from the coded bitstream.

[0109] As further shown in FIG. 13, in operation 1304, process 1300 includes determining whether to enable the CfL intra prediction mode for the coded bitstream.

[0110] As further shown in FIG. 13, in operation 1306, process 1300 includes determining a color format corresponding to the coded bitstream. In an embodiment, the color format can be any of the color formats described above.

[0111] As further shown in FIG. 13, in operation 1308, process 1300 includes obtaining, from the coded bitstream, a syntax element indicating a downsampling filter used for the CfL intra prediction mode, based on a determination that the CfL intra prediction mode is enabled and based on the determined color format. In an embodiment, the downsampling filter can be any filter from the filters described with respect to Examples 1 to 5, and the filters corresponding to Equation 4 and FIG. 10. In an embodiment, the syntax element can be the high-level flag or index described above. In an embodiment, the syntax element can be signaled in at least one of a sequence header, a sequence parameter set, a picture parameter set, an adaptive parameter set, a video parameter set, a slice header, a picture header, a frame header, a tile header, a coding tree unit header, a superblock header, or a block having a predetermined block size.

[0112] As further shown in FIG. 13, in operation 1310, process 1300 includes selecting a downsampling filter to be used for a chroma block in the CfL intra prediction mode from among a plurality of downsampling filters.

[0113] As further shown in FIG. 13, in operation 1312, process 1300 includes determining luma sample positions associated with the current chroma block based on the selected downsampling filter.

[0114] As further shown in FIG. 13, in operation 1314, process 1300 includes downsampling a plurality of luma samples at luma sample positions, where the pixels within each downsampled luma sample are co-located with corresponding pixels within the current chroma block. In an embodiment, operation 1310 may be performed using any one or more of the above-described equations (4) through (13).

[0115] As further shown in FIG. 13, in operation 1316, process 1300 includes reconstructing the current chroma block based at least on the plurality of downsampled luma samples.

[0116] In an embodiment, the CfL intra prediction mode is included in a plurality of cross-component coding tools, and based on all of the plurality of cross-component coding tools being disabled, the syntax element may not be signaled.

[0117] In an embodiment, based on the color format being a 4:2:0 color format or a 4:2:2 color format, the syntax element may not be signaled.

[0118] In an embodiment, based on the color format being a 4:4:4 color format or a 4:0:0 color format, the syntax element may not be signaled.

[0119] In an embodiment, the downsampling filter may include at least one of a 6-tap filter, a 5-tap filter, and a 3-tap filter.

[0120] In an embodiment, the downsampling filter may include a 4-tap filter, and the luma sample positions from among the plurality of luma sample positions may be co-located with the current chroma block.

[0121] FIG. 13 shows exemplary blocks of process 1300, but in some implementations, process 1300 may include additional blocks, fewer blocks, different blocks, or blocks in a different arrangement than those shown in FIG. 13. Additionally or alternatively, two or more of the blocks of process 1300 may be executed in parallel.

[0122] Furthermore, the proposed method may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium to perform one or more of the proposed methods.

[0123] The techniques of the embodiments of the present disclosure described above can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, FIG. 14 shows a computer system (900) suitable for implementing an embodiment of the disclosed subject matter.

[0124] The computer software can be coded using any suitable machine code or computer language that can be the subject of assembly, compilation, linking, or similar mechanisms, and can create code that includes instructions that can be executed directly by a computer central processing unit (CPU), a graphics processing unit (GPU), etc., or through interpretation, microcode execution, etc.

[0125] The instructions can be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0126] The components shown in FIG. 14 for the computer system (900) are illustrative in nature and are not intended to suggest any limitations regarding the scope or functionality of the computer software implementing the embodiments of the present disclosure. Also, the configuration of the components should not be construed as having any dependency or requirement regarding any one or combination of the components illustrated in the exemplary embodiments of the computer system (900).

[0127] The computer system (900) may include a specific human interface input device. Such a human interface input device may respond to input by one or more human users through, for example, tactile input (keystrokes, swipes, movement of a data glove, etc.), audio input (voice, clapping, etc.), visual input (gestures, etc.), and olfactory input (not shown). Also, the human interface input device can be used to capture a specific medium, such as audio (voice, music, ambient sound, etc.), images (scanned images, photographic images obtained from a still image camera, etc.), and video (2D video, 3D video including stereoscopic video, etc.), which is not necessarily directly related to conscious human input.

[0128] The human interface input device may include one or more of a keyboard (901), a mouse (902), a trackpad (903), a touch screen (910), a data glove, a joystick (905), a microphone (906), a scanner (907), and a camera (908) (only one of each is shown).

[0129] The computer system (900) may also include certain human interface output devices. Such human interface output devices can stimulate the senses of one or more human users, for example, through tactile output, acoustics, light, and smell / taste. Such human interface output devices include tactile output devices (e.g., tactile feedback by a touch screen (910), data glove, or joystick (905), although there may be tactile feedback devices that do not function as input devices). For example, such devices can be audio output devices (such as speakers (909), headphones (not shown), etc.), visual output devices (regardless of whether each has a touch screen input function and also regardless of whether each has a tactile feedback function, some of which can output two-dimensional visual output or output beyond three dimensions through means such as stereoscopic image output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown), including screens (910) such as CRT screens, LCD screens, plasma screens, OLED screens, etc.) and printers (not shown).

[0130] The computer system (900) can also include human-accessible storage devices and their associated media, such as optical media or similar media (921) including CD / DVD ROM / RW (920) with CD / DVDs, thumb drives (922), removable hard drives or solid state drives (923), legacy magnetic media such as tapes and floppy disks (not shown), and special ROM / ASIC / PLD-based devices such as security dongles (not shown).

[0131] One of ordinary skill in the art should also understand that the term "computer-readable medium" does not include transmission media, carrier waves, or other transient signals when used in connection with the presently disclosed subject matter.

[0132] The computer system (900) can also include an interface to one or more communication networks. The network can be, for example, wireless, wired, optical. The network can further be local, wide area, metropolitan, vehicle and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet (registered trademark), wireless LAN, cellular networks including GSM (registered trademark), 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV and terrestrial broadcast TV, vehicle and industrial networks including CANBus, etc. A particular network generally requires attachment to a particular general-purpose data port or peripheral bus (949) and an external network interface adapter (such as a USB port of the computer system (900)), while others are generally integrated into the core of the computer system (900) by attachment to the system bus described later (such as an Ethernet (registered trademark) interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (900) can communicate with other entities. Such communication can be, for example, one-way receive-only (such as broadcast TV), one-way transmit-only (such as from a particular CANbus to a particular CANbus device) or two-way to other computer systems using a local or wide area digital network. Such communication can include communication to a cloud computing environment (955). As described above, a particular protocol and protocol stack can be used in each of these networks and network interfaces.

[0133] The aforementioned human interface device, human-accessible storage device and network interface (954) can be attached to the core (940) of the computer system (900).

[0134] The core (940) can include one or more central processing units (CPUs) (941), a graphics processing unit (GPU) (942), a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) (943), a hardware accelerator (944) for specific tasks, etc. These devices can be connected through a system bus (948) together with a read-only memory (ROM) (945), a random access memory (RAM) (946), an internal large-capacity storage such as an internal non-user-accessible hard drive, SSD, etc. (947). In some computer systems, the system bus (948) can be accessed in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be attached directly to the system bus (948) of the core or via a peripheral bus (949). The architecture of the peripheral bus includes PCI, USB, etc. A graphics adapter (950) may be included in the core (940).

[0135] The CPU (941), GPU (942), FPGA (943), and accelerator (944) can execute specific instructions that can be combined to form the above-mentioned computer code. The computer code can be stored in the ROM (945) or RAM (946). Also, temporary data can be stored in the RAM (946), while permanent data can be stored, for example, in the internal large-capacity storage (947). Through the use of cache memory that can be closely associated with one or more CPUs (941), GPUs (942), large-capacity storage (947), ROM (945), RAM (946), etc., fast storage and retrieval for any of the memory devices can be enabled.

[0136] A computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and the computer code can be specially designed and constructed for the purposes of this disclosure, or they can be of the kind well-known and available to those having skill in the computer software arts.

[0137] By way of example and not limitation, the architecture corresponding to computer system (900) and in particular core (940) can provide functionality as a result of a processor (including CPU, GPU, FPGA, accelerator, etc.) executing software embodied on one or more tangible computer-readable media. Such computer-readable media can be associated with user-accessible mass storage as introduced above, as well as specific storage of core (940) of a non-transitory nature such as internal core mass storage (947) or ROM (945). The software implementing various embodiments of the present disclosure can be stored on such devices and executed by core (940). The computer-readable media can include one or more memory devices or chips according to specific needs. The software can cause core (940) and in particular the processor therein (including CPU, GPU, FPGA, etc.) to define data structures stored in RAM (946) and modify such data structures according to processes defined by the software, thereby executing specific processes or specific parts of specific processes described herein. Additionally or alternatively, the computer system can provide functionality as a result of being embodied in logic hardware wires or otherwise in a circuit (e.g., accelerator (944)), which can operate instead of or together with software to execute specific processes or specific parts of specific processes described herein. References to software include logic and, if necessary, vice versa. References to computer-readable media can include circuits (such as integrated circuits (ICs), etc.) storing software for execution, circuits embodying logic for execution, or both as appropriate. The present disclosure encompasses any suitable combination of hardware and software.

[0138] Embodiments of the present disclosure may be used separately or combined in any order. Further, each of the embodiments (and methods thereof) may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium.

[0139] The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit implementations to the exact form disclosed. Modifications and variations are possible in light of the above disclosure, or may be obtained from practice of the implementations.

[0140] As used herein, the term component is intended to be broadly construed as hardware, firmware, or a combination of hardware and software.

[0141] Combinations of features are recited in the claims and / or disclosed herein, but these combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features may not be specifically recited in the claims and / or may be combined in ways not disclosed herein. Each of the dependent claims listed below may depend directly on only one claim, but the disclosure of possible implementations includes each dependent claim in combination with all other claims in the claim set.

[0142] Elements, operations or instructions used in this specification should not be construed as important or essential unless so explicitly described. Also, as used in this specification, the articles "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more". Further, as used in this specification, the term "set" is intended to include one or more items (e.g., related items, unrelated items, combinations of related and unrelated items, etc.) and may be used interchangeably with "one or more". When only one item is intended, the term "one" or similar language is used. Also, as used in this specification, the terms "has", "have", "having", etc. are intended to be open-ended terms. Further, the phrase "based on" is intended to mean "at least in part" "based on" unless otherwise specified.

[0143] Although the present disclosure describes some non-limiting exemplary embodiments, there are changes, substitutions and various alternative equivalents within the scope of the present disclosure. Thus, it will be understood by those skilled in the art that although not explicitly shown or described herein, various systems and methods that embody the principles of the present disclosure and are thus within the spirit and scope of the present disclosure can be devised.

Claims

Claim 1 A method for performing Chroma from Luma (CfL) intra prediction, the method being executed by at least one processor, receiving a current chroma block from a coded bitstream; determining whether a CfL intra prediction mode is enabled for the coded bitstream; determining a color format corresponding to the coded bitstream; obtaining, from the coded bitstream, a syntax element indicating a downsampling filter used for the CfL intra prediction mode based on a determination that the CfL intra prediction mode is enabled and based on the determined color format; selecting the downsampling filter to be used for the chroma block in the CfL intra prediction mode from among a plurality of downsampling filters; determining luma sample positions associated with the current chroma block based on the selected downsampling filter; downsampling a plurality of luma samples at the luma sample positions, wherein pixels within each downsampled luma sample are co-located with corresponding pixels within the current chroma block; reconstructing the current chroma block based at least on the plurality of downsampled luma samples; A method comprising the above steps. Claim 2 The CfL intra prediction mode is included in a plurality of cross-component coding tools, and the syntax element is not signaled based on all of the plurality of cross-component coding tools being disabled. The method according to claim 1. The method according to claim 1. Claim 3 The syntax element is not signaled based on the color format not being a 4:2:0 color format or a 4:2:2 color format. The method according to claim 1. The method according to claim 1. Claim 4 The syntax element is not signaled based on the color format being a 4:4:4 color format or a 4:0:0 color format. The method according to claim 1. The method according to claim 1. Claim 5 The downsampling filter includes at least one selected from a 6-tap filter, a 5-tap filter, and a 3-tap filter. The method according to claim 1.

6. The downsampling filter includes a 4-tap filter. A certain luma sample position among the luma sample positions is co-located with the current chroma block. The method according to claim 1.

7. The syntax element is signaled by at least one selected from a sequence header, a sequence parameter set, a picture parameter set, an adaptive parameter set, a video parameter set, a slice header, a picture header, a frame header, a tile header, a coding tree unit header, a super block header, or a block having a predetermined block size. The method according to claim 1.

8. A device for performing chroma from luma (C f L) intra prediction, the device comprising: at least one memory configured to store program code; at least one processor configured to access the program code and operate as instructed by the program code; and the program code includes: reception code configured to cause the at least one processor to receive a current chroma block from a coded bitstream; first determination code configured to cause the at least one processor to determine whether a C f L intra prediction mode is enabled for the coded bitstream; second determination code configured to cause the at least one processor to determine a color format corresponding to the coded bitstream; acquisition code configured to cause the at least one processor to acquire, based on a determination that the C f L intra prediction mode is enabled and based on the determined color format, a syntax element indicating a downsampling filter used for the C f L intra prediction mode from the coded bitstream. Selection code configured to cause the at least one processor to select, from among a plurality of downsampling filters, the downsampling filter to be used for a chroma block in the CfL intra prediction mode; Third determination code configured to cause the at least one processor to determine a luma sample position associated with the current chroma block based on the selected downsampling filter; Downsampling code configured to cause the at least one processor to downsample a plurality of luma samples at the luma sample position, wherein pixels within each downsampled luma sample are co-located with corresponding pixels within the current chroma block; Reconstruction code configured to cause the at least one processor to reconstruct the current chroma block based at least on the plurality of downsampled luma samples; A device comprising the same. **Claim 9** The CfL intra prediction mode is included in a plurality of cross-component coding tools, Based on all of the plurality of cross-component coding tools being disabled, the syntax element is not signaled. The device according to claim 8. **Claim 10** Based on the color format not being the 4:2:0 color format or the 4:2:2 color format, the syntax element is not signaled. The device according to claim 8. **Claim 11** Based on the color format being the 4:4:4 color format or the 4:0:0 color format, the syntax element is not signaled. The device according to claim 8. **Claim 12** The downsampling filter includes at least one of a 6-tap filter, a 5-tap filter, and a 3-tap filter. The device according to claim 8. **Claim 13** The downsampling filter includes a 4-tap filter, A certain luma sample position among the luma sample positions is co-located with the current chroma block. The device according to claim 8. **Claim 14** The syntax element is signaled by at least one of a sequence header, a sequence parameter set, a picture parameter set, an adaptive parameter set, a video parameter set, a slice header, a picture header, a frame header, a tile header, a coding tree unit header, a super block header, or a block having a predetermined block size. The device according to claim 8. Claim 15 A non-transitory computer-readable medium storing instructions that, when executed by one or more processors of a device for performing chroma from luma (CfL) intra prediction, cause the one or more processors to execute the method according to any one of claims 1 to 7.